waypointjobs

Coral Bricks AI

Distributed Systems Engineer, GPU Infrastructure

northern, KY

Check who can apply and the requirements below before continuing.

About this opportunity

Coral Bricks AI lists this Distributed Systems Engineer, GPU Infrastructure opportunity in northern, Kentucky. Review the employer’s description below for duties, qualifications and application requirements.

Job description

Engineering San Francisco or remote - Full-time

Own the clusters, GPU fleet, and production systems that turn inference research into a reliable service

About Coral Bricks

Our mission is to make frontier intelligence affordable and accessible to everyone. Frontier models are finally here - but almost nobody can afford to use them freely. People have token anxiety: they meter every call, ration every context window, and settle for weaker models because the best ones are priced out of everyday use.

We're building the inference platform that ends that, starting with the workloads that feel the squeeze hardest: research and coding agents that swarm across multiple models, plan, call tools for hours, and reason over big context. Classic LLM serving was never built for them - rate limits that throttle real workloads, queues that stretch a 20-minute job into a 4-hour one, costs that grow with every agent turn. Same models, same prompts - many times the tokens per second at a fraction of the cost.

The team is small, technical, and shipping. We also build in the open: a lot of the day-to-day happens in our Discord, where the developers building on Coral Bricks tell us what broke, compare numbers with us, and push on what we work on next.

The role

You'll own the operational systems behind our inference platform: the clusters, GPU fleet, deployment machinery, and control plane that keep models available and traffic moving. When research produces a faster serving technique or a new model drops, you'll turn it into a repeatable, observable, production launch.

This role is distinct from our inference research role. You won't be measured on inventing a new attention kernel. You'll be measured on whether we can provision capacity, place workloads, ship changes, recover from failures, and operate a growing fleet without heroics.

This is a founding-team role with broad ownership. You'll work across cloud infrastructure, distributed systems, networking, storage, deployment, and the serving layer where they meet.

What you’ll work on

Own our GPU clusters and fleet across cloud providers: capacity, provisioning, machine images, drivers, networking, storage, health, and cost.

Build the control-plane systems that place workloads, manage capacity, drain and replace unhealthy nodes, and recover cleanly from failures.

Turn model launches into a reliable process: bring up new weights, validate serving configurations, roll out safely, watch production behavior, and roll back when needed. You'll hear how a launch landed from the developers in our Discord, not only from the dashboards.

Build deployment and release systems for inference servers and the services around them, with fast feedback and clear failure modes.

Create the observability we need to operate the fleet: metrics, logs, traces, dashboards, alerts, and tools that make incidents diagnosable instead of mysterious.

Improve reliability at every layer - autoscaling, load balancing, failover, backpressure, graceful degradation, and capacity planning.

Automate recurring operational work so the fleet can grow faster than the team operating it.

You probably have

Strong backend or distributed-systems fundamentals and experience owning production services end to end.

Experience with Linux, containers, networking, and at least one major cloud platform. You can debug across application, host, and infrastructure boundaries.

Good instincts around reliability: staged rollouts, observability, failure isolation, incident response, and simple systems that are easy to operate.

Comfort working from symptoms to root cause. A failed launch, an unhealthy node, or a latency spike is a systems problem to investigate, not a ticket to hand off.

A high work ethic and excitement about early-stage startups. The pace is fast, the problems are open-ended, and everyone does a bit of everything.

A bias toward shipping and automation. You fix the immediate problem, then build the mechanism that keeps it from becoming routine work.

Bonus

Experience operating GPU or accelerator fleets, including NVIDIA or AMD drivers, topology, health checks, and failure modes.

Experience with Kubernetes, Nomad, Slurm, ECS, or another cluster scheduler - especially if you've had to work below its happy path.

Familiarity with vLLM, SGLang, TensorRT-LLM, PyTorch distributed, NCCL, or other model-serving and collective-communication systems.

Experience with multi-cloud capacity, bare-metal provisioning, or scarce-resource scheduling.

You've built an internal platform, scheduler, deployment system, or piece of infrastructure that other engineers trusted in production.

Compensation

$120,000-$200,000 base salary, plus 0.25%-$2.0% equity. Where you land depends on experience, and cash and equity move together - take less of one and we'll weight the other.

Equity vests over four years with a one-year cliff. Health, dental, and vision coverage, and flexible time off.

Founding engineers shape the platform, the technical direction, and the team we build around it.

#J-18808-Ljbffr

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

American National Red Cross

WhatJobs

Building Engineer

parkhill, PA

Salary not specified

The American National Red Cross seeks a Building Engineer to support our Facilities & Maintenance team. This role ensures safe, efficient operati…

Listing review due 2026-10-06View job

Hillcrest HealthCare System

WhatJobs

Facilities Engineer

tulsa, OK

Salary not specified

Hillcrest HealthCare System in Tulsa, OK is seeking a Facilities Engineer to support safe, reliable operations across our hospitals and clinics. …

Listing review due 2026-10-06View job

American National Red Cross

WhatJobs

Building Engineer

benson, PA

Salary not specified

The American National Red Cross seeks a Building Engineer to support our Facilities & Maintenance team. This role ensures safe, efficient operati…

Listing review due 2026-10-06View job

American National Red Cross

WhatJobs

Building Engineer

tire hill, PA

Salary not specified

The American National Red Cross seeks a Building Engineer to support our Facilities & Maintenance team. This role ensures safe, efficient operati…

Listing review due 2026-10-06View job

American National Red Cross

WhatJobs

Building Engineer

st. michael, PA

Salary not specified

The American National Red Cross seeks a Building Engineer to support our Facilities & Maintenance team. This role ensures safe, efficient operati…

Listing review due 2026-10-06View job

American National Red Cross

WhatJobs

Building Maintenance Engineer

johnstown, PA

Salary not specified

The Building Maintenance Engineer at the American National Red Cross supports mission-driven operations by ensuring safe, reliable, and efficient…

Listing review due 2026-10-06View job