waypointjobs

Andromeda Cluster

Customer Reliability Engineer

san francisco, CA

Check who can apply and the requirements below before continuing.

About this opportunity

Andromeda Cluster lists this Customer Reliability Engineer opportunity in san francisco, California. Review the employer’s description below for duties, qualifications and application requirements.

Job description

Site Reliability Engineer - AI Infrastructure

Location: Global Remote / San Francisco · Full-Time

About Andromeda

Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers.

We began with a single managed cluster — but it filled almost instantly. Since then, we’ve been quietly building the systems, network, and orchestration layer that makes the world’s AI infrastructure more accessible.

Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it’s needed most. Our platform routes training and inference jobs across global supply, unlocking flexibility and efficiency in one of the fastest-growing markets on earth.

Our long-term vision is to build the liquidity layer for global AI compute — a marketplace that moves the infrastructure and workloads powering AGI not dissimilar to the flows of capital in the world's financial markets.

We are expanding to new frontiers to find the brightest that work in AI infrastructure, research and engineering.

What You’ll Do

Provision, configure, and operate Kubernetes-based clusters for customers across multiple providers.

Build automation and tooling to streamline cluster deployments and integrations.

Debug customer issues across networking, storage, scheduling, and system layers.

Improve reliability and scalability of both training and inference infrastructure.

Design and implement monitoring, alerting, and observability for critical systems.

Collaborate with engineering and product teams to plan and deliver infrastructure for new services.

Participate in on-call and incident response, leading postmortems and reliability improvements.

What We’re Looking For

5+ years experience in SRE, DevOps, or infrastructure engineering roles.

Strong Linux systems and networking fundamentals.

Deep experience with Kubernetes and container orchestration at scale.

Proficiency with Infrastructure-as-Code (Terraform, Helm, Ansible, etc.).

Strong automation and scripting skills (Python, Go, or Bash).

Experience with observability stacks (Prometheus, Grafana, Loki, Datadog, etc.).

Track record of operating production systems and leading incident response.

Nice to Have

Exposure to ML/AI infrastructure or GPU-based systems (CUDA, Slurm, Triton, etc.).

Familiarity with high-performance networking (InfiniBand, NVLink) or distributed storage (VAST, Weka, Ceph).

Customer-facing support or consulting experience.

Why You’ll Love It Here

This is a builder’s role. You’ll have ownership and autonomy to shape how our systems run, working directly with customers and providers while building the foundation for reliable, scalable AI infrastructure.

#J-18808-Ljbffr

Worksite address

san francisco, CA, 94199, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

U.S. Army Corps of Engineers

USAJOBS

Interdisciplinary Waterways Maintenance Chief

Portland, OR

$114,684.00 – $149,091.00 per yearfull time

About the Position: As Chief of the Waterways Maintenance Section, the incumbent exercises full technical, administrative, and managerial authori…

Closes 2026-10-16View job

U.S. Army Corps of Engineers

USAJOBS

Interdisciplinary

Baltimore, MD

$102,415.00 – $133,142.00 per yearfull time

About the Position: You will be responsible for the environmental assessment of hazardous, toxic, and radiological waste (HTRW) sites and Militar…

Closes 2026-10-15View job

American Honda Motor Co., Inc.

WhatJobs

Senior Product Quality & Root Cause Engineer

haw river, NC

Salary not specified

What Makes a Honda, is Who makes a Honda Honda has a clear vision for the future, and it’s a joyful one.  We are looking for individuals with t…

Listing review due 2026-10-06View job

Avantor

WhatJobs

Process Engineer

carpinteria, CA

See pay details in description

The Opportunity: NuSil (apart of Avantor) is seeking a Process Engineer to be responsible for all phases of silicone products manufacturing…

Listing review due 2026-10-06View job

GE Vernova

WhatJobs

Lead Application Engineer

boston, MA

See pay details in description

Job Description Summary The Lead Application Engineer is an established leader in their respective engineering team. They will drive busine…

Listing review due 2026-10-06View job