This job is closed
Applications are no longer available for this announcement. Explore current related opportunities below.
Job description
Competitive salary
Competitive salary
Plus meaningful equity
All roles
San Francisco, CA
Site Reliability Engineer
San Francisco, CAFull-timeMid to SeniorOn-site
Zof AI is hiring for this role in San Francisco, CA. This is a full-time opportunity for candidates who want to contribute directly to the development of ambitious AI products in a high-performance environment.
Must be able to run infrastructure for large agent workloads and use AI tools to automate operational work.
About This Role
Zof AI is seeking a Site Reliability Engineer to run the infrastructure that lets fleets of sandboxed agents execute customer code safely and cheaply. This role owns the execution layer of our control plane: Kubernetes and container orchestration, CI/CD pipelines, hard isolation for untrusted code, observability, and the cost controls that keep large agent fleets affordable. If you have worked as a Site Reliability Engineer, Platform Engineer, Cloud Engineer, or Infrastructure Engineer, this is that discipline at Zof AI. The ideal candidate has operated production infrastructure at scale and treats security, reliability, and cost per agent run as constraints they personally own.
Responsibilities
Design and operate the sandboxed environments where agents reproduce defects and validate fixes.
Own Kubernetes, container, and compute infrastructure end to end.
Build CI/CD pipelines that let engineers ship safely many times a day.
Harden isolation boundaries so untrusted customer code stays inside its sandbox.
Instrument fleet health with metrics, logs, tracing, and alerting that catch failures early.
Drive down cost per agent run through scheduling, autoscaling, and capacity work.
Automate provisioning, deployment, rollback, and environment management.
Partner with engineers to make infrastructure fast and safe to build on.
Requirements
Experience running production infrastructure on a major cloud platform.
Working knowledge of Kubernetes, containers, and orchestration.
Experience building CI/CD pipelines, deployment automation, and infrastructure as code.
Familiarity with observability, monitoring, and on-call practice.
Judgment about security, reliability, and cost trade-offs.
Daily use of AI tools to automate operational and engineering work.
Clear written and verbal communication.
Comfort operating in a fast-moving environment.
Nice to have
Experience with sandboxing or multi-tenant isolation tooling such as gVisor or Firecracker.
Experience with Terraform, Pulumi, or similar infrastructure as code tooling.
Experience running large batch or job-based workloads cost efficiently.
Experience in early-stage infrastructure or platform teams.
DevOpsKubernetesCI/CDCloud InfrastructureReliability
What we provide in San Francisco
MacBook Pro
Premium AI development tools
Cursor Ultra
Claude Code Ultra
OpenAI Codex Max or equivalent advanced AI tooling
Access to a high-performance AI product environment
Close collaboration with leadership, engineering, and customers
Opportunity to work in the San Francisco AI ecosystem
Wellness and productivity support where applicable
Competitive startup environment
High ownership
Direct product impact
Benefits may depend on role and final offer terms.
#J-18808-Ljbffr
Worksite address
san francisco, CA, 94199, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.