waypointjobs

fal - Features & Labels

Software Engineer, Site Reliability

san francisco, CA

Check who can apply and the requirements below before continuing.

About this opportunity

fal - Features & Labels lists this Software Engineer, Site Reliability opportunity in san francisco, California. Review the employer’s description below for duties, qualifications and application requirements.

Job description

fal is the generative media ecosystem powering the next generation of AI products. We build the infrastructure, tools, and model access that teams need to move from idea to production, and do it at scale without compromise. For developers and enterprises, fal is the foundation that makes generative media not just possible, but practical: a unified platform where high-performance inference, orchestration, and observability come together to unlock new categories of AI-native products.

As generative media reshapes industries across a market projected to grow by hundreds of billions over the next decade, fal is becoming the ecosystem that ambitious teams build on.

About this role

You are a seasoned SRE who keeps production infrastructure running at scale. You own the reliability and availability of customer-facing systems — from Kubernetes clusters to deployment pipelines to the networking layer that connects it all. You think in SLOs, automate ruthlessly, and treat every incident as a chance to make the system better.

What you'll do:

Own and operate our Kubernetes infrastructure: cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads

Build and maintain CI/CD pipelines and deployment infrastructure

Leverage AI to an extreme level to automate analysis and resolution of production issues, and improve software development speed, reliability and maintainability

Build dashboards, alerting, and anomaly detection across our systems

Define and enforce SLOs and build out incident response processes

Manage and improve our networking, load balancing, and service mesh configurations

Drive reliability improvements across the stack through automation, runbooks, and chaos engineering

Qualifications/nice to have:

5+ years experience in managing critical production systems and software development workflows

Strong production experience setting up and operating Kubernetes at scale, using infrastructure-as-code (Terraform, Ansible)

Deep knowledge of Linux networking, container networking (CNI plugins, VXLAN, BGP), and DNS

Experience building CI/CD systems and GitOps workflows (FluxCD, ArgoCD)

Proficiency in Python and either Go or Bash for tooling and automation

Strong experience with logging, monitoring and alerting (Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog)

Excellent communication and ability to drive technical decisions across teams

Self-starter who executes quickly, takes ownership, and constantly seeks improvement

Experience with managing GPU and AI/ML workloads

Experience with kernel-based monitoring and routing (eBPF, XDP)

Experience with security tooling (Falco, Coroot, SIEM)

Experience with bare metal Kubernetes networking (Calico, Cilium, MetalLB)

Experience with distributed storage systems (Ceph, Longhorn, etc.)

What we offer at fal

Interesting and challenging work

A lot of learning and growth opportunities

We are currently hiring in downtown San Francisco.

Health, dental, and vision insurance (US)

Regular team events and offsites

U.S. EQUAL EMPLOYMENT OPPORTUNITY INFORMATION:

fal provides equal employment opportunities to applicants and employees without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status, disability, or any other classification protected by applicable law.

#J-18808-Ljbffr

Worksite address

san francisco, CA, 94199, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

UF Health

WhatJobs

Information Technology Specialist

gainesville, FL

Salary not specified

UF Health is seeking an ECMO Specialist to provide expert support to critically ill adult and pediatric patients requiring extracorporeal life su…

Listing review due 2026-10-07View job

Govcio LLC

WhatJobs

Senior Dynamics Developer

washington, DC

$175,000.00 /Yr

Overview: GovCIO is looking for an experienced Senior CRM developer to support the continued enhancement and operations of a Microsoft Dynamics C…

Listing review due 2026-10-06View job

Govcio LLC

WhatJobs

Senior Application Developer

fairfax, VA

$165,000.00 /Yr

Overview: GovCIO is currently hiring a highly skilled Senior Application Developer with an active Secret clearance to design, build, and modify a…

Listing review due 2026-10-06View job

Govcio LLC

WhatJobs

Software Developer- SME

quantico, VA

$195,000.00 /Yr

Overview: GovCIO is currently hiring for a Software Developer-SME   to design, code, test, and maintain software applications and systems in supp…

Listing review due 2026-10-06View job

Govcio LLC

WhatJobs

Sr. Instructional Developer

fort meade, MD

$120,000.00 /Yr

Overview: GovCIO is currently hiring for an Instructional Developer  to design and develop training materials for personnel . This position will…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Senior Embedded Systems Developer

laurel, MD

See pay details in description

Description Do you love working on a motivated team to solve complex problems in innovative ways? Do you enjoy creating embedded prototypes i…

Listing review due 2026-10-06View job