waypointjobs

Acceler8 Talent

Machine Learning Engineer (Inference)

san francisco, CA

Check who can apply and the requirements below before continuing.

About this opportunity

Acceler8 Talent lists this Machine Learning Engineer (Inference) opportunity in san francisco, California. Review the employer’s description below for duties, qualifications and application requirements.

Job description

ML Inference Engineer

Build the infrastructure that makes cutting-edge AI fast enough to work at scale.

We’re hiring anML Inference Engineer for a Stanford-spun AI startup in San Francisco that has already grown to8-figure revenue .

The team is rebuilding itsLLM inference stack from the ground up , solving challenging systems problems around GPU performance, distributed compute and real-time model serving.

This role is for engineers who enjoy going deep onperformance, infrastructure, and optimisation .

The role

Build the infrastructure that serveslarge-scale LLM workloads

Push the limits oflatency, throughput and GPU efficiency

Design distributed inference acrosssingle and multi-GPU systems

Improve GPU scheduling, orchestration and resource utilisation

Profile and remove bottlenecks acrosscompute, memory and networking

Scale production workloads acrossKubernetes and GPU clusters

Make low-level architecture decisions wheremilliseconds matter

What we're looking for

StrongPython and/or C++

Experience withdistributed systems or high-performance computing

Knowledge ofLLM inference and model serving

Experience optimising GPU-heavy workloads

Exposure toCUDA, NCCL or Triton

Strong understanding ofPyTorch and modern ML infrastructure

Experience withvLLM, TensorRT-LLM, SGLang or similar

Knowledge of techniques such asquantisation, batching, KV caching and parallelism

Why join?

Tackle genuinely difficultAI infrastructure problems

Work on systems operating atreal production scale

Join a fast-growing company already at8-figure revenue

Significant ownership over anew inference architecture

Work at the intersection ofLLMs, GPUs and distributed systems

#J-18808-Ljbffr

Worksite address

san francisco, CA, 94199, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

American Honda Motor Co., Inc.

WhatJobs

Senior Product Quality & Root Cause Engineer

haw river, NC

Salary not specified

What Makes a Honda, is Who makes a Honda Honda has a clear vision for the future, and it’s a joyful one.  We are looking for individuals with t…

Listing review due 2026-10-06View job

Avantor

WhatJobs

Process Engineer

carpinteria, CA

See pay details in description

The Opportunity: NuSil (apart of Avantor) is seeking a Process Engineer to be responsible for all phases of silicone products manufacturing…

Listing review due 2026-10-06View job

GE Vernova

WhatJobs

Lead Application Engineer

boston, MA

See pay details in description

Job Description Summary The Lead Application Engineer is an established leader in their respective engineering team. They will drive busine…

Listing review due 2026-10-06View job