About this opportunity
AI Breaking Wire lists this Machine Learning Engineer, Inference Infrastructure opportunity in san francisco, California. Review the employer’s description below for duties, qualifications and application requirements.
Job description
About the Role
We are seeking a Machine Learning Engineer to join our Inference Infrastructure team. You will build and optimize the high-throughput, low-latency distributed systems that power models like GPT-4 and Sora for millions of users worldwide.
Responsibilities
Architect, build, and scale low-latency distributed inference serving systems for massive generative models.
Optimize GPU utilization, memory management, and kernel execution for state-of-the-art transformer architectures.
Collaborate with research teams to ensure smooth transition of new model architectures into production.
Monitor system performance, troubleshoot bottlenecks, and implement robust reliability measures.
Requirements
BS, MS, or Ph.D. in Computer Science or related technical field.
3+ years of industry experience building large-scale distributed systems or ML infrastructure.
Deep proficiency in C++ and Python.
Extensive experience with CUDA, Triton, or deep learning hardware accelerators.
Familiarity with distributed training and inference frameworks (vLLM, TensorRT-LLM, Megatron).
Benefits
Top-tier compensation including equity.
Full medical, dental, and vision coverage with zero employee contribution options.
Unlimited paid time off and flexible hybrid work policy.
Catered daily lunches and wellness stipends.
#J-18808-Ljbffr
Worksite address
san francisco, CA, 94199, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.