About this opportunity
Brahma Consulting Group lists this Machine Learning Engineer, GPU Performance opportunity in san francisco, California. Review the employer’s description below for duties, qualifications and application requirements.
Job description
Brahma is conducting this search on behalf of a client.
We're a small, well-funded founding team rebuilding the training and inference stack for generative video and image models. Today's stack was built for language models. We're co-designing across GPU kernels, distributed systems, and the models themselves to make these workloads dramatically faster and cheaper. Our inference engine is live with customers in generative media and robotics.
This is an early engineering hire. You'll report directly to the CEO and work alongside a founding team with deep expertise in distributed systems, kernel optimization, cloud infrastructure, and ML research.
What you'll do
Optimize GPU performance for training and inference on image and video generation workloads
Profile and remove bottlenecks at the kernel, memory, system, and cluster level using Nsight and related tools
Write CUDA and Triton kernels that ship to production
Build distributed inference and training engines for diffusion models across multiple GPUs and nodes
Own communication performance: NCCL, RDMA over InfiniBand or RoCE, and disaggregated serving
Build benchmarking and regression harnesses so performance gains hold in production
What we're looking for
1+ years working on deep learning inference or training systems, or distributed systems
Hands-on experience with CUDA, Triton, PyTorch internals, or GPU profiling
Strong CS fundamentals and a drive to go deep on hard technical problems
You'd rather make a model ten times faster than train one
#J-18808-Ljbffr
Worksite address
san francisco, CA, 94199, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.