About this opportunity
AI Breaking Wire lists this Scale Low-Latency ML Inference Engineer (GPU/CUDA) opportunity in northern, Kentucky. Review the employer’s description below for duties, qualifications and application requirements.
Job description
AI Breaking Wire seeks a Machine Learning Engineer to join our Inference Infrastructure team. You will build and optimize the high-throughput, low-latency distributed systems that power models like GPT-4 and Sora for millions of users worldwide.
The role emphasizes C++ and Python proficiency, GPU optimization with CUDA and Triton, and collaboration with research teams to productionize new architectures. This hybrid position is based in San Francisco.
#J-18808-Ljbffr
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.