About this opportunity
NVIDIA lists this Senior LLM Inference: GPU Kernel Optimization opportunity in austin, Texas. Review the employer’s description below for duties, qualifications and application requirements.
Job description
NVIDIA is seeking a Sr. Inference Engineer to push LLM inference performance through GPU kernel optimization.
You will help develop silicon-measured benchmarking, model-level performance projection tooling, and agentic optimization systems, collaborating across compiler, hardware, and framework teams to surface bottlenecks and deliver measurable gains. You will work on GPU kernel microbenchmarking, end-to-end model performance analysis, and agentic optimization, shaping production inference
#J-18808-Ljbffr
Worksite address
austin, TX, 78716, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.