About this opportunity
Hewlett Packard Enterprise lists this Senior LLM Inference Runtime Engineer opportunity in durham, North Carolina. Review the employer’s description below for duties, qualifications and application requirements.
Job description
Hewlett Packard Enterprise is seeking an experienced software engineer to design and implement core components of the LLM runtime, focusing on engine integration, batching, and KV cache management. You will partner with distributed teams to optimize latency, throughput, and execution across deployments.
The ideal candidate has 8+ years in software engineering with 1-2+ years in LLM inference runtimes, and strong expertise in Go, Python, Kubernetes, and deep engine internals.
#J-18808-Ljbffr
Worksite address
durham, NC, 27703, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.