waypointjobs

Hobbsnews

Senior Software Engineer, Inference

northern, KY

Check who can apply and the requirements below before continuing.

About this opportunity

Hobbsnews lists this Senior Software Engineer, Inference opportunity in northern, Kentucky. Review the employer’s description below for duties, qualifications and application requirements.

Job description

Senior Software Engineer, InferenceThis role has been designed as 'Hybrid' with a requirement that you will work on average 2 days per week from an HPE office.

Who We Are:

Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today’s complex world.Our culture thrives onfinding new and better ways to accelerate what’s next.We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs.We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you.Open up opportunities with HPE.

HPE's Private Cloud AI organization is seeking a Senior Software Engineer to build and evolve the model runtime within HPE AI Essentials, the inference platform used by enterprises to operate large language models on infrastructure they own, including air-gapped and sovereign environments. The core engineering challenge in this domain is not model deployment but sustained execution efficiency: achieving low tail latency and high GPU utilization on customer-owned hardware of varying generation and configuration. In this role you will design and implement key components of that runtime – engine integration, batching, KV cache management, and distributed execution – together with the Kubernetes orchestration layer that supports it. The primary work location is as listed, but could be any other HPE site location in the US; however, remote work options will be considered.

Responsibilities

Design, implement, and own major components of the LLM serving deployment, including engine integration, continuous batching, KV cache management and reuse, and quantized execution

Partner with inference engineering teams and contribute to improving time-to-first-token, inter-token latency, throughput per GPU, and P95/P99 tail latency

Build and operate distributed execution capabilities, including disaggregated prefill/decode, tensor and pipeline parallelism, and KV cache offload across GPU memory, host memory, and RDMA-attached storage

Evaluate emerging runtimes, quantization schemes, speculative decoding, and mixture-of-experts serving, and make well-supported recommendations on adoption

Contribute to the orchestration layer supporting the runtime, including model admission, GPU scheduling and partitioning, cache-aware request routing, and autoscaling

Triage and resolve customer issues end-to-end, identifying root causes and improving systems and processes to prevent recurrence

Provide insightful code and design reviews, mentor team members, and lead by example on engineering practices within the team

Knowledge and Skills

Required

Familiar with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, TGI, or NVIDIA NIM, including modification of engine internals

Strong understanding of inference internals, including continuous batching, paged attention, KV cache reuse and prefix caching, chunked prefill, quantization, and speculative decoding

Working knowledge of tensor and pipeline parallelism, NCCL collective operations, and the GPU memory hierarchy and interconnect characteristics that govern them

Advanced proficiency in Kubernetes platform architectures, including operators, custom resources, controllers, and scheduling

Strong programming proficiency in Go and Python, with the ability to read, debug, and profile C++/CUDA using tools such as Nsight

Familiar with debugging/profiling multi-tier application workloads such as RAG, Agents

Excellent analytical, debugging, and problem-solving abilities

Preferred

Upstream contribution to vLLM, SGLang, TensorRT-LLM, llm-d, LMCache, or KServe

Disaggregated prefill/decode serving, or KV cache offload and reuse at scale

RDMA, GPUDirect Storage, InfiniBand, or RoCE

MIG, fractional GPU allocation, and multi-tenant GPU isolation

On-premises, air-gapped, or regulated enterprise software delivery

Experience and Education

Minimum of 8 years of experience in Software Engineering, including 1-2+ years working directly on LLM inference runtimes or production model serving

Degree in Computer Science or related field

Accessibility

HPE is committed to creating an inclusive and accessible workplace and encourages applications from all qualified individuals, including those with disabilities. If you believe you require accommodation during any stage of the application or interview process, please submit your request by completing our secure form linked here.

Note: This option is reserved for applicants needing assistance/reasonable accommodation related to a disability.

What We Can Offer You:

Health & Wellbeing

We strive to provide our team members and their loved ones with a comprehensive suite of benefits that supports their physical, financial and emotional wellbeing.

Personal & Professional Development

We also invest in your career because the better you are, the better we all are. We have specific programs catered to helping you reach any career goals you have — whether you want to become a knowledge expert in your field or apply your skills to another division.

Unconditional Inclusion

We are unconditionally inclusive in the way we work and celebrate individual uniqueness. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good.

Let's Stay Connected:

Follow @HPECareers on Instagram to see the latest on people, culture and tech at HPE.

Job: Engineering

Job Level: TCP_04

The expected salary/wage range for this position is provided below. Actual offer may vary from this range based upon geographic location, work experience, education/training, and/or skill level.

– United States of America: Annual Salary USD 144,000 - 273,000 in Colorado // 137,000 - 315,000 in North Carolina & Texas

The listed salary range reflects base salary. Variable incentives may also be offered.

Information about employee benefits offered in the US can be found at

The estimated job application period closure is December ; this timeline is provided for transparency and internal planning purposes.

HPE is an Equal Employment Opportunity/ Veterans/Disabled/LGBT employer. We do not discriminate on the basis of race, gender, or any other protected category, and all decisions we make are made on the basis of qualifications, merit, and business need. Our goal is to be one global team that is representative of our customers, in an inclusive environment where we can continue to innovate and grow together.

Hewlett Packard Enterprise is EEO Protected Veteran/ Individual with Disabilities.

HPE will comply with all applicable laws related to employer use of arrest and conviction records, including laws requiring employers to consider for employment qualified applicants with criminal histories.

#J-18808-Ljbffr

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

Govcio LLC

WhatJobs

Senior Dynamics Developer

washington, DC

$175,000.00 /Yr

Overview: GovCIO is looking for an experienced Senior CRM developer to support the continued enhancement and operations of a Microsoft Dynamics C…

Listing review due 2026-10-06View job

Govcio LLC

WhatJobs

Senior Application Developer

fairfax, VA

$165,000.00 /Yr

Overview: GovCIO is currently hiring a highly skilled Senior Application Developer with an active Secret clearance to design, build, and modify a…

Listing review due 2026-10-06View job

Govcio LLC

WhatJobs

Software Developer- SME

quantico, VA

$195,000.00 /Yr

Overview: GovCIO is currently hiring for a Software Developer-SME   to design, code, test, and maintain software applications and systems in supp…

Listing review due 2026-10-06View job

Govcio LLC

WhatJobs

Sr. Instructional Developer

fort meade, MD

$120,000.00 /Yr

Overview: GovCIO is currently hiring for an Instructional Developer  to design and develop training materials for personnel . This position will…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Senior Embedded Systems Developer

laurel, MD

See pay details in description

Description Do you love working on a motivated team to solve complex problems in innovative ways? Do you enjoy creating embedded prototypes i…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Oracle E-Business Suite (EBS) Developer

laurel, MD

See pay details in description

Description Are you a skilled problem-solver who blends technical expertise with creativity to build effective and efficient solutions? If s…

Listing review due 2026-10-06View job