waypointjobs

Career Techniques

Sr HPC Hardware Engineer

dallas, TX

Check who can apply and the requirements below before continuing.

About this opportunity

Career Techniques lists this Sr HPC Hardware Engineer opportunity in dallas, Texas. Review the employer’s description below for duties, qualifications and application requirements.

Job description

RESPONSIBILITIES

Design, configure, and manage a high-performance compute fleet comprising large-scale GPU (NVIDIA V100/A100/H200/GB200) and CPU nodes across the firm's infrastructure.

Own the full firmware and BIOS lifecycle across the HPC/AI fleet — from establishing baselines and validation through rollout, compliance, and ongoing maintenance.

Lead troubleshooting of hardware components including CPUs, GPUs, DPUs, NVSwitches, NICs, memory, PSUs, and BMCs; drive component replacement and configuration remediation.

Automate health checks, onboarding workflows, and recurring hardware issue remediation to accelerate safe deployment and reduce recovery time.

Validate and operationalize next-generation AI platforms (e.g., NVL72 / Grace Blackwell) from day one, ensuring stability, performance readiness, and production fitness.

Collaborate with vendors on firmware and hardware issues, providing clear reproduction cases, diagnostic logs, and business impact to drive timely resolution.

Perform hardware performance analysis, tuning, and capacity planning to ensure reliable scale-out of the compute environment.

Define and implement security hardening best practices for hardware infrastructure, maintaining platform integrity across the fleet.

Leverage Infrastructure as Code (IaC) methodologies and scripting to drive efficient, repeatable, and scalable infrastructure management.

Mentor junior engineers, act as a subject matter expert for infrastructure-related escalations, and champion a culture of continuous improvement across the team.

REQUIREMENTS

Bachelor’s degree in Electrical Engineering, Computer Engineering, or a related field, or equivalent hands‑on experience.

8+ years of experience managing large-scale HPC or AI compute infrastructure in a production environment.

Deep knowledge of server hardware architecture, including processors, memory, storage, networking, power systems, and thermal management.

Hands‑on experience with bare-metal provisioning, firmware and BIOS lifecycle management, and hardware automation tools such as Ansible, Puppet, or Chef.

Proficiency with Redfish API and BMC/IPMI tooling (iDRAC, iLO) for remote hardware management and diagnostics.

Demonstrated ability to troubleshoot and resolve complex hardware issues across GPU and CPU nodes, including NVIDIA-SMI and GPU diagnostics.

Experience with hardware monitoring platforms, performance tuning, and capacity planning at scale.

Familiarity with Linux-based environments and scripting proficiency in Python, Bash, or PowerShell for infrastructure automation.

Experience with OpenStack (particularly Ironic) or equivalent cloud/bare-metal provisioning platforms is strongly preferred.

Strong cross-functional communication skills and proven ability to collaborate effectively with software, networking, and vendor teams.

Prior technical leadership experience, including mentoring engineers and driving team-wide best practices.

#J-18808-Ljbffr

Worksite address

dallas, TX, 75215, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Space Systems Mechanical Engineer

laurel, MD

See pay details in description

Description Do you want to design and build unique space structures and spacecraft for NASA missions that enable groundbreaking scientific disco…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

System Security Engineer

laurel, MD

See pay details in description

Description Are you looking for an opportunity to utilize your technical skills to solve complex, real-world problems? If so, we're looking …

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Advanced Reentry Mission Engineer

laurel, MD

See pay details in description

Description Are you interested in hypersonic and reentry system design and prototyping? Do you want to make contributions to next generation …

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Thermal and EO/IR Modeling and Simulation Engineer

laurel, MD

See pay details in description

Description Are you looking for a unique opportunity to impact significant advances to the nation's groundbreaking integrated air and missile de…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Network Effects Engineer

laurel, MD

See pay details in description

Description Do you want to perform advanced research, development, and test & evaluation of communications systems and network technologies that…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

System Realization and Resilience Engineer

laurel, MD

See pay details in description

Description Are you passionate about applying system engineering principles to influence the development and resilience of future strategic weap…

Listing review due 2026-10-06View job