This job is closed
Applications are no longer available for this announcement. Explore current related opportunities below.
Job description
Compensation: Competitive Base Salary + Performance Bonus
Overview
Our client is seeking an HPC & AI Solutions Architect to lead the technical design, integration, and delivery of high-performance computing and AI infrastructure solutions.
This is a highly technical, customer-facing role focused on designing scalable architectures across GPU/CPU compute, storage, networking, Kubernetes, orchestration, and security . The position spans the full solution lifecycle from technical discovery and workload analysis through proof-of-concept, deployment, and ongoing optimization.
The ideal candidate brings deep HPC and AI infrastructure expertise, strong hands-on system design and performance tuning experience, and the ability to translate complex customer requirements into scalable, production-ready solutions.
Key Responsibilities
Customer Engagement & Technical Discovery
Work directly with customers to understand workload requirements, performance targets, and technical objectives.
Lead technical discovery sessions focused on application behavior, bottlenecks, scalability, and infrastructure requirements.
Serve as a trusted technical advisor throughout the solution lifecycle.
Design end-to-end HPC and AI architectures across compute, storage, networking, orchestration, and security.
Recommend hardware and software solutions aligned with performance, scalability, reliability, and efficiency goals.
Develop architecture blueprints, integration plans, and technical documentation.
Design solutions supporting GPU-intensive AI/ML, LLM, and advanced compute workloads.
Performance & Workload Optimization
Support proof-of-concept, benchmarking, and performance-validation initiatives.
Perform workload profiling, system tuning, and infrastructure optimization.
Identify bottlenecks across compute, storage, networking, and orchestration layers.
Recommend improvements that increase workload performance, scalability, and resilience.
Implementation & Delivery
Provide technical leadership during deployment and integration.
Partner with customers and internal Engineering, Product, and Operations teams throughout implementation.
Support solutions from architecture through production deployment and optimization.
Troubleshoot complex infrastructure and workload issues during delivery.
Technical Leadership
Maintain expertise across emerging HPC, AI, GPU, storage, networking, and orchestration technologies.
Build relationships with technology partners across GPU, networking, and storage ecosystems.
Contribute to reference architectures, reusable design patterns, and technical best practices.
Lead customer workshops, architecture reviews, and technical presentations.
Required Qualifications
Strong technical expertise across:
GPU and CPU architectures
NVIDIA / CUDA ecosystem
Slurm and Kubernetes
InfiniBand, RDMA, and RoCE
Lustre, GPFS / Spectrum Scale, Ceph, VAST, or similar storage platforms
Kubernetes and container orchestration
Identity, encryption, and infrastructure security
Strong Linux systems knowledge, including tuning and performance analysis.
Experience translating workload requirements into detailed technical architectures.
Experience with proof-of-concept, benchmarking, or workload optimization.
Strong customer-facing communication and presentation skills.
Ability to work effectively with engineering, product, operations, and executive stakeholders.
Preferred Experience
AI/ML, LLM, GPU, or HPC workloads.
NVIDIA GPU infrastructure.
Automation and Infrastructure-as-Code.
Workload migration and performance engineering.
Next-generation GPU and high-speed interconnect technologies.
Bachelor's or Master's degree in Computer Science, Engineering, Physics, or related field.
Relevant cloud, Linux, networking, Kubernetes, or security certifications.
#J-18808-Ljbffr
Worksite address
dallas, TX, 75215, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.