About this opportunity
Oho Group lists this Principal Inference Engineer opportunity in san francisco, California. Review the employer’s description below for duties, qualifications and application requirements.
Job description
We’re supporting an advanced-compute company building a new hardware and software platform for AI workloads.
They’re looking for an inference specialist who can optimize the complete path from transformer models through execution engines, compilers, runtimes and kernels to multi-accelerator systems.
What you’ll work on
Architect high-performance transformer inference on a new compute platform
Optimize model execution, graph transformations and runtime behavior
Develop or guide performance-critical C++, CUDA or Triton components
Improve attention, GEMM, MoE and other critical execution paths
Design KV-cache, batching, memory-management and decoding strategies
Optimize tensor, pipeline and expert parallelism
Analyze multi-device and multi-node inference performance
Drive improvements in latency, throughput, utilization and cost per token
What we’re looking for
Principal, Distinguished or equivalent senior technical scope
Direct optimization of transformer or generative-AI inference
Hands-on low-level implementation in C++, CUDA, Triton or similar
Deep expertise across multiple connected layers of the inference stack
Strong profiling and performance-debugging skills
Experience taking optimizations beyond isolated kernels into complete systems
#J-18808-Ljbffr
Worksite address
san francisco, CA, 94199, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.