waypointjobs

Lever, Inc.

Senior Distributed Systems Engineer

sunnyvale, CA

Check who can apply and the requirements below before continuing.

About this opportunity

Lever, Inc. lists this Senior Distributed Systems Engineer opportunity in sunnyvale, California. Review the employer’s description below for duties, qualifications and application requirements.

Job description

About the Institute of Foundation Models

The Institute of Foundation Models (IFM) designs and operates ultra-scale GPU supercomputing systems to train next-generation foundation models. We believe performance, fault tolerance, and scalability are co-designed across model architecture, communication systems, runtime, and hardware topology.

This role sits at the core of that effort — driving communication performance, distributed reliability, and cross-layer optimization for large-scale training workloads.

The Mission

We are looking for a deeply technical engineer to co-design and optimize the communication stack for large-scale distributed training, including hybrid parallelism and Mixture-of-Experts (MoE) workloads.

This is not a network operations role. This is a systems-level engineering position focused on performance engineering, distributed debugging, and communication-runtime co-design.

Design and optimize expert-parallel and hybrid-parallel communication patterns

Drive high-performance hierarchical collectives for MoE workloads

Co-design runtime orchestration with communication topology awareness

Reduce tail latency and improve determinism across thousands of GPUs

Architect fault-tolerant distributed execution under real‑world cluster failures

Core Technical Scope

Communication-compute overlap and topology-aware collective optimization

Deep debugging of NCCL, RDMA, and custom communication layers

Hybrid expert parallel strategies in modern large-scale MoE systems

Elastic and resilient distributed job orchestration concepts

Congestion analysis and routing optimization across InfiniBand/RoCE fabrics

Microbenchmarking and performance modeling for communication-heavy workloads

Expected Technical Depth

Hybrid expert parallel communication for Mixture-of-Experts training

Scaling behavior under network pressure

Distributed orchestration for elastic, large-scale training

Fault detection and recovery in distributed GPU workloads

Cross-layer bottlenecks: GPU NIC PCIe NVSwitch Fabric Scheduler

Required Background

Experience optimizing distributed training at 1,000+ GPU scale (or equivalent depth)

Hands‑on expertise with RDMA, InfiniBand, RoCE, and GPUDirect RDMA

Deep familiarity with NCCL and/or UCX internals

Strong systems programming ability (C/C++, Rust, or Go)

Strong familiarity with modern model training frameworks such as PyTorch

Ability to troubleshoot and profile training performance issues related to communication bottlenecks

Ability to translate research ideas into production‑grade optimizations

Experience debugging distributed hangs, desynchronization, and performance regressions

What We Mean by "Hardcore"

You can explain why an communication degrades at scale and how to fix it

You have improved real cluster throughput via communication redesign

You can trace a distributed hang across ranks and identify the root cause

You are comfortable working at the boundary between hardware and runtime

Application Requirements

Include a link to your GitHub (required)

Provide links to relevant distributed systems, HPC, or large-scale training projects

Include a list of publications and/or public technical reports (if applicable)

Describe the hardest distributed debugging problem you solved

#J-18808-Ljbffr

Worksite address

sunnyvale, CA, 94087, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Space Systems Mechanical Engineer

laurel, MD

See pay details in description

Description Do you want to design and build unique space structures and spacecraft for NASA missions that enable groundbreaking scientific disco…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

System Security Engineer

laurel, MD

See pay details in description

Description Are you looking for an opportunity to utilize your technical skills to solve complex, real-world problems? If so, we're looking …

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Advanced Reentry Mission Engineer

laurel, MD

See pay details in description

Description Are you interested in hypersonic and reentry system design and prototyping? Do you want to make contributions to next generation …

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Thermal and EO/IR Modeling and Simulation Engineer

laurel, MD

See pay details in description

Description Are you looking for a unique opportunity to impact significant advances to the nation's groundbreaking integrated air and missile de…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Network Effects Engineer

laurel, MD

See pay details in description

Description Do you want to perform advanced research, development, and test & evaluation of communications systems and network technologies that…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

System Realization and Resilience Engineer

laurel, MD

See pay details in description

Description Are you passionate about applying system engineering principles to influence the development and resilience of future strategic weap…

Listing review due 2026-10-06View job