waypointjobs

Cerence

Senior Principal AI Engineer

Remote — United States (see country and timezone requirements)

Check who can apply and the requirements below before continuing.

This job is closed

Applications are no longer available for this announcement. Explore current related opportunities below.

Job description

A Moving Experience.

What You Will Work On

Design andoperatedistributed training systems for large neural networks(autoregressive, diffusion,State Space Modelsetc.)across GPU clusters

Optimisemulti‑node, multi‑GPU execution tomaximizethroughput andutilization

Diagnose&resolve bottlenecks across compute, memory, and network

Improve training stability and fault tolerance at scale

Partner with research and applied ML teams to productionizelarge‑modeltraining pipelines

Core Responsibilities

Distributed Training Infrastructure

Build andoptimizeGPU cluster orchestration using:

Slurm

Kubernetes

Ray

RunAI

Ensure efficient scheduling, isolation, and fairness across training workloads

Communication & Networking

Optimizeand debug distributed communication using:

NCCL

RDMA

InfiniBand

NVLink

Minimizenetworking bottlenecks that dominateend‑to‑endtraining time

Training Frameworks

Scale large-model training using:

PyTorchDistributed

Megatron‑LM

DeepSpeed

Ownmulti‑nodelaunch configurations, failure recovery, and performance tuning

Memory & Performance Optimization

Apply advanced memory optimization techniques:

Activation checkpointing

ZeRO(Stage 1–3) and offload strategies

Balance compute, memory, and communication to push model size and batch scale

What Success Looks Like

GPUutilizationconsistently stays high (>80–90%)

Training scales cleanly from single node to dozens or hundreds of GPUs

Communication overhead is minimized and predictable

Large training jobs run stably for days or weeks without failure

New models can be trained faster, larger, and more reliably than before

Required Experience & Skills

Strongly Required

Deephands‑onexperience with distributed systems or ML systems

Experience runninglarge‑scaleworkloads on GPU clusters

Production experience withPyTorchdistributed training

Strong understanding of parallelism strategies (data, tensor, pipeline parallelism)

Low‑levelunderstanding of GPU communication and networking

Critical Technical Skills

GPU orchestration:Slurm, Kubernetes, Ray,RunAI

Communication libraries: NCCL, RDMA, InfiniBand,NVLink

Training frameworks:PyTorchDistributed,Megatron‑LM,DeepSpeed

Memoryoptimisation: activation checkpointing,ZeROoffload techniques

Common ProblemsYou’llBe Solving

Many teams fail at scale because:

GPUutilizationis low despite large clusters

Networking and communication dominate training time

Training jobs crash or become unstable at large scale

You will be explicitly focused oneliminatingthese failure modes.

Ideal Background

This role is a strong fit for individuals who have worked as:

ML Systems Engineer

Distributed Systems Engineer

AI Infrastructure Engineer

HPC Engineer transitioning into ML

Experience working with large language models or foundation models is a strong plus, but deep systemsexpertiseis valued over pure model architecture experience.

Why This Role Matters

Without robust distributed training infrastructure, progress on largemodelsstalls. This role directly enables:

Larger models

Faster iteration cycles

More reliable research-to-production pipelines

You will be building the foundation that makeslarge‑scaleAI possible.

Cerence Inc. (Nasdaq: CRNC and ) is the global industry leader in creating unique, moving experiences for the automotive world. Spun out from Nuance in October 2019, Cerence is a new, independent company that has quickly gained traction as a leader in the automotive voice assistant space, working with all of the world’s leading automakers – from Ford and Fiat Chrysler to Daimler, Audi and BMW to Geely and SAIC – to transform how a car feels, responds and learns. Its track record is built on more than 20 years of industry experience and leadership and more than 500 million cars on the road today across more than 70 languages.

As Cerence looks to the future and continues an ambitious growth agenda, we need someone to join the team and help build the future of voice and AI in cars. This is an exciting opportunity to join Cerence’s passionate, dedicated, global team and be a part of meaningful innovation in a rapidly growing industry.

EQUAL OPPORTUNITY EMPLOYER

Cerence is firmly committed to Equal Employment Opportunity (EEO) and to compliance with all federal, state and local laws that prohibit employment discrimination on the basis of age, race, color, gender, gender identity, gender expression, sex, sex stereotyping, pregnancy, national origin, ancestry, religion, physical or mental disability, medical condition, marital status, citizenship status, sexual orientation, protected military or veteran status, genetic information and other protected classifications. Cerence Equal Employment Opportunity Policy Statement.

All prospective and current Employees need to remain vigilant when it comes to executing security policies in the workplace. This includes:

- Following workplace security protocols and training programs to familiarize with the ways to maintain a safe workplace.

- Following security procedures to report any suspicious activity.

- Having respect for corporate security procedures to allow those procedures to be effective.

- Adhering to company's compliance and regulations.

- Encouraging to follow a zero tolerance for workplace violence.

- Basic knowledge of information security and data privacy requirements (e.g., how to protect data & how to be handling this data).

- Demonstrative knowledge of information security through internal training programs.

Originally posted on Himalayas

Who can apply

Eligible countries: United States. Accepted UTC offsets: UTC-10, UTC-9, UTC-8, UTC-7, UTC-6, UTC-5, UTC+14. Review the full description for employer-specific work authorization, residency and schedule requirements.

Explore related searches

Current related jobs

Spreedly

Jobicy

Senior Database Engineer

Remote — USA

Salary not specifiedRemote

About Us: Spreedly is the Intelligent Payments Platform for businesses that want to own their payments strategy. Founded in 2007 and headquarter…

Last received from source 2026-10-07View job

Quince

Jobicy

Paid Social Video Producer

Remote — USA

Salary not specifiedRemote

ABOUT QUINCE Quince is a destination for builders, creators, innovators, and operators who want to come together and challenge the status quo. O…

Last received from source 2026-10-07View job

Eight Sleep

Jobicy

Senior Backend Engineer

Remote — Canada, USA

Salary not specifiedRemote

Join the Sleep Fitness Movement At Eight Sleep, we’re on a mission to fuel human potential through optimal sleep. As the world’s first sleep fit…

Last received from source 2026-10-07View job

Stripe

Jobicy

Credit Operations Manager

Remote — USA

Salary not specifiedRemote

Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies - from the world’s largest enterpr…

Last received from source 2026-10-07View job