waypointjobs

Blackrock-Neurotech

AI/ML Training Performance Engineer

salt lake city, UT

Check who can apply and the requirements below before continuing.

About this opportunity

Blackrock-Neurotech lists this AI/ML Training Performance Engineer opportunity in salt lake city, Utah. Review the employer’s description below for duties, qualifications and application requirements.

Job description

Build the systems that expand human capability

At Blackrock Neurotech, we’ve spent decades making the impossible possible – helping people move, speak, and reconnect with the world when they otherwise could not. We’ve seen that restoring function restores more than ability. It restores independence, identity, and agency.

Today, we are building the next generation of human capability: brain-computer interfaces that are designed to be safe, scalable, and trusted in the real world. Our work is not only about reconnecting people to what was lost, but about expanding what is possible – creating a seamless interface between human intent and technology.

This is foundational work in a category-defining field. You will help build the infrastructure for a future where neural interfaces are invisible, reliable, and deeply human-centered.

Working at Blackrock Neurotech means:

Owning meaningful, high-impact problems at the frontier of science and engineering

Building alongside experienced, thoughtful peers across disciplines

Solving technically complex challenges grounded in real human outcomes

Contributing to a culture that values rigor, clarity, and long-term thinking over noise

The Role

The ML Training Performance Engineer will own the efficiency and scalability of training models on GPU and cloud infrastructure. You will turn available compute into faster, more capable experiments as our model training scales in complexity and compute requirements.

As a hands‑on individual contributor on a small research team, you will work across the training stack, from Python and model execution to GPU kernels, distributed communication, and runtime environments. Partnering with model researchers, data engineers, and infrastructure and IT teams, you will identify and implement performance improvements while preserving numerical correctness and scientific intent.

You will have significant ownership over how we measure, optimize, and scale training performance, establishing the baselines, tooling, and technical approaches that will support our AI/ML work as it grows.

What You’ll Do

Own training performance across single‑GPU, multi‑GPU, and multi‑node workloads, establishing reproducible baselines for throughput, memory use, utilization, time to target quality, and cost

Profile the full training path to distinguish compute, memory, communication, CPU, and I/O bottlenecks and prioritize changes with measurable end‑to‑end impact

Optimize tensor layouts, precision, memory allocation, activation checkpointing, operator fusion, and execution graphs to fit larger or longer‑context models within available resources

Write, tune, and validate custom GPU kernels using CUDA, Triton, HIP/ROCm, or the appropriate platform tools when existing implementations limit performance

Improve Python training scripts, framework and compiler settings, batching, gradient accumulation, and optimizer execution while preserving intended training behavior

Design and tune distributed training strategies, including data, tensor, pipeline, or sharded parallelism, based on model structure, memory limits, and interconnect topology

Partner with model researchers on hardware‑aware architecture and hyperparameter changes, measuring their effects on convergence, model quality, and compute requirements

Coordinate with the neural data infrastructure engineer on prefetching, pinned memory, host‑to‑device transfer, and I/O overlap so data delivery keeps pace with training

Work with infrastructure and IT on GPU selection, cloud instance configurations, networking, drivers, containers, scheduling, and capacity planning as training needs and compute capacity scale

Build robust checkpoint, restart, and recovery workflows and performance regression checks that keep long‑running experiments reproducible and productive

Communicate benchmark evidence, numerical tradeoffs, scaling limits, and resource recommendations clearly to researchers and organizational stakeholders

What You Bring

Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field, or equivalent practical experience with 3+ years of relevant industry experience OR Master’s degree with 1+ years of relevant industry experience OR PhD in a related field

Demonstrated experience improving the performance of substantial deep learning training workloads, with measured gains in speed, memory efficiency, or compute costI

A measurement‑driven approach to performance optimization, using profiling and benchmarking to validate meaningful end‑to‑end improvements

Exceptional programming ability in Python and C++ or a comparable systems language, with strong debugging, testing, and performance analysis practices

Deep understanding of GPU execution, including memory hierarchies, memory coalescing, thread blocks, warps or wavefronts, occupancy, synchronization, and bandwidth limits

Hands‑on experience developing and profiling GPU kernels with CUDA or HIP/ROCm, and the ability to diagnose correctness and performance at the hardware level

Deep knowledge of PyTorch or an equivalent framework, including automatic differentiation, computation graphs, tensor storage, compilation, and mixed‑precision training

Strong understanding of deep learning architectures and the underlying computations that drive training performance

Experience with distributed training, collective communication, sharding, and the interaction between model partitioning and GPU interconnects

Strong understanding of numerical stability and the ability to validate gradients, convergence, and model quality after performance changes

Experience configuring and diagnosing Linux‑based GPU environments, containers, cloud compute, and high‑throughput storage or networking

Ability to collaborate closely with researchers and infrastructure teams and make clear tradeoffs between implementation effort, performance reliability, and scientific value

Experience with Triton, compiler optimization, advanced GPU profiling tools, multiple accelerator generations, long‑sequence or multimodal models, neural time series, or large‑scale model training is a plus Working Location:

This is an on‑site role based at Blackrock Neurotech's headquarters in Salt Lake City, Utah. Occasional travel may be required.

How We Work

We are a small, experienced team working on consequential problems.

We take ownership of outcomes and follow through with clarity and accountability

We prioritize sustained, high‑quality work over performative urgency

We value rigor, sound judgement and thoughtful decision‑making

We collaborate deliberately: low ego, high trust and high context

This is a high‑ownership role, but it is not an "always-on" one. We expect strong work and our people to have a life outside of it.

#J-18808-Ljbffr

Worksite address

salt lake city, UT, 84193, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

U.S. Army Corps of Engineers

USAJOBS

Interdisciplinary Waterways Maintenance Chief

Portland, OR

$114,684.00 – $149,091.00 per yearfull time

About the Position: As Chief of the Waterways Maintenance Section, the incumbent exercises full technical, administrative, and managerial authori…

Closes 2026-10-16View job

U.S. Army Corps of Engineers

USAJOBS

ENGINEERING TECHNICIAN (CIVIL)

Philadelphia, PA

$44,887.00 – $106,982.00 per yearfull time

About the position: Serves as cost estimator responsible for the preparation of budget, control and detailed final estimates of cost pertaining t…

Closes 2026-10-19View job

U.S. Coast Guard

USAJOBS

NAVAL ARCHITECT

Washington, DC

$121,785.00 – $158,322.00 per yearfull time

This vacancy is for a GS-0871-13, NAVAL ARCHITECT located in the Department of Homeland Security, U.S. Coast Guard, USCG MARINE SAFETY CENTER in …

Closes 2026-10-14View job

U.S. Coast Guard

USAJOBS

NAVAL ARCHITECT

Washington, DC

$121,785.00 – $158,322.00 per yearfull time

This vacancy is for a GS-0871-13, NAVAL ARCHITECT located in the Department of Homeland Security, U.S. Coast Guard, USCG MARINE SAFETY CENTER in …

Closes 2026-10-14View job

U.S. Army Corps of Engineers

USAJOBS

Interdisciplinary

Baltimore, MD

$102,415.00 – $133,142.00 per yearfull time

About the Position: You will be responsible for the environmental assessment of hazardous, toxic, and radiological waste (HTRW) sites and Militar…

Closes 2026-10-15View job

U.S. Army Corps of Engineers

USAJOBS

Student Trainee (Engineering and Architecture)

Lakewood, CO

$36,464.00 – $89,470.00 per yearfull time

About the Position: Position(s) will be filled under the Department of the Army Pathways Internship "Indefinite" Program. Click here for more inf…

Closes 2026-10-12View job