waypointjobs

Confidential

MLOps Engineer, LLM Systems

Elko, MN

Check who can apply and the requirements below before continuing.

About this opportunity

Confidential lists this MLOps Engineer, LLM Systems opportunity in Elko, Minnesota. Review the employer’s description below for duties, qualifications and application requirements.

Job description

Role Overview

Help develop foundational large language models by creating and evaluating rigorous ML systems work that improves frontier AI training data. This hands-on infrastructure role focuses on GPU kernels, performance profiling and trace analysis, accelerated and distributed workload debugging, and high-throughput inference serving.

Key Responsibilities

Design challenging, domain-relevant MLOps and ML systems tasks in GPU kernels, profiling, debugging, and inference serving, then produce accurate, well-structured solutions.

Evaluate technical tasks and solutions, providing clear written feedback that can withstand detailed review.

Support research and engineering teams in closing knowledge gaps and improving model performance across ML systems, training infrastructure, and framework-level subjects.

Create detailed guidelines and evaluation rubrics for kernel optimization, profiler-output interpretation, distributed-systems reasoning, and serving throughput and latency trade-offs.

Partner with subject matter experts to maintain consistent, accurate training data.

Qualifications

At least 2 years of hands-on professional experience in ML systems, ML infrastructure, model serving, or GPU and accelerator performance engineering.

Experience in at least one of the following areas, with experience across multiple areas strongly preferred: custom GPU kernel development or optimization using CUDA, Triton, or Pallas; profiling and trace analysis using Kineto, torch.profiler, Nsight, XLA, or JAX profiler; debugging distributed or accelerator-bound workloads; or serving LLMs at scale using vLLM, SGLang, TensorRT-LLM, Ray Serve, KV cache, paged attention, or continuous batching.

Production experience with JAX and/or PyTorch. Framework-level expertise in custom operators, distributed training with FSDP, DDP, DeepSpeed, or Megatron, or compiler and graph-level work is preferred.

Familiarity with A100, H100, B200, or TPU accelerators, including the ability to assess throughput, latency, and memory trade-offs.

Demonstrable career progression, strong written communication, and the ability to explain complex technical decisions clearly.

This is a systems-focused position, not an applied modeling or data science role.

Work Terms

Hourly W-2 employment with placement on an extended workforce team supporting a leading AI lab.

Full-time, 40-hour-per-week weekday commitment. Candidates must be able to work without conflicting engagements or other concurrent work commitments.

Available to candidates located in Canada, the United Kingdom, or the United States.

Compensation

$90 to $120 per hour.

Equal Employment Opportunity

Employment decisions are made without discrimination based on race, religion, color, national origin, sex, including pregnancy, childbirth, reproductive-health decisions, or related medical conditions, sexual orientation, gender identity or expression, age, protected veteran status, disability, genetic information, political views or activity, or any other legally protected characteristic.

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

Annapurna Labs (u.s.) Inc.

WhatJobs

PD Engineer, Annapurna Labs

cupertino, CA

Salary not specified

As a member of the Cloud-Scale Machine Learning Acceleration team you'll be responsible for the design and optimization of Hardware in our data c…

Last received from source 2026-10-07View job

Amazon.com Services Llc - A57

WhatJobs

Senior Automation Engineer

suffolk, VA

Salary not specified

Operations is at the heart of Amazon's business. We are known for our speed, accuracy, and exceptional service. Our buildings deliver tens of tho…

Last received from source 2026-10-07View job

Annapurna Labs (u.s.) Inc.

WhatJobs

DFT Design Engineer, Machine Learning Acceleration

austin, TX

Salary not specified

Custom SoCs (System on Chip) are at the heart of AWS Machine Learning servers. As a member of the Cloud-Scale Machine Learning Acceleration team,…

Last received from source 2026-10-07View job