waypointjobs

Realmlabs

Software Engineer, ML Infrastructure

sunnyvale, CA

Check who can apply and the requirements below before continuing.

About this opportunity

Realmlabs lists this Software Engineer, ML Infrastructure opportunity in sunnyvale, California. Review the employer’s description below for duties, qualifications and application requirements.

Job description

Role Overview

We are hiring a Founding ML Infrastructure Engineer to own the end-to-end deployment, optimization, and operation of our suits of models in production.

This is a core founding role focused on building and operating production-grade LLM systems . You will apply deep knowledge of model internals to deploy, optimize, and run modern LLMs at scale , owning performance end-to-end across latency, throughput, and reliability .

You will design and operate the full ML serving stack from model artifacts to GPU execution, and work closely with Product and ML teams to ensure our models can support high QPS, strict SLAs, and production correctness .

This role is ideal for someone who deeply understands how LLMs work internally but chooses to specialize in making them fast, stable, and production-ready .

About Realm Labs

Realm Labs is an AI trust and security startup. We help enterprises detect, debug, and prevent AI’s misbehaviors in production. We are backed by top VCs and serve some of the most iconic global enterprises.

Key Responsibilities

Own the end-to-end LLM inference stack , including:Model loading and execution

GPU utilization and memory efficiency

Runtime performance tuning

Production deployment and scaling

Design and operate high-performance LLM serving systems using technologies such as:vLLM, TensorRT / TensorRT-LLM, Triton Inference Server, SGLang

Optimize inference across:Latency

Throughput (QPS)

GPU memory footprint

Cost efficiency

Work hands-on with PyTorch and TensorFlow models , including:Model graph understanding

Attention mechanisms, KV cache behavior, batching strategies

Precision tradeoffs (FP16, BF16, INT8, etc.)

Build and maintain production-grade GPU services :

Multi-model serving

Autoscaling strategies

Fault isolation and graceful degradation

Collaborate with application and platform teams to:Define serving APIs

Ensure correctness and safety of outputs

Debug production issues end-to-end

Build a reproducible model training and versioning system for customer deployments

Establish best practices for:Model versioning

Rollouts and rollbacks

Performance benchmarking

Production validation

Expected Qualifications

5+ years of professional experience in ML infrastructure, systems engineering, or production ML roles.

Strong software engineering fundamentals; ability to write robust, maintainable production code .

Deep hands-on experience with LLM inference infrastructure , including:PyTorch (required)

TensorFlow (working knowledge)

Proven experience with GPU inference optimization , including:TensorRT / TensorRT-LLM

vLLM

Triton Inference Server

SGLang or similar serving runtimes

Strong understanding of LLM internals , such as:Transformer architectures

Attention and KV caching

Batching, streaming, and token-level generation

Experience running ML systems in production with high traffic and SLAs.

Comfortable working in Linux-based, cloud production environments .

Preferred Qualifications

Experience deploying LLMs on Kubernetes and GPU clusters.

Familiarity with CUDA, NCCL , or low-level GPU performance concepts.

Experience with:Model sharding and parallelism strategies

Multi-GPU inference

Streaming inference systems

Knowledge of observability for ML systems (metrics, latency breakdowns, GPU monitoring).

Experience working at startups or owning systems with minimal abstraction layers.

Additional Information

This is a founding, high-ownership role with direct impact on core product capabilities.

You will be expected to build, run, and own systems end-to-end .

The role may include limited on-call responsibilities aligned with production ownership.

Compensation & Benefits

Market aligned compensation and benefits.

Founding engineer equity (Equity is a significant component of this role and will be discussed ).

Medical, Dental, Vision, Life insurance, 401-K, In-office lunch etc.

Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and candidate. But if we make you an offer, we will make all reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.

#J-18808-Ljbffr

Worksite address

sunnyvale, CA, 94087, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Senior Embedded Systems Developer

laurel, MD

See pay details in description

Description Do you love working on a motivated team to solve complex problems in innovative ways? Do you enjoy creating embedded prototypes i…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Oracle E-Business Suite (EBS) Developer

laurel, MD

See pay details in description

Description Are you a skilled problem-solver who blends technical expertise with creativity to build effective and efficient solutions? If s…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Senior Software Engineer

laurel, MD

See pay details in description

Description Are you passionate about building solutions for our greatest national security challenges? Are you searching for engaging work wi…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Reverse Engineer and AI Workflow Developer

laurel, MD

See pay details in description

Description Are you passionate about reverse engineering complex software, firmware, and hardware? Do you enjoy developing agentic AI workflo…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

AFSIM Developer

laurel, MD

See pay details in description

Description Are you passionate about building high-fidelity simulations that shape the future of U.S. Space Force capabilities and force design …

Listing review due 2026-10-06View job