waypointjobs

Google

Staff Software Engineer, ML Performance, GPU

sunnyvale, CA

Check who can apply and the requirements below before continuing.

About this opportunity

Google lists this Staff Software Engineer, ML Performance, GPU opportunity in sunnyvale, California. Review the employer’s description below for duties, qualifications and application requirements.

Job description

In most instances, this position requires in-person interviews as part of the hiring process.

Minimum qualifications

Bachelor's degree or equivalent practical experience.

8 years of experience in software development.

5 years of experience with ML design and ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).

Experience with modern GPU architectures, memory hierarchies, and performance bottlenecks.

Experience with low-level GPU programming (CUDA, Triton, CUTLASS, etc.) and performance engineering techniques.

Experience with modern LLMs and their deployment on AI accelerators.

Preferred qualifications

Master's degree or PhD in Engineering, Computer Science, or a related technical field.

8 years of experience with data structures and algorithms.

3 years of experience in a technical leadership role leading project teams and setting technical direction.

3 years of experience working in a complex, matrixed organization involving cross-functional, or cross-business projects.

Experience in hardware-aware algorithm design and compiler stacks (e.g., OpenXLA), tailoring large-scale ML models and distributed systems for peak performance across accelerator hardware.

About The Job

Google's software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search. We're looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to Google's needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward.

While known for pioneering work with TPUs, GPUs are an equally vital and rapidly expanding frontier within Google's ML infrastructure. GPUs are indispensable to Google's ever-evolving landscape for strategic, pragmatic, and performance-driven reasons - ensuring top performance for our ML models, adapting to ML workloads, achieving results, and influencing next-gen GPU architectures via partnerships.

Core ML's GPU Performance team is responsible for optimizing, modeling, and evaluating GPU systems for comparative analysis and benchmarking for internal and external ML workloads. Our team's focus on performance analysis and optimization identifies opportunities in Google production and research ML workloads and lands optimizations to entire fleet. We evaluate current and future ML workloads and runs performance/total cost of ownership simulations to collect roofline estimates and guide decision-making for the hardware teams.

Behind everything our users see online is the architecture built by the Technical Infrastructure team to keep it running. From developing and maintaining our data centers to building the next generation of Google platforms, we make Google's product portfolio possible. We're proud to be our engineers' engineers and love voiding warranties by taking things apart so we can rebuild them. We keep our networks up and running, ensuring our users have the best and fastest experience possible.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $ - $ (USD) + 20% bonus target + equity + benefits

Responsibilities

Identify and maintain LLM training and serving benchmarks; use them to identify performance opportunities, drive XLA:GPU/Triton performance and guide XLA releases.

Partner with product teams (e.g., Google DeepMind) to onboard, optimize, and scale LLMs and machine learning models on GPU hardware.

Conduct architecture-level simulations, performance benchmarking, and roofline analyses using tools like TRT-LLM, vLLM, and SGLang to guide system designs.

Analyze fleet-wide performance and efficiency metrics to identify bottlenecks and engineer scalable optimizations across Google's infrastructure.

Research and implement model/data efficiency techniques, tooling, and profiling mechanisms to improve workload performance and training efficiency.

Google is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also Google's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form.

#J-18808-Ljbffr

Worksite address

sunnyvale, CA, 94087, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Reverse Engineer and AI Workflow Developer

laurel, MD

See pay details in description

Description Are you passionate about reverse engineering complex software, firmware, and hardware? Do you enjoy developing agentic AI workflo…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

AFSIM Developer

laurel, MD

See pay details in description

Description Are you passionate about building high-fidelity simulations that shape the future of U.S. Space Force capabilities and force design …

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Senior Software Engineer/Architect - TS SCI cleared

laurel, MD

See pay details in description

Description Are you passionate about building solutions for our greatest national security challenges? Are you searching for engaging work wi…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Senior Software Engineer

laurel, MD

See pay details in description

Description Are you passionate about building solutions for our greatest national security challenges? Are you searching for engaging work wi…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Senior Embedded Systems Developer

laurel, MD

See pay details in description

Description Do you love working on a motivated team to solve complex problems in innovative ways? Do you enjoy creating embedded prototypes i…

Listing review due 2026-10-06View job