waypointjobs

ServiceNow

Staff Machine Learning Systems & Reliability Engineer (Moveworks)

mountain view, CA

Check who can apply and the requirements below before continuing.

About this opportunity

ServiceNow lists this Staff Machine Learning Systems & Reliability Engineer (Moveworks) opportunity in mountain view, California. Review the employer’s description below for duties, qualifications and application requirements.

Job description

We are building AI-enabled product capabilities that improve through data, feedback, and real-world use

We need the production systems that make those capabilities dependable: repeatable delivery, measurable quality, controlled learning loops, and reliable operation at scale

We’re looking for a hands-on Staff Engineer who can move machine-learning models, agentic workflows, and self-learning approaches from promising prototypes into secure, observable, continuously deployable production systems

This role sits at the intersection of ML systems, platform engineering, and site reliability engineering

You will partner with ML, data, product, and infrastructure teams to create a paved path from experimentation to production—and take ownership of how those systems perform and evolve once deployed

Design and build the production path for the complete ML lifecycle: data and feature preparation, training, experiment tracking, evaluation, artifact and model management, serving, monitoring, feedback collection, and retraining

Build continuous-delivery workflows for models, prompts, agent workflows, data dependencies, and supporting services. Establish automated quality, safety, performance, and compatibility checks

Implement safe rollout patterns such as shadow traffic, canaries, progressive delivery, feature flags, versioned artifacts, automated rollback, and operational kill switches

Turn self-learning approaches into controlled production feedback loops. Build systems for collecting outcomes, validating feedback, maintaining lineage, triggering model refreshes, comparing candidates, and promoting changes under explicit guardrails

Define and operate SLIs, SLOs, alerts, and error budgets across infrastructure, data pipelines, inference services, model quality, and product behavior

Connect model analytics and product telemetry with traditional operational signals so teams can understand whether a problem originates in infrastructure, data, model behavior, or the surrounding product

Improve the scalability, availability, latency, and cost efficiency of distributed training, inference, and data-processing workloads. Own capacity planning and resource optimization, including GPU resources where applicable

Participate in production ownership across the service lifecycle: architecture reviews, deployment, on-call, incident response, blameless postmortems, and systemic remediation

Build self-service platforms and automation that reduce operational toil and shorten the time required for ML engineers and data scientists to reach production

Apply LLMs or agentic automation to evaluation, troubleshooting, and operational workflows where they produce reliable, measurable improvements

Establish practical standards for cloud infrastructure, Kubernetes, infrastructure as code, observability, security, and compliance

Provide technical leadership across ML, data, product, and platform teams, mentoring engineers and influencing architecture without relying on formal authority

Benefits

Generous family leave

Matched donations

Annual learning stipends

Flexible PTO

Competitive retirement plan

Paid volunteer time

Experience distinguishing service-health problems from data-quality or model-quality problemsPractical understanding of the ML lifecycle—including training, evaluation, model deployment, serving, monitoring, versioning, and retraining—and the ability to collaborate effectively with applied ML engineers or researchersStrong software-engineering skills in Python and at least one production systems language such as Go, Java, C++, or RustFamiliarity with SRE practices such as SLIs/SLOs, error budgets, sustainable on-call, incident management, and blameless postmortemsExcellent technical judgment and communication skills, especially when navigating ambiguity and coordinating across teams during production incidentsA track record of Staff-level technical ownership, typically gained through 7+ years of experience in software engineering, platform engineering, SRE, production engineering, or ML infrastructureExperience designing, operating, and troubleshooting distributed production systems, including failure analysis, capacity planning, and performance optimizationA strong automation and internal-customer mindset: you build platforms that are reliable, understandable, and pleasant for other engineers to useHands-on experience with cloud infrastructure, containers and Kubernetes, infrastructure as code, CI/CD, and modern observability

#J-18808-Ljbffr

Worksite address

mountain view, CA, 94039, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

American Honda Motor Co., Inc.

WhatJobs

Senior Product Quality & Root Cause Engineer

haw river, NC

Salary not specified

What Makes a Honda, is Who makes a Honda Honda has a clear vision for the future, and it’s a joyful one.  We are looking for individuals with t…

Listing review due 2026-10-06View job

Avantor

WhatJobs

Process Engineer

carpinteria, CA

See pay details in description

The Opportunity: NuSil (apart of Avantor) is seeking a Process Engineer to be responsible for all phases of silicone products manufacturing…

Listing review due 2026-10-06View job

GE Vernova

WhatJobs

Lead Application Engineer

boston, MA

See pay details in description

Job Description Summary The Lead Application Engineer is an established leader in their respective engineering team. They will drive busine…

Listing review due 2026-10-06View job