About this opportunity
ServiceNow lists this Staff Machine Learning Systems & Reliability Engineer (Moveworks) opportunity in mountain view, California. Review the employer’s description below for duties, qualifications and application requirements.
Job description
We are building AI-enabled product capabilities that improve through data, feedback, and real-world use
We need the production systems that make those capabilities dependable: repeatable delivery, measurable quality, controlled learning loops, and reliable operation at scale
We’re looking for a hands-on Staff Engineer who can move machine-learning models, agentic workflows, and self-learning approaches from promising prototypes into secure, observable, continuously deployable production systems
This role sits at the intersection of ML systems, platform engineering, and site reliability engineering
You will partner with ML, data, product, and infrastructure teams to create a paved path from experimentation to production—and take ownership of how those systems perform and evolve once deployed
Design and build the production path for the complete ML lifecycle: data and feature preparation, training, experiment tracking, evaluation, artifact and model management, serving, monitoring, feedback collection, and retraining
Build continuous-delivery workflows for models, prompts, agent workflows, data dependencies, and supporting services. Establish automated quality, safety, performance, and compatibility checks
Implement safe rollout patterns such as shadow traffic, canaries, progressive delivery, feature flags, versioned artifacts, automated rollback, and operational kill switches
Turn self-learning approaches into controlled production feedback loops. Build systems for collecting outcomes, validating feedback, maintaining lineage, triggering model refreshes, comparing candidates, and promoting changes under explicit guardrails
Define and operate SLIs, SLOs, alerts, and error budgets across infrastructure, data pipelines, inference services, model quality, and product behavior
Connect model analytics and product telemetry with traditional operational signals so teams can understand whether a problem originates in infrastructure, data, model behavior, or the surrounding product
Improve the scalability, availability, latency, and cost efficiency of distributed training, inference, and data-processing workloads. Own capacity planning and resource optimization, including GPU resources where applicable
Participate in production ownership across the service lifecycle: architecture reviews, deployment, on-call, incident response, blameless postmortems, and systemic remediation
Build self-service platforms and automation that reduce operational toil and shorten the time required for ML engineers and data scientists to reach production
Apply LLMs or agentic automation to evaluation, troubleshooting, and operational workflows where they produce reliable, measurable improvements
Establish practical standards for cloud infrastructure, Kubernetes, infrastructure as code, observability, security, and compliance
Provide technical leadership across ML, data, product, and platform teams, mentoring engineers and influencing architecture without relying on formal authority
Benefits
Generous family leave
Matched donations
Annual learning stipends
Flexible PTO
Competitive retirement plan
Paid volunteer time
Experience distinguishing service-health problems from data-quality or model-quality problemsPractical understanding of the ML lifecycle—including training, evaluation, model deployment, serving, monitoring, versioning, and retraining—and the ability to collaborate effectively with applied ML engineers or researchersStrong software-engineering skills in Python and at least one production systems language such as Go, Java, C++, or RustFamiliarity with SRE practices such as SLIs/SLOs, error budgets, sustainable on-call, incident management, and blameless postmortemsExcellent technical judgment and communication skills, especially when navigating ambiguity and coordinating across teams during production incidentsA track record of Staff-level technical ownership, typically gained through 7+ years of experience in software engineering, platform engineering, SRE, production engineering, or ML infrastructureExperience designing, operating, and troubleshooting distributed production systems, including failure analysis, capacity planning, and performance optimizationA strong automation and internal-customer mindset: you build platforms that are reliable, understandable, and pleasant for other engineers to useHands-on experience with cloud infrastructure, containers and Kubernetes, infrastructure as code, CI/CD, and modern observability
#J-18808-Ljbffr
Worksite address
mountain view, CA, 94039, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.