waypointjobs

Apple

Observability SRE Manager, Apple Services Engineering

seattle, WA

Check who can apply and the requirements below before continuing.

About this opportunity

Apple lists this Observability SRE Manager, Apple Services Engineering opportunity in seattle, Washington. Review the employer’s description below for duties, qualifications and application requirements.

Job description

Observability SRE Manager, Apple Services EngineeringPeople at Apple don't just build products, they craft the kind of experiences that have revolutionized entire industries. The diverse collection of our people and their ideas inspire innovation in everything we do. Imagine what you could do here! Join Apple, and help us leave the world better than we found it. The Apple Services Engineering (ASE) team builds and provides systems and infrastructure that fuel Apple's services (such as iCloud, iTunes, Siri, and Maps). We're looking for a senior SRE leader to own and evolve this platform: the metrics, logging, tracing, and alerting infrastructure that underpins operational excellence across Apple Services Engineering.You'll set technical direction for reliability and operational excellence while mentoring engineers, driving automation, and partnering closely with software, infrastructure, finance, and product teams to uphold the reliability of the platform whilst shipping improvements that matter at Apple scale. This is a senior leadership position that defines where our observability platform reliability, scalability and performance goes next, how our SRE practice evolves, including how AI reshapes it, and how we build the team and partnerships to get there. You will lead engineers solving reliability and scale problems few organizations encounter, integrating monitoring seamlessly across disparate infrastructures, hardware, software, application, and network layers, at a scale built to reach every user on the planet. You'll build a team culture that makes SRE sustainable, rewarding, and central to how Apple ships services. The successful candidate has a strong aptitude for both technical leadership and people management, with the ability to context-switch between strategic planning and tactical execution. You should be comfortable building and scaling teams, driving complex cross-functional initiatives, and thriving under pressure, while creating an inclusive, high-trust team culture where engineers do their best work. We believe AI will fundamentally reshape how SRE is practiced, from anomaly detection and root-cause analysis to capacity planning and toil elimination, and we're looking for a leader who shares that conviction and can drive that transformation across the organization.ResponsibilitiesTechnical & Operational LeadershipOwn the reliability, availability, and performance of Apple's observability platform (metrics, logging, tracing, alerting) and the self-service capabilities built on top of itLead staging and production environments for the observability platform with the goal of maximizing availabilityEnsure the platform accurately monitors the health of every application and piece of infrastructure across the Apple ecosystem, the "central nervous system" that other engineering teams and incident responders reach for firstDefine and drive the strategic roadmap for observability infrastructure in partnership with SRE, engineering, and product stakeholdersEstablish and refine SRE practices including SLOs, error budgets, capacity planning, scale testing, disaster recovery, and change managementGuide deep dives into systemic and latent reliability issues spanning the full stack (hardware, software, application, and network), partnering with software and systems engineers to drive fixes to resolutionDrive incident response, post-incident reviews, and systemic improvements that reduce operational toilChampion automation to eliminate manual processes through tooling, self-service platforms, and APIs for internal customersManage on-call rotations and ensure sustainable, well-supported operational coverageDrive standardization of monitoring and troubleshooting methodology across embedded SREs and services throughout the organizationRepresent the SRE organization in design reviews and operational readiness exercises for new and existing servicesPartner with software engineering and architecture teams to influence system design for reliability, scalability, and operabilityBalance technical debt reduction with feature development to maintain platform healthAI & ModernizationDefine and execute a clear AI strategy for the SRE organization, identifying high-impact opportunities where AI/ML tooling can reduce toil, accelerate root-cause analysis, and improve reliability outcomesDrive adoption of AI-assisted tooling (copilots, intelligent runbooks, LLM-based diagnostics, anomaly detection) into day-to-day SRE workflowsBuild a culture where engineers actively experiment with AI tools and modern approaches to solve operational problemsPeople Leadership & Org BuildingLead, grow, and mentor a team of Site Reliability Engineers, conducting regular 1:1s, performance reviews, and career development discussionsHire and build out the SRE team, with a desire to develop engineers to meet both their career goals and the organization's goalsBuild and scale a high-performing team through coaching, clear expectations, and psychological safetyFoster an inclusive team culture and mentor diverse talentChampion engineering best practices for code quality, system design, and operational excellence across the broader organizationCross-Functional & Executive CommunicationBuild strong partnerships across Apple, negotiating priorities and aligning on shared goals with an Apple-first mindsetRepresent the SRE perspective in cross-functional planning, translating reliability requirements into architecture and investment decisionsCommunicate effectively at the executive level, presenting strategy, trade-offs, and progress to senior leadershipDrive standardization and best practices across the observability and infrastructure landscapeMinimum Qualifications5+ years of engineering management experience leading SRE, infrastructure, or observability/monitoring teamsExperience hiring and leading engineers, with a desire to build, grow, and mentor a teamDeep understanding of observability systems and practices: metrics, logging, tracing, alerting, SLOs, error budgets, and fault analysis at scaleStrong systems background, comfortable troubleshooting across the full stack (network, OS, container runtime, application)Experience operating large-scale, multi-tenant distributed systems in production, including Kubernetes environmentsPractical, solid knowledge of shell/bash scripting and at least one higher-level production language (Python preferred; Go, Java, or Scala also valued)Demonstrated experience applying AI/ML tooling or LLM-based solutions to improve SRE or infrastructure operationsTrack record of building high-performing teams through coaching, clear expectations, and psychological safetyDemonstrated ability to drive cross-functional initiatives to completion and communicate at the executive levelBachelor's or Master's degree in Computer Science, Engineering, or related field, or equivalent experiencePreferred QualificationsDeep familiarity with the Prometheus ecosystem and cloud-native observability stacks (Thanos, Splunk, OpenTelemetry, or similar)Experience with third-party cloud platforms (AWS, GCP, or Azure) and infrastructure as code (Terraform, Ansible)Comfortable with open-source configuration management and orchestration tools (Helm, Puppet, Spinnaker)Demonstrable knowledge of TCP/IP, HTTP, web application security, and multi-tier web application architecturesExperience running infrastructure as an internal managed service with defined SLAsFamiliarity with microservices architecture and container orchestration with Kubernetes at scaleBackground in capacity planning, performance engineering, or infrastructure architectureTrack record of driving cultural and process transformation within SRE organizationsExperience building or deploying AI-powered operational tooling (AIOps, intelligent alerting, automated diagnostics)Developing and delivering multi-mode communications tailored to the unique needs of different audiencesAnticipating and balancing the needs of multiple stakeholdersMaking sense of complex, high-quantity, and sometimes contradictory information to solve problems effectivelyRebounding from setbacks and adversity when facing difficult situationsKnowing the most effective and efficient processes to get things done, with a focus on continuous improvementPay & BenefitsAt Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $225,600 and $338,400, and your base pay will depend on your skills, qualifications, experience, and location. Apple employees also have the opportunity to become an Apple shareholder through participation in Apple's discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple's Employee Stock Purchase Plan. You'll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses — including tuition. Additionally, this role might be

Worksite address

seattle, WA, 98101, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

AMN Healthcare

WhatJobs

Gastroenterologist

pensacola, FL

See pay details in description

Job Description & Requirements Gastroenterologist StartDate: ASAP Pay Rate: $ - $ A highly respected healthcare organization with over 60 y…

Listing review due 2026-10-07View job

AMN Healthcare

WhatJobs

Neurologist

yuma, AZ

See pay details in description

Job Description & Requirements Neurologist StartDate: ASAP Available Shifts: M-F no call Pay Rate: $ - $ AMN Healthcare has partnered with a …

Listing review due 2026-10-07View job

AMN Healthcare

WhatJobs

OBGYN-Wisconsin

baldwin, WI

Base Salary: $390,000

Job Description & Requirements OBGYN-Wisconsin StartDate: ASAP Pay Rate: $ OB/GYN Opportunity – Western Wisconsin  Just East of the Twin C…

Listing review due 2026-10-07View job