waypointjobs

Practice By Numbers, Inc.

Sr. Site Reliability Engineer

bellevue, WA

Check who can apply and the requirements below before continuing.

About this opportunity

Practice By Numbers, Inc. lists this Sr. Site Reliability Engineer opportunity in bellevue, Washington. Review the employer’s description below for duties, qualifications and application requirements.

Job description

This is an engineering-first Senior SRE role.

We’re looking for senior engineers who have:

Built and shipped significant backend systems and/or distributed platforms

Owned services end-to-end in production (design → launch → on-call → reliability improvements)

Led incident response and driven durable follow-ups

Improved reliability by writing software and changing system design—not by adding manual process

You’ll partner closely with product engineering to ensure reliability is designed in from day one, while also building the tooling and platforms that make operating services safer and easier for every engineer.

Engineers here own services end-to-end—from design to production reliability.

Important: This is not a system administrator role. We are explicitly hiring an engineering leader in reliability. Engineering degree is an absolute requirement (BS/MS in CS/CE/EE or closely related engineering field).

What You’ll Do

Own reliability outcomes for critical services: availability, latency, incident rate, and recovery time.

Design and build reliable, scalable distributed systems that support mission-critical healthcare workflows.

Define and operationalize SLOs/SLIs and error budgets; drive adoption across teams and use them to prioritize work.

Lead incident response for high-severity issues; improve on-call effectiveness and reduce alert fatigue.

Run blameless postmortems and ensure follow-ups are implemented, measured, and stick.

Write software to eliminate operational toil: automation, self-service tooling, guardrails, and developer platforms.

Raise the bar on observability (metrics/logs/traces), alerting strategy, and operational readiness.

Improve resilience through capacity planning, load testing, performance tuning, and failure testing.

Mentor engineers (SRE and product engineers) on reliability practices, debugging, and production ownership.

Drive cross-team improvements like production readiness reviews, release safety (progressive delivery), and standard runbooks.

What We’re Looking For

Required

Engineering degree is mandatory: BS/MS in Computer Science, Computer Engineering, Electrical Engineering, or a closely related engineering field.

6+ years experience in software engineering, SRE, infrastructure/platform engineering, or related.

Strong programming skills in Go, Python, Java, or similar (production-quality code).

Proven experience building and operating production backend services or distributed systems.

Meaningful experience in on-call rotations, incident leadership, and post-incident improvement execution.

Strong debugging ability across complex systems: latency, saturation, cascading failures, dependency issues.

Experience with cloud infrastructure (AWS preferred, GCP/Azure acceptable).

Strong Signal

You’ve owned reliability for customer-facing services with clear, measurable improvements (e.g., higher availability, lower MTTR).

You’ve built internal platforms/tooling that made other engineers faster and reduced operational burden.

You’ve worked in an SRE culture with SLOs, error budgets, and blameless postmortems.

You’ve led multi-quarter reliability initiatives spanning multiple teams/services.

Technologies We Work With (Examples)

Cloud: AWS

Containers: Docker, Kubernetes

Infrastructure as Code: Terraform

Observability: Prometheus, Grafana, OpenTelemetry

Languages: Go, Python, TypeScript

CI/CD: GitHub Actions

(Experience with everything isn’t required—strong fundamentals and learning velocity matter most.)

What This Role Is Not

System administration / IT ops / helpdesk

Manual server patching as a primary responsibility

A \"click-ops\" cloud operator role

This is a senior engineering role focused on software-driven reliability and platform engineering.

Why Join PBN

Build and operate mission-critical healthcare infrastructure that supports real patient workflows.

High impact: reliability work directly improves customer trust and revenue-critical operations.

Small team with high ownership, autonomy, and ability to influence architecture.

Strong engineering culture focused on automation, simplicity, and measurable outcomes.

#J-18808-Ljbffr

Worksite address

bellevue, WA, 98009, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Maritime Robotics Systems Engineer

laurel, MD

See pay details in description

Description Do you have a passion for maritime robotics, autonomy, and AI? Do you like to take ownership of technical problems, seek creative…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Senior Data Fusion and Tracking Systems Engineer

laurel, MD

See pay details in description

Description Are you excited about applying your knowledge, skills, and talents to solve some of the nation's most complex defense challenges? …

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Applied Signal Processing and AI/ML Engineer

laurel, MD

See pay details in description

Description Are you an engineer who enjoys turning messy sensor data into actionable information? Have you developed signal processing or AI…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Senior Guidance and Control Engineer

laurel, MD

See pay details in description

Description Are you passionate about collaborating with experienced engineers to design and develop the next generation of weapon systems? Ca…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Wireless Communications Algorithm Engineer

laurel, MD

See pay details in description

Description Do you have experience solving real world problems related to wireless communications? Are you searching for meaningful work to s…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Systems Engineer / Analyst

laurel, MD

See pay details in description

Description Are you a "systems thinker" with systems engineering experience who wants to have impact on nationally important defense programs? …

Listing review due 2026-10-06View job