About this opportunity
ECS lists this Cloud Site Reliability Engineer (SRE) opportunity in arlington, Virginia. Review the employer’s description below for duties, qualifications and application requirements.
Job description
Job Description
Everforth ECS is seeking a Cloud Site Reliability Engineer (SRE) to work in our Arlington, VA office/remotely. We believe the job of an SRE is to engineer the cloud to run itself. That means writing software and automation that lets systems detect and recover from failure on their own, rather than relying on someone to notice an alert and manually fix it. When something breaks, self-healing comes first, deep root-cause debugging happens after service is restored, not instead of it. We’re looking for someone who automates the operational task by default, not documents the runbook for doing it by hand.
About The Role
This role owns reliability and operational readiness for production systems across our federal cloud platform (AWS GovCloud, IL5 zero-trust). You’ll define what “reliable enough” looks like for our services, build the automation that gets us there, and do it all on an infrastructure-as-code (IaC) foundation.
Responsibilities
Self-Healing Operations
Design and build automated remediation so systems detect, respond to, and recover from failure without manual intervention
Shift the team’s posture from “is it running, how do we fix it” to “how do we make it fix itself”
Automate service restoration first; investigate root cause after
Uptime Goals & Reliability
Define reasonable, data-driven SLOs and error budgets for critical services alongside the teams that own them
Use live metrics to decide what’s “reliable enough” and where to invest next
Infrastructure
Enforce infrastructure-as-code and configuration-as-code, no manual tech change
Own Terraform standards and reusable modules adopted across programs
Drive a containerization-first approach with production-scale Kubernetes (multi-tenancy, security policies, advanced scheduling)
Set CI/CD and pipeline-as-code standards, including progressive delivery
Observability & Incidents
Build monitoring, logging, alerting, and tracing (Datadog, Splunk) that gives automation the signal it needs to self-correct
Own the incident framework: escalation, restoration, root cause analysis, and post-incident review that closes the loop with more automation
Collaboration & Leadership
Partner with development and contractor teams leads to embed reliability and automation across the software
Mentor engineers toward this same automation-first philosophy
Support ATO/RMF and FedRAMP High compliance as it relates to infrastructure and automation
Salary Range
$130,000 - $180,000
Required Skills
Bachelor’s degree in Computer Science, Information Technology, or related field (or equivalent practical experience)
5+ years of SRE experience (or equivalent), with demonstrated technical leadership
10 years of general work experience
Track record building self-healing/auto-remediating systems, not just dashboards
Jenkins experience
Expert AWS knowledge, GovCloud experience strongly preferred
Deep Kubernetes and Terraform expertise at production scale
Strong software engineering background (Python and/or Go)
Experience operating observability platforms (Grafana, Splunk, Prometheus, Loki, etc.)
Proven incident command and postmortem experience
Strong communication skills across technical and federal leadership audiences
Ability to obtain/maintain required government clearance or suitability (CAC/PIV as applicable)
US Citizenship
Desired Skills
Federal government experience
Active security clearance
EEO Statement
ECS Federal LLC is an equal opportunity employer and does not discriminate or allow discrimination on the basis any characteristic protected by law. All qualified applicants will receive consideration for employment without regard to disability, status as a protected veteran or any other status protected by applicable federal, state, or local jurisdiction law.
Everforth ECS
Everforth ECS is the federal segment of Everforth, a $4B global organization with over 10,000 employees. Our nearly 3,500 professionals deliver advanced technology solutions in data and AI, cybersecurity, and enterprise transformation, serving defense, intelligence, and federal civilian agencies. Our work powers mission-critical outcomes, strengthens technology partnerships, and creates meaningful opportunities for our people. We are defined by a commitment to excellence in delivery, a culture of innovation, and an environment where talent can thrive and grow.
We Value
Attracting and developing top talent and high-performing teams
Fostering a culture that is engaging, accountable, and mission-driven
Meet the challenge. Make a difference with Everforth ECS!
#J-18808-Ljbffr
Worksite address
arlington, VA, 22201, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.