waypointjobs

Ses Corporation

Reliability Engineer

lincoln, MA

Check who can apply and the requirements below before continuing.

About this opportunity

Ses Corporation lists this Reliability Engineer opportunity in lincoln, Massachusetts. Review the employer’s description below for duties, qualifications and application requirements.

Job description

This role supports the U.S. Air Force Cloud One Architecture and Common Shared Services contract and currently has an opening for a Reliability Engineer . The Reliability Engineer is responsible for ensuring the availability, performance, scalability, and resiliency of mission‑critical systems. This role applies software engineering principles to infrastructure and operations, with a strong emphasis on automation, monitoring, incident response, and continuous reliability improvement. The reliability engineer serves as the bridge between development, operations, and platform teams to ensure production systems consistently meet defined service level objectives (SLOs) while supporting rapid, safe delivery of new capabilities.

Location

This position will be hybrid remote. Candidates will be required to work onsite as needed. Candidates preferred to be located near Hanscom AFB (Boston, MA).

Requirements

System Reliability & Availability

Design, implement, and maintain highly available, fault-tolerant systems in cloud and hybrid environments

Define, measure, and report Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets

Identify reliability risks and implement mitigation strategies across the system lifecycle

Conduct capacity planning and performance modeling to ensure systems scale to meet demand

Monitoring, Observability & Alerting

Implement and manage monitoring, logging, and tracing solutions to provide full system observability

Define actionable alerting thresholds that minimize noise and enable rapid incident detection

Analyze trends and metrics to proactively identify potential reliability issues

Incident Response & Problem Management

Participate in on‑call rotations and lead incident response activities for production systems

Coordinate troubleshooting efforts across development, infrastructure, and security teams

Conduct post‑incident reviews (PIRs) and develop corrective and preventive action plans

Track recurring issues and ensure root causes are resolved

Automation & Engineering Excellence

Automate operational tasks to reduce manual intervention and operational risk

Develop scripts, tools, and services that improve system reliability and reduce mean time to recovery (MTTR)

Promote "automation over toil" and standardize operational workflows

Reliability‑Focused Engineering

Participate in architecture and design reviews with an emphasis on reliability, resiliency, and recoverability

Validate disaster recovery (DR) and business continuity plans; test failover mechanisms

Support chaos engineering, fault injection testing, and resilience validation where appropriate

Collaboration & Governance

Partner with DevOps, Platform, and Security teams to ensure reliability aligns with delivery and compliance objectives

Document system reliability standards, runbooks, and operational procedures

Support compliance and audit activities (e.g., FedRAMP, FISMA, internal operational controls)

Required Skills

Bachelors and eight (8) years or more of experience; Masters and six (6) years or more of experience. Additional experience may be accepted in lieu of degree

Active Secret clearance at a minimum required to start

US citizenship required

Experience with cloud platforms (AWS, Azure, OCI, or GCP), including managed services

Experience with containerized environments (Docker, Kubernetes)

Familiarity with CI/CD pipelines and deployment automation

SLOs and error budgets

Capacity modeling and performance testing

Strong understanding of:

Distributed systems and high‑availability architectures

Linux/Windows system administration

Networking fundamentals (DNS, TCP/IP, load balancing)

Hands‑on experience with:

Monitoring and observability tools (e.g., Prometheus, Grafana, ELK/Elastic, Datadog, Azure Monitor)

Infrastructure as Code (Terraform, ARM, CloudFormation)

Scripting or programming languages (Python, Bash, Go, PowerShell, or similar)

Experience supporting incident management and on‑call operations

Preferred Skills

Experience with USAF Cloud One or Platform 1.

Experience with Zero Trust Architecture

Cloud certifications in AWS, Azure, Google, or Oracle clouds

Benefits

SES provides a competitive salary and the following benefits:

Medical

Dental

Vision

AD&D

STD

LTD

Company paid Life Insurance

401k with employer contribution

Paid Time Off

Pet Insurance

#J-18808-Ljbffr

Worksite address

lincoln, MA, 01773, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Space Systems Mechanical Engineer

laurel, MD

See pay details in description

Description Do you want to design and build unique space structures and spacecraft for NASA missions that enable groundbreaking scientific disco…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

System Security Engineer

laurel, MD

See pay details in description

Description Are you looking for an opportunity to utilize your technical skills to solve complex, real-world problems? If so, we're looking …

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Advanced Reentry Mission Engineer

laurel, MD

See pay details in description

Description Are you interested in hypersonic and reentry system design and prototyping? Do you want to make contributions to next generation …

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Thermal and EO/IR Modeling and Simulation Engineer

laurel, MD

See pay details in description

Description Are you looking for a unique opportunity to impact significant advances to the nation's groundbreaking integrated air and missile de…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Network Effects Engineer

laurel, MD

See pay details in description

Description Do you want to perform advanced research, development, and test & evaluation of communications systems and network technologies that…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

System Realization and Resilience Engineer

laurel, MD

See pay details in description

Description Are you passionate about applying system engineering principles to influence the development and resilience of future strategic weap…

Listing review due 2026-10-06View job