Availability awaiting confirmation
We are waiting for a fresh update from the source. This page preserves the last received job details; current availability is not confirmed.
Job description
Seeking a Principal Engineer to own the resiliency strategy and reference architecture for a large-scale AWS environment. This individual will establish availability standards, lead business continuity and disaster recovery initiatives, and drive high-availability architecture across infrastructure, containers, data platforms, and engineering teams.
Required Skills & Experience
12+ years of engineering experience, including 7+ years owning large-scale resiliency or high-availability architecture.
Deep hands-on AWS experience across multi-account, multi-region environments, networking, IAM, Route 53, load balancing, and Global Accelerator.
Proven ownership of multi-AZ/multi-region architecture, BC/DR strategy, automated failover, RTO/RPO, and recovery testing.
Expertise in the AWS Well-Architected Framework and production container platforms, including EKS and ECS/Fargate.
Strong data resiliency experience with Aurora/RDS and at least one of DynamoDB, ElastiCache, or S3 replication.
Advanced Terraform/IaC and infrastructure CI/CD experience, including state management, guardrails, and drift detection.
Hands-on experience with chaos engineering, observability, SLOs, error budgets, and incident response.
Strong cross-functional leadership with the ability to balance availability, cost, risk, and operational complexity.
Key Responsibilities
- Own AWS resiliency strategy, availability standards, and multi-AZ/multi-region architecture.
- Define appropriate active-active, active-passive, warm-standby, and pilot-light patterns.
- Lead BC/DR planning, including RTO/RPO targets, automated failover, runbooks, DR testing, and game days.
Conduct AWS Well-Architected Reviews and drive remediation efforts.
- Build resiliency across ECS/Fargate, EKS, Aurora, RDS, - DynamoDB, ElastiCache, and S3.
- Automate recovery using IaC, self-healing, drift detection, and fault injection.
- Establish SLOs, error budgets, health checks, dependency mapping, and observability standards.
- Lead availability incident response and convert recurring failures into architectural improvements.
- Partner with engineering, product, finance, and business leaders to balance reliability, cost, and operational complexity.
#J-18808-Ljbffr
Worksite address
austin, TX, 78716, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.