Job description
Were looking for a skilled Site Reliability Engineer (SRE) to join our globally distributed team, with a focus on maintaining system reliability and supporting operations over the weekend. In this role, youll ensure high availability and performance across Elwoods cloud-based EMS and PMS platforms, built on AWS and GCP.
Youll take ownership of incident response, system monitoring, and automation efforts, playing a key role in keeping our infrastructure resilient and efficient. This is a highly visible position that works closely with engineering and client-facing teams to resolve production issues and continuously improve system reliability. The ideal candidate has strong cloud infrastructure experience, a background in DevOps or SRE roles, and a passion for building scalable, fault-tolerant systems.
Originally posted on Himalayas
Who can apply
Eligible countries: United States. Accepted UTC offsets: UTC-10, UTC-9, UTC-8, UTC-7, UTC-6, UTC-5, UTC+14. Review the full description for employer-specific work authorization, residency and schedule requirements.