Job description
At Oracle Health, we’re building the next generation of reliable, AI-powered healthcare infrastructure.
Our Clinical AI Assistant platform supports healthcare professionals operating in real-world clinical environments where reliability, scalability, and operational excellence are critical. We’re looking for experienced Site Reliability Engineers who want significant ownership, difficult technical challenges, and the opportunity to influence how large-scale AI systems are operated.
This is a hands-on engineering role focused on solving hard distributed systems problems, improving platform resilience, and building intelligent operational capabilities at scale.
What You’ll Own
Lead reliability engineering efforts for large-scale cloud-native healthcare platforms
Design and operate highly available distributed systems supporting AI-driven services
Build automation, self-healing systems, and intelligent operational tooling
Drive improvements across scalability, observability, deployment safety, and incident response
Lead complex production investigations and engineer durable long-term fixes
Develop AIOps capabilities including anomaly detection, predictive scaling, and automated remediation
Partner with software and platform teams to improve architecture, resiliency, and operational readiness
Influence engineering standards across Kubernetes, CI/CD, infrastructure as code, and cloud operations
Mentor engineers and help raise operational engineering maturity across the organization
Originally posted on Himalayas
Who can apply
Eligible countries: United States. Accepted UTC offsets: UTC-10, UTC-9, UTC-8, UTC-7, UTC-6, UTC-5, UTC+14. Review the full description for employer-specific work authorization, residency and schedule requirements.