About this opportunity
Ethos Group lists this Site Reliability Engineer opportunity in irving, Texas. Review the employer’s description below for duties, qualifications and application requirements.
Job description
Ethos Group is seeking a talented and proactive Site Reliability Engineer (SRE) to join our growing technology team.
This role is ideal for an engineer who is passionate about building highly available, scalable, and reliable systems while partnering closely with development and infrastructure teams to improve platform performance, automation, and operational excellence.
As a Site Reliability Engineer, you will play a critical role in maintaining system uptime, optimizing application performance, automating operational processes, and supporting a modern cloud-based technology environment.
Key Responsibilities
Design, implement, and maintain reliable, scalable, and secure cloud infrastructure
Monitor application and system performance to ensure maximum availability
Automate operational tasks and deployment processes
Respond to incidents, troubleshoot production issues, and drive root cause analysis
Develop and maintain monitoring, alerting, and observability solutions
Partner with software engineering teams to improve application reliability and performance
Support CI/CD pipelines and deployment automation initiatives
Create operational documentation, runbooks, and best practices
Participate in on-call rotations and incident response activities
Continuously identify opportunities for system improvements and increased efficiency
Qualifications
Bachelor's degree in Computer Science, Information Technology, or related field, or equivalent experience
Experience in Site Reliability Engineering, DevOps, Systems Engineering, or Cloud Infrastructure roles
Strong understanding of Linux and cloud technologies
Direct hands-on experience with AWS, Azure, or Google Cloud Platform
Experience writing and maintaining infrastructure-as-code tools such as Azure resource templates, Terraform, or CloudFormation
Direct hands-on experience supporting CI/CD pipelines and automation tools
Direct hands-on experience with production grade Kubernetes deployments
Experience with monitoring and observability platforms
Strong scripting skills in Python, PowerShell, Bash, or similar languages
Excellent troubleshooting and problem-solving abilities
Strong communication and collaboration skills
Preferred Qualifications
Deep knowledge of networking, security, and cloud architecture principles
Experience with incident management and root cause analysis
Relevant cloud, Kubernetes, or DevOps certifications
#J-18808-Ljbffr
Worksite address
irving, TX, 75084, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.