About this opportunity
The Mice Groups, Inc. lists this Site Reliability Engineer opportunity in austin, Texas. Review the employer’s description below for duties, qualifications and application requirements.
Job description
Job Description
We are seeking an experienced Site Reliability Engineer (SRE) with strong data center and bare-metal infrastructure experience. This role focuses on maintaining and improving the reliability of data center infrastructure through monitoring, automation, troubleshooting, and incident response.
You will work across servers, hardware, power/cooling infrastructure, monitoring, and automation , partnering with infrastructure, hardware, and data center operations teams to improve reliability and operational efficiency.
Key Responsibilities
Monitor and improve data center infrastructure using Prometheus, Grafana, and Splunk , including server health and power/cooling telemetry.
Develop Python and Shell scripts to automate troubleshooting, incident response, alert management, and other operational processes.
Maintain and improve NetBox or similar data center inventory systems, including device, rack, and infrastructure information.
Build and maintain Grafana dashboards to monitor server health, infrastructure performance, capacity, and power/cooling metrics.
Use SQL and Splunk queries to troubleshoot infrastructure issues, analyze system data, and identify potential bottlenecks.
Troubleshoot bare-metal servers, hardware, and data center infrastructure , including issues involving IPMI, PDUs, and power feeds.
Participate in incident response and on-call support , including troubleshooting, root cause analysis, mitigation, and resolution of infrastructure issues.
Develop and maintain runbooks and operational documentation for common server, hardware, power, cooling, and facility-related issues.
Work with software, hardware, and data center operations teams to support reliable infrastructure deployments and ongoing operations.
Qualifications
8+ years of experience in SRE, Production Operations, Data Center Infrastructure, or a similar role.
Strong hands‑on experience with bare‑metal servers, hardware troubleshooting, provisioning, and data center infrastructure .
Experience with NetBox or similar DCIM/inventory tools .
Strong SQL experience and experience working with REST APIs .
Hands‑on experience with Prometheus, Grafana, and Splunk .
Strong understanding of IPMI, out‑of‑band server management, PDUs, and power distribution .
Practical understanding of data center power and cooling systems , including HVAC, liquid cooling, hot/cold aisle containment, and air handling.
Strong Python and Shell scripting skills with experience automating on‑premises infrastructure.
Bachelor’s degree in Computer Science, Computer Engineering, or a related technical field preferred .
#J-18808-Ljbffr
Worksite address
austin, TX, 78716, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.