waypointjobs

The Mice Groups, Inc.

Software Engineer SRE

austin, TX

Check who can apply and the requirements below before continuing.

About this opportunity

The Mice Groups, Inc. lists this Software Engineer SRE opportunity in austin, Texas. Review the employer’s description below for duties, qualifications and application requirements.

Job description

Location: Austin, TX (On Site 40 hours per week)

Contract Duration: 12+ month

Pay Rate: $60-$70/hourly (W2)

Job Description

We are seeking an experienced Sr. Data Center Site Reliability Engineer to automate operations and maximize the uptime, efficiency, and scalability of data center, facility power/cooling infrastructure, and software automation . In this role, you will manage, monitor, and optimizing both server reliability and the critical power and cooling infrastructure that sustains our distributed production systems.

Key Responsibilities

Enhance data center observability, logging, and alerting solutions using Grafana, Splunk, and Prometheus, building dashboards that correlate server health, network telemetry, facility power and cooling performance.

Develop automation scripts for hardware incident triage, alert noise reduction, log correlation, and operational workflows, converting recurring manual bare-power/cooling infrastructure investigation patterns into reusable tooling.

Maintain our NetBox data center inventory, building automated pipelines via APIs to track physical infrastructure, rack layouts, and cable topologies.

Build and tune Grafana dashboards with complex queries spanning multiple data sources (including Prometheus metrics) for server health visualization, bare-metal hardware bottleneck identification, and data center capacity monitoring using power feed and cooling infrastructure metrics.

Utilize Splunk and relational databases for infrastructure analytics, writing extensive SQL queries and SPL queries to troubleshoot server production issues, identify infrastructure bottlenecks, and surface environmental insights via IPMI interfaces into dashboards.

Lead incident response and on-call rotations for high-severity data center infrastructure events, directing triage, root cause analysis, mitigation, and resolution for both server-level and facility-level power feed or environmental anomalies.

Develop and maintain runbooks, hardware operational playbooks, and process documentation for common facility, power feed, and server failure scenarios, standardizing infrastructure SOPs across the SRE organization.

Collaborate closely with development, hardware engineering, and facility operations teams to integrate observability best practices into the infrastructure lifecycle and embed monitoring into new compute, storage, power and cooling system rollouts.

Qualifications

Experience: 8+ years of experience in site reliability engineering, production operations, or data center infrastructure management operations.

Education: Bachelor's Degree in Computer Science, Computer Engineering, or a related technical field is highly preferred.

Bare-Metal & Hardware Expertise: Deep hands-on experience troubleshooting, provisioning, and managing enterprise power, bare-metal hardware and server architectures.

Inventory & Asset Management: Strong proficiency using NetBox (or similar DCIM tools) for managing rack space, device lifecycle, and asset tracking.

Data & API Capabilities: Extensive SQL experience (e.g., PostgreSQL, MySQL) for querying relational data infrastructure and deep familiarity consuming/building RESTful APIs to integrate infrastructure tools.

Monitoring & Tooling: Strong expertise with Prometheus for metrics collection, Grafana for visualization, and Splunk for enterprise logging.

Infrastructure Protocols: Proficient with IPMI and server out-of-band management protocols, alongside a strong understanding of data center PDU management and power feed architecture.

Facilities Knowledge: Practical understanding of data center physical infrastructure, specifically power feed distribution systems and cooling infrastructure (e.g., HVAC, liquid cooling, hot/cold aisle containment, air handling units).

Automation: Strong scripting capabilities (Python, Shell) and experience managing infrastructure across highly distributed on-premise environments.

We are an equal opportunity employer and value diversity at The Mice Groups Inc. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Pursuant to the Los Angeles Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

#J-18808-Ljbffr

Worksite address

austin, TX, 78716, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Reverse Engineer and AI Workflow Developer

laurel, MD

See pay details in description

Description Are you passionate about reverse engineering complex software, firmware, and hardware? Do you enjoy developing agentic AI workflo…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

AFSIM Developer

laurel, MD

See pay details in description

Description Are you passionate about building high-fidelity simulations that shape the future of U.S. Space Force capabilities and force design …

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Senior Software Engineer/Architect - TS SCI cleared

laurel, MD

See pay details in description

Description Are you passionate about building solutions for our greatest national security challenges? Are you searching for engaging work wi…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Senior Software Engineer

laurel, MD

See pay details in description

Description Are you passionate about building solutions for our greatest national security challenges? Are you searching for engaging work wi…

Listing review due 2026-10-06View job

Johns Hopkins Applied Physics Laboratory (APL)

WhatJobs

Senior Embedded Systems Developer

laurel, MD

See pay details in description

Description Do you love working on a motivated team to solve complex problems in innovative ways? Do you enjoy creating embedded prototypes i…

Listing review due 2026-10-06View job