waypointjobs

Jobvite, Inc.

Cloud Site Reliability Engineer II

ann arbor, MI

Check who can apply and the requirements below before continuing.

About this opportunity

Jobvite, Inc. lists this Cloud Site Reliability Engineer II opportunity in ann arbor, Michigan. Review the employer’s description below for duties, qualifications and application requirements.

Job description

Come join our passionate team! Barracuda is a leading cybersecurity company providing complete protection against complex threats. Our platform protects email, data, applications, and networks with innovative solutions, and a managed XDR service, to strengthen cyber resilience. Hundreds of thousands of IT professionals and managed service providers worldwide trust us to protect and support them with solutions that are easy to buy, deploy, and use.

We know a diverse workforce adds to our collective value and strength as an organization. Barracuda Networks is proud to be an Equal Opportunity Employer, committed to equal employment opportunity and equitable compensation regardless of race, gender, religion, sex, sexual orientation, national origin, or disability.

Envision yourself at Barracuda

As a Cloud Site Reliability Engineer II on the CloudOps Platform team, you will focus on the observability, monitoring, dashboards, and automation systems that power Barracuda’s next-generation multi-Tenant Kubernetes platform. While your primary mission centers on delivering deep operational visibility, reliable telemetry pipelines, and proactive alerting, you will also play a key role in improving the broader Kubernetes platform running across AWS and Azure.

Our team values collaborative knowledge sharing, automated reliability, and modern engineering practices. We actively embrace AI-assisted workflows (Claude Code, OpenCode, Codex CLI) to accelerate development, diagnostics, and routine platform maintenance. In this role, you will build and operate centralized observability stacks (Grafana, Loki, Mimir, Tempo), create actionable dashboards, automate telemetry via GitOps and Terragrunt, and partner with internal engineering teams to optimize application reliability in production.

Tech Stack Exposure:

Observability & Monitoring: Grafana, Prometheus / Mimir, Loki, Tempo, OpenTelemetry / Grafana Alloy, Sloth (SLOs), Alertmanager

Orchestration & Compute: Kubernetes (AWS EKS, Azure AKS), Helm, Kustomize

IaC & Cloud Provisioning: Terragrunt, Terraform

GitOps & CI/CD: ArgoCD, GitHub Actions

Public Clouds: Amazon Web Services (AWS), Microsoft Azure

Languages & Scripting: Python, Bash (Go is a plus)

AI Developer Tooling: Claude Code, OpenCode, Codex CLI, GitHub Copilot

What you’ll be working on

Operating, scaling, and automating our centralized LGTM telemetry infrastructure (Loki for logs, Mimir/Prometheus for metrics, Tempo for distributed tracing, and Grafana for unified visualization).

Designing intuitive, high-impact Grafana dashboards and executive health overviews for platform services, Kubernetes clusters, and tenant workloads.

Establishing reliable alerting strategies, SLO/SLI tracking (via Sloth), and notification routing to detect and resolve degradation before it impacts production systems.

Automating the deployment of log collectors, metric exporters, and monitoring agents across multi-cluster EKS and AKS environments using GitOps (ArgoCD) and Terragrunt.

Directly contributing to core MTK Kubernetes platform health, performance tuning, and infrastructure modernization.

Leveraging modern AI coding tools (Claude Code, OpenCode, Codex CLI) to build automation, diagnostic tooling, and operational scripts.

Collaborating closely with internal product and tenant teams to assist with observability onboarding, distributed tracing instrumentation, and performance troubleshooting.

What you bring to the role

2–4 years of experience working with public cloud infrastructure (AWS and/or Azure) with a strong passion for observability, monitoring, and systems reliability.

1–2+ years of hands‑on experience deploying, operating, or troubleshooting containerized workloads in Kubernetes (EKS/AKS).

Practical experience configuring, operating, or building dashboards in modern observability stacks (Grafana, ELK, Splunk, etc).

Working knowledge of Infrastructure as Code using Terraform and/or Terragrunt, and GitOps delivery workflows (ArgoCD or Flux).

Solid scripting skills in Python or Bash for system automation, telemetry pipelines, and operational tooling (Go is a plus).

Curiosity and eagerness to leverage AI coding agents (Claude Code, OpenCode, Codex CLI, GitHub Copilot) in everyday engineering workflows.

Strong analytical troubleshooting instincts, clear communication skills, and a collaborative team mindset.

What you’ll get from us

A team where you can voice your opinion, make an impact, and where you and your experience are valued. Internal mobility – there are opportunities for cross training and the ability to attain your next career step within Barracuda.

Equity, in the form of non-qualifying options

High-quality health benefits

Retirement Plan with employer match

Career-growth opportunities

Flexible Time Off and Paid Time Off benefits

Volunteer opportunities

Job ID -

#LI-hybrid

#J-18808-Ljbffr

Worksite address

ann arbor, MI, 48113, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

State of Florida

WhatJobs

PROFESSIONAL ENGINEER II - 37010105

tallahassee, FL

See pay details in description

Requisition No: Agency: Environmental Protection Working Title: PROFESSIONAL ENGINEER II - Pay Plan: Career Service Position Number: Salary…

Last received from source 2026-10-08View job

L3Harris Technologies, Inc.

WhatJobs

Specialist, Logistics Engineer

cape canaveral af station, FL

Salary not specified

Job Title: Specialist, Logistics Engineer Job Code: 45593 Job Location : Cape Canaveral, Florida Job Schedule: 9/80: Employees work 9 out of evry…

Last received from source 2026-10-08View job

L3Harris Technologies, Inc.

WhatJobs

Senior Specialist, Mechanical Engineer Design

huntsville, AL

See pay details in description

Job Title: Senior Specialist, Mechanical Engineer – Design Job Code: 45652 Job Location: Huntsville, AL or Sacramento, CA; On-site Job Schedule: …

Last received from source 2026-10-08View job

L3Harris Technologies, Inc.

WhatJobs

Specialist, Reliability and Systems Safety Engineer

canoga park, CA

See pay details in description

Job Title: Specialist, Reliability and Systems Safety Engineer Job Code: 45666 Job Location: Onsite at our Canoga Park, CA Facility Job Schedule:…

Last received from source 2026-10-08View job

L3Harris Technologies, Inc.

WhatJobs

Senior Specialist, Structural Engineer

canoga park, CA

See pay details in description

Job Title: Senior Specialist, Structural Engineering Job Code: 45376 Job Location: Canoga Park, CA Job Schedule: 9/80: Employees work 9 out of ev…

Last received from source 2026-10-08View job

Saab

WhatJobs

Senior Systems Engineer

dewitt, NY

See pay details in description

Job Description: Saab is seeking a self-motivated, experienced, enthusiastic Systems Engineer interested in designing, deploying, integrating…

Last received from source 2026-10-08View job