waypointjobs

Barracuda

Cloud Site Reliability Engineer II

ann arbor, MI

Check who can apply and the requirements below before continuing.

About this opportunity

Barracuda lists this Cloud Site Reliability Engineer II opportunity in ann arbor, Michigan. Review the employer’s description below for duties, qualifications and application requirements.

Job description

Come join our passionate team! Barracuda is a leading cybersecurity company providing complete protection against complex threats. Our platform protects email, data, applications, and networks with innovative solutions, and a managed XDR service, to strengthen cyber resilience. Hundreds of thousands of IT professionals and managed service providers worldwide trust us to protect and support them with solutions that are easy to buy, deploy, and use.

We know a diverse workforce adds to our collective value and strength as an organization. Barracuda Networks is proud to be an Equal Opportunity Employer, committed to equal employment opportunity and equitable compensation regardless of race, gender, religion, sex, sexual orientation, national origin, or disability.

Envision yourself at Barracuda

As a Cloud Site Reliability Engineer II on the CloudOps Platform team, you will focus on the observability, monitoring, dashboards, and automation systems that power Barracuda’s next-generation multi-Tenant Kubernetes platform. While your primary mission centers on delivering deep operational visibility, reliable telemetry pipelines, and proactive alerting, you will also play a key role in improving the broader Kubernetes platform running across AWS and Azure.

Tech Stack Exposure

Observability & Monitoring: Grafana, Prometheus / Mimir, Loki, Tempo, OpenTelemetry / Grafana Alloy, Sloth (SLOs), Alertmanager

Orchestration & Compute: Kubernetes (AWS EKS, Azure AKS), Helm, Kustomize

IaC & Cloud Provisioning: Terragrunt, Terraform

GitOps & CI/CD: ArgoCD, GitHub Actions

Public Clouds: Amazon Web Services (AWS), Microsoft Azure

Languages & Scripting: Python, Bash (Go is a plus)

AI Developer Tooling: Claude Code, OpenCode, Codex CLI, GitHub Copilot

What You’ll Be Working On

Operating, scaling, and automating our centralized LGTM telemetry infrastructure (Loki for logs, Mimir/Prometheus for metrics, Tempo for distributed tracing, and Grafana for unified visualization).

Designing intuitive, high-impact Grafana dashboards and executive health overviews for platform services, Kubernetes clusters, and tenant workloads.

Establishing reliable alerting strategies, SLO/SLI tracking (via Sloth), and notification routing to detect and resolve degradation before it impacts production systems.

Automating the deployment of log collectors, metric exporters, and monitoring agents across multi-cluster EKS and AKS environments using GitOps (ArgoCD) and Terragrunt.

Directly contributing to core MTK Kubernetes platform health, performance tuning, and infrastructure modernization.

Leveraging modern AI coding tools (Claude Code, OpenCode, Codex CLI) to build automation, diagnostic tooling, and operational scripts.

Collaborating closely with internal product and tenant teams to assist with observability onboarding, distributed tracing instrumentation, and performance troubleshooting.

What You Bring To The Role

2–4 years of experience working with public cloud infrastructure (AWS and/or Azure) with a strong passion for observability, monitoring, and systems reliability.

1–2+ years of hands‑on experience deploying, operating, or troubleshooting containerized workloads in Kubernetes (EKS/AKS).

Practical experience configuring, operating, or building dashboards in modern observability stacks (Grafana, ELK, Splunk, etc).

Working knowledge of Infrastructure as Code using Terraform and/or Terragrunt, and GitOps delivery workflows (ArgoCD or Flux).

Solid scripting skills in Python or Bash for system automation, telemetry pipelines, and operational tooling (Go is a plus).

Curiosity and eagerness to leverage AI coding agents (Claude Code, OpenCode, Codex CLI, GitHub Copilot) in everyday engineering workflows.

Strong analytical troubleshooting instincts, clear communication skills, and a collaborative team mindset.

What You’ll Get From Us

A team where you can voice your opinion, make an impact, and where you and your experience are valued. Internal mobility – there are opportunities for cross training and the ability to attain your next career step within Barracuda.

Equity, in the form of non-qualifying options

High-quality health benefits

Retirement Plan with employer match

Career-growth opportunities

Flexible Time Off and Paid Time Off benefits

Volunteer opportunities

#J-18808-Ljbffr

Worksite address

ann arbor, MI, 48113, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

American Honda Motor Co., Inc.

WhatJobs

Senior Product Quality & Root Cause Engineer

haw river, NC

Salary not specified

What Makes a Honda, is Who makes a Honda Honda has a clear vision for the future, and it’s a joyful one.  We are looking for individuals with t…

Listing review due 2026-10-06View job

Avantor

WhatJobs

Process Engineer

carpinteria, CA

See pay details in description

The Opportunity: NuSil (apart of Avantor) is seeking a Process Engineer to be responsible for all phases of silicone products manufacturing…

Listing review due 2026-10-06View job

GE Vernova

WhatJobs

Lead Application Engineer

boston, MA

See pay details in description

Job Description Summary The Lead Application Engineer is an established leader in their respective engineering team. They will drive busine…

Listing review due 2026-10-06View job