waypointjobs

Kintsugi AI

Senior Site Reliability Engineer (DevOps)

Remote — United States (see country and timezone requirements)

Check who can apply and the requirements below before continuing.

Job description

About Kintsugi

Kintsugi is revolutionizing sales tax automation with our AI-powered platform designed specifically for e-commerce and SaaS businesses. Our solution reduces tax preparation time by 75 percent and cuts compliance costs by 50 percent, allowing finance teams to focus on strategic initiatives rather than routine calculations. As we continue to grow and disrupt the tax automation space, we're building a world-class team to help us scale with purpose.

The Role

We're looking for a Senior Site Reliability Engineer (DevOps) to help scale and harden the infrastructure that powers Kintsugi. This role sits at the intersection of software engineering and operations: you'll keep production reliable under real load, and you'll build the tooling and automation that reduce the manual work of doing that, rather than absorbing more of it yourself as we grow.

You'll work closely with Platform Engineering, Product, and QA to design resilient architectures, improve deployment pipelines, and build the internal tools and guardrails that let the team move quickly without sacrificing stability. You'll operate across our managed Kubernetes and AWS-hosted data layer (Postgres, Redis, networking), and you'll be a key force in shaping the reliability and developer-experience foundation of our engineering org.

We're an agentic-coding-first team — coding agents already do real engineering and operations work here, not just autocomplete on the side. We expect this role to build and extend that practice, not just adopt it.

What You'll Do

Own the reliability of production infrastructure running on managed Kubernetes and AWS, keeping a high-traffic system up and catching issues before customers do

Work agent-first day to day: build, debug, and automate using agentic coding workflows as your default mode, not a fallback tool

Build internal tools and automation that eliminate recurring manual work (toil) for the team, rather than just documenting runbooks around it

Develop and operate monitoring, alerting, and observability systems (metrics, tracing, logging) across the stack

Partner with engineering teams to design for reliability and performance from the start, not bolt it on after incidents

Automate infrastructure management through infrastructure-as-code, and improve CI/CD pipelines and local developer workflows

Lead and evolve incident response practices, including postmortems and blameless learning

Optimize infrastructure for cost efficiency while maintaining high availability and security standards

Contribute to security, compliance, and disaster recovery efforts as the platform scales

Support developer enablement: build and improve the in-house tooling, local dev workflows, and internal platforms other engineers rely on

What We're Looking For

5-8 years in Site Reliability Engineering, DevOps, or Infrastructure Engineering roles, with real ownership of a production system at meaningful scale — someone who drives reliability and tooling initiatives rather than waiting to be assigned them

Fluent working agent-first day to day — directing coding agents to do real engineering work, not just occasional autocomplete

Strong foundation in AWS-hosted data and networking services (RDS/Postgres, ElastiCache/Redis, VPC/networking) and experience running workloads on managed Kubernetes

A track record of building tools, not just running playbooks: scripts, services, or internal platforms that removed manual work for a team

Hands-on experience with CI/CD pipelines and infrastructure-as-code (e.g., Terraform, CloudFormation)

Expertise in observability stacks (metrics, tracing, logging) and modern monitoring practices

Familiarity with security and compliance in cloud environments (SOC 2, GDPR, etc. a plus)

A collaborative mindset with a passion for empowering developers to move fast safely

Experience with (or strong interest in) developer enablement — building the internal tooling and platforms that make other engineers more productive

Nice to have: experience operating across multiple cloud providers, as our infrastructure footprint expands

Originally posted on Himalayas

Who can apply

Eligible countries: United States. Accepted UTC offsets: UTC-10, UTC-9, UTC-8, UTC-7, UTC-6, UTC-5, UTC+14. Review the full description for employer-specific work authorization, residency and schedule requirements.

Ready for your next step?Apply on the official website
Apply on Himalayas ↗

Explore related searches

Current related jobs

infisical

Jobicy

Senior Full Stack Engineer

Remote — Brazil, Canada, Europe, USA

Salary not specifiedRemote

Infisical is the open source security infrastructure platform that engineers use for secrets management, certificates, and privileged access mana…

Listing review due 2026-10-07View job

Spotify

Jobicy

Data Scientist - Music Mission

Remote — USA

Salary not specifiedRemote

The Music Mission enables music creators to grow, engage, and monetize their fan bases on Spotify. Central to the Music Mission's vision is the d…

Listing review due 2026-10-07View job

Spotify

Jobicy

Data Scientist - Music Promotion

Remote — USA

Salary not specifiedRemote

The Music Mission enables Music creators to grow, engage & monetize their fan bases on Spotify. Central to the Music Mission's vision is the deve…

Listing review due 2026-10-07View job

infisical

Jobicy

Strategic Finance

Remote — Canada, USA

Salary not specifiedRemote

Infisical is the open source security infrastructure platform that engineers use for secrets management, certificates, and privileged access mana…

Listing review due 2026-10-07View job