waypointjobs

Movable Ink

Lead Site Reliability Engineer

new york, NY

Check who can apply and the requirements below before continuing.

About this opportunity

Movable Ink lists this Lead Site Reliability Engineer opportunity in new york, New York. Review the employer’s description below for duties, qualifications and application requirements.

Job description

Movable Ink scales content personalization for marketers through data-activated content generation and AI decisioning. The world’s most innovative brands rely on Movable Ink to maximize revenue, simplify workflow and boost marketing agility. Headquartered in New York City with close to 600 employees, Movable Ink serves its global client base with operations throughout North America, Central America, Europe, Australia, and Japan.

As one of our Lead Site Reliability Engineers, you will combine hands‑on technical expertise with strategic technical leadership across infrastructure and software development. You will own the design and evolution of major systems within our multi‑cloud, multi‑region, active‑active content serving platform that serves upwards of 25 Billion requests daily. Through a combination of architectural vision, cross‑team collaboration and mentorship, you will help drive the reliability initiatives and define the technical strategy that scales our platform to 50 Billion requests per day and beyond.

Responsibilities

Define and drive the automation strategy for infrastructure tooling, establishing standards that minimize manual work, increase performance and reduce incident frequency and severity of incidents

Own the design, reliability and evolution of core platform applications, mentoring team members on best practices and ensuring systems meet long‑term business objectives

Architect and lead the logging platform strategy, driving its design and balancing availability, retention and cost optimization

Establish capacity planning and performance management frameworks, proactively identifying scaling opportunities and guiding teams through complex troubleshooting scenarios

Lead cross‑functional reliability initiatives with SRE and service engineering teams, influencing architectural decisions and championing practices that ensure resilient service delivery

Demonstrate a high level of autonomy in anticipating, identifying, and addressing systemic weaknesses and opportunities for platform improvement without direct supervision.

Qualifications

Proven track record in Site Reliability or Software Engineering, designing, building, and owning scalable, resilient services with a focus on long‑term reliability strategy

Deep expertise in architecting and operating complex distributed systems such as Apache Pulsar, Apache Kafka, Grafana Loki, ScyllaDB/Cassandra, with the ability to guide teams through distributed system challenges

Designing and owning automation strategies to manage services at scale, with expertise in establishing performance analysis frameworks and mentoring others on diagnostics and resolution

Deep, hands‑on experience (6+ years) in Site Reliability or Software Engineering, specifically leading and shaping multi‑cloud architecture and strategy (AWS and GCP).

Experience architecting and leading large-scale observability platforms, including defining observability standards and SLO frameworks. We use Prometheus and Thanos with Grafana Alloy, Loki and Tempo

Experience leading on‑call excellence, including driving improvements to monitoring and alerting strategies, automating runbooks and mentoring team members on incident response best practices. Every member of the SRE team does a week long on‑call rotation

Expert‑level proficiency with infrastructure as code, including defining IaC standards and patterns across teams. We use Terraform and Chef

Advanced Kubernetes expertise, including cluster architecture design, multi‑tenancy strategies, and guiding teams on container orchestration best practices. We use EKS and GKE

Proficiency in multiple programming languages with the ability to design and review code that meets reliability standards. We use NodeJS, Golang, Ruby, Python and shell scripting

Advanced Linux systems expertise, with the ability to diagnose complex system‑level issues and mentor others on performance tuning and troubleshooting

The base pay range for this position is $184,200 - 240,000 /year, which can include additional bonus depending on the position ultimately offered, in addition to a full range of medical, financial, and/or other benefits. The base pay offered may vary depending on job‑related knowledge, skills, and experience.

Studies have shown that women, communities of color, and historically underrepresented people are less likely to apply to jobs unless they meet every single qualification. We are committed to building a diverse and inclusive culture where all Inkers can thrive. If’re excited about the role but don’t meet all of the abovementioned qualifications, we encourage you to apply. Our differences bring a breadth of knowledge and perspectives that makes us collectively stronger.

We welcome and employ people regardless of race, color, gender identity or expression, religion, genetic information, parental or pregnancy status, national origin, sexual orientation, age, citizenship, marital status, ethnicity, family or marital status, physical and mental ability, political affiliation, disability, Veteran status, or other protected characteristics. We are proud to be an equal opportunity employer.

#J-18808-Ljbffr

Worksite address

new york, NY, 10261, US

Who can apply

Review the original listing for work authorization, qualifications and employer requirements.

Ready for your next step?Apply on the official website
Apply on WhatJobs ↗

Explore related searches

Current related jobs

Faith Technologies

WhatJobs

Structural Engineer III

fox crossing, WI

Salary not specified

Faith Technologies is seeking a Structural Engineer III to support complex electrical and construction projects. In this role, you’ll perform str…

Last received from source 2026-10-06View job

Amazon Web Services, Inc.

WhatJobs

Data Center Controls Engineer

manassas, VA

Salary not specified

Amazon Web Services (AWS) seeks a Data Center Controls Engineer to design, implement, and optimize control systems that power AWS’s global data c…

Last received from source 2026-10-06View job

FM

WhatJobs

Senior Product Certification Engineer

glocester, RI

Salary not specified

FM seeks an Approvals Advanced Engineer to lead complex product testing and certification projects. In this role, you’ll interpret standards, des…

Last received from source 2026-10-06View job

Amazon Web Services, Inc.

WhatJobs

Senior DevOps Engineer – Cloud Consulting

herndon, VA

Salary not specified

Amazon Web Services, Inc. seeks a Senior Delivery Consultant DevOps to lead complex cloud and DevOps engagements within AWS ProServe for public s…

Last received from source 2026-10-06View job

Amazon Web Services, Inc.

WhatJobs

Data Center Commissioning Engineer

umatilla, OR

Salary not specified

Amazon Web Services (AWS) seeks a Commissioning Engineer to ensure reliable startup, testing, and performance of critical infrastructure across A…

Last received from source 2026-10-06View job

Amazon Web Services, Inc.

WhatJobs

Data Center Controls Engineer

culpeper, VA

Salary not specified

Amazon Web Services is seeking a Controls Engineer to design, implement, and optimize advanced data center control systems that ensure availabili…

Last received from source 2026-10-06View job