waypointjobs

CXM

Site Reliability Engineer

Remote — Worldwide (see timezone requirements)

Check who can apply and the requirements below before continuing.

Job description

Join our Platform & Production Reliability team and help ensure the reliability, performance, and availability of our mission-critical trading systems. As an Application Site Reliability Engineer (SRE), you will own the day-to-day reliability of our .NET/C# services running on Windows, starting with our in-house liquidity bridge that connects MetaTrader trading servers to external liquidity providers. Over time, you will expand your impact across related trading and back-office services.

This is a hands-on role for an engineer who enjoys solving production challenges, improving observability, automating operations, and building resilient systems where uptime directly impacts customer experience.

Position Details

TeamPlatform & Production Reliability

LocationRemote (Americas, LatAm preferred)

Working HoursAmericas time zones (UTC-3 to UTC-8)

On-callRotation aligned with the London trading day

Employment TypeFull-time, Permanent

Experience LevelMid-Level (3–5 years)

Technology Stack.NET/C#, Windows Server, AWS, Aurora PostgreSQL, Prometheus, Grafana, Terraform

About the Role

Our trading platform powers every customer interaction, making reliability a first-class product concern. You will be responsible for maintaining and improving the operational reliability of our .NET/C# services on Windows, ensuring they remain highly available, observable, and resilient.

You'll collaborate closely with software engineers to improve monitoring, deployment safety, automation, fault isolation, and incident response, while driving continuous improvements in platform reliability and operational excellence.

What You'll Do

Participate in the on-call rotation for production trading systems and lead incident response during service disruptions.

Investigate production incidents, perform root cause analysis, and implement preventive actions to eliminate recurring issues.

Build and maintain Grafana dashboards, Prometheus alerts, and operational health views across applications, infrastructure, and databases.

Instrument .NET services to improve telemetry, metrics, logging, and visibility into service health and customer impact.

Define, implement, and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.

Troubleshoot issues across:

.NET/C# applications

Windows Server

Aurora PostgreSQL databases

AWS infrastructure

CI/CD pipelines and deployments

Improve deployment safety, release automation, and rollback strategies.

Partner with developers to improve application operability, resilience, and fault isolation.

Automate operational tasks through scripting and infrastructure automation.

Create and maintain runbooks, operational documentation, and incident response procedures.

Continuously improve monitoring, alert quality, automation, and platform reliability.

Required Technical Skills

.NET & Windows

Strong experience debugging and supporting .NET/C# applications in production.

Hands-on experience with Windows Server environments.

Scripting & Automation

Strong PowerShell scripting skills.

Experience with Python or Bash.

Observability

Experience with Grafana, Prometheus, and Loki (or equivalent monitoring and observability tools).

Solid understanding of metrics, logging, tracing, and alerting best practices.

CI/CD & DevOps

Experience with modern CI/CD pipelines.

Knowledge of deployment strategies, release automation, and rollback mechanisms.

Cloud & Infrastructure

Experience working with AWS.

Hands-on experience with Terraform or other Infrastructure as Code (IaC) tools.

Databases

Experience troubleshooting and supporting Aurora PostgreSQL or other relational database platforms.

Reliability Engineering

Practical experience with:

SLIs & SLOs

Error Budgets

Incident Response

Root Cause Analysis (RCA)

Alert Design

Production Operations

Preferred Qualifications

Experience supporting high-availability or low-latency financial or trading systems.

Familiarity with MetaTrader environments or financial technology platforms.

Experience with distributed systems and microservices.

Knowledge of OpenTelemetry or similar observability frameworks.

Exposure to Docker, Kubernetes, or containerized environments.

Why Join Us?

Work on mission-critical trading infrastructure that directly impacts customers.

Solve challenging reliability and scalability problems in a real-time environment.

Build world-class observability, automation, and deployment practices.

Collaborate with experienced engineers in a modern engineering culture.

Influence reliability strategy and engineering best practices across the platform.

If you're passionate about production engineering, automation, and building reliable systems at scale, we'd love to hear from you.

Originally posted on Himalayas

Who can apply

The source lists worldwide eligibility. Accepted UTC offsets: UTC-11, UTC-10, UTC-9.5, UTC-9, UTC-8, UTC-7, UTC-6, UTC-5, UTC-4, UTC-3.5, UTC-3, UTC-2, UTC-1, UTC+0, UTC+1, UTC+2, UTC+3, UTC+3.5, UTC+4, UTC+4.5, UTC+5, UTC+5.5, UTC+5.75, UTC+6, UTC+6.5, UTC+7, UTC+8, UTC+8.75, UTC+9, UTC+9.5, UTC+10, UTC+10.5, UTC+11, UTC+12, UTC+12.75, UTC+13, UTC+14. Review the full description for employer-specific work authorization, residency and schedule requirements.

Ready for your next step?Apply on the official website
Apply on Himalayas ↗

Explore related searches

Current related jobs

Red Wine and Blue

Himalayas

General Interest Application

Remote — United States (see country and timezone requirements)

Salary not specifiedfull timeRemote

WHO THE HECK ARE WE? Red Wine & Blue is a national community of over 600,000 diverse suburban women working together to defeat extremism, one fr…

Listing expires 2026-12-04View job

Carle Health

Himalayas

Finance Systems Analyst

Remote — United States (see country and timezone requirements)

$26.41 – $44.10 per hourfull timeRemote

OverviewThe Finance Systems Analyst assists with supporting assigned finance/accounting applications for the enterprise such as costing, producti…

Listing expires 2026-12-04View job

Kyowa Kirin

Himalayas

Scientific Relations Manager

Remote — Worldwide (see timezone requirements)

Salary not specifiedfull timeRemote

OverviewWE PUSH THE BOUNDARIES OF MEDICINE. LEAPING FORWARD TO MAKE PEOPLE SMILE At Kyowa Kirin International (KKI), our purpose is to make peop…

Listing expires 2026-12-04View job

Belden, Inc

Himalayas

Solution Consultant - Cybersecurity Practice (US)

Remote — United States (see country and timezone requirements)

$125,000.00 – $160,000.00 per yearfull timeRemote

Innovation Starts With YouPropel your career at Belden, where innovation creates possibilities—for our people, our customers, and the communities…

Listing expires 2026-12-04View job