About this opportunity
Applied Intelligence Consulting (Singapore) lists this SRE – Distributed Systems (Mandarin Required) opportunity in san francisco, California. Review the employer’s description below for duties, qualifications and application requirements.
Job description
Site Reliability Engineer – Distributed Systems (Mandarin Required)
Location:
Palo Alto, California
Experience
3–7 years
Language Requirement
Fluent Mandarin Chinese and professional English are mandatory.
Candidates must be able to conduct technical discussions and collaborate effectively in Mandarin with engineering and infrastructure teams in China, while also operating professionally in English.
About The Role
Our client is a rapidly scaling global consumer internet platform serving hundreds of millions of users worldwide. We are seeking a Site Reliability Engineer (SRE) with strong experience in supporting and engineering large-scale, highly available distributed systems. This is not a traditional IT operations or infrastructure support position. You will work on the reliability architecture behind a massive consumer platform, solving engineering challenges across multi-region systems, global traffic, high availability, disaster recovery, fault isolation, capacity, and large-scale incident response. As the platform expands internationally, you will help build the infrastructure and reliability capabilities required to operate critical services consistently across global regions.
What You\'ll Work On
Design and improve multi-region infrastructure architecture
Build High Availability (HA) and Disaster Recovery (DR) capabilities
Design automated failover and fault-isolation mechanisms
Build and operate infrastructure supporting international production environments
Develop and improve global traffic routing and scheduling
Improve release, deployment, configuration management, and service-governance systems
Build reliability frameworks across Metrics, Logging, and Distributed Tracing
Drive incident response, Root Cause Analysis (RCA), and postmortems
Solve complex production issues across large-scale distributed systems
Improve capacity planning, scalability, and system resilience
Optimize network and data architecture for global deployment
Build automation and reliability tooling to reduce operational overhead
The role specifically covers deployment, monitoring, configuration, service governance, and traffic scheduling alongside reliability engineering and incident response.
What We\'re Looking For
3–7 years of professional experience in SRE, Reliability Engineering, Platform Engineering, or Infrastructure Engineering
Experience supporting large-scale internet or consumer-facing production systems
Strong understanding of distributed systems and high-availability architecture
Hands-on experience with:
High Availability / fault tolerance
Incident management
Capacity planning and scaling
Production troubleshooting
Disaster Recovery and failover
Strong Linux and networking fundamentals
Experience with infrastructure technologies such as MySQL, Redis, and Kafka
Familiarity with Kubernetes, Service Mesh, and cloud-native infrastructure
Experience with observability across Metrics, Logging, and Distributed Tracing
Programming ability in at least one of Go, Python, or Java
Experience building automation tooling, reliability platforms, or infrastructure systems
Multi-Region / Global Infrastructure Experience
Experience designing or operating cross-region distributed systems is particularly valuable, including:
Multi-region deployment
Global traffic routing and scheduling
Data replication and synchronization
Disaster Recovery
Automated failover
Fault isolation
Global / international infrastructure
Particularly Relevant Backgrounds
We are especially interested in engineers who have worked on reliability and infrastructure for high-traffic consumer internet products in areas such as:
Site Reliability Engineering
Production Engineering
Large-scale Distributed Systems
Global Infrastructure
Traffic Infrastructure
Platform Engineering
High Availability / Disaster Recovery
Cloud Infrastructure
Reliability Platforms
Infrastructure Automation
Experience from environments involving large user populations, significant traffic volumes, and business-critical online services is particularly relevant.
Nice To Have
Experience with AWS, GCP, Azure, or other multi-cloud environments
Familiarity with global traffic technologies such as DNS, GSLB, Anycast, or Global Load Balancing
Experience with Chaos Engineering
Participation in failure drills / Game Days
Experience with automated recovery systems
Experience operating infrastructure across multiple countries or regions
Why This Role
You will have direct exposure to the infrastructure supporting a consumer platform with hundreds of millions of users, solving real-world distributed systems and reliability challenges at significant scale. The role provides the opportunity to work deeply across global architecture, multi-region systems, HA/DR, traffic infrastructure, and production reliability, while helping build international infrastructure from the ground up.
#J-18808-Ljbffr
Worksite address
san francisco, CA, 94199, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.