Availability awaiting confirmation
We are waiting for a fresh update from the source. This page preserves the last received job details; current availability is not confirmed.
Job description
Senior Site Reliability Engineerensures the reliability, availability, and performance of large-scale software systems through a blend of software engineering and systems administration. Key responsibilities involve automating operational tasks,improving observability, andcontributing to incident management, while also collaborating with developmentand technologyteams to build more reliable and scalable applications.
Join Altium as a Senior Site Reliability Engineer to ensure the reliability and performance of the Altium Cloud Platforms.
Key Responsibilities:
Understanding how an Altium Cloud Platform works
Pioneer improvements in observability, including logging, monitoring, and application performance management (APM), ensuring system reliability and proactive issue detection.
Develop and implement reliability frameworks and patterns that standardize and elevate the resilience of our SaaS products across multiple regions and environments.
Cultivate a shared responsibility model where the SRE team collaborates with and educates engineering teams on reliability best practices.
Contribute to incident response and management, ensuring rapid resolution, clear stakeholder communication, and post-incident analysis for continuous improvement.
Participate in system design consulting, platform management,infrastructureupgradesand capacity planning.
Partner closely with engineering and development teams to enhance product stability, observability, and manageability through best practices in reliability engineering.
Partner closely with DevOps/Operations, drive automation initiatives, promote Infrastructure as Code (IaC), and streamline deployment processes to improve operational efficiency and scalability.
Champion Service-Oriented Organization (SOO) principles to ensure accountability and clarity in service ownership.
Participate in on-call rotation and drive operational improvements after incidents.
6+ years in SRE, DevOps or related role in a large-scale environment
3+ years professional experience in software development
Software development experience(ideally working with and as a .NET developer)
Strong understanding of SDLC, microservice and HA architecture
Observability - NewRelic, ELK, Grafana, PagerDuty, OTEL or similar
Experience with Kubernetes clusters in production setting, AWS, IOC
Experience with operational tasks
Knowledge of CI-CD tooling Jenkins, Gitlab, GitHub, ArgoCD or similar
Knowledge of IaaC Terraform, Ansible
Basic knowledge of networking fundamentals
Experience with relational databases (mysql, postgres) as a plus
MUST be a US permanent resident (Citizen or Green Card)
Altium Limited, a part of the Renesas Group and headquartered in San Diego, California, is a global software company accelerating the pace of electronics innovation. We are redefining electronic product creation in a software-defined world with our industry-first cloud-based platform that unites every stakeholder and phase ofelectronicsdevelopment.
From startups to world’s technology giants,our digital platforms give more power to PCB designers, supply chain, and manufacturing, letting them collaborate as never before. At Altium, our teams are empowered to innovate, collaborate globally, and help create the future of electronics development.
We believe in rewarding our employees with a competitive benefits package alongside their salary. More information will be provided during the hiring process.
#J-18808-Ljbffr
Worksite address
california, MO, 65018, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.