About this opportunity
GCS Recruitment lists this Sr. Dev Ops Manager opportunity in san jose, California. Review the employer’s description below for duties, qualifications and application requirements.
Job description
Senior DevOps Manager
Hybrid · 3 days/week onsite · Full-time
About the role
We're looking for a seasoned Senior DevOps Manager to lead a global operations team delivering world‑class infrastructure performance and reliability. This is a high‑impact leadership role responsible for driving operational excellence, fostering cross‑functional collaboration, and implementing scalable project‑management practices to support enterprise customers.
As a key member of the engineering leadership team, you'll ensure 24/7 operational stability across the infrastructure while continuously improving processes, systems, and team capabilities to meet evolving business needs.
What you'll do
Lead and mentor Technical Operations engineers across multiple time zones
Build a collaborative team culture centred on knowledge sharing, innovation, and operational excellence
Own 24/7 operational stability - incident response, escalation procedures, and post-incident reviews
Drive incident management including alert management, outage response, and root cause analysis (RCA/CAR)
Implement and maintain SRE practices: SLOs, error budgets, and reliability engineering
Establish robust monitoring and alerting using APM tools and diagnostic dashboards
Lead technical project delivery with clear timelines, resource allocation, and stakeholder communication
Drive automation initiatives - runbook automation, deployment automation, and infrastructure-as-code
Manage infrastructure automation using Terraform, Kubernetes, and cloud platforms (GCP/AWS)
Provide executive reporting on operational metrics, project status, and team performance
Drive cost-optimisation and resource-planning initiatives
What you'll bring
8+ years in technical operations, with 4+ years leading Technical Operations, SRE, or infrastructure teams
Proven track record developing and mentoring high-performing teams
Strong project management skills (Agile/Kanban, JIRA)
Excellent communication, including presenting to executive stakeholders
Deep SRE experience - incident management, monitoring, and reliability engineering
Infrastructure automation expertise (Terraform, Kubernetes, Docker, CI/CD)
Cloud platform proficiency (GCP/AWS), including networking, security, and cost management
Monitoring and observability experience (Prometheus, Grafana, APM, log aggregation)
24/7 operations experience, including on-call and global team coordination
Change management and compliance experience (SOC 2, security reviews, audits)
Linux, networking protocols, and security fundamentals
Nice to have
Cloud certifications (GCP Professional, AWS Solutions Architect) or SRE credentials
Enterprise networking, wireless infrastructure, IoT, or telecoms experience
Security and compliance expertise (vulnerability management, regulatory frameworks)
Degree in Computer Science, Engineering, or a related field
Multi-region operations experience across time zones
What's on offer
Comprehensive medical, dental, and vision plans
Life and accidental death insurance
401(k) and company incentive plan
Paid holidays and vacation, plus additional leave options
Hybrid working (3 days/week onsite)
#J-18808-Ljbffr
Worksite address
san jose, CA, 95199, US
Who can apply
Review the original listing for work authorization, qualifications and employer requirements.