openleverjobgether
Manager, Cloud Services and Site Reliability
Jobgether
LocationCanada
EmploymentFull-time
Posted2026-08-20T02:30:54.574000+00:00
Last observed2026-08-26 21:51:40.410433
Job idjobgether-jobgether:lever:e6da1d55-fd93-4629-a952-d7c5bc76eae9
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Manager, Cloud Services and Site Reliability based in Canada. This leadership role is responsible for the reliability, availability, scalability, and operational excellence of high-volume, business-critical SaaS applications. You will lead and develop an SRE team while providing technical direction across cloud infrastructure and production operations. The role partners closely with engineering, product, platform, and security teams to strengthen resilient and scalable services. You’ll drive improvements in incident response, automation, monitoring, service health, and operational readiness. Using data and operational insights, you’ll identify reliability risks and turn them into measurable improvements. This is an opportunity to shape engineering practices, reduce operational toil, and build a culture of ownership and continuous improvement. You’ll work in a collaborative, remote-friendly environment where technical expertise and people leadership are equally valued. Lead, coach, mentor, and develop a high-performing SRE team, establishing clear expectations and fostering ownership, collaboration, accountability, and continuous improvement. Drive reliability engineering practices across critical cloud services, including SLOs, SLIs, monitoring, alerting, capacity planning, and service health reporting. Partner with engineering and platform teams to improve the architecture, scalability, resilience, maintainability, and operational performance of cloud-based systems. Own and continuously improve incident management practices, including major incident coordination, post-incident reviews, root-cause analysis, and follow-up actions. Champion automation, tooling, and engineering practices that reduce manual operational work and improve consistency and scalability. Analyze operational data, service metrics, and risk indicators to identify reliability gaps, prioritize improvements, and communicate progress to technical and business stakeholders. Partner with security and engineering teams to support secure, compliant, and operationally mature production environments. Contribute to disaster recovery, infrastructure automation, CI/CD, cost optimization, and multi-cloud initiatives where relevant. Improve documentation, operational readiness, and service management practices across teams. Evaluate tools, technologies, and vendors that can strengthen service reliability and operational effectiveness. Influence cross-functional stakeholders and help establish a culture focused on resilient systems, measurable service health, and continuous improvement. Requirements: 5+ years of experience in SRE, DevOps, infrastructure, cloud operations, or a related technical operations discipline, including experience leading or managing technical teams. Strong understanding of cloud platforms, distributed systems, production operations, and modern site reliability engineering principles. Hands-on experience implementing or improving SLOs, SLIs, monitoring, alerting, incident response, and post-incident review processes. Demonstrated ability to hire, mentor, coach, and develop engineers while creating a healthy, inclusive, and accountable team culture. Strong communication and stakeholder-management skills, with the ability to explain complex technical concepts to engineers, product teams, and business leaders. Proven track record of using operational data, structured problem-solving, and technical insight to improve reliability and team effectiveness. Experience with infrastructure automation, CI/CD, disaster recovery, cost optimization, or multi-cloud environments. Ability to influence operational change across teams and drive adoption of improved processes, documentation, and reliability practices. Strong judgment and prioritization skills, with the ability to balance immediate operational needs against longer-term reliab
This page is generated from the committed OpenOpps static snapshot. Use the source posting or apply link for the employer's current canonical posting state.