openleverjobgether
Senior Site Reliability Engineer
Jobgether
LocationIndia
EmploymentFull-time
Posted2026-08-26T02:20:09.013000+00:00
Last observed2026-08-26 21:51:40.410433
Job idjobgether-jobgether:lever:d30aeed6-1dd1-435d-92a1-ddca4082c479
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in India. As a Senior Site Reliability Engineer II, you will help architect and operate secure, highly available cloud infrastructure supporting business-critical healthcare applications. You will take strategic ownership of AWS environments, driving reliability, scalability, performance, and cost optimization. The role combines hands-on engineering with technical leadership across Kubernetes, CI/CD, observability, and infrastructure automation. You will strengthen cloud operations through Infrastructure as Code, proactive monitoring, and resilient deployment practices. Security and compliance are central, with responsibility for maintaining rigorous HIPAA, GDPR, and SOC 2 standards. You will also mentor engineers, lead complex initiatives, and influence technical strategy across cross-functional teams. This is an opportunity to make a direct impact on healthcare technology while working in a collaborative, innovation-focused environment. Cloud Architecture & Reliability: Design, deploy, and continuously improve secure, scalable, and fault-tolerant AWS infrastructure, with a focus on availability, resilience, performance, and cost efficiency. Infrastructure Operations: Manage and optimize AWS services including EC2, S3, Lambda, and RDS while improving resource utilization and operational efficiency. Observability: Enhance and maintain monitoring and observability capabilities using DataDog, enabling proactive issue detection, performance analysis, and deep visibility across cloud environments. CI/CD & Deployment: Lead the evolution of Jenkins-based deployment pipelines, improving automation, release reliability, deployment velocity, and engineering confidence. Kubernetes: Manage and optimize containerized environments, improving scalability, consistency, resilience, and deployment practices. Infrastructure as Code: Champion automation through Terraform, Ansible, and related technologies to reduce manual effort, standardize infrastructure, and improve operational efficiency. Security & Compliance: Ensure cloud operations and infrastructure adhere to stringent security and regulatory requirements, including HIPAA, GDPR, and SOC 2. Technical Leadership: Lead complex engineering initiatives, influence the technical roadmap, and make architecture decisions that strengthen long-term platform reliability. Mentorship: Coach and mentor engineers, share technical expertise, encourage strong engineering practices, and foster a collaborative culture. Problem Resolution & Continuous Improvement: Proactively identify reliability risks, investigate complex incidents, and architect durable solutions that prevent recurring operational issues. Requirements: Experience: 9–12 years of professional experience in Site Reliability Engineering, Cloud Engineering, or a closely related discipline, with demonstrated ownership of large-scale AWS environments. AWS Expertise: Strong hands-on knowledge of AWS services such as EC2, S3, Lambda, and RDS, including experience with cost optimization and resource management. SRE Tooling: Proven experience with Kubernetes, DataDog, Jenkins, and modern cloud-native operational practices. Automation & Coding: Strong scripting capabilities in Python or Bash and professional experience with Infrastructure as Code tools, particularly Terraform. Cloud & DevOps: Strong understanding of cloud architecture, deployment automation, CI/CD, containerization, scalability, availability, and production operations. Security & Compliance: Experience implementing secure cloud operations and working with regulated environments or compliance frameworks is highly valuable. Problem Solving: Strong analytical and troubleshooting skills, combined with an ownership mindset and the ability to design systems that proactively prevent failures. Leadership: Demonstrated ability
This page is generated from the committed OpenOpps static snapshot. Use the source posting or apply link for the employer's current canonical posting state.