openleverjobgether
Senior Site Reliability Engineer
Jobgether
LocationNetherlands
EmploymentFull-time
Posted2026-08-24T10:35:39.326000+00:00
Last observed2026-08-26 21:51:40.410433
Job idjobgether-jobgether:lever:343eae02-b034-4cce-90c4-3b611b0a33e1
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in Netherlands. This is a senior-level SRE opportunity focused on solving complex reliability, infrastructure, and platform challenges in a fully remote environment. You will take ownership of high-impact projects, from solution discovery through delivery, while helping shape platform architecture and long-term reliability strategy. Your work will span Kubernetes, cloud infrastructure, infrastructure as code, observability, CI/CD, and operational excellence. You will also play a key role in establishing strong SLOs, alerting, incident response, and reliability practices across engineering. AI is embedded into the way the team works, with an emphasis on building practical, reusable AI workflows that improve engineering productivity and safety. The role offers significant autonomy, technical influence, and opportunities to mentor engineers while collaborating asynchronously across a global organization. Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust and maintainable technical solutions. Contribute to platform architecture, infrastructure tooling, technical roadmaps, and engineering priorities, advocating for initiatives that improve reliability and developer experience. Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards . Use operational and incident metrics to identify systemic reliability issues and influence the team's technical strategy. Resolve cross-team infrastructure and platform requests while turning recurring problems into reusable solutions, automation, documentation, and runbooks. Operate and scale production Kubernetes environments and associated container infrastructure. Build and manage cloud infrastructure using AWS or comparable cloud platforms , with strong emphasis on reliability, scalability, and operational efficiency. Develop and maintain infrastructure as code using Terraform and support automated deployment workflows through modern CI/CD platforms. Use AI natively in infrastructure, operations, and development workflows, creating reusable prompts, skills, tooling, and agentic workflows that improve team-wide productivity and reliability. Design infrastructure and systems with AI-assisted engineering in mind, including clean interfaces, strong observability, secure-by-default patterns, CI protections, and review guardrails. Mentor less-senior engineers through actionable feedback, technical guidance, hiring, onboarding, and RFC discussions. Collaborate with Security on infrastructure hardening, threat mitigation, and defensive security practices. Contribute to infrastructure capacity planning, performance optimization, and cost efficiency. Participate in incident response and on-call rotations, helping maintain high standards of platform availability and reliability. Requirements Solid professional experience in Site Reliability Engineering, DevOps, Platform Engineering , or a closely related discipline. Strong hands-on experience operating and scaling Kubernetes in production, including Docker and the wider container ecosystem. Proven experience designing, building, and managing production cloud infrastructure using AWS or a comparable cloud provider . Strong practical expertise with Terraform and infrastructure-as-code principles. Hands-on experience with reliability engineering frameworks, including SLOs, SLIs, error budgets, alerting, and incident management . Strong observability experience with technologies such as OpenTelemetry, Grafana, Prometheus , or equivalent platforms. Experience designing and operating CI/CD pipelines using GitLab CI, GitHub Actions, or similar technologies. Comfortable working with Golang, Bash, and scripting , with broader programming experie
This page is generated from the committed OpenOpps static snapshot. Use the source posting or apply link for the employer's current canonical posting state.