openleverjobgether
Senior Site Reliability Engineer
Jobgether
LocationUS
EmploymentFull-time
Posted2026-08-25T08:33:12.967000+00:00
Last observed2026-08-26 21:51:40.410433
Job idjobgether-jobgether:lever:9c54b709-582c-46f5-9d42-9e15fd841dc9
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in United States. This is an opportunity to join a critical AI Hardware SRE team responsible for the reliability of next-generation dedicated AI infrastructure. You will help scale and optimize high-density hardware and software environments across regional data centers. The role combines automation, observability, infrastructure engineering, networking, and real-time incident response. You will build Python-based tooling, infrastructure-as-code utilities, telemetry pipelines, and intelligent monitoring solutions. Your work will directly improve uptime, performance, scalability, and operational efficiency for business-critical systems. You will collaborate with engineering teams, infrastructure vendors, and field technicians to solve complex reliability challenges. This role is ideal for an experienced SRE who thrives on ownership, ambiguity, automation, and production-scale infrastructure. Develop and scale robust Python-based tooling, infrastructure-as-code utilities, and automation frameworks to eliminate operational toil and streamline fleet-wide provisioning. Build automated workflows and API integrations across corporate ticketing systems to accelerate resolution of hardware and network incidents. Apply modern AI and LLM-based development tools to improve technical execution, automate scripting, and evaluate complex infrastructure systems. Work with advanced private cloud and compute technologies to improve availability, latency, scalability, and overall health across high-density hardware environments. Design and implement telemetry pipelines, Prometheus and Grafana dashboards, and AI-driven anomaly detection for bare-metal and virtualized infrastructure. Define operational KPIs, monitoring standards, telemetry baselines, alerting thresholds, and operational readiness criteria for new services and infrastructure deployments. Participate in a 24x7x365 on-call rotation, leading real-time incident response and managing high-severity service disruptions through automated PagerDuty and Slack workflows. Develop detailed technical runbooks, lead incident response bridges, and drive blameless post-mortems that identify systemic improvements and prevent recurring issues. Partner with infrastructure vendors and coordinate on-site field technicians to support hardware reliability, break-fix activities, and uptime objectives. Collaborate across engineering and infrastructure teams to identify reliability gaps, establish best practices, and deliver production-grade solutions to ambiguous technical challenges. Requirements 5+ years of relevant Site Reliability Engineering, infrastructure engineering, systems engineering, or related experience, along with a Bachelor’s degree in Computer Science or a related technical field. Exceptional proficiency in Python and experience developing scalable operational tooling, API integrations, automation frameworks, and infrastructure utilities. Hands-on experience with modern observability technologies such as Prometheus, Grafana, OpenTelemetry, and Loki, as well as familiarity with time-series monitoring and telemetry systems. Strong understanding of advanced networking concepts, including high-bandwidth routing and switching, BGP, and dual-stack IPv4/IPv6 environments. Experience designing and launching new services with clear operational readiness requirements, telemetry baselines, monitoring strategies, and alerting thresholds. Extensive experience creating technical runbooks, leading complex incident response processes, and conducting comprehensive, blameless post-mortems. Strong understanding of distributed infrastructure, high-density compute environments, private cloud technologies, and large-scale content or infrastructure delivery challenges. Ability to leverage AI-assisted development tools and LLM-based approaches to ac
This page is generated from the committed OpenOpps static snapshot. Use the source posting or apply link for the employer's current canonical posting state.