openleverjobgether
Senior Database Reliability Engineer
Jobgether
LocationUS
EmploymentFull-time
Posted2026-08-26T17:37:57.109000+00:00
Last observed2026-08-26 21:51:40.410433
Job idjobgether-jobgether:lever:6b93878c-3242-4255-b599-d89f3a34f5d0
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Database Reliability Engineer based in the United States. This is a high-impact engineering role responsible for the reliability, performance, availability, and cost efficiency of production databases powering a modern healthcare technology platform. You will take end-to-end ownership of database operations across EHR, data, and AI workloads, with a strong focus on preventing issues rather than simply reacting to them. The role combines deep database engineering with observability, automation, incident response, architecture, and cloud cost optimization. You will partner closely with engineering teams to improve query quality, establish safe migration practices, and ensure scalable database architectures. You will also play a key role in protecting sensitive healthcare information within a HIPAA-compliant environment. This is an opportunity to influence platform architecture, engineering standards, and operational practices as the organization continues to scale. The ideal candidate is a database expert who enjoys solving complex production problems and building systems that remain reliable under growing workloads. Own the reliability, performance, availability, and operational health of production databases running on AWS Aurora MySQL across EHR, Data, and AI workloads. Manage database observability end to end, maintaining the metrics and alerting pipeline into Datadog and integrating after-hours database alerts into the DevOps on-call rotation. Establish automated safeguards to identify and terminate long-running or runaway queries and provide immediate visibility into database activity across instances. Investigate database performance issues by analyzing execution plans, optimizing queries, re-indexing where appropriate, and reducing unnecessary database load and latency. Review application and EHR queries before production release, serving as a performance gate to prevent inefficient workloads from reaching production. Educate engineering teams on database hygiene, query optimization, and practices that improve reliability and performance. Create and maintain comprehensive database runbooks so first responders can resolve incidents quickly and consistently. Architect and continuously optimize Aurora reader topologies and read-routing strategies based on actual workload requirements. Own database cost efficiency through instance right-sizing, reserved-capacity and Savings Plan strategies, and appropriate storage tiering. Manage replication health, backups, restore testing, failover procedures, and disaster recovery capabilities. Establish safe, repeatable standards for database schema changes and migrations across engineering teams. Partner with platform architecture stakeholders to evaluate and determine the best technical approach for new and existing database workloads. Handle protected health information responsibly while maintaining database operations within a HIPAA-compliant environment. Requirements: 6+ years of experience in database reliability engineering, database administration, database engineering, or a closely related discipline, including ownership of production systems at scale. Deep expertise in MySQL, including query optimization, execution-plan analysis, indexing strategies, and replication; hands-on Aurora MySQL experience is strongly preferred. Experience operating large, multi-reader Aurora or RDS clusters at terabyte scale, including read-routing and connection-management strategies. Strong knowledge of database observability and monitoring technologies such as Datadog, Percona Monitoring and Management (PMM), Performance Insights, Prometheus, or Grafana. Proficiency with Python, Bash, or a comparable scripting language for automation, along with practical experience using infrastructure-as-code tools such as Terraform. Strong AWS operational knowledge, particular
This page is generated from the committed OpenOpps static snapshot. Use the source posting or apply link for the employer's current canonical posting state.