openleverjobgether
Senior Software Engineer — Infra Agent Systems
Jobgether
LocationIndia
EmploymentFull-time
Posted2026-08-26T02:12:23.461000+00:00
Last observed2026-08-26 21:51:40.410433
Job idjobgether-jobgether:lever:d1e314bb-5c2a-4aba-be1c-8b81fc38cae5
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Software Engineer — Infra Agent Systems based in India. This is a high-impact engineering role focused on building production AI agents that operate and automate large-scale GPU infrastructure. You’ll design systems that diagnose hardware failures, investigate incidents, gather evidence, and support remediation across complex infrastructure environments. The role spans AI agent systems, distributed services, knowledge graphs, retrieval, orchestration, and developer tooling. You’ll own systems end to end, from architecture and implementation through deployment, observability, and production operations. You’ll collaborate across infrastructure, datacenter, and engineering teams to turn operational knowledge into reliable automation. This is an opportunity to help shape how autonomous AI systems can safely and intelligently operate real-world infrastructure at massive scale. Design and build production AI agent systems capable of diagnosing, investigating, and supporting remediation of infrastructure issues across large-scale GPU environments. Develop the distributed services, orchestration frameworks, knowledge graphs, retrieval systems, and supporting infrastructure that power AI agents. Build fleet intelligence capabilities that combine telemetry, infrastructure state, operational knowledge, and historical incidents to improve agent decision-making. Integrate agent systems with observability, incident management, ticketing, fleet inventory, source control, communication platforms, and internal infrastructure through reliable APIs. Own services throughout their lifecycle, including architecture, implementation, testing, deployment, monitoring, reliability, and production support. Improve agent quality and reliability through evaluations, retrieval optimization, better tools, and continuous feedback from production environments. Convert insights and knowledge generated through production use into reliable, reviewed software, workflows, and automation. Contribute to the evolution of platform architecture and engineering practices as autonomous infrastructure capabilities scale. Requirements Bachelor’s degree or equivalent professional experience in Computer Science, Engineering, or a related technical field. 5+ years of professional experience building production backend systems, distributed systems, infrastructure platforms, or similarly complex software. Strong systems design capabilities and demonstrated experience taking significant systems from initial architecture through production. Deep expertise in at least one relevant area, such as AI agent systems, orchestration, tool use, evaluation, grounding, knowledge graphs, graph data modeling, search, retrieval, ranking, RAG, or semantic search. Strong backend engineering skills, including API design, service boundaries, data modeling, and integrations across complex technical environments. Experience with Kubernetes, GitOps practices such as ArgoCD, infrastructure-as-code, and cloud platforms. Proficiency in one or more relevant programming languages, such as Go, TypeScript, Python, or Rust, with the ability to work across multiple languages when required. Strong analytical and problem-solving abilities, with an interest in solving ambiguous and technically challenging infrastructure problems. Ability to own systems in production, balancing engineering quality, reliability, operational requirements, and delivery speed. Experience with GPU infrastructure, datacenters, bare-metal environments, hardware failure modes, BMC/IPMI, or cluster schedulers is a plus. Experience with graph databases, event-driven systems, messaging platforms such as NATS or Kafka, or observability tools such as Prometheus and Grafana is advantageous. Experience building evaluation frameworks or improving the reliability and quality of LLM-powered systems is also a plus.
This page is generated from the committed OpenOpps static snapshot. Use the source posting or apply link for the employer's current canonical posting state.