openleverjobgether
Staff Software Engineer, Alerting Platform
Jobgether
LocationIndia
EmploymentFull-time
Posted2026-08-21T02:17:17.486000+00:00
Last observed2026-08-26 21:51:40.410433
Job idjobgether-jobgether:lever:3cc6828b-ec3e-49eb-9dec-c6fe8c5bf654
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Software Engineer, Alerting Platform based in India. This is a high-impact technical leadership role focused on building a reliable, intelligent alerting platform that helps customers distinguish meaningful signals from overwhelming operational noise. You will shape the architecture of an alerting engine spanning detection, correlation, delivery, and downstream notification. The role combines deep software engineering with systems thinking across security, monitoring, telemetry, and network data. You will work across multiple engineering teams, influencing critical architectural decisions while remaining hands-on where your expertise is most valuable. You will also help establish the technical foundations for AI- and LLM-driven investigation, troubleshooting, and remediation capabilities. This is an opportunity to influence platform strategy at organizational scale while mentoring engineers and raising engineering standards. Define the architecture for high-fidelity alerting across rule-based thresholds, machine-learning anomaly scores, correlation logic, and customer-defined alert rules. Own the reliability of the alerting pipeline end to end, from alert evaluation through Kafka-based delivery and downstream notification systems. Establish robust guarantees around webhook delivery, idempotency, reliability, and sustained-load performance. Drive correlation strategies that transform security tooling, monitoring telemetry, syslog, OpenTelemetry, and network protocol data into meaningful network and system topologies and dependency maps. Make topology and dependency context usable for AI and LLM-driven reasoning across investigations, troubleshooting, correctness analysis, and remediation workflows. Establish the technical foundations for generating alert definitions from normalized data models using LLM-powered capabilities. Provide architectural leadership across multiple teams and workstreams, stepping into complex initiatives where deep technical expertise is required. Balance strategic investments in detection and correlation capabilities with the long-term scalability, reliability, and maintainability of the alerting platform. Make organization-level architecture and technology decisions and communicate technical direction effectively to engineering and business stakeholders. Mentor engineers across the teams you support, promote strong engineering practices, and raise the technical quality of the overall platform. Contribute to additional strategic initiatives and projects that support the success and evolution of the engineering organization. Requirements: Demonstrated experience designing, building, or operating high-reliability alerting, detection, or notification systems in production at significant scale. Strong experience implementing and operating both rule-based and ML-driven detection or event-processing logic against real-world production traffic. Advanced experience with Apache Kafka, particularly producing and consuming event streams for reliable downstream processing and delivery. Strong understanding of telemetry and protocol data, including security tooling output, monitoring telemetry, syslog, OpenTelemetry, NetFlow or sFlow, SNMP, ICMP, and firewall logs. Ability to transform diverse technical data sources into meaningful network and system topologies, dependency maps, and operational context. Proven experience building systems where webhook reliability, idempotency, delivery guarantees, and performance under sustained load are critical. Expert-level proficiency in Go and/or Python, with strong experience across container orchestration and multi-cloud environments. Demonstrated ability to make organization-level architectural and technology decisions rather than focusing solely on individual implementation tasks. Strong systems-thinking, problem-solving, and technical com
This page is generated from the committed OpenOpps static snapshot. Use the source posting or apply link for the employer's current canonical posting state.