openleverjobgether
Staff Applied AI Engineer, Product & Agent Performance
Jobgether
LocationUS
EmploymentFull-time
Posted2026-08-21T09:39:19.744000+00:00
Last observed2026-08-26 21:51:40.410433
Job idjobgether-jobgether:lever:7088ff00-c5b7-4f6b-84ed-79a40603bd98
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Applied AI Engineer, Product & Agent Performance based in the United States. This is a staff-level individual contributor role focused on making AI agents reliable, measurable, and production-ready in complex healthcare workflows. You will shape agent behavior across prompting, retrieval, context, memory, tool use, evaluation, and human escalation. Your work will directly influence how AI systems perform across real-world, high-impact use cases serving healthcare organizations. You will establish rigorous evaluation practices that identify failure modes, quantify risk, and guide model and product release decisions. The role combines deep technical execution with product judgment, requiring you to translate production evidence into practical improvements. You will work closely with Product and Engineering to build AI systems that are accurate, steerable, transparent, cost-efficient, and trustworthy. This is an opportunity to define a scalable product-layer AI performance discipline in an environment where safety and responsible innovation matter. Design, implement, and continuously improve agent behavior across live, long-horizon, multi-turn, and multi-agent workflows. Architect retrieval and context strategies that deliver the right source data to models in the right structure while keeping agents grounded in reliable information. Design memory and state-management approaches for multi-turn and multi-agent experiences, determining what information should be retained, summarized, or discarded. Develop prompt and context templates using few-shot examples, structured formats, reasoning scaffolding, and other techniques to create consistent agent behavior. Improve agent performance through experimentation with prompting, tool-use strategies, retrieval, and context construction rather than relying on assumptions. Build production-representative evaluation suites and regression checks to measure accuracy, reliability, regressions, failure modes, edge cases, latency, and cost. Create evaluation rubrics, quality heuristics, and performance thresholds that account for the severity and business or safety impact of failures, not simply their frequency. Design and validate escalation mechanisms that route uncertain or high-risk cases to human review while maintaining safe and consistent behavior. Establish cost-aware approaches to AI performance, balancing accuracy, reliability, latency, context efficiency, and tool-call usage. Baseline existing behavior, conduct comparative evaluations, and assess model or system changes to make evidence-based go/no-go recommendations before customer release. Maintain product-level AI documentation, including model cards, intended-use guidance, limitations, known failure modes, and performance information. Partner closely with Product and Engineering to ensure agentic systems are not only capable but also steerable, trustworthy, transparent, and scalable. Translate production failures and performance evidence into clear diagnoses, experiments, fixes, and actionable recommendations for cross-functional teams. Requirements 8+ years of production software engineering experience, including at least 3 years of hands-on ownership of ML, LLM, or agentic systems in production. Professional experience working with AI systems in healthcare, finance, or another regulated environment where reliability, safety, and transparency are important. Demonstrated ability to diagnose agent failures and determine whether improvements should come from instructions, retrieval, context, memory, tool use, or other system components. Strong understanding of how to evaluate AI failures based on severity, risk, and cost rather than frequency alone. Hands-on experience designing and implementing RAG architectures and production-grounded evaluation frameworks. Experience developing fallback mechanisms,
This page is generated from the committed OpenOpps static snapshot. Use the source posting or apply link for the employer's current canonical posting state.