Senior AI Engineer
ecolab
Job Description
- Lead the design and implementation of responsible AI, safety, and quality engineering practices for GenAI and agentic AI products
- Define and operationalize AI validation strategies covering functional correctness, factual reliability, hallucination risk, retrieval quality, prompt safety, agent behavior, tool use, and failure handling
- Establish test approaches for LLM-powered applications, RAG systems, agentic workflows, multi-step reasoning, tool orchestration, and context-sensitive AI interactions
- Design and implement evaluation frameworks for relevance, safety, groundedness, consistency, latency, token usage, and business outcome alignment
- Work with AI engineers and architects to embed guardrails, prompt controls, model usage boundaries, escalation paths, fallback strategies, and safety-oriented design patterns
- Drive adoption of governance and auditability practices across AI solution development, including development of automated test scripts, enable traceability, review checkpoints, risk controls, and evidence collection
- Partner with engineering and platform teams to implement AI observability, logging, trace analysis, usage monitoring, cost awareness, and incident diagnostics
- Review solution designs to ensure they account for responsible AI, data sensitivity, model behavior risks, compliance expectations, and operational resilience
- Guide teams on how to test and validate agentic AI systems, including state management, tool-calling reliability, context integrity, multi-agent coordination, and autonomous decision boundaries
- Contribute to the definition of engineering standards for prompt lifecycle management, evaluation automation, red-teaming, adversarial testing, and regression prevention
- Collaborate with product, process, and engineering stakeholders to balance innovation speed with risk management, trust, and enterprise readiness
- Mentor engineers and quality professionals on AI testing, safety validation, responsible AI engineering practices, and production monitoring approaches
- Contribute reusable assets such as test harnesses, prompt evaluation templates, governance checklists, safety review frameworks, red-team patterns, and validation accelerators