Senior AI Engineer
sapiens
Job Description
Agent architecture & orchestration
- Design and implement agentic systems capable of multi-step reasoning, planning, tool use, and workflow execution against complex, document-intensive implementation processes
- Build stateful workflows using LangGraph or equivalent — including branching, retries, self-correction, human-in-the-loop checkpoints, and reusable orchestration patterns
- Engineer for long-horizon reliability — multi-step task completion, recovery from compounding errors, planning under uncertainty, and robust tool use when individual steps fail
- Build the reasoning behind high-stakes implementation decisions — criteria-grounded outputs, structured review patterns, and auditable rationales that delivery consultants can act on and defend
Retrieval, grounding & context engineering
- Develop end-to-end RAG pipelines: ingestion, chunking, embeddings, vector and hybrid retrieval, reranking, contextual compression, and grounding strategies
- Engineer memory and context management — conversational state, persistent memory, retrieval-aware context assembly, and token-efficient context selection
- Apply MCP-style tool and context interfaces so agents access the right information at the right time across enterprise knowledge repositories, document sources, and structured configuration data
Reliability, evaluation & safety
- Implement observability and tracing for prompts, tool calls, retrieval quality, agent traces, failures, drift, latency, and production behaviour
- Apply guardrails, safety controls, and failure-handling to reduce hallucinations in agents whose outputs practitioners act on directly in live client settings
- Evaluate agents at trajectory and task level — multi-step task success, failure-mode and regression analysis, sandboxed test environments — alongside retrieval and generation quality metrics, automated checks, and human review
Integration & production craft
- Build integrations with internal and external tools, APIs, enterprise systems, databases, and model providers so agents operate reliably within real delivery workflows
- Deliver production-quality Python code with strong practices in testing, CI/CD, logging, versioning, and documentation; make architecture decisions that balance quality, reliability, latency, cost, and model risk
- Translate ambiguous, high-complexity implementation processes into robust system logic and reusable AI patterns; stay current with advances in agentic systems and translate research into practical engineering decisions