GenAI Lead Engineer
persistent
Job Description
We are looking for a GenAI/LLM Engineer with strong Python engineering skills and proven experience building production-grade Retrieval-Augmented Generation (RAG) systems using LlamaIndex and/or LangChain, and integrating vector databases (Pinecone preferred). The role will focus on designing scalable RAG pipelines, implementing advanced retrieval strategies, and building MCP/tool-calling connectors to expose enterprise APIs as agent tools for read/write operations.
Key Responsibilities
1) RAG Pipeline Design & Development
- Design and develop end-to-end RAG pipelines, including:
- Data ingestion
- Document parsing and preprocessing
- Chunking strategies
- Embedding generation
- Indexing into Pinecone (preferred)
- Retrieval and response generation
- Build production-ready semantic retrieval solutions and continuously improve relevance/grounding quality.
- Implement and optimize advanced retrieval strategies, including semantic search and retrieval tuning.
2) Agent Tooling & MCP Integrations
- Build and integrate MCP connectors to expose internal/external system APIs as agent-callable tools (read/write).
- Contribute to agent orchestration patterns including:
- Intent routing (e.g., deciding between RAG vs MCP vs workflow)
- Tool selection and execution sequencing
- Agent reliability patterns (fallbacks, retries, observability)
3) Security, Reliability & Performance
- Apply security controls and handle authentication/authorization tokens, ensuring safe access to enterprise systems.
- Optimize AI/ML workflows for performance, scalability, and reliability (latency, throughput, cost, robustness).
- Ensure seamless deployment and integration across environments in collaboration with platform/DevOps teams.
4) Cross-functional Collaboration
- Work closely with product, backend, data engineering, and platform teams to ensure successful integration and delivery.
- Contribute to design discussions, technical documentation, and best practices for GenAI application engineering.