AI Engineer
shamrockus
Job Description
+ Build and deploy LLM-powered agents for enterprise data validation — reading specs, reasoning about business rules, identifying failure modes, and generating structured outputs
+ Design and own evaluation frameworks: automated test suites, LLM-as-judge pipelines, regression detection, and benchmarks that track whether our agents are improving
+ Build RAG pipelines that work reliably on real enterprise data — messy schemas, inconsistent formats, mixed structured and unstructured content
+ Integrate AI systems with enterprise infrastructure (SAP, Snowflake, Databricks, Postgres, REST APIs) with attention to latency, data residency, and compliance
+ Design agentic workflows with tool use, multi-step reasoning, and deterministic guardrails
+ Build observability tooling: trace agent reasoning, track output reliability, and detect hallucinations or drift in production
+ Work directly with FDSEs to understand real deployment failures and translate them into system improvements
The Stack:
+ Languages: Python (primary), Go, Node.js
+ AI/ML: LLMs (Claude, GPT-4, Command R+), RAG, vector databases, embeddings, fine-tuning
+ Evaluation: LLM-as-judge, automated eval pipelines, custom benchmarks
+ Data: Snowflake, Databricks, Postgres
+ Infra: containers, Kubernetes / ECS / Cloud Run
+ Tools: LangChain, LlamaIndex, OpenAI / Anthropic APIs, LangSmith
Compensation & Logistics:
Salary: INR 30 - 45 Lakhs (mid) / 70 Lakhs - 1 Cr+ (senior) depending on experience
Equity: Early-stage equity grant
Requirements
-
Production builder: you’ve shipped LLM-powered features real users depend on and debugged them when they broke
-
LLM practitioner: you understand hallucinations, retrieval failures, context limits, and what it takes to make agents deterministic enough for enterprise use
-
Systems thinker: you design for latency, failure modes, retry logic, and observability before features
-
Enterprise-aware: data residency, compliance, audit trails, and deterministic guardrails are first-class design constraints for you
Background That Maps Well:
-
3+ years in AI/ML or backend engineering with strong AI exposure
-
Hands-on production experience with LLM APIs (Anthropic, OpenAI, Cohere)
-
Experience designing evaluation frameworks: automated evals, regression tests, or LLM-as-judge pipelines
-
Strong Python; experience with LangChain, LlamaIndex, or similar agentic frameworks
-
Familiarity with RAG architectures: chunking, embedding models, vector DBs, retrieval quality
-