Full-Stack AI Engineer
roche
Job Description
Agentic & GenAI Application Development:
- Design and build advanced AI agentic systems, state machines, search-based conversational systems that solve complex business problems.
- Develop workflows leveraging Large Foundational Multimodal Models to process and reason across text, audio, and video modalities.
- Implement Model Context Protocol (MCP) servers/clients to standardize context exchange between agents, data sources, and external tools.
- Collaborate with AI Architects, Product Owners, and fellow developers to integrate AI capabilities into scalable, fair, and ethical end-user applications focusing on relevance and real-time performance.
Full-Stack Engineering & Agentic SDLC:
- Leverage AI coding agents (e.g., Claude Code) daily to accelerate full-stack development cycles, maintaining high productivity across frontend, backend, and infrastructure tasks.
- Take end-to-end accountability for features: write high-quality, production-ready Python (and occasionally TypeScript) code with comprehensive testing and documentation.
- Manage the DevOps/MLOps lifecycle: containerize applications using Docker, configure CI/CD pipelines, and architect high-throughput, reliable cloud-native solutions on AWS/multicloud.
Data Science, EDA & Strategy:
- Perform thorough Exploratory Data Analysis (EDA) to understand dataset characteristics, uncover patterns, detect biases, and identify data quality issues.
- Use statistical and visualization techniques to inform feature engineering, model selection, and optimization of foundation model-based applications.
- Design robust data pipelines to curate, preprocess, and structure diverse datasets that maximize LLM effectiveness and reduce bias.
Algorithm Development & Optimization:
- Design, customize, optimize, and fine-tune LLM-based and traditional AI algorithms for specific use cases (e.g., text generation, summarization, AI agents, sequence modeling).
- Lead advanced prompt engineering strategies, utilizing zero-shot, few-shot and other paradigms to optimize model outputs without extensive fine-tuning.
- Implement pre-generative AI models (e.g., classification, clustering, regression) when they provide a more efficient, interpretable, or cost-effective solution compared to LLMs.
- Optimize model inference speed, reduce latency (cold start reduction, caching strategies), and manage resource usage across cloud architectures.
Evaluation, Observability & Continuous Improvement:
- Conduct rigorous experimentation (A/B testing) and implement automatic metric pipelines (e.g. BLEU/ROUGE, RAG retrieval accuracy, human rating frameworks, etc.) to evaluate generative and multimodal systems.
- Implement real-time algorithms monitoring and observability practices, ensuring visibility into pipelines behavior, drift detection, and anomaly identification using telemetry tools.
- Translate complex technical results into clear, actionable insights for stakeholders, driving data-driven decision-making.