MLOps Engineer
jobylon
Job Description
- Deploy, version, and manage ML model endpoints in production — including LLM and computer vision pipelines — ensuring reliability across multi-region AWS infrastructure.
- Build and maintain CI/CD pipelines for model delivery, incorporating automated testing, staged rollouts, and safe rollback procedures.
- Monitor model and system health: track performance metrics, detect drift, manage alerts, and lead incident response when things go wrong.
- Optimise cloud infrastructure costs and resource usage across SageMaker, Bedrock, Lambda, and related AWS services.
- Manage experiment tracking and model/data versioning, maintaining clean lineage across training, validation, and deployment stages.
- Write and maintain infrastructure-as-code to keep our environments reproducible, auditable, and scalable.
What are we looking for?
- 4–5 years of hands-on experience in software engineering, data engineering, or ML engineering, with meaningful exposure to production systems beyond notebooks or coursework.
- Solid end-to-end understanding of ML workflows — data preparation, training, validation, deployment, and monitoring — and the real-world challenges of keeping models reliable in production (latency, versioning, rollback, drift).
- Hands-on AWS experience, including core services (S3, SQS, Lambda) and ideally SageMaker and Bedrock; comfort with containerisation and infrastructure-as-code.
- Strong Python skills, solid version control practices, and a working understanding of CI/CD concepts and production observability (logs, metrics, traces, dashboards).
- Full professional proficiency in English.