Staff ML Learning Engineer
adp
Job Description
AWS SageMaker Operations & Orchestration
• Design and implement SageMaker Pipelines with complex DAG authoring for end-to-end ML workflows including data preprocessing, training, evaluation, and deployment
• Configure and optimize SageMaker Processing Jobs, Training Jobs (PyTorch/HuggingFace containers), and Managed Spot Instances for cost-effective model training
• Manage SageMaker Model Registry with versioning, lineage tracking, and approval workflows for model governance
• Implement sophisticated endpoint deployment strategies including blue/green traffic routing, canary deployments, and A/B testing configurations
• Configure and maintain SageMaker Model Monitor for continuous data quality monitoring, model quality assessment, bias detection, and feature drift alerting
GenAI Model Lifecycle Management
• Build operational frameworks for Amazon Bedrock model deployments including provisioned throughput management, guardrail configuration, and usage monitoring
• Implement automated evaluation pipelines for foundation model outputs with quality gates and human-in-the-loop review workflows
• Design prompt versioning systems integrated with model registry for reproducible GenAI deployments
• Monitor GenAI model performance including latency, cost per inference, hallucination rates, and guardrail trigger metrics
Model Lifecycle Automation
• Develop automated retraining trigger systems based on data drift, performance degradation, scheduled intervals, or manual triggers
• Implement evaluation gates as code with configurable thresholds for metrics (F1, precision, recall, AUC, etc.) before model promotion
• Build automated model promotion workflows from dev → staging → prod with approval gates and rollback capabilities
• Design rollback automation with traffic shifting strategies to quickly revert to previous model versions upon performance degradation
Training Data Engineering
• Build AWS Glue ETL pipelines for training corpus assembly, transformation, and feature engineering at scale
• Design dataset versioning systems using S3 + manifest files ensuring reproducibility and lineage tracking across model training runs
• Establish reproducibility patterns including seed management, environment pinning, and deterministic data splitting
Container Engineering & Registry Management
• Build optimized Docker containers for training and inference workloads with multi-stage builds, layer caching, and security scanning
• Manage Amazon ECR repositories with lifecycle policies, image scanning, and vulnerability remediation workflows
• Implement bring-your-own-container (BYOC) patterns for SageMaker supporting custom frameworks and dependencies
• Optimize container startup times and resource utilization for cost-effective inference
ML Observability & Monitoring
• Design comprehensive CloudWatch dashboards for model health including accuracy, latency, throughput, error rates, and drift metrics
• Implement custom CloudWatch metrics for business-specific KPIs (e.g., clinical accuracy per condition, false positive rates for critical alerts)
• Build alerting systems for accuracy regression, data drift, concept drift, and model staleness with appropriate escalation paths
• Create observability frameworks that integrate with AWS X-Ray for end-to-end request tracing from API call through model inference
CI/CD for ML Artifacts
• Design and implement GitHub Actions or AWS CodePipeline workflows for automated ML artifact testing, validation, and deployment
• Build multi-stage promotion pipelines (dev → staging → prod) with automated testing gates and manual approval checkpoints
• Implement artifact versioning and lineage tracking for models, datasets, feature transformations, and deployment configurations
• Create integration testing frameworks for ML APIs including performance benchmarking and regression testing
Infrastructure & Compliance
• Collaborate with Full Stack AI/ML Engineers to de