Data Engineering Professional II

takeda

Bengaluru, India 4 Years Exp Posted 15d ago

Job Description

ML Lifecycle & Pipeline Automation

  • Design, build, and operate end-to-end ML pipelines (data ingestion → feature engineering → training → validation → deployment → monitoring) using Databricks (Delta Lake, MLflow, Unity Catalog, Feature Store, Workflows/Jobs) and AWS services.
  • Implement CI/CD for ML and data assets (e.g., GitHub Actions, GitLab CI, or Jenkins), including automated testing, environment promotion (dev → test → prod), and reproducible builds.
  • Stand up and maintain model registries, model versioning, and artifact lineage so every deployed model is traceable to its data, code, and configuration.

Cloud & Platform Engineering (AWS)

  • Build and manage ML infrastructure on AWS — e.g., SageMaker, Bedrock, S3, Lambda, ECS/EKS, Step Functions, ECR, IAM, CloudWatch — using Infrastructure as Code (Terraform or CloudFormation/CDK).
  • Integrate Databricks with AWS securely (Unity Catalog governance, cross-account access, VPC/networking, KMS encryption, secrets management).
  • Optimize compute and cost (cluster policies, autoscaling, spot strategy, job orchestration) without compromising performance or compliance.

Production Monitoring & Reliability

  • Implement model and data monitoring: drift detection, data-quality checks, performance/SLA tracking, and automated alerting/retraining triggers.
  • Establish observability and incident-response practices for ML services; participate in on-call/runbook ownership as needed.
  • Maintain feature stores and data contracts to ensure consistency between training and serving.

 

Regulated-Environment & Compliance Engineering

  • Build ML systems that meet GxP expectations and support Computer System Validation (CSV) / Computer Software Assurance (CSA), GAMP 5, 21 CFR Part 11, and data-integrity (ALCOA+) requirements.
  • Implement audit trails, electronic records/signatures controls, access controls, and change-management workflows suitable for validated environments.
  • Handle PII/PHI and sensitive R&D data in line with HIPAA, GDPR, and internal privacy/data-governance policies (de-identification, anonymization, role-based access).
  • Author and maintain technical documentation, validation deliverables, and SOP-aligned procedures; partner with Quality/QA and Regulatory on audits and inspections.

Collaboration & Enablement

  • Work under the guidance of Director, Solution Engineering/Solution Architect to produce artifacts and deliverables that adhere to best practices at Takeda.
  • Partner with data scientists to productionize models (including LLM/GenAI and RAG applications) and to translate research code into robust, maintainable services.
  • Contribute reusable templates, accelerators, and self-service tooling that raise the engineering bar across teams.
  • Promote MLOps best practices, mentor peers, and document standards.

Required Qualifications

  • Bachelor’s degree in Computer Science, Engineering, Data Science, or a related field (or equivalent practical experience).
  • 4+ years of hands-on experience in MLOps, ML engineering, data engineering, or DevOps for data/ML systems.
  • Strong Databricks experience: Delta Lake, MLflow, Unity Catalog, Jobs/Workflows, and Spark (PySpark).
  • Strong AWS experience across compute, storage, and IAM, plus at least one ML service (SageMaker and/or Bedrock).
  • Proficiency in Python for production code (packaging, testing, typing), plus solid SQL.
  • Experience building CI/CD pipelines and using Git-based workflows.
  • Working knowledge of containerization (Docker) and orchestration (Kubernetes/EKS or ECS).
  • Experience with Infrastructure as Code (Terraform, CloudFormation, or CDK).
    • Understanding of ML lifecycle concepts: experiment tracking, model registry, feature stores, and model monitoring/drift.

Similar Openings for You