Principal Data Engineer

kmart

Bengalor 10 Years Exp Posted 1h ago

Job Description

  • Cloud Platforms — Multi-cloud fluency across AWS, GCP, and Azure ; with strong grasp of compute, storage, networking, IAM, and FinOps principles
  • Data Warehousing & SQL — Experience with cloud data warehouses (Snowflake, BigQuery, Databricks SQL); query engines (Trino, Spark SQL); advanced SQL and data modeling patterns (lakehouse, data vault, dimensional modeling)
  • Programming Languages — Proficient in Python/Java/Scala for data engineering, ML pipelines, and distributed systems; SQL for advanced analytics and performance tuning
  • Streaming & Messaging — Hands-on experience with Kafka, Kafka Streams, KSQLDB, and Kafka Connect (Confluent preferred); real-time processing using Apache Flink or Spark Structured Streaming; schema management with Avro or Protobuf
  • Data Quality, Transformation & Governance — dbt for transformation and testing; data quality frameworks (Great Expectations, Soda); open table formats (Iceberg, Delta Lake); data cataloging and lineage (DataHub, OpenLineage); GDPR and CCPA compliance
  • DevOps & Orchestration — Containerization with Docker and Kubernetes; CI/CD pipelines (Jenkins, GitHub Actions); workflow orchestration using Airflow, Prefect, or Dagster; infrastructure as code with Terraform or Pulumi
    • Knowledge working with ML & AI Platform — ML lifecycle management with MLflow or Kubeflow; feature store; model serving (Ray Serve, Triton, BentoML); LLM/GenAI platforms including RAG pipelines, vector databases, and AI gateways

Similar Openings for You