Data Engineer

citi

pune NM Years Exp Posted 6h ago

Job Description

  • Data Pipeline Architecture & Development: Design, develop, and deploy scalable batch and real-time end-to-end data pipelines (ETL/ELT) using Python, PySpark, Spark SQL, and Databricks(Must have).
  • Cloud Infrastructure Integration: Deploy and maintain Databricks workspaces on cloud environments (AWS or GCP). Manage secure integrations with cloud storage (S3/GCS), access controls (IAM), secrets management (Vault/KMS), and serverless query engines.
  • Performance Optimization & Tuning: Diagnose and resolve performance bottlenecks in Spark clusters, SQL queries, and Databricks jobs. Optimize storage layouts using Delta Lake properties (e.g., Z-Ordering, partitioning, and vacuuming).
  • Data Quality & Governance: Implement automated data validation frameworks, data quality monitoring, and metadata management solutions utilizing Databricks Unity Catalog to ensure strict compliance with internal data governance policies and external financial regulations (such as BCBS 239).
  • Technical Leadership & Mentorship: Act as a technical lead within an Agile/Scrum environment. Lead peer code reviews, enforce coding standards, and mentor junior and mid-level data engineers (C10/C11).
  • DevOps & CI/CD: Establish and maintain automated CI/CD pipelines (using Jenkins, GitLab, or GitHub Actions) for packaging and deploying data engineering artifacts (dbt, Spark jobs, Databricks workflows).
  • Collaboration: Partner with Data Science teams to operationalize machine learning models, and work with business intelligence developers to build efficient semantic layers for reporting.

Technical Qualifications (Must-Haves)

  • Programming Languages: Strong, production-grade proficiency in Python (including standard libraries, pandas, and testing frameworks like pytest) and advanced SQL (including window functions, CTEs, and query optimization).
  • Distributed Computing: Deep hands-on experience with Apache Spark (PySpark) for processing multi-terabyte datasets in a distributed cluster environment.
  • Unified Lakehouse Platforms: Minimum 3 years of hands-on experience developing within Databricks. Expert knowledge of Delta Lake ACID transactions, Delta Live Tables (DLT), Unity Catalog, and Databricks Workflows is required.
  • Cloud Platforms: Extensive experience deploying Databricks within either Amazon Web Services (AWS) or Google Cloud Platform (GCP). Proficiency in cloud-native components (AWS S3, EC2, IAM, EMR, Athena, Redshift OR GCP GCS, Compute Engine, IAM, Dataproc, BigQuery) is a strict requirement.
    • Data Modeling: Solid understanding of data warehousing concepts, including dimensional modeling (Star and Snowflake schemas), slow-changing dimensions (SCDs), and Medallion (Bronze/Silver/Gold) architecture design.

Similar Openings for You