Senior Data Engineer
caterpillar
Job Description
- Design, develop, and maintain scalable batch and streaming data pipelines using PySpark, Azure Databricks, and Microsoft Fabric.
- Architect and implement Medallion Architecture across Bronze, Silver, and Gold layers for enterprise data processing and analytics.
- Build robust ETL/ELT solutions for data ingestion, transformation, validation, reconciliation, and delivery across multiple source systems.
- Develop and optimize PySpark and Spark SQL workloads for high-volume structured, semi-structured, and unstructured data.
- Design and maintain Lakehouse and data lake solutions using Azure Data Lake Storage Gen2, Delta Lake, Microsoft Fabric OneLake, Fabric Lakehouse, and Warehouse.
- Implement integration solutions using Azure Data Factory, Fabric Data Factory, data pipelines, notebooks, and Dataflows Gen2.
- Design secure and governed data-sharing solutions across workspaces, domains, business units, and approved external consumers.
- Implement reusable data products and cross-workspace sharing patterns using OneLake, OneLake shortcuts, Lakehouse, Warehouse, and semantic models.
- Contribute to Microsoft Fabric capacity planning, workspace-to-capacity assignment, workload monitoring, utilization analysis, and performance optimization.
- Monitor Fabric workloads using available capacity and workload metrics, identify resource contention, and recommend workload or scheduling improvements.
- Design domain-aligned Fabric workspace structures with appropriate separation for development, testing, production, security, and ownership boundaries.
- Implement data quality controls, monitoring, observability, lineage, error handling, reconciliation, and auditability across data pipelines.
- Apply security best practices using managed identities, role-based access control, workspace roles, row-level or object-level controls where applicable, and secure secrets management.
- Integrate data engineering solutions with Git-based source control and CI/CD pipelines for automated testing and deployment.
- Optimize performance, scalability, reliability, and cost across Azure Databricks and Microsoft Fabric workloads.
- Collaborate with Data Architects, Product Owners, Business Analysts, Data Scientists, QA engineers, governance teams, and platform teams in an Agile/Scrum environment.
- Provide technical leadership, conduct design and code reviews, establish engineering standards, and mentor data engineers.