Data Engineer
prodapt
Job Description
1. Databricks Development
- Design and develop scalable data engineering solutions using Databricks.
- Build batch and near-real-time data ingestion and transformation pipelines.
- Develop data processing workflows using PySpark, Spark SQL and Python.
- Implement data models using Delta Lake and follow Medallion Architecture principles.
- Develop reusable and optimized data pipelines for large datasets.
- Implement data quality, validation, error handling and reconciliation processes.
- Optimize Spark jobs, SQL queries, Delta tables and cluster configurations for performance and cost.
2. Palantir Foundry to Databricks Migration
- Analyze existing Palantir Foundry pipelines, datasets, transformations and business logic.
- Map existing Foundry capabilities to appropriate Databricks technologies and patterns.
- Re-engineer Foundry data pipelines and transformations using Databricks, PySpark, SQL and Delta Lake.
- Identify opportunities to simplify, modernize and improve existing functionality rather than performing a one-to-one migration.
- Perform source-to-target data mapping and migration validation.
- Compare outputs between Palantir and Databricks to ensure functional and data accuracy.
- Support migration of complex data transformations and large-scale data comparison processes.
- Work with architects and business teams to resolve gaps between the existing Foundry implementation and the target Databricks solution.
3. Data Engineering & Platform Capabilities
- Implement data pipelines using Databricks Lakeflow / Spark Declarative Pipelines, where appropriate.
- Work with Unity Catalog for data governance, access control, lineage and discovery.
- Implement incremental processing, CDC and SCD patterns where required.
- Build data quality and observability capabilities.
- Implement CI/CD and deployment practices for Databricks workloads.
- Work with Databricks Asset Bundles and Git-based development practices.
- Develop and maintain Databricks Workflows/Jobs and production operational processes.
4. Integration & Application Support
- Integrate Databricks with enterprise data sources, APIs, databases, cloud storage and downstream applications.
- Work with structured and unstructured data sources.
- Support data requirements for dashboards, analytics and downstream applications.
- Collaborate with AI/ML teams where data pipelines are required to support AI and GenAI use cases.
5. Performance & Production Support
- Analyze and resolve production data pipeline issues.
- Perform root-cause analysis of data quality and processing failures.
- Optimize workloads for performance, reliability and cost.
- Establish appropriate monitoring, logging and alerting mechanisms.
- Participate in code reviews and enforce engineering standards.