Data Engineer
darwinbox
Job Description
-
Design, develop, and maintain scalable batch and streaming data pipelines.
-
Build reliable ingestion frameworks supporting CDC, incremental processing, and near real-time data movement.
-
Develop optimized Spark and SQL workloads with a focus on scalability, reliability, and cost efficiency.
-
Build and maintain curated datasets and data marts to support analytics, reporting, and AI use cases.
-
Work with large-scale structured and semi-structured datasets using modern lakehouse architectures.
-
Collaborate with Product, Engineering, Analytics, and Data Science teams to translate business requirements into robust data solutions.
-
Ensure data quality through validation frameworks, monitoring, and automated testing.
-
Troubleshoot production issues, optimize pipeline performance, and continuously improve platform reliability.
-
Participate in architecture discussions, code reviews, and engineering best practices.
-
Contribute to reusable frameworks, platform tooling, and technical documentation.
Minimum Qualifications
-
Bachelor's degree in Computer Science, Information Technology, or a related engineering discipline.
-
2–4 years of professional experience in Data Engineering, Data Platform Engineering, or a related software engineering role.
-
Strong proficiency in SQL, including query optimization, window functions, joins, and analytical queries.
-
Strong programming skills in Python for building production-grade data pipelines and automation.
-
Hands-on experience with PySpark or Apache Spark for distributed data processing.
-
Good understanding of Spark fundamentals, including partitions, shuffles, caching, and distributed execution.
-
Experience designing and maintaining batch and/or streaming data pipelines.
-
Understanding of ETL/ELT, Change Data Capture (CDC), and incremental data processing.
-
Familiarity with modern Data Lake/Lakehouse concepts such as the Medallion Architecture (Bronze, Silver, Gold).
-
Good understanding of relational databases, data modeling, and database fundamentals.
-
Familiarity with columnar data formats such as Parquet, Avro, or ORC, and concepts like partitioning and schema evolution.
-
Strong debugging, analytical, and problem-solving skills.
-
Experience writing clean, maintainable, and well-tested production code.
-
Strong communication skills and the ability to collaborate effectively in cross-functional Agile teams.
Preferred Qualifications
-
Experience with one or more cloud-native data platforms such as Databricks, EMR, Snowflake, BigQuery, or Redshift.
-
Exposure to distributed query engines such as Trino or Presto.
-
Experience with streaming technologies such as Kafka, Spark Structured Streaming, or equivalent.
-
Familiarity with cloud platforms, preferably AWS, including services such as S3, EMR, Glue, Lambda, or IAM.
-
Exposure to ingestion frameworks such as AWS DMS, Kafka Connect, or similar technologies.
-
Familiarity with data governance and metadata management platforms such as Unity Catalog.
-
Experience using BI and self-service analytics tools such as Superset, Tableau, or Power BI.
-
Understanding of data quality, monitoring, and observability best practices.
-
Exposure to modern table formats such as Delta Lake, Apache Iceberg, or Apache Hudi.
-
Databricks or AWS Data Engineering certifications are an added advantage.
-