Data Engineering
ripplehire
Job Description
Detailed Job Description Job Responsibilities: Responsible for ingesting data from multiple sources into a Data Lake using Java-based tools and frameworks such as Apache Spark, Spring Batch, and connectors for relational and NoSQL databases. Write transformation logic using Spark (Java API) and store processed data in HDFS or cloud storage solutions (e.g., AWS S3, GCP Storage, or on-prem equivalents). Responsible for writing stored procedures for relational databases (e.g., PostgreSQL, MySQL, or Oracle). Develop/re-develop existing ETL code and workflows into the modernized data pipeline. Experience with Batch and Realtime data processing and transformations. Minimum Job Requirements: Must have experience in writing SQL queries and Spark code using Java. Must have experience with Java frameworks for data processing and integration (e.g., Spring, Apache Spark, JDBC). Should be familiar with Data Warehousing concepts and data modeling techniques like star schema. Preferred Job Requirements: Good to have knowledge of Delta Lake or similar transactional storage formats in Spark. Knowledge of various Slowly Changing Dimensions (SCD) types. Basics of workflow orchestration tools (e.g., Apache Airflow, Oozie) and trigger mechanisms. Strong client communication skills. Educational Requirement - BE/BTech Experience Level (Min-Max) - 5+ years