Senior Spark / PySpark Data Engineer

webspiders

kolkata 5 Years Exp Posted 11h ago

Job Description

  • Design, develop, and optimize large-scale ETL/ELT pipelines using Apache Spark and PySpark.
  • Develop scalable data transformation and processing solutions using PySpark and Python.
  • Build distributed data-processing applications capable of handling large volumes of data.
  • Develop reusable and maintainable Spark/PySpark frameworks and data-processing components.
  • Optimize Spark jobs for performance, scalability, memory utilization, and execution efficiency.
  • Work with complex transformations, joins, aggregations, partitioning, and large datasets.
  • Implement data validation, quality checks, error handling, and monitoring within data pipelines.
  • Work with AWS data services including EMR, Glue, S3, and Redshift.
  • Develop data pipelines supporting data lakes, warehouses, analytics, and downstream applications.
  • Troubleshoot production data pipeline and Spark processing issues.
  • Identify and resolve performance bottlenecks in Spark/PySpark workloads.
  • Collaborate with Data Engineering, Cloud, AI/ML, and Product teams to deliver reliable data solutions.

Must-Have Skills:

  • 5+ years of hands-on experience in Data Engineering.
  • Strong hands-on experience with Apache Spark.
  • Strong hands-on experience with PySpark.
  • Strong programming experience in Python.
  • Proven experience developing and optimizing large-scale ETL/ELT pipelines.
  • Strong understanding of distributed computing and data-processing concepts.
  • Experience working with large datasets and complex data transformations.
  • Strong understanding of Spark performance optimization and tuning.
  • Experience with cloud-based data engineering, preferably AWS.
    • Experience with Amazon S3 and at least one AWS data-processing service such as EMR or Glue.

Similar Openings for You