Data Scientist
cummins
Job Description
-
Design, develop, and automate distributed data ingestion and transformation solutions using data from relational, event-based, semi-structured, and unstructured sources.
-
Build reliable, scalable, and efficient ETL/ELT data pipelines using appropriate tools, technologies, and scripting languages.
-
Design and implement data quality, validation, monitoring, and alerting frameworks to identify and resolve data integrity issues.
-
Implement data governance practices covering metadata, data access, retention, compliance, and security.
-
Design and implement physical data models, including database structures, indexing, and table relationships, to support performance and scalability.
-
Develop and operate large-scale data storage and processing solutions across cloud and distributed data platforms, including data lakes, warehouses, and lakehouse environments.
-
Optimize data pipelines, Spark workloads, databases, and cloud infrastructure for performance, reliability, scalability, and cost efficiency.
-
Integrate data from a variety of enterprise applications and source systems and support real-time and event-driven data processing.
-
Develop automation for common and repeatable data preparation, integration, deployment, and platform-management activities to minimize manual and error-prone processes.
-
Implement CI/CD and DevOps practices to support automated deployment, testing, and release management.
-
Participate in troubleshooting, testing, validation, and continuous improvement of data pipelines and platform solutions.
-
Ensure data platforms and solutions meet applicable quality, governance, security, compliance, and regulatory requirements.
-
Collaborate with data scientists, analysts, architects, IT teams, and business stakeholders to understand requirements and deliver effective data solutions.
-
Document data solutions, processes, designs, and technical information to support knowledge transfer and operational effectiveness.
-
Apply Agile development methodologies such as Scrum and Kanban to deliver data engineering initiatives.
-
Provide technical leadership and mentor less experienced team members, promoting engineering excellence and collaboration.
-