Lead Data Engineer
biospace
Job Description
-
Design and implement scalable data collection, storage, and processing pipelines to support enterprise-wide data needs.
-
Implement and maintain data governance frameworks and data quality checks within data pipelines to ensure compliance and reliability.
-
Build and optimize data models and data marts to support self-service analytics and reporting tools such as Tableau, Looker, and Power BI
-
Partner with data scientists to operationalize models by integrating them into production-grade pipelines, ensuring scalability, performance, and maintainability.
-
Collaborate with cross-functional stakeholders (business, product, analytics, and engineering teams) to translate business requirements into scalable data solutions and prioritize data initiatives.
-
Provide technical leadership and architectural guidance for data platform design, ensuring alignment with enterprise standards and long-term data strategy.
-
Lead design reviews, code reviews, and data architecture discussions to ensure best practices, reusability, and high-quality deliverables.
-
Collaborate with platform, DevOps, and security teams to ensure secure, cost-effective, and scalable data infrastructure.
-
Influence data roadmap and strategy by identifying opportunities for data platform enhancement, automation, and cost optimization.
Who USP is Looking For?
The successful candidate will have a demonstrated understanding of our mission, commitment to excellence through inclusive and equitable behaviors and practices, ability to quickly build credibility with stakeholders, along with the following competencies and experience:
Education
Bachelor’s degree in relevant field (e.g. Engineering, Analytics or Data Science, Computer Science, Statistics) or equivalent experience.
Experience
-
7+ years of experience in big data technologies such as Python, PySpark, and SQL for processing structured, semi-structured, and unstructured data.
-
Strong experience with AWS data services including Redshift, S3, Glue, Lambda, EventBridge, Postgres, Neo4j (Azure/GCP equivalents such as ADLS, Synapse, ADF acceptable).
-
Experience in building batch, micro-batch, and streaming pipelines (real-time / near real-time) using Lambda/Kappa architectures.
-
Hands-on expertise in designing and delivering enterprise-scale data platforms, including data lakehouse, data warehouse, data lake, and data marts.
-
Strong understanding and hands-on implementation of data modeling techniques including - Data Vault 2.0, Dimensional Modeling, Knowledge Graphs, One Big Table (OBT) approaches (Certification in at least one area preferred)
-
Experience with medallion architecture and metadata-driven data pipeline frameworks.
-
Strong expertise in data governance frameworks, including- Data discovery, Data quality, Data security, Hands-on experience with DQ tools such as Great Expectations, Pydantic, etc.
-
Strong SQL and programming skills for data transformation, modeling, and analysis.
-
Hands-on experience in building and maintaining complex ETL/ELT pipelines and managing day-to-day data operations.
-
Experience with workflow orchestration tools such as Airflow (2+ years or equivalent).
-
Knowledge of streaming/event-driven architectures and modern data processing patterns.
-
Good understanding of dashboarding and visualization techniques with tools like Tableau, Power BI, or equivalent.
-
Experience with Agile (SAFe) methodologies, CI/CD pipelines, and modern deployment practices for data platforms.
-
Exposure to AI/ML concepts, with familiarity in Generative AI patterns (e.g., RAG, chunking techniques) as an added advantage.
-