Data Engineer

vanguardjobs

Hyderabad 5 Years Exp Posted 2d ago

Job Description

  • Build and operate production data pipelines on AWS — Glue (PySpark), Step Functions, Lambda, Athena, SNS, and S3 — from ingestion through transformation to reconciled, reportable output.
  • Own operational support for production data pipelines, including monitoring, incident management, PagerDuty response, root cause analysis, and driving timely resolution of production issues while continuously improving platform reliability and observability.
  • Develop and refactor transformation logic in PySpark, including migrating legacy Pandas/SQL logic and performance-tuning large-scale Spark jobs.
  • Implement data reconciliation and sign-off — design recon checks (including parity between old and new implementations), investigate discrepancies to root cause, and produce the evidence the business needs to trust a run.
  • Translate complex business and financial logic (asset classification, valuation and pricing rules, methodology calculations) into correct, testable, well-documented code.
  • Take features through production readiness — pre-checks, access/role setup, environment promotion, and clear runbooks for ongoing support.
  • Contribute to our cloud lakehouse migration — modelling refined tables, building repeatable ETL, and lifting reporting workloads onto a modern Delta/Databricks platform.
    • Keep running pipelines healthy — monitor scheduled jobs, triage failures (partitioning, schema/data-type issues, backfills), and keep periodic (month-end / quarter-end) processing on time.

Similar Openings for You