Data Engineer
vanguardjobs
Job Description
- Build and operate production data pipelines on AWS — Glue (PySpark), Step Functions, Lambda, Athena, SNS, and S3 — from ingestion through transformation to reconciled, reportable output.
- Own operational support for production data pipelines, including monitoring, incident management, PagerDuty response, root cause analysis, and driving timely resolution of production issues while continuously improving platform reliability and observability.
- Develop and refactor transformation logic in PySpark, including migrating legacy Pandas/SQL logic and performance-tuning large-scale Spark jobs.
- Implement data reconciliation and sign-off — design recon checks (including parity between old and new implementations), investigate discrepancies to root cause, and produce the evidence the business needs to trust a run.
- Translate complex business and financial logic (asset classification, valuation and pricing rules, methodology calculations) into correct, testable, well-documented code.
- Take features through production readiness — pre-checks, access/role setup, environment promotion, and clear runbooks for ongoing support.
- Contribute to our cloud lakehouse migration — modelling refined tables, building repeatable ETL, and lifting reporting workloads onto a modern Delta/Databricks platform.
- Keep running pipelines healthy — monitor scheduled jobs, triage failures (partitioning, schema/data-type issues, backfills), and keep periodic (month-end / quarter-end) processing on time.