Lead Data Engineer

firmable

Remote 6 Years Exp Posted 11d ago

Job Description

Data Sourcing and ETL architecture

  • End-to-end pipeline design: extraction, normalisation, deduplication, validation, and load; plus cost, performance, and reliability of the whole layer

  • Sourcing layer standards: coverage, accuracy, schema design; the rule-vs-LLM standard codified so the team applies it without you

  • Cost and performance optimisation: cost ceilings with the token math behind them, incremental processing, recovery design, scheduling

Extraction systems and frameworks

  • Extractor framework: patterns, abstractions, and tooling other engineers ship into; new extractors are fast to build and reliable to run

  • Hard-source extractors: anti-bot defences, JS-heavy rendering, schema drift, low-quality structure; plus the proxy and IP rotation strategy behind them

  • Agentic extraction pipelines: rule-based triage, LLM escalation, structured-output validation, retries, human-review queues

LLM infrastructure and observability

  • LLMs as production systems: versioned prompts, labelled eval sets, measured precision and recall, prompt versioning you can roll back, judges debugged on real data

  • Eval and observability scaffolding: eval frameworks, prompt versioning, traces, drift detection when a vendor silently updates a model

    • Model-choice playbook: cheap models for classification, stronger models for nuanced extraction, frontier models for hard edge cases; revised as model economics shift

Similar Openings for You