Lead Data Engineer
firmable
Job Description
Data Sourcing and ETL architecture
-
End-to-end pipeline design: extraction, normalisation, deduplication, validation, and load; plus cost, performance, and reliability of the whole layer
-
Sourcing layer standards: coverage, accuracy, schema design; the rule-vs-LLM standard codified so the team applies it without you
-
Cost and performance optimisation: cost ceilings with the token math behind them, incremental processing, recovery design, scheduling
Extraction systems and frameworks
-
Extractor framework: patterns, abstractions, and tooling other engineers ship into; new extractors are fast to build and reliable to run
-
Hard-source extractors: anti-bot defences, JS-heavy rendering, schema drift, low-quality structure; plus the proxy and IP rotation strategy behind them
-
Agentic extraction pipelines: rule-based triage, LLM escalation, structured-output validation, retries, human-review queues
LLM infrastructure and observability
-
LLMs as production systems: versioned prompts, labelled eval sets, measured precision and recall, prompt versioning you can roll back, judges debugged on real data
-
Eval and observability scaffolding: eval frameworks, prompt versioning, traces, drift detection when a vendor silently updates a model
-
Model-choice playbook: cheap models for classification, stronger models for nuanced extraction, frontier models for hard edge cases; revised as model economics shift
-