Data Engineer
firmable
Job Description
-
LLM-in-the-pipeline systems — extraction, enrichment, entity resolution, and semantic validation steps that run over millions of records a day with structured outputs, retries, and human-review fallbacks
-
Eval harnesses — labelled sets, scorers, and regression suites that gate every prompt or model change; precision/recall tracked per check, not vibes
-
Rules vs. LLMs — deterministic checks (dbt tests, data contracts, SQL) wherever structure allows; LLMs where semantic judgement is needed; the discipline to know which is which
-
Cost and drift — token budgets per pipeline, model routing (cheap models for classification, stronger ones for hard cases), drift detection on vendor updates
-
Observability — every LLM call logged with prompt version, model, cost, latency, and decision, alongside standard pipeline alerting that surfaces problems before they cascade
-
Core pipelines and warehouse — Airflow orchestration, dbt models across staging to mart, Snowflake performance and cost over billions of rows, AWS infrastructure as code
-
Matching and deduplication — embeddings and retrieval patterns for company and people entity resolution across 13 markets
-