GCP Data Engineer
lighthouse
Job Description
- Data pipeline engineering: Design and implement scalable batch and streaming ETL/ELT pipelines using Dataflow, Apache Beam or Spark, BigQuery, and Cloud Storage.
- Data integration: Integrate data from SQL Server, Couchbase, APIs, files, and event-driven sources into governed data warehouses, data lakes, or lakehouse platforms.
- Data transformation and modeling: Transform raw data into trusted, analytics-ready datasets using SQL, Python, dbt, and Dataflow; apply sound dimensional and analytical modeling practices.
- Orchestration: Build, schedule, monitor, and support workflows using Cloud Composer and appropriate Google Cloud services.
- Performance and cost optimization: Monitor and tune pipelines, queries, storage patterns, and compute usage for reliability, throughput, scalability, and cost efficiency.
- Data quality and governance: Implement validation rules, reconciliation controls, observability, lineage, metadata, anomaly detection, and data-quality standards using Dataplex and related capabilities.
- Security and compliance: Apply least-privilege access, secure data handling, encryption, auditability, retention, and privacy controls in alignment with organizational and regulatory requirements.
- DevOps and automation: Provision and manage Google Cloud resources with Terraform; create automated build, test, deployment, and release workflows using Cloud Build and source-control practices.
- Event-driven solutions: Use Pub/Sub, Cloud Functions, and Datastream where applicable for near-real-time ingestion, change data capture, and event-driven processing.
- Collaboration and delivery: Partner with data scientists, analysts, application teams, security, platform engineering, and business stakeholders to translate requirements into dependable data solutions.
- Documentation and standards: Maintain architecture diagrams, data mappings, runbooks, operational procedures, deployment documentation, and coding standards.
- Innovation: Track relevant Google Cloud data, DevOps, and AI platform developments and assess their practical application to enterprise data-engineering use cases.