I will build full stack open source production data pipeline with cicd automation
Data Engineer
About this Gig
Are you looking for flexible data engineering support whether it's a
quick fix or a full pipeline build? I offer three engagement levels:
Hourly Support Quick debugging, small Spark/ETL tasks, or consultation
Part-Time Engagement Focused 20-hour pipeline or dbt model development
Full Project Build Complete end-to-end lakehouse setup (50+ hours)
I am a Data Engineer with hands-on experience building a full self-hosted
data lakehouse on Kubernetes processing ~5 million rows using Apache Spark,
Apache Iceberg, Trino, dbt-Trino, and Google BigQuery.
What I deliver:
Apache Spark ETL pipelines ingestion, cleansing, transformation
Data quality checks using PyDeequ
dbt-Trino Gold layer models Star Schema, Data Vault, OBT, Marts
Dagster orchestration and CI/CD automation
Full documentation (Standard & Premium tiers)
Message me before ordering so we can discuss your specific requirements.
Warehouse Platform:
Snowflake
•
BigQuery
•
Databricks
Project Type:
New Build
My Portfolio
FAQ
Q: Do you work with cloud platforms like AWS or GCP?
A: Yes — I support Google BigQuery, AWS S3, and any S3-compatible storage like MinIO.
Q: Can you work with my existing data infrastructure?
A: Yes — I can integrate with your existing databases, cloud storage, or data warehouse.
Q: Do you provide documentation with the code?
A: Yes — all packages include full code documentation and a delivery summary report.
Q: Can you handle real-time streaming pipelines?
A: Yes — I work with Apache Flink for streaming and Spark Structured Streaming for micro-batch processing.
Q: What if I need changes after delivery?
A: Revisions are included in all packages. Additional revisions are available as an extra.

