I will build data pipelines that
About this Gig
I build production-style ETL pipelines using PySpark for transformation, Airflow for orchestration, and Delta Lake for storage the same stack I use in my day-to-day data engineering work.
What's covered:
- Bronze-to-Silver (raw-to-cleaned) transformation logic in PySpark
- Airflow DAG design with proper scheduling and dependency handling
- Delta Lake table setup with versioning
- Optional: Trino/Metabase dashboard on top of the cleaned data
- Fully in-house, self-hosted pipeline architecture no dependency on third-party managed ETL platforms (Fivetran, Airbyte Cloud, etc.) if you'd rather own the full stack
I'll ask about your data volume and source format before starting pipeline design changes a lot between a 10k-row CSV and a streaming source, and I'll scope honestly rather than reuse a generic template.
FAQ
Do you work with cloud data sources (S3, GCS, Azure Blob)?
Yes — specify your source in requirements.
Can you fix an existing broken pipeline instead of building new?
Yes, message me first with details — pricing differs from new-build gigs.
Do I need my own Airflow instance?
No, I can set one up (local/Docker) as part of delivery, or work within yours if you provide access.
Can you build this without relying on managed third-party ETL tools?
Yes — I build fully in-house/self-hosted pipelines (PySpark + Airflow + Delta Lake) rather than wiring together tools like Fivetran or Airbyte Cloud. Available at the Premium tier.
