I will build etl,elt data pipelines with python, spark, airflow, and dbt
Data and AI Engineer, Python, PySpark, Airflow and Cloud ETL
About this Gig
About this gig:
Need your data cleaned, transformed, and flowing reliably from source to destination? I build ETL/ELT pipelines using modern data engineering tools Python, PySpark, SQL, Apache Airflow, dbt, and Databricks following the Medallion Architecture (Bronze/Silver/Gold) for clean, production-style data layers.
Recent work includes a pipeline that ingests live data via API, orchestrates it through Airflow with automated failure alerts, transforms it through 10+ dbt models with 60+ data quality tests, and surfaces it in a Power BI dashboard plus a separate Databricks pipeline processing e-commerce data end-to-end with PySpark.
What I offer:
- ETL/ELT pipeline design and development (batch)
- Data cleaning, validation, and quality testing
- Medallion Architecture implementation (Bronze/Silver/Gold)
- Orchestration with Apache Airflow
- Data transformation with dbt, SQL, PySpark
- Databricks pipeline development
- API-based data ingestion
- Dashboard integration (Power BI)
Why work with me:
- Real, verifiable project work every deliverable is backed by a public GitHub repo
- Clear communication, realistic timelines
- Clean, documented, production-style code not throwaway scripts
Destination Platform:
PostgreSQL
•
Amazon S3
•
Other
Tools & Platforms:
Azure Data Factory
•
Other
My Portfolio
FAQ
What kind of data pipelines can you build?
I build ETL/ELT pipelines that extract data from APIs, databases, flat files (CSV/Excel), or websites. I clean, transform, and test that data using Python, PySpark, and dbt, and then deliver it in a structured format to your database or data warehouse (like PostgreSQL or Databricks).
I'm not technical, can you still help me?
Absolutely. Most of my clients aren't engineers. Tell me your end goal in plain English (e.g., "I need my daily sales data and inventory pulled into one place"), and I will handle the technical architecture and explain it simply.
Can you automate a pipeline to run daily or weekly?
Yes. I specialize in using Apache Airflow to orchestrate and automate workflows. Depending on your package, I can set up scheduled runs, task dependencies, and automated email alerts if any part of the pipeline fails.
Can you help with web scraping?
Yes. I can build web scrapers using Python to extract data from public websites, ensuring the project follows legal requirements and site rules. I can also integrate directly with APIs, which I always recommend first for better stability.
What is the "Medallion Architecture"?
It is a best-practice data design pattern that organizes data into three layers: Bronze (raw data), Silver (cleaned and filtered data), and Gold (business-ready data). I use this to ensure your final reports are built on highly reliable, tested data.
Can this integrate with BI tools?
Yes. I design the final layer of your data pipeline specifically so that tools like Power BI, Tableau, or other analytics platforms can connect directly and query the data efficiently.
What is AI data enrichment?
AI data enrichment means using AI or LLM tools to add more meaning and structure to raw, messy data. For example, unstructured customer feedback can be automatically categorized by sentiment, or specific entities (like names or locations) can be extracted from large text blocks.
Can you add AI or LLM enrichment to my data?
Yes. I can integrate AI/LLM workflows into your pipeline to handle tasks like summarization, classification, tagging, and converting unstructured text into clean, queryable database fields before it hits your dashboard.
Do you use data quality testing?
Yes. I strongly believe data quality should be baked in, not bolted on. I use dbt to implement automated tests (checking for nulls, duplicates, or accepted values) so your pipeline alerts you to bad data before it breaks your reports.

