I will build production pyspark, python and sql data pipelines
Data and Backend Engineer
About this Gig
I have spent 13+ years building data pipelines at Goldman Sachs, Walmart Global Tech, Ola and two venture-backed startups. Pipelines I have owned move terabytes a day. Now I will build yours.
WHAT YOU GET
- A working PySpark / Python / SQL pipeline, not a notebook dump
- Idempotent and re-runnable, so it is safe to replay after a failure
- Handles nulls, duplicates, late data and schema drift instead of crashing on them
- Readable commented code, plus a short video walking you through it
I WORK WITH
PySpark, Pandas, Polars, SQL (Postgres, MySQL, Redshift, BigQuery), Airflow, dbt, Kafka, Iceberg, AWS (S3, EMR, Glue, Lambda)
TYPICAL JOBS
- Messy CSV/JSON/Parquet turned into a clean modeled table
- A Spark job or query that needs to actually finish
- A scheduled extract from an API or database
- Window functions, dedup logic, incremental loads, CDC
BEFORE YOU ORDER
Message me with a data sample and what the output should look like. I will tell you which package fits, or that I am not the right person for it. Scoping is free.
I am in IST and work overlapping US and EU hours.
FAQ
Can you work with my data if it's confidential?
Yes. I'll work from a schema-only sample or synthetic data matching your structure, and I don't retain anything after delivery. Happy to sign an NDA before you send anything.
How big can the data be?
Comfortably into the terabytes — I've run pipelines processing ~3TB/day. Large jobs I run on Spark or DuckDB rather than pandas, so they actually finish.
Do I need Spark or just SQL?
Message me. Most of the time the answer is "you don't need Spark," and I'll tell you that even though the Spark version would be a bigger order.
Can I maintain this after delivery?
hat's the point. Commented code, a README, and a video walkthrough. If your team can read Python, they can own it.

