I will build production pyspark, python and sql data pipelines

India

I speak English, Hindi

Data and Backend Engineer

13+ years building data platforms and backend systems — Goldman Sachs, Walmart Global Tech, Ola, and two funded startups where I led teams of 6-12 engineers. What I actually do: build pipelines that ...
About this Gig

I have spent 13+ years building data pipelines at Goldman Sachs, Walmart Global Tech, Ola and two venture-backed startups. Pipelines I have owned move terabytes a day. Now I will build yours.


WHAT YOU GET

- A working PySpark / Python / SQL pipeline, not a notebook dump

- Idempotent and re-runnable, so it is safe to replay after a failure

- Handles nulls, duplicates, late data and schema drift instead of crashing on them

- Readable commented code, plus a short video walking you through it


I WORK WITH

PySpark, Pandas, Polars, SQL (Postgres, MySQL, Redshift, BigQuery), Airflow, dbt, Kafka, Iceberg, AWS (S3, EMR, Glue, Lambda)


TYPICAL JOBS

- Messy CSV/JSON/Parquet turned into a clean modeled table

- A Spark job or query that needs to actually finish

- A scheduled extract from an API or database

- Window functions, dedup logic, incremental loads, CDC


BEFORE YOU ORDER

Message me with a data sample and what the output should look like. I will tell you which package fits, or that I am not the right person for it. Scoping is free.


I am in IST and work overlapping US and EU hours.

Destination Platform:

Google BigQuery

Amazon Redshift

ClickHouse

Tools & Platforms:

Hevo Data

Kafka Connect

Debezium