I will build production ready data pipelines with apache spark and kafka
About this Gig
I build data pipelines that hold up in production, not just on a happy path.
Twelve years in banking and financial services means I've spent most of my time on the unglamorous parts: duplicate records, late arriving data, schema changes that break everything downstream, and jobs that need to be safely re-runnable. That's the difference between a pipeline that works in a demo and one you can leave running.
What I work with: Apache Spark, Kafka, Avro, Apache Iceberg, PostgreSQL, Java, Python, on GCP and Kubernetes.
Typical jobs I take on:
- Batch or streaming ETL from source systems into a warehouse or lake
- Turning fragile scripts into proper Spark jobs
- Kafka consumers and producers with sensible error and retry handling
- Fixing pipelines that are slow, flaky, or silently dropping data
I explain what I'm doing in plain English, and I'll tell you honestly if your problem needs a simpler solution than the one you asked for.
Message me before ordering with your data sources and rough volume, and I'll tell you which package fits or whether I'm the wrong person for the job.
Tools & Platforms:
Other

