I will build real time data pipelines with apache kafka, spark and aws
About this Gig
PIPELINES THAT SURVIVE PRODUCTION, NOT JUST THE DEMO.
I build batch and streaming data pipelines that run every day without someone babysitting them.
WHAT I DELIVER
- Streaming ingestion: Kafka, Kinesis, RabbitMQ, NiFi, CDC from your database
- PySpark / Spark SQL transformations, tuned - not the default config
- Data lake on S3 with Apache Iceberg: schema evolution, time travel, MERGE upserts
- Warehouse loads into Redshift, Snowflake, Athena or BigQuery
- SCD Type 2 history, deduplication, idempotent re-runs, late-arriving data handling
- Airflow DAGs, retries, alerting and CloudWatch/Grafana dashboards
- CI/CD, tests and documentation so the thing is maintainable
TRACK RECORD
3.5+ years on a European telecom's streaming platform: 100+ production pipelines, 0.6 TB/day (18 TB/month), millions of users across 8 countries. 50%+ compute time cut across 100+ PySpark jobs. 60% fewer job failures. 99.9%+ uptime.
I also migrated a Redshift workload to Apache Iceberg: 40%+ faster queries at lower spend.
Send me your source, your destination and your volume. You'll get a fixed scope and a real timeline before you order.
My Portfolio
FAQ
Batch or streaming — which do I need?
If nothing in your business reacts within the hour, batch is cheaper and simpler and I'll tell you so. Streaming earns its cost for fraud, live personalisation, alerting and real-time dashboards.
Can you work inside our AWS account?
Yes, with a scoped IAM role. I can also build in my own sandbox against sample data and hand you deployable scripts, if you'd rather not grant access at all.
Do you fix existing broken pipelines?
Yes — and it's usually cheaper than a rebuild. Order my audit gig first, or message me with the job code and I'll quote a custom fix.
