I will build automated etl data pipelines using pyspark and python

India

I speak English
Data Engineer who architected and built a healthcare data platform from scratch, now processing 5M+ patient records and 10+ TB of clinical data sourced from EHR systems including eClinicalWorks (eCW),...
About this Gig

Need scalable, high-performance data processing for your business?

I am a professional Data Engineer specializing in PySpark, Python, SQL, and Cloud Data Architectures. I help businesses transform raw, unstructured datasets into clean, reliable, and actionable data pipelines.

What I can do for you:

  • Build robust batch and real-time ETL/ELT pipelines using PySpark.
  • Process massive CSV, Parquet, JSON, or SQL datasets with high performance.
  • Integrate PySpark with Databricks, AWS (S3, Glue, Redshift), GCP, and Azure.
  • Optimize slow PySpark code (repartitioning, caching, broadcast joins, out-of-memory error fixes).
  • Clean, validate, and structure dirty datasets for analytics or Machine Learning.

Why choose me?

  • Production-grade, fully commented code.
  • Comprehensive setup instructions.
  • 100% data security and confidentiality.

Destination Platform:

PostgreSQL

•

MySQL

•

Microsoft SQL Server

Tools & Platforms:

AWS Glue DataBrew

•

Azure Data Factory

•

Other