I will build or optimize pyspark databricks data pipelines

United States

I speak English, Hindi

FullStack Agentic AI Agentic Systems Design and Automation, MLOps

Full-Stack AI/ML Engineer and Applied Statistician with 10 years across data, ML, and enterprise AI. I build agentic workflows, RAG systems, predictive models, data pipelines, APIs, and internal tools...
About this Gig

Need a PySpark or Databricks pipeline that is reliable, scalable, and understandable?


I will review, build, fix, or optimize a bounded data pipeline using PySpark, Spark SQL, Delta Lake, and Azure Databricks where appropriate.


I can address slow jobs, excessive shuffles, join skew, partitioning, pruning, window functions, temporal features, incremental processing, Delta Lake, data-quality checks, and Python UDF refactoring.


Your delivery can include corrected code, performance findings, pipeline implementation, validation checks, benchmark guidance, documentation, and a technical handoff.


My recent work includes distributed processing over billions of package-network records in Azure Databricks without Python UDFs. I focus on explainable, maintainable improvements.


Packages cover defined jobs, sources, and transformations. Cloud charges, production credentials, platform administration, and unrelated dashboards are excluded unless included in a custom offer.


Please message me before ordering with the code, schemas, data scale, current runtime, and desired outcome.

Langugae:

English

Hindi

Technical expertise:

Apache Spark

Databricks

Snowflake

Expertise:

Data Pipelines

ETL Development

Industry:

Data analytics

Legal

Marketing & advertising

My Portfolio