I will build or optimize pyspark databricks data pipelines
FullStack Agentic AI Agentic Systems Design and Automation, MLOps
About this Gig
Need a PySpark or Databricks pipeline that is reliable, scalable, and understandable?
I will review, build, fix, or optimize a bounded data pipeline using PySpark, Spark SQL, Delta Lake, and Azure Databricks where appropriate.
I can address slow jobs, excessive shuffles, join skew, partitioning, pruning, window functions, temporal features, incremental processing, Delta Lake, data-quality checks, and Python UDF refactoring.
Your delivery can include corrected code, performance findings, pipeline implementation, validation checks, benchmark guidance, documentation, and a technical handoff.
My recent work includes distributed processing over billions of package-network records in Azure Databricks without Python UDFs. I focus on explainable, maintainable improvements.
Packages cover defined jobs, sources, and transformations. Cloud charges, production credentials, platform administration, and unrelated dashboards are excluded unless included in a custom offer.
Please message me before ordering with the code, schemas, data scale, current runtime, and desired outcome.
My Portfolio
FAQ
Can you optimize an existing PySpark job?
Yes. I can review execution logic, joins, shuffles, skew, partitioning, caching, window functions, UDF usage, reads/writes, and job structure. Actual gains depend on data, cluster configuration, and the existing implementation.
Can you guarantee a specific performance improvement?
No. I will identify bottlenecks and validate improvements when representative execution access is available, but runtime depends on data distribution, cluster resources, concurrency, storage, and platform configuration.
Do you work only with Azure Databricks?
My strongest recent experience is Azure Databricks and PySpark, but Spark concepts also apply to other managed Spark environments. Confirm your platform before ordering.
Do you need access to production data?
Not necessarily. Sanitized samples, schemas, execution plans, Spark UI screenshots, logs, and reproducible test data are often sufficient. Never send unrestricted production credentials through Fiverr requirements.
Are cloud and cluster charges included?
No. The buyer pays Databricks, cloud, storage, database, and other infrastructure charges. Use a cost-controlled non-production environment whenever possible.
Will you provide source code and documentation?
Yes. Deliverables follow the selected package and can include notebooks, Python modules, SQL, configuration examples, tests, findings, README instructions, and an operations runbook.
Can you build an end-to-end platform?
This Gig covers bounded pipelines and connected jobs. A full data platform, organization-wide migration, governance program, or continuous production operation requires discovery and a custom milestone offer.
What counts as a revision?
A revision corrects delivered work against agreed requirements. New sources, transformations, notebooks, platforms, schemas, or changed business logic are additional scope.

