I will build etl data pipelines with pyspark and databricks
About this Gig
Hello, I am Arbind Sah, and I can help you build scalable ETL pipelines, process Big Data, or optimize cloud-based systems.
I specialize in Python, SQL, and cloud-native tools to streamline data integration, automation, and analytics for businesses of all sizes.
What I Can help you with:
-ETL & Data Pipelines Design and optimize ETL/ELT workflows using PySpark and Databricks for reliable, production-ready data transformation.Solutions Utilize Redshift, BigQuery, Spark, and Databricks to handle massive datasets efficiently.
-Cloud Infrastructure: Build and manage infrastructure on AWS (EC2, RDS, S3, Lambda) and GCP, including containerized deployments with Docker.
-Database Optimization & Security Improve query performance, implement robust security, and ensure compliance with industry standards.
-Data Migration & Schema Reconciliation Migrate and consolidate data across PostgreSQL, MySQL and cloud platforms, handling type mismatches, null values, and complex schema mapping.
With a strong focus on cloud computing, data automation, and analytics, I ensure scalable, secure, and high-performance data solutions.
Connect with me to optimize your data workflows for maximum efficiency!
My Portfolio
FAQ
What information do you need from me to get started?
Ideally your source data (or a sample), the target system/format, and what "success" looks like for the migration or pipeline. If you're not sure, message me and I'll help you figure out the scope.
Can you work with messy or inconsistent data?
Yes — handling nulls, type mismatches, and inconsistent formats is a core part of what I do, not an edge case I avoid.
Do you provide documentation with the delivered pipeline?
Yes, all pipelines come with clear documentation so your team can maintain and extend the work after delivery.
What if my data source or target isn't listed in your specialties?
Message me with details first — I work primarily with PostgreSQL, MySQL, PySpark/Databricks, and AWS/GCP, but I'm happy to discuss other stacks before you order.
Can you fix or optimize an existing pipeline instead of building a new one?
Yes — debugging and optimizing existing ETL pipelines is something I regularly do, not just building from scratch.
How do you handle large datasets?
I use PySpark and Databricks for distributed processing, so large-scale data is handled efficiently rather than choking a single script.
How long does a typical project take?
It depends on data volume and complexity — for a clear timeline, message me with your project details before ordering.
Do you sign NDAs?
Yes, I'm happy to sign an NDA if your project requires one — just send it over before we start.

