I will load hugging face and kaggle datasets into your database
Experienced Data Engineer ! Data Pipeline Specialist ! Data Migration
About this Gig
Need a public dataset inside your own database? I'll pull it from Hugging Face, Kaggle or any open-data source, clean it, and load it into your database, data warehouse or cloud storage, ready to query.
What I do:
Extract from Hugging Face, Kaggle, data.gov, AWS Open Data, BigQuery public datasets, APIs, and CSV, JSON or Parquet files
Clean and transform: correct data types, flatten nested JSON, remove duplicates and broken rows
Load into PostgreSQL, MySQL, SQL Server, MongoDB, Supabase, BigQuery, Snowflake, Redshift, Databricks, S3, GCS or Azure
Stream large datasets in batches, so size is rarely a problem
Incremental loads and scheduled refresh with Airflow, GitHub Actions or cron
Text datasets embedded and loaded into pgvector, Pinecone or Qdrant for AI and RAG
You get:
- Query-ready tables in your storage
- Clean, reusable Python code
- A README and a validation report (row counts, schema)
I check every dataset's license and flag any restrictions before loading.
Not sure which package fits? Send me the dataset link and your destination before ordering, and I'll recommend one.

