I will prepare your dataset for llm fine tuning or rag ingestion

Pakistan

I speak Urdu, English

Fix the cause, not the symptom

I’m a Full Stack Developer specializing in AI, automation, and modern web applications. I build AI assistants, LLM integrations, RAG systems, SaaS platforms, React/Next.js apps, backend APIs, Telegram...
About this Gig

I clean and prep your dataset for LLM fine-tuning or RAG ingestion exported in the format your trainer expects (JSONL / Parquet / HuggingFace / OpenAI / Axolotl / LLaMA-Factory).


What you get:

Missing values handled, dedup (exact + fuzzy), outliers treated

Label audit + stratified train/val/test splits

Text cleaning: HTML strip, Unicode normalize, language detection

Format export for HuggingFace Datasets, OpenAI JSONL, Axolotl, or LLaMA-Factory

Data validation report (Great Expectations)


Stack: Pandas, NumPy, Polars, LangChain, LlamaIndex, HuggingFace Datasets.


Dataset size: 50K rows (Basic) / 500K rows (Standard) / 5M rows (Premium). RAG chunking + embedding export included on Premium.


Message me with one sentence about your use case I'll reply within 2 hours with a fixed quote.

Programming Language:

Python

AI Model Frameworks & Tools:

Hugging Face Transformers

PyTorch

Data Type:

Text

AI Engine:

GPT

Llama

Other