I will fine tune an open source llm on your custom dataset
AI Engineer, RAG Systems, LLM FineTuning, LangChain Python
About this Gig
Are you looking to fine-tune an open-source LLM on your
own data? I build custom fine-tuning pipelines that
actually improve model performance.
I recently completed EdgeAlign - a full SFT to DPO alignment pipeline on Qwen3-0.6B using LoRA adapters and 4-bit NF4 quantization.
Results:
- BLEU score: 3.70 (baseline) to 10.22 (+176% improvement)
- BERTScore: 0.7675 (baseline) to 0.8149
- Ran 10 systematic trials varying LoRA rank, learning rate, and DPO beta
- Both LoRA adapters released on HuggingFace
What you will get:
- Custom fine-tuning on your dataset (SFT and/or DPO)
- LoRA adapters (lightweight, easy to deploy)
- Evaluation metrics (BLEU, BERTScore, loss curves)
- Clean, documented Python code
- HuggingFace Model Card (Premium)
- Deployed model on HuggingFace Hub (Premium)
Models I can fine-tune:
- Qwen (0.6B, 1.8B, 7B)
- LLaMA (1B, 3B, 8B)
- Mistral (7B)
- Phi-3 / Phi-2
- Gemma (2B, 7B)
- Any open-source model on HuggingFace
What I need from you:
- Your dataset (JSONL, CSV, or plain text)
- Task description (instruction-following, summarization,
classification, Q&A, etc.)
- Preferred base model (or I recommend one for you)
Message me before ordering to di
Programming Language:
Python
Data Type:
Text
•
Tabular Data
AI Engine:
GPT
•
Llama
•
Falcon
•
Langchain
•
Other
My Portfolio
FAQ
What format does my dataset need to be in?
Ideally JSONL with instruction/input/output fields, or CSV. I can help you reformat raw data into the right structure if needed - just share a sample first.
What if my dataset is small (under 500 samples)?
Small datasets are fine for LoRA fine-tuning. Message me with your sample count and I will advise on the best approach.
Will I own the fine-tuned model?
Yes - you get full ownership of the LoRA adapters and all code. Premium package includes deployment to your own HuggingFace account.
Can you fine-tune GPT-4 or Claude?
No - those are closed-source models and cannot be fine-tuned this way. I work with open-source models on HuggingFace (LLaMA, Qwen, Mistral, Gemma, Phi etc.)
What evaluation metrics do you provide?
BLEU score, BERTScore, training/validation loss curves, and sample output comparisons against baseline. Premium package includes a full ablation study.
