I will fine tune llama, qwen or mistral llms on your data with lora

S
sidharth1
S
sidharth1
Sidharth P

About this gig

Need an open-source LLM that really knows your domain, format or style? I will fine tune it on your data and show you, with numbers, how much better it got.


I am an AI researcher (ICML 2026, IJCNLP-AACL 2025) and I fine tune and evaluate open models on multi-GPU H100 setups as part of my research.


What you get:

  • Base model and method recommendation (LoRA, QLoRA, full SFT or GRPO)
  • Data cleaning and formatting into chat or instruction templates
  • Clean, reproducible training code and config
  • Trained LoRA adapter, or merged weights on higher packages
  • Before and after evaluation on a held-out set
  • Optional fast serving setup with vLLM or SGLang


Models: Llama, Qwen, Mistral, Gemma, DeepSeek and other Hugging Face models.


Message me your dataset and goal before ordering and I will suggest the right base model, method and package.

Get to know Sidharth P

Sidharth P

AI Safety Researcher and LLM Fine Tuning Expert

  • FromIndia
  • Member sinceApr 2017
  • Languages

    Telugu, English, Hindi
AI researcher in LLM safety and NLP. Research Intern at AI4Bharat (first author of IndicBERT-v3, open multilingual encoder LLMs) and Research Fellow at SPAR. Papers at ICML 2026 and IJCNLP-AACL 2025, plus arXiv work on memory attacks against LLM agents. Hands-on with LoRA/QLoRA and full fine-tuning (SFT, GRPO), multi-GPU training on H100s, vLLM/SGLang inference and LLM-as-judge evaluation. I help teams fine-tune open-source LLMs, build evaluation pipelines, red-team LLM agents and reproduce ML papers. Message me your goal and I will reply with a clear plan.

My Portfolio

Other AI Development Services I Offer