I will deploy, finetune, and optimize local llms and vlms
About this gig
High-Performance AI, Optimized for Your Infrastructure
Looking to run powerful open-source models without breaking the bank on cloud costs? I specialize in making Large Language Models (LLMs) and Vision-Language Models (VLMs) faster, lighter, and smarter.
️ What I Do:
- Custom Fine-Tuning: Tailoring models (Llama, Mistral, Qwen, Phi) to your specific dataset using advanced techniques like LoRA/QLoRA (PEFT).
- Quantization: Reducing model size (GGUF, AWQ, EXL2) to run efficiently on consumer hardware or smaller cloud instances without sacrificing accuracy.
- Inference Optimization: Setting up ultra-fast inference frameworks (vLLM, TensorRT-LLM, Ollama) for maximum throughput.
Why Choose Me?
- End-to-End Expertise: From raw dataset preparation to production-ready deployment.
- Cost-Conscious: Focused on reducing your production VRAM footprint and compute costs.
- Clean, Documented Code: Full delivery with deployment scripts.
Get to know Akash Bhansali
Whatever it Takes
- FromIndia
- Member sinceJan 2021
- Avg. response time5 hours
- Last delivery2 years
Languages
English, Hindi, German
FAQ
Which models do you work with?
I work with all major open-source LLMs and VLMs, including Llama 3, Mistral, Qwen, Phi-3, and Llava, across frameworks like Hugging Face, PyTorch, and DeepSpeed.
Do I need to provide the dataset for fine-tuning?
Yes, ideally a clean dataset in JSON/CSV format. However, if you need help with dataset formatting or preprocessing, we can include that in the project scope.
What is quantization and does it reduce performance?
Quantization compresses model weights from high precision to lower precision (e.g., 32‑bit → 8‑bit), cutting memory usage. When implemented carefully, the performance impact is minimal and inference speed often improves.
Do you offer parameter‑efficient fine‑tuning (LoRA/QLoRA)?
Yes! PEFT methods freeze most parameters and tune only small adapters, reducing compute requirements while delivering strong results.
How do you handle data privacy?
Your data remains confidential. I’ll sign an NDA if needed and only use your datasets for training and evaluation.
Who owns the fine‑tuned model?
You retain full rights to the fine‑tuned weights, adapters and quantized models. I do not resell or reuse your custom model.

