I will deploy your llm for production with vllm on your cloud GPU


About this gig
A model that works in a notebook and one that survives 1,000 concurrent users are two different projects. I do the second one.
I'm a senior ML-infrastructure engineer, 7 years of production reliability underneath. Most recently: serving infrastructure handling 50,000+ concurrent tasks at sub-100ms for 1,000+ users, built on vLLM and SGLang with GPU Kubernetes node pools, at 35% lower infra cost than when I started.
What you get:
- Your open-weight model (Llama, Mistral, Qwen, DeepSeek) served with vLLM
- Deployed in YOUR cloud account: you keep the keys, data, and endpoint
- GPU autoscaling and batching tuned for inference, not web traffic
- Load-test results, so you know real capacity before your users find it
- Monitoring dashboards; runbook on Standard/Premium
If your GPU budget and latency target don't match, I'll say so before you spend a cent on compute.
Basic: one model, one GPU node, load-tested.
Standard: adds autoscaling, tuning, dashboards.
Premium: multi-model, cost pass, handover.
Send your model and traffic details; the requirements form covers the rest.
Get to know Joshua
Senior Platform Engineer
- FromNigeria
- Member sinceJul 2026
Languages
English
My Portfolio
FAQ
Which models can you deploy?
Open-weight models: Llama, Mistral, Qwen, DeepSeek, Gemma, and most anything on Hugging Face that vLLM supports. If you mean OpenAI/Anthropic APIs, you don't need serving infrastructure, and I'll say so rather than sell you this gig.
Do I need my own cloud account and GPUs?
Yes, everything deploys in your account (AWS, Azure, GCP, or on-prem), so you keep full control. If you don't have GPU quota yet, I'll tell you exactly what to request and which instance types fit your model and budget.
What latency and throughput should I expect?
It depends on model size, quantization, and GPU choice, which is why every delivery includes load-test numbers for YOUR setup rather than generic promises. Sub-100ms first-token latency is achievable for many mid-size models on the right hardware.
Is my data and model private?
Everything runs inside your account and your network. I work through scoped access you grant and revoke, and nothing is copied out. NDAs are fine.
What counts as a revision?
Tuning what was scoped: batch sizes, autoscaling thresholds, endpoint config. A different model, a second environment, or new tooling is new scope, quoted as an add-on.

