I will build mixture of experts, cot reasoning, quantization llms


About this gig
Most people building "LLM projects" are calling OpenAI's API and wrapping it in a chatbot UI. That's not LLM engineering that's plumbing.
I work under the hood: architecture, reasoning, and optimization. If you need something beyond a prompt wrapper, this is the gig.
What I actually build:
Mixture of Experts (MoE) sparse expert routing, gating networks, load-balancing losses, custom MoE layers built from scratch or fine-tuned on top of existing architectures
Chain-of-Thought Reasoning CoT/ToT prompting pipelines, self-consistency decoding, reasoning-distillation into smaller models, step-verification and reward-guided reasoning
Quantization INT8/INT4 quantization, GPTQ, AWQ, QLoRA, GGUF conversion for local/edge deployment, benchmarking accuracy-vs-speed tradeoffs before you commit to a setup
Also covered: fine-tuning (LoRA/QLoRA), model compression, inference optimization, and deployment on constrained hardware.
Why this matters:
A model that reasons well but can't run on your hardware is useless. A quantized model that's fast but dumb is worse. I build for the actual constraint you're solving cost, latency, accuracy, or all three not just whatever's trending on Twitter.
Get to know Awais J
Applied AI Engineer
- FromPakistan
- Member sinceAug 2026
- Avg. response time1 hour
Languages
Urdu, English, French, German, Arabic, Chinese

