I will quantize your ai model


About this gig
Are your cloud GPU bills getting out of hand, or is your fine-tuned model too heavy to run on edge hardware?
I am a machine learning systems engineer specializing in low-bit quantization, inference optimization, and edge deployment. I take large language, vision, and mixture-of-experts (MoE) models and shrink their memory footprint by 50% to 80% without destroying accuracy.
Whether you need a 70B model squeezed onto consumer GPUs, a 7B running on an Apple Silicon Mac/iPhone, or a calibrated GGUF file for local corporate compliance, I will build the exact runtime pipeline you need.
What I Specialize In:
* Formats & Runtimes: GGUF (llama.cpp), MLX (Apple Silicon Mac/iOS), EXL2, AWQ, GPTQ, bitsandbytes (NF4/FP8).
* Model Architectures: Llama 3, Qwen 2.5 / 3.8, Mistral, Gemma, MoE models (OLMoE, Mixtral), and Vision-Language Models (VLMs).
* Target Hardware: Single NVIDIA GPUs (L4, A10, RTX 4090), Apple Silicon (M-series unified memory, iOS), and CPU inference.
* Precision & Quality:Advanced calibration (imatrix), activation-aware quantization, and deterministic quality verification (Perplexity & downstream evals).
Get to know Kaleb C
Senior Data AI Leader
- FromUnited States
- Member sinceMay 2024
Languages
English

