I will set up colibri to run kimi glm deepseek and large ai models locally

T
techyesa
T
techyesa
Khan

About this gig

Want to run a very large AI model on your own computer without sending every prompt to a cloud API?


I will install and configure Colibri, the open-source inference engine that streams Mixture-of-Experts model weights from SSD. On compatible hardware, it can run kimi, Deepseek v4 flash and GLM-5.2 744B locally.


Depending on your package, I can:

  • check hardware and storage compatibility
  • install or build Colibri on Linux, Windows, or macOS
  • configure one supported model
  • tune CPU, RAM, GPU, cache, and NVMe placement
  • enable CLI chat, the web dashboard, or a local API
  • run diagnostics and a performance benchmark


Important: the recommended GLM-5.2 int4 model needs about 372 GB of free storage. A GPU is optional, but a fast NVMe SSD matters. Slower systems may generate below one token per second, so I will not promise a speed before checking your hardware.


Please message me before ordering with your OS, CPU, RAM, GPU, disk type, free space, and internet speed. I will confirm compatibility and the right package.

Get to know Khan

Khan

Building Your Future: AI Websites, Automation, Chatbots and SaaS

5.0(3)
  • FromPakistan
  • Member sinceJun 2025
  • Avg. response time14 hours
  • Last delivery8 months
  • Languages

    English, Urdu
I'm an AI researcher and engineer who turns ideas and research papers into working AI applications. I build n8n workflow automations, custom AI agents, RAG and knowledge graph systems, and API integrations. I also work with PyTorch, LangGraph, semantic search, GGUF model quantization, and Streamlit MVPs. Send me your project requirements, and I'll help you choose and build the right solution.

My Portfolio