I will set up colibri to run kimi glm deepseek and large ai models locally


About this gig
Want to run a very large AI model on your own computer without sending every prompt to a cloud API?
I will install and configure Colibri, the open-source inference engine that streams Mixture-of-Experts model weights from SSD. On compatible hardware, it can run kimi, Deepseek v4 flash and GLM-5.2 744B locally.
Depending on your package, I can:
- check hardware and storage compatibility
- install or build Colibri on Linux, Windows, or macOS
- configure one supported model
- tune CPU, RAM, GPU, cache, and NVMe placement
- enable CLI chat, the web dashboard, or a local API
- run diagnostics and a performance benchmark
Important: the recommended GLM-5.2 int4 model needs about 372 GB of free storage. A GPU is optional, but a fast NVMe SSD matters. Slower systems may generate below one token per second, so I will not promise a speed before checking your hardware.
Please message me before ordering with your OS, CPU, RAM, GPU, disk type, free space, and internet speed. I will confirm compatibility and the right package.
Get to know Khan
Building Your Future: AI Websites, Automation, Chatbots and SaaS
- FromPakistan
- Member sinceJun 2025
- Avg. response time14 hours
- Last delivery8 months
Languages
English, Urdu
My Portfolio
FAQ
Can Colibri run GLM-5.2 744B without a GPU?
Yes. Colibri supports CPU-only inference. A GPU can help on supported hardware, but the model can run without one. Storage speed, RAM, CPU, and configuration all affect performance.
What hardware do I need for GLM-5.2?
The official quick start lists about 16 GB RAM as the minimum, 24 GB or more as recommended, and about 380 GB of free disk space for the int4 model. A fast NVMe SSD is strongly recommended.
Will it run fast on my computer?
I cannot promise a speed before checking your system. Colibri trades memory use for storage traffic. A slow or shared disk can produce well below one token per second. I will benchmark the completed setup and report the measured result.
Which operating systems do you support?
Linux, Windows 10 or 11, and macOS are supported. Available acceleration options depend on the buyer's hardware and drivers.
Is the model download included?
Standard and Premium include downloading and configuring one supported model on your machine. You must provide enough storage and internet bandwidth and accept the applicable model license. I do not resell model weights.
Can you connect Colibri to my app or workflow?
Premium includes OpenAI-compatible and Anthropic Messages local APIs. Connecting it to a specific app, n8n workflow, or custom interface may require a custom offer depending on the scope.
Can you connect it to Claude Code?
Yes. Colibri 1.4.0 provides an Anthropic Messages endpoint that can be used by Claude Code. Large coding-agent prompts can take a long time to prefill on a CPU and SSD setup, so I will test it and explain the measured performance instead of promising cloud-like speed.
Can you set up a model other than GLM-5.2?
Colibri currently supports GLM-5.2, Inkling, Kimi K3, and OLMoE. Their memory and storage needs differ. Send your hardware details and preferred model before ordering.
Do I need to share passwords?
Do not send personal passwords in Fiverr chat. We will agree on a safe setup method that stays within Fiverr's rules. Replace any temporary credentials after delivery.
