I will setup deepseek harness and local ai agent with llama cpp


About this gig
Turn your GPU into a private AI agent no per-token fees.
>
DeepSeek Harness + llama.cpp is the most powerful open-source local AI stack. Harness gives you an agent framework with plugin support; llama.cpp gives you inference that's 2-5x faster than Ollama.
>
I'll set it all up for you in one session.
>
What you get:
- DeepSeek Harness (dsh) installed and configured
- llama.cpp compiled from source with GPU acceleration
- llama-server running as a production REST API
- The right model sized for your VRAM
>
Stack: DeepSeek Harness + llama-server + GGUF models (Qwen, Gemma, DeepSeek, Llama, Mistral...)
>
What I need: Your OS, GPU specs, and which model you want.
Get to know Lucas L.
Remote IT Support Windows Printers Virtual Machines Local AI Setup
- FromArgentina
- Member sinceNov 2025
- Avg. response time1 hour
- Last delivery1 month
Languages
English, Spanish
My Portfolio
FAQ
Will this work on my hardware?
Share your GPU (NVIDIA/AMD/Apple Silicon) or CPU-only setup before ordering. I'll confirm
Which model should I run?
I'll recommend one based on your VRAM. Larger models = better quality, more memory needed
Ollama is easier. Why llama.cpp?
llama.cpp is 2-5x faster and gives you full control. Harness works with both, but the llama.cpp backend is the performance winner.

