I will build, integrate or debug local llm and vlm systems


About this gig
I build, integrate, troubleshoot and optimize local AI systems based on LLMs and VLMs.
I work directly with local model runtimes and infrastructure, including llama.cpp/ggml, GGUF, Qwen, Llama, Mistral, DeepSeek, Hugging Face, Ollama, CUDA, Python and C++.
I can help with:
- Local LLM/VLM installation and deployment
- llama.cpp and GGUF configuration or debugging
- GPU/CPU inference and CUDA optimization
- Python integration and custom AI workflows
- LoRA/QLoRA training and model adaptation
- OCR, document AI, audio and data-processing pipelines
- Internal/private AI tools and automation
- Diagnostics, telemetry, testing and performance issues
- Database and application integration
One implementation means one clearly defined AI component, integration or deployment task. More complex systems can be divided into multiple implementations.
Because local AI projects vary greatly in hardware, models and scope, please contact me before ordering if your project is complex or requires custom development.
Get to know Daniel G
Python AIML Developer
- FromPoland
- Member sinceAug 2026
- Avg. response time2 hours
Languages
English, Polish
My Portfolio
FAQ
Should I contact you before placing an order?
Yes, especially for custom or complex projects. Local AI systems vary significantly depending on the model, hardware, operating system and required integrations. Send me a short description of your project first so I can confirm the appropriate package and scope.
Which local AI models and runtimes do you work with?
I work with local LLM and VLM systems including Qwen, Llama, Mistral, DeepSeek, GGUF models, llama.cpp/ggml, Hugging Face and Ollama, as well as custom Python and C++ integrations.
Can you troubleshoot an existing local LLM setup?
Yes. I can diagnose problems involving llama.cpp, GGUF models, CUDA/GPU offloading, CPU/GPU inference, model configuration, Python integrations, performance issues and local AI pipelines.
Can you help me choose the right model for my hardware?
Yes. I can evaluate your available CPU, GPU, VRAM, RAM and use case, then recommend an appropriate model size, quantization and deployment configuration.
Do you provide LoRA or QLoRA training and model adaptation?
Yes, when the project scope and available hardware make it appropriate. Training requirements depend heavily on the model, dataset, GPU memory and expected result, so please contact me before ordering training work.
Can the solution run completely locally without sending data to external AI APIs?
Yes. I specialize in local AI systems and can design solutions where models and data remain on your own hardware when the selected model and infrastructure support this.

