I will audit your rag system and reduce hallucinations

R
royxforge
R
royxforge
Sourav Roy

About this gig

I will audit your RAG system and show you exactly where it fails with numbers.


I'm an AI engineer with 3 research papers and production RAG systems in healthcare, education, and sales. I built an open-source evaluation framework that scores 6 dimensions: faithfulness, hallucination rate, retrieval precision, answer relevance, context coverage, and a label-free confidence metric (UCM) that needs zero ground-truth labels.


WHAT YOU GET:

- Evaluation run on your real queries (OpenAI, Anthropic, Llama, or any provider)

- Hallucination rate + faithfulness scores, ranked by severity

- Root causes: chunking, retrieval, prompts, reranking

- Prioritized fix roadmap


Standard adds a full 6-dimension benchmark report. Premium implements the fixes and re-runs evaluation to show before/after numbers.


Works with LangChain, LlamaIndex, FastAPI, vector DBs (Pinecone, Chroma, FAISS, Weaviate).


Send your repo/docs and I'll start within 24h.


Get to know Sourav Roy

Sourav Roy

AIML Engineer

  • FromIndia
  • Member sinceJun 2026
  • Avg. response time1 hour
  • Languages

    Bengali, Hindi, English, Italian
I am an AI/ML Engineer with three published research papers and experience building production systems in healthcare, education, and sales. I architect end-to-end pipelines involving LLM orchestration, RAG, and neural model calibration. My focus is on creating robust, scalable AI infrastructure that bridges the gap between applied research and real-world reliability.