I will fix, debug, and optimize your rag pipeline or ai chatbot

C
chaudhry_az33m
C
chaudhry_az33m
Muhammad Azeem

About this gig

Your RAG pipeline or LLM chatbot works in the demo but fails with real users. Wrong chunks retrieved, hallucinated answers, slow responses, rising API costs. I fix that.

I am an AI engineer maintaining a production RAG system over thousands of documents, live on Google Cloud and serving real users daily. I have dealt with every failure mode this gig covers.

What I fix:

  • Poor retrieval: chunking strategy, embeddings, hybrid search, reranking
  • Hallucinations: grounding, citations, prompt structure
  • Latency and cost: caching, model choice, batching
  • Broken ingestion: PDFs, scraped data, messy documents
  • No visibility: evaluation setup so quality is measurable

Stacks: LangChain, LlamaIndex, custom pipelines, OpenAI, Gemini, Claude, Pinecone, Qdrant, pgvector, Vertex AI Vector Search, FastAPI, GCP, AWS.

How it works: you share your repo or describe your setup, I audit it, and you get a written diagnosis with prioritized fixes. On Standard and Premium I implement the fixes and prove the improvement with before and after tests.

You own all code. NDA on request. Message me before ordering so I can confirm scope.

Get to know Muhammad Azeem

Muhammad Azeem
  • FromPakistan
  • Member sinceJun 2025
  • Avg. response time1 hour
  • Languages

    English
Hi, I'm Azeem, a Python & AI Engineer specializing in backend development, intelligent automation, and production-ready AI systems. I build scalable APIs with FastAPI, agentic AI workflows, RAG applications, large-scale web scrapers, document processing pipelines, vector search solutions, and cloud-native applications on Google Cloud. My expertise includes LLM integration, workflow automation, REST APIs, data engineering, and scalable backend architecture. I focus on writing clean, maintainable code, clear communication, and delivering reliable software that creates long-term value.

My Portfolio