I will tune pinecone qdrant or opensearch for rag

S
spoidey
S
spoidey
ali shah

About this gig

Bad answers in a RAG app are usually a retrieval problem, not a model problem. I configure, tune and debug your vector database Pinecone, Qdrant or OpenSearch so the right chunks reach the prompt.


WHAT I DO

- Schema & ingestion: embedding dimension, similarity metric, HNSW parameters, clean upsert/delete pipelines in Python or TypeScript

- SDK wiring: LangChain, LlamaIndex or raw client, with retries and idempotent indexing

- Retrieval quality: metadata filtering, multi-tenant namespace isolation, top-k and score thresholds, hybrid BM25 + vector, reranking

- Pinecone: serverless vs pod, namespaces per tenant, index cost tuning

- Qdrant: cloud or self-hosted Docker, payload schema, quantization, snapshot strategy

- OpenSearch: k-NN on your existing cluster without losing keyword search

- Debug & migrate: empty or irrelevant results, latency, runaway spend, ChromaPinecone/Qdrant migrations verified with query suites


If chunking is the real cause, I show you the evidence instead of invoicing you for index tuning. No chat UI, no prompt engineering, no fine-tuning.


Before ordering, send: DB, embedding model, vector count and one failing query.

Get to know ali shah

ali shah

AI LLM Engineer RAG, Agents, LLM Ops Shopify Hydrogen

  • FromPakistan
  • Member sinceAug 2026
  • Languages

    Urdu, English
5 years shipping production software (Systems Limited, then Contour Software) as a backend + AI engineer. I build the unglamorous parts that make AI products work: RAG retrieval that actually returns the right chunk, evals that catch hallucinations before customers do, LiteLLM gateways and Langfuse traces so you know your per-user cost, and vLLM on your own GPU. Also custom Shopify Hydrogen storefronts. .

My Portfolio

Other AI Development Services I Offer