I will test and evaluate your rag or llm application

F
fafnir_kyloth
F
fafnir_kyloth
Antonio

About this gig

Is your RAG or LLM application producing unsupported, irrelevant or inconsistent answers? I will evaluate its behavior and provide evidence-based recommendations for improvement.

I can assess:

Answer relevance, clarity and completeness

Grounding in retrieved context

Hallucinations and unsupported claims

Retrieval quality and source usefulness

Citation consistency

Failure patterns and fallback behavior

Latency and cost indicators when data is available

Chunking, context and model configuration decisions

Depending on the package, I will review supplied test cases, create a structured evaluation set, analyze retrieval and generation failures, and use automated metrics when they are technically appropriate. You will receive a clear report separating observations, evidence, limitations and prioritized recommendations.

Please provide access to a safe test environment or exported application results, retrieved contexts, expected answers when available, and a description of the intended users and use case. Remove API keys, production credentials and private customer data.

This Gig evaluates an existing system. Implementing major changes, rebuilding the RAG pipeline, production deployment a

Get to know Antonio

Antonio

Python Backend Applied AI Developer

  • FromBrazil
  • Member sinceSep 2023
  • Avg. response time1 hour
  • Languages

    English, Portuguese, Spanish
Python backend and applied AI developer focused on practical, maintainable solutions. I build and troubleshoot FastAPI and Flask APIs, RAG pipelines, local LLM applications, chatbots and Python automations. My portfolio includes a medical RAG assistant, an AI helpdesk triage API and a local coding assistant with OAuth. I value clear scope, readable code, automated tests, documentation and honest communication.

My Portfolio

Other AI Development Services I Offer

Related tags