I will build an ai agent evaluation harness with regression tests

J
janak_g808
J
janak_g808
Janak G

About this gig

Your agent passed the demo. Then a prompt changed, a model version rolled forward, and nobody found out until a user did.


I build the evaluation harness that catches it first.


WHAT YOU GET

- A failure taxonomy written for YOUR agent, not a generic checklist

- A runnable eval harness with named metrics: task completion, tool-call correctness, argument correctness, faithfulness, latency, cost

- Regression tests that fail on the exact behaviours you care about

- A technical report you can forward to your CTO

- Everything in your repo, your framework, your ownership


WHO THIS IS FOR

Teams already running an LLM agent who cannot answer: "did that change make it worse?"


HOW IT WORKS

1. You send your agent's repo or API, plus 3 failures that annoy you

2. I reproduce them and map the failure surface

3. I build and run the harness, then hand it over with a walkthrough


I do not build agents. I make existing ones testable.


Tell me your stack before you order and I will confirm fit, or tell you honestly if you do not need this yet.


Get to know Janak G

Janak G

AI Automation Engineer

  • FromIndia
  • Member sinceJul 2026
  • Avg. response time1 hour
  • Languages

    English
I build AI systems that answer your customers while you sleep. Chatbots trained on your own documents, so every answer comes from your material and links back to the page it came from. Automations that clear repetitive work. AI agents that qualify leads. Trained in Data Science and AI. I ship production systems, not demos - every order includes docs, a video walkthrough, and 14 days of free fixes. I reply within one hour during working hours. Delivery 2-4 days. No fluff, no ghosting. Send me the task stealing your time. Free 15-min consult - I'll tell you honestly if AI can solve it.

My Portfolio

Other AI Development Services I Offer