I will evaluate and red team your llm or ai product outputs


About this gig
I evaluate and red-team AI/LLM outputs for accuracy, safety, and failure modes, using rubric-based scoring, not vibes.
Background: I do this professionally for multiple AI research organizations (under NDA), assessing outputs for instruction-following, factual accuracy, reasoning, and safety, and comparing competing model responses to surface hallucinations and edge-case failures.
Evidence, not adjectives: EndpointDrift, one of my public builds, is a clinical-trial auditor validated against a 200-trial frozen corpus with 358 gold labels and 101 tests. OSINT Fusion runs an adversarial verification swarm across 1,400+ tests. Write-ups at thisistravissmith.com.
What you get: a rubric matched to your product's actual failure modes, a scored pass over your outputs, a written report of specific failures (hallucinations, unsafe completions, reasoning breaks, instruction misses) with examples, and a plain-language severity ranking.
Good fits: pre-launch AI features, chatbot/agent QA, comparing model or prompt versions, guardrail spot-checks.
Send me sample outputs (or sandbox access) and what "good" looks like, and I'll scope the pass.
Get to know Travis Smith
Agentic AI Engineer for LLM Evaluation, RAG, and Custom AI Builds
- FromUnited States
- Member sinceJul 2026
- Avg. response time1 hour
Languages
English

