I will audit your ai benchmark or llm eval numbers for statistical validity
About this Gig
You'll get a clear verdict on whether your benchmark or eval number survives the checks it skipped before you
publish it. I take one headline result (a leaderboard rank, an LLM-as-judge win-rate, a "beats baseline by X%," a
scaling-law exponent) and test it adversarially: multiple-comparison correction, statistical power, judge bias, and
reproduction from your raw data.
You get one of two honest answers: it survives here's the corrected statistic you can cite or it's within noise,
here's why and the one-line fix.
What sets this apart is proof, not promises. I wrote the book on this failure mode (Measured, Not Believed), built the
open-source tool that runs the checks (evalgate), and publicly reproduced three real overclaims MT-Bench,
RewardBench, and a published scaling law. Every verdict I give is reproducible from your data; I show the work. And if
your number is real, I certify it an auditor that only ever finds fault is a cynic, not an auditor.
FAQ
What if my number turns out to be fine?
Then I certify it and you get a citable clean bill of health. That's a valid, common outcome — not a failed engagement. An auditor that only ever finds fault is a cynic; if your result holds, I show you it holds.
Is my data confidential?
Yes. Reports are private and nothing is published. A redacted or anonymized slice of your data is usually enough to reproduce the number, so you rarely need to share anything sensitive.
Do you just run a tool on it?
No. The individual checks are textbook; the value is running all of them adversarially and having the judgment to tell a real confound from an artifact of how the data was built. You get reasoning, not just output.
Which package do I need?
If you have one number you're about to publish, start with Basic or Standard. If you're auditing a launch post, model card, or several claims at once, choose Premium. Not sure? Message me the one number you're least sure about.
Can you wire these checks into our CI?
Yes — that's an ongoing monthly retainer rather than a one-off Catalog project. Message me with your stack and eval cadence and I'll scope it separately.

