I will audit your ai benchmark or llm eval numbers for statistical validity

Finland

I speak English, Finnish

Trading bot expert

I am a Quantitative Engineer building custom trading bots and API pipelines. I don't sell "get-rich-quick" scripts. I build solid, reliable trading architecture. My Python code follows the strict V2....
About this Gig

 You'll get a clear verdict on whether your benchmark or eval number survives the checks it skipped before you

 publish it. I take one headline result (a leaderboard rank, an LLM-as-judge win-rate, a "beats baseline by X%," a

 scaling-law exponent) and test it adversarially: multiple-comparison correction, statistical power, judge bias, and

 reproduction from your raw data.


 You get one of two honest answers: it survives here's the corrected statistic you can cite or it's within noise,

 here's why and the one-line fix.


 What sets this apart is proof, not promises. I wrote the book on this failure mode (Measured, Not Believed), built the

 open-source tool that runs the checks (evalgate), and publicly reproduced three real overclaims MT-Bench,

 RewardBench, and a published scaling law. Every verdict I give is reproducible from your data; I show the work. And if

 your number is real, I certify it an auditor that only ever finds fault is a cynic, not an auditor.