I will audit and stress test your llm benchmark or ai evaluation

India

I speak English

AI Evaluation and Cybersecurity QA Specialist

I audit LLM evaluations, AI-agent benchmarks, and coding tasks for verifier loopholes, unreliable scoring, false positives, weak adversarial coverage, reproducibility issues, and packaging defects. Cl...
About this Gig

Need confidence that your LLM evaluation, AI-agent benchmark, or coding task measures what you intend? I will independently audit the task specification, scoring logic, verifier, test coverage, and delivery package for loopholes, false positives, brittle assumptions, reproducibility failures, and weak adversarial coverage.


You will receive:

- A prioritized findings report

- Concrete exploit or failure cases

- Reproduction steps

- Risk and impact ratings

- Practical repair recommendations


Standard and Premium include deeper adversarial testing. Premium can include tested fixes and clean packaging when the scope permits.


I work only with client-owned, public, or explicitly authorized material. Results are not guaranteed because a sound evaluation may contain no defects; you are buying the agreed audit depth and evidence-backed report.

Testing application:

Software

Device:

PC

Mac

Linux