I will audit and stress test your llm benchmark or ai evaluation
AI Evaluation and Cybersecurity QA Specialist
About this Gig
Need confidence that your LLM evaluation, AI-agent benchmark, or coding task measures what you intend? I will independently audit the task specification, scoring logic, verifier, test coverage, and delivery package for loopholes, false positives, brittle assumptions, reproducibility failures, and weak adversarial coverage.
You will receive:
- A prioritized findings report
- Concrete exploit or failure cases
- Reproduction steps
- Risk and impact ratings
- Practical repair recommendations
Standard and Premium include deeper adversarial testing. Premium can include tested fixes and clean packaging when the scope permits.
I work only with client-owned, public, or explicitly authorized material. Results are not guaranteed because a sound evaluation may contain no defects; you are buying the agreed audit depth and evidence-backed report.
Testing application:
Software
Device:
PC
•
Mac
•
Linux
FAQ
What do you need to start?
Please send the authorized task specification, relevant files or ZIP, verifier and tests, expected behavior, known constraints, and your deadline.
Will you guarantee that you find defects?
No. I guarantee the agreed audit depth and an evidence-backed report, not that defects exist. A sound evaluation may pass the audit.
Can you implement fixes?
Yes. Standard includes repair guidance. Premium can include tested fixes and clean packaging when the scope and access permit.
What materials can you audit?
I work only with client-owned, public, or explicitly authorized materials. I do not bypass access controls or test systems without authorization.
