I will benchmark and optimize your ai prompts with measurable results


About this gig
I will evaluate, benchmark, and optimize your AI prompts or LLM workflows to improve reliability, consistency, and output quality.
I will review your existing prompt, test it against example inputs, identify weaknesses or failure patterns, and provide optimized prompt improvements based on structured evaluation and benchmarking.
Depending on the package selected, services may include:
- Prompt analysis and optimization
- Benchmark testing using datasets
- Tone and instruction-following evaluation
- Consistency and reliability analysis
- JSON/schema validation checks
- Before vs after comparisons
- Optimization recommendations
- Summary or detailed professional reports
My process uses a proprietary evaluation framework designed to measure improvements with evidence-based results rather than trial-and-error prompt editing.
This service is suitable for:
- ChatGPT prompts
- Claude and Gemini workflows
- AI assistants
- Customer support bots
- AI automation workflows
- Structured JSON output prompts
- Production LLM systems
The goal is to help your AI systems perform more reliably across real-world inputs and edge cases.
Note, test cases are limited to 5 for Basic, 20 for standard and 50 for premium.
Get to know Ray Gill
Data Driven AI Evaluation and Optimization
- FromUnited Kingdom
- Member sinceMay 2026
Languages
English
FAQ
What kinds of issues can you help improve?
Examples include: Inconsistent responses Hallucinations Poor instruction following Weak reasoning Tone/style problems Invalid JSON formatting Excessive verbosity Reliability issues across different inputs
Do you provide reports?
Yes. Depending on the package selected, reports can include: Benchmark findings Before vs after comparisons Optimization recommendations Evaluation summaries Reliability analysis
Can you test prompts against datasets or multiple scenarios?
Yes. Higher-tier packages can include structured benchmark testing across multiple test cases and edge cases.
Do you guarantee perfect AI outputs?
No AI system can be guaranteed to perform perfectly in all situations. The goal is to significantly improve reliability, consistency, and overall output quality using structured evaluation and optimization methods.
Which AI models do you support?
I can work with prompts and workflows for: OpenAI ChatGPT Anthropic Claude Google Gemini
What do you need from me?
Please provide: Your current prompt or workflow The goal of the AI system Example inputs and outputs (if available) Any formatting, tone, or business requirements Details of any issues you are experiencing

