I will evaluate and benchmark your llm system and ai outputs

S
samanazharr
S
samanazharr
Saman A.

About this gig

As a Machine Learning Engineer specializing in NLP and symbolic AI, I provide comprehensive evaluation, benchmarking, and red-teaming for your Large Language Models (LLMs) and RAG pipelines. I help engineering teams identify edge-case failures, eliminate hallucinations, and enforce strict security guardrails.


What I Offer:

  • LLM Benchmark & Performance Auditing: Evaluating model accuracy, latency, context usage, and reasoning capabilities against customized testing rubrics.
  • Hallucination & Edge-Case Detection: Stress-testing system prompts, RAG retrieval accuracy, and factual consistency under adversarial conditions.
  • Red Teaming & Security Testing: Identifying vulnerabilities like prompt injection attacks, jailbreaks, and system prompt leaks.
  • Schema & Output Validation: Verifying strict JSON/Pydantic structure compliance, function calling execution, and API integration reliability.
  • Detailed Audit Report & Action Plan: Comprehensive failure analysis complete with actionable prompt engineering fixes and guardrail recommendations

Get to know Saman A.

Saman A.

Data Annotation and LLM QA Expert

  • FromPakistan
  • Member sinceFeb 2026
  • Avg. response time1 hour
  • Languages

    English, Urdu
I am a Machine Learning Engineer specializing in NLP, with hands-on experience building, evaluating, and deploying production-grade LLM systems. My background spans the full pipeline: from dataset curation and prompt engineering to security auditing, hallucination mitigation, and containerized deployment. Recently, my work focuses on LLM Evaluation & Alignment: refining model outputs through rigorous fact-checking, benchmark testing, prompt injection defense, and human-in-the-loop quality control.

My Portfolio