I will improve your llm evaluation and ai reliability

K
katri_t
K
katri_t
Ekaterina T

About this gig

If you're deploying LLMs in production and don't have a real evaluation pipeline, you're flying blind on reliability. I design evaluation systems: benchmarking, judge calibration, hallucination analysis, and testing pipelines that catch failures before your users do.


This is the layer most AI products skip and the one that determines whether your system is trustworthy at scale.


Background: independent AI measurement research lab, published research on LLM-as-judge evaluation dynamics and failure modes.

Get to know Ekaterina T

Ekaterina T

AI Assistant and Chatbot Developer, Customer Support Automation, LLM Systems

  • FromRussia
  • Member sinceOct 2024
  • Avg. response time1 hour
  • Languages

    English, Russian, Serbian, French, German, Finnish
AI Systems Builder and Machine Learning Engineer specializing in LLM applications, RAG architectures, and operational intelligence. I translate operational bottlenecks into practical AI-powered solutions, bridging business needs and technology. https://www.linkedin.com/in/ekaterina-taratuta/ ETSystemsAI https://www.linkedin.com/company/111727786 LLM Measurement Lab https://www.linkedin.com/company/112245236 Google Scholar: https://scholar.google.com/citations?hl=ru&user=SplaIk8AAAAJ

My Portfolio