Vahit Feryad / evaluation science
$ applied research · AI evaluation & reliability

Measuring whether AI measurements can be trusted

AI models are ranked by other AI models and by small panels of people. I study when those scores hold up and when they don't: judge bias, rater reliability and the statistical resolution of leaderboards. Every result is open and reproducible.

PhDsignal processing & ML
10+ yrsresearch & industry
230+Scholar citations
Fig. 1  judge swap test6 models · 200 items
#1#2#3#4#5#6 true winner
rank of the true #1 model#1
Same answers, different grader. The answers of six models were held fixed; only the grader changed. The model that is best by the answer key drops as far as 5th under three LLM judges from three labs. Data: leaderboard-audit.
01findings

Consistent is not the same as correct

A judge that gives the same answer every time looks reliable. These results show why that check is not enough.

−38%
change in a published benchmark's mean score when one stated voting rule moves from 3-of-4 to unanimous. 19 of 35 models change rank.
judge-ceiling · ASR benchmark, 39/39 scores reproduced
6.0 pts
bias against one system from a judge that agreed with itself 98.3% of the time. Repeat grading could not reveal it; only a second grader did.
leaderboard-audit · Holm-corrected p = .005
r = 0.35
best automatic judge's agreement with human listeners on voice identity, the TTS dimension where judges are weakest.
judge-ceiling · public TTS judge leaderboard, 8 × 7
Fig. 2  per-model judge biaspoints vs answer key
System P: +4.2 points (too generous) +4.2system P System Q: -6.0 points (too harsh) −6.0system Q −6−30+3+6 too harshtoo generous
One judge, two directions. Same judge, same temperature, 98.3% self-consistent across three gradings.
Fig. 3  rater-panel calculatorSpearman–Brown
ρk = kρ1 / (1 + (k − 1)ρ1)
panel reliability ρk0.73
raters needed for ρk ≥ 0.806
How many listeners do you need? Panel reliability rises with raters, but with diminishing returns. Sizing the panel from measured single-rater reliability avoids paying for raters you don't need.
02research

Open studies

Each project regenerates its published numbers in CI from shipped data, states its limits, and reports negative results next to positive ones.

real dataLLM judges
6 models · 200 GSM8K items
3 judges from 3 labs
124 tests · $0.49 API

Leaderboard Audit

Does an LLM grader treat every model the same?
Method
Hold every model's answers fixed, vary only the grader; compare against a deterministic answer key with exact nonparametric tests.
Finding
The answer-key #1 ranked 2nd, 4th and 5th under the judges. With 200 items the minimum detectable effect is 5.1–7.6 points: two tiers, not six ranks.
#1 → #5
true winner under judge swap
real dataspeechhuman panels
arXiv:2608.19936 reproduced
TTS judge board 8×7
109 tests

Judge-Ceiling

Is the human yardstick behind a leaderboard reliable enough to rank on?
Method
Reproduce published scores from released data; derive the reliability bound each judge correlation implies (ρ1 ≥ r²); test rater screens on simulations with a known answer.
Finding
One stated threshold cut the mean by 38% and reranked 19 of 35 models. A two-sigma outlier gate overstated reliability by 44%. One result contradicted my own hypothesis and is reported.
−38%
mean score, one rule change
preprintagentsedge
Qwen2.5-3B Q4_K_M
~1,000 episodes / arm
4 CPU · 8 GB

LongHaul-Bench

Does an agent's accumulated memory improve it, or slowly rot it?
Method
Seed-reproducible long-horizon protocol with corruption dose-response and an inert control arm; Holm-corrected tests; public CI rerun under resource caps.
Finding
Correct recall +9.4 points, misleading recall −19.8. Plausible corruption hurts; crude corruption does not. The same memory rule gave Llama-3.2 +24.6 and Qwen +0.8. Limit: an LLM-free heuristic beats every agent arm.
−19.8
points from misleading memory
PyPIagents
747 tests
~100k events in ~3 s

Runopsy

Where did an agent run start going wrong, not just where it stopped?
Method
Record the run, localize the likely failure onset with evidence, test the cause by counterfactual replay in a disposable sandbox.
Finding
94.4% top-1, 100% top-3 on a labelled synthetic suite (baselines 22.2%, 50%). 0% top-1 on expert-labelled TRAIL / Who&When traces, published as a limit.
94.4%
top-1 onset localization
simulationvoice agents
40-turn sessions
Llama-3.1-8B judge

VoiceHaul

What happens to voice agents across long conversations?
Finding
The agent the judge panel rated most empathetic failed every 40-turn conversation; the calibrated one failed none. Estimates how many human ratings one judge rating is worth.
all failed
long sessions of the agent judges rated "most empathetic"
03methods

How I check a number

Three questions, asked of every reported score.

Q1 · BIAS

Does the judge treat every system the same?

Per-model bias against an independent second grader, not just repeat grading of the same judge.

biasm = scorejudge(m) − scoreref(m)
Q2 · RESOLUTION

Is the gap bigger than the test can see?

A minimum detectable effect for every reported difference; tiers instead of ranks when models cannot be separated.

report a gap only if |Δ| > MDE0.8, α
Q3 · RELIABILITY

Is the human yardstick itself reliable?

Single-rater reliability, panel size and the effect of screening rules, measured rather than assumed.

ρ1 ≥ r2  (judge–human correlation bound)
04audits

Evaluation audits for teams

The same methods, applied to your benchmark or internal eval, under NDA. Useful before you publish a claim, launch a release or size a human-rating study.

Process2-week pilot → monthly
  1. day 1
    Scoping call
    Pick the one number your decisions rely on most.
  2. week 1
    Reproduce & test
    Judge bias, rater reliability, minimum detectable effect.
  3. week 2
    Verdict
    Plain-language report, reproducible code, review call.
  4. monthly
    Release checks
    Before each release goes public. Cancel any month.
Outputone verdict per claim
"New model is better than v2"✓ supported
"Ranked #1 for voice quality"! needs more data
"Beats competitor X on naturalness"✗ not supported
Illustrative example. Each verdict comes with its reason and the fix. One-page overview (PDF) →
05about

Vahit Feryad, PhD

Applied research engineer. PhD in electrical and electronics engineering (signal processing and machine learning), 10+ years across research and industry, from embedded and edge systems to LLM agents and speech models.

LLM-as-judgebenchmark statisticsrater reliabilityspeech / TTS / ASR evallong-horizon agentsedge & quantized inferenceInspect AIPyTorchTensorRT
  • LongHaul-Bench: A Reproducible Benchmark for Long-Horizon Reliability of Self-Improving LLM Agents Under Industrial Edge Constraints. Preprint, 2026.
  • Efficient Design of Energy Disaggregation Model with BERT-NILM Trained by AdaX Optimization Method for Smart Grid. Energies, 2021.
  • New Design of a Supervised Energy Disaggregation Model Based on the Deep Neural Network for a Smart Grid. Energies, 2019.
  • Multi Objective Clustering for Wireless Sensor Networks. Expert Systems with Applications, 2016.

full record on google scholar →

  • 2024–now
    Data Analytics Manager
    Turkuvaz Media Digital, Istanbul
  • 2022–2023
    Computer Vision Engineer (contractor)
    SMT SportsMEDIA Technology, Durham NC · remote
  • 2017–2022
    AI Researcher
    ASIS Automation and Fueling Systems
  • 2016–2017
    R&D Engineer
    Surteks Automotive
  • 2015–2016
    Electronics & Communication Engineer
    Delphisonic
  • 2014–2021
    PhD, Electrical & Electronics Eng.
    Karadeniz Technical University
$ contact

Have a number you need to trust?

Research collaborations, evaluation audits and applied evaluation work. Based in Istanbul, available remotely for teams in the US and Europe.