Why
The PoC builds a small eval set and an LLM-judge, then measures the judge's own biases — the meta-evaluation that keeps a scoreboard honest.
How it works
Not yet built.
Workspace Index › Dev Notes › Evals — you cannot improve what you do not measure, judge included
#206PoC
LLM evaluation ranges from exact-match benchmarks to using a model as a judge, and the judge itself has biases (length, position, self-preference) that must be measured before its scores are trusted.
The PoC builds a small eval set and an LLM-judge, then measures the judge's own biases — the meta-evaluation that keeps a scoreboard honest.
Not yet built.
LLM 평가는 정확 일치 벤치마크부터 모델을 심판으로 쓰는 것까지 걸쳐 있으며, 심판 자체에 편향(길이, 위치, 자기 선호)이 있어 그 점수를 믿기 전에 측정해야 합니다.
이 PoC는 작은 평가 집합과 LLM 심판을 만든 뒤 심판 자신의 편향을 측정합니다 — 점수판을 정직하게 유지하는 메타 평가입니다.
아직 만들지 않음.