Workspace IndexDev Notes › Evals — you cannot improve what you do not measure, judge included

#206PoC

Evals — you cannot improve what you do not measure, judge included

LLM evaluation ranges from exact-match benchmarks to using a model as a judge, and the judge itself has biases (length, position, self-preference) that must be measured before its scores are trusted.

Not yet scoped.

Why

The PoC builds a small eval set and an LLM-judge, then measures the judge's own biases — the meta-evaluation that keeps a scoreboard honest.

How it works

Not yet built.

← All Dev Notes · Workspace Index · Top ↑

평가 — 측정하지 않는 것은 개선할 수 없다, 심판까지 포함해

LLM 평가는 정확 일치 벤치마크부터 모델을 심판으로 쓰는 것까지 걸쳐 있으며, 심판 자체에 편향(길이, 위치, 자기 선호)이 있어 그 점수를 믿기 전에 측정해야 합니다.

아직 범위 미정.

이 PoC는 작은 평가 집합과 LLM 심판을 만든 뒤 심판 자신의 편향을 측정합니다 — 점수판을 정직하게 유지하는 메타 평가입니다.

동작 방식

아직 만들지 않음.

← 전체 개발 노트 · 워크스페이스 인덱스 · 맨 위 ↑