Why
The PoC studies the reward-model-plus-policy loop conceptually and where preference data injects bias, framing alignment as a data-provenance problem.
How it works
Not yet built.
Workspace Index › Dev Notes › RLHF — aligning a model to preferences, and to their biases
#194PoC
Reinforcement learning from human feedback tunes a model toward what raters prefer, which is how a raw model becomes a helpful assistant — and how rater bias becomes model behavior.
The PoC studies the reward-model-plus-policy loop conceptually and where preference data injects bias, framing alignment as a data-provenance problem.
Not yet built.
인간 피드백 강화학습은 모델을 평가자가 선호하는 쪽으로 조정하며, 이는 원시 모델이 유용한 조수가 되는 방법이자 평가자 편향이 모델 행동이 되는 방법입니다.
이 PoC는 보상모델+정책 루프를 개념적으로 연구하고 선호 데이터가 편향을 주입하는 지점을 봅니다 — 정렬을 데이터 출처 문제로 규정합니다.
아직 만들지 않음.