Workspace IndexDev Notes › RLHF — aligning a model to preferences, and to their biases

#194PoC

RLHF — aligning a model to preferences, and to their biases

Reinforcement learning from human feedback tunes a model toward what raters prefer, which is how a raw model becomes a helpful assistant — and how rater bias becomes model behavior.

Not yet scoped.

Why

The PoC studies the reward-model-plus-policy loop conceptually and where preference data injects bias, framing alignment as a data-provenance problem.

How it works

Not yet built.

← All Dev Notes · Workspace Index · Top ↑

RLHF — 모델을 선호에, 그리고 그 편향에 정렬하기

인간 피드백 강화학습은 모델을 평가자가 선호하는 쪽으로 조정하며, 이는 원시 모델이 유용한 조수가 되는 방법이자 평가자 편향이 모델 행동이 되는 방법입니다.

아직 범위 미정.

이 PoC는 보상모델+정책 루프를 개념적으로 연구하고 선호 데이터가 편향을 주입하는 지점을 봅니다 — 정렬을 데이터 출처 문제로 규정합니다.

동작 방식

아직 만들지 않음.

← 전체 개발 노트 · 워크스페이스 인덱스 · 맨 위 ↑