Why
The PoC studies the self-critique loop and where a principle set decides behavior, framing the constitution as the reviewable seat of the model's values.
How it works
Not yet built.
Workspace Index › Dev Notes › Constitutional AI — alignment from written principles, not just raters
#203PoC
Constitutional AI has a model critique and revise its own outputs against a written set of principles, reducing reliance on human labels — and moving the value judgment into an auditable document.
The PoC studies the self-critique loop and where a principle set decides behavior, framing the constitution as the reviewable seat of the model's values.
Not yet built.
헌법적 AI는 모델이 적힌 원칙 집합에 대해 자기 출력을 비평·수정하게 하여 인간 라벨 의존을 줄이고 — 가치 판단을 감사 가능한 문서로 옮깁니다.
이 PoC는 자기 비평 루프와 원칙 집합이 행동을 결정하는 지점을 연구하여, 헌법을 모델 가치의 검토 가능한 자리로 규정합니다.
아직 만들지 않음.