Workspace IndexDev Notes › Constitutional AI — alignment from written principles, not just raters

#203PoC

Constitutional AI — alignment from written principles, not just raters

Constitutional AI has a model critique and revise its own outputs against a written set of principles, reducing reliance on human labels — and moving the value judgment into an auditable document.

Not yet scoped.

Why

The PoC studies the self-critique loop and where a principle set decides behavior, framing the constitution as the reviewable seat of the model's values.

How it works

Not yet built.

← All Dev Notes · Workspace Index · Top ↑

헌법적 AI — 평가자만이 아니라 적힌 원칙에서 오는 정렬

헌법적 AI는 모델이 적힌 원칙 집합에 대해 자기 출력을 비평·수정하게 하여 인간 라벨 의존을 줄이고 — 가치 판단을 감사 가능한 문서로 옮깁니다.

아직 범위 미정.

이 PoC는 자기 비평 루프와 원칙 집합이 행동을 결정하는 지점을 연구하여, 헌법을 모델 가치의 검토 가능한 자리로 규정합니다.

동작 방식

아직 만들지 않음.

← 전체 개발 노트 · 워크스페이스 인덱스 · 맨 위 ↑