Workspace IndexDev Notes › Multimodal — one model that reads images and text together

#208PoC

Multimodal — one model that reads images and text together

Vision-language models take pixels and tokens in the same context, enabling screenshot understanding and document parsing — and inheriting prompt-injection risk through images, not just text.

Not yet scoped.

Why

The PoC runs a VLM on a screenshot task and shows both the capability and an image-borne injection, connecting multimodal power to its new attack surface.

How it works

Not yet built.

← All Dev Notes · Workspace Index · Top ↑

멀티모달 — 이미지와 텍스트를 함께 읽는 한 모델

비전-언어 모델은 픽셀과 토큰을 같은 컨텍스트에 받아 스크린샷 이해와 문서 파싱을 가능하게 하며 — 텍스트만이 아니라 이미지를 통한 프롬프트 인젝션 위험도 물려받습니다.

아직 범위 미정.

이 PoC는 스크린샷 과제에서 VLM을 돌려 능력과 이미지로 실린 인젝션을 함께 보이며, 멀티모달의 힘을 그 새 공격 표면으로 연결합니다.

동작 방식

아직 만들지 않음.

← 전체 개발 노트 · 워크스페이스 인덱스 · 맨 위 ↑