Why
The PoC runs a VLM on a screenshot task and shows both the capability and an image-borne injection, connecting multimodal power to its new attack surface.
How it works
Not yet built.
Workspace Index › Dev Notes › Multimodal — one model that reads images and text together
#208PoC
Vision-language models take pixels and tokens in the same context, enabling screenshot understanding and document parsing — and inheriting prompt-injection risk through images, not just text.
The PoC runs a VLM on a screenshot task and shows both the capability and an image-borne injection, connecting multimodal power to its new attack surface.
Not yet built.
비전-언어 모델은 픽셀과 토큰을 같은 컨텍스트에 받아 스크린샷 이해와 문서 파싱을 가능하게 하며 — 텍스트만이 아니라 이미지를 통한 프롬프트 인젝션 위험도 물려받습니다.
이 PoC는 스크린샷 과제에서 VLM을 돌려 능력과 이미지로 실린 인젝션을 함께 보이며, 멀티모달의 힘을 그 새 공격 표면으로 연결합니다.
아직 만들지 않음.