M23.8 CONNECT THE MECHANISM
Check whether answers actually follow the supplied media
Two models both score 75%. One invents objects that aren't there; the other misses ones that are. Learn the tests that tell them apart before a user finds out.
LESSON OVERVIEW14 min lesson
Lesson overview
Two models both score 75%. One invents objects that aren't there; the other misses ones that are. Learn the tests that tell them apart before a user finds out.
What you’ll explore
- Multimodal evaluation should distinguish perception, cross-modal alignment, reasoning, and generation; contrastive examples and modality ablations expose shortcuts and unsupported answers.
GO TO THE SOURCE
Original explanations, connected to the research.
Evaluating Object Hallucination in Large Vision-Language Models (POPE; Li et al., 2023)Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality (Thrush et al., 2022)Object Hallucination in Image Captioning (CHAIR; Rohrbach et al., 2018)MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models (Fu et al., 2023)MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark (Yue et al., 2023)Suggest a correction
A precise note can make an explanation better.
Choose the scene and describe what needs attention. Download a feedback file to share through a channel you already use. This page does not send feedback or connect you with a reviewer.