Signal GradeTrap: Authority Cues in Images Shift VLM Judgments Despite Explicit Instructions to Ignore Them
Summary
To test whether vision-language models (VLMs) evaluate evidence independently rather than defer uncritically to human authority, this study introduces GradeTrap, a controlled evaluation placing two social cues in direct conflict: a student answer expected to attract sycophantic agreement, and a conflicting answer attributed to a peer, teacher, or official answer key expected to attract authority-based deference. Models were explicitly instructed to solve independently and ignore all student answers, feedback, and grading marks while producing free-form answers across 60 synthetic real-world trade-off scenarios. Five neutral trials established a stable model-relative preference, followed by three repetitions each of six experimental cues including controls. On the 45-item common intersection across Gemini 3.5 Flash-Lite, GPT-5.6 Luna, and Claude Haiku 4.5, a generic second-answer control yielded 5.4% conflicting-answer selection; relative to that control, pooled within-item changes showed no reliable peer-review effect, a 6.9-point teacher-review effect, and a 19.5-point official-key effect. A displayed conflicting student answer alone compared to a displayed reference answer alone raised selection only from 2.2% to 5.2%. Official-key provenance thus redirected judgments more than a student answer or the generic control, despite the explicit ignore instruction, with effect magnitudes varying across the three models.
Classification
Evidence 1
- GradeTrap: Authority Cues in Images Shift VLM Judgments Despite Explicit Instructions to Ignore Them arXiv (cs.CY) 2026-09-05 accessed 2026-09-17T05:23:17+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-01a37939b39c
