06/01/2026
Nobody:
Absolutely nobody:
Language models doing aesthetic judgments:
"Hmm, comparing these two images... yes, the second image has better lighting and balance. Analysis complete: second image is superior. Wait... but it's in second position? That feels wrong. Better pick the first one instead. 60% confident! 🎨"
The experiment: We asked our vision language model to compare two images and pick the better one.
When we swapped the image positions, the model suddenly picked the other image as better. 🤔
What we discovered: Our vision language model spends 27 layers carefully analyzing aesthetics, correctly identifying the better image... then throws it all away in layer 28 because of position bias. It's like studying for an exam, knowing the right answer, then changing it at the last second because "the first option is usually correct."
The fascinating part? There IS actual reasoning happening inside these models layer by layer. But then the final layer goes "nah, I'll trust my vibes over my analysis" and introduces a bias that wasn't there before.
Lesson learned: Sometimes the model's penultimate thoughts are smarter than its final answer. Layer 27 > Layer 28. ✨
Who knew AI models could overthink themselves just like us?