I am once again reminded that LLMs don’t really “see” anything
Chicken nuggets and Dzigurski

I shared the above image with ChatGPT with the following prompt: “I need to write about this marvelous scene.” I was curious about what the AI would have to say.
I was immediately disappointed by the response: a long paean to the USS Indianapolis speech from Jaws. It’s a marvelous scene, to be sure, but it’s not the scene in question.
The still depicts the railway baron Mr. Morton from Sergio Leone’s masterpiece Once Upon a Time in the West, in the railroad car that has become his prison, staring at a seascape by Serbian-American artist Alexander Dzigurski. The scene is worth its own discussion (it sets up Morton as a sympathetic villain and is paid off in an hauntingly poetic way later in the film), but the main point here is that the still does not really look like Jaws at all. Sure, the image is low resolution (I pulled it from the Bodega Bay Heritage Gallery’s page on Dzigurski), but the only similarity between the two scenes is that they both prominently feature sweaty men.
I know that “ha ha, AI failed to meet some arbitrary test I gave it” became a tired genre of post years ago, but I think this incident is worth writing about precisely because AI has become so adept at multimodality. I have regained the ability to be disappointed by the models because they fail so infrequently now (at many ordinary tasks, anyway).
Image comprehension—combined with the ability to explain humour—was something that bowled me over when OpenAI released their first multimodal model (GPT-4). I still remember the famous “chicken nugget map” from the GPT-4 Technical Report.

Humour was supposed to be hard for robots! There was a whole 2015 episode of On the Media about it! And okay, the explanation that the humour comes from the unexpected juxtaposition between the majestic description and the very ordinary nuggies is a bit generic, but isn’t all humour ultimately based on incongruity?
Anyway, where was I again? I guess this episode just serves as a reminder that as much as we are tempted to anthropomorphize these models, the more human-like their outputs seem in light of advancing capabilities, they do not think like us. As we have mentioned previously on this blog, a multimodal model can smash visual understanding benchmarks without ever having access to the underlying images it is supposed to be understanding.
Whatever these models are doing, it is not what we mean when we say that we can see.
