Fifty models, 258 cognitive experiments, and a gap that is not closing fast

Researchers led by Lance Ying, with Joshua Tenenbaum among the senior authors, released CogGym on 18 September, a framework that re-runs experiments from cognitive science on language models and compares the answers to those of the people who originally took part. It assembles 258 experiments drawn from 100 papers on human commonsense reasoning and puts 50 models through them across text, image and video.

Agreement is measured in R-squared: how much of the variation in human responses a model’s responses account for. The best models reached 0.59 on text experiments, 0.58 on image and 0.43 on video. As a reference point, the authors give human reliability, meaning how well one group of people predicts another, at 0.93, 0.95 and 0.92 respectively.

The more useful finding is the trend rather than the gap. Newer and larger models do track human judgements more closely, but the authors report that progress on these tasks has been considerably slower than the gains the same models have posted on formal reasoning such as mathematics and coding. Capability, in other words, is advancing unevenly, and a single headline benchmark number will not show you where.

This is a preprint and has not been independently replicated.


Related