SyncAI.news, a Varaisys broadcasting
A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities
FG

Faiz Ghifari Haznitrama, Faeyza Rishad Ardi, Alice Oh

· 1 min read

ResearcharXiv cs.AI

A Neuropsychologically Grounded Evaluation of LLM Cognitive Abilities

arXiv:2603.02540v2 Announce Type: replace Abstract: Large language models (LLMs) display a unified "general factor" of capability across 10 benchmarks (a finding confirmed by our factor analysis of 156 models), yet they still struggle with simple, trivial tasks for humans. This is because current benchmarks focus on task completion, failing to probe the foundational cognitive abilities that highlight these behaviors. We address this by introducing the NeuroCognition benchmark, grounded in three adapted neuropsychological tests targeting distinct foundational cognitive components: Raven's Progressive Matrices (abstract relational reasoning), Spatial Working Memory (goal-directed spatial updating), and the Wisconsin Card Sorting Test (cognitive flexibility). Our evaluation reveals that while models perform strongly on text, their performance degrades for images and with increased complexity. Comparison with a human baseline shows that LLMs and humans fail at different parts of the same tasks. Furthermore, we observe that complex reasoning is not universally beneficial, whereas simple, human-like strategies yield partial gains. We also find that NeuroCognition correlates positively with standard general-capability benchmarks, while still measuring distinct cognitive abilities beyond them. Overall, NeuroCognition emphasizes where current LLMs align with human-like intelligence and where they lack core adaptive cognition, showing the potential to serve as a verifiable, scalable source for improving LLMs.

Original source

This story was published by arXiv cs.AI and written by Faiz Ghifari Haznitrama, Faeyza Rishad Ardi, Alice Oh. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News