Summary:
A new study in Current Biology from researchers at York University reveals a fundamental computational divide between biological vision and artificial intelligence. By testing motion aftereffect illusions in humans, primates, and artificial neural networks, the team found that while primate inferior temporal (IT) cortex dynamically shifts object position codes based on recent visual history, current AI vision models remain rigid and pixel-bound. The findings introduce a novel benchmark for NeuroAI, showing that human-compatible machine vision must incorporate dynamic, history-dependent perceptual computations.
Key Facts:
- The Illusion Divide: When exposed to a motion aftereffect, human observers perceive a stationary object as shifted in the opposite direction. Primate inferior temporal (IT) cortex neurons mirrored this subjective shift, but current state-of-the-art AI vision networks failed to reproduce it.
- History-Dependent Perception: Biological vision continuously reshapes its spatial representations based on recent temporal experience, treating perceptual shifts as an adaptive feature rather than a flaw. AI systems, by contrast, remain strictly tethered to static pixel coordinates.
- New NeuroAI Benchmark: The research establishes that building AI systems that work safely and intuitively alongside humans requires training them on dynamic perceptual computations rather than solely optimizing for static pixel accuracy.
Source: York University
Human eyes and primate visual brains do not operate like passive digital cameras. Instead of continuously cataloging raw, objective pixel coordinates, biological vision relies on adaptive computations that actively integrate immediate visual history to make sense of a dynamic world.
Sometimes, this historical context leads to systematic perceptual quirks. In the classic “motion aftereffect” illusion, adapting to continuous directional motion causes a subsequent stationary object to appear displaced in the opposite direction. The physical light hitting the retina has not shifted, yet the conscious perception of where the object sits in space has moved.
Rather than an engineering defect, neuroscientists increasingly view these perceptual “mistakes” as signatures of optimal, energy-efficient biological computations.
Now, an investigation by researchers at York University asks a critical question for the future of artificial intelligence: If AI vision models are meant to interact with and understand the world like humans, should they reproduce these exact same perceptual illusions?
Published in the journal Current Biology, the study demonstrates that leading artificial vision networks completely lack the history-dependent spatial flexibility shared by humans and non-human primates.
โTodayโs AI vision systems are impressive, but they still do not always see the world the way we do. This study captures the promise of NeuroAI and what it can do when neuroscience and artificial intelligence are brought together. By using smart experiments to reveal the computations biological vision uses and AI still lacks, we can use those insights to build better, more brain-like artificial systems,โ said senior author Kohitij Kar, Ph.D., Canada Research Chair in Visual Neuroscience and Assistant Professor in the Faculty of Science at York University.
Tracking Neural Coordinates in the Primate IT Cortex
Current deep artificial neural networks (ANNs) excel at spatial object recognition, often matching or surpassing humans at classifying objects in static frames. However, they evaluate each image primarily through feedforward, physical pixel properties without the continuous temporal adaptation characteristic of animal vision.
To pin down where artificial and biological vision diverge, the York team, led by graduate researcher and first author Elizaveta Yakubovskaya, paired human psychophysics with electrophysiological recordings from the primate inferior temporal (IT) cortex, a higher-order visual area known for object recognition.
The investigators exposed human observers and macaque monkeys to motion adaptation and presented stationary targets to induce the position-shift illusion.
โCombining recordings from primate visual cortex with human perception experiments, we used motion adaptation to induce a visual illusion and make a stationary object appear slightly shifted in position, then asked whether the brain and AI showed the same effect,” explained Yakubovskaya. “Human observers reported the illusion, and neural representations of position in the primate inferior temporal (IT) cortex shifted in the same direction, even though the image itself had not changed.โ
When the researchers tested standard artificial vision networks on the identical paradigm, the models showed no such adaptation. The AIโs internal spatial representations remained anchored to the objective, unshifted pixel coordinates, failing to reflect the perceptual reality experienced by the biological visual system.
Bridging the NeuroAI Gap for Human-Machine Alignment
The findings highlight that the primate IT cortex does not merely encode static identity; it represents object position in coordinates aligned with conscious perception. This discovery provides an essential new computational benchmark for evaluating dynamic, recurrent vision models.
As autonomous systems, robotic assistants, and computer vision algorithms become integrated into everyday environments, such as driving alongside human motorists or interpreting real-time medical scans alongside clinicians, discrepancies in how humans and machines perceive spatial relations could lead to dangerous misalignments.
โThere is a growing question in AI about whether increasingly capable systems will become more like us or increasingly different from us,โ noted Dr. Kar, who is also an investigator with York’s Centre for Vision Research and the Connected Minds initiative.
โIf we want AI that works with humans and understands the world in more human-compatible ways, we cannot focus only on whether it gets the right answer. We also need to understand the computations that produce human perception and behavior. Neuroscience gives us a way to discover those computations and, potentially, build them into AI.โ
Editorial Notes:
- This article was edited by a Neuroscience News editor.
- Journal paper reviewed in full.
- Additional context added by our staff.
About this AI and visual neuroscience Research:
- Media Contact:ย Sandra McLean
- Source:ย York University
- Image Credit:ย Image credited to Neuroscience News
- Original Research is Open Access:ย Current Biology (Sept 22, 2026). โThe macaque IT cortex but not current artificial vision networks encode object position in perceptually aligned coordinates.โ Authors: Elizaveta Yakubovskaya, Hamidreza Ramezanpour, Matteo Dunnhofer, and Kohitij Kar.
- DOI:ย 10.1016/j.cub.2026.09.019
Abstract
The macaque IT cortex but not current artificial vision networks encode object position in perceptually aligned coordinates
Efficient interaction with the visual world requires not only object identification but also localization of where objects are in space. While spatial (โwhereโ) processing has classically been attributed to dorsal stream pathways, recent work has shown that object position can also be decoded from ventral stream responses, including the inferior temporal (IT) cortex.
However, because object position in these paradigms is coupled to pixel-based location, it has remained unclear whether ventral stream position signals are perceptually meaningful or instead reflect incidental inheritance from retinotopic inputs.
Here, we address this question by leveraging a visual illusion, the motion aftereffect, to dissociate perceived object position from retinal location while holding visual input constant. Combining intracortical recordings in macaque IT with matched human psychophysics, we show that motion adaptation induces direction-opponent biases in IT population codes for object position that mirror human perceptual reports, despite unchanged pixel-level input.
Motion adaptation reshapes IT representational geometry that likely contributes to these perceptual biases. Extending these findings to artificial vision systems, we observe that feedforward, recurrent, and state-of-the-art video-based neural networks fail to exhibit adaptation-induced position shifts, despite accurately encoding object position.
Interestingly, imposing empirically derived IT-based transformations on model features is sufficient to simulate the effect, revealing adaptation-driven representational warping as a missing computational ingredient in artificial vision systems.
Together, these results identify IT as a candidate locus of perceptually aligned spatial coding, reveal adaptation-driven representational restructuring as a mechanism linking neural dynamics to perception, and expose a principled gap between biological and artificial vision.

