OpenAI finds CLIP has 'Halle Berry' neurons like the human brain
Curated by the Inblix editorial team
OpenAI researchers have discovered that CLIP, their vision model matching ResNet-50 performance, contains multimodal neurons that fire on the same concept across wildly different formats. A single neuron might light up for a photograph of a spider, the text ‘spider,’ and an illustration of Spider-Man. This echoes a famous 2005 neuroscience finding where a neuron in a human patient responded to photos, sketches, and the name ‘Halle Berry.’
The presence of these neurons offers a glimpse into why CLIP handles abstraction so well, nailing tough benchmark datasets like ImageNet Sketch and ObjectNet where objects appear as cartoons or statues rather than straightforward photos. The discovery ties directly to the model’s organizational principle. The highest layers of CLIP don’t store rigid visual features. Instead, they arrange information as a loose semantic collection of ideas, a mechanism that seems to mirror both synthetic and natural vision systems.
Using feature visualization and dataset examples, the team found the majority of neurons in a scaled-up CLIP RN50x4 were interpretable. The categories that popped out spanned an extraordinary range: facial expressions, religious iconography, geographical regions, art styles, and even neurons that seem to count. Some neurons responded to images showing evidence of digital alteration. The breadth is startling and feels like a high-level map of the human visual lexicon encoded in weights.
But the analysis also highlights a stubborn mystery about how models truly represent information. CLIP can geolocate photos with eerie precision, sometimes down to a specific San Francisco neighborhood like Twin Peaks. Yet despite deliberate searching, the researchers could not find a single ‘San Francisco’ neuron. The concept doesn’t decompose neatly into units like ‘California’ plus ‘city.’ That information is in there somewhere, but its distributed nature defies the clean, one-neuron-one-concept narrative. The finding is a useful reminder that interpretability, even when it works beautifully, still only gets you a partial view of what these systems know.
💡 Key Takeaways
- CLIP neurons can respond to the same concept presented as a photo, text, or drawing, mirroring multimodal neurons documented in the human brain fifteen years ago.
- The model's high-level layers organize visual information as loose semantic collections of ideas rather than rigid feature detectors, explaining its robustness to abstract and unconventional images.
- Not all knowledge is neatly compartmentalized: CLIP can identify specific city neighborhoods but researchers could not locate a single neuron responsible for that information.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.