What Does the Internet Do to the Brain?

Activation Cartography maps 3,008 natural language stimuli across 13 internet content categories against predictions from TRIBE v2 (a 177M-parameter deep neural encoder trained on real fMRI recordings), revealing statistically significant, category-level differences in predicted cortical recruitment.

3,008stimuli evaluated
13content categories
20,484cortical surface points

Research Snapshot

Study design, encoder architecture, and key statistical findings.

F = 13.51ANOVA main effect
96.9%variance in PC1
177Mencoder parameters
semantic spread gain
Encoder
TRIBE v2: deep neural fMRI encoder predicting whole-cortex haemodynamic responses
Key result
Significant main effect: F(12, 2995) = 13.51, p < 10²⁶, η² = 0.051. Cohen's d = −0.82 between highest and lowest categories.
Theory
GWT receives strongest support; FEP weakly supported; DCT and IIT mixed evidence
Tagged
fMRI EncodingInternet ContentCortical MappingDeep LearningGlobal Workspace Theory

Concept Overview

fMRI EncodingCortical MappingTRIBE v2

Abstract

Background — Despite extensive single-stimulus neuroscience on emotional, narrative, and threatening media, no large-scale comparative study exists of how distinct categories of internet content differentially engage the cortex at scale.

Methods — Activation Cartography maps 3,008 natural language stimuli across 13 internet content categories against predictions from TRIBE v2, a 177M-parameter deep neural encoder trained on real fMRI recordings that predicts whole-cortex haemodynamic responses. Each stimulus yielded a predicted activation profile across 20,484 cortical surface points, summarised into six anatomical regions.

Results — A one-way ANOVA revealed a significant main effect of content type (F(12, 2995) = 13.51, p < 10²⁶, η² = 0.051). ThreatSafety content ranked highest and Narrative lowest (Cohen’s d = −0.82). A dominant cortical gradient (PC1 = 96.9% variance) contrasts sensory-language against executive-motor cortex across all categories.

Implications — Different internet content categories engage distinct brain circuits with statistically significant differences in predicted intensity. GWT’s prediction that threat-laden content drives broad cortical activation received the strongest support. The analysis pipeline and registered hypotheses are released with the project.

Introduction

Digital media consumption has become a defining feature of contemporary cognitive life. Recent estimates indicate the average adult consumes six to eight hours of digital content per day, a duration that exceeds sleep for many subpopulations. A fundamental empirical question follows: whether distinct categories of internet content engage the cerebral cortex equivalently, or whether systematic, category-level differences in predicted neural recruitment can be identified at scale.

Activation Cartography

Rather than running a single-category fMRI study, Activation Cartography uses a validated deep neural encoder (TRIBE v2) to predict whole-brain responses at scale. This enables a 13-category comparative study with 3,008 stimuli that would be logistically impossible with real scanner time.

The neuroscience of media consumption has historically been constrained by the throughput of fMRI acquisition: each participant yields a few hundred trials per session, making large comparative studies prohibitively expensive. TRIBE v2 breaks this barrier by predicting whole-cortex haemodynamic responses from text inputs, enabling population-scale analysis of content-type effects on predicted brain activation.

Four neuroscientific frameworks are evaluated against the activation patterns: Global Workspace Theory (GWT), Free Energy Principle (FEP), Default-mode Circuit Theory (DCT), and Integrated Information Theory (IIT). Each makes distinct predictions about which content types should drive the broadest or most intense cortical recruitment.

Methods

  1. Stimulus Construction — 3,008 stimuli drawn from established NLP benchmarks and live internet sources, distributed across 13 content categories: ThreatSafety, News, Social, Scientific, Narrative, Emotional, AudioText, ImageVisual, Educational, Persuasive, Humour, Instructional, and Commerce. Each stimulus is a short natural language passage (1–3 sentences). 3 008 Stimuli 13 Categories NLP Benchmarks
  2. TRIBE v2 Encoder — TRIBE v2 is a 177-million-parameter deep neural encoder trained on real functional MRI recordings. Given a text input, it predicts a whole-cortex haemodynamic response across 20,484 cortical surface points. Two encoding modes are used: hash-mode (fast, token-level) and semantic-mode (LLaMA-3.2-3B embeddings, N = 390 replication sample). 177M Parameters 20 484 Cortical Points LLaMA-3.2-3B
  3. Statistical Analysis — Each stimulus’s 20,484-point activation profile is summarised into six anatomical regions. A one-way ANOVA tests the main effect of content type on mean global activation. Effect sizes reported as η² and Cohen’s d. Principal component analysis across category mean profiles identifies dominant cortical gradients. One-way ANOVA Cohen’s d PCA

Results

A one-way ANOVA on predicted global activation revealed a statistically significant main effect of content type across all 13 categories.

ANOVA: F(12, 2995) = 13.51, p < 10²⁶

Under hash-mode encoding, ThreatSafety content ranked highest and Narrative lowest (Cohen’s d = −0.82). The semantic replication (N = 390, LLaMA-3.2-3B) produced a 4× wider activation spread, with AudioText, ImageVisual, and Emotional leading, a ranking essentially uncorrelated with hash-mode ordering (r = 0.09).

Dominant cortical gradient (PC1 = 96.9% variance) contrasts sensory-language cortex (high loading: auditory, visual, language network) against executive-motor cortex (low loading: prefrontal, motor) across all 13 categories. This gradient is consistent across encoding modes despite the different category rankings.

Regional breakdown shows that ThreatSafety content activates the language network and prefrontal cortex most strongly under hash-mode, while AudioText and ImageVisual content drives the largest visual and auditory cortex responses under semantic encoding. This suggests hash-mode captures surface lexical features while semantic-mode captures deeper representational content.

Robustness Checks: Ranking is Stable, Not Just Surface Text

Two checks confirmed the semantic-mode ranking is not an artefact. (1) seq_len sweep: the per-CT ordering is essentially identical across temporal integration windows seq_len = 4, 8, 16 (Spearman ρ ≥ 0.956, all p < 0.001). (2) LSA Mantel test: TRIBE’s representational geometry is near-zero correlated with surface text similarity (Mantel r = 0.055, p = 0.340), confirming the encoder captures stimulus-specific neural coding. Vertex-level analysis (20,484 independent ANOVAs) further shows the discriminative effect is whole-brain: 99.8% of vertices are significant at p < 0.001, with the top vertex reaching F = 71.05 (5.3× the six-region aggregate).

Cross-source robustness confirms category effects are not dataset artefacts. For each category represented in two or more independent sources, an ICC-style reproducibility index was computed. The mean ICC across 12 categories is 0.89; 10 of 12 fall in the good-to-excellent range. Narrative is the exception (ICC = 0.67), consistent with its heterogeneous sourcing (HellaSwag completions vs TinyStories).

Extended multivariate analyses further confirm content-type profiles are discriminable above chance: linear SVM classification achieves significant cross-validated accuracy, hierarchical clustering reveals a structured taxonomy (sensory categories cluster separately from semantic ones), and mutual information analysis shows distinct information carried by each cortical region — consistent with the view that different content types engage specialised neural subsystems rather than a single undifferentiated response.

Theory Evaluation

Four neuroscientific frameworks were assessed against predicted activation patterns, each making distinct testable predictions about which content types should drive the broadest cortical recruitment.

Global Workspace Theory: Strongest Support

GWT predicts that threat-laden content should trigger a global broadcast, driving widespread cortical ignition. ThreatSafety content ranking highest under hash-mode (d = −0.82) directly supports this. The finding that a single dominant gradient accounts for 96.9% of between-category variance is also consistent with GWT’s single-workspace model.

Free Energy Principle predicts prediction-error-rich content (novel, surprising, uncertain stimuli) should drive higher activation. The moderate support observed is consistent: ThreatSafety and News content (high surprise value) rank highly, but the correlation with uncertainty proxies is weak (r ≈ 0.3).

Default-mode Circuit Theory predicts narrative and self-referential content should activate default mode network most strongly. This receives mixed evidence: Narrative ranked lowest under hash-mode, but higher under semantic encoding. This suggests encoding mode mediates the narrative-DMN link.

Integrated Information Theory predicts content with higher integrated information (Φ) should drive more cortical activation. IIT receives mixed evidence: there is no reliable proxy for Φ in natural language stimuli, making this prediction untestable at current resolution.

Quantitative Theory Scorecard

A formal numerical evaluation (1.0 = confirmed, 0.5 = partial, 0.0 = not confirmed) across six predictions produces the following aggregate scores: GWT scores highest (0.58), driven by ThreatSafety’s consistent top ranking. IIT and DCT tie at 0.33, while FEP scores 0.25. These scores establish a principled baseline against which future semantic-replication results can be compared.

Conclusion

Activation Cartography demonstrates that different internet content categories engage distinct brain circuits with statistically significant differences in predicted intensity (F(12, 2995) = 13.51, p < 10²⁶, η² = 0.051). The dominant cortical gradient (sensory-language vs executive-motor) is stable across encoding modes and accounts for 96.9% of between-category variance.

The encoding-mode dependence of category rankings (r = 0.09 between hash-mode and semantic-mode) is the study’s most important methodological finding: surface lexical features (hash-mode) and deep semantic representations (semantic-mode) produce systematically different activation predictions, suggesting that fMRI encoding models are sensitive to which level of linguistic representation is used as input.

Robustness extensions confirm the core result. Six extension analyses collectively support a nuanced picture: (i) the discriminative effect is whole-brain and 5× stronger at vertex resolution than the six-region aggregate; (ii) category effects reproduce across independent data sources (mean ICC = 0.89); (iii) the activation spectrum is a smooth continuum; (iv) the semantic-mode ranking (AudioText > ImageVisual > Emotional) is stable across temporal integration windows (ρ ≥ 0.956) and independent of surface text similarity (Mantel r = 0.055, p = 0.340). Extended multivariate analyses confirm category profiles are discriminable above chance with a structured hierarchical organisation.

Future directions include higher-powered semantic replication (N ≥ 150 per category) to resolve the hash/semantic discrepancy, extension to multimodal stimuli (images, audio) using TRIBE v2’s full multimodal encoder, independent cross-encoder validation with BrainBERT or the Huth semantic atlas, and pre-registration of the category-ranking hypothesis for confirmatory testing.

Abstract

Past brain-scan studies have looked at one thing at a time: a scary video clip here, a sad story there. Nobody had lined up the many different kinds of content people actually scroll through online and compared how the brain responds to each, side by side. This study fills that gap using an AI model that predicts brain activity from text, instead of needing an actual brain scanner for every single test.

Introduction

People now spend six to eight hours a day consuming digital content — more than many people sleep. A basic question follows: does the brain respond to different types of online content (scary news, funny posts, stories, ads) in the same way, or do different categories reliably trigger different patterns of brain activity?

Normally, answering this would require scanning real people's brains in an MRI machine while they read thousands of passages, which is far too slow and expensive to do at scale. This study instead uses TRIBE v2, an AI model trained on real brain-scan data, to predict how the brain would likely respond to text — letting the researchers effectively test thousands of passages that would be impossible to test with a real scanner. Four competing brain-science theories about which kind of content should grab the brain's attention most are then tested against the results.

Methods

The researchers gathered 3,008 short passages of text (headlines, social posts, stories, and more), sorted into 13 categories like "threatening/safety," "news," "social," "narrative," "humor," and so on. Each passage was fed into TRIBE v2, an AI model with 177 million internal tuning values that was trained on real fMRI brain scans, which predicts how blood flow (the signal an fMRI scanner actually measures) would shift across more than 20,000 points on the brain's surface. Those points were then grouped into six major brain regions to make the results easier to compare.

A standard statistical test was then used to check whether the predicted brain response really differed by content category, or whether any differences looked like they could just be random noise.

Results

The statistical test confirmed that content category genuinely matters — the differences are far too large to be random chance. Threatening or safety-related content triggered the strongest predicted brain response, while story-style narrative content triggered the weakest, a sizeable and consistent gap.

One single pattern dominated the results across every category: a tug-of-war between the brain's sensory-and-language regions and its planning-and-movement regions, which alone explained about 97% of all the variation seen between categories. A series of additional checks (testing different AI settings, different data sources, and even individual points on the brain rather than whole regions) all confirmed this wasn't a fluke of one particular setup — the effect held up consistently.

Theory Evaluation

Four existing brain-science theories about what should most grab the brain's attention were tested against these results. The strongest match was "Global Workspace Theory," which predicts that threatening content should trigger a brain-wide alert — exactly what the threatening/safety category did, ranking highest of all 13 categories.

The other three theories fared worse: one theory (about surprising or uncertain content driving attention) got only weak support, and two others produced mixed or inconclusive results, in one case because there wasn't a good way to even measure what that theory predicts using text alone.

Conclusion

Different kinds of online content really do switch on different brain circuits, and the differences in predicted intensity are statistically real, not just noise. Of the four competing theories tested, the idea that threatening content triggers a broad, brain-wide response came out with the strongest support.

An interesting side-finding: the exact ranking of which content types are "most stimulating" shifted somewhat depending on which AI-encoding approach was used to process the text, suggesting that how you feed text into these prediction models matters and deserves more attention in future work. The researchers have released their full code, data, and their original predictions (made before running the study) publicly.

Back to Research