Google DeepMindP042026-08-26full textvisual-general-intelligenceworld-modelsvideo-generationembodied-aiscaling

Visual General Intelligence: A White Paper

A multi-lab group of computer vision researchers each argue, from their own specialty, why and how learning from vision (not just language) might lead to general intelligence.

It maps the competing bets AI labs are placing on vision versus language as the next route to AGI, before any of those bets have resolved.

Hirokatsu Kataoka · Yoshihiro Fukuhara · Yonglong Tian · Shangzhe Wu · Oishi Deb · Ryousuke Yamada · Christian Rupprecht · Jianyuan Wang · Kohsuke Ide · Koichi Namekata · Xianzheng Ma · Yiming Chen · Robert Geirhos · Aditi Raghunathan · … — meet the researchers →

How much of this do you want?
Orient me keeps four things: the abstract, the method, the claim↔evidence panel, and where it leads next. Everything adds constructs, the model table, every reported statistic, the discussion framing and the style moves. Switching hides nothing permanently and never changes what a section says — it only changes how many are on screen.
Abstract

Two readings, equal authority

How to choose: The paper’s words is verbatim — use it when you need to quote, or to judge how they write. Plain language is a paraphrase written for comprehension — use it when you want the idea fast. Neither is a summary of the other; they are two doors into the same room.

“This paper reconsiders intelligence from a vision-centered perspective and examines whether intelligence emerging from visual experience and learning may provide a pathway toward AGI. In the language domain, beginning with the introduction of the Transformer architecture, the GPT series has demonstrated transfer to unseen tasks through autoregressive language modeling on web-scale text combined with aggressive scaling. This raises a natural question, namely, what capabilities and forms of intelligence can emerge from visual modalities such as images, videos, and geometry? In this paper, we discuss whether visual intelligence can serve as a pathway toward AGI, referred to in this paper as visual general intelligence (VGI), by bringing together contributors from diverse standpoints and affiliations. Our aim is not to offer a single definition of visual intelligence, but to clarify the principles that computer vision should pursue in the AGI era, the visual input modalities, the benchmarks, the learning paradigms, and the relationship between vision, when taken as the core, and other modalities such as language.”

Constructs

What this paper defines

Every definition below is the paper’s own sentence, with its locator. The plain gloss is a reading aid and is marked as one.

Visual general intelligence (VGI)

“Rather, it is a research agenda that asks what vision can understand, predict, and generalize, either before being coupled with language or while interacting with it.”1. Introduction

In plain terms: The idea that general intelligence might come from learning to see, not only from learning language.

Spatial AI

“I defined Spatial AI as the capability for artificial devices to build genuinely rich but efficient representations of their surroundings which enable them to interact with them in the generally intelligent ways people can.”2.6 What is the Computational Structure of Spatial AI?

In plain terms: A device's ability to build a compact but rich internal map of its surroundings that it can use to act intelligently, the way people do.

Algorithmic creativity

“In both cases, the goal is to produce solutions that are coherent, distinct, and not simply reproduced from training examples. We refer to this operational notion as algorithmic creativity.”2.2 Creativity as a Test of Visual General Intelligence

In plain terms: A measurable stand-in for real-world creativity: can a system produce outputs that make sense, differ meaningfully from each other, and aren't just copied from training data.

Combinational creativity

“Combinational creativity identifies unfamiliar connections among familiar elements, as in analogy, wordplay, and the discovery of connections between previously separate ideas.”2.2 Creativity as a Test of Visual General Intelligence

In plain terms: Creativity that comes from spotting an unexpected link between two already-familiar things, like a good analogy or pun.

Exploratory creativity

“Exploratory creativity constructs new patterns subject to a collection of rules or constraints, as in designing problems, proteins, mechanisms, or narratives.”2.2 Creativity as a Test of Visual General Intelligence

In plain terms: Creativity that comes from building something genuinely new while still obeying a set of rules, like designing a new mechanism.

Structure (of the physical world)

“It helps to be specific about what this structure of the physical world consists of.”2.8 Seeing the Physical World via Code

In plain terms: The underlying facts about a scene (its objects, their real physical properties, pose, relations, and how they change) as opposed to just what it looks like in pixels; the paper breaks this into five parts: entities, intrinsics, extrinsics, relations, and dynamics.

Method

What they actually did

Each step is a synthesis. Open any step to see the paper’s own sentence it was derived from, with its locator — so nothing here floats free of the source.

How the white paper was assembled: from a workshop, to ten independently authored perspectives, to a summary table, to a cross-cutting discussion, ending in a conclusion that treats disagreement itself as the finding.
Click any box to open it.
  1. The paper originated from a workshop discussion rather than a single research project.
    Trace this step to the paper
    “This paper originates from discussions at the CVPR 2026 VGI Workshop [85].”Footnote 1 (heading of Section 2)
  2. The editors deliberately chose to compile multiple viewpoints instead of adjudicating a single definition of VGI.
    Trace this step to the paper
    “The views presented here are not intended to provide a single definition of VGI, but rather to summarize multiple perspectives that open diverse possibilities for vision research toward general intelligence.”Footnote 1 (heading of Section 2)
  3. Ten named contributors each wrote a self-contained, first-person perspective section (2.1 through 2.10).
    Trace this step to the paper
    “2.1. Video models are visual foundation models (Robert Geirhos)”Section 2.1 heading
  4. The editors built a single table condensing each contributor's position into one comparable sentence.
    Trace this step to the paper
    “Table 1 provides a concise summary of their central positions.”3.1 Summary of Perspectives
  5. The ten individual positions were grouped into broad thematic directions for the summary discussion.
    Trace this step to the paper
    “Taken as a whole, the positions can be broadly grouped into the following directions: scaling and generative modeling; creativity, prediction, imagination, and reconstruction; continual and lifelong learning; …”3.1 Summary of Perspectives
  6. A discussion section cross-references specific contributors against each other on named open questions (can intelligence emerge from vision alone; scaling versus structure; continual learning and evaluation).
    Trace this step to the paper
    “Before asking whether visual intelligence can lead to general intelligence, a more immediate question is whether intelligence can emerge from visual experience itself.”3.2 Discussion
  7. The conclusion explicitly refuses to merge the contributors' views into one forecast, presenting the plurality itself as the paper's output.
    Trace this step to the paper
    “Rather than forcing these positions into one unified forecast, this paper presents several plausible and complementary conclusions.”4. Conclusion
The models under study

Exactly what was run, and how

What they reported — and what they left out

The paper runs no experiments of its own and reports no configuration details (temperature, deployment, reasoning effort, sampling) for any model; it only names existing systems in passing during discussion, e.g. Veo 3, GPT-2/GPT-3/ChatGPT, CLIP, Flamingo, BLIP/BLIP-2, LLaVA, SORA, DINO, JEPA, NanoBanana, and MASt3R-SLAM, without stating how any of them were run.

Results

The numbers they report

Vision evolved far earlier in the history of life than language did.

roughly half a billion years

See it in the paper
“In the history of Earth, advanced visual systems, including camera-type eyes and compound eyes, date back to the Cambrian period roughly half a billion years ago …”1. Introduction

In earlier work by one of the contributors, a video model (Veo 3) performed a wide range of visual tasks without being trained on any of them specifically.

See it in the paper
“Without any task-specific training, the model can perform a wide range of visual tasks simply through image-to-video generation: edge detection, object segmentation, keypoint localization, super-resolution, image editing, style transfer, and even early forms of visual reasoning such as maze solving and graph traversal.”2.1 Video models are visual foundation models

On controlled algorithmic-creativity tasks, multi-token/diffusion-style training produced more diverse and original outputs than standard next-token training.

See it in the paper
“On the controlled tasks, multitoken approaches based on teacherless training and diffusion generated more diverse and original solutions than conventional next-token learning.”2.2 Creativity as a Test of Visual General Intelligence

One contributor estimates vision may need orders of magnitude more effective compute and data than language pretraining before useful visual units emerge.

1,000x or even 10,000x

See it in the paper
“Some forms of visual intelligence may require several orders of magnitude more effective compute and data processing than language pre-training, perhaps 1,000x or even 10,000x in some settings, before the right abstractions emerge naturally.”2.9 From Visual Models to Vision-Native Intelligence

Real solar oscillation periods can be confused with spurious instrument artifacts unless carefully disentangled.

11 years and 5 minutes (real); 160 minutes and 24 hours (spurious)

See it in the paper
“the sun oscillates at 11 years [38] and 5 minutes, but spurious signals have been found at 160 minutes [10] and 24 hours [43].”2.5 Visual Intelligence and Discovery

The dominant AI processor has shifted from CPUs to GPUs over roughly the last decade and a half.

past 15 years

See it in the paper
“The move from CPUs to GPUs over the past 15 years as the dominant processors of AI is only the beginning of a trend.”2.6 What is the Computational Structure of Spatial AI?

A classic developmental-psychology experiment found that only kittens whose own movement controlled their visual input developed normal visually guided behavior.

See it in the paper
“In Held and Hein’s experiment, pairs of kittens received closely matched visual input, but only the kittens whose own movements controlled that input developed normal visually guided behavior [39].”2.7 Visual Intelligence is Embodied Intelligence
Claim ↔ evidence

What they assert, beside what they showed

Left is the claim in the paper’s own words. Right is the data offered for it. Where the two do not fully meet, a gold band names the distance.

The claim

Generative video models, without any special training, already function as general-purpose visual foundation models.

“Veo 3, a video model trained for entertainment, happens to be a visual foundation model.”

The evidence

“Without any task-specific training, the model can perform a wide range of visual tasks simply through image-to-video generation: edge detection, object segmentation, keypoint localization, super-resolution, image editing, style transfer, and even early forms of visual reasoning such as maze solving and graph traversal.”

2.1 Video models are visual foundation models
Mind the gap: The evidence is the author's own prior single-model demonstration on Veo 3; it supports a proof-of-concept for one system, while the claim is stated as a general property of video models.
The claim

The learning objective used matters for whether a model produces creative (diverse, original) visual outputs.

“Our results suggest that the learning objective matters for acquiring such behavior.”

The evidence

“On the controlled tasks, multitoken approaches based on teacherless training and diffusion generated more diverse and original solutions than conventional next-token learning.”

2.2 Creativity as a Test of Visual General Intelligence
Mind the gap: The evidence comes from 'minimal algorithmic tasks' the authors describe as 'loose abstractions of open-ended real-world tasks,' so the paper itself signals distance between the controlled setting and real visual creativity.
The claim

Rendering a visually convincing outcome does not show that a model has represented the physical properties (mass, contact, friction) that produced it.

“A video model that renders a convincing falling cup has shown that it knows what falling cups look like; it has not shown that mass, contact, or friction appear anywhere inside it.”

The evidence

“The gap is hard to detect when evaluation rewards perceptual quality, and becomes evident as soon as one wants to intervene, edit, or act [25, 87].”

2.8 Seeing the Physical World via Code
Mind the gap: This is presented as a conceptual/definitional argument rather than a reported experimental comparison; the cited works [25, 87] are gestured at for empirical grounding but not quoted directly.
The claim

Coupling perception with an agent's own self-generated action matters for developing genuine visual intelligence.

“Classic evidence from biological vision supports the importance of coupling perception with self-generated action.”

The evidence

“In Held and Hein’s experiment, pairs of kittens received closely matched visual input, but only the kittens whose own movements controlled that input developed normal visually guided behavior [39].”

2.7 Visual Intelligence is Embodied Intelligence
Mind the gap: The evidence is a decades-old experiment in biological kitten vision, offered as an analogy for artificial visual systems; the generalization from biological vision to machine learning is asserted, not directly tested here.
The claim

Despite their differences, the ten contributed perspectives share a common premise about what VGI is not.

“They nevertheless share an important premise: visual general intelligence should not be understood simply as a more accurate image recognition system or as a visual encoder attached to a language model.”

The evidence

“The perspectives presented in this paper approach visual general intelligence (VGI) from complementary directions. Table 1 provides a concise summary of their central positions.”

3.1 Summary of Perspectives
Mind the gap: The 'evidence' is the paper's own curated set of ten invited opinion essays rather than independent data; convergence among people convened for a workshop on this exact topic is expected and does not by itself establish the premise for the wider field.
Discussion & after

How they frame it, and what they want next

Their framing

The paper frames visual general intelligence as an open, plural research agenda rather than a settled claim, explicitly declining to pick one winning theory or offer a single definition. Contributors present their views as personal 'bets' in first person, and the synthesis in Sections 3 and 4 treats the resulting disagreement among ten experts as itself the paper's central, intended conclusion rather than an unresolved weakness.

Register: The writing is deliberately tentative throughout, dense with hedge words like 'may,' 'might,' 'perhaps,' and 'I believe,' explicit self-labeling of claims as personal 'bets,' and a closing move that presents unresolved disagreement as the paper's intended finding rather than a gap to be filled.

Where they hedge

“It may therefore be premature to ask whether VGI has already become a path to general intelligence.”4. Conclusion
“I do not know whether this is the missing ingredient in current AI systems.”2.9 From Visual Models to Vision-Native Intelligence
“Happy to be wrong, I like simplicity.”2.3 A learning system that learns from one datum (footnote 3)
“These capabilities remain fragmented, computationally expensive, limited in temporal extent, or dependent on carefully selected data and training conditions.”4. Conclusion
“I do not have an insightful answer here, but believe the right framing of the problem is that of reinforcement learning (RL).”2.4 Intelligence will be multimodal, generative, and efficient

What they say it means

  • Future VGI benchmarks will need to measure many different capacities at once rather than a single accuracy number.
    the paper’s words
    “VGI benchmarks should therefore evaluate transfer, continual adaptation, active observation, persistent world knowledge, creativity, physical consistency, and efficiency.”3.2 Discussion
  • Computer vision's central object of study may shift from building better task-specific tools to studying how intelligence itself emerges from visual experience.
    the paper’s words
    “The transition from visual foundation models to visual intelligence would therefore mark a new phase of computer vision: from constructing increasingly capable visual functions to investigating how intelligence itself may emerge from visual experience.”3.2 Discussion
  • Vision has been undervalued in current AI interfaces not because it matters less than language, but because systems aren't yet built around it as a primary channel.
    the paper’s words
    “Vision is not behind language in importance; it is behind language in current interface visibility.”2.9 From Visual Models to Vision-Native Intelligence
  • Autonomous systems such as robots or lab assistants will need to perceive the world for themselves rather than depend on people to narrate it.
    the paper’s words
    “A robot, a laboratory assistant, or a medical system cannot rely on humans to continuously describe the state of the world [20].”2.9 From Visual Models to Vision-Native Intelligence

What they call for next

  • Treat learner-driven data curation (reinforcement learning) as the right framework for vision, as it already is in robotics and language.
    the paper’s words
    “I hope the vision community will embrace it as has robotics and natural language.”2.4 Intelligence will be multimodal, generative, and efficient
  • Move visual question-answering research beyond passively answering questions about a fixed input.
    the paper’s words
    “I hope future work on VGI will move beyond static image or video question answering.”2.9 From Visual Models to Vision-Native Intelligence
  • Evaluate whether a visual system knows what it still needs to look at, not only what it can infer from what it has already been shown.
    the paper’s words
    “We should ask not only what a model can infer from a given visual input, but whether it knows what it needs to see next.”2.9 From Visual Models to Vision-Native Intelligence
  • Determine what evidence would actually convince the field that intelligence is emerging from vision, not just how to build bigger visual models.
    the paper’s words
    “The question left to the field is therefore not only how to build more capable visual models, but what evidence would convince us that intelligence has begun to arise from vision – and how that intelligence might eventually contribute to general intelligence.”4. Conclusion

Limitations they state

“These capabilities remain fragmented, computationally expensive, limited in temporal extent, or dependent on carefully selected data and training conditions.”4. Conclusion
“there remains substantial room to disentangle whether a capability arises from training data, task design, decoders, prompts, model size, or their combinations.”1. Introduction
“I do not have an insightful answer here, but believe the right framing of the problem is that of reinforcement learning (RL).”2.4 Intelligence will be multimodal, generative, and efficient
“Present-day agents lean heavily on the abstractions handed to them and on a language model’s prior knowledge of how everyday objects are built and how they work, and they degrade as those priors are withdrawn [29].”2.8 Seeing the Physical World via Code
For your own writing

Moves worth stealing

Frames speculative claims explicitly as personal 'bets' rather than settled findings, inviting disagreement instead of asserting certainty.

“I will frame them as “bets”; some may turn out to be wrong, while some may viewed as obvious.”

Signs each section with a single named contributor writing in first person, so it is clear which claims belong to whom instead of blending into one authorial voice.

“2.1. Video models are visual foundation models (Robert Geirhos)”

Uses a compact summary table to reduce ten heterogeneous, paragraph-length positions to one comparable sentence each before discussing them together.

“Table 1 provides a concise summary of their central positions.”

Explicitly declines to converge on one definition and instead argues that the disagreement itself is the paper's most important result.

“This plurality is itself one of the central conclusions of the paper.”

Closes on an open question rather than a resolved claim, deliberately leaving the field's verdict undetermined.

“The question left to the field is therefore not only how to build more capable visual models, but what evidence would convince us that intelligence has begun to arise from vision – and how that intelligence might eventually contribute to general intelligence.”
Connected

Where else this leads

Same people

Published alongside it

The nearest publications in time, across all three labs.

What this page was built from

The full paper text (abstract, Sections 1-4, acknowledgments, and reference list) was read in full; this is a discussion/position white paper synthesizing ten invited first-person perspectives rather than an empirical study, so it contains no unified experimental method, model configuration details, or self-reported statistical results in the conventional sense.