High-end Spatial & Human Intelligence For Physical AI
Powering AI with embodied spatial and human data and the behavioral-realism benchmark, from environment to behavior.
THE LANDSCAPE
Physical AI Data Forces a Tradeoff
Natural Human Behavior or Reliable World State. VLGE brings both together.
Physics Simulation, Robot Teleop
Isaac, Habitat, PARTNR, Open-X, ALOHA
Exact state, but scripted, policy-driven, or mediated through a robot.
Generated Video
Cosmos, GR00T-Dreams
Scalable visuals, but actions and state are inferred, not ground truth
Human Ego-Video and Physiology
Ego4D, Ego-Exo4D, sEMG Band, Tesla ego rigs
Real humans, but occluded, one-shot, and disconnected from exact world state.
VLGE
Real human behavior + exact world state + repeatable as a distribution
MARKET : WHY NOW
The Human Data Bottleneck
The bottleneck is moving from compute to real-world, spatial human data. What’s still missing is a reliable human reference for how people actually move, decide, interact, and transact.
TAM:
$23bn
The embodied‑AI market in 2030.MarketsandMarkets SE 9427, June 2025
TRACTION
$10M+
Data order pipeline
5M+
Licensable hours a month
40K+
Human players a month
10K+
Worlds live
22 Apr 2025
Physical Intelligence
π0.5
5 Aug 2025
Google DeepMind
Genie 3
18 Feb 2026
NVIDIA
EgoScale · GR00T N1.7
4 Sep 2026
World Labs
Atlas
WHY EVERYONE HAS ONE SAMPLE
The Physical World Cannot Be Re-run.
More hours do not solve this. The problem is repeatability.
3,670 Hours
Meta Ego4D2021
1,286 Hours
Meta Ego-Exo4D2024
Fleet Scale
Tesla Ego Rigs2025
20,854 Hours
NVIDIA GR00T N1.72026
Real-world repetition ≠ controlled repetition.
Every human demonstration starts from a different world state.
When an agent behaves differently, was it the policy, or the world?
We hold the world fixed and measure the distribution of behavior.
Everyone else
One human demonstration per world state
Human1 run
AgentMultiple runs
VLGE
Multiple humans and the agent in the exact same world state
HumansMultiple runs
Identical world state
AgentMultiple runs
Identical world state
Capturing Human Behavior At Scale
WHAT WE HAVE
Spatial, Human Intelligence For Physical AI
Bring Consented and Branded Human And Spatial. Excellence Into The Center Of AI
For decades we built the world’s highest‑fidelity cultural, physical and digital spaces for museums, luxury houses and fashion weeks.
BUILDING OF URL WORLDS
How worlds/spaces are built.
Capturing real world
Diverse, Unique and Rare spaces at scale.
HOW PEOPLE MOVE
Games, Shopping, Events, Social.
human talent
Real Bodies, Instrumented, In Real Spaces,
WHAT WE BUILT
Humans Act Inside A World We Own.
The moat is state access, not data volume.
State is read out, not estimated
Joints, contacts, and object poses come from the engine, with no label estimation.
The scenario is re-instantiable
Same scene. Same goal. Different person. As many times as needed.
Every signal shares one clock
Trajectory, hesitation, dwell, decision and attention are frame-aligned in a single schema.
~400KHours
of human play, every month
Human Build and Play
trainshuman play trains the agent
scales ×Nre-runs, keeps human-like data
Agent Build and Play
Prompting
COLLECTING BEHAVIOR DISTRIBUTION
One Scenario: Many Human Behaviors
One trajectory cannot capture how people actually behave. Averaging human behavior can erase the different strategies humans actually use.
One Context: c
Many humans move through it. The result is not a path, it is a distribution over behaviors.
𝒮(π) = 𝔼c [ D( P(φ(Bπ) | c) ‖ P(φ(H) | c) ) ]
ONE HUMAN SESSIONone clock one schema frame-aligned
01Locomotion & explorationDoes it move through a space like a human?
02ManipulationDoes it grasp — and fail — like a human?
03Attention & intentDoes it look where a human looks, before acting?
04TransferDoes a policy trained on this improve a robot?
Pose Streams
3D joints rotation (6D / quat) timestamps
Action Streams
Discretized actions (token ids)
Hand Interaction Streams
Contact, grasp, object states (event list)
Frame-Level Feature Map
Compute low-level motion and interaction features per frame.
F ∈ ℝT×6
Session-Level Feature Map
Aggregate frame features into higher-level behavioral features over the session.
Kinematic(12)
Action(13)
Temporal(8)
Exploration(8)
Manipulation(20)
Progression(18)
ϕ ∈ ℝ79
Session Embedding
Normalize to obtain a fixed-length session representation.
z1
z2
z3
z4
⋮
z78
z79
z = ϕ − μHσH
Evaluating Agent Behavior
THE QUESTION
Did The Agent Behave Like A Human?
To answer it we need P(behaviour | scenario): The distribution of what people actually do in the same scenario, not one reference trajectory.
The Human Reference
Same scenario. Many people. Trajectory, hesitation, decisions, attention.
P(ϕ(H) | c)
The Agent
Run repeatedly in the matched scenario Measure the behavioral signals.
P(ϕ(Bπ) | c)
Distributional Distance
Compare the two behavioral distributions
D(P(ϕ(Bπ) | c) ‖ P(ϕ(H) | c))
FIRST RESULTS
Agent Distance From Human Behavior
Measured relative to natural human-to-human variation.
THE BENCHMARK
The VLGE Human-Grounded Benchmark
Trained VLGE’s agents on real human behavior, used as a reference for AI labs to measure their own agents or world models against it.
01
GROUND TRUTH
Real Human CaptureTrajectory, EMG, EEG
02
THE REFERENCE
VLGE Human-Grounded AgentsThe Real-Human Distribution
03
ANY LAB SUBMITS
Their Agents, world modelsPolicy, IL, or world
04
OUTPUTS
Behavioral-realism scoreS(Π) + Per-signal gap
Contender
Optimizes for
Task success
Behavioral realism S(Π) ↓
Reads as
Real humansreference
Being human
—
0.00 · floor
The standard
VLGE human-grounded agent
Human behavior
✓
low
Indistinguishable
Imitation / task policy
Task completion
✓
medium
Succeeds, moves unhuman
World-model agent
Visual / physical plausibility
✓
high
Looks right, behaves off
Scripted / random
—
✗
highest
Non-human
PERMUTATION
Optimize The Behavior That Matters
Optimize your agent in any dimensions.
Modular Behavioral Evaluation
Example 1: Robotics Lab
Example 2: VLA / World Model Lab
Example 3: General Agent Evaluation
BENCHMARK COMPARISON
Human role
Human reference
Score type
Score resolution
Validity evidence
Frontier labs
NVIDIA, Google DeepMind, Meta
SIMA 2 · PARTNR · Meta Motivo preference study · Ego-Exo4D proficiency · WAGIBench · Cosmos world models · RoboLab / Isaac Lab-Arena · Gemini Robotics 1.5 evals · ASIMOV 2.0