A dancer with a prosthetic leg, arched back on a reflective floor

High-end Spatial &
Human Intelligence
For Physical AI

Powering AI with embodied spatial and human data and the behavioral-realism benchmark, from environment to behavior.

THE LANDSCAPE

Physical AI Data Forces a Tradeoff

Natural Human Behavior or Reliable World State. VLGE brings both together.

Physics Simulation, Robot Teleop

Isaac, Habitat, PARTNR, Open-X, ALOHA

Exact state, but scripted, policy-driven, or mediated through a robot.

Generated Video

Cosmos, GR00T-Dreams

Scalable visuals, but actions and state are inferred, not ground truth

Human Ego-Video and Physiology

Ego4D, Ego-Exo4D, sEMG Band, Tesla ego rigs

Real humans, but occluded, one-shot, and disconnected from exact world state.

VLGE

Real human behavior + exact world state + repeatable as a distribution

Recorded from humans Recorded from robots Generated or simulated Hours (log scale) 100,00010,0001,000100 Open X-Embodiment reports 1M+ trajectories across 22 embodiments, counted in episodes rather than hours. Tesla has not disclosed hours. NVIDIA’s robot-native base is 88 hours. The 827 h above it is that same 88 h re-generated; the 6,500 h is simulation.
MARKET : WHY NOW

The Human Data Bottleneck

The bottleneck is moving from compute to real-world, spatial human data.
What’s still missing is a reliable human reference for how people actually move, decide, interact, and transact.

TAM:
$23bn

The embodied‑AI market in 2030.MarketsandMarkets SE 9427, June 2025

TRACTION
$10M+
Data order pipeline
5M+
Licensable hours a month
40K+
Human players a month
10K+
Worlds live
22 Apr 2025
Physical Intelligence
π0.5
5 Aug 2025
Google DeepMind
Genie 3
18 Feb 2026
NVIDIA
EgoScale · GR00T N1.7
4 Sep 2026
World Labs
Atlas
WHY EVERYONE HAS ONE SAMPLE

The Physical World Cannot Be Re-run.

More hours do not solve this. The problem is repeatability.

3,670 Hours
Meta Ego4D2021
1,286 Hours
Meta Ego-Exo4D2024
Fleet Scale
Tesla Ego Rigs2025
20,854 Hours
NVIDIA GR00T N1.72026

Real-world repetition ≠ controlled repetition.

Every human demonstration starts from a different world state.

When an agent behaves differently, was it the policy, or the world?

We hold the world fixed and measure the distribution of behavior.

Everyone else
One human demonstration per world state
Human1 run
A box, a mug and a plant on a table
AgentMultiple runs
VLGE
Multiple humans and the agent in the exact same world state
HumansMultiple runs
Identical world state
AgentMultiple runs
Identical world state

Capturing Human Behavior At Scale

WHAT WE HAVE

Spatial, Human Intelligence For Physical AI

Bring Consented and Branded Human And Spatial. Excellence Into The Center Of AI

For decades we built the world’s highest‑fidelity cultural, physical and digital spaces for museums, luxury houses and fashion weeks.

A digital meadow
A digital corridor
A digital pavilion of arches
The Louvre pyramid at night, rebuilt digitally
A digital courtyard
A pink house in a digital garden
The VLGE world-building interface

BUILDING OF URL WORLDS

How worlds/spaces are built.

A fashion show set under a glass roof
A department store dome
A luxury storefront
A high-rise apartment interior

Capturing real world

Diverse, Unique and Rare spaces at scale.

A digital city street

HOW PEOPLE MOVE

Games, Shopping, Events, Social.

A person wearing a sensor cap

human talent

Real Bodies, Instrumented, In Real Spaces,

WHAT WE BUILT

Humans Act Inside A World We Own.

The moat is state access, not data volume.

State is read out, not estimated

Joints, contacts, and object poses come from the engine, with no label estimation.

The scenario is re-instantiable

Same scene. Same goal. Different person. As many times as needed.

Every signal shares one clock

Trajectory, hesitation, dwell, decision and attention are frame-aligned in a single schema.

~400KHours

of human play, every month

Human Build and Play
A person building and playing inside a VLGE world
trainshuman play trains the agent
scales ×Nre-runs, keeps human-like data
Agent Build and Play
Prompting
A person typing a prompt on a laptop
The VLGE agent executing inside the same world
COLLECTING BEHAVIOR DISTRIBUTION

One Scenario: Many Human Behaviors

One trajectory cannot capture how people actually behave.
Averaging human behavior can erase the different strategies humans actually use.

One Context: c

Many humans move through it. The result is not a path, it is a distribution over behaviors.

𝒮(π) = 𝔼c [ D( P(φ(Bπ) | c) ‖ P(φ(H) | c) ) ]

ONE HUMAN
SESSION
one clock
one schema
frame-aligned
01Locomotion & explorationDoes it move through a space like a human?
02ManipulationDoes it grasp — and fail — like a human?
03Attention & intentDoes it look where a human looks, before acting?
04TransferDoes a policy trained on this improve a robot?
Pose Streams
3D joints
rotation (6D / quat)
timestamps
Action Streams
Discretized actions
(token ids)
Hand Interaction Streams
Contact, grasp,
object states
(event list)
Frame-Level Feature Map
Compute low-level motion and interaction features per frame.
F ∈ ℝT×6
Session-Level Feature Map
Aggregate frame features into higher-level behavioral features over the session.
Kinematic(12)
Action(13)
Temporal(8)
Exploration(8)
Manipulation(20)
Progression(18)
ϕ ∈ ℝ79
Session Embedding
Normalize to obtain a fixed-length session representation.
z1
z2
z3
z4
⋮
z78
z79
z = ϕ − μHσH

Evaluating Agent Behavior

THE QUESTION

Did The Agent Behave Like A Human?

To answer it we need P(behaviour | scenario):
The distribution of what people actually do in the same scenario, not one reference trajectory.

The Human Reference

Same scenario. Many people.
Trajectory, hesitation, decisions, attention.

P(ϕ(H) | c)

The Agent

Run repeatedly in the matched scenario
Measure the behavioral signals.

P(ϕ(Bπ) | c)

Distributional Distance

Compare the two
behavioral distributions

D(P(ϕ(Bπ) | c) ‖ P(ϕ(H) | c))
Isometric apartment with many human trajectories, dwell points, object interactions and hesitations in blue
The same apartment with the agent's repeated trajectories, dwell points, interactions and hesitations in orange
FIRST RESULTS

Agent Distance From Human Behavior

Measured relative to natural human-to-human variation.

THE BENCHMARK

The VLGE Human-Grounded Benchmark

Trained VLGE’s agents on real human behavior, used as a reference for AI labs to measure their own agents or world models against it.

01
GROUND TRUTH
Real Human CaptureTrajectory, EMG, EEG
02
THE REFERENCE
VLGE Human-Grounded AgentsThe Real-Human Distribution
03
ANY LAB SUBMITS
Their Agents, world modelsPolicy, IL, or world
04
OUTPUTS
Behavioral-realism scoreS(Π) + Per-signal gap
ContenderOptimizes forTask successBehavioral realism S(Π) ↓Reads as
Real humansreferenceBeing human—
0.00 · floor
The standard
VLGE human-grounded agentHuman behavior✓
low
Indistinguishable
Imitation / task policyTask completion✓
medium
Succeeds, moves unhuman
World-model agentVisual / physical plausibility✓
high
Looks right, behaves off
Scripted / random—✗
highest
Non-human
PERMUTATION

Optimize The Behavior That Matters

Optimize your agent in any dimensions.

Modular Behavioral Evaluation
Example 1: Robotics Lab
Example 2: VLA / World Model Lab
Example 3: General Agent Evaluation
BENCHMARK COMPARISON
Human role
Human reference
Score type
Score resolution
Validity evidence
Frontier labs
NVIDIA, Google DeepMind, Meta
SIMA 2 · PARTNR · Meta Motivo preference study · Ego-Exo4D proficiency · WAGIBench · Cosmos world models · RoboLab / Isaac Lab-Arena · Gemini Robotics 1.5 evals · ASIMOV 2.0
Repeated rollouts1 of 9
Yes1 of 9
Preference / Elo2 of 9
Trajectory4 of 9
Agreement with humans2 of 9
Human-likeness studies
HumanTracker (HumanScore) · Motion Turing Test
Repeated rollouts1 of 2
Partial
Scalar score
Trajectory
Agreement with humans
Real-robot and world-model evaluation
RoboArena · RBench / RoVid-X · AutoEval · RoboWorld · SimplerEnv
Human raters2 of 5
Partial1 of 5
Preference / Elo1 of 5
Trajectory1 of 5
Agreement with humans2 of 5
Robot companies
1X, AgiBot, X-Humanoid
EWMBench · RoboMIND · AgiBot World Challenge · 1XWM
Human raters1 of 4
No
Fidelity score2 of 4
Trajectory2 of 4
Agreement with humans1 of 4
Academic simulation suites
RoboCasa · RoboTwin 2.0 · CALVIN · VLABench · EmbodiedBench · ALFRED / TEACh / DialFRED · LIBERO · BEHAVIOR-1K · GenManip · GRUtopia · Habitat navigation suite · HumanoidBench
Training demos5 of 12
No
Scalar score3 of 12
Sub-task / phase4 of 12
Correlation with real2 of 12
Standards and governance
EmbodiedGovBench · EAI Bench (YD/T 6770-2026)
None
No
Composite index
Sub-task / phase1 of 2
Stated, not measured1 of 2
VLGE
Repeated rollouts
Yes
Distribution distance
Micro-behaviour
None
Full Table Link

Train Agents with Expert Behavior

THE REFERENCE LADDER

Human Is Not The Ceiling.

T1

Ordinary human

in the engine
“Is it human?”
First-person view: hands holding a pot of soup in a simulated kitchen
30 people · same task · natural variation
T2

Expert human

in the engine
“How does an expert do it?”
First-person view: a chef's hands holding a pot of soup in the simulated kitchen
An expert cohort, same task,
higher skill ceiling
T3

Sensory-rich human

in the real world
“What are they sensing and intending?”
EEG
166 Hz
Gaze
166 Hz
sEMG
166 Hz
Gaze · EEG · sEMG · 166 Hz
T4

Digital twin

same task, both worlds
“How well does it transfer?”
Simulation
Simulated kitchen, hands holding a pot
Real (Ego-cam)
Real kitchen from an egocentric camera, hands holding a pot
Identical task · Parallel data
Real-to-sim alignment
A Distance
A Position On A Scale
PRIME AGENTS

From Experts, To Prime Agents.

A chef cracking an egg into a bowl, with hand landmarks traced
The chef's eyes, close up
The chef's hands, close up
EYE GAZE &
ATTENTION
HANDS &
FINGERS
AUDIO
BRAIN
ACTIVITY
Transfer Skills From
Expert to Agents
A humanoid robot taking a soufflé out of an oven
Soufflé
Scrambled eggs on toast
Eggs Benedict
Omelette
Omelette with salad and a coffee

EXPERT TALENT

World-class human expertise
ChefsCulinary techniques
AthletesMovement & performance
ArtistsCreative process
DesignersSpatial & product design
PerformersMusic, dance & expression
SpecialistsDomain-specific skills
REAL EXPERTS. REAL SKILLS.
REAL-WORLD SCENARIOS.
Portrait of an expert in a blue shirt
Portrait of a chef in whites
Portrait of an artist
Portrait of a designer in tortoiseshell glasses
Portrait of an athlete in a blazer
Portrait of a performer

Perpetually Licensable
Expertise

EXPERTISE CAPTURE

Several Domains, One Schema.

The world is instrumented, not the person. Every stream lands on one clock.

415
Fields in a single hand recordBeyond x, y, z: grip, contact, object state, and outcome.
9
Streams, one clockBody, hands, gaze, actions, sensors, scene: Frame-aligned.
0
Pipeline changes per new domainChef or climber, the schema stays the same.
TRAINData licensed for training
BENCHThe subset with re-instantiable world state that can score an agent
STREAM
CHEF
SPORT
USED FOR
INTEGRATION

Your Model. Your World.

Two adapters is the whole integration. Your weights never leave your infrastructure.

VLGE 3D ENGINE
A robot hand reaching for a mug on a wooden table, rendered in the VLGE engine
RGB
RGB view
Depth
Depth view
Segmentation
Segmentation view
State
State view: object poses and hand skeleton
World state exactly re-instantiated(from model's predicted action)
Observation Packet (from VLGE)
RGB
Depth
Segmentation
State
Standardized observation packet(VLGE format)
Predicted Action (from model)
Model's action representation(e.g., joint commands, end-effector Δpose, discrete action)
↓
Mapped to VLGE action space(standardized execution format)
CUSTOMER MODEL
World Model(e.g., video-based, diffusion, ...)
VLA / Policy(e.g., OpenVLA, RT-2, ...)
Custom Stack(runs in your infrastructure)

Full-episode loop (default)Repeat until the scenario ends
Step-by-step API (optional)Single-step interaction
Closed-loop execution over full episodes (N rollouts)Repeat for different variations (environments, objects, tasks, agents, etc.)
WHAT IT PRODUCES
Human (the floor)
Your agent (N rollouts)
Expert (the ceiling)

Platform Summary

PROPRIETARY HUMAN DATA

Expert skill captured as multimodal data
Body MotionFull-body kinematics
HandsFinger motion + manipulation
ContactTouch + force
GazeAttention & visual intent
Cognitive SignalsDecision states + errors + retries
EnvironmentScene + object context
MULTIMODAL. HIGH-FIDELITY.
REAL-WORLD CONTEXT.

EXPERT IP

Human expertise becomes
structured, reusable intelligence

CaptureMulti-modal
fusion
StructureBehavior
representation
LearnSkill
modeling
FROM HUMAN PERFORMANCE
TO MACHINE CAPABILITY.

LICENSABLE INTELLIGENCE

One performance. Reusable in perpetuity.
LICENSEProprietary training data
for AI & robotics
BENCHMARKMeasure AI vs. expert
human behavior
TRAINTask-specific agents
on expert skills
SCALEExpert skills across
machines and environments
IMMORTAL SKILLS.
REAL-WORLD IMPACT.

The Future Of AGI Will Be
Shaped By Human Experience.