Skip to the lesson
Robotics Lab Explore Robotics / Physical AI / Theory / T5

Physical AI · Theory

T5 — World Models and Symbolic State

How a robot turns observations into a structured representation of the world.

  • Lesson 5
  • Physical AI Theory
  • Representation
  • About 11 minutes to read

In T4 we ended with observations. This lesson asks what the robot does with them: how it holds what it currently believes about the world, in a form it can reason over.

Learning objectives

After this lesson, you should be able to:

  • Explain why a Physical AI system needs a world model.
  • Represent a simple environment using objects, properties and relationships.
  • Explain the difference between an observation and a world-model fact.
  • Understand predicates and symbolic state.
  • Distinguish robot state, environment state and task state.
  • Explain how observations update a world model.
  • Explain why the world model can be incomplete or wrong.
  • Explain how planners use world-model information.

The one-sentence idea

A world model is the robot’s structured, continuously updated representation of the parts of the world that matter for its decisions.

Structured
Organized into things a program can query, not a pile of readings. Example: “which object is red?”
Continuously updated
Revised as the robot observes and acts. Example: the bottle moves, and the model follows.
Representation
A stand-in for the world, not the world itself. Example: a fact that says the bottle is on the table.
Relevant
Limited to what the task needs. Example: the bottle’s location, not the wallpaper.
Decisions
The purpose of the whole thing. Example: whether to grasp now or look again.

The robot does not need to represent every detail of reality. It needs the information required for its current tasks, and no more.

Why a robot needs a world model

Return to the recurring request: “Bring the red bottle from the kitchen to me.” The robot cannot reason about this from raw sensor data. Pixels and ranges are too low-level and fragmented for task reasoning. What it needs looks more like this:

robot_location = kitchen
red_bottle = detected
red_bottle_location = table
user_location = living_room
path_to_table = available
gripper = empty

Each line is something a planner can use directly. A world model is the layer that turns fragmentary observations into that structured form.

  1. Physical world
  2. Sensors
  3. Observations
  4. World model
  5. Planner
  6. Action
  7. Physical world
The world model sits between observations and the planner, and the action changes the world it describes.

What is a world model?

A world model is an internal representation of the relevant entities, properties, relationships and state in the environment. Systems implement it in many ways:

  • Structured objects
  • Symbolic predicates
  • Scene graphs
  • Occupancy grids
  • Geometric maps
  • Semantic maps
  • Learned representations
  • Combinations

No one of these is universally correct, and real robots often combine several: a grid for free space, symbols for objects and tasks. This lesson focuses on a simple symbolic world model, because it makes the ideas easy to see and to reason about.

What goes into a world model

Four kinds of content appear again and again:

Objects

  • robot
  • red_bottle
  • table
  • door
  • user

Properties

  • color = red
  • movable = true
  • open = false

Relationships

  • on(red_bottle, table)
  • near(robot, table)
  • inside(robot, kitchen)

State

  • gripper_empty
  • door_closed
  • bottle_grasped
  • task_complete
A world model: objects, their properties, the relationships between them, and state, including the state of the task.

Objects and properties

A robot usually represents things as entities with attributes. Two examples:

red_bottle
color: red · type: bottle · movable: true · graspable: true
blue_cup
color: blue · type: cup · movable: true

This abstraction matters because the planner no longer needs camera pixels to ask “which object is red?” It queries the world model and gets red_bottle. The perception work happened earlier; the reasoning works with the result.

Relationships between objects

Objects alone are not enough, because robot tasks are relational. Some typical relationships:

on(red_bottle, table)
inside(robot, kitchen)
near(red_bottle, blue_cup)
next_to(table, wall)
holding(robot, red_bottle)

Consider the tasks “pick up the bottle on the table,” “move the box next to the shelf,” and “bring the object inside the room.” Each depends on a relationship, not just an identity. Relationships chain together into a small graph, sometimes called a scene graph:

  1. red_bottle
  2. on
  3. table
  4. inside
  5. kitchen
A tiny scene graph. From it the robot can infer that the bottle is in the kitchen.

Facts and predicates

A simple way to write all of this down is with predicates. A predicate is a template for a property or relationship, with slots for objects: predicate(object1, object2). Filled in, it becomes a statement: on(red_bottle, table), holding(robot, red_bottle), open(door).

Predicate
A relationship or property template, such as on(?, ?).
Fact
A predicate that is currently believed to be true, such as on(red_bottle, table).
State
The collection of relevant facts describing the current situation.

So a world state can be a set of facts:

{ inside(robot, kitchen),
  on(red_bottle, table),
  open(door),
  empty(gripper) }

Notice the careful wording: a fact is believed true, not known true. We come back to that. We stay informal here and do not go into formal logic.

Robot state and environment state

Robot state
Robot pose, velocity, gripper state, battery, joint state.
Environment state
Object locations, obstacles, doors, room states, object relationships.

Practical systems often combine these, along with task state (next section), into one task-relevant representation:

  • Robot state
  • Environment state
  • Task state
  1. Physical AI world model
Three kinds of state combine into one world model.

Task state

The world model also records progress. For the goal “bring the red bottle to the user,” the task state might be:

goal = bring(red_bottle, user)
bottle_found = true
robot_at_bottle = true
bottle_grasped = false
robot_at_user = false
task_complete = false

Task state answers a different question from physical state: where are we in the process of completing the goal? It is what lets a robot resume sensibly after a failure, instead of starting from scratch. Later lessons on planning and behavior trees use it heavily.

From observations to world state

This is one of the most important steps. An observation does not become a fact automatically; it has to be interpreted.

  1. Sensors
  2. Measurements
  3. Observations
  4. Interpretation
  5. World model update
  6. New belief / state
  7. Planning
From sensors through interpretation to a world-model update, which feeds planning.
Observations and the world-model updates they may cause
ObservationPossible world-model update
Camera: red bottle detected on the tableon(red_bottle, table)
Robot has entered the kitcheninside(robot, kitchen)
Gripper sensor indicates an object is heldholding(robot, red_bottle)

The word “possible” is deliberate. The system decides whether the evidence is good enough to update the model. And this never stops: the robot is not building the world model once and then forgetting about it.

A world model is not ground truth

This connects straight back to T2: the world model belongs on the belief side of the line, not the truth side. Say the bottle really is on the table, and the robot’s model says so too. Then someone moves it to the shelf. The truth is now on(red_bottle, shelf), but the model still says on(red_bottle, table) until a new observation arrives. A world model can become stale.

True world

The bottle is on the shelf.

Robot’s world model

The bottle is on the table.

The true world ≠ the robot’s world model, because of:

  • Noise
  • Missing observations
  • Incorrect perception
  • Stale information
  • Environmental changes
The model and the world can disagree, for several ordinary reasons.

Errors enter from perception mistakes, stale observations, wrong inference, failed actions and a changing environment. None of them is exotic. This prepares the ground for later lessons on checking whether a plan is still valid.

Uncertainty in a world model

Not every fact deserves equal trust. Compare red_bottle_location = table with high confidence against the same fact with low confidence, or a model holding competing hypotheses: 70% table, 30% shelf.

A simple symbolic system may store definite facts. More capable ones attach extra information to each fact:

  • Confidence
  • Probability
  • Timestamp
  • Source
  • Uncertainty

A timestamp tells you how old a fact is, and a source tells you whether it came from a camera or an assumption. No Bayesian mathematics is needed for the key idea: the robot’s world model can contain uncertain knowledge.

Dynamic world models

The world changes: people move, doors open and close, objects are carried away, obstacles appear, the robot itself changes location, and a successful grasp changes the state of the object. So the model is a sequence, not a snapshot.

  1. World model at t0
  2. Observation or action
  3. World model at t1
  4. Observation or action
  5. World model at t2
Each observation or action produces the next version of the world model.

A Physical AI system must keep reconciling its internal model with new observations, and with the effects of its own actions.

One complete example

Take “Bring the red bottle from the kitchen to the user.” The initial world model:

inside(robot, kitchen)
on(red_bottle, table)
inside(user, living_room)
empty(gripper)
door_open(kitchen_door)

The goal is delivered(red_bottle, user). Watch the model change:

How the world model changes during the red-bottle task
EventWorld-model update
Observation: the robot sees the red bottleNo major change; confidence may rise
Action: navigate_to(red_bottle)near(robot, red_bottle)
Action: grasp(red_bottle), then verification succeedsholding(robot, red_bottle); empty(gripper) removed
Action: navigate_to(user)near(robot, user)
Action: give the bottle to the userdelivered(red_bottle, user); task_complete = true

Notice that holding is added only after the grasp is verified, not when the command returns. That is the lesson from Lab 09, applied to the world model. This is a simplified, conceptual picture; real systems keep much richer state.

World models and planning

A planner keeps asking questions about the state. Can I grasp the object? Is it reachable? Is the door open? Where is the robot? Which actions have already happened? Every answer comes from the world model, and each action’s preconditions (what must be true to do it) are checked against it.

  1. World model
  2. Preconditions
  3. Planner
  4. Action
  5. Observation
  6. World model update
The planner checks preconditions against the world model, acts, and the next observation updates the model.

Planning operates on a representation of the world, not directly on raw sensor measurements.

Planning algorithms themselves come in later lessons.

World models in Physical AI

Place the world model on the stack from T3:

  1. Sensors
  2. Observations
  3. World model / belief
  4. Task planning
  5. Skills
  6. Motion
  7. Control
  8. Robot
  9. Sensors
The world model is the bridge in the middle of the stack.

The world model is the bridge between what the robot observes and what it decides to do. That is why it is central to Physical AI.

Common misconceptions

Misconception 1 “The world model is a copy of the real world.”

Correction It is a task-relevant representation, and it can be incomplete, uncertain or wrong.

Misconception 2 “The world model is created once.”

Correction It must be continuously updated as the robot observes and acts.

Misconception 3 “Every fact in the world model is guaranteed to be true.”

Correction World-model facts are beliefs based on observations and inference.

Misconception 4 “Objects are enough; relationships aren’t important.”

Correction Many robot tasks depend on relationships such as on, inside, near, holding and next_to.

Misconception 5 “The planner can work directly from camera images.”

Correction Some modern systems can reason directly from rich perceptual representations, but explicit or implicit world representations are still needed for many planning and decision-making tasks.

Engineering takeaways

  1. A robot needs a structured representation of its relevant environment.
  2. A world model represents objects, properties, relationships and state.
  3. Observations are used to create and update the world model.
  4. Symbolic predicates provide a simple way to represent facts.
  5. Robot state, environment state and task state can all matter.
  6. The world model is not ground truth.
  7. World-model information can be uncertain or stale.
  8. Dynamic environments require continuous updates.
  9. Planning depends on the robot’s current world model.
  10. A world model connects perception to reasoning and action.

The robot does not plan directly from reality. It plans from its current representation of reality.

Knowledge check

Three conceptual questions. Write an answer, then reveal the explanation. Your answers stay in your browser.

Question 1

A camera detects a red bottle on a table. What kind of information could the world model store from this observation?

Explanation

A fact such as on(red_bottle, table), along with relevant object properties and possibly confidence and source information. The property side might record the bottle’s color, type and that it is movable. The extra information says how much to trust the fact: a confidence, a timestamp, and that it came from the camera.

Question 2

The robot’s world model says on(red_bottle, table), but a human has moved the bottle to a shelf. Is the world model necessarily correct?

Explanation

No. It may be stale until new observations update it. A world-model fact is a belief based on past observations, not a guarantee. Until the robot looks again, its model and the world disagree, and a plan built on the old fact may fail.

Question 3

Why does a planner need a world model?

Explanation

It needs a representation of the current relevant state to determine which actions are possible and which sequence can achieve the goal. Each action has preconditions, such as an empty gripper before a grasp, and those can only be checked against a representation of the world. Raw sensor measurements do not answer such questions directly.

What’s next

We now have a representation of the world. But what happens when the robot cannot fully observe that world?

Next lesson

T6 — Belief Under Uncertainty

T6 looks at incomplete information, competing hypotheses, confidence, how beliefs are updated, and how a robot decides what to do when it cannot be sure.

Nothing below is required to finish this lesson.