Skip to the lesson
Robotics Lab Explore Robotics / Physical AI / Theory / T2

Physical AI · Theory

T2 — State, Observation and Belief

Your robot never sees the world directly. It sees measurements and builds an understanding from them.

  • Lesson 2
  • Physical AI Theory
  • Builds on T1
  • About 15 minutes to read

In T1, we saw that a physical agent cannot directly access the complete state of the world. This lesson makes that idea precise.

Learning objectives

After this lesson, you should be able to:

  • Define state in a robotics system.
  • Distinguish world state from robot state.
  • Explain what an observation is, and why it is not the same as state.
  • Distinguish true state from estimated state.
  • Explain partial observability.
  • Explain uncertainty and belief.
  • Describe the basic idea of state estimation.
  • Explain why these concepts matter for Physical AI agents.

The central skill is keeping five terms apart: true state, sensor measurement, observation, state estimate and belief. They sound alike and they are different.

The one-sentence idea

A robot acts on what it believes about the world, not on perfect access to the world itself.

The world has a definite condition, but the robot can only read it through sensors. Those readings are turned into an estimate, the estimate carries some doubt, and that estimate, with its doubt, is what the robot uses to decide and act. The chain looks like this:

  1. True world stateHow things actually are. Never directly available.
  2. SensorsTurn physical quantities into measurements.
  3. ObservationsWhat the robot receives: readings and detections.
  4. State estimationCombines observations with earlier knowledge.
  5. BeliefWhat the robot treats as plausible, and how strongly.
  6. DecisionChooses what to do from the belief.
  7. ActionChanges the world, and the chain starts again.
From the true world state to action. The robot never touches the first box, only the ones after the sensors.

This is a conceptual model, not necessarily a literal software pipeline. Some systems merge these steps and some split them further. The distinctions still hold, because they describe what is known, not how the code is organized.

What is state?

State is the information needed to describe the relevant condition of a system at a given time. For a robot, it is the answer to “what do I need to know right now to decide what to do?” The word “relevant” matters, because the useful state depends on the problem:

A mobile robot driving around
Position x, y, heading θ and velocity.
A manipulation task
Robot pose, joint angles, gripper state, object poses and how objects relate to each other.
A task-level agent
Robot location, objects, task progress, what it is holding and which actions have completed.

There is no single universal state representation. The useful state is the state required to make the decisions we care about. A robot choosing wheel speeds needs velocity; a robot choosing which room to visit does not.

Robot state vs world state

World state is everything relevant that is actually true in the environment. Robot state is the relevant internal physical condition of the robot itself: pose, velocity, joint positions, gripper width. Robot state is a part of the world state, but it is useful to name it separately because robots sense themselves differently from how they sense everything else.

robot = kitchen
red_box = table
blue_box = shelf
door = closed
gripper = empty

A Physical AI system usually needs a broader representation that holds both: where the robot is and what it is doing, plus what is around it. The diagram shows where that representation comes from, and the catch.

  1. World state
    • Robot
    • Objects
    • Environment
    • Relations
  2. Sensorsa narrow, noisy view
  3. Observations
    • Camera
    • LiDAR
    • IMU
    • Encoders
The robot does not receive the world state itself. It receives only what its sensors report. (LiDAR is a laser scanner that measures distance; an IMU, or inertial measurement unit, measures acceleration and rotation rate; encoders count wheel or joint rotation.)

What is an observation?

An observation is information the robot obtains through its sensors or other available measurements. It is evidence about the world. Different sensors produce very different kinds of evidence:

What each sensor produces
SensorWhat it produces
CameraPixels
LiDARRanges, or a point cloud
IMUAcceleration and angular velocity
Wheel encodersWheel rotation
Force sensorMeasured force
GPS (satellite positioning)A position measurement
Gripper sensorWidth, force or contact

An observation is a measurement. It is not automatically a complete explanation of the world. A note on vocabulary: a raw measurement is what the sensor outputs (the pixels), while a processed observation is what software makes of it (“red box detected”). People use “observation” for both. Later lessons follow the path from one to the other; here, what matters is that either way it is evidence, not the state.

State vs observation

This is one of the most important distinctions in the lesson. Take a red box.

True state

  • red_box = (2.0, 1.5)
  • red_box = on_table
  • red_box = not_in_gripper

Camera observation

  • red_box detected
  • pixel position = (421, 237)
  • confidence = 0.82

The observation does not equal the state. A pixel position is not a position in the room, and a detection with 0.82 confidence is not a guarantee that the box exists. Whether the box is on the table or in the gripper is not stated at all. A second example: the true distance to an obstacle is 3.2 m, and the LiDAR measures 3.1 m.

State ≠ measurement.

Measurements provide evidence about the state. The robot has to infer the state from the evidence.

Example: robot and red box

Return to the scenario from T1. A robot stands in the kitchen. A red box is on the table, a blue bottle is on a shelf, and the target area is in another room. Its camera detects red_box with confidence 0.91. Compare what is true with what the robot has:

What is true compared with what the robot has
QuestionTrue stateWhat the robot has
Where is the box?(2.00, 1.50)An estimate such as (2.04, 1.46)
Is it on the table?YesA detection with confidence 0.91
Is it graspable?Depends on its shape and what is around itNot known yet
Is something hiding part of it?PerhapsOnly a hint, from a partial outline
Is the gripper empty?YesA gripper sensor reading

The robot combines observations over time and keeps an internal representation. That is the difference between two questions: what is true? and what does the robot currently know? A Physical AI agent can only act on the second.

True state vs estimated state

The true state is not available directly. The robot gets observations, and from them it computes an estimated state, its current best guess at the relevant state.

true robot position:       x = 5.00 m
estimated robot position:  x ≈ 4.86 m

The estimate is off by 14 cm, and the robot does not know that. The estimate can be wrong. That is not necessarily a software bug; it is a fundamental consequence of imperfect sensing and inference. The goal of engineering is not to make estimates perfect but to make them good enough, and to know how far to trust them.

Why sensors can’t reveal everything

Every sensor gives only a particular view of reality, and each view has blind spots:

Typical limits of common sensors
SensorTypical limits
CameraOcclusion, lighting, depth ambiguity, limited field of view
LiDARLimited range, occlusion, sparse measurements, reflective or light-absorbing surfaces
IMUBias and drift; errors build up when acceleration is integrated over time
Wheel odometryWheel slip, uneven surfaces, encoder errors
GPSReflections off buildings, poor availability indoors, limited accuracy

None of these is a defect to be engineered away. A camera cannot see through a box and a LiDAR cannot measure past its range. These are the limits of the measurement itself.

Hidden state

Hidden state is relevant information about the world that cannot be directly observed at the current moment. The robot sees a box. It may not directly know whether the box is empty, heavy or fragile, whether it is attached to something, or even whether it is where the detector says it is.

Hidden state still matters. A heavy box changes how the arm should lift, and a fragile one changes how hard the gripper should squeeze. The robot cannot read these off the sensors, so it must infer them, test them by acting carefully, or accept the risk. Much of Physical AI decision making is deciding what to do about what the robot cannot see.

Uncertainty

Uncertainty is the robot’s doubt about the state: how far the estimate might be from the truth. Suppose the robot detects a red bottle once with confidence 0.93 and once with confidence 0.51. It should not treat those two situations identically. Uncertainty comes from several places:

  • sensor noise,
  • ambiguous observations (is that a bottle or a vase?),
  • occlusion,
  • model error (the robot’s idea of how it moves is slightly wrong),
  • environmental change,
  • unknown objects, and
  • delayed measurements.

The point of tracking uncertainty is that it should change decisions. With high confidence, the robot can act. With low confidence, it can inspect again, gather another observation, or ask for clarification. One caution: a detector’s “confidence” score is often not a true probability, so treat it as a useful signal, not as a precise measure. Later lessons build the act, inspect or ask decision on this foundation.

Belief

A belief is the robot’s current representation of which states of the world are plausible, given the information it has observed. It is the new idea in this lesson, and it is simpler than it sounds. Suppose the robot sees a partly hidden box. Its belief might be:

70%
the box is on the table.
30%
the box is behind the table edge.

The robot is not claiming to know. It is keeping both possibilities alive, with weights. Belief can be represented in many ways: a single estimated value, a value with a confidence, a probability distribution, a set of hypotheses, an occupancy grid (a map of which cells are probably free or occupied), or a cloud of weighted guesses called particles. This lesson does not require any one of them. Belief is a conceptual idea about uncertainty, and the representation is an engineering choice.

Note that “belief” here is a technical term for a representation of uncertainty. It has nothing to do with consciousness or human-like opinions.

State estimation

State estimation is the process of using observations and prior information to estimate the state of a system. Its shape is the same everywhere:

  1. Previous estimate
  2. New observation
  3. State estimation
  4. Updated estimate
Estimation combines what the robot already believed with what it has just observed. The result becomes the next “previous estimate.”

Examples include robot localization (where am I?), object tracking (where is that thing, as it moves?), velocity estimation and sensor fusion (merging several sensors). The best-known tools are the Kalman filter, the extended Kalman filter and the particle filter. We deliberately skip their mathematics. What matters here is the problem they solve: turning a stream of noisy, partial measurements into a steady, honest estimate. The Localization & SLAM track (coming soon) goes deeper; for a hands-on first look, see Where Am I?.

Prediction and correction

Many estimators work in two repeating steps. Prediction asks: given what I knew before and how the system moves, where do I expect it to be now? Correction asks: given the new sensor measurement, how should I update my estimate?

  1. Previous estimate
  2. Predict
  3. Predicted state + new observation
  4. Correct
  5. Updated estimate
Predict from motion, then correct with a measurement. The updated estimate starts the next cycle.

A simple example: the robot estimates x = 2.0 m. After the wheels turn, it predicts x = 2.8 m. A sensor measurement says x ≈ 2.6 m. The updated estimate lands somewhere between the prediction and the measurement, closer to whichever one is more certain. If the wheels tend to slip, it moves toward the measurement. If the sensor is noisy, it stays closer to the prediction.

Multiple sensors and multiple observations

Robots combine information because no sensor is enough alone, and their weaknesses differ:

Wheel odometry
Good short-term motion information, but it drifts.
IMU
High-rate motion information, but it has bias and drift.
LiDAR and camera
Information about the surroundings, but intermittent or noisy.
  • Odometry
  • IMU
  • LiDAR
  • Camera
  1. State estimation
  2. Robot belief
Several sensors feed one estimate, which feeds the robot’s belief.

Combining them can give a better estimate than any one alone, because one sensor’s strength covers another’s weakness. More sensors do not guarantee a better result: they bring calibration, synchronization, noise and fusion problems of their own. For measurement details, continue with the Sensors track.

Why this matters for Physical AI

An agent constantly needs answers to questions like these: Where am I? Where is the object? What am I holding? Is the door open? Did my action succeed? Has the environment changed? None of them can be answered from perfect knowledge. The agent has to reason from observations, estimates, beliefs, memory and models. In the Physical AI loop from T1, the update step is now explicit:

  1. Observe
  2. Update belief
  3. Decide
  4. Act
  5. Observe again
  6. Update belief
  7. Continue or recover
The Physical AI loop with belief made explicit. Each observation changes what the robot believes.

“Did my action succeed?” is a question about state, answered by observation, not by the action’s return value. That is exactly the situation in Lab 09.

Connection to planning and decision making

Plans are built from state. Say the goal is to bring the red box to the shelf. The planner assumes the red box is on the table, the robot is in the kitchen and the path is clear. Now suppose the box has been moved, a door has closed, or the robot’s own location estimate was wrong. The plan was valid for a world that no longer exists.

  1. Sensors
  2. Perception
  3. State / belief
  4. Planning
  5. Action
  6. New observation
Planning sits on top of state and belief, and every action produces a new observation that can invalidate the plan.

So the cycle is: represent the state, plan, act, observe, update the state or belief, and replan if necessary. A planner is only as good as the picture of the world beneath it. Later lessons on planning, execution and replanning build on this.

Common misconceptions

Misconception 1 “The sensor tells the robot the state.”

Correction A sensor produces measurements. The robot may need estimation or inference to determine the state.

Misconception 2 “The estimated state is the true state.”

Correction An estimate can be uncertain or wrong.

Misconception 3 “Belief means the robot is conscious or has human-like beliefs.”

Correction In robotics, belief is a technical representation of uncertainty about possible states.

Misconception 4 “If the robot cannot observe something, it doesn’t matter.”

Correction Unobserved state can still affect planning and the outcome of actions.

Misconception 5 “More sensors always mean perfect state estimation.”

Correction More sensors give more information, but also add calibration, synchronization, noise and fusion challenges.

Engineering example: localization in a hallway

A robot drives down a straight hallway. Here is one cycle of estimation, step by step.

  1. Step 1 — Start

    The initial estimate is x = 0 m, with little doubt.

  2. Step 2 — Predict

    The robot drives forward. From the wheel rotation, odometry predicts x = 5.2 m. But a wheel slipped on the way, so the prediction is too optimistic.

  3. Step 3 — Observe

    The LiDAR sees a wall at a known position. Working back from that, the measurement suggests x ≈ 4.8 m.

  4. Step 4 — Notice the discrepancy

    The two numbers disagree by 0.4 m. Neither is simply “the truth.” The prediction may have drifted, but the measurement could also be off, for example if the wall was mismatched.

  5. Step 5 — Correct

    The robot blends them by how much it trusts each. If odometry is uncertain by about ±0.4 m and the LiDAR by about ±0.1 m, the LiDAR deserves 16 times the weight, and the updated estimate is about 4.82 m. If the two were equally trusted, it would be 5.0 m. The estimate also keeps a smaller uncertainty than either input.

Neither source is trusted blindly, and neither is thrown away. This is the core of real localization, which scales the same idea to many sensors, many dimensions and continuous motion. The Where Am I? lab lets you watch odometry drift and landmarks correct it, and the Localization & SLAM track (coming soon) covers it in depth.

Engineering takeaways

  1. State describes the relevant condition of the system and the world.
  2. Sensors provide observations, not perfect state.
  3. The true state may be hidden.
  4. State estimates are inferred from observations.
  5. Estimates can be uncertain.
  6. Belief represents which states the robot considers plausible.
  7. State estimation combines information over time.
  8. Prediction and correction are fundamental estimation ideas.
  9. Planning depends on the robot’s current state and belief.
  10. Physical AI continuously updates its understanding as it interacts with the world.

A robot does not need perfect knowledge of the world. It needs a sufficiently good and uncertainty-aware understanding to make the next safe and useful decision.

Knowledge check

Three questions about reasoning, not vocabulary. Write an answer, then reveal the explanation. Your answers stay in your browser.

Question 1

A LiDAR measurement says an obstacle is 4.8 m away, while the robot’s predicted position suggests it should be 5.2 m away. Is either number automatically the true state?

Explanation

No. They are pieces of evidence with different uncertainties, and state estimation combines them into an updated estimate. The prediction comes from a motion model that can drift. The measurement has its own noise and can be wrong. The robot weighs each by how far it trusts it, and keeps track of how uncertain the result remains.

Question 2

A robot’s camera cannot currently see a red box because it is behind an obstacle. Does that mean the box no longer exists in the world model?

Explanation

No. The box is still part of the possible world state, even though it is currently unobserved. “Not seen” is not the same as “not there.” The robot can keep a belief about where the box probably is, and that belief can steer it to go and look again.

Question 3

What is the difference between an observation, an estimated state and a belief?

Explanation

An observation is a sensor-derived measurement. An estimated state is the robot’s current best estimate of the relevant state. A belief is a representation of uncertainty, or of the possible states, given the available information. The observation is the evidence, the estimate is the best single answer drawn from the evidence, and the belief keeps the whole picture, including how unsure the robot is.

What’s next

You now know what the robot is trying to know. Next, we need to understand how raw sensor measurements become useful observations.

Next lesson

T3 — The Physical AI Stack at a Glance

T3 connects sensors, perception, world models, planning, skills, motion, control, verification and recovery into one system-level architecture, so you can see where state and belief sit among the other parts.

Each idea here has a hands-on home elsewhere on FixTheRobot. Nothing below is required to finish this lesson.

  • Track
    Sensors

    Want to understand LiDAR measurements in detail? Continue with the Sensors track.

  • Lab
    Teach the Robot to Sense

    LiDAR, camera and wheel encoders: what each one actually reports.

  • Lab
    Where Am I?

    Watch odometry drift and landmarks correct it, as in the hallway example.

  • Lab
    Build the Robot’s World

    Turn sensor data into an occupancy grid, one way to represent belief about space.

  • Lab 09
    Don’t Trust Your Actions Available now

    Use two imperfect sensors to decide whether a grasp really worked.

  • Lab 02
    World Model as Typed State In development

    Build the structured world representation planning relies on.

  • Lab 03
    Partial Observability In development

    Turn noisy, partial observations into a belief.

  • Track
    Perception and Localization & SLAM Coming soon

    How measurements become observations, and how robots localize in depth.