Physical AI · Theory
T2 — State, Observation and Belief
Your robot never sees the world directly. It sees measurements and builds an understanding from them.
In T1, we saw that a physical agent cannot directly access the complete state of the world. This lesson makes that idea precise.
Learning objectives
After this lesson, you should be able to:
- Define state in a robotics system.
- Distinguish world state from robot state.
- Explain what an observation is, and why it is not the same as state.
- Distinguish true state from estimated state.
- Explain partial observability.
- Explain uncertainty and belief.
- Describe the basic idea of state estimation.
- Explain why these concepts matter for Physical AI agents.
The central skill is keeping five terms apart: true state, sensor measurement, observation, state estimate and belief. They sound alike and they are different.
The one-sentence idea
A robot acts on what it believes about the world, not on perfect access to the world itself.
The world has a definite condition, but the robot can only read it through sensors. Those readings are turned into an estimate, the estimate carries some doubt, and that estimate, with its doubt, is what the robot uses to decide and act. The chain looks like this:
- True world stateHow things actually are. Never directly available.
- SensorsTurn physical quantities into measurements.
- ObservationsWhat the robot receives: readings and detections.
- State estimationCombines observations with earlier knowledge.
- BeliefWhat the robot treats as plausible, and how strongly.
- DecisionChooses what to do from the belief.
- ActionChanges the world, and the chain starts again.
This is a conceptual model, not necessarily a literal software pipeline. Some systems merge these steps and some split them further. The distinctions still hold, because they describe what is known, not how the code is organized.
What is state?
State is the information needed to describe the relevant condition of a system at a given time. For a robot, it is the answer to “what do I need to know right now to decide what to do?” The word “relevant” matters, because the useful state depends on the problem:
- A mobile robot driving around
- Position
x,y, headingθand velocity. - A manipulation task
- Robot pose, joint angles, gripper state, object poses and how objects relate to each other.
- A task-level agent
- Robot location, objects, task progress, what it is holding and which actions have completed.
There is no single universal state representation. The useful state is the state required to make the decisions we care about. A robot choosing wheel speeds needs velocity; a robot choosing which room to visit does not.
Robot state vs world state
World state is everything relevant that is actually true in the environment. Robot state is the relevant internal physical condition of the robot itself: pose, velocity, joint positions, gripper width. Robot state is a part of the world state, but it is useful to name it separately because robots sense themselves differently from how they sense everything else.
robot = kitchen red_box = table blue_box = shelf door = closed gripper = empty
A Physical AI system usually needs a broader representation that holds both: where the robot is and what it is doing, plus what is around it. The diagram shows where that representation comes from, and the catch.
- World state
- Robot
- Objects
- Environment
- Relations
- Sensorsa narrow, noisy view
- Observations
- Camera
- LiDAR
- IMU
- Encoders
What is an observation?
An observation is information the robot obtains through its sensors or other available measurements. It is evidence about the world. Different sensors produce very different kinds of evidence:
| Sensor | What it produces |
|---|---|
| Camera | Pixels |
| LiDAR | Ranges, or a point cloud |
| IMU | Acceleration and angular velocity |
| Wheel encoders | Wheel rotation |
| Force sensor | Measured force |
| GPS (satellite positioning) | A position measurement |
| Gripper sensor | Width, force or contact |
An observation is a measurement. It is not automatically a complete explanation of the world. A note on vocabulary: a raw measurement is what the sensor outputs (the pixels), while a processed observation is what software makes of it (“red box detected”). People use “observation” for both. Later lessons follow the path from one to the other; here, what matters is that either way it is evidence, not the state.
State vs observation
This is one of the most important distinctions in the lesson. Take a red box.
True state
red_box = (2.0, 1.5)red_box = on_tablered_box = not_in_gripper
Camera observation
red_box detectedpixel position = (421, 237)confidence = 0.82
The observation does not equal the state. A pixel position is not a position in the room, and a detection with 0.82 confidence is not a guarantee that the box exists. Whether the box is on the table or in the gripper is not stated at all. A second example: the true distance to an obstacle is 3.2 m, and the LiDAR measures 3.1 m.
State ≠ measurement.
Measurements provide evidence about the state. The robot has to infer the state from the evidence.
Example: robot and red box
Return to the scenario from T1. A robot stands in the kitchen. A red box is on the table, a blue bottle is on a shelf, and the target area is in another room. Its camera detects red_box with confidence 0.91. Compare what is true with what the robot has:
| Question | True state | What the robot has |
|---|---|---|
| Where is the box? | (2.00, 1.50) | An estimate such as (2.04, 1.46) |
| Is it on the table? | Yes | A detection with confidence 0.91 |
| Is it graspable? | Depends on its shape and what is around it | Not known yet |
| Is something hiding part of it? | Perhaps | Only a hint, from a partial outline |
| Is the gripper empty? | Yes | A gripper sensor reading |
The robot combines observations over time and keeps an internal representation. That is the difference between two questions: what is true? and what does the robot currently know? A Physical AI agent can only act on the second.
True state vs estimated state
The true state is not available directly. The robot gets observations, and from them it computes an estimated state, its current best guess at the relevant state.
true robot position: x = 5.00 m estimated robot position: x ≈ 4.86 m
The estimate is off by 14 cm, and the robot does not know that. The estimate can be wrong. That is not necessarily a software bug; it is a fundamental consequence of imperfect sensing and inference. The goal of engineering is not to make estimates perfect but to make them good enough, and to know how far to trust them.
Why sensors can’t reveal everything
Every sensor gives only a particular view of reality, and each view has blind spots:
| Sensor | Typical limits |
|---|---|
| Camera | Occlusion, lighting, depth ambiguity, limited field of view |
| LiDAR | Limited range, occlusion, sparse measurements, reflective or light-absorbing surfaces |
| IMU | Bias and drift; errors build up when acceleration is integrated over time |
| Wheel odometry | Wheel slip, uneven surfaces, encoder errors |
| GPS | Reflections off buildings, poor availability indoors, limited accuracy |
None of these is a defect to be engineered away. A camera cannot see through a box and a LiDAR cannot measure past its range. These are the limits of the measurement itself.
Hidden state
Hidden state is relevant information about the world that cannot be directly observed at the current moment. The robot sees a box. It may not directly know whether the box is empty, heavy or fragile, whether it is attached to something, or even whether it is where the detector says it is.
Hidden state still matters. A heavy box changes how the arm should lift, and a fragile one changes how hard the gripper should squeeze. The robot cannot read these off the sensors, so it must infer them, test them by acting carefully, or accept the risk. Much of Physical AI decision making is deciding what to do about what the robot cannot see.
Uncertainty
Uncertainty is the robot’s doubt about the state: how far the estimate might be from the truth. Suppose the robot detects a red bottle once with confidence 0.93 and once with confidence 0.51. It should not treat those two situations identically. Uncertainty comes from several places:
- sensor noise,
- ambiguous observations (is that a bottle or a vase?),
- occlusion,
- model error (the robot’s idea of how it moves is slightly wrong),
- environmental change,
- unknown objects, and
- delayed measurements.
The point of tracking uncertainty is that it should change decisions. With high confidence, the robot can act. With low confidence, it can inspect again, gather another observation, or ask for clarification. One caution: a detector’s “confidence” score is often not a true probability, so treat it as a useful signal, not as a precise measure. Later lessons build the act, inspect or ask decision on this foundation.
Belief
A belief is the robot’s current representation of which states of the world are plausible, given the information it has observed. It is the new idea in this lesson, and it is simpler than it sounds. Suppose the robot sees a partly hidden box. Its belief might be:
- 70%
- the box is on the table.
- 30%
- the box is behind the table edge.
The robot is not claiming to know. It is keeping both possibilities alive, with weights. Belief can be represented in many ways: a single estimated value, a value with a confidence, a probability distribution, a set of hypotheses, an occupancy grid (a map of which cells are probably free or occupied), or a cloud of weighted guesses called particles. This lesson does not require any one of them. Belief is a conceptual idea about uncertainty, and the representation is an engineering choice.
Note that “belief” here is a technical term for a representation of uncertainty. It has nothing to do with consciousness or human-like opinions.
State estimation
State estimation is the process of using observations and prior information to estimate the state of a system. Its shape is the same everywhere:
- Previous estimate
- New observation
- State estimation
- Updated estimate
Examples include robot localization (where am I?), object tracking (where is that thing, as it moves?), velocity estimation and sensor fusion (merging several sensors). The best-known tools are the Kalman filter, the extended Kalman filter and the particle filter. We deliberately skip their mathematics. What matters here is the problem they solve: turning a stream of noisy, partial measurements into a steady, honest estimate. The Localization & SLAM track (coming soon) goes deeper; for a hands-on first look, see Where Am I?.
Prediction and correction
Many estimators work in two repeating steps. Prediction asks: given what I knew before and how the system moves, where do I expect it to be now? Correction asks: given the new sensor measurement, how should I update my estimate?
- Previous estimate
- Predict
- Predicted state + new observation
- Correct
- Updated estimate
A simple example: the robot estimates x = 2.0 m. After the wheels turn, it predicts x = 2.8 m. A sensor measurement says x ≈ 2.6 m. The updated estimate lands somewhere between the prediction and the measurement, closer to whichever one is more certain. If the wheels tend to slip, it moves toward the measurement. If the sensor is noisy, it stays closer to the prediction.
Multiple sensors and multiple observations
Robots combine information because no sensor is enough alone, and their weaknesses differ:
- Wheel odometry
- Good short-term motion information, but it drifts.
- IMU
- High-rate motion information, but it has bias and drift.
- LiDAR and camera
- Information about the surroundings, but intermittent or noisy.
- Odometry
- IMU
- LiDAR
- Camera
- State estimation
- Robot belief
Combining them can give a better estimate than any one alone, because one sensor’s strength covers another’s weakness. More sensors do not guarantee a better result: they bring calibration, synchronization, noise and fusion problems of their own. For measurement details, continue with the Sensors track.
Why this matters for Physical AI
An agent constantly needs answers to questions like these: Where am I? Where is the object? What am I holding? Is the door open? Did my action succeed? Has the environment changed? None of them can be answered from perfect knowledge. The agent has to reason from observations, estimates, beliefs, memory and models. In the Physical AI loop from T1, the update step is now explicit:
- Observe
- Update belief
- Decide
- Act
- Observe again
- Update belief
- Continue or recover
“Did my action succeed?” is a question about state, answered by observation, not by the action’s return value. That is exactly the situation in Lab 09.
Connection to planning and decision making
Plans are built from state. Say the goal is to bring the red box to the shelf. The planner assumes the red box is on the table, the robot is in the kitchen and the path is clear. Now suppose the box has been moved, a door has closed, or the robot’s own location estimate was wrong. The plan was valid for a world that no longer exists.
- Sensors
- Perception
- State / belief
- Planning
- Action
- New observation
So the cycle is: represent the state, plan, act, observe, update the state or belief, and replan if necessary. A planner is only as good as the picture of the world beneath it. Later lessons on planning, execution and replanning build on this.
Common misconceptions
Misconception 1 “The sensor tells the robot the state.”
Correction A sensor produces measurements. The robot may need estimation or inference to determine the state.
Misconception 2 “The estimated state is the true state.”
Correction An estimate can be uncertain or wrong.
Misconception 3 “Belief means the robot is conscious or has human-like beliefs.”
Correction In robotics, belief is a technical representation of uncertainty about possible states.
Misconception 4 “If the robot cannot observe something, it doesn’t matter.”
Correction Unobserved state can still affect planning and the outcome of actions.
Misconception 5 “More sensors always mean perfect state estimation.”
Correction More sensors give more information, but also add calibration, synchronization, noise and fusion challenges.
Engineering example: localization in a hallway
A robot drives down a straight hallway. Here is one cycle of estimation, step by step.
-
Step 1 — Start
The initial estimate is
x = 0 m, with little doubt. -
Step 2 — Predict
The robot drives forward. From the wheel rotation, odometry predicts
x = 5.2 m. But a wheel slipped on the way, so the prediction is too optimistic. -
Step 3 — Observe
The LiDAR sees a wall at a known position. Working back from that, the measurement suggests
x ≈ 4.8 m. -
Step 4 — Notice the discrepancy
The two numbers disagree by 0.4 m. Neither is simply “the truth.” The prediction may have drifted, but the measurement could also be off, for example if the wall was mismatched.
-
Step 5 — Correct
The robot blends them by how much it trusts each. If odometry is uncertain by about ±0.4 m and the LiDAR by about ±0.1 m, the LiDAR deserves 16 times the weight, and the updated estimate is about 4.82 m. If the two were equally trusted, it would be 5.0 m. The estimate also keeps a smaller uncertainty than either input.
Neither source is trusted blindly, and neither is thrown away. This is the core of real localization, which scales the same idea to many sensors, many dimensions and continuous motion. The Where Am I? lab lets you watch odometry drift and landmarks correct it, and the Localization & SLAM track (coming soon) covers it in depth.
Engineering takeaways
- State describes the relevant condition of the system and the world.
- Sensors provide observations, not perfect state.
- The true state may be hidden.
- State estimates are inferred from observations.
- Estimates can be uncertain.
- Belief represents which states the robot considers plausible.
- State estimation combines information over time.
- Prediction and correction are fundamental estimation ideas.
- Planning depends on the robot’s current state and belief.
- Physical AI continuously updates its understanding as it interacts with the world.
A robot does not need perfect knowledge of the world. It needs a sufficiently good and uncertainty-aware understanding to make the next safe and useful decision.
Knowledge check
Three questions about reasoning, not vocabulary. Write an answer, then reveal the explanation. Your answers stay in your browser.
Question 1
A LiDAR measurement says an obstacle is 4.8 m away, while the robot’s predicted position suggests it should be 5.2 m away. Is either number automatically the true state?
Explanation
No. They are pieces of evidence with different uncertainties, and state estimation combines them into an updated estimate. The prediction comes from a motion model that can drift. The measurement has its own noise and can be wrong. The robot weighs each by how far it trusts it, and keeps track of how uncertain the result remains.
Question 2
A robot’s camera cannot currently see a red box because it is behind an obstacle. Does that mean the box no longer exists in the world model?
Explanation
No. The box is still part of the possible world state, even though it is currently unobserved. “Not seen” is not the same as “not there.” The robot can keep a belief about where the box probably is, and that belief can steer it to go and look again.
Question 3
What is the difference between an observation, an estimated state and a belief?
Explanation
An observation is a sensor-derived measurement. An estimated state is the robot’s current best estimate of the relevant state. A belief is a representation of uncertainty, or of the possible states, given the available information. The observation is the evidence, the estimate is the best single answer drawn from the evidence, and the belief keeps the whole picture, including how unsure the robot is.
What’s next
You now know what the robot is trying to know. Next, we need to understand how raw sensor measurements become useful observations.
Go deeper
Each idea here has a hands-on home elsewhere on FixTheRobot. Nothing below is required to finish this lesson.
- TrackSensors
Want to understand LiDAR measurements in detail? Continue with the Sensors track.
- LabTeach the Robot to Sense
LiDAR, camera and wheel encoders: what each one actually reports.
- LabWhere Am I?
Watch odometry drift and landmarks correct it, as in the hallway example.
- LabBuild the Robot’s World
Turn sensor data into an occupancy grid, one way to represent belief about space.
- Lab 09Don’t Trust Your Actions Available now
Use two imperfect sensors to decide whether a grasp really worked.
- Lab 02World Model as Typed State In development
Build the structured world representation planning relies on.
- Lab 03Partial Observability In development
Turn noisy, partial observations into a belief.
- TrackPerception and Localization & SLAM Coming soon
How measurements become observations, and how robots localize in depth.
Get notified when the next lesson and new labs are released.
One email per new lesson or lab.