Physical AI · Theory
T1 — What Is Physical AI, and Why Is It Hard?
Intelligence changes when it has to act in the real world.
Understand how intelligence becomes action in the physical world.
Learning objectives
After this lesson, you should be able to:
- Define Physical AI in your own words.
- Explain the difference between AI, robotics, embodied AI and Physical AI.
- Describe the perception–reasoning–action loop.
- Explain why physical systems operate under uncertainty.
- Explain why closed-loop feedback is essential.
- Identify the major layers of a Physical AI system.
This is a reading lesson: no code, no simulation. It assumes basic robotics but nothing about Physical AI. World models, behavior trees and task planning appear only at a high level here; later lessons study them in depth.
The one-sentence idea
Physical AI is intelligence that must perceive, reason, and act through a physical body while dealing with uncertainty, limited information, time constraints, and real-world consequences.
This sentence anchors the lesson, so we take it apart. Each term gets a plain explanation and a robotics example.
- Physical
- The system acts through a body in a world ruled by physics: objects have mass, floors have friction, nothing teleports. Example: a mobile robot cannot drive through a table, however confident its plan.
- AI
- Methods that let a machine interpret, decide or learn instead of replaying a fixed script. Example: recognizing a bottle, or choosing the next sub-task.
- Perceive
- Turning raw sensor data into usable information. Example: a camera frame becomes “a bottle, about 1.2 m ahead, slightly left.”
- Reason
- Deciding what to do from what the system believes and what it wants. Example: approaching the bottle from the side because a box blocks the front.
- Act
- Producing motion or force through wheels, joints, grippers or rotors. Example: closing a gripper to a chosen force.
- Uncertainty
- Nothing the system knows is exact: measurements are noisy, models approximate, outcomes not guaranteed. Example: the bottle looks 1.2 m away but may be 1.15 m or 1.28 m.
- Consequences
- Actions change the real world, and some changes cannot be undone. Example: a bottle that has been dropped stays dropped.
What is Physical AI?
Much of the AI you have used works on information: classify an image, translate a sentence, summarize a document, generate code, answer a question. You give it an input, and it returns an output. Nothing about the answer changes the question.
A Physical AI system has to interact with an environment: a robot navigating a warehouse, an autonomous vehicle, a drone, a humanoid manipulating an object, a delivery robot. Its output is not an answer. It is motion, and motion changes the world the system is about to look at again.
Digital system
- Input
- Computation
- Output
Physical AI
- Environment
- Sense
- Understand
- Decide
- Act
- Environment changes
- Sense again
The environment is part of the computation loop.
In a Physical AI system, your next input depends on your last output. A translation model’s poor word choice in sentence three leaves sentence four unaffected. If a robot nudges a cup while reaching for a bottle, the cup is now somewhere else, and every later decision must account for that.
Physical AI vs conventional AI
Conventional AI does touch the world: trading systems act on markets, software agents click through websites. The difference is one of degree, and the degree is large. A typical software AI system works on a relatively well-defined input and produces an output. A physical agent copes with all of this at once:
- Incomplete information. Part of the world is hidden.
- Noisy measurements. Sensors report approximations.
- Changing environments. People move, doors close.
- Delayed information. Data is old by the time it is used.
- Continuous actions. Not a choice among a few labels, but speeds, forces and paths.
- Imperfect actuators. Commanded motion is not achieved motion.
- Unexpected events. Things go wrong in unlisted ways.
- Physical constraints. Reach, speed, payload and balance limit what is possible.
Compare two problems that both involve a red bottle.
Image classifier
- Image
- “red bottle”
Robot
- Camera
- Detect the bottle
- Determine which bottle
- Estimate its location
- Navigate
- Reach
- Grasp
- Verify the grasp
- Recover if the grasp failed
The second problem is fundamentally different, not just longer, for four reasons.
- The answer is not the goal. The classifier is finished when it says “red bottle.” The robot is finished only when the world has changed.
- Every step can fail, and errors travel. A slightly wrong position estimate makes the reach miss; the miss makes delivery impossible.
- Steps change the situation. After bumping the table, the robot is no longer solving the problem it started with.
- There is no reset. A model can be queried again for free. A dropped bottle stays on the floor.
Robotics, AI, embodied AI and Physical AI
These terms are used loosely, sometimes interchangeably, and people draw the lines in different places. This course uses the working definitions below. None is “better” than another; they describe different things, and they overlap.
Robotics
Engineering machines that sense, compute and act in the physical world.
Example: an industrial arm repeating a fixed welding path.
Artificial intelligence
Methods for machines to perform tasks involving perception, reasoning, learning or decision making.
Example: a medical-image classifier. It has no body.
Embodied AI
Intelligence whose behavior is grounded in interaction with an environment through an embodiment.
Example: an agent learning to walk in a physics simulator; the embodiment can be simulated.
Physical AI
A practical class of intelligent systems whose AI capabilities are connected to physical sensing and action, where decisions must operate under real-world constraints.
Example: a warehouse robot that perceives shelves, plans a route, picks an item and checks that it holds it.
A robotic system does not automatically constitute sophisticated Physical AI. A welding arm replaying a recorded path is good robotics, but the “intelligence” was a human’s, frozen in advance.
An AI model does not become Physical AI simply because someone puts it on a robot. A vision model whose output never influences motion, or whose conclusions are never checked against what happened, sits next to a robot rather than being connected to it.
The important property is the closed connection between these steps:
- Perception
- Understanding
- Decision
- Physical action
- Feedback
- Perception
The Physical AI loop
This is the most important diagram in the lesson. The world is sensed, the data interpreted, a model updated, a decision made, the robot acts, and the world changes. Then it starts again.
- WorldObjects, people, floors, light: the source of all information and the target of all actions.
- SensorsCameras, LiDAR (laser scanners that measure distance), wheel encoders, force sensors. They give measurements, not truths.
- PerceptionTurns measurements into descriptions: “a bottle, here, 63% confident.”
- World / state modelThe robot’s running estimate of what exists, where, and in what state, including itself and task progress.
- Reason / planChooses the next step from the goal and the estimate.
- ActSends commands to navigation, arm and gripper.
- RobotExecutes with real motors, delays and limits.
- WorldDifferent now, because the robot acted.
- Sense againThe robot must check what its action really did.
Every block can be wrong, and the line returning to the world is what makes this a loop rather than a pipeline. A pipeline ends. A loop keeps checking.
A useful picture, not a required architecture
This is a conceptual decomposition; real systems do not all contain these exact modules. Some merge blocks, for example one neural network mapping camera images straight to motor commands. The loop still exists there, with its difficulties moved inside the model. Use the diagram to ask, “which part of my system is doing this job?”
The loop runs at several speeds at once
A walking humanoid corrects its balance hundreds of times per second, while its task planner decides what to do next every few seconds. Real robots nest loops at different timescales, and the fast ones protect the slow ones.
- Fast control loop
- Hundreds of times per second: keeps motors and joints on target.
- Perception loop
- Tens of times per second: updates the state from each camera or LiDAR frame.
- Local planning loop
- A few times per second: steers around obstacles.
- Task planning loop
- Seconds to minutes: chooses the next sub-task and when to change plan.
These rates are orders of magnitude, not rules. Later lessons treat each loop; for now, remember there is more than one clock.
A concrete example: bring the red bottle
We will use one task throughout the Physical AI curriculum:
“Bring the red bottle from the kitchen to the user.”
A person does this without thinking. A robot must establish a great deal.
-
Step 1 — Understand the goal
The command is human language. To plan, the robot needs a structured goal such as
Bring(red_bottle, user). Even this hides questions: which bottle, if two are red? Where is “the user”? Answering them is called grounding. -
Step 2 — Understand the world
Where are the bottle, the user and the robot? The robot answers from its world model, a belief built from earlier observations. That belief can be stale or empty: the bottle may never have been seen.
-
Step 3 — Plan
From the goal and the belief, the robot produces steps:
- Find the bottle
- Navigate to it
- Grasp it
- Verify the grasp
- Navigate to the user
- Place it or hand it over
- Verify completion
Verification is part of the plan itself, not an afterthought.
-
Step 4 — Act
Navigation becomes wheel commands; grasping becomes joint motion. Here the plan meets friction, slip, delays and tolerances.
-
Step 5 — Observe
After closing the gripper, the robot asks: did I actually grasp the bottle? Its evidence is how far the gripper closed, whether it feels contact, what a camera sees. None is perfect. It also watches the wider world.
-
Step 6 — Recover
If the grasp failed: retry, reposition, re-observe or replan. The right choice depends on why it failed. A slight miss may need a retry; a bottle that rolled away needs fresh perception first.
-
Step 7 — Verify the final goal
The robot should not conclude “I executed all my commands.” It should establish that the requested world state has actually been achieved: the bottle is with the user, not on the floor.
Why the physical world is hard
Each challenge is easy to state; the trouble is that they arrive together, in one robot. For each: what it is, an example, and what it forces the agent to do. Several get their own sections afterwards.
1 Uncertainty
- What it is
- Knowledge is approximate.
- Example
- The bottle is at 1.20 m, give or take a few centimeters.
- Consequence
- Decisions must tolerate error and carry confidence.
2 Partial observability
- What it is
- Only part of the world is visible.
- Example
- The bottle is behind a cereal box.
- Consequence
- Act without knowing, remember, or go and look.
3 Dynamic environments
- What it is
- The world changes without the robot’s help.
- Example
- A person walks through the kitchen mid-task.
- Consequence
- Plans go stale; they must be checked, not just followed.
4 Sensor noise
- What it is
- Measurements deviate from the truth.
- Example
- A depth camera returns holes on shiny surfaces.
- Consequence
- Readings must be filtered and cross-checked.
5 Actuator imperfections
- What it is
- Achieved motion differs from commanded motion.
- Example
- A wheel slips on a wet floor.
- Consequence
- A command is not an outcome; measure the outcome.
6 Timing and latency
- What it is
- Information is old when used; actions take time.
- Example
- An obstacle moved since the frame was captured.
- Consequence
- Reason about the present from a past snapshot.
7 Physical constraints
- What it is
- Physics limits what is possible.
- Example
- The shelf is beyond the arm’s reach.
- Consequence
- Plans must be physically feasible, not just logical.
8 Long-horizon tasks
- What it is
- Useful tasks take many steps; any can fail.
- Example
- 20 steps at 95% reliability each succeed end to end only about 36% of the time.
- Consequence
- Small failure rates compound; recovery is not optional.
9 Unexpected failures
- What it is
- Things fail in unanticipated ways.
- Example
- The bottle is slippery, or the wireless link drops.
- Consequence
- Detect failure, recover generically, stop safely.
10 Safety
- What it is
- The robot can hurt people or things.
- Example
- An arm moving near a person.
- Consequence
- Limits are enforced independently of the AI’s decisions.
Uncertainty
Uncertainty means that the system’s knowledge is approximate and can be wrong. Suppose the robot sees a bottle. It still does not know, with absolute certainty:
- Is this the correct bottle?
- Exactly where is it?
- Can the gripper reach it?
- Will the grasp succeed?
- Has the object moved since I last looked?
To think clearly about this, separate two easily confused things: the world state, how things really are, and the robot’s estimate of the world state, what it believes.
- True stateHow the world really is
- Sensorsnoise, limited view
- ObservationWhat the sensors reported
- Estimationmemory, models, assumptions
- Estimated state / beliefWhat the robot concludes
- True state: whatever is actually the case. The robot has no direct access to it.
- Observation: what the sensors reported at one moment, such as a detection saying “bottle at (1.2, 0.4).”
- Estimated state, or belief: what the robot concludes by combining observations with memory and assumptions, ideally with a measure of confidence.
Two consequences follow. The robot decides from its belief, so belief errors become action errors. And a well-built system tracks not only “where is the bottle?” but “how sure am I?” The next lesson, T2, develops these distinctions carefully.
Partial observability
A robot cannot see everything. A cereal box hides the bottle behind it. The next room is outside the camera’s field of view. A LiDAR scan is noisy and has limited range.
The world may contain information that the robot does not currently possess.
This matters most for planning, because the central trap is confusing not seen with not there. If the bottle is missing from the latest frame, the robot cannot conclude it has vanished; it may be hidden, out of range or in another room.
The agent must treat unseen parts of the world as unknown, not empty, and it has options: go and look, move for a better view, use what it remembers, ask a person, or act cautiously and check again.
Compactly: world state ≠ observed state. Later lessons formalize this; for now you only need the intuition that sensors give a partial, momentary view.
Time and latency
A robot never acts on the world as it is now, only as it was when the sensor captured its data. Latency is the delay between something happening and the system being able to respond, and it builds up in four places:
- Sensor latency. A camera exposes, reads out and transmits each frame.
- Computation latency. Detection, state updates and planning take time.
- Communication latency. Data crosses buses, networks and software layers.
- Actuator latency. Motors ramp up; grippers take time to close.
A concrete case: the camera sees an obstacle at position A, and by the time the planner uses that observation, the obstacle is at position B. If a person walks at 1.2 m/s and the delay from camera to command is 300 ms, they have moved about 0.36 m, enough to turn a safe path into a collision.
A robot is always acting on information that may already be old. Good systems timestamp data so they know when an observation was true, predict forward instead of assuming the world froze, and run safety-critical loops fast, close to the hardware.
Actions have physical consequences
In software, a wrong answer gives a bad output you can inspect and discard. In the physical world, a wrong action can cause a collision, a dropped object, damaged hardware, unsafe motion, or a failed task that leaves the world worse than before.
That is why Physical AI needs four things digital systems often avoid. Constraints limit what the robot may do. Verification checks what actually happened. Recovery handles what did not work. Safety mechanisms stop harm even when everything else has failed.
When SUCCESS does not mean success
Here is a pattern you will meet repeatedly. A robot’s API (the set of commands its software exposes) returns:
grasp("red_box") → SUCCESS
Is the object in the gripper? Not necessarily. The status usually means only that the command ran: the gripper closed and raised no fault. It may have closed on air, or the object may have slipped. Three different facts hide behind one word:
| Level | The question | For a grasp |
|---|---|---|
| Command executed | Did the robot do what it was told? | The gripper closed without a fault. |
| World changed as intended | Is the effect we wanted actually there? | The bottle is in the gripper. |
| Goal achieved | Is the task really done? | The bottle has been delivered to the user. |
Reporting the first as if it were the third is a classic source of robot failures. How would a robot find out the truth? That is an engineering question, and Lab 09, linked at the end, makes it hands-on.
Closed-loop intelligence
There are two basic ways to carry out a plan.
Open loop
- Plan
- Execute
- Done (assumed)
If reality differs from the plan, nobody notices.
Closed loop
- Plan
- Act
- Observe
- Verify
- Update
- Continue, recover or replan
If reality differs from the plan, the mismatch is detected and handled.
If you have tuned a PID controller (a feedback controller that corrects a command in proportion to its error), you know closed-loop control at the signal level; see Lab 02 and the oscillating-robot lab. Physical AI applies the same principle one level up: not only “is the wheel at the right speed?” but “is the bottle actually in the gripper?”
- Navigation
- Odometry drifts, so the robot keeps comparing where it thinks it is with what its sensors see, and replans when a path is blocked.
- Grasping
- The robot checks for contact and retries when the object is not where it expected.
- Manipulation
- While pushing a box, the robot watches it and corrects the push if it starts to rotate.
- Humanoid walking
- A gait cannot be planned once and replayed. Balance is corrected from sensor feedback hundreds of times per second, or the robot falls.
Closed loop is essential everywhere a robot meets reality. Carry forward this distinction: executing a plan versus continuously checking whether reality still matches the plan. Closed loop does not mean replanning constantly. It means checking, so that when plan and world disagree, the system notices.
Physical AI is a systems problem
Physical AI is not simply “put a large language model (LLM) on a robot.” An LLM can help interpret a command or propose steps, but alone it cannot localize the robot, avoid a chair, grip an object, or know whether a grasp worked. It is one component. A working system needs many parts to cooperate:
- Perception
- State estimation
- World representation
- Task reasoning
- Planning
- Behavior execution
- Motion planning
- Control
- Sensing
- Verification
- Recovery
- Safety
Because the parts depend on each other, a failure in one layer spreads. Two examples:
A perception failure
- Bad perception
- Wrong world model
- Wrong plan
- Wrong action
- Task failure
A verification failure
- Bad action verification
- False belief that the task succeeded
- The planner continues
- Compounded failure
This changes how you debug. The useful question is rarely “is the AI bad?” but “which layer was wrong, and why did the layers after it not catch it?” That diagnostic habit is the thread running through FixTheRobot.
A complete Physical AI stack, at a glance
This diagram puts the pieces on one page, as a mental map for the rest of the track.
Conceptual Physical AI stack
Decide and act (top to bottom)
- Human / goal
- Language / task understanding
- Grounding
- Task planning
- Skills / behavior
- Motion planning
- Control
- Robot
Sense and know (bottom to top)
The world model feeds grounding and task planning on the left.
- Sensors
- Perception
- World model
Feedback: the robot’s actions change the world, the sensors see the change, and the cycle repeats.
This is not a mandatory architecture; different robots and research systems combine these components differently. Other FixTheRobot tracks already cover much of the map, so Physical AI does not repeat them. It focuses on how the capabilities are connected into an intelligent closed-loop system:
- Robot Foundations introduces sensing, moving, localizing and planning.
- Control covers the lower layers: making motion do what was commanded.
- Sensors covers how robots measure the world, including LiDAR point clouds.
- Perception, Localization & SLAM, and Navigation are coming soon.
Common misconceptions
Five wrong ideas appear again and again, because each sounds reasonable.
Misconception 1 “Physical AI is just an LLM controlling a robot.”
Correction LLMs can be one component. Physical AI also includes perception, state estimation, planning, execution, feedback, safety and recovery.
Misconception 2 “If the robot’s API says SUCCESS, the task succeeded.”
Correction An action result is not necessarily evidence that the intended world-state change occurred.
Misconception 3 “A robot knows the world because its sensors see it.”
Correction Sensors provide observations, not perfect knowledge of the world.
Misconception 4 “Planning is enough.”
Correction The physical world can change after planning, so execution must be monitored and plans may need to be revised.
Misconception 5 “More intelligence means putting everything into one neural network.”
Correction Physical AI systems often combine learned models with explicit planning, control, constraints and verification.
Engineering takeaways
If you remember only ten things, make them these.
- Physical AI connects intelligence to physical action.
- Sensors provide observations, not perfect knowledge.
- The robot maintains an estimate of the world.
- Planning determines what should happen.
- Control and motion planning determine how actions are physically executed.
- Physical environments introduce uncertainty and failure.
- Closed-loop feedback is fundamental.
- Action success is not necessarily task success.
- Verification and recovery are first-class components.
- Physical AI is a complete systems problem.
Knowledge check
Three questions about reasoning, not recall. Write an answer, then reveal the explanation. Your answers stay in your browser.
Question 1
A robot receives grasp("red_box") → SUCCESS, but its wrist camera cannot find the box in the gripper. What should the robot conclude?
Explanation
The command succeeded at the API level, but the intended physical outcome has not been verified. “SUCCESS” says the gripper command ran, not that the box is held. The camera points the other way, so the robot should treat the grasp as unconfirmed and not carry on as if it held the box.
A careful answer adds that one sensor can also be wrong: gather more evidence, then retry, reposition or replan if the box really is not held.
Question 2
Why can’t a robot simply plan once and execute the entire plan without observing the environment again?
Explanation
The environment and the robot’s own state can differ from the assumptions used during planning. The plan came from an approximate, out-of-date belief. Since then an object may have moved, a wheel slipped, a path been blocked, or a grasp failed. Without fresh observations, the robot cannot discover that its plan no longer fits the world.
Question 3
What is the fundamental difference between Plan → Execute and Plan → Act → Observe → Verify → Continue/Recover?
Explanation
The second is closed loop: it uses feedback to account for uncertainty and changes in the physical world. The first assumes every step works and never looks back. The second checks each step against reality and continues, recovers or replans when something is off.
What’s next
Now you have the mental model. Next, we need to understand what the robot actually knows about the world.
Where this appears in the labs
No lab is required to finish this lesson, but each idea here becomes a hands-on problem in the Physical AI labs.
- Lab 09Don’t Trust Your Actions Available now
A robot’s grasp command reports SUCCESS even when the object was not grasped. The lab turns this lesson’s verification idea into an engineering problem.
- Lab 01Open Loop vs Closed Loop In development
The same robot with and without feedback.
- Lab 02World Model as Typed State In development
Build the structured world picture that planning relies on.
- Lab 03Partial Observability In development
Turn noisy, partial observations into a belief.
Ready to see closed-loop verification in practice? Lab 09 is free and runs in your browser.
Open Lab 09Get notified when the next lesson and new labs are released.
One email per new lesson or lab.