Physical AI · Theory
T3 — The Physical AI Stack at a Glance
How perception, reasoning, planning, control, and feedback become one intelligent robotic system.
In T1 we saw why the physical world is hard, and in T2 we saw what a robot actually knows. This lesson steps back and draws the whole map.
Learning objectives
After this lesson, you should be able to:
- Describe the major layers of a Physical AI system.
- Explain the purpose of each layer.
- Distinguish task planning from motion planning and control.
- Explain how information flows from sensors to decisions and actions.
- Explain why Physical AI is a closed-loop system.
- Identify where perception, SLAM, navigation and control fit in the larger architecture.
The goal is the map, not the details. Each layer gets its own lessons or tracks, and none is taught in depth here.
The one-sentence idea
A Physical AI system connects high-level goals to physical action through a hierarchy of perception, representation, planning, execution, control, and feedback.
A hierarchy here means that each layer works at one level of detail. The top layers think in goals (“the bottle is with the user”). The bottom layers think in motor commands. Each layer takes a decision from above and makes it more concrete for the layer below, while information about the world flows back up. Feedback ties the two directions into a loop.
Why we need a stack
Suppose a person says: “Bring the red bottle from the kitchen to me.” A robot cannot solve this with a single operation. It has to settle a series of different questions:
- What does the command mean?
- Which object is the red bottle?
- Where is it?
- Where is the robot?
- How should it get there?
- How should it grasp the bottle?
- How should it carry it?
- What if the grasp fails?
- What if the environment changes?
- How does it know the task is complete?
These are different kinds of problems, solved at different levels of detail and at different speeds. The engineering answer is a stack: split the job into layers, give each layer one responsibility, and define what passes between them.
The Physical AI stack
Read the diagram from the top: it shows how a goal becomes motion. Then follow the return path at the bottom, where the robot’s sensors read the result and feed it back.
Conceptual Physical AI stack
- HumanGives a goal or command.
- Task understanding & groundingTurns words into a goal tied to real objects and places.
- PerceptionTurns sensor measurements into useful observations.
- World model / beliefHolds what the robot currently believes about the world.
- Task planningDecides what should happen, as a sequence of steps.
- Skills / behaviorRuns reusable actions such as navigate and grasp.
- Motion planningFinds a feasible, collision-free movement.
- ControlMakes the actuators follow the desired motion.
- RobotMoves in the real world.
- SensorsMeasure the result of the robot’s motion.
- FeedbackBack to perception and the world model, and the loop repeats.
This is a conceptual architecture, not a mandatory one. Real robots merge layers, skip some, or reorder them. A learned policy might fold several layers into one neural network, and many layers run at once rather than in sequence. The map is still useful because it lets you ask, for any problem, “which layer is responsible for this?”
Layer 1: Human goal / task
The question it answers: What does the human want?
The stack starts with an objective, usually in words: “Bring the red bottle to the user.” Three terms are easy to blur. A command is what the human says. A goal is the world state that would satisfy them (the bottle is with the user). A task is the work needed to reach that goal. The system must turn the first into a machine-understandable form of the second. A later lesson covers this in detail.
Layer 2: Grounding and task understanding
The question it answers: Which real things and places do these words mean?
Language is not tied to the physical world until something ties it. “Red bottle” must refer to an actual object, and “to me” must refer to a person or a place. This takes four ideas: object grounding (which object?), spatial grounding (which place?), task interpretation (what does “bring” involve?) and ambiguity (what if two bottles are red, so the robot should ask?). Language models can help here, but they are one tool, and a later lesson looks at them properly.
Layer 3: Perception
The question it answers: What is out there, according to the sensors?
Perception converts sensor measurements into useful observations. A camera yields objects, LiDAR (a laser scanner) yields geometry, an IMU (inertial measurement unit) yields motion measurements, and encoders yield wheel or joint measurements.
- Sensors
- Measurements
- Perception
- Useful observations
Go deeper: the Sensors track covers how robots measure the world. The Perception track is coming soon, and T4 follows this step in detail.
Layer 4: World model / belief
The question it answers: What does the robot currently believe?
This layer connects directly to T2. Observations pass through state estimation and reasoning to produce the robot’s current understanding of the world:
robot = kitchen red_bottle = table gripper = empty door = open
This picture may carry uncertainty: the bottle is probably on the table, not certainly. Everything above this layer reads the world model, not the raw sensors, so it is the information that higher-level reasoning actually runs on.
Layer 5: Task planning
The question it answers: What should the robot do?
Given a goal and the current belief, task planning produces a sequence: find the bottle, navigate to it, grasp it, verify the grasp, navigate to the user, place the bottle, verify completion. It works with goals, actions, preconditions (what must be true before an action, such as “the gripper is empty” before grasping), effects (what is true after) and task sequences. Planning algorithms are not taught here.
Layer 6: Skills / behavior
The question it answers: How is each step carried out?
A skill is a reusable robot capability with a clear interface, such as navigate(), grasp(), place(), open() or close(). A high-level planner should not command individual motors. It calls skills, and each skill uses the lower-level system.
- Planner
- Skill
- Lower-level robotics system
Skills are often organized with behavior trees, a way to arrange skills, conditions and fallbacks so that a failed step triggers a defined response. We only name them here; a later lesson explains how they work.
Layer 7: Motion planning
The question it answers: What physical movement should the robot make?
Task planning decides what to do: “Navigate to the kitchen.” Motion planning decides how the body should move: which collision-free path to take. For the task “reach the table,” the motion planner chooses a route around the chair. The two are separate jobs, solved separately and at different speeds.
Go deeper: path-planning methods belong to the Navigation track, which is coming soon.
Layer 8: Control
The question it answers: How do I make the actuators do it?
Planning produces a desired motion. Control makes the robot physically follow it, correcting for friction, load and error using feedback.
- Desired trajectory
- Controller
- Actuators
- Robot
For example, a planner says “move forward at 0.5 m/s.” The controller works out the motor commands needed to actually achieve that speed, and adjusts them as the wheels slip or the load changes.
Go deeper: the Control labs teach how controllers are tuned and made safe.
Layer 9: Physical robot
The question it answers: What actually happens?
At the bottom, the system meets hardware: motors, wheels, joints, grippers and sensors. Physical execution brings noise, delay, friction, dynamics, actuator limits and unexpected events. This is where the idealized plan meets reality, and why the layers above must keep checking what really happened.
Feedback: the stack is a loop
This is the most important idea in the lesson. The stack is not simply input, then AI, then robot. It is a loop in which every action produces new observations.
Not this
- Input
- AI
- Robot
A straight line from command to motion, with no check on the result.
This
- Goal
- Reason
- Act
- Observe
- Update
- Reason again
Every action is checked, and the system reasons again.
Inside the loop, each cycle looks like this:
- Plan
- Act
- Observe
- Update world model
- Verify
- Continue, recover or replan
- Act again
Physical AI is closed-loop intelligence.
This is the loop from T1, now placed on the stack. The observation step feeds the belief you met in T2, and the verify step is exactly what Lab 09 asks you to build.
One complete example
Follow “Bring the red bottle from the kitchen to the user” through the whole stack.
| Step | Layer | What happens |
|---|---|---|
| 1 | Human | Gives the command. |
| 2 | Task understanding | Determines the goal: the red bottle is with the user. |
| 3 | Grounding | Identifies the physical red bottle, and where “the user” is. |
| 4 | Perception | Detects the bottle and the environment from sensor data. |
| 5 | World model | Represents the robot’s location, the bottle’s location, obstacles and the user’s location. |
| 6 | Task planning | Creates the plan: navigate, grasp, verify, navigate, place, verify. |
| 7 | Skills | The skill executor runs the steps in order. |
| 8 | Motion planning | Generates the physical paths and arm movements. |
| 9 | Control | Executes the desired motion. |
| 10 | Sensors | Observe the result. |
| 11 | World model | Updates: is the bottle in the gripper? |
| 12 | Recovery | If the grasp failed: recover, retry or replan. |
| 13 | Verification | Establishes whether the goal was actually achieved. |
Steps 10 to 13 are the ones a plan-then-execute design leaves out, and they are what make the system robust.
Where existing robotics topics fit
Many of these layers are already subjects on FixTheRobot. They form a pipeline of capabilities:
- Sensors
- Perception
- Localization / SLAM
- World understanding
- Navigation / motion planning
- Control
- Robot
Physical AI sits above and across these capabilities. They are not isolated subjects. Physical AI is where the system learns to coordinate them toward goals. Here is how the tracks map to the stack, with nothing duplicated:
- Robot Foundations
- Start here: sensing, moving, localizing and planning for beginners.
- Control
- Layer 8: making motion follow a command.
- Sensors
- The measurements that feed Layer 3.
- Perception, Localization & SLAM, Navigation
- Layers 3, 4 and 7 in depth. All three are coming soon.
- Physical AI
- The coordination across all layers, toward goals.
Common misconceptions
Misconception 1 “Physical AI is just an LLM connected to a robot.”
Correction A large language model (LLM) can be one component, for example in grounding or task planning. It is not the complete system. The other layers still have to perceive, plan motion, control and verify.
Misconception 2 “Task planning and motion planning are the same.”
Correction Task planning decides what to do, as a sequence of steps. Motion planning decides how the body should move to carry out one step.
Misconception 3 “The planner controls the motors.”
Correction There is a hierarchy: task planning calls skills, skills use motion planning, and control drives the actuators. Each layer hands a more concrete request downward.
Misconception 4 “If the robot executes the planned actions, the task is complete.”
Correction Executing actions is not the same as achieving the goal. The robot must observe the result and verify it, then recover or replan if needed.
Misconception 5 “More intelligence belongs in the highest-level AI model.”
Correction Different layers solve different engineering problems. Keeping a robot upright, avoiding a collision and verifying a grasp are often best handled by dedicated methods, not by a single high-level model.
Engineering takeaways
- Physical AI is a complete system, not a single model.
- Sensors provide observations.
- Perception extracts useful information.
- A world model represents the robot’s current understanding.
- Task planning determines what should happen.
- Skills and behaviors execute tasks.
- Motion planning determines physically feasible movement.
- Control makes the robot follow the desired behavior.
- Feedback continuously updates the system.
- Verification, recovery and replanning make the system robust.
Knowledge check
Three conceptual questions. Write an answer, then reveal the explanation. Your answers stay in your browser.
Question 1
What is the difference between task planning, motion planning and control?
Explanation
Task planning decides what to do, motion planning decides how the body should move, and control makes the actuators follow that motion. For example: “navigate to the kitchen” (task), a collision-free path around the chair (motion plan), and the wheel commands that follow the path despite slip (control). Each hands a more concrete request to the next.
Question 2
Why does the Physical AI stack need feedback instead of simply executing a plan from beginning to end?
Explanation
Because the world and the robot’s execution can differ from what the plan assumed. Sensors are noisy, actuators are imperfect and the environment changes. Feedback lets the system observe what actually happened, update its world model, and continue, recover or replan, instead of acting on a plan that no longer fits reality.
Question 3
A robot’s planner believes the bottle is on the table, but a new camera observation shows that the table is empty. Which part of the system should be updated before continuing?
Explanation
The robot’s world model, or belief, should be updated, which may trigger replanning. The plan was built on a belief that the bottle was on the table. Once the observation contradicts it, the belief must change first. Then the planner can decide whether the old plan is still valid or whether to search for the bottle.
What’s next
You now have the map of the Physical AI system. The next lessons will zoom into each part of this architecture.
Get notified when the next lesson and new labs are released.
One email per new lesson or lab.