Physical AI · Theory
T7 — Command → Goal → Task → Skill → Motion
A robot cannot directly execute most human commands. It has to translate them into increasingly executable representations.
T3 introduced the Physical AI stack, T5 explained how the robot represents the world, and T6 explained how uncertainty affects that representation. This lesson follows a human intention down into executable robot behavior.
Learning objectives
After this lesson, you should be able to:
- Explain why natural-language commands are not directly executable.
- Distinguish command, goal, task, skill and motion.
- Explain how a high-level intention becomes executable behavior.
- Decompose a simple command into tasks.
- Understand the difference between a skill and a motion.
- Explain why this hierarchy is useful in Physical AI.
- Identify where perception, planning and control fit into the hierarchy.
The one-sentence idea
A Physical AI system turns human intent into increasingly concrete representations until the robot can execute physical motion.
The levels, from most abstract to most concrete, are these. One example, “Bring me the red bottle,” runs through the whole lesson.
- Human command
- Goal
- Task
- Skill
- Motion
- Control
- Physical action
Why a robot can’t just execute a command
The user says: “Bring me the red bottle.” What exactly should the robot do? The sentence leaves out almost everything the robot needs:
- where the bottle is
- where the user is
- how to reach the bottle
- how to grasp it
- which route to take
- how to carry it
- where exactly to put it
- what to do if the bottle is not visible
- what to do if grasping fails
The command expresses intent, not an action sequence. That is the whole difference between two things:
What the user wants
The red bottle, delivered to them.
How the robot does it
Find it, route to it, grasp it, carry it, hand it over, and handle what goes wrong.
Command
A command is the human-provided instruction, or expression of intent. Some examples: “Bring me the red bottle.” “Open the door.” “Go to the kitchen.” “Pick up the cup.” “Clean the table.”
Commands are written for people, who fill in the gaps without noticing. So a command can be ambiguous (which cup?), incomplete (clean it how?), underspecified (go to the kitchen, by what route?) and context-dependent (“the table” means different tables in different rooms). T8 covers how a robot ties words to the real world and when it should ask. Here, we assume the command is clear enough to follow.
Goal
A goal describes the desired outcome, not the procedure. It is a state of the world that would satisfy the human.
Command
What the human says.
“Bring me the red bottle.”
“Open the door.”
Goal
The desired world state.
The red bottle is delivered to the user.
The door is in an open state.
Symbolic forms from T5 can express goals: holding(user, red_bottle), at(red_bottle, user) or open(door). Not every goal has to be written this way, and a system might hold it in a different representation. What matters is that a goal says where we want to end up, and leaves the route open.
Task
A task is a meaningful unit of work needed to achieve a goal. For “Bring me the red bottle” one might identify:
- Locate the bottle.
- Navigate to the bottle.
- Grasp the bottle.
- Transport the bottle.
- Navigate to the user.
- Deliver the bottle.
Tasks say what must happen, without fixing the exact physical trajectory. They can also nest: “grasp the bottle” has subtasks of its own, such as approaching, closing the gripper and lifting. How finely you slice the work is a design choice, and another designer might merge “transport” and “navigate to the user” into one. In the worked example below we use five tasks. Choosing and ordering tasks is the job of task planning, which a later lesson (T10) covers in depth.
Skill
A skill is a reusable capability that the robot knows how to execute. Typical skills:
- NavigateTo(location)
- Detect(object)
- Approach(object)
- Grasp(object)
- Place(object, location)
- Open(door)
- Follow(person)
Skills bridge abstract tasks and low-level robot behavior. The distinction to keep:
Task
What needs to be accomplished, in this situation.
“Pick up the red bottle.”
Skill
A reusable capability for that kind of task.
Grasp(red_bottle)
A skill is an abstraction, not a single motor command. Inside it there may be perception (find the bottle), motion planning (choose an approach), control (drive the arm), feedback (watch what happens) and failure detection (notice the slip).
Motion
Motion is the physical movement needed to execute a skill. What it looks like depends on the robot and the skill:
- NavigateTo
- A base trajectory, velocity commands, turning, and avoiding obstacles.
- Grasp
- An arm trajectory, end-effector movement and gripper motion.
- A humanoid walking
- Footstep placement, body motion, joint trajectories and balance control.
So for the skill “grasp the bottle,” the motion is to move the arm and end-effector (the hand or tool at the arm’s tip) toward the bottle, close the gripper, and follow the required trajectory. Motion planning methods come in a later lesson.
Control
A motion is only a desire until something makes the hardware follow it. Control does that, despite dynamics, disturbances and modeling errors.
- Goal
- Task
- Skill
- Motion / trajectory
- Control
- Actuators
- Robot
If you have worked through the Control labs, this is the layer you were tuning. Typical techniques include:
- PID
- Feedback that corrects a command from the error, its history and its trend.
- MPC (model predictive control)
- Choosing the next few control actions by predicting the outcome with a model.
- Whole-body control
- Coordinating all of a humanoid’s joints to achieve several goals at once.
- Impedance control
- Making contact behave like a spring and damper, so touching things stays gentle.
None of these is taught here.
Complete decomposition: the red bottle
Here is the central example: one command, made more concrete at every level.
- Command“Bring me the red bottle.”
- GoalThe red bottle is delivered to the user.
- Tasks1. Locate the red bottle. 2. Navigate to it. 3. Grasp it. 4. Navigate to the user. 5. Deliver it.
- SkillsSearch, Detect, NavigateTo, Approach, Grasp, Carry, Place or HandOver.
- MotionBase trajectory, arm trajectory, gripper motion, and the final handover motion.
- ControlVelocity control, joint control and gripper control.
One command can have many valid plans
The hierarchy does not fix a single sequence of actions. The robot might take either of these routes:
Plan A
- Kitchen
- Bottle
- Living room
- User
Plan B
- Kitchen
- Bottle
- Hallway
- User
Both satisfy the goal. The goal constrains the outcome; planning chooses one feasible way to achieve it. Which one, and how, is the subject of a later lesson (T10).
Skills are not fixed motions
NavigateTo(kitchen) does not mean “always send these exact wheel velocities.” The motion the skill produces depends on:
- Robot location
- Obstacles
- The map
- Goal location
- Robot dynamics
- Current belief
Likewise Grasp(red_bottle) does not imply one fixed arm trajectory. If the bottle is lying down, or tucked behind a cup, the arm must move differently. The skill stays the same while the motion adapts to the actual physical situation, which is exactly what makes a skill a skill, and not a recording.
Where perception fits
The robot cannot execute Grasp(red_bottle) unless it knows enough: where the bottle is, whether it is reachable, how it is oriented, and whether something blocks it. Those answers come from the world model and belief of T5 and T6.
- Perception
- World model / belief
- Task / skill decision
- Motion
- Control
So the hierarchy is not a one-way street. It looks like a chain from command to control, but information keeps flowing back up from the world.
Feedback closes the loop
- Command
- Goal
- Task
- Skill
- Motion
- Robot
- Sensors
- Observation
- World model / belief
After an action, the world may have changed. The action may have failed, the object may have moved, and the robot may be somewhere unexpected. The robot may then need to reconsider the task, or the plan, and not just carry on down the list. This is the loop from T6, applied to the whole hierarchy. Later lessons on execution and recovery build on it.
A failure example
The task is “grasp the red bottle,” carried out by the skill Grasp(red_bottle). The robot executes the arm motion and the gripper closes. But the bottle slips out. Three different things happened, and they must not be confused:
| Level | What it means | In the example |
|---|---|---|
| Command success | The call to the skill was accepted and returned SUCCESS | Yes |
| Skill execution | The skill ran: the arm moved and the gripper closed | Yes |
| Task success | The intended effect happened: the bottle is held | No |
The first two succeeded while the task failed. This is why verification matters, and it is the subject of T13, “Command Success ≠ Task Success.” You can also practice the idea in Lab 09.
Why the hierarchy matters
Splitting the work into levels buys a lot:
- Abstraction
- Modularity
- Reuse
- Debugging
- Planning
- Verification
- Recovery
- Easier integration across robots
For example, a high-level task can call NavigateTo(kitchen) without caring whether the robot has differential drive, Ackermann steering, four legs or two. The skill’s implementation changes from robot to robot, while the task above it can stay almost the same. And when something fails, the levels tell you where to look: was the command misread, the plan wrong, the skill faulty, or the motion poorly controlled?
Connection to the Physical AI stack
This lesson is the bridge between human intent and physical execution in the stack from T3:
- Human / goal
- Command understanding
- Grounding
- World model / belief
- Task planning
- Skills / behavior
- Motion planning
- Control
- Robot
- Sensors
- Feedback
Common misconceptions
Misconception 1 “A command is already a plan.”
Correction A command specifies intent. A plan determines how to achieve it.
Misconception 2 “A task and a skill are the same.”
Correction A task is something that needs to be accomplished. A skill is a reusable capability for accomplishing a class of tasks.
Misconception 3 “A skill is just one motor command.”
Correction A skill may contain perception, planning, control and feedback.
Misconception 4 “Motion is the same as planning.”
Correction Planning determines what should happen. Motion is the physical trajectory or behavior used to execute it.
Misconception 5 “Once the robot starts executing, the hierarchy no longer matters.”
Correction Feedback can cause the robot to revise the task, the skill or the plan.
Misconception 6 “There is only one correct decomposition of a command.”
Correction Multiple task sequences and motion strategies can achieve the same goal.
Engineering takeaways
- Human commands express intent, not complete robot procedures.
- Goals describe desired outcomes.
- Tasks describe meaningful units of work.
- Skills provide reusable robot capabilities.
- Motion is the physical realization of skills.
- Control turns desired behavior into actuator-level actions.
- Perception and feedback continuously influence execution.
- Good Physical AI connects these levels without confusing their responsibilities.
Physical AI turns human intent into progressively more concrete representations (command, goal, task, skill, motion, control) while feedback continuously connects execution back to perception and decision-making.
Knowledge check
Three conceptual questions. Write an answer, then reveal the explanation. Your answers stay in your browser.
Question 1
Given the command “Bring me the red bottle,” what is the goal?
Explanation
The goal is the desired world state: the red bottle has been delivered to the user, for instance holding(user, red_bottle) or at(red_bottle, user). The goal is not the list of steps (find, grasp, carry). It says where things should end up, and it leaves the route open.
Question 2
What is the difference between a task and a skill?
Explanation
A task is something that needs to be accomplished in a particular situation. A skill is a reusable capability for accomplishing a class of tasks. “Pick up the red bottle” is a task. Grasp(object) is a skill that can carry out that task and many similar ones. The task says what, and the skill is the reusable know-how for doing it.
Question 3
Given the skill NavigateTo(kitchen), what belongs to the skill, and what belongs to the resulting motion?
Explanation
The skill is the reusable capability and the logic behind it; the motion is the actual movement it produces in this situation. The skill covers things like localizing, choosing a path, avoiding obstacles, monitoring progress and detecting failure. The motion is the specific trajectory and velocities that result, which differ with the robot’s location, the obstacles, the map and its belief.
What’s next
So far we assumed the command was clear. Humans often give commands that contain ambiguity or missing information.
Related content
Nothing below is required to finish this lesson.
- LessonT3: The Physical AI Stack at a Glance
The whole stack this hierarchy runs through.
- LessonT5: World Models and Symbolic State
How goals and states can be written down.
- LessonT6: Belief Under Uncertainty
Why skills must cope with what the robot is unsure of.
- TrackRobot Foundations
Sensing, moving, localizing and planning for beginners.
- TrackControl
The layer that makes motion happen.
- PlanningPlanning Coming later
Path planning arrives with the Navigation track; task planning with a later theory lesson.
- Lab 04Command → Structured Task In development
Turn a command into structured data a robot can use.
- Lab 05Ground and Clarify In development
Resolve ambiguous references, and ask when you must.
Get notified when the next lesson and new labs are released.
One email per new lesson or lab.