Skip to the lesson
Robotics Lab Explore Robotics / Physical AI / Theory / T7

Physical AI · Theory

T7 — Command → Goal → Task → Skill → Motion

A robot cannot directly execute most human commands. It has to translate them into increasingly executable representations.

  • Lesson 7
  • Physical AI Theory
  • Hierarchy
  • About 11 minutes to read

T3 introduced the Physical AI stack, T5 explained how the robot represents the world, and T6 explained how uncertainty affects that representation. This lesson follows a human intention down into executable robot behavior.

Learning objectives

After this lesson, you should be able to:

  • Explain why natural-language commands are not directly executable.
  • Distinguish command, goal, task, skill and motion.
  • Explain how a high-level intention becomes executable behavior.
  • Decompose a simple command into tasks.
  • Understand the difference between a skill and a motion.
  • Explain why this hierarchy is useful in Physical AI.
  • Identify where perception, planning and control fit into the hierarchy.

The one-sentence idea

A Physical AI system turns human intent into increasingly concrete representations until the robot can execute physical motion.

The levels, from most abstract to most concrete, are these. One example, “Bring me the red bottle,” runs through the whole lesson.

  1. Human command
  2. Goal
  3. Task
  4. Skill
  5. Motion
  6. Control
  7. Physical action
Each level is more concrete than the one above it.

Why a robot can’t just execute a command

The user says: “Bring me the red bottle.” What exactly should the robot do? The sentence leaves out almost everything the robot needs:

  • where the bottle is
  • where the user is
  • how to reach the bottle
  • how to grasp it
  • which route to take
  • how to carry it
  • where exactly to put it
  • what to do if the bottle is not visible
  • what to do if grasping fails

The command expresses intent, not an action sequence. That is the whole difference between two things:

What the user wants

The red bottle, delivered to them.

How the robot does it

Find it, route to it, grasp it, carry it, hand it over, and handle what goes wrong.

The human supplies the what. The robot has to supply the how.

Command

A command is the human-provided instruction, or expression of intent. Some examples: “Bring me the red bottle.” “Open the door.” “Go to the kitchen.” “Pick up the cup.” “Clean the table.”

Commands are written for people, who fill in the gaps without noticing. So a command can be ambiguous (which cup?), incomplete (clean it how?), underspecified (go to the kitchen, by what route?) and context-dependent (“the table” means different tables in different rooms). T8 covers how a robot ties words to the real world and when it should ask. Here, we assume the command is clear enough to follow.

Goal

A goal describes the desired outcome, not the procedure. It is a state of the world that would satisfy the human.

Command

What the human says.

“Bring me the red bottle.”

“Open the door.”

Goal

The desired world state.

The red bottle is delivered to the user.

The door is in an open state.

A command is what is said. A goal is the state of the world that would make it true.

Symbolic forms from T5 can express goals: holding(user, red_bottle), at(red_bottle, user) or open(door). Not every goal has to be written this way, and a system might hold it in a different representation. What matters is that a goal says where we want to end up, and leaves the route open.

Task

A task is a meaningful unit of work needed to achieve a goal. For “Bring me the red bottle” one might identify:

  1. Locate the bottle.
  2. Navigate to the bottle.
  3. Grasp the bottle.
  4. Transport the bottle.
  5. Navigate to the user.
  6. Deliver the bottle.

Tasks say what must happen, without fixing the exact physical trajectory. They can also nest: “grasp the bottle” has subtasks of its own, such as approaching, closing the gripper and lifting. How finely you slice the work is a design choice, and another designer might merge “transport” and “navigate to the user” into one. In the worked example below we use five tasks. Choosing and ordering tasks is the job of task planning, which a later lesson (T10) covers in depth.

Skill

A skill is a reusable capability that the robot knows how to execute. Typical skills:

  • NavigateTo(location)
  • Detect(object)
  • Approach(object)
  • Grasp(object)
  • Place(object, location)
  • Open(door)
  • Follow(person)

Skills bridge abstract tasks and low-level robot behavior. The distinction to keep:

Task

What needs to be accomplished, in this situation.

“Pick up the red bottle.”

Skill

A reusable capability for that kind of task.

Grasp(red_bottle)

The task is one job. The skill is the reusable know-how that does that job and similar ones.

A skill is an abstraction, not a single motor command. Inside it there may be perception (find the bottle), motion planning (choose an approach), control (drive the arm), feedback (watch what happens) and failure detection (notice the slip).

Motion

Motion is the physical movement needed to execute a skill. What it looks like depends on the robot and the skill:

NavigateTo
A base trajectory, velocity commands, turning, and avoiding obstacles.
Grasp
An arm trajectory, end-effector movement and gripper motion.
A humanoid walking
Footstep placement, body motion, joint trajectories and balance control.

So for the skill “grasp the bottle,” the motion is to move the arm and end-effector (the hand or tool at the arm’s tip) toward the bottle, close the gripper, and follow the required trajectory. Motion planning methods come in a later lesson.

Control

A motion is only a desire until something makes the hardware follow it. Control does that, despite dynamics, disturbances and modeling errors.

  1. Goal
  2. Task
  3. Skill
  4. Motion / trajectory
  5. Control
  6. Actuators
  7. Robot
Control sits between desired motion and the actuators.

If you have worked through the Control labs, this is the layer you were tuning. Typical techniques include:

PID
Feedback that corrects a command from the error, its history and its trend.
MPC (model predictive control)
Choosing the next few control actions by predicting the outcome with a model.
Whole-body control
Coordinating all of a humanoid’s joints to achieve several goals at once.
Impedance control
Making contact behave like a spring and damper, so touching things stays gentle.

None of these is taught here.

Complete decomposition: the red bottle

Here is the central example: one command, made more concrete at every level.

  1. Command“Bring me the red bottle.”
  2. GoalThe red bottle is delivered to the user.
  3. Tasks1. Locate the red bottle. 2. Navigate to it. 3. Grasp it. 4. Navigate to the user. 5. Deliver it.
  4. SkillsSearch, Detect, NavigateTo, Approach, Grasp, Carry, Place or HandOver.
  5. MotionBase trajectory, arm trajectory, gripper motion, and the final handover motion.
  6. ControlVelocity control, joint control and gripper control.
The same intention at six levels of detail, from what was said to what the actuators do.

One command can have many valid plans

The hierarchy does not fix a single sequence of actions. The robot might take either of these routes:

Plan A

  1. Kitchen
  2. Bottle
  3. Living room
  4. User

Plan B

  1. Kitchen
  2. Bottle
  3. Hallway
  4. User
Both plans reach the same goal by different routes.

Both satisfy the goal. The goal constrains the outcome; planning chooses one feasible way to achieve it. Which one, and how, is the subject of a later lesson (T10).

Skills are not fixed motions

NavigateTo(kitchen) does not mean “always send these exact wheel velocities.” The motion the skill produces depends on:

  • Robot location
  • Obstacles
  • The map
  • Goal location
  • Robot dynamics
  • Current belief

Likewise Grasp(red_bottle) does not imply one fixed arm trajectory. If the bottle is lying down, or tucked behind a cup, the arm must move differently. The skill stays the same while the motion adapts to the actual physical situation, which is exactly what makes a skill a skill, and not a recording.

Where perception fits

The robot cannot execute Grasp(red_bottle) unless it knows enough: where the bottle is, whether it is reachable, how it is oriented, and whether something blocks it. Those answers come from the world model and belief of T5 and T6.

  1. Perception
  2. World model / belief
  3. Task / skill decision
  4. Motion
  5. Control
Information flows into every decision the hierarchy makes.

So the hierarchy is not a one-way street. It looks like a chain from command to control, but information keeps flowing back up from the world.

Feedback closes the loop

  1. Command
  2. Goal
  3. Task
  4. Skill
  5. Motion
  6. Robot
  7. Sensors
  8. Observation
  9. World model / belief
The hierarchy runs downward; feedback runs back up through sensors, observation and belief.

After an action, the world may have changed. The action may have failed, the object may have moved, and the robot may be somewhere unexpected. The robot may then need to reconsider the task, or the plan, and not just carry on down the list. This is the loop from T6, applied to the whole hierarchy. Later lessons on execution and recovery build on it.

A failure example

The task is “grasp the red bottle,” carried out by the skill Grasp(red_bottle). The robot executes the arm motion and the gripper closes. But the bottle slips out. Three different things happened, and they must not be confused:

Command success, skill execution and task success compared
LevelWhat it meansIn the example
Command successThe call to the skill was accepted and returned SUCCESSYes
Skill executionThe skill ran: the arm moved and the gripper closedYes
Task successThe intended effect happened: the bottle is heldNo

The first two succeeded while the task failed. This is why verification matters, and it is the subject of T13, “Command Success ≠ Task Success.” You can also practice the idea in Lab 09.

Why the hierarchy matters

Splitting the work into levels buys a lot:

  • Abstraction
  • Modularity
  • Reuse
  • Debugging
  • Planning
  • Verification
  • Recovery
  • Easier integration across robots

For example, a high-level task can call NavigateTo(kitchen) without caring whether the robot has differential drive, Ackermann steering, four legs or two. The skill’s implementation changes from robot to robot, while the task above it can stay almost the same. And when something fails, the levels tell you where to look: was the command misread, the plan wrong, the skill faulty, or the motion poorly controlled?

Connection to the Physical AI stack

This lesson is the bridge between human intent and physical execution in the stack from T3:

  1. Human / goal
  2. Command understanding
  3. Grounding
  4. World model / belief
  5. Task planning
  6. Skills / behavior
  7. Motion planning
  8. Control
  9. Robot
  10. Sensors
  11. Feedback
The stack, with this lesson’s five levels running through the middle of it.

Common misconceptions

Misconception 1 “A command is already a plan.”

Correction A command specifies intent. A plan determines how to achieve it.

Misconception 2 “A task and a skill are the same.”

Correction A task is something that needs to be accomplished. A skill is a reusable capability for accomplishing a class of tasks.

Misconception 3 “A skill is just one motor command.”

Correction A skill may contain perception, planning, control and feedback.

Misconception 4 “Motion is the same as planning.”

Correction Planning determines what should happen. Motion is the physical trajectory or behavior used to execute it.

Misconception 5 “Once the robot starts executing, the hierarchy no longer matters.”

Correction Feedback can cause the robot to revise the task, the skill or the plan.

Misconception 6 “There is only one correct decomposition of a command.”

Correction Multiple task sequences and motion strategies can achieve the same goal.

Engineering takeaways

  1. Human commands express intent, not complete robot procedures.
  2. Goals describe desired outcomes.
  3. Tasks describe meaningful units of work.
  4. Skills provide reusable robot capabilities.
  5. Motion is the physical realization of skills.
  6. Control turns desired behavior into actuator-level actions.
  7. Perception and feedback continuously influence execution.
  8. Good Physical AI connects these levels without confusing their responsibilities.

Physical AI turns human intent into progressively more concrete representations (command, goal, task, skill, motion, control) while feedback continuously connects execution back to perception and decision-making.

Knowledge check

Three conceptual questions. Write an answer, then reveal the explanation. Your answers stay in your browser.

Question 1

Given the command “Bring me the red bottle,” what is the goal?

Explanation

The goal is the desired world state: the red bottle has been delivered to the user, for instance holding(user, red_bottle) or at(red_bottle, user). The goal is not the list of steps (find, grasp, carry). It says where things should end up, and it leaves the route open.

Question 2

What is the difference between a task and a skill?

Explanation

A task is something that needs to be accomplished in a particular situation. A skill is a reusable capability for accomplishing a class of tasks. “Pick up the red bottle” is a task. Grasp(object) is a skill that can carry out that task and many similar ones. The task says what, and the skill is the reusable know-how for doing it.

Question 3

Given the skill NavigateTo(kitchen), what belongs to the skill, and what belongs to the resulting motion?

Explanation

The skill is the reusable capability and the logic behind it; the motion is the actual movement it produces in this situation. The skill covers things like localizing, choosing a path, avoiding obstacles, monitoring progress and detecting failure. The motion is the specific trajectory and velocities that result, which differ with the robot’s location, the obstacles, the map and its belief.

What’s next

So far we assumed the command was clear. Humans often give commands that contain ambiguity or missing information.

Next lesson

T8 — Grounding and When to Ask

Take “Bring me the bottle.” Which bottle? T8 explains how Physical AI connects language to the actual world, and how it decides when it needs clarification.

Nothing below is required to finish this lesson.

  • Lesson
    T3: The Physical AI Stack at a Glance

    The whole stack this hierarchy runs through.

  • Lesson
    T5: World Models and Symbolic State

    How goals and states can be written down.

  • Lesson
    T6: Belief Under Uncertainty

    Why skills must cope with what the robot is unsure of.

  • Track
    Robot Foundations

    Sensing, moving, localizing and planning for beginners.

  • Track
    Control

    The layer that makes motion happen.

  • Planning
    Planning Coming later

    Path planning arrives with the Navigation track; task planning with a later theory lesson.

  • Lab 04
    Command → Structured Task In development

    Turn a command into structured data a robot can use.

  • Lab 05
    Ground and Clarify In development

    Resolve ambiguous references, and ask when you must.