Physical AI · Theory
T8 — Grounding and When to Ask
Understanding language is not enough. The robot must connect it to the actual entities and situations in the physical world.
T5 introduced the world model, T6 explained uncertainty and belief, and T7 showed how commands become goals, tasks and skills. This lesson explains how a command is connected to the actual world, and when the robot should ask for clarification.
Learning objectives
After this lesson, you should be able to:
- Explain what grounding means in Physical AI.
- Distinguish language understanding from grounding.
- Ground references such as objects, locations and people to the world.
- Recognize ambiguity in human commands.
- Understand when a command contains insufficient information.
- Explain why the robot should sometimes ask rather than guess.
- Understand how belief and grounding interact.
- Explain the tradeoff between asking and acting.
One example runs through the lesson: “Bring me the bottle.” The robot has to work out which bottle, where it is, whether it can see it, and whether it knows enough to go.
- Human command
- Interpret
- Ground to the world
- Check whether the information is sufficient
Act
The command is grounded well enough.
Ask
It is not, and the difference matters.
The one-sentence idea
Grounding means connecting what a person says to what actually exists and matters in the robot’s world.
When the robot cannot reliably determine what the user means, asking is often better than guessing.
What is grounding?
Grounding is connecting the symbols and references in language to actual things in the robot’s world. Language is made of words that point at things. The robot operates in a physical environment, so something has to tie the words to that environment. The user says “Bring me the red bottle,” and the robot must connect “red bottle” to a real object:
- Language: “red bottle”
- Grounding
- World entity: red_bottle_01
Grounding can involve many kinds of references:
- Objects
- People
- Locations
- Spatial relationships
- Actions
- Properties
- Task references
Understanding vs grounding
These sound alike and are different. Suppose there are two red bottles in the room.
Language understanding
“I understand that the user is referring to a red bottle.”
Grounding
“I know which physical red bottle the user means.”
Language understanding ≠ grounding.
Grounding to objects
References are mapped onto the objects in the world model from T5. Say the user asks, “Pick up the red bottle,” and the world model holds red_bottle_01, blue_bottle_01, red_bottle_02 and cup_01. Matching color and type leaves a set of candidates:
- “red bottle”
- Candidate objects
- red_bottle_01 · red_bottle_02
If only one red bottle exists, grounding is straightforward. With two, it becomes ambiguous, and the robot must find out whether any additional context settles it.
Grounding to locations
Places in language also need to connect to places in the world. “Go to the kitchen.” “Put it on the table.” “Take it to the room.” “Move it near the sofa.” Each must map to an entity or region in the world model:
- “kitchen”
room_kitchen- “table”
table_01- “near the sofa”
- A spatial relation involving
sofa_01
Some are easy and some are not. “The room” could be any room, and “near” is a matter of degree.
Grounding spatial relationships
People constantly describe things by where they are: on the table, next to the chair, inside the cabinet, near the door, behind the sofa, in front of the robot. These connect to the relations of the symbolic world model:
on(red_bottle, table) inside(cup, cabinet) near(robot, sofa)
Grounding therefore often requires reasoning about relationships, not just names. “The bottle on the table” picks out the bottle for which on(bottle, table) holds. And some relations depend on a viewpoint: “in front of the robot” depends on where the robot is facing, while “behind the sofa” depends on which side the speaker means.
Grounding people
People are referred to by words that shift with the situation: “me,” “you,” “the person,” “that guy,” “the user.” In “Bring it to me,” the robot must ground “me” to whoever is speaking, which is the current user. In “Give it to the person near the door,” it must connect the words to a particular person using a spatial clue. Both are the same kind of problem as grounding an object: find the entity in the world that the words point to.
Ambiguity
Ambiguity is one of the central problems in Physical AI. If the robot sees a red bottle, a blue bottle and a green bottle, “Bring me the bottle” leaves open which one. If there are several tables, “Put it on the table” does too. Ambiguity can arise from:
- multiple matching objects,
- missing information,
- vague references,
- uncertain perception,
- a changing environment,
- an incomplete world model.
Types of ambiguity
- Referential
- “Pick up the bottle.” There are several bottles.
- Spatial
- “Put it on the table.” There are several tables.
- Attribute
- “Bring the large cup.” Two cups might both count as large, depending on what perception reports.
- Temporal
- “Bring the bottle I used earlier.” The robot may not know which earlier event is meant.
- Contextual
- “Put it there.” What does “there” refer to?
Knowing the type helps decide what to do. A referential ambiguity can often be answered with a short question; a contextual one may need the user to point.
Grounding under uncertainty
Grounding is not always a yes or no. As in T6, the robot may have high confidence that one object is meant, two plausible candidates, incomplete observations, or stale information. Say the user asks for “the red bottle,” and the robot’s belief about which one is meant looks like this:
Is 85% enough? That depends on the task consequences, safety, the cost of being wrong, and the ability to verify later. If the wrong bottle only means a short detour that can be corrected on arrival, the robot may well proceed. If it means moving something fragile or dangerous, it may not. There is no universal confidence threshold.
When should the robot ask?
The core decision is between acting and asking. The robot should consider asking when:
- multiple interpretations are plausible,
- the difference between them matters,
- acting on the wrong one could cause failure,
- it lacks information the task requires,
- the user’s intent cannot be reliably grounded, or
- the cost of asking is lower than the cost of being wrong.
Say the user asks, “Bring me the bottle,” and two bottles exist. The robot should ask, and ask specifically: “Which bottle do you mean, the red one or the blue one?” A good question offers the candidates, so the user can answer in a word.
When should it not ask?
This is equally important. A robot that asked about every minor doubt would be unusable. Compare:
| The user says | The situation | Decision |
|---|---|---|
| “Go to the kitchen.” | Only one kitchen exists | Do not ask |
| “Bring me the red bottle.” | One red bottle, clearly detected | Usually do not ask |
| “Put the bottle on the table.” | One relevant table in the task context | Probably do not ask |
| “Bring me the bottle.” | A red and a blue bottle | Ask |
Ask when ambiguity materially affects the task, not whenever uncertainty exists.
Ask vs infer from context
Before asking, a robot can try to use context. Two red bottles exist, but earlier the user said, “The bottle on the kitchen table.” That settles it, with no question needed.
- Command
- Context
- World model
- Grounded interpretation
Grounding is therefore contextual. It is more than keyword matching: the same words can point at different things depending on what was said before, what the robot is doing and where it is.
Grounding and the world model
The world model supplies the entities and relationships that language is grounded against. Suppose the user says, “Put the bottle on the table.” The model contains:
Objects: red_bottle_01, blue_bottle_01, table_01, table_02 Relations: on(red_bottle_01, kitchen_counter) inside(robot, kitchen) inside(table_01, kitchen) inside(table_02, living_room)
“The bottle” matches two objects, and “the table” matches two tables. But the relations carry information. The robot is in the kitchen, and only table_01 is in the kitchen, so if the task is happening in the kitchen, one table is the better reading. The world model does more than store objects; it is what makes it possible to narrow the candidates, and to notice when it cannot.
A complete example
The user says, “Bring me the bottle.” The world model holds red_bottle_01 on the kitchen counter and blue_bottle_01 on table_01.
Step 1 — Interpret
The robot understands the request: fetch a bottle and bring it to the user.
Step 2 — Ground
“The bottle” maps to two candidates,
red_bottle_01andblue_bottle_01. “Me” maps to the current user.Step 3 — Check
Nothing in the context favors either bottle, and fetching the wrong one wastes a trip.
Step 4 — Ask
“Which bottle do you mean, the red one or the blue one?”
Step 5 — Ground again, then act
The user answers, “The red one.” The reference now maps to
red_bottle_01, and the robot continues down the hierarchy from T7: goal, tasks, skills, motion.
Common misconceptions
Misconception 1 “If the robot understands the sentence, it knows what to do.”
Correction Understanding a sentence is not grounding it. The robot also has to know which real object, place or person the words refer to.
Misconception 2 “The robot should ask whenever it is uncertain.”
Correction Ask when ambiguity materially affects the task. Asking about every doubt makes the robot unusable.
Misconception 3 “It is better to guess than to interrupt the user.”
Correction When the interpretations differ and a wrong guess is costly, a short question is cheaper than a failed task.
Misconception 4 “Grounding is just matching keywords to object names.”
Correction Grounding uses relationships, context and the state of the world, not only names.
Misconception 5 “There is a confidence level above which the robot should always act.”
Correction The right threshold depends on the consequences, the safety risk and whether the result can be verified or undone.
Engineering takeaways
- Grounding connects language to entities, places, relationships and people in the world.
- Language understanding is not the same as grounding.
- The world model provides what language is grounded against.
- Ambiguity can be referential, spatial, attribute, temporal or contextual.
- Grounding can be uncertain, and belief tells you how uncertain.
- Ask when ambiguity materially affects the task.
- Do not ask when context and the world model already settle it.
- Asking and acting are both decisions, with costs on each side.
A good robot does not guess when it cannot tell what the user means, and does not interrogate the user when it can.
Knowledge check
Three conceptual questions. Write an answer, then reveal the explanation. Your answers stay in your browser.
Question 1
There are two red bottles in the room, and the user says, “Bring me the red bottle.” The robot understands the sentence perfectly. Has it grounded the command?
Explanation
No. Understanding the sentence is not the same as knowing which physical bottle is meant. The words map to two candidate objects, so the reference is still ambiguous. Grounding is complete only when the robot knows which of the two the user means, perhaps after using context or asking.
Question 2
The robot is 85% sure the user means red_bottle_01 and 15% sure it is red_bottle_02. Should it always act?
Explanation
Not necessarily. Whether 85% is enough depends on the consequences of being wrong, the safety risk, and whether the result can be verified or undone. If a mistake is cheap and easy to correct, acting may be right. If it is costly or hard to reverse, asking may be better. There is no universal threshold.
Question 3
The user says “Go to the kitchen,” and the robot’s world model contains exactly one kitchen. Should the robot ask which kitchen?
Explanation
No. With only one kitchen, the reference is already grounded, so asking would just be an annoying interruption. The principle is to ask when ambiguity materially affects the task, not whenever any uncertainty exists. A robot that asks about everything is as unhelpful as one that guesses about everything.
What’s next
The robot can now connect a command to the actual world, and decide whether to ask. Next we look at the reusable capabilities that carry the grounded task out.
Related content
Nothing below is required to finish this lesson.
- LessonT5: World Models and Symbolic State
The entities and relations language is grounded against.
- LessonT6: Belief Under Uncertainty
Why grounding comes with degrees of confidence.
- LessonT7: Command → Goal → Task → Skill → Motion
What happens once the command is grounded.
- TrackRobot Foundations
Sensing, moving, localizing and planning for beginners.
- TrackPerception Coming soon
How a robot comes to know which objects are in front of it.
- Lab 04Command → Structured Task In development
Turn a command into structured data.
- Lab 05Ground and Clarify In development
Resolve ambiguous references, and ask when you must.
Get notified when the next lesson and new labs are released.
One email per new lesson or lab.