As a baseline, the BToM model is compared against a cue-based "MotionHeuristic" model which only takes into account which objects the agent is moving towards / away from.
Why is the MotionHeuristic model unable to produce human-like inferences about the agent's desires, using the scenario in Figure 1 as an example?
Why is MotionHeuristic better at producing human-like inferences about the agent's beliefs (as Figure 5d shows)?
How might the scenarios be modified to "break" the MotionHeuristic, so that it no longer produces human-like inferences about the agent's beliefs?