A live model of the room
Vision, depth, tactile and proprioceptive streams are fused into a continuously updated estimate of the environment. It holds up when lighting shifts, when objects move, and when the scene stops matching anything in the training set.