Full description
For embodied agents, such as robots, tracking objects in their surroundings through visual observation is essential — a task, referred to as Visual Object Tracking (VOT). For instance, during a rearrangement task, a robot may need to track objects, as part of the scene change understanding process, to accurately restore them to their original states. Classic Multiple Object Tracking (MOT) datasets typically focus on tracking moving, single-class object instances in a video from a fixed viewpoint, limiting their applicability to embodied AI tasks. In embodied AI tasks, objects belong to multiple classes, are often static, and are observed from continuously changing viewpoints, as the camera is mounted on a moving robotic agent. This ego-centric perspective in M3T dataset introduces unique challenges, such as frequent attention shifts; large camera motions, causing frequent object disappearances; and object manipulations, leading to occlusions, rapid changes in object scale, pose, or appearance. The proposed M3T dataset is specifically designed for the 2D scene understanding stage of embodied AI tasks. M3T expands the scope of traditional MOT datasets to accommodate the complexities of ego-centric visual exploration and static, multi-class object tracking in dynamic environments.
Notes
M3T dataset includes the largest amount of scenes (1,048 tracking sequences), generated using the Ai2Thor simulator. The M3T dataset offers the highest class diversity, featuring 42 indoor object types. Unlike other datasets, which primarily track vehicles or pedestrians in mostly static scenes, 41 of M3T's classes are interactable objects. These objects can change location or state, enabling tracking in dynamic, constantly evolving indoor environments.
User Contributed Tags
Login to tag this record with meaningful keywords to make it easier to discover
- DOI : 10.25958/YQ3N-FY41
