Comparison of Detection Models for Unconstrained Objects for Humanoid Robots
Abstract
Localization for soccer-playing humanoid robots is essential for planning game strategies. Visual localization is commonly achieved by detecting landmarks of the soccer-field, nevertheless, this is difficult due to disturbances caused by motion blur during gait. Moreover, landmarks can have multiple valid bounding boxes that makes difficult the labeling for training. In this work, we compare three state-of-the-art methods for detecting landmarks: YOLO, RT-DETR and Deformable-DETR. We used the TORSO-21 dataset, augmented with data taken during Robocup Humanoid League 2025. We compare the performance using precision, recall and the angular localization error. Our results showed that Deformable-DETR overcome RT-DETR and YOLO in precision but not in recall for all classes. Nevertheless, Deformable-DETR has the lowest angular localization error for all classes but also the highest computing time consumption.
Keywords
Detection models, unconstrained objects, humanoid robots.