Real-Time Recognition of Unseen Destinations for Drone Delivery: A Multi-Modal and Generative Approach

Victoria Eugenia Vázquez Meza, Delia Irazú Hernández-Farías, Jose Martinez-Carranza

Abstract


We investigate the application of generative models to assist artificial agents, such as delivery drones or service robots, in visualising unfamiliar destinations solely based on textual descriptions. To this end, we propose a methodology aimed at comparing generated images from textual descriptions against images captured with a camera onboard a drone.
Our methodology involves using clustering to organize thousands of generated images, thus enabling a rapid comparison. For this purpose, we employ visual embeddings such as CLIP to numerically represent the images and make them comparable at a semantic level. However, when two candidate images are similar in semantic terms, it is useful to assess which one more closely resembles the target image visually.
Therefore, we also explore the use of visual features obtained with ResNet. Then, we combine them through a measure that scores the similarity between CLIP embeddings and ResNet features of the images to be compared. We achieve an accuracy of 0.80 with an average operation frequency of 18 Hz for online image processing, demonstrating the feasibility of our approach for recognizing previously unseen places for real-time drone applications.

Keywords


Place recognition, drone delivery, generative model, stable diffusion, CLIP.

Full Text: PDF