Visual-Prompt Guided Wildlife Instance-Level Recognition

Visual-Prompt Guided Wildlife Instance-Level Recognition

视觉提示引导的野生动物实例级识别

Abstract: Fine-grained wildlife re-identification remains a challenging area in research. Current state-of-the-art approaches apply a detection and re-identification pipeline. We propose a one-stage end-to-end detection and re-identification model that performs identity searching within the latent space. We adopt DINOv2 for robust spatial geometry and MegaDescriptor for wildlife re-identification. We enhance latent queries with prompt re-identification features. A detection decoder queries the scene latent space to establish object boundaries around the target identity. Preliminary findings reflect a competitive mean average precision score of 30.584% compared to the state-of-the-art two stage approach of 44.89%. Qualitative results depict effective bounding and identification of animal identities.

摘要: 细粒度野生动物重识别(Re-ID)仍然是研究中的一个挑战性领域。目前最先进的方法通常采用“检测+重识别”的流水线模式。我们提出了一种单阶段端到端检测与重识别模型,该模型在潜在空间(latent space)内执行身份搜索。我们采用 DINOv2 来获取稳健的空间几何信息,并使用 MegaDescriptor 进行野生动物重识别。我们通过提示重识别特征来增强潜在查询。检测解码器通过查询场景潜在空间,在目标身份周围建立对象边界。初步研究结果显示,该模型的平均精度均值(mAP)为 30.584%,与目前最先进的两阶段方法(44.89%)相比具有竞争力。定性结果表明,该模型能够有效地对动物身份进行边界框标注与识别。


Authors: Mufhumudzi Muthivhi, Jiahao Huo, Terence van Zyl, Fredrik Gustafsson 作者: Mufhumudzi Muthivhi, Jiahao Huo, Terence van Zyl, Fredrik Gustafsson

Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) 学科分类: 计算机视觉与模式识别 (cs.CV);人工智能 (cs.AI);机器学习 (cs.LG)

Cite as: arXiv:2608.18246 [cs.CV] 引用格式: arXiv:2608.18246 [cs.CV]