Asset Details
MbrlCatalogueTitleDetail
Do you wish to reserve the book?
Animal-CLIP: A Dual-Prompt Enhanced Vision-Language Model for Animal Action Recognition
by
Li, Yongxiang
, Ma, Zhanyu
, Liang, Kongming
, He, Zhongjiang
, Jing, Yinuo
, Zhang, Ruxu
, Sun, Hao
in
Accuracy
/ Activity recognition
/ Animals
/ Artificial Intelligence
/ Automation
/ Biodiversity
/ Computer Imaging
/ Computer Science
/ Computer vision
/ Datasets
/ Ecosystems
/ Enhanced vision
/ Image Processing and Computer Vision
/ Knowledge
/ Language
/ Large language models
/ Pattern Recognition
/ Pattern Recognition and Graphics
/ Prompt engineering
/ Semantics
/ Video
/ Vision
2025
Hey, we have placed the reservation for you!
By the way, why not check out events that you can attend while you pick your title.
You are currently in the queue to collect this book. You will be notified once it is your turn to collect the book.
Oops! Something went wrong.
Looks like we were not able to place the reservation. Kindly try again later.
Are you sure you want to remove the book from the shelf?
Animal-CLIP: A Dual-Prompt Enhanced Vision-Language Model for Animal Action Recognition
by
Li, Yongxiang
, Ma, Zhanyu
, Liang, Kongming
, He, Zhongjiang
, Jing, Yinuo
, Zhang, Ruxu
, Sun, Hao
in
Accuracy
/ Activity recognition
/ Animals
/ Artificial Intelligence
/ Automation
/ Biodiversity
/ Computer Imaging
/ Computer Science
/ Computer vision
/ Datasets
/ Ecosystems
/ Enhanced vision
/ Image Processing and Computer Vision
/ Knowledge
/ Language
/ Large language models
/ Pattern Recognition
/ Pattern Recognition and Graphics
/ Prompt engineering
/ Semantics
/ Video
/ Vision
2025
Oops! Something went wrong.
While trying to remove the title from your shelf something went wrong :( Kindly try again later!
Do you wish to request the book?
Animal-CLIP: A Dual-Prompt Enhanced Vision-Language Model for Animal Action Recognition
by
Li, Yongxiang
, Ma, Zhanyu
, Liang, Kongming
, He, Zhongjiang
, Jing, Yinuo
, Zhang, Ruxu
, Sun, Hao
in
Accuracy
/ Activity recognition
/ Animals
/ Artificial Intelligence
/ Automation
/ Biodiversity
/ Computer Imaging
/ Computer Science
/ Computer vision
/ Datasets
/ Ecosystems
/ Enhanced vision
/ Image Processing and Computer Vision
/ Knowledge
/ Language
/ Large language models
/ Pattern Recognition
/ Pattern Recognition and Graphics
/ Prompt engineering
/ Semantics
/ Video
/ Vision
2025
Please be aware that the book you have requested cannot be checked out. If you would like to checkout this book, you can reserve another copy
We have requested the book for you!
Your request is successful and it will be processed during the Library working hours. Please check the status of your request in My Requests.
Oops! Something went wrong.
Looks like we were not able to place your request. Kindly try again later.
Animal-CLIP: A Dual-Prompt Enhanced Vision-Language Model for Animal Action Recognition
Journal Article
Animal-CLIP: A Dual-Prompt Enhanced Vision-Language Model for Animal Action Recognition
2025
Request Book From Autostore
and Choose the Collection Method
Overview
Animal action recognition has a wide range of applications. With the rise of visual-language pretraining models (VLMs), new possibilities have emerged for action recognition. However, while current VLMs perform well on human-centric videos, they still struggle with animal videos. This is primarily due to the lack of domain-specific knowledge during model training and more pronounced intra-class variations compared to humans. To address these issues, we introduce Animal-CLIP, a specialized and efficient animal action recognition framework built upon existing VLMs. To address the lack of domain-specific knowledge in animal actions, we leverage the extensive expertise of large language models (LLMs) to automatically generate external prompts, thereby expanding the semantic scope of labels and enhancing the model’s generalization capability. To effectively integrate external knowledge into the model, we propose a knowledge-enhanced internal prompt fine-tuning approach. We design a text feature refinement module to reduce potential recognition inconsistencies. Furthermore, to address the high intra-class variation in animal actions, a novel category-specific prompting method is introduced to generate adaptive prompts to optimize the alignment between text and video features, facilitating more precise partitioning of the action space. Experimental results demonstrate that our method outperforms six previous action recognition methods across three large-scale multi-species, multi-action datasets and exhibits strong generalization capability on unseen animals.
Publisher
Springer US,Springer Nature B.V
Subject
This website uses cookies to ensure you get the best experience on our website.