Follow

CoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing. (arXiv:2401.12264v1 [eess.AS]) arxiv.org/abs/2401.12264

· · feed2toot · 0 · 0 · 0
Sign in to participate in the conversation
CleverLibre Social

CleverLibre Social is an inclusive social instance for open discussion, learning, and community.
All cultures welcome.
Hate speech and harassment strictly forbidden.