About This Project
We introduce Vidi, a family of Large Multimodal Models (LMMs) for a wide range of video understanding and editing (VUE) scenarios. The models focus on temporal retrieval, spatio-temporal grounding, and maintaining robust open-ended video QA performance.
Tags
Reviews & Ratings
Share your experience
User Reviews (0)
No reviews yet. Be the first to share your experience!