VGBench: Evaluating Large Language Models on Vector Graphics Understanding and Generation Paper • 2407.10972 • Published Jul 15 • 1
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models Paper • 2410.10818 • Published 27 days ago • 14
Making Large Multimodal Models Understand Arbitrary Visual Prompts Paper • 2312.00784 • Published Dec 1, 2023 • 2
Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want Paper • 2403.20271 • Published Mar 29 • 3
Split & Merge: Unlocking the Potential of Visual Adapters via Sparse Training Paper • 2312.02923 • Published Dec 5, 2023 • 1