Quick answer

Paper2025-07-07•Source ↗•10 attns7,668 checkouts

Claim

jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval

Authors

Discuss with Grok

Michael Günther·

Saba Sturua·

Mohammad Kalim Akram·

Isabelle Mohr·

Andrei Ungureanu·

Bo Wang·

Sedigheh Eslami·

Scott Martens·

Maximilian Werk·

Nan Wang·

Han Xiao

ABSTRACT

We introduce jina-embeddings-v4, a 3.8 billion parameter multimodal embedding model that unifies text and image representations through a novel architecture supporting both single-vector and multi-vector embeddings in the late interaction style. The model incorporates task-specific Low-Rank Adaptation (LoRA) adapters to optimize performance across diverse retrieval scenarios, including query-document retrieval, semantic text similarity, and code search. Comprehensive evaluations demonstrate that jina-embeddings-v4 achieves state-of-the-art performance on both single-modal and cross-modal retrieval tasks, with particular strength in processing visually rich content such as tables, charts, diagrams, and mixed-media formats. To facilitate evaluation of this capability, we also introduce Jina-VDR, a novel benchmark specifically designed for visually rich image retrieval.

#deep-learning/month/202507 #machine-learning 📋 Awesome List: multimodal #deep-learning/year/2025 #multimodal #deep-learning #machine-learning/year/2025 #machine-learning/month/202507

Review Snapshot

Explore ratings

0.0

★★★★★

0 ratings

5 star

4 star

3 star

2 star

1 star

Recommendation

recommend this content.

Review this content

Share your opinion to help other learners triage faster.

Write a review

Invite a reviewer

Invite someone by email to share an invited review for jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval.

Author Inquiries

Public questions about this content. Attendemia will route your question to the author. Vote on the most important ones. No guarantee of response.

Post an inquiry

Sort by: Most helpful