Unified Text + Visual Document Reranker by LightOn
Fuente del modelo
Extracto de la fuente
Unified Text + Visual Document Reranker by LightOn
Fuentes
1 fuenteVerificado 28 ago
Artefactos del modelo
1 artefactoExtractos de fuentes
2 extractos--- license: apache-2.0 language: - en - fr base_model: - Qwen/Qwen3.5-0.8B pipeline_tag: image-text-to-text library_name: transformers tags: - reranker - cross-encoder - multimodal - text-ranking - document-reranking - vidore - beir datasets: - vidore/colpali_train_set - lightonai/embeddings-fine-tuning --- <div align="center"> <img src="rerank.png" alt="Reranking: the second stage that puts the best text and image candidates on top" width="560"> [](https://lighton.ai) [](https://www.linkedin.com/company/lighton/) [](https://x.com/LightOnIO) 📝 [Blog post](https://huggingface.co/blog/lightonai/lighton-rerank) </div> <h1 align="center">LightOn-rerank-LW-0.8B</h1> <h3 align="center">Unified Text + Visual Document Reranker by LightOn</h3> <p align="center"> <a href="https://huggingface.co/lightonai/LightOn-rerank-PW-0.8B">PW-0.8B</a> | <a href="https://huggingface.co/lightonai/LightOn-rerank-LW-0.8B">LW-0.8B</a> | <a href="https://huggingface.co/lightonai/LightOn-rerank-PW-2B">PW-2B</a> | <a href="https://huggingface.co/lightonai/LightOn-rerank-LW-2B">LW-2B</a> | <a href="https://huggingface.co/lightonai/LightOn-rerank-PW-4B">PW-4B</a> | <a href="https://huggingface.co/lightonai/LightOn-rerank-LW-4B">LW-4B</a> </p> --- ## About the LightOn-rerank family Production retrieval pipelines usually need two rerankers: one for text passages and one for visual documents (PDF pages, slides, scans). **LightOn-rerank** models are *unified* cross-encoder rerankers: a single model scores both text passages and document page images against a query, on top of any first-stage retriever (BM25, dense embeddings, or ColPali-family late-interaction models). The models are built on Qwen3.5 backbone (hybrid linear + full attention) and jointly fine-tuned on text and visual reranking data with mixed-modality batches (LoRA, merged into the released weights). Training data is English-only; French performance transfers zero-shot from the multilingual backbone. The family comes in two scoring flavours × three sizes (0.8B / 2B / 4B): - **PW (pointwise)**: each candidate is scored independently. The model judges whether the document answers the query, and the...
Source context: 205 downloads · 3 likes · Pipeline image-text-to-text · Library transformers · Repo lightonai/LightOn-rerank-LW-0.8B