Falcon OCR is a 300M parameter early-fusion vision-language model for document OCR. Given an image, it can produce plain text, LaTeX for formulas, or HTML for tables, depending on the requested output format.
Modellquelle
Quellenauszug
Falcon OCR is a 300M parameter early-fusion vision-language model for document OCR. Given an image, it can produce plain text, LaTeX for formulas, or HTML for tables, depending on the requested output format.
Quellen
1 QuelleVerifiziert 27. Aug.
Modellartefakte
1 ArtefaktQuellenauszüge
2 Auszüge--- pipeline_tag: image-to-text library_name: transformers tags: - falcon - ocr - vision-language - document-understanding license: apache-2.0 --- <div style="width: 480px; text-align: left;"> <img src="https://cdn-uploads.huggingface.co/production/uploads/663c9939c1b4f7297c4ae6f6/YIuxzgDiV5T2ZuSB4bam9.png" alt="Falcon OCR Logo" style="max-width: 100%; height: auto;"> </div> # Falcon OCR Falcon OCR is a 300M parameter early-fusion vision-language model for document OCR. Given an image, it can produce plain text, LaTeX for formulas, or HTML for tables, depending on the requested output format. Most OCR VLM systems are built as a pipeline with a vision encoder feeding a separate text decoder, plus additional task-specific glue. Falcon OCR takes a different approach: a single Transformer processes image patches and text tokens in a shared parameter space from the first layer, using a hybrid attention mask where image tokens attend bidirectionally and text tokens decode causally conditioned on the image. We built it this way for two practical reasons. First, it keeps the interface simple: one backbone, one decoding path, and task switching through prompts rather than a growing set of modules. Second, a 0.3B model has a lower latency and cost footprint than 0.9B-class OCR VLMs, and in our vLLM-based serving setup this translates into higher throughput, often 2–3× faster depending on sequence lengths and batch configuration. To our knowledge, this is one of the first attempts to apply this early-fusion single-stack recipe directly to competitive document OCR at this scale. ### Links - Code and inference engine: [https://github.com/tiiuae/Falcon-Perception](https://github.com/tiiuae/Falcon-Perception) - Tech report: [https://arxiv.org/pdf/2603.27365](https://arxiv.org/pdf/2603.27365) - Perception model: `tiiuae/falcon-perception` - vLLM/Docker: [https://ghcr.io/tiiuae/falcon-ocr:latest](https://ghcr.io/tiiuae/falcon-ocr:latest) ## Quickstart ### Installation ```bash pip install "torch>=2.5" transformers pillow einops ``` Falcon OCR requires PyTorch 2.5 or newer for FlexAttention. The first call may be slower as `torch.compile` builds optimized kernels. ### Single-Image OCR ```python import torch from PIL import Image from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained( "tiiuae/Falcon-OCR", trust_remote_code=...
Source context: 3 downloads · 0 likes · Pipeline image-to-text · Library transformers · Repo beaupi/Falcon-OCR-oQ8