AssameseOCR is a vision-language model for Optical Character Recognition (OCR) of printed Assamese text. Built on Microsoft's Florence-2-large foundation model with a custom character-level decoder, it achieves...
Model source
Source excerpt
AssameseOCR is a vision-language model for Optical Character Recognition (OCR) of printed Assamese text. Built on Microsoft's Florence-2-large foundation model with a custom character-level decoder, it achieves...
Sources
1 sourceVerified Sep 10
Model artifacts
1 artifactSource excerpts
2 excerpts--- language: - asm # Assamese ISO 639-1 code license: apache-2.0 base_model: microsoft/Florence-2-large-ft tags: - vision - ocr - assamese - northeast-india - indic-languages - character-recognition - florence-2 - vision-language datasets: - darknight054/indic-mozhi-ocr metrics: - accuracy - character_error_rate library_name: transformers pipeline_tag: image-to-text model-index: - name: AssameseOCR results: - task: type: image-to-text name: Optical Character Recognition dataset: name: Mozhi Indic OCR (Assamese) type: darknight054/indic-mozhi-ocr config: assamese split: test metrics: - type: accuracy value: 94.67 name: Character Accuracy verified: false - type: character_error_rate value: 5.33 name: Character Error Rate (CER) verified: false --- # AssameseOCR **AssameseOCR** is a vision-language model for Optical Character Recognition (OCR) of printed Assamese text. Built on Microsoft's Florence-2-large foundation model with a custom character-level decoder, it achieves 94.67% character accuracy on the Mozhi dataset. ## Model Details ### Model Description - **Developed by:** MWire Labs - **Model type:** Vision-Language OCR - **Language:** Assamese (অসমীয়া) - **License:** Apache 2.0 - **Base Model:** microsoft/Florence-2-large-ft - **Architecture:** Florence-2 Vision Encoder + Custom Transformer Decoder ### Model Architecture ``` Image (768×768) ↓ Florence-2 Vision Encoder (frozen, 360M params) ↓ Vision Projection (1024 → 512 dim) ↓ Transformer Decoder (4 layers, 8 heads) ↓ Character-level predictions (187 vocab) ``` **Key Components:** - **Vision Encoder:** Florence-2-large DaViT architecture (frozen) - **Decoder:** 4-layer Transformer with 512 hidden dimensions - **Tokenizer:** Character-level with 187 tokens (Assamese chars + English + digits + symbols) - **Total Parameters:** 378M (361M frozen, 17.5M trainable) ## Training Details ### Training Data - **Dataset:** [Mozhi Indic OCR Dataset](https://huggingface.co/datasets/darknight054/indic-mozhi-ocr) (Assamese subset) - **Training samples:** 79,697 word images - **Validation samples:** 9,945 word images - **Test samples:** 10,146 word images - **Source:** IIT Hyderabad CVIT ### Training Procedure **Hardware:** - GPU: NVIDIA A40 (48GB VRAM) - Training time: ~8 hours (3 epochs) **Hyperparameters:** - Epochs: 3 - Batch size: 16 - Learning rate: 3e-4 - Optimizer: AdamW...
Source context: 0 downloads · 1 likes · Pipeline image-to-text · Library transformers · Repo MWirelabs/assamese-ocr