This repository contains a deep learning model designed for Khmer and English Optical Character Recognition (OCR). It utilizes a ResNet backbone for spatial feature extraction, a bidirectional LSTM for sequence...
Fuente del modelo
Extracto de la fuente
This repository contains a deep learning model designed for Khmer and English Optical Character Recognition (OCR). It utilizes a ResNet backbone for spatial feature extraction, a bidirectional LSTM for sequence...
Fuentes
1 fuenteVerificado 6 ago
Artefactos del modelo
2 artefactosExtractos de fuentes
3 extractos--- language: - km - en base_model: tharas/resnet_bilstm_ctc_100k pipeline_tag: image-to-text tags: - ocr - khmer - crnn - resnet - ctc --- # Khmer OCR — ResNet + BiLSTM + CTC This repository contains a deep learning model designed for Khmer and English Optical Character Recognition (OCR). It utilizes a ResNet backbone for spatial feature extraction, a bidirectional LSTM for sequence modeling, and Connectionist Temporal Classification (CTC) loss for alignment-free text recognition. ## Model Details ### Model Description - **Model type:** `khm_ocr_general_document` - **Language(s):** Khmer (`km`), English (`en`). *Note: The model exhibits higher accuracy and optimization patterns for Khmer character compositions.* - **Training State:** Trained from scratch on a comprehensive dataset of printed text lines. - **Architecture:** - **CNN:** ResNet blocks processing grayscale document line images downscaled to a fixed height. - **RNN:** 2-layer Bidirectional LSTM tracking character transitions. - **Classifier:** Linear layer mapping features to vocabulary indices decoded via greedy CTC. - **Charactor Error Rate:** 0.005589 - **Word Error Rate:** 0.045868 - **Training**: - **batch_size:** 16 - **device:** "cuda" - **epochs:** 30 - **learning_rate:** 0.001 - **num_workers:** 4 - **optimizer:** "adam" - **scheduler_factor:** 0.5 - **scheduler_patience:** 5 ### Out-of-Scope Use While this model handles varied font distributions well, it is strictly an line-level text recognizer. * **Optimal on:** Printed text documents featuring standard degradation, light blur, or low levels of scanner noise. * **Fails on:** Heavily underlined text paths, highly blurred captures, handwritten manuscripts, or unwarped, rotated document images. Ensure input crops are straight and baseline-aligned before inference. ### Recommendations This architecture is under active development. While it successfully handles clean, isolated line crops, performance drops on extreme layouts. Preprocessing elements (like layout analysis and text-line deskewing) should be executed prior to feeding imagery to this network. --- ## How to Get Started with the Model ### File Layout Requirements Before running predictions, ensure your repository or local folder contains your model files and characters map structured as follows: ```text . ├── predict.py └── khmer_ocr_model_CRNN/ ├── bes...
Source context: 0 downloads · 0 likes · Pipeline image-to-text · Repo Thareah/khmer_ocr_model_CRNN
Source context: 0 downloads · 0 likes · Pipeline image-to-text · Repo Thareah/khmer_ocr_model_CRNN