> MVP model. Production upgrade: swap to ppocrv5 variant (same interface, > better accuracy). See config.json → architecturevariant for programmatic detection.
Modellquelle
Quellenauszug
MVP model. Production upgrade: swap to ppocrv5 variant (same interface, > better accuracy). See config.json → architecturevariant for programmatic detection.
Quellen
1 QuelleVerifiziert 3. Sept.
Modellartefakte
1 ArtefaktQuellenauszüge
2 Auszüge--- library_name: pytorch license: mit language: - th - en tags: - ocr - text-recognition - thai-id-card - crnn - ctc - on-device - mobile - english-ocr - name-recognition pipeline_tag: image-to-text --- # Thai ID Nano OCR — English OCR Reader (SimpleCRNN (MVP)) > **MVP model.** Production upgrade: swap to `ppocrv5` variant (same interface, > better accuracy). See `config.json` → `architecture_variant` for programmatic detection. CTC-based text recognition model for Thai National ID card **english** fields, designed for on-device inference at 30fps on mobile. | Metric | Value | |--------|-------| | Architecture | SimpleCRNN (MVP) | | Variant | `crnn` | | ExactMatch | 97.6% | | CharAccuracy | 98.2% | | Parameters | 3,048,762 | | Vocab size | 58 | | Best epoch | 75 | ## Quick Start ```python from huggingface_hub import hf_hub_download model_path = hf_hub_download("chayuto/thai-id-ocr-crnn-english-reader", "model.pt") vocab_path = hf_hub_download("chayuto/thai-id-ocr-crnn-english-reader", "vocab.txt") config = hf_hub_download("chayuto/thai-id-ocr-crnn-english-reader", "config.json") ``` ## Architecture **SimpleCRNN** — CNN (4-layer) + BiLSTM (2-layer) + CTC decoder. ``` Input: [B, 3, 48, 320] (RGB, normalized to [-1, 1]) → CNN: 32→64→128→256 channels, BatchNorm+ReLU, MaxPool(2,2)×3 → AdaptiveAvgPool2d((1, None)) → T=40 time steps → BiLSTM: hidden=256, layers=2, dropout=0.1 → Linear(512 → 58) → CTC decode (blank=0, collapse repeats) Output: Unicode string ``` ## Field Details - **Zone:** `text_eng_zone` (romanized Thai names) - **Charset:** A-Z, a-z, `.,-'` and space (57 chars + CTC blank) - **Case-sensitive** — Title Case names (Mr. Somchai Sombun) ## Input Preprocessing ```python import cv2 import numpy as np def preprocess(img_path, height=48, max_width=320): img = cv2.imread(img_path) h, w = img.shape[:2] ratio = height / h new_w = min(int(w * ratio), max_width) img = cv2.resize(img, (new_w, height)) # Pad to max_width with white if new_w < max_width: pad = np.full((height, max_width - new_w, 3), 255, dtype=np.uint8) img = np.concatenate([img, pad], axis=1) # Normalize to [-1, 1] img = img.astype(np.float32) / 255.0 img = (img - 0.5) / 0.5 return np.transpose(img, (2, 0, 1)) # CHW ``` ## CTC Decoding ```python def ctc_decode(indices, vocab_chars, blank_idx=0): chars, prev = [], -1 for idx in indices: if idx...
Source context: 10 downloads · 0 likes · Pipeline image-to-text · Library pytorch · Repo chayuto/thai-id-ocr-crnn-english-reader