Fine-tuned Florence-2 model for UI icon recognition in desktop applications.
Fuente del modelo
Extracto de la fuente
Fine-tuned Florence-2 model for UI icon recognition in desktop applications.
Fuentes
1 fuenteVerificado 27 ago
Artefactos del modelo
1 artefactoExtractos de fuentes
2 extractos--- license: mit base_model: microsoft/Florence-2-base datasets: - custom tags: - florence-2 - icon-caption - ui-detection - omniparser language: - en pipeline_tag: image-to-text --- # Florence-2 Icon Caption (Fine-tuned) Fine-tuned Florence-2 model for **UI icon recognition** in desktop applications. Based on [OmniParser-v2.0](https://huggingface.co/microsoft/OmniParser-v2.0) icon_caption weights, further fine-tuned on 12.8k curated icon samples from 163 desktop applications (including WeChat, Photoshop, VS Code, Figma, etc.) ## Key Features - **Functional icon recognition**: Trained only on learnable UI elements (buttons, tools, nav icons), excluding avatars/thumbnails/decorative elements - **Clean, standardized labels**: 2-5 word functional descriptions like `search button`, `settings gear`, `chats nav icon` - **163 app coverage**: Adobe suite, Microsoft Office, WeChat, Slack, Chrome, and 150+ more - **Chinese app support**: WeChat, DingTalk, Feishu, QQ, Bilibili, etc. ## Performance | Model | Val Loss | Exact Match | Output Quality | |-------|----------|-------------|----------------| | OmniParser (baseline) | - | 0% | Verbose, generic ("a loading or buffering indicator") | | **This model** | **1.329** | **18.8%** | Concise, functional ("settings gear", "search button") | ## Training Pipeline 1. **YOLO detection** → crop icons from 750+ screenshots across 163 apps 2. **Claude annotation** → send original screenshot + icon grid to Claude for context-aware labeling 3. **Smart filtering** → skip avatars, thumbnails, video frames, line numbers (unlearnable elements) 4. **Label standardization** → normalize synonyms (close window button → close button) 5. **Frequency filtering** → remove labels appearing < 3 times 6. **Full parameter training** → vision tower unfrozen, lr=3e-6, 15 epochs ## Usage ```python from transformers import AutoProcessor, AutoModelForCausalLM, AutoConfig from safetensors.torch import load_file from huggingface_hub import hf_hub_download from pathlib import Path from PIL import Image import torch # Load processor processor = AutoProcessor.from_pretrained("microsoft/Florence-2-base", trust_remote_code=True) # Load model structure from OmniParser config config_path = hf_hub_download("microsoft/OmniParser-v2.0", "icon_caption/config.json") config = AutoConfig.from_pretrained(str(Path(config_path).parent), trust_remote...
Source context: 23 downloads · 0 likes · Pipeline image-to-text · Repo josley/florence-2-icon-caption