This model is a BLIP image captioning model fine-tuned using LoRA (Low Rank Adaptation) on the RICO Screen2Words dataset.
Model source
Source description
This model is a BLIP image captioning model fine-tuned using LoRA (Low Rank Adaptation) on the RICO Screen2Words dataset.
The goal of the model is to generate textual descriptions of mobile user interface screenshots. The model takes an image of a UI screen as input and produces a caption describing the interface.
Sources
1 sourceVerified Sep 3
Model artifacts
1 artifactSource excerpts
2 excerptsThis model can be used to generate captions describing mobile UI screenshots. It can help automatically summarize UI layouts and components present in mobile application interfaces.
Example use cases include:
The model is not designed for:
Example inference code:
from transformers import pipeline
pipe = pipeline(
"image-to-text",
model="suhas2468/rico-blip-lora-model"
)
result = pipe("example_image.png")
print(result)
The model was fine-tuned using the RICO Screen2Words dataset.
This dataset contains mobile UI screenshots paired with captions describing the interface.
Training was performed using Hugging Face Transformers, PEFT, and PyTorch.
The model was fine-tuned using LoRA (Low Rank Adaptation).
LoRA enables efficient fine-tuning by adding small trainable adapters to the base model while keeping most of the original model parameters frozen. This reduces GPU memory usage and speeds up training.
Training was performed on:
GitHub repository containing the training notebook and code:
--- library_name: transformers tags: - transformers - image-to-text --- # RICO BLIP LoRA Model ## Model Description This model is a BLIP image captioning model fine-tuned using **LoRA (Low Rank Adaptation)** on the **RICO Screen2Words dataset**. The goal of the model is to generate textual descriptions of mobile user interface screenshots. The model takes an image of a UI screen as input and produces a caption describing the interface. --- ## Model Details * **Developed by:** suhas2468 * **Model type:** Vision-Language Model (Image Captioning) * **Language(s):** English * **License:** Apache 2.0 * **Finetuned from model:** Salesforce/blip-image-captioning-base --- ## Model Sources * **Repository:** https://github.com/suhassh2/orange-multimodal-slm * **Base Model:** https://huggingface.co/Salesforce/blip-image-captioning-base * **Dataset:** https://huggingface.co/datasets/rootsautomation/RICO-Screen2Words --- ## Intended Use ### Direct Use This model can be used to generate captions describing mobile UI screenshots. It can help automatically summarize UI layouts and components present in mobile application interfaces. Example use cases include: * UI understanding * UI documentation * Automated interface description ### Out-of-Scope Use The model is not designed for: * general image captioning outside UI screenshots * safety-critical applications * decision-making systems --- ## How to Use the Model Example inference code: ```python from transformers import pipeline pipe = pipeline( "image-to-text", model="suhas2468/rico-blip-lora-model" ) result = pipe("example_image.png") print(result) ``` --- ## Training Details ### Training Data The model was fine-tuned using the **RICO Screen2Words dataset**. This dataset contains mobile UI screenshots paired with captions describing the interface. Dataset link: https://huggingface.co/datasets/rootsautomation/RICO-Screen2Words --- ### Training Setup * **GPU:** NVIDIA Tesla T4 * **Epochs:** 1 * **Batch size:** 4 * **Learning rate:** 2e-5 * **Training samples:** 500 Training was performed using Hugging Face **Transformers**, **PEFT**, and **PyTorch**. --- ## Training Method The model was fine-tuned using **LoRA (Low Rank Adaptation)**. LoRA enables efficient fine-tuning by adding small trainable adapters to the base model while keeping most of the original model parameters frozen....
Source context: 0 downloads · 0 likes · Pipeline image-to-text · Library transformers · Repo suhas2468/rico-blip-lora-model