This model is the MatCha model, fine-tuned on Chart2text-statista dataset. This fine-tuned checkpoint might be better suited for chart summarization task.
Source du modèle
Extrait de la source
This model is the MatCha model, fine-tuned on Chart2text-statista dataset. This fine-tuned checkpoint might be better suited for chart summarization task.
Sources
1 sourceVérifié 11 sept.
Artefacts du modèle
1 artefactExtraits de sources
2 extraits--- language: - en - fr - ro - de - multilingual inference: false pipeline_tag: visual-question-answering license: apache-2.0 tags: - matcha --- # Model card for MatCha - fine-tuned on Chart2text-statista <img src="https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/model_doc/matcha_architecture.jpg" alt="drawing" width="600"/> This model is the MatCha model, fine-tuned on Chart2text-statista dataset. This fine-tuned checkpoint might be better suited for chart summarization task. # Table of Contents 0. [TL;DR](#TL;DR) 1. [Using the model](#using-the-model) 2. [Contribution](#contribution) 3. [Citation](#citation) # TL;DR The abstract of the paper states that: > Visual language data such as plots, charts, and infographics are ubiquitous in the human world. However, state-of-the-art visionlanguage models do not perform well on these data. We propose MATCHA (Math reasoning and Chart derendering pretraining) to enhance visual language models’ capabilities jointly modeling charts/plots and language data. Specifically we propose several pretraining tasks that cover plot deconstruction and numerical reasoning which are the key capabilities in visual language modeling. We perform the MATCHA pretraining starting from Pix2Struct, a recently proposed imageto-text visual language model. On standard benchmarks such as PlotQA and ChartQA, MATCHA model outperforms state-of-the-art methods by as much as nearly 20%. We also examine how well MATCHA pretraining transfers to domains such as screenshot, textbook diagrams, and document figures and observe overall improvement, verifying the usefulness of MATCHA pretraining on broader visual language tasks. # Using the model You should ask specific questions to the model in order to get consistent generations. Here we are asking the model whether the sum of values that are in a chart are greater than the largest value. ```python from transformers import Pix2StructProcessor, Pix2StructForConditionalGeneration import requests from PIL import Image processor = Pix2StructProcessor.from_pretrained('google/matcha-chart2text-statista') model = Pix2StructForConditionalGeneration.from_pretrained('google/matcha-chart2text-statista') url = "https://raw.githubusercontent.com/vis-nlp/ChartQA/main/ChartQA%20Dataset/val/png/20294671002019.png" image = Image.open(requests.get(url, stream=Tr...
Source context: 140 downloads · 10 likes · Pipeline visual-question-answering · Library transformers · Repo google/matcha-chart2text-statista