LLaVAction: evaluating and training multi-modal large language models for action recognition
Fuente del modelo
Extracto de la fuente
LLaVAction: evaluating and training multi-modal large language models for action recognition
Fuentes
1 fuenteVerificado 21 ago
Artefactos del modelo
1 artefactoExtractos de fuentes
2 extractos--- license: cc-by-nc-sa-4.0 language: - en base_model: - lmms-lab/llava-onevision-qwen2-0.5b-ov pipeline_tag: video-text-to-text tags: - Action - Video - MQA - multimodal - MLLMs - LLaVAction metrics: - accuracy library_name: transformers --- # LLaVAction-0.5B <div align="center"> <h2>LLaVAction: evaluating and training multi-modal large language models for action recognition </h2> [Shaokai Ye](https://yeshaokai.github.io/)<sup>1**</sup> [Haozhe Qi](https://people.epfl.ch/haozhe.qi)<sup>1**</sup> [Alexander Mathis](https://mathislab.org/)<sup>1</sup><sup>†</sup> [Mackenzie Weygandt Mathis](https://www.mackenziemathislab.org/mackenziemathis)<sup>1</sup><sup>†</sup><sup>‡</sup> <sup>1</sup> EPFL <sup>**</sup> First authors <sup>†</sup> Senior Authors <sup>‡</sup> Corresponding Author \[[arXiv Paper](arxiv.org/abs/2503.18712)\] \[[Project Page](https://mmathislab.github.io/llavaction/)\] \[[Github Repo](https://github.com/AdaptiveMotorControlLab/LLaVAction)\] </div> ## Model Summary The LLaVAction-0.5B model is trained on EPIC-KITCHENS-100-MQA, based on Qwen2 language model with a context window of 32K tokens. - **Project Page**: [https://mmathislab.github.io/llavaction/](https://mmathislab.github.io/llavaction/) - **Paper**: For more details, please check our [paper](https://arxiv.org/abs/tbd) - **Repository**: [https://github.com/AdaptiveMotorControlLab/LLaVAction](https://github.com/AdaptiveMotorControlLab/LLaVAction) - **Point of Contact**: [Mackenzie Mathis](https://people.epfl.ch/mackenzie.mathis) - **Languages**: English - ## Useage ### Intended use The model was trained on EPIC-KITCHENS-100-MQA. It's intended to be used on videos that are similar to EPIC-KITCHENS-100. ### Generation We provide the simple generation process for using our model. For more details, you could refer to our [Github](https://github.com/AdaptiveMotorControlLab/LLaVAction). ```python !pip install llavaction from llavaction.model.builder import load_pretrained_model from llavaction.mm_utils import get_model_name_from_path, process_images, tokenizer_image_token from llavaction.constants import IMAGE_TOKEN_INDEX, DEFAULT_IMAGE_TOKEN, DEFAULT_IM_START_TOKEN, DEFAULT_IM_END_TOKEN, IGNORE_INDEX from llavaction.conversation import conv_templates, SeparatorStyle from PIL import Image import requests import copy import torch i...
Source context: 39 downloads · 3 likes · Pipeline video-text-to-text · Library transformers · Repo MLAdaptiveIntelligence/LLaVAction-0.5B