ProtGPT3-MSA is a multiple-sequence, homolog-conditioned autoregressive protein language model. It is part of the ProtGPT3 family, an open-source suite of promptable and aligned protein language models for protein...
Source du modèle
Extrait de la source
ProtGPT3-MSA is a multiple-sequence, homolog-conditioned autoregressive protein language model. It is part of the ProtGPT3 family, an open-source suite of promptable and aligned protein language models for protein...
Sources
1 sourceVérifié 26 sept.
Artefacts du modèle
1 artefactExtraits de sources
3 extraits--- library_name: transformers tags: - biology - protein-language-model - protein-generation - msa - multiple-sequence-alignment - few-shot-prompting - homolog-conditioned-generation - causal-lm - mixture-of-experts - transformers --- # Model Card for ProtGPT3-MSA ## Model Description ProtGPT3-MSA is a multiple-sequence, homolog-conditioned autoregressive protein language model. It is part of the [ProtGPT3 family](https://huggingface.co/collections/AI4PD/protgpt3-family), an open-source suite of promptable and aligned protein language models for protein sequence generation. Unlike the single-sequence ProtGPT3 models, ProtGPT3-MSA can be prompted with sets of homologous protein sequences, enabling few-shot, family-conditioned protein generation without task-specific fine-tuning. At inference, users can provide homologous protein sequences as context and generate additional family-consistent sequences. ProtGPT3-MSA was trained to autoregressively predict sets of 16 concatenated protein sequences, separated by a special token `<s>` (i.e., marking the protein boundaries). Therefore, at inference, the model should be prompted with at most 15 concatenated protein sequences. - For more details on how to use ProtGPT3-MSA check out our [colab](https://colab.research.google.com/drive/1HZFLUkRIhjUJdbQyvJC8ftio_ZHNL7kI?usp=sharing#scrollTo=zwWWcxwkPm6c). - For a quick usage of the model for generating new sequences by prompting it with a fasta file of homologous sequences check out [ProtGPT3-MSA API](https://huggingface.co/spaces/AI4PD/ProtGPT3-MSA). ### Model Modalities 1. **Aligned vs unaligned mode**:ProtGPT3-MSA has been trained to process concatenated sets of homologs in both "aligned" (i.e., the homologs are passed aligned with gap tokens) and "unaligned" mode via special `<gap>` and `<no_gap>` tokens, which should be placed at the start of the concatenated protein sequences to select the modality. 2. **N-to-C vs C-to-N**:ProtGPT3-MSA has been trained to process concatenated homologs in both N-to-C and C-to-N directions, via two special "directional" tokens, "1" for N-to-C and "2" for C-to-N which which should be placed at the start of the concatenated protein sequences (i.e., before the gap token) to select the direction. We provide some examples below. ## Uses ## How to Get Started with the Model Install dependencies: ```bash pip install transfo...
Source context: 4110 downloads · 3 likes · Pipeline text-generation · Library transformers · Repo AI4PD/ProtGPT3-MSA
Source context: 684 downloads · 2 likes · Pipeline text-generation · Library transformers · Repo AI4PD/ProtGPT3-MSA