Vision Transformers that predict a Pokemon's crowd "smash" fraction (0-100) from a single image -- i.e. how attractive the internet finds it. Trained on official artwork, in-game sprites, and Safebooru fan-art, with...
Source du modèle
Description de la source
Vision Transformers that predict a Pokemon's crowd "smash" fraction (0-100) from a single image -- i.e. how attractive the internet finds it. Trained on official artwork, in-game sprites, and Safebooru fan-art, with labels from aggregate votes on pokesmash.xyz.
Code, docs, and full reproduction guides: https://github.com/byrte1024/SmashOrTransformer
ViT-Small/16 @ 224 + a scalar regression head, fine-tuned with soft-label BCE. Each holds + + .
Sources
1 sourceVérifié 10 sept.
Artefacts du modèle
3 artefactsExtraits de sources
2 extraitstimm.ptmodel_stateconfigmetrics| File | Sources | Spearman (all_avg) | Notes |
|---|---|---|---|
vit_small_mixed_v1.pt | portrait + in-game + booru | 0.770 | recommended |
vit_small_portraits_v1.pt | portrait + in-game | 0.690 | sprite-only baseline |
vit_small_mixed_v2.pt | + heavy booru aug | 0.734 | deprecated (regression) |
Spearman is a fair cross-evaluation on a common held-out set of 102 Pokemon;
*.calibration.json are the isotonic calibration maps (mixed_v2 has none).
git clone https://github.com/byrte1024/SmashOrTransformer && cd SmashOrTransformer
uv sync
uv run python download_models.py # fetches vit_small_mixed_v1 into runs/
uv run python -m model.infer --checkpoint runs/vit_small_mixed_v1/checkpoints/best.pt img.png
other -- the training data includes third-party fan-art and official assets
that are not ours to relicense. Weights are provided for research use.
vit_small_portraits_v1.pt
pt · 82,7 MB · SHA-256 5eb115bc8317…c563 · Hugging Face
--- license: other library_name: timm tags: - pytorch - timm - vision-transformer - image-regression - pokemon pipeline_tag: image-classification --- # Smash or Transformer -- ViT-Small checkpoints Vision Transformers that predict a Pokemon's crowd "smash" fraction (0-100) from a single image -- i.e. how attractive the internet finds it. Trained on official artwork, in-game sprites, and Safebooru fan-art, with labels from aggregate votes on pokesmash.xyz. **Code, docs, and full reproduction guides:** https://github.com/byrte1024/SmashOrTransformer ## Checkpoints `timm` ViT-Small/16 @ 224 + a scalar regression head, fine-tuned with soft-label BCE. Each `.pt` holds `model_state` + `config` + `metrics`. | File | Sources | Spearman (all_avg) | Notes | |------|---------|:---:|-------| | `vit_small_mixed_v1.pt` | portrait + in-game + booru | **0.770** | recommended | | `vit_small_portraits_v1.pt` | portrait + in-game | 0.690 | sprite-only baseline | | `vit_small_mixed_v2.pt` | + heavy booru aug | 0.734 | deprecated (regression) | Spearman is a fair cross-evaluation on a common held-out set of 102 Pokemon; `*.calibration.json` are the isotonic calibration maps (mixed_v2 has none). ## Usage ```bash git clone https://github.com/byrte1024/SmashOrTransformer && cd SmashOrTransformer uv sync uv run python download_models.py # fetches vit_small_mixed_v1 into runs/ uv run python -m model.infer --checkpoint runs/vit_small_mixed_v1/checkpoints/best.pt img.png ``` Dataset: [supernovayuli/smash-or-transformer-data](https://huggingface.co/datasets/supernovayuli/smash-or-transformer-data) ## License `other` -- the training data includes third-party fan-art and official assets that are not ours to relicense. Weights are provided for research use.
Source context: 0 downloads · 0 likes · Pipeline image-classification · Library timm · Repo supernovayuli/smash-or-transformer