libraryname: pytorch license: other tags: string-atlas tokenizer-interface fineweb-edu realistic-bpe5
Model source
Source description
Final 1B wiki-qa RealisticBPE5 specialist checkpoint trained on the wiki-qa 20B packed FineWeb-Edu stream.
HF repo: cminst/realistic-bpe5-wiki-qa-specialist-1b-20b.
Source run directory name: R5FW20B-spec-wiki_qa_bpe32-1B-20B-tillicum.
Data bundle: cminst/realistic-bpe5-wiki-qa-20b.
Sources
1 sourceVerified Jul 17
Model artifacts
1 artifactSource excerpts
2 excerptsmodel.pt: final model weightsconfig.json: training/model configuration saved by train_v3.pyresults.json: final evaluation metrics and logsstep_*/: intermediate model-only checkpoints (model.pt, config.json, results.json, and string_atlas.json when present)tokenizers/*.json: tokenizer JSON fileswiki_qa_bpe32wiki_qa_bpe32: BPB 0.8559 at step 19083{
"anchor": "wiki_qa_bpe32",
"anchor_weight": 1.0,
"completed_steps": 19083,
"ctx_len": 2048,
"d_model": 2048,
"data_format": "packed_tokens",
"input_mode": "full",
"lr_decay_steps": 19083,
"model_size": "1B",
"n_heads": 32,
"n_layers": 20,
"output_mode": "full",
"run_name": "R5FW20B-spec-wiki_qa_bpe32-1B-20B-tillicum",
"tokens_processed": 20000205312,
"vocab_sizes": {
"wiki_qa_bpe32": 32000
}
}
--- library_name: pytorch license: other tags: - string-atlas - tokenizer-interface - fineweb-edu - realistic-bpe5 --- # RealisticBPE5 Wiki-QA Specialist 1B 20B Final 1B wiki-qa RealisticBPE5 specialist checkpoint trained on the wiki-qa 20B packed FineWeb-Edu stream. HF repo: `cminst/realistic-bpe5-wiki-qa-specialist-1b-20b`. Source run directory name: `R5FW20B-spec-wiki_qa_bpe32-1B-20B-tillicum`. Data bundle: `cminst/realistic-bpe5-wiki-qa-20b`. ## Files - `model.pt`: final model weights - `config.json`: training/model configuration saved by `train_v3.py` - `results.json`: final evaluation metrics and logs - `step_*/`: intermediate model-only checkpoints (`model.pt`, `config.json`, `results.json`, and `string_atlas.json` when present) - `tokenizers/*.json`: tokenizer JSON files - optimizer and DDP resume state are intentionally excluded ## Tokenizers - `wiki_qa_bpe32` ## Final Metrics - `wiki_qa_bpe32`: BPB `0.8559` at step `19083` ## Config Snapshot ```json { "anchor": "wiki_qa_bpe32", "anchor_weight": 1.0, "completed_steps": 19083, "ctx_len": 2048, "d_model": 2048, "data_format": "packed_tokens", "input_mode": "full", "lr_decay_steps": 19083, "model_size": "1B", "n_heads": 32, "n_layers": 20, "output_mode": "full", "run_name": "R5FW20B-spec-wiki_qa_bpe32-1B-20B-tillicum", "tokens_processed": 20000205312, "vocab_sizes": { "wiki_qa_bpe32": 32000 } } ```
Source context: 38 downloads · 0 likes · Library pytorch · Repo cminst/realistic-bpe5-wiki-qa-specialist-1b-20b