Source du modèle
Extrait de la source
Artefacts du modèle
1 artefactadapter_model.safetensors
safetensors · 369 MB · SHA-256 16b76011c0f9…5f02
Extraits de sources
3 extraits--- license: cc-by-4.0 language: - en base_model: google/gemma-4-E2B tags: - natural-language-autoencoder - nla - interpretability - mechanistic-interpretability - gemma - consumer-gpu - peft - lora library_name: peft pipeline_tag: text-generation --- ## Evaluation across released versions  Content-fidelity doc-level retrieval and reconstruction round-trip cosine across the released AV versions. The verbalizer's content-surfacing is domain-sensitive: at chance on out-of-domain news, modestly but significantly above chance in-domain. Round-trip cosine is structural-projection dominated, not a faithfulness metric. Regenerate with `make_nla_eval_figure.py` as new versions / evaluations land. # Gemma-4-E2B NLA AV (Activation Verbalizer) — v0.0.1 LoRA adapter for `google/gemma-4-E2B` that takes a 1536-dimensional residual-stream activation captured at layer 23 and produces a natural-language explanation of what the activation represents. This is the **first non-Anthropic-team open-source NLA Activation Verbalizer** released publicly. Trained end-to-end on a single 4 GB consumer GPU (NVIDIA GTX 1650 Ti Max-Q) following a customized variation of the methodology of Fraser-Taliente, Kantamneni, Ong et al. 2026 ([Transformer Circuits](https://transformer-circuits.pub/2026/nla/)). Pairs with the matched [`Solshine/gemma-4-e2b-nla-L23-ar-v0_0_1`](https://huggingface.co/Solshine/gemma-4-e2b-nla-L23-ar-v0_0_1) reconstructor. ## How to use ```python from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig from peft import PeftModel import numpy as np import torch BASE = "google/gemma-4-E2B" AV_REPO = "Solshine/gemma-4-e2b-nla-L23-av-v0_0_1" # Injection convention INJECTION_TOKEN_ID = 249568 # ㊗ INJECTION_LEFT_NEIGHBOR_ID = 236813 # < INJECTION_RIGHT_NEIGHBOR_ID = 954 # > INJECTION_CHAR = chr(0x3297) D_MODEL = 1536 INJECTION_SCALE = float(np.sqrt(D_MODEL)) # = 39.2; matches Gemma-4-E2B token-embed norm PROMPT = ( "You are a meticulous AI researcher conducting an important investigation " "into activation vectors from a language model. Your overall task is to " "describe the semantic content of that activation vector.\n\n" "We will pass the vector enclosed in <concept> tags into your context. " "You must then...
Source context: 7 downloads · 1 likes · Pipeline text-generation · Library peft · Repo Solshine/gemma-4-e2b-nla-L23-av-v0_0_1
Source context: 32 downloads · 0 likes · Pipeline text-generation · Library peft · Repo Solshine/gemma-4-e2b-nla-L23-av-v0_0_1