We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Fonte do modelo
Trecho da fonte
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Fontes
1 fonteVerificado 17 de jul.
Artefatos de modelo
1 artefatoadapter_model.safetensors
safetensors · 4,13 GB · SHA-256 5abaaed207cb…3d95
Trechos de fonte
3 trechos--- license: llama3.1 base_model: meta-llama/Llama-3.1-8B-Instruct language: - en library_name: transformers pipeline_tag: text-generation tags: - prompt-injection - prompt-injection-defense - agent - tool-calling - agentdojo - dpo - drip - security --- # Llama-3.1-8B-Instruct · DRIP (4-role / tool-calling) A **prompt-injection-hardened** version of [`meta-llama/Llama-3.1-8B-Instruct`](https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct), trained with **DRIP** (*Defending Prompt Injection via Token-wise Representation Editing and Residual Fusion*). This is the **4-role / tool-calling** variant (`TextTextText-4roles`), built for **agentic** settings where injections hide inside **tool outputs** rather than the user turn. Chat format: **`system` → `user` → `tool` (untrusted) → `assistant`** (the untrusted segment uses Llama's native `ipython` role). - 📦 **Code:** https://github.com/lindsey98/PromptInjection - 📊 **Data:** [Zenodo 10.5281/zenodo.20603331](https://doi.org/10.5281/zenodo.20603331) ## What DRIP does DRIP adds two architectural modifications on top of the base model so that adversarial instructions hidden in untrusted content are treated as inert data: - **Token-wise de-instruction shift** — moves the representation of data/tool tokens away from directive semantics. - **Residual re-instruction fusion** — a residual path that keeps generation anchored on the legitimate top-level (system/user) instruction. The fuse pipeline assigns trusted/untrusted slots internally via `expert_labels`, so the model knows which tokens came from the untrusted `tool` role. ## Training | | | |---|---| | Base model | `meta-llama/Llama-3.1-8B-Instruct` | | Objective | DPO | | Architecture | DRIP fuse (`LlamaForCausalLMDRIP`) | | Delimiter | `TextTextText-4roles` | | Training data | `datasets/alpaca_injecagent_dpo_combined.json` — **~20,162** DPO pairs (~20K clean Alpaca + ~1K InjecAgent tool-call pairs) | | Epochs | 1 | The ~1K InjecAgent pairs familiarize the model with the tool-calling format and with injections planted in tool observations; clean Alpaca is included to match Meta SecAlign's benign training mix for a fair comparison. See the [AgentDojo README](https://github.com/lindsey98/PromptInjection/blob/main/testing/agentdojo/README.md) for how the data is built. ## How to use > ⚠️ This checkpoint is **not** a drop-in `AutoModelForCausalLM`....
Source context: 0 downloads · 0 likes · Pipeline text-generation · Library transformers · Repo Kelsey98/Llama-3.1-8B-Instruct-TextTextText-4roles-toolcall-drip
Source context: 0 downloads · 0 likes · Pipeline text-generation · Library transformers · Repo Kelsey98/Llama-3.1-8B-Instruct-TextTextText-4roles-toolcall-drip