This repository contains the LoRA adapter for the Direct Preference Optimization (DPO) stage of Qwen3-1.7B-UltraChat-SFT.
Fuente del modelo
Extracto de la fuente
This repository contains the LoRA adapter for the Direct Preference Optimization (DPO) stage of Qwen3-1.7B-UltraChat-SFT.
Fuentes
1 fuenteVerificado 8 ago
Artefactos del modelo
1 artefactoExtractos de fuentes
2 extractos--- library_name: peft tags: - qwen3 - trl - dpo - alignment - preference-optimization - lora - conversational - text-generation-inference license: apache-2.0 datasets: - argilla/ultrafeedback-binarized-preferences-cleaned - argilla/distilabel-capybara-dpo-7k-binarized - argilla/distilabel-intel-orca-dpo-pairs base_model: - ayushshah/Qwen3-1.7B-UltraChat-SFT pipeline_tag: text-generation --- # Qwen3-1.7B Chat LoRA Adapter This repository contains the **LoRA adapter** for the Direct Preference Optimization (DPO) stage of [Qwen3-1.7B-UltraChat-SFT](https://huggingface.co/ayushshah/Qwen3-1.7B-UltraChat-SFT). The adapter aligns the SFT model using human preference datasets to improve response quality, reasoning, helpfulness, and truthfulness while preserving the conversational capabilities learned during supervised fine-tuning. View the [merged model](https://huggingface.co/ayushshah/Qwen3-1.7B-Chat).<br> View the [GitHub](https://github.com/AyushShahh/Qwen3-1.7B-Post-Training/tree/main) repo for Post-training pipeline. > **Note:** This repository contains only the DPO adapter weights. It must be loaded on top of the <u>Qwen3-1.7B-UltraChat-SFT</u> model. ## Training **Alignment Method**: Direct Preference Optimization (DPO)<br> **Frameworks**: TRL, Unsloth, PEFT (LoRA) ## LoRA Configuration Applied modules: `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj` Configuration: - LoRA Rank: 32 - LoRA Alpha: 64 - LoRA Dropout: 0 ## Evaluation Compared to the SFT model, this DPO adapter brings improvement to the model while introducing a small trade-off in strict instruction-following performance. | Benchmark | SFT | DPO | Δ | |-----------|-----|-----|---| | MMLU-Pro | 37.91 | 39.86 | **+1.95** | | GSM8K | 71.27 | 72.48 | **+1.21** | | HellaSwag | 62.28 | 62.26 | -0.02 | | TruthfulQA <sub>mc2</sub> | 51.62 | 53.09 | **+1.47** | ARC Challenge | 44.97 | 43.26 | -1.71 | | IFEval <sub>strict prompt</sub> | 33.09 | 30.13 | -2.96 | Blind **A/B Evaluation** tests with **LLM-as-a-judge** and **Human-In-The-Loop** calibration preferred the DPO model over the SFT model in majority of the runs. ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer from peft import PeftModel base_model = AutoModelForCausalLM.from_pretrained( "ayushshah/Qwen3-1.7B-UltraChat-SFT" ) model = PeftModel.from_pretrained( base_model, "...
Source context: 48 downloads · 0 likes · Pipeline text-generation · Library peft · Repo ayushshah/Qwen3-1.7B-Chat-LoRA