A LoRA fine-tune of HopCoder-Mini-9B (a Qwen3.5 multimodal model) that teaches the model to emit native tool-call blocks in a compact XML-like format instead of JSON-based function-calling schemas. Trained on a Modal H200 GPU in BF16 precision using a blend of xLAM and Hermes function-calling data plus 1,920 targeted CLI-tool examples.
These are the real CLI agent tools the adapter was specifically trained on:
ask_user_question — Show interactive questions in the CLI
todo_write — Create or update a structured task list
read_file — Read a UTF-8 text file
search_code — Search source files for a text or regex pattern
glob — Find files by glob pattern
grep_search — Search file contents for a regex pattern
edit — Replace text in a file with new content
run_shell_command — Execute a shell command and return output
Training hyperparameters
Parameter
Value
Learning rate
1e-4
Epochs
1.0
Train batch size
4
Eval batch size
4
Gradient accumulation
4
Effective batch size
16
LR scheduler
Cosine
Warmup ratio
0.05
Weight decay
0.01
Max grad norm
1.0
Optimizer
AdamW (fused)
Precision
BF16 + TF32
Gradient checkpointing
Disabled
Dataloader workers
Training metrics
Metric
Value
Total steps
453
Train loss
0.0034
Eval loss
0.0237
Tool prompt format
The system prompt uses compact tool signatures instead of verbose JSON schemas:
Find files by glob pattern (e.g., **/*.py).
Search file contents for a regex pattern.
Required parameters have no suffix; optional parameters are marked with ?
Type annotations are compact (e.g., array[object{label,description}])
Descriptions are truncated to 120 characters (tools) / 60 characters (parameters)
How to Use
Installation
pip install torch transformers peft
Quick start
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor
from peft import PeftModel
MODEL_ID = "TaimoorSiddiqui/Hopcoder-Mini-9B"
ADAPTER_ID = "TaimoorSiddiqui/Hopcoder-Mini-9B-Native-ToolCall-LoRA-H200"
processor = AutoProcessor.from_pretrained(MODEL_ID, trust_remote_code=True)
tokenizer = processor.tokenizer
if tokenizer.pad_token_id is None:
tokenizer.pad_token = tokenizer.eos_token
model = AutoModelForImageTextToText.from_pretrained(
MODEL_ID,
trust_remote_code=True,
dtype=torch.bfloat16,
device_map="auto",
)
model = PeftModel.from_pretrained(model, ADAPTER_ID)
model.eval()
Generating a tool call
SYSTEM_PROMPT = (
"Use the provided tools whenever the request requires one.\n\n"
"For a tool request, emit only complete native tool-call blocks. "
"Never emit a function name as a top-level tag. Never leave unmatched "
"parameter, function, or tool_call tags. Arrays and objects inside "
"parameter blocks must be valid JSON. Do not use Markdown fences.\n\n"
)
TOOLS_XML = (
" \n"
" "
"Find files by glob pattern (e.g., **/*.py). \n"
" "
)
messages = [
{"role": "system", "content": SYSTEM_PROMPT + TOOLS_XML},
{"role": "user", "content": "Find all Python files in the project."},
]
prompt = processor.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=False,
)
inputs = tokenizer(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)
with torch.inference_mode():
outputs = model.generate(
**inputs,
max_new_tokens=384,
do_sample=False,
repetition_penalty=1.05,
eos_token_id=tokenizer.eos_token_id,
pad_token_id=tokenizer.pad_token_id,
)
generated = outputs[0, inputs["input_ids"].shape[1]:]
print(tokenizer.decode(generated, skip_special_tokens=True).strip())
Generation parameters
Parameter
Value
max_new_tokens
384
do_sample
False (greedy)
repetition_penalty
1.05
enable_thinking
False
Benchmark
The benchmark evaluates the adapter on 40 cases:
30 targeted cases — 5 per CLI tool x 6 tools (ask_user_question, todo_write, glob, grep_search, edit, run_shell_command)
10 general cases — xLAM-style queries with 10 different tools (get_weather, search_flights, calculate_mortgage, send_email, book_restaurant, get_stock_price, create_event, translate_text, get_directions, set_reminder)
Validation criteria
Each generated tool call is validated on:
Complete tool-call block — must contain at least one valid block
No extra prose — no text outside tool-call blocks
No markdown fences — no code blocks
Balanced tags — matching counts of opening/closing tags for tool_call, function, and parameter
Correct function name — the called function matches the expected one
Valid JSON — array/object parameter values must be valid JSON
Required parameters — all required parameters must be present
No top-level function tags — function name must not appear as a standalone XML tag
Running the benchmark
The benchmark runs on Modal with an H200 GPU:
python -m modal run --detach hopcoder_benchmark.py
The benchmark script (hopcoder_benchmark.py) is included in this repository.
Limitations
Function confusion — The model occasionally confuses similar tools (e.g., search_code vs grep_search, read_file vs edit)
Missing parameters — Rare cases of omitting required parameters for complex tools
Single-turn only — The adapter was trained on single-turn examples; multi-turn conversations may require additional fine-tuning
CLI-focused — The 8 targeted tools are CLI agent tools; the adapter has not been tested with real-world API tools
Compact schema format — The system prompt uses a compact XML-like tool signature format, not standard JSON schemas. This may not be compatible with all tool-calling frameworks