Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps
A Hugging Face guide presents a public, inexpensive recipe for improving a small model's structured-output compliance. The procedure fine-tunes LiquidAI's LFM2.5-350M with Group Relative Policy Optimization (GRPO) using the TRL library and evaluates it on the IFStruct benchmark. The full run uses around 500 samples and 100 training steps, sized for a free-tier Colab or Kaggle GPU, with code available on GitHub.
On the IFStruct benchmark, performance rose from 22.6% to 29.7%. The guide notes the pipeline is not the one used to train the RL model described in the IFStruct blog, and does not aim to recreate that benchmark score, but to show how task-specific fine-tuning of smaller models can improve performance.
The base-model evaluation serves LFM2.5-350M locally through llama.cpp using the BF16 GGUF, reproducing a local score of 22.6%, described as close to the 21.1% reported in the IFStruct blog, and used as the baseline for the same serving stack. Evaluation ran on a MacBook Pro with an Apple M5 Max and 36 GB of unified memory. The full evaluation covers 2000 samples.
Training data is nvidia/Nemotron-RL-instruction_following-structured_outputs. To close distribution gaps with IFStruct, 40% of prompts get a fenced code block instruction appended and a disjoint 20% become top-level-array tasks with a required item count. A LoRA adapter with r=16 and alpha=32 targets LFM-specific modules, training about 6M parameters, roughly 1.66% of the model. Three reward functions are combined with weights [1.0, 0.5, 2.0]: json_format_reward, field_count_reward and schema_validation_reward. Training used 100 steps, 8 generations per prompt group, learning rate 5e-5, temperature 1.1 and beta 0.01.
After merging the adapter, converting to BF16 GGUF and rerunning evaluation, JSON pass rates rose from 18.0% to 31.9% while YAML stayed roughly flat (27.2% to 27.5%). The guide states this remains below the Qwen3.5-2B score of 33.15%.
Based on reporting from the original publisher. Visit the source for full context and later updates.