What it is
Qwen3-4B, fine-tuned with SFT (QLoRA) and then DPO on its own harvested failures, to turn messy expense text into policy-checked JSON. It’s measured by a six-system, five-metric evaluation harness against zero-shot, few-shot, and a frontier-model ceiling.
Key result
SFT taught the policy reasoning that prompting alone couldn’t. Policy accuracy went from 27% to 55%, and extraction nearly matched the frontier model.
| System | Expense F1 | Policy accuracy |
|---|---|---|
| Qwen3-4B zero-shot | 0.980 | 27.1% |
| Qwen3-4B 5-shot | 0.988 | 34.3% |
| + SFT (QLoRA), answered set | 1.000 | 55.5% |
| Frontier model (ceiling) | 1.000 | 90.4% |
The DPO and quantized GGUF runs are in progress. Built with Python, PyTorch, Unsloth, TRL, and Transformers.