Research · 2026

ExpenseLM

Fine-tuning a 4B model to turn messy expense text into policy-checked JSON. SFT more than doubled policy accuracy.

View on GitHub

Role
Researcher
Years
2026
Links

What it is

Qwen3-4B, fine-tuned with SFT (QLoRA) and then DPO on its own harvested failures, to turn messy expense text into policy-checked JSON. It’s measured by a six-system, five-metric evaluation harness against zero-shot, few-shot, and a frontier-model ceiling.

Key result

SFT taught the policy reasoning that prompting alone couldn’t. Policy accuracy went from 27% to 55%, and extraction nearly matched the frontier model.

SystemExpense F1Policy accuracy
Qwen3-4B zero-shot0.98027.1%
Qwen3-4B 5-shot0.98834.3%
+ SFT (QLoRA), answered set1.00055.5%
Frontier model (ceiling)1.00090.4%

The DPO and quantized GGUF runs are in progress. Built with Python, PyTorch, Unsloth, TRL, and Transformers.