Project · 2026

Warrant

Checks whether an AI's fix actually fixed the bug. A failure replayer with a cheating detector.

View on GitHub

Role
Creator
Years
2026
Status
Open source
Links

The problem

When one AI writes both the fix and the test, the test doesn’t check the fix. It restates it. If the model misunderstood the bug, both are wrong in the same direction, and everything looks green.

How it works

Warrant records a real failure, then does four things, with no AI judgment involved: replay, revert, vary, scan.

  1. Replay the failure to get a baseline rate (say, 7 out of 10).
  2. Apply the fix and replay again.
  3. Undo the fix and replay again. If the failure doesn’t come back, something else fixed it.
  4. Vary the inputs, and scan the diff for cheating: an edited test, a special-cased input, a deleted assertion, a swallowed error.

The honest part

Warrant started as a startup thesis for a broad agent control plane. I concluded that most companies already cover that with guardrails, human review, and tests, and the problem wasn’t acute enough to be a business. So I cut it down to the one question nobody else answers and released it open-core under Apache-2.0.