A coding agent should be judged by the work it completes. Here is our proposed evaluation framework, published before the results.
We have not run or published an Ablitron benchmark. The graphic on the homepage is illustrative and represents no measurements.
The question
Does an abliterated model complete legitimate coding tasks more reliably than its original checkpoint when both run in the same agent environment?
We want to distinguish changes in model behavior from changes in the agent, tools, prompts, and inference settings. Comparisons should hold those conditions constant whenever possible.
What we plan to measure
| Measure | Evidence |
|---|---|
| Task completion | Task-specific checks against the final repository state. |
| Instruction adherence | A rubric for explicit constraints, scope, and output requirements. |
| Unnecessary refusal | Human-reviewed refusal labels on clearly legitimate tasks. |
| Unrequested changes | Diff review for edits outside the requested scope. |
| Cost and time | Recorded usage, price assumptions, and end-to-end runtime. |
| Recovery | Whether the agent corrects a failed check within the run budget. |
A reproducible run
- Freeze the task, repository commit, environment, and acceptance criteria.
- Record the model checkpoint, quantization, prompt, tools, and sampling settings.
- Use the same time and tool budgets for paired comparisons.
- Repeat runs to measure variability; report failures and timeouts.
- Publish permitted task data, run logs, costs, and the scoring method.
How we will report results
We plan to show absolute task counts, denominators, uncertainty, and per-task outcomes. We will identify who built and funded the evaluation, including our interest in Ablitron, and explain limitations rather than presenting one universal “best model” score.
What this will not prove
A narrow coding test does not establish overall model safety, capability, or suitability for every repository. Public task contamination, small samples, judging errors, and provider changes can all affect results.
See the build roadmap