Active-SWE
A software engineering benchmark for proactive bug fixing task, involving patch generation, test generation and test validation.
Leaderboard
All model evaluations reported on this leaderboard were conducted using Claude Code.
Evaluated a new model on Active-SWE? Contact haobinli.gm@gmail.com or pingdeng001@gmail.com to have your results reviewed and added to the leaderboard.
Model Performance
Cost Efficiency
Cost methodology
Cost is calculated for each evaluation instance over the complete Recorded → Potential trajectory using all requests from the raw Claude logs. All input, cache-read, cache-creation, and output tokens are included, with duplicate streamed events removed by message ID. The table reports the median cost across instances; the Judge stage is excluded. Official on-demand prices are used, without cache, Batch, free-tier, promotional, or gateway discounts.
Green line: The Pareto cost-efficiency frontier, not a trend line. A model joins the frontier only when it outperforms every cheaper model under the selected score.
Dataset
Task Formulation
Proactive bug-fixing task evaluates the ability to proactively discover and fix bugs in real-world software engineering without issue reports.
About
Citation
If you found the Active-SWE helpful for your work, please cite as follows:
Disclaimer
Active-SWE is intended for research purposes only. Models evaluated on Active-SWE may produce unexpected or incorrect results. We are not responsible for any damages caused by the use of Active-SWE, including but not limited to any loss of profit, data, or use of data.