Active-SWE

Active-SWE

A software engineering benchmark for proactive bug fixing task, involving patch generation, test generation and test validation.

Leaderboard

Evaluation protocol

All model evaluations reported on this leaderboard were conducted using Claude Code.

Submit new results

Evaluated a new model on Active-SWE? Contact haobinli.gm@gmail.com or pingdeng001@gmail.com to have your results reviewed and added to the leaderboard.

Model Performance

Active-SWE · Resolved Performance across independently evaluated models
Closed-source Open-source

Cost Efficiency

Active-SWE · Resolved vs. cost Median cost per instance · price increases from left to right
Efficiency frontier Closed-source Open-source

Cost methodology

Cost is calculated for each evaluation instance over the complete Recorded → Potential trajectory using all requests from the raw Claude logs. All input, cache-read, cache-creation, and output tokens are included, with duplicate streamed events removed by message ID. The table reports the median cost across instances; the Judge stage is excluded. Official on-demand prices are used, without cache, Batch, free-tier, promotional, or gateway discounts.

Green line: The Pareto cost-efficiency frontier, not a trend line. A model joins the frontier only when it outperforms every cheaper model under the selected score.

Dataset

Task Formulation

Proactive bug-fixing task evaluates the ability to proactively discover and fix bugs in real-world software engineering without issue reports.

About

Citation

If you found the Active-SWE helpful for your work, please cite as follows:


    

Disclaimer

Active-SWE is intended for research purposes only. Models evaluated on Active-SWE may produce unexpected or incorrect results. We are not responsible for any damages caused by the use of Active-SWE, including but not limited to any loss of profit, data, or use of data.