SWE-agent
SWE-agent/SWE-agent View on GitHub20,119 starsPythonUpdated Aug 2026
The research agent that solves real GitHub issues - and shows along the way how to measure agents in a comparable fashion at all.
Most agents look good in a demo. Whether they are actually any good only shows on a task nobody arranged for them - real, already closed GitHub issues, for instance, whose solution is known and whose tests you can run.
That is exactly the subject here. SWE-agent is an agent that works on such issues, and at the same time the showcase for a claim that is easy to underestimate when building agents: the tool surface is not a detail. A file editor that hands the model line numbers and context back after every change moves the success rate more than the tenth sentence in the system prompt. Anyone designing their own tools will find the most convincing argument here for shaping them from the model's point of view rather than the programmer's.
A second reason to look: the setup of the evaluation. One container per task, fixed starting conditions, a test run as the verdict - no model jury. That is the template for measuring your own agent instead of admiring it.
- eval
- benchmark
- research
Community rating
Sign in to rate this project.
Sign inDiscussion· no posts yet
Our comment agent reads every new post, says thanks or recommends related content.
Be the first voice - what do you think?
Sign in to join the discussion.
Sign in