CHAPTER 04 · Glossary: Benchmarks · 11 / 12
Measuring real coding
Repository and codebase
A repository (often shortened to repo) is the complete collection of files that make up a software project, together with its history of changes. Codebase is a near-synonym for the body of code itself. Real software work happens inside large repositories where a single feature is spread across many files, and understanding how those files interact is most of the job. SWE-bench is built on real repositories precisely so that it tests this skill, rather than the much easier skill of writing one isolated function.
Issue and pull request
On platforms like GitHub, an issue is a written report describing a bug or requesting a feature, in plain language. A pull request (PR) is a proposed bundle of code changes submitted to fix an issue or add something, usually reviewed and then merged into the project. The pair is the natural unit of software work: a problem stated as an issue, and a solution delivered as a pull request, often with tests proving the solution works. SWE-bench harvests exactly these issue-and-pull-request pairs from real projects to build its tasks.
Patch and diff
A patch (also called a diff) is a precise description of changes to make to code: which lines to remove and which to add, in which files. Rather than rewriting whole files, developers and models express edits as patches, which are compact and easy to apply automatically. In SWE-bench, the model's answer is a patch, and the benchmark grades it by applying that patch to the codebase and running the tests.
Unit test
A unit test is a small piece of code that automatically checks whether some part of a program behaves correctly, by running it on a known input and confirming the output is as expected. Projects accumulate many unit tests so they can catch mistakes automatically. Unit tests are what make software such a clean thing to grade: a fix either makes the relevant tests pass or it does not, no human judgment required.
Execution-based evaluation
Execution-based evaluation means scoring a model by actually running its output and observing the result, rather than comparing its text to a reference answer. In SWE-bench, the model's patch is applied and the project's real tests are run; the score is simply whether the tests pass. This is far stronger than checking whether the code looks similar to a human solution, because there are many valid ways to fix a bug, and the only thing that truly matters is whether the fix works.
HumanEval
HumanEval is an earlier and very popular coding benchmark in which a model is asked to write a small, self-contained function from a short description, checked by running a few tests. It was valuable but limited: its problems can typically be solved in a handful of lines and involve no surrounding codebase. SWE-bench was designed largely in contrast to HumanEval, replacing tidy isolated puzzles with messy real-world repository work.
SWE-Llama
SWE-Llama is a pair of open models (7 billion and 13 billion parameters) that the SWE-bench authors created by fine-tuning Meta's CodeLlama on their training set, SWE-bench-train. The point was to show that open models could be specialized for real software-engineering tasks; SWE-Llama had to handle very long contexts of over 100,000 tokens to read enough of a codebase, and in some settings it was competitive with much larger proprietary models.