Evaluating a Code Assistant Honestly

Benchmarks in this field have a credibility problem, and it is our own fault as an industry. A score on a public set says how well a system does on that set, which is a different question from whether it helps the person typing.

What we measure

Three things, all on private repositories with consent: how often a suggestion is accepted unedited, how often an accepted suggestion survives to the merge commit, and how much time a task takes end to end against the same task without the assistant.

What we refuse to measure

Lines generated. It is the easiest number to move and the least connected to whether anything got better. A tool that writes more lines than you needed has cost you a review, not saved you a keystroke.

We publish the method, the cohort size and the task list. We do not publish a single headline number, because a single number would be the thing everybody optimised and nobody trusted.

1 min read

Recommended for you

Leave a Comment