A Benchmark Should Outrun the Tools It Grades
We build a scanner and the benchmark that grades it, and we say so. The benchmark is built to cover more than any single tool, so a score reflects real capability, not a test shaped to flatter.
We build both a code scanner and the benchmark that grades it, and we do not hide that they come from the same place. The obvious worry writes itself: a house benchmark drifts to cover exactly what the house tool does well, the score climbs, and it stops meaning anything. So we build the benchmark to do the opposite of flatter.
Wider than the tool, on purpose
A benchmark shaped to a tool’s strengths is a mirror. It tells you what you already knew and hides the rest. Ours is built to reach past any single scanner, into languages and code shapes that our own tool does not fully handle yet. When TheAuditor has a gap, BenchProctor is designed to land on it and show it, not to route around it. A grader that cannot embarrass its own team is not a grader.
Honesty by construction, not by distance
The reason this works is not that the two products pretend to be strangers. They share an origin and we say so. The honesty is built into the artifacts. The graded code carries no labels, no hint about which line is the planted weakness. The answer key lives separately from the code a scanner sees. So even the people who build both cannot quietly tune one to the other, because the thing you would tune toward is not in front of the tool at all.
Why one roof is a discipline
Running a scanner and its grader under one roof is only a conflict of interest if you build the grader to win. Build it to outrun the tool instead, and the same roof becomes the opposite: the one place willing to measure your product against a bar it did not get to lower. That is the standard we hold across everything at Code Reality Labs.