You Cannot Grade Your Own Homework
Most AI coding setups are a closed room where the model does the work and then declares it correct. Closing the loop takes three separate moves: know the code, act on it, and prove the result against something you did not grade yourself.
There is a quiet flaw in almost every AI coding setup, and it hides in plain sight. The model writes the code. The model reviews the code. The model decides the code is correct. Then it reports that the work is done. Read that back slowly. The thing being graded and the thing doing the grading are the same thing. That is not a feedback loop. It is an echo.
An agent that marks its own exam will tend to pass. Not because it is dishonest, but because confidence and correctness come out of the same machinery, and that machinery has no outside reference to check itself against. It feels sure. Sure is not the same as right, and you have shipped enough confident, fluent, wrong code to know the difference in your bones.
A real loop fixes this, and a real loop has three moves that have to stay separate: know the code, act on it, and prove the result against something you did not grade yourself. Miss any one of them and you are back in the closed room. Here is what each move takes, and which part of the Code Reality Labs stack does it.
Know the code cold
Before an agent can act well, it has to actually know what it is working on: the real call graph, the true data flow, where the framework boundaries sit, what a change will touch two modules over. Skip that and the agent invents a picture, and the invented picture is where the wrong answer is born.
TheAuditor hands the agent a deterministic account of the code instead of a plausible story assembled from a few file reads. Curator adds the other half of knowing: the history and the decisions that led here, kept by whether they are still true rather than by how recent they are, on your own GPU. One gives the agent the code as it actually is. The other gives it the memory a teammate would already carry in their head. Together they mean the agent starts every task knowing the system instead of guessing at it.
Act on it without handing over control
Knowing is not enough. The agent has to act, and acting is where cost and risk actually live. This is the move where a careless setup burns your budget on boilerplate and lets an agent do something irreversible while you are looking the other way.
Warden is the cockpit. It treats facts as facts instead of flattening them into prose, watches a fleet of parallel sessions, surfaces the one that is blocked waiting on you, holds the turn while you approve a plan, and tracks spend to the decimal. Arbiter sits over the top and routes each task to the cheapest model that can actually do it, keeps a local tier that costs nothing for mechanical and routine work, recovers a run that a crash or a throttled account would otherwise lose, and answers to a code on your phone from a machine that opens no network port by default. You act at scale, and you stay the one in control.
Prove it against something you did not grade
Here is the move the closed room never makes. Knowing and acting still leave you holding a verdict the system produced about itself. To trust that verdict, you need a check the system did not author and cannot memorize.
That is BenchProctor. It scores any code scanner, ours included, against an answer key it cannot read off a filename. The benchmark is combinatorial, rotating, and adversarial, and it is growing more cross-language with each release, so a scanner cannot easily overfit to it, and it is scored on an honest metric that rewards real detection rather than a number a tool can inflate by flagging everything. It is open under Apache-2.0, because a yardstick you have to take on faith is not a yardstick. The whole point is the independence: you did not write the answer key, the tool cannot see it, so the score it comes back with actually means something. That is exactly what grading your own homework can never give you.
An open loop drifts, a closed loop corrects
Put the three moves together and the reason for their order becomes clear. When the writer grades the writer, small errors survive, get built on, and compound quietly until something breaks in production. When an outside check that you did not create sits at the end of the loop, those errors get caught, recorded, and corrected, and the correction carries into the next session instead of evaporating overnight. The gap between the two is the gap between a system that drifts and one that gets sharper the longer you run it.
None of this requires all five products at once. Each move earns its place alone, and you can start with the one that hurts most today. But notice what the third move buys you that the first two cannot buy for themselves. Knowing and acting make the agent good. Proving, against something you did not grade, is the only part that turns the verdict into something you can actually bank on.