Notes

An LLM Judge as a Deploy Gate

In summer 2026 I worked as an AI and ML engineering intern at Tresic. One of the clearest pieces of that work was an LLM-as-a-judge evaluation harness built to act as a production deploy gate. The question was not whether a model could produce a strong demo. It was whether the evidence was strong enough to trust a candidate in a real rollout.

That distinction mattered because the harness was meant to sit in the path of a decision, not after it as a post-hoc report. A report can summarize what already happened. A gate has to say yes or not yet before the system moves. If evaluation only appears once the choice is emotionally made, it becomes decoration.

The context was open-weight models considered for use across regions. Open-weight changes the trust problem. You inherit more of the behavior surface yourself, and you cannot lean as hard on a vendor narrative about what the model will do. Multi-region raises the bar again: the same candidate has to hold up under different deployment conditions, constraints, and failure costs, not just look good in one comfortable setting.

The failure mode I cared about most was the false positive. Demos hide risk by showing the best case. Single pass-rates hide risk by collapsing a distribution into one number. Both can make a system look ready when the evidence underneath is thin.

False positives are expensive in a specific way. They do not only ship a weaker model. They burn reviewer trust. After one confident pass that later falls apart, every later score looks like marketing. People stop reading the harness as a decision tool and start reading it as theater.

So the design bias was toward high precision and failing closed. I would rather withhold a pass than publish a score that looks strong and then collapses. A delayed deploy is recoverable. A burned review process is harder to rebuild.

The first concrete move was to calibrate a quality threshold on purpose. A number is not meaningful because it is high. It is meaningful because someone decided what good enough means for this job, before looking for a flattering result.

The second move was a hard split between development and decision measurement. Tuning, prompt work, and exploratory runs stayed on one side. The score that counted ran blind against held-out examples. Without that split, it is too easy to reward yourself for fitting the set you already saw.

The third move was to stop treating a point estimate as enough. I used bootstrap confidence intervals and paid attention to the lower bound. A promising mean can still sit on thin evidence. If the lower bound does not clear the bar, the claim is not ready, even if the center of the interval looks fine.

Blocked versus passed, as a decision rule, felt less like a trophy and more like a sentence about evidence. Pass meant the held-out result and the lower bound cleared the calibrated threshold with enough support behind them. Block meant the mean could look encouraging while the interval, the split, or the threshold still said not yet.

That sounds simple in retrospect. In practice it is uncomfortable, because the pressure in the room usually points the other way. People want the green light. The harness exists to keep the green light expensive.

I write about the same loop more generally in What I Mean by Evaluation: define what good means, measure held-out behavior, understand uncertainty, and be willing to say the result is not ready. The Tresic gate is where that habit had to carry a real deploy decision instead of staying abstract.

What I would tighten next is the part that turns a blocked result into a useful next action. A fail-closed gate is only half the system. The other half is making the miss diagnosable: which failure modes dominated, where the threshold was brittle, and what evidence would actually move the lower bound.

I would also push harder on keeping the decision rule boring and stable. If the threshold or the split keeps moving every time a result is inconvenient, the harness stops being a gate and becomes a negotiation.

And I would keep resisting the urge to summarize with one headline number. The useful artifact is the decision boundary: what we required, what we measured blind, how uncertain we were, and why that was or was not enough.

That is the kind of evaluation work I want to keep doing. Not scoreboards for their own sake, and not demos that persuade by highlight reel, but systems that make trust contingent on evidence and keep the costly pass hard to earn.