Title graphic reading: A zero is not a missing value. Validating AI output at the persist boundary.

I wanted to write this one because it covers the Diamond Elite bugs I would most want other teams to know about. In every AI application there’s a point where a model’s response stops being a response and becomes data. Before that point, a bad answer is an annoyance the user can retry. After it, a bad answer is a row in your database that will be averaged, charted, trended, and used to tell someone something false about themselves for months.

In Diamond Elite that point matters more than usual, because the output is a score. A parent uploads a swing and gets a number from 1 to 10. That number feeds a weekly average, a season progress chart, a slump-detection baseline, and a forward-looking development forecast. It doesn’t scroll away like a chat response. It’s a measurement, and measurements build on each other.

I learned that the hard way, twice. This post covers both failures and the validator we built to stop them.

Session summary showing average, best and worst swing scores and a trend

The aggregates a single bad score contaminates: session average, best, worst and trend. Demo data, synthetic player names.

Failure One: Repairing Truncation

Language models occasionally emit JSON that arrives malformed, whether it’s a trailing comma, an unescaped character, or a stray token. The standard fix is a repair utility: parse, and if parsing fails, apply a few structural fixes and try again. We had one, and it worked well. It closed unbalanced braces and brackets and recovered responses that would otherwise have been thrown away.

For a while, it also quietly made things up.

Here’s how it happened. A full swing analysis is generated under a token ceiling. Every so often a response is long enough to hit that ceiling, and the model stops mid-object, say four phases into a seven-phase breakdown, partway through a sentence, with no overall score emitted yet. The JSON is structurally incomplete. The repair utility does exactly what it was built to do: it closes the open braces, produces syntactically valid JSON, and hands back a parsed object.

That object then got stamped with a schema version and saved, and from that point on it looked exactly like a complete analysis. Four phases instead of seven, a breakdown ending mid-thought, and, most importantly, no score, which brings me to the second failure.

The fix is one check, placed before anything else touches the response: if the model stopped because it hit the token cap, throw. Don’t parse it, don’t repair it, and don’t persist it. The stop reason is available on the response, and it tells you clearly whether the model finished or was cut off. Our validator exposes this as a single assertion that runs before the JSON is trusted, and it raises a typed error that the client turns into an honest “that got cut off, please try again.”

The way I think about it: a JSON repair utility is a transport-layer tool, and truncation is not a transport problem. Repair makes sense for a response that’s complete but malformed. It’s actively dangerous for a response that’s well-formed but unfinished, because it turns a loud failure into a silent one. Check completeness first, then repair, never the other way around.

Failure Two: The Zero That Wasn’t There

The second failure is smaller and more common, and in some ways I think it’s worse.

Side-by-side comparison: defaulting a missing score to zero silently corrupts averages and forecasts, while leaving it null fails visibly and gets rejected.

When an analysis came back without a score, either because of the truncation above or because the model simply left the field out, the code did the ordinary defensive thing and defaulted it. A missing score became 0.

Zero is a terrible default for a 1-to-10 rating, and the reason is that it looks fine. It’s a number, it has the right type, and it passes every schema check that only asks “is this numeric.” Then it flows downstream into a weekly average and drags it down by a whole grade. It sets a slump-detection baseline that triggers alerts for a slump that never happened. It anchors a development forecast to a data point that doesn’t exist.

A null would have shown up right away. Charts would have gapped, averages would have excluded it, and someone would have noticed within a day. A 0 spread invisibly, as a bad measurement instead of a missing one. Personally, I’d take no number over a wrong number every time.

The rule we follow now: never default a measurement. If the model didn’t produce a score, the correct value is null and the correct behavior is to reject the analysis, not to invent a floor for it. Defaults are fine for configuration and preferences, but in my opinion they don’t belong on observations.

What the Validator Actually Does

What came out of both incidents is a single module that sits at the persist boundary. Every AI scoring path in the system, including hitting analysis, pitching analysis, catching analysis, and automated game-data analysis, goes through it before anything reaches the database. It does four things:

Four validator checks: assert completeness first, clamp ratings to 1 to 10, re-derive the label from five allowed values, and require all seven phases.

  • Asserts completeness. It checks the stop reason and throws a typed truncation error before the JSON is parsed or repaired.
  • Clamps ratings into range. Any score is coerced to a finite number, rounded to one decimal, and bounded to 1–10. A non-finite or non-numeric value becomes null, which fails validation instead of turning into zero.
  • Re-derives labels rather than trusting them. The response includes a human-facing label: Excellent, Great, Good, Average, or Needs Work. If the returned label isn’t one of those exact values, it gets recomputed from the numeric score against fixed thresholds. That way the model can’t invent its own grade vocabulary, and the label can never disagree with the number it’s supposed to describe.
  • Requires structural completeness. A hitting analysis must cover all seven phases. Four phases isn’t a partial success, it’s a failure, and it gets treated as one.

The whole module is under two hundred lines, and I’d argue it’s the most valuable code in the application.

Why This Generalizes Past Sports

Everything above applies to the AI features enterprises are actually building, without changes. Substitute “swing score” for “risk rating,” “invoice confidence,” “sentiment score,” “document classification,” or “extracted contract value,” and the failure modes are the same.

An extraction pipeline that defaults a missing contract value to zero will under-report a portfolio. A classification step that persists a truncated response will silently drop categories. A scoring model whose label is trusted instead of derived will eventually produce a record where the number says 8 and the label says “Poor,” and whoever finds it will lose confidence in the entire system, and they’d be right to.

The uncomfortable part is that none of this gets caught by the checks teams usually build. Schema validation passes, because the types are right. Monitoring stays quiet, because no exception was thrown. Evals on the model itself pass, because the model wasn’t really the problem. The bad data comes entirely from the code between a valid-looking response and a database write, which is the one layer I’ve found most AI pilots treat as plumbing.

Getting Started

If you have an AI feature writing to a system of record, here’s where I’d start:

  1. Find your persist boundary and name it. There should be exactly one place where model output becomes stored data. If there are five, chances are you’ll only validate three of them.
  2. Check the stop reason before you parse. Every major model API reports whether generation finished or hit a limit. Truncated output should never reach a repair step.
  3. Audit every default applied to model output. Search your code for the null-coalescing and fallback patterns around AI responses. Each one could be silently turning “unknown” into “wrong.”
  4. Re-derive anything derivable. If a label, category, or band can be computed from a value you already validated, compute it. Don’t count on the model staying consistent with itself across fields.
  5. Clamp every numeric to its real range. Not just its type. A confidence score of 1.4 and a rating of -3 should both be impossible by the time they reach storage.

The next post in this series leaves the model behind and goes into the part that taught me the most per line of code: video. Files that won’t seek, thumbnails that never generate, and a 4.3 GB upload from an ordinary phone.

If you’re putting AI output into a system of record and want the boundary reviewed before it turns into a data-quality problem, feel free to contact us.

CB5 Solutions is a Microsoft Solutions Partner specializing in Microsoft Security, Data & AI, Modern Workplace, and Azure Infrastructure.