Most consultancies that talk about AI have never had to run any. There is a meaningful difference between designing a Copilot readiness roadmap for a client and being the person who finds out that a vision model returned malformed JSON overnight and quietly poisoned a week of user-facing averages. The first teaches you the architecture. The second teaches you what the architecture actually costs.
So we built a product. Diamond Elite is an AI analytics platform for baseball and softball — a parent or coach uploads a video of a swing, and the system returns a phase-by-phase mechanical breakdown, a score, annotated frames, and drills specific to what it saw. It also ingests game data, builds scouting reports, tracks player development across a season, and generates highlight reels. It is live, it has real families using it, and it has taught us more about production AI engineering than any readiness assessment we have ever run.
This is the first of five posts about how it is built and what building it taught us. This one covers the shape of the system, and the single most expensive infrastructure lesson we learned.
The Stack Is Deliberately Boring
Diamond Elite is a TypeScript monorepo, roughly 190,000 lines across a client and a server. The client is React 19 on Vite 7 with Tailwind and react-router. The server is Express 4 on Node, running tsx in development and compiled JavaScript in production. The AI layer is the Anthropic SDK calling Claude Sonnet. Pose estimation for the skeleton overlay is MediaPipe’s PoseLandmarker, loaded from a CDN and run entirely in the browser. Reports are generated with PDFKit.
There is no message queue, no Kubernetes, no service mesh, no GraphQL layer. For a long stretch there was no database at all — just JSON files on disk.
That last decision is the one engineers react to, so it is worth defending. The product began as a question: could a vision model usefully critique a swing at all? Nobody knew. Answering that required a video upload path, a frame capture path, a prompt, and somewhere to put the result. A schema migration framework would have contributed nothing to the answer. A swings.json file contributed everything, because it was writable in an afternoon and trivially inspectable with a text editor when something looked wrong.
The product now runs on SQLite, with a dedicated migration path from those original JSON stores. But it moved when the file-based approach started costing more than it saved — concurrent writes, query patterns that wanted indexes, per-tenant scoping that wanted a real WHERE clause. It did not move because file-based storage was embarrassing. Infrastructure you adopt before you have felt the pain of not having it is infrastructure you will maintain for reasons you cannot articulate.
Surface Area Grows Faster Than You Plan For
What did surprise us was breadth. The server now has more than eighty route modules. The client has over ninety pages. The original scope — upload a swing, get feedback — is maybe six of those files.
The rest arrived because real users have adjacent problems that become obvious only once they are using the thing. Someone analyzing swings wants to compare two side by side with synced playback. Someone comparing swings wants to see the whole season. Someone tracking a season wants scouting data on the opponent. Someone with scouting data wants a report to hand the coach. Someone with a report wants it on a schedule. Each step is a reasonable half-day. Collectively they are a platform.
The architectural consequence is that the boring stack earns its keep here too. Eighty route modules that all follow the same Express pattern, share the same helper libraries, and compile under one tsc invocation are eighty modules a small team can still reason about. Eighty microservices would not be.
The Lesson That Cost Us Two Months
Here is the failure we would most like other teams to avoid, because it is silent and it is not rare.
Production is a Docker build deployed on Railway, triggered by a push to the main branch. The client build step is tsc -b && vite build. The server step is tsc. A healthcheck endpoint confirms the container is alive. It is a clean pipeline, and for two months in 2026 it was completely broken without anyone noticing.
The cause was the client’s root tsconfig.json. It is a project-references shell — it declares an empty files array and delegates to referenced sub-projects. That is a perfectly normal TypeScript pattern. The problem is what it does to the obvious verification command. Running tsc --noEmit against that config type-checks nothing at all, and exits zero. Every time.
So the pre-push check passed. It always passed. It would have passed on an empty repository. Meanwhile eight genuine type errors accumulated — including a page importing a component that had never been written, which would have crashed at runtime — and every production build failed. Because the deploy failed rather than deploying something broken, the previous container kept right on serving traffic. The site was up. The site was simply two months stale, and the signal that should have said so was a green check mark that meant nothing.
Three things came out of that:
- The verification command must be the build command. Client type-checking is now tsc -b, matching exactly what Docker runs. A check that does not reproduce the build is not a check.
- A green result from a command you have never seen fail deserves suspicion. Introduce a deliberate error and confirm your check goes red. If it stays green, you do not have a check, you have a ritual.
- Deploy success needs positive confirmation, not the absence of alarms. The healthcheck now returns an explicit version string that gets bumped on notable releases. “Is the site up” and “is the site running the code I just wrote” are different questions, and only the second one was ever in doubt.
There is a related shell trap worth naming, because it hid the broken typecheck twice: in a pipeline like npm run typecheck | head, the exit status you read back belongs to head, not to the typecheck. Use ${PIPESTATUS[0]}, or do not pipe.
What Good Looks Like Now
Deploys land in under a minute and are verified by comparing the built asset hash against what production is actually serving — identical hash, identical code, no inference required. The type check is the build. The healthcheck names its own version. Three deploys in a single night with zero failed builds is now an ordinary evening rather than a notable one.
None of that required new infrastructure. It required removing one false signal and replacing it with a true one.
Getting Started
If you are running an internal application or an AI pilot and want to avoid the same trap, three concrete steps:
- Break your CI on purpose. Introduce a type error, a failing test, a lint violation. Confirm each check actually goes red. Any check that stays green is telling you nothing, and you are budgeting trust against it.
- Make your verification command byte-identical to your build command. Not equivalent — identical. The gap between “what CI runs” and “what the build runs” is exactly where silent failures live.
- Give every deployed artifact a version it can report. A healthcheck that returns “alive” answers the least interesting question. Have it return the release string, and check that instead.
We build these systems for ourselves before we recommend them to anyone. The next post in this series goes inside the part everybody asks about: how a video of a swing actually becomes structured coaching feedback from a vision model.
If you are standing up an AI application and want the architecture reviewed by people who have run one in production, contact us.
CB5 Solutions is a Microsoft Solutions Partner specializing in Microsoft Security, Data & AI, Modern Workplace, and Azure Infrastructure.
