Case 04 · Sovereign AI lab
Foundation fine-tune stack
An end-to-end post-training platform for a national-scale model — preference flywheels, eval CI, and cost-aware routing that delivered 43% domain lift and 6× cheaper inference versus the base fleet.
Context
A sovereign AI lab needed to move from research checkpoints to a governed production fleet serving public-sector and strategic industry workloads. They had strong researchers and weak release engineering. Fine-tunes lived on laptops. Eval was a spreadsheet. Cost was someone else’s problem until the invoice arrived.
The problem, precisely
Domain lift without catastrophic forgetting. Preference data that did not encode political or safety debt. Inference routing that could choose the right expert without a PhD on-call. And a marketing / communications layer that could explain capability honestly to ministries without overclaiming.
What we built
A post-training platform spanning data contracts, RLHF/DPO pipelines, mixture-of-experts routing, and continuous evaluation gates in CI. Every candidate model failed closed if safety, domain, or regression suites moved the wrong way. Distillation and speculative decoding cut serving cost without silently eroding quality.
Architecture highlights
- Preference data flywheel with reviewer QA and toxicity screens.
- MoE routing with cost/latency multi-objective selection.
- Eval suites as merge gates — not dashboards after the fact.
- Model cards + evidence packs for governance reviewers.
- Comms toolkit for accurate external capability claims.
The hard parts
Catastrophic forgetting showed up late: domain suites rose while general reasoning quietly fell. We added “canary general” evals to every domain train job and blocked merges that traded away core capability beyond agreed budgets.
Preference data politics were real. We separated style, safety, and task preferences into distinct datasets with different owners — preventing a single committee from encoding contested norms into every token.
Rollout
Nine months from platform charter to production fleet. Internal dogfood, restricted pilots, then broader ministry workloads. External messaging was version-locked to eval snapshots so marketing could not outrun measured capability.
Outcomes
43% lift on domain evaluation suites versus the prior base. Roughly 6× lower inference cost via routing and distillation. A release cadence measured in weeks, not quarters — with security sign-off embedded rather than appended.
What we’d repeat
Make eval the product. Tie communications to eval snapshots. And treat cost as a model quality dimension, not an ops afterthought.