The eval matrix and smarter model routing
infrastructure
What's new
A pre-deploy gate with test personas and a runbook for adding scenarios, then a thirty-six case weighted matrix scoring four aspects, run out of process against four different models.
Impact
The comparison fed directly into cost work: simple turns now route to a small model while the larger one keeps the rest.
Upgrade
No action required. Suvi deployments are managed, so this reached your instance automatically.