What we fix
Six ways live AI quietly degrades.
Start with two weeks, not a retainer.
- 01
Two-week audit
We build an eval set from your real cases, trace a week of traffic, and price every step of the pipeline. Fixed fee, no commitment beyond it.
deliverable: scored baseline - 02
Prioritised fixes
A ranked backlog by impact against effort, then four to six weeks executing the top of it — each change measured against the baseline before it ships.
deliverable: measured deltas - 03
Ongoing optimisation
A light monthly cycle: evals on every change, cost review, drift alerts, and a written report of what moved. Cancel any month.
deliverable: monthly report
Six levers, pulled in order of evidence.
Nothing changes without a before-and-after number against your own case set.
Quality evaluationsA graded case set built from your real traffic, run on every prompt, model or retrieval change.240-case baseline
Model routingCheap model for the easy majority, escalation to a stronger one only where the score justifies it.-38% cost
Prompt & context tuningTighter instructions, structured output, better retrieval — measured, not vibed.+6.4pt accuracy
CachingSemantic and exact-match caching for repeated questions, with sane invalidation.p95 3.4s → 1.3s
ObservabilityTraces per run: input, context, decision, tool calls, cost, outcome — searchable by case.100% traced
Compliance controlsRedaction before inference, retention policy, access boundaries, exportable audit trail.auditor-ready logs