After Codex finishes a task boundary, independent verification checks the original request against the diff, tests, run results, and any needed semantic judgment. The executing model's claim of success is not evidence. Verification uses a fixed tier outside the economic routing loop.
Did using less frontier capacity still complete the task correctly, and did the full delivery cost — including Jev, corrections, and cache effects — actually go down?
Failure returns to the same session
- Specific failure facts return to the same Codex session for correction.
- Default is at most two correction cycles, then Root takes over.
- Repeated defects, permission issues, or scope blowups can escalate earlier.
What Router Compass records
Router Compass links each call's Jev Choice to the actual model and effort, usage, cache behavior, latency, and fallback reason, then associates those facts with the task's verification result.
- Unknown usage stays
UNKNOWN, never zero. - Model switches can lose prompt-cache reuse, so cost must be measured.
- Historical replay can estimate counterfactual prices; it does not prove quality.
Evidence levels
| Evidence | What it can establish |
|---|---|
| Production observation | Actual model mix, usage, and task verification |
| Historical replay | Estimated prices under explicit assumptions — a counterfactual, not quality proof |
| Fixed-Terra control | Whether equivalent acceptance and complete accounting show real savings with maintained quality |
P0 gate before any savings claim
The project does not claim general savings or automatic routing after installation until a real Codex CLI A→B→A switch within one tool loop passes across the four tiers, covering authentication, actual model and effort, tool-call IDs, streaming events, cancellation, continuation, and compaction — plus a controlled comparison.