- Dataset
- CodeRouterBench, 4,781 rows in total; the source probing / id_test split provides 1,425 held-out rows. Hash fraction and seed do not determine this split.
- Correctness threshold
- 0.5. The same threshold applies to every router, including the oracle.
- Confidence interval
- 2,000 paired bootstrap resamples, seed 0, 95% interval. Paired means the comparison is made on the same row, not between the means of two independent samples.
- Model pool
- All 8 models qualify; checkpoint
sft-f23aa90cd7c68f27e9b702a0, pool fingerprint 3520b40db3b55b1e. - Latency
- Qwen/Qwen3-0.6B on cpu, float32, 8 candidates, 25 iterations plus 5 warmups. No GPU.
- Cost
- Billed against actual token counts at each provider's published list prices, not estimates; the oracle and cheapest baselines use the same price table.