One shared set (compare/github-questions/Q1…Q30.md) — title, id, and a ## Question only; no hints, no answer key. Real GitHub research from single-hit lookups → deep multi-hop reads. Same questions, same frozen refs for every tool — only the CLI differs.
Octocode is the anchor (npx octocode tools …). Baselines: gh (plain), gh+RTK (rtk gh …), gh+Headroom (gh → Headroom compressor). Each is a separate pairwise matchup.
0 Preflight verify + pin tools. 1 Answer a fresh isolated agent per (question, arm, pass), leanest legal path. 2 Judge grade blind. 3 Summarize validate + aggregate. ≥3 passes.
Wrappers log each call: total = model-in + model-out (Unicode code points, both directions; primer excluded, failed calls counted). Never self-reported — recomputed & re-hashed by sumlog.py --strict. Chars ≈ tokens (not a direct token/latency/cost measure).
Blind X/Y (randomized, seed=42), tool identity hidden. Establishes ground truth, then scores correctness 0–10, depth 1–5, workflow 1–5. Different model family (gpt-5.5) from runners. Correctness-first: leaner never beats more-correct.
Question is the unit. Headline = per-question ratio geometric mean; median = robust central; pooled total shown only with top-contributor share + leave-one-out. ≥3 passes + 95% bootstrap CI. Public suite = orientation, not a shipping gate.