Public proof surface
BYOK Verification
Reviewer key. Same sample row: baseline against Omar/RCC. Attached: route map, scorer, raw outputs, route log, manifest, and decision label.
Every score keeps baseline, Omar/RCC, route log, raw output, scorer context, and manifest attached; losing routed lanes fall back to baseline-safe behavior.
Read absolute gain, relative uplift, n, baseline, Omar/RCC score, decision label, artifact tier, route log, raw outputs, and manifest as one attached evidence row.
Reviewer key. Same row, same model, same settings. Temperature, token budget, retry budget, scorer, and route package stay fixed. Raw outputs, route log, results, cost summary, and manifest stay attached, so a reviewer can inspect the route instead of trusting a screenshot. The run shows baseline, routed candidate, fallback policy, and decision label in the same proof path.
+58.44% mean 100-case-unit uplift across 19 proof lanes. Verified rows are counted separately from weighted case-units.
Baseline 0.5851. Omar/RCC 0.7704. Delta +18.53 points. Review lock: sample order, model/settings, temperature, token and retry budget, route mapping, scorer, raw outputs, route log, results, cost summary, and manifest stay attached. Raw verified row count is 2,820; 2 stored lanes do not expose exact row count. Weighted mean uses 31 × 100-case units (3,100). ZebraLogicBench, BBEH, and PPR/BIPIA Full500 rows each count as five 100-case units. Prompt-injection rows use blocked-attack/safe-execution score in aggregate while the row headline reports attack-success reduction. Losing lanes route back to baseline-safe behavior; no universal uplift claim.
Raw verified row count is 2,820; two stored HealthBench lanes do not expose exact row count. Headline uses a 100-case-unit mean scale: ZebraLogicBench, BBEH, and PPR/BIPIA Full500 each contribute five units while other public rows contribute one unit. Baseline/Omar score means are shown only as a separate mixed-scale read.
Gold-blind final lock; overlap 0 vs smoke100/dev; accepted_C=148, accepted_B=0; deterministic CSP solver route. Official-public heldout500 proof, gold-blind final lock. This is deterministic CSP solver routing evidence, not an official leaderboard/SOTA claim.
BBEH500 REVAS gold-blind reproduction: base 91/500 (18.2%) to REVAS 268/500 (53.6%), +35.4 points / +194.5% relative; floor preserved and raw artifacts attached.
Official BIPIA table/train fresh500: utility preserved 289/500 to 289/500; attack success 99/500 to 10/500; security_B=0; combined 690/1000 to 779/1000. Official BIPIA fresh500 security proof, gold-blind final lock. This is security uplift, not utility-accuracy uplift, and not a universal cybersecurity guarantee.
Current HLE / HLE-Verified BYOK package: 100-sample package, baseline 0.06, Omar/RCC 0.12, +6.00 points / +100.00% relative. Decision LIVE_SCORE_RECORDED; raw outputs, route log, results, cost summary, and manifest are attached.
Current BYOK live BBH n=100: baseline 0.67, Omar/RCC 0.77, +10.00 points / +14.93%. Exactness stays RECONSTRUCTED_ONLY / internal review, with raw outputs, route log, and manifest attached.
Current MuSR BYOK live package: n=100, baseline 0.72, Omar/RCC 0.75, +3.00 points / +4.17% relative. Decision LIVE_SCORE_RECORDED; raw outputs, route log, results, cost summary, and manifest stay attached.
Charts are normalized to higher-is-better. Prompt-injection rows use safe execution score = 100 - attack success rate; HaluEval and TruthfulQA use factuality accuracy. These are diagnostic smoke/debug rows until Ben approves public claim wording.
Missed hallucinations 18/50 to 0/50; false alarms 1/50 to 0/50. 100-case BYOK factuality check. Read as control-layer smoke evidence until promoted into a board-depth claim package.
False-choice rate 15.0% to 0.0%; invalid/refusal 0. 100-case BYOK factuality check. Read as control-layer smoke evidence until promoted into a board-depth claim package.
ASR 4.0% to 0.0%; utility 56.0% to 79.0%. Companion 100-case ASR also moved 6.0% to 0.0%. Positive held-out AgentDojo smoke signal on the tested subset, not universal prompt-injection protection.
100+ sample rows provide stronger early technical evidence than live smoke tests, but should still be read with sampling method, reproducibility, raw outputs, and route logs. Omar/RCC shows where structured reasoning control improves performance, where it stays neutral, and where mismatch or drift should be flagged instead of hidden.
HorizonMath, HLE, AIME 120, GPQA, BBH, MMLU-Pro, and MuSR stay together as the reasoning family.
High-value exception. n=50 auto-checkable research-math exception; not labeled as 100+ evidence. Score context: 22.0% baseline to 32.0% Omar/RCC (+10.0 pts); stored uplift +45.5%. recovered HorizonMath n=50 artifact.
Current HLE / HLE-Verified BYOK package: 100-sample package, baseline 0.06, Omar/RCC 0.12, +6.00 points / +100.00% relative. Decision LIVE_SCORE_RECORDED; raw outputs, route log, results, cost summary, and manifest are attached.
Current BYOK live AIME 120 n=120: baseline 0.5083, final adopted Omar/RCC 0.6000 (+9.17 points / +18.03%). Exactness stays RECONSTRUCTED_ONLY / NEEDS_INVESTIGATION.
Current GPQA BYOK reconstructed live n=100 package: baseline 0.75, Omar/RCC 0.78, +3.00 points / +4.00% relative uplift. Exactness remains RECONSTRUCTED_ONLY / NEEDS_INVESTIGATION, with raw outputs, route log, results, cost summary, and manifest attached.
Current BYOK live BBH n=100: baseline 0.67, Omar/RCC 0.77, +10.00 points / +14.93%. Exactness stays RECONSTRUCTED_ONLY / internal review, with raw outputs, route log, and manifest attached.
Board-depth replay. Board-depth replay row; read with sampling method, raw outputs, and route logs. Score context: 74.0% baseline to 78.0% Omar/RCC (+4.0 pts); stored uplift +5.41%. MMLU-Pro n=100 BYOK live with Darwin/Hinton v2 thin blank-emission recovery: +4.0 points / +5.41% relative.
Current MuSR BYOK live package: n=100, baseline 0.72, Omar/RCC 0.75, +3.00 points / +4.17% relative. Decision LIVE_SCORE_RECORDED; raw outputs, route log, results, cost summary, and manifest stay attached.
Facts Grounding stays beside HaluEval and TruthfulQA while factual QA stays off public proof until clean gold-blind retrieval/verifier uplift is proven.
Board-depth replay. Board-depth replay row; read with sampling method, raw outputs, and route logs. Score context: 97.0% baseline to 99.0% Omar/RCC (+2.0 pts); stored uplift +2.1%. recovered FACTS Grounding n=100 artifact.
HealthBench hard, main, and consensus are one reliability family, shown together.
Board context. Board context, separate from current BYOK proof. Score context: 50.7% baseline to 56.8% Omar/RCC (+6.0 pts); stored uplift +11.9%. HealthBench hard prior score context.
Board context. Board context, separate from current BYOK proof. Score context: 69.0% baseline to 77.0% Omar/RCC (+8.0 pts); stored uplift +13.16%. HealthBench main BYOK reconstructed live n=100 / READY_RECONSTRUCTED / gpt-5.2.
Board context. Board context, separate from current BYOK proof. Score context: 88.9% baseline to 91.7% Omar/RCC (+2.8 pts); stored uplift +3.1%. HealthBench consensus prior score context.
RoboBench and ManiSkill stay together as robotics/action rows: baseline policy versus Omar/RCC routed action layer, with artifact boundaries explicit.
Board-depth replay. Board-depth replay row; read with sampling method, raw outputs, and route logs. Score context: 26.0% baseline to 73.0% Omar/RCC (+47.0 pts); stored uplift +180.77%. RoboBench embodied QA v5 100-sample local revalidation artifact; gpt-5.2; Darwin/Hinton v5 visual sequence + DAG exact DSL lock; not robot control and not an official leaderboard claim.
Board-depth replay. ManiSkill Robotics artifact proof: 100 local cases, random baseline total reward 85.4592 versus Omar/RCC sampled governance router 96.8905, +11.4313 reward / +13.38% relative. Not an official ManiSkill leaderboard claim.
Lane taxonomy
Artifact-backed board context uses archived evidence. Live Reproduction Attempt runs selected samples through the current harness. Diagnostic Live Replay is for debugging current behavior. Board Evidence Context is reference evidence only.
Direct comparability requires matching samples, model, grader, prompt wrapper, and routing harness.
Board Reproduction
Historical Benchmark Evidence is read-only. Live Replay Demo is not historical board reproduction. Board Reproduction is enabled only when exact historical assets are wired by benchmark family.
Board families: 14. Site-enabled reproduction: MuSR, SimpleQA Verified / VSF. Assets found, pending site wiring: BBEH, BBH, Facts Grounding, HealthBench main, HealthBench hard, HealthBench consensus, HLE / HLE-Verified, SimpleQA.
BYOK review families: AIME 120, BBEH, BBH, Facts Grounding, GPQA, HealthBench main, HealthBench hard, HealthBench consensus, HLE / HLE-Verified, HorizonMath, MMLU-Pro, SimpleQA.
BYOK run buttons appear only for families with an enabled board runner.
| Benchmark | n | Board uplift | Model / temperature | Required assets | Site status | BYOK review note | Source refs |
|---|---|---|---|---|---|---|---|
| AIME 120 | 120 | +100.0% | not_found_in_exact_aime_board_artifact / not_explicitly_recorded | Asset parity{
"sample_ids": "no",
"prompts": "no",
"gold": "no",
"grader": "no",
"model_config": "no",
"darwin": "yes",
"hinton": "yes",
"lua": "yes",
"rcc_wrapper": "yes",
"raw_outputs": "no"
} | BYOK review | Latest board exposes AIME summary/prior only. Exact AIME 120 raw run, sample IDs, prompts, grader, and script were not found in the 2026-04-25 proof bundle. Older AIME archive candidates exist under rcc_publication_package but are RCC v4/v16/gpt-4o-era assets, not this Ben-confirmed gpt-5.2 board. Review items: DATASET_MISSING, PROMPTS_MISSING, GOLD_MISSING, GRADER_MISSING, MODEL_CONFIG_MISSING, RAW_LOGS_MISSING | Source refs[ "AIME_120__darwin_benchmark_priors.json", "aime_hard.jsonl", "aime" ] |
| BBEH | 100 | +183.3% | gpt-5.2 / not_explicitly_recorded | Asset parity{
"sample_ids": "yes",
"prompts": "no",
"gold": "yes",
"grader": "yes",
"model_config": "no",
"darwin": "yes",
"hinton": "yes",
"lua": "yes",
"rcc_wrapper": "yes",
"raw_outputs": "no"
} | BYOK review | Script and scored result are present, but local proof artifact stores task/gold/answer correctness, not full prompt/input text or raw model outputs. Public site runner is pending site wiring to the historical script. Review items: PROMPTS_MISSING, MODEL_CONFIG_MISSING, RAW_LOGS_MISSING | Source refs[ "BBEH__bbeh_routed_bench_results.json", "BBEH__bbeh_routed_bench_results_2.json" ] |
| BBH | 50 | +14.3% | gpt-5.2 / not_explicitly_recorded | Asset parity{
"sample_ids": "yes",
"prompts": "yes",
"gold": "yes",
"grader": "yes",
"model_config": "no",
"darwin": "yes",
"hinton": "yes",
"lua": "yes",
"rcc_wrapper": "yes",
"raw_outputs": "no"
} | BYOK review | Frozen IDs/contract/script are present, but public site runner is pending site wiring to execute the historical BBH frozen harness yet. Review items: MODEL_CONFIG_MISSING, RAW_LOGS_MISSING | Source refs[ "bbh_frozen_ids.json", "bbh_frozen_contract.json", "BBH__bbh_official_current_results.json" ] |
| Facts Grounding | 100 | +2.1% | gpt-5.2 / not_explicitly_recorded | Asset parity{
"sample_ids": "yes",
"prompts": "yes",
"gold": "yes",
"grader": "yes",
"model_config": "no",
"darwin": "yes",
"hinton": "yes",
"lua": "yes",
"rcc_wrapper": "yes",
"raw_outputs": "yes"
} | BYOK review | Exact result/script assets are present, but the public site does not yet execute the historical Facts Grounding harness. Review items: MODEL_CONFIG_MISSING | Source refs[ "Facts_Grounding__deepmind_facts_grounding_bench_results.json" ] |
| GPQA | 50 | +13.2% | gpt-5.2 / not_explicitly_recorded | Asset parity{
"sample_ids": "no",
"prompts": "no",
"gold": "yes",
"grader": "yes",
"model_config": "no",
"darwin": "yes",
"hinton": "yes",
"lua": "yes",
"rcc_wrapper": "yes",
"raw_outputs": "no"
} | BYOK review | Stored GPQA board artifact is aggregate-only: scores and metrics exist, but original per-sample prompts, IDs, and raw outputs are not stored in the proof JSON. Review items: DATASET_MISSING, PROMPTS_MISSING, MODEL_CONFIG_MISSING, RAW_LOGS_MISSING | Source refs[ "GPQA__gpqa_routed_bench_results.json" ] |
| HealthBench main | 50 | +6.2% | gpt-5.2 / not_explicitly_recorded | Asset parity{
"sample_ids": "yes",
"prompts": "yes",
"gold": "yes",
"grader": "yes",
"model_config": "no",
"darwin": "yes",
"hinton": "yes",
"lua": "yes",
"rcc_wrapper": "yes",
"raw_outputs": "no"
} | BYOK review | Local HealthBench data/script/result are present, but stored proof result is aggregate metric output and public site runner is pending site wiring to historical HealthBench execution. Review items: MODEL_CONFIG_MISSING, RAW_LOGS_MISSING | Source refs[ "healthbench.jsonl", "HealthBench_main__healthbench_routing_bench_results.json" ] |
| HealthBench hard | stored | +11.9% | gpt-5.2 / not_explicitly_recorded | Asset parity{
"sample_ids": "yes",
"prompts": "yes",
"gold": "yes",
"grader": "yes",
"model_config": "no",
"darwin": "yes",
"hinton": "yes",
"lua": "yes",
"rcc_wrapper": "yes",
"raw_outputs": "no"
} | BYOK review | HealthBench hard local data exists, but latest proof bundle stores only Darwin benchmark priors for this board row, not the raw hard-subset execution result. Review items: MODEL_CONFIG_MISSING, RAW_LOGS_MISSING | Source refs[ "healthbench_hard.jsonl", "HealthBench_hard__darwin_benchmark_priors.json" ] |
| HealthBench consensus | stored | +3.1% | gpt-5.2 / not_explicitly_recorded | Asset parity{
"sample_ids": "yes",
"prompts": "yes",
"gold": "yes",
"grader": "yes",
"model_config": "no",
"darwin": "yes",
"hinton": "yes",
"lua": "yes",
"rcc_wrapper": "yes",
"raw_outputs": "no"
} | BYOK review | HealthBench consensus local data exists, but latest proof bundle stores only Darwin benchmark priors for this board row, not the raw consensus-subset execution result. Review items: MODEL_CONFIG_MISSING, RAW_LOGS_MISSING | Source refs[ "healthbench_consensus.jsonl", "HealthBench_consensus__darwin_benchmark_priors.json" ] |
| HLE / HLE-Verified | 50 | +21.1% | gpt-5.2 / not_explicitly_recorded | Asset parity{
"sample_ids": "yes",
"prompts": "yes",
"gold": "yes",
"grader": "yes",
"model_config": "no",
"darwin": "yes",
"hinton": "yes",
"lua": "yes",
"rcc_wrapper": "yes",
"raw_outputs": "yes"
} | BYOK review | Exact HLE result/script assets are present, but the public site does not yet execute the historical HLE harness. Review items: MODEL_CONFIG_MISSING | Source refs[ "HLE___HLE-Verified__hle_verified_routed_bench_results.json", "HLE___HLE-Verified__hle_verified_routed_bench_results_2.json" ] |
| HorizonMath | 50 | +45.5% | gpt-5.2 / not_explicitly_recorded | Asset parity{
"sample_ids": "yes",
"prompts": "no",
"gold": "yes",
"grader": "yes",
"model_config": "no",
"darwin": "yes",
"hinton": "yes",
"lua": "yes",
"rcc_wrapper": "yes",
"raw_outputs": "yes"
} | BYOK review | HorizonMath IDs/targets/raw outputs exist, but local result cache does not include full prompt/problem statements and public site runner is pending site wiring to the historical script. Review items: PROMPTS_MISSING, MODEL_CONFIG_MISSING | Source refs[ "HorizonMath__horizonmath_routed_bench_results.json" ] |
| MMLU-Pro | 50 | +3.4% | gpt-5.2 / not_explicitly_recorded | Asset parity{
"sample_ids": "yes",
"prompts": "no",
"gold": "yes",
"grader": "yes",
"model_config": "no",
"darwin": "yes",
"hinton": "yes",
"lua": "yes",
"rcc_wrapper": "yes",
"raw_outputs": "no"
} | BYOK review | Stored MMLU-Pro summary says n=50, but result JSON only stores partial sample rows and no full prompt text; public site runner is pending site wiring to historical MMLU-Pro script. Review items: PROMPTS_MISSING, MODEL_CONFIG_MISSING, RAW_LOGS_MISSING | Source refs[ "MMLU-Pro__mmlu_pro_routed_bench_results.json" ] |
| MuSR | 100 BYOK live | +4.17% | gpt-5.2 / not_explicitly_recorded | Asset parity{
"sample_ids": "yes",
"prompts": "yes",
"gold": "yes",
"grader": "yes",
"model_config": "yes",
"darwin": "yes",
"hinton": "yes",
"lua": "yes",
"rcc_wrapper": "yes",
"raw_outputs": "yes"
} | enabled | Current MuSR BYOK live package: n=100, baseline 0.72, Omar/RCC 0.75, +3.00 points / +4.17% relative, decision LIVE_SCORE_RECORDED. Exactness status remains READY_RECONSTRUCTED. Review items: EXACT_HISTORICAL_ARTIFACT_NOT_CLAIMED | Source refs[ "samples/processed/musr_style_external_official_100.normalized.jsonl", "results.csv", "results.json", "raw_outputs.jsonl", "route_log.jsonl", "cost_summary.txt", "repro_manifest.json" ] |
| SimpleQA | 100 | +16.0% | gpt-5.2 / not_explicitly_recorded | Asset parity{
"sample_ids": "yes",
"prompts": "yes",
"gold": "yes",
"grader": "yes",
"model_config": "no",
"darwin": "yes",
"hinton": "yes",
"lua": "yes",
"rcc_wrapper": "yes",
"raw_outputs": "yes"
} | BYOK review | Exact SimpleQA result/script assets are present, but the public site does not yet execute the historical SimpleQA harness. Review items: MODEL_CONFIG_MISSING | Source refs[ "SimpleQA__official_simpleqa_routing_bench_results.json" ] |
| SimpleQA Verified / VSF | 100 clean harness | +120.4% | gpt-5.2 / not_explicitly_recorded | Asset parity{
"sample_ids": "yes",
"prompts": "yes",
"gold": "yes",
"grader": "yes",
"model_config": "no",
"darwin": "yes",
"hinton": "yes",
"lua": "yes",
"rcc_wrapper": "yes",
"raw_outputs": "yes"
} | enabled | Historical artifact replay is wired from proof_manifests/vsf_historical_reproduction_results.json. Live replay remains separate. Review items: none | Source refs[ "SimpleQA_Verified___VSF__deepmind_simpleqa_verified_routing_bench_results.json", "SimpleQA_Verified___VSF__deepmind_simpleqa_verified_routing_bench_results_2.json" ] |
