BYOK is the proof.
Same model, same sample boundary, reviewer key. The package keeps raw outputs, route log, scorer context, results, and manifest attached.
Open EvidenceSame model, same sample boundary, reviewer key. The package keeps raw outputs, route log, scorer context, results, and manifest attached.
Open EvidencePages, emails, documents, and retrieved text stay as evidence until Omar grants command authority.
Read the BoundaryRoutes, fallbacks, and tool use wait for source checks before execution.
Request ReviewReviewer key. Same row, same model, same settings. Temperature, token budget, retry budget, scorer, and route package stay fixed. Raw outputs, route log, results, cost summary, and manifest stay attached, so a reviewer can inspect the route instead of trusting a screenshot. The run shows baseline, routed candidate, fallback policy, and decision label in the same proof path.
+58.44% mean 100-case-unit uplift across 19 proof lanes. Verified rows are counted separately from weighted case-units.
Baseline 0.5851. Omar/RCC 0.7704. Delta +18.53 points. Review lock: sample order, model/settings, temperature, token and retry budget, route mapping, scorer, raw outputs, route log, results, cost summary, and manifest stay attached. Raw verified row count is 2,820; 2 stored lanes do not expose exact row count. Weighted mean uses 31 × 100-case units (3,100). ZebraLogicBench, BBEH, and PPR/BIPIA Full500 rows each count as five 100-case units. Prompt-injection rows use blocked-attack/safe-execution score in aggregate while the row headline reports attack-success reduction. Losing lanes route back to baseline-safe behavior; no universal uplift claim.
Raw verified row count is 2,820; two stored HealthBench lanes do not expose exact row count. Headline uses a 100-case-unit mean scale: ZebraLogicBench, BBEH, and PPR/BIPIA Full500 each contribute five units while other public rows contribute one unit. Baseline/Omar score means are shown only as a separate mixed-scale read.
Gold-blind final lock; overlap 0 vs smoke100/dev; accepted_C=148, accepted_B=0; deterministic CSP solver route. Official-public heldout500 proof, gold-blind final lock. This is deterministic CSP solver routing evidence, not an official leaderboard/SOTA claim.
BBEH500 REVAS gold-blind reproduction: base 91/500 (18.2%) to REVAS 268/500 (53.6%), +35.4 points / +194.5% relative; floor preserved and raw artifacts attached.
Official BIPIA table/train fresh500: utility preserved 289/500 to 289/500; attack success 99/500 to 10/500; security_B=0; combined 690/1000 to 779/1000. Official BIPIA fresh500 security proof, gold-blind final lock. This is security uplift, not utility-accuracy uplift, and not a universal cybersecurity guarantee.
Current HLE / HLE-Verified BYOK package: 100-sample package, baseline 0.06, Omar/RCC 0.12, +6.00 points / +100.00% relative. Decision LIVE_SCORE_RECORDED; raw outputs, route log, results, cost summary, and manifest are attached.
Current BYOK live BBH n=100: baseline 0.67, Omar/RCC 0.77, +10.00 points / +14.93%. Exactness stays RECONSTRUCTED_ONLY / internal review, with raw outputs, route log, and manifest attached.
Current MuSR BYOK live package: n=100, baseline 0.72, Omar/RCC 0.75, +3.00 points / +4.17% relative. Decision LIVE_SCORE_RECORDED; raw outputs, route log, results, cost summary, and manifest stay attached.
Charts are normalized to higher-is-better. Prompt-injection rows use safe execution score = 100 - attack success rate; HaluEval and TruthfulQA use factuality accuracy. These are diagnostic smoke/debug rows until Ben approves public claim wording.
Missed hallucinations 18/50 to 0/50; false alarms 1/50 to 0/50. 100-case BYOK factuality check. Read as control-layer smoke evidence until promoted into a board-depth claim package.
False-choice rate 15.0% to 0.0%; invalid/refusal 0. 100-case BYOK factuality check. Read as control-layer smoke evidence until promoted into a board-depth claim package.
ASR 4.0% to 0.0%; utility 56.0% to 79.0%. Companion 100-case ASR also moved 6.0% to 0.0%. Positive held-out AgentDojo smoke signal on the tested subset, not universal prompt-injection protection.
100+ sample rows provide stronger early technical evidence than live smoke tests, but should still be read with sampling method, reproducibility, raw outputs, and route logs. Omar/RCC shows where structured reasoning control improves performance, where it stays neutral, and where mismatch or drift should be flagged instead of hidden.
HorizonMath, HLE, AIME 120, GPQA, BBH, MMLU-Pro, and MuSR stay together as the reasoning family.
High-value exception. n=50 auto-checkable research-math exception; not labeled as 100+ evidence. Score context: 22.0% baseline to 32.0% Omar/RCC (+10.0 pts); stored uplift +45.5%. recovered HorizonMath n=50 artifact.
Current HLE / HLE-Verified BYOK package: 100-sample package, baseline 0.06, Omar/RCC 0.12, +6.00 points / +100.00% relative. Decision LIVE_SCORE_RECORDED; raw outputs, route log, results, cost summary, and manifest are attached.
Current BYOK live AIME 120 n=120: baseline 0.5083, final adopted Omar/RCC 0.6000 (+9.17 points / +18.03%). Exactness stays RECONSTRUCTED_ONLY / NEEDS_INVESTIGATION.
Current GPQA BYOK reconstructed live n=100 package: baseline 0.75, Omar/RCC 0.78, +3.00 points / +4.00% relative uplift. Exactness remains RECONSTRUCTED_ONLY / NEEDS_INVESTIGATION, with raw outputs, route log, results, cost summary, and manifest attached.
Current BYOK live BBH n=100: baseline 0.67, Omar/RCC 0.77, +10.00 points / +14.93%. Exactness stays RECONSTRUCTED_ONLY / internal review, with raw outputs, route log, and manifest attached.
Board-depth replay. Board-depth replay row; read with sampling method, raw outputs, and route logs. Score context: 74.0% baseline to 78.0% Omar/RCC (+4.0 pts); stored uplift +5.41%. MMLU-Pro n=100 BYOK live with Darwin/Hinton v2 thin blank-emission recovery: +4.0 points / +5.41% relative.
Current MuSR BYOK live package: n=100, baseline 0.72, Omar/RCC 0.75, +3.00 points / +4.17% relative. Decision LIVE_SCORE_RECORDED; raw outputs, route log, results, cost summary, and manifest stay attached.
Facts Grounding stays beside HaluEval and TruthfulQA while factual QA stays off public proof until clean gold-blind retrieval/verifier uplift is proven.
Board-depth replay. Board-depth replay row; read with sampling method, raw outputs, and route logs. Score context: 97.0% baseline to 99.0% Omar/RCC (+2.0 pts); stored uplift +2.1%. recovered FACTS Grounding n=100 artifact.
HealthBench hard, main, and consensus are one reliability family, shown together.
Board context. Board context, separate from current BYOK proof. Score context: 50.7% baseline to 56.8% Omar/RCC (+6.0 pts); stored uplift +11.9%. HealthBench hard prior score context.
Board context. Board context, separate from current BYOK proof. Score context: 69.0% baseline to 77.0% Omar/RCC (+8.0 pts); stored uplift +13.16%. HealthBench main BYOK reconstructed live n=100 / READY_RECONSTRUCTED / gpt-5.2.
Board context. Board context, separate from current BYOK proof. Score context: 88.9% baseline to 91.7% Omar/RCC (+2.8 pts); stored uplift +3.1%. HealthBench consensus prior score context.
RoboBench and ManiSkill stay together as robotics/action rows: baseline policy versus Omar/RCC routed action layer, with artifact boundaries explicit.
Board-depth replay. Board-depth replay row; read with sampling method, raw outputs, and route logs. Score context: 26.0% baseline to 73.0% Omar/RCC (+47.0 pts); stored uplift +180.77%. RoboBench embodied QA v5 100-sample local revalidation artifact; gpt-5.2; Darwin/Hinton v5 visual sequence + DAG exact DSL lock; not robot control and not an official leaderboard claim.
Board-depth replay. ManiSkill Robotics artifact proof: 100 local cases, random baseline total reward 85.4592 versus Omar/RCC sampled governance router 96.8905, +11.4313 reward / +13.38% relative. Not an official ManiSkill leaderboard claim.
Private review window
After the public packet, Omar can be tested against a real failure case: agent drift, brittle evaluation, tool routing, or reasoning instability. The output is a bounded review with artifacts, not a sales deck.
For teams with a real failure case, request a private review.
www.omaragi.com
This website uses a security service to protect against malicious bots. This page is displayed while OmarAGI verifies you are human.