Omar OS-1

présence, not assistantTalk to LuaProof of reliability

Enter VR Room
Open OS1 BrowserTap the corner

Benchmark replayBYOK Harness Verification

Reviewer key. Same row, same model, same settings. Temperature, token budget, retry budget, scorer, and route package stay fixed. Raw outputs, route log, results, cost summary, and manifest stay attached, so a reviewer can inspect the route instead of trusting a screenshot. The run shows baseline, routed candidate, fallback policy, and decision label in the same proof path.

Score Summary

+58.44% mean 100-case-unit uplift across 19 proof lanes. Verified rows are counted separately from weighted case-units.

Baseline 0.5851. Omar/RCC 0.7704. Delta +18.53 points. Review lock: sample order, model/settings, temperature, token and retry budget, route mapping, scorer, raw outputs, route log, results, cost summary, and manifest stay attached. Raw verified row count is 2,820; 2 stored lanes do not expose exact row count. Weighted mean uses 31 × 100-case units (3,100). ZebraLogicBench, BBEH, and PPR/BIPIA Full500 rows each count as five 100-case units. Prompt-injection rows use blocked-attack/safe-execution score in aggregate while the row headline reports attack-success reduction. Losing lanes route back to baseline-safe behavior; no universal uplift claim.

Proof rows19 Verified rows2,820 Stored lanes2 Weighted units3,100 Mean lane uplift+58.44%

Raw verified row count is 2,820; two stored HealthBench lanes do not expose exact row count. Headline uses a 100-case-unit mean scale: ZebraLogicBench, BBEH, and PPR/BIPIA Full500 each contribute five units while other public rows contribute one unit. Baseline/Omar score means are shown only as a separate mixed-scale read.

Benchmark board 19 displayed lanes move from live BYOK proof into stored board context: baseline path, Omar/RCC path, gain, relative uplift, sample depth, and attached proof files stay visible row by row.
BOARD-DEPTH REPLAY 500 official-public heldout / deterministic CSP reasoning / REVAS floor-safe solver routing ZebraLogicBench
+148 rows / +29.6pp 276→424
Baseline accuracy: 55.2% REVAS final accuracy: 84.8%

Gold-blind final lock; overlap 0 vs smoke100/dev; accepted_C=148, accepted_B=0; deterministic CSP solver route. Official-public heldout500 proof, gold-blind final lock. This is deterministic CSP solver routing evidence, not an official leaderboard/SOTA claim.

+148 rows / +29.6pp 276→424
Run 500 official-public heldout
BOARD-DEPTH REPLAY 500 samples / gpt-4o Run revas_bbeh500_full_gold_blind_20260630 BBEH
+194.51% +35.40 pts
Baseline accuracy: 0.1820 REVAS final accuracy: 0.5360

BBEH500 REVAS gold-blind reproduction: base 91/500 (18.2%) to REVAS 268/500 (53.6%), +35.4 points / +194.5% relative; floor preserved and raw artifacts attached.

+194.51% +35.40 pts
Decision GOLD_BLIND_REVAS_ACCEPTED_B0
BOARD-DEPTH REPLAY 500 / prompt-injection security / authority boundary PPR/BIPIA
+89.9% attack defense 99→10 attacks
Baseline blocked: 80.2% REVAS blocked: 98.0%

Official BIPIA table/train fresh500: utility preserved 289/500 to 289/500; attack success 99/500 to 10/500; security_B=0; combined 690/1000 to 779/1000. Official BIPIA fresh500 security proof, gold-blind final lock. This is security uplift, not utility-accuracy uplift, and not a universal cybersecurity guarantee.

+89.9% attack defense 99→10 attacks
Run 500
BOARD-DEPTH REPLAY 100 samples / gpt-4o Run run_1779989838150_31a83b3945e2 HLE / HLE-Verified
+100.00% +6.00 pts
Baseline accuracy: 0.0600 Omar/RCC accuracy: 0.1200

Current HLE / HLE-Verified BYOK package: 100-sample package, baseline 0.06, Omar/RCC 0.12, +6.00 points / +100.00% relative. Decision LIVE_SCORE_RECORDED; raw outputs, route log, results, cost summary, and manifest are attached.

+100.00% +6.00 pts
Decision LIVE_SCORE_RECORDED
BOARD-DEPTH REPLAY 100 samples / gpt-4o Run bbh-100-byok-live-2026-05-28 BBH
+14.93% +10.00 pts
Baseline accuracy: 0.6700 Omar/RCC accuracy: 0.7700

Current BYOK live BBH n=100: baseline 0.67, Omar/RCC 0.77, +10.00 points / +14.93%. Exactness stays RECONSTRUCTED_ONLY / internal review, with raw outputs, route log, and manifest attached.

+14.93% +10.00 pts
Decision LIVE_SCORE_RECORDED
BOARD-DEPTH REPLAY 100 samples / gpt-5.2 Run musr-100-byok-live-2026-05-31 MuSR
+4.17% +3.00 pts
Baseline accuracy: 0.7200 Omar/RCC accuracy: 0.7500

Current MuSR BYOK live package: n=100, baseline 0.72, Omar/RCC 0.75, +3.00 points / +4.17% relative. Decision LIVE_SCORE_RECORDED; raw outputs, route log, results, cost summary, and manifest stay attached.

+4.17% +3.00 pts
Decision LIVE_SCORE_RECORDED
Control-layer smoke results Prompt-injection and hallucination checks now sit beside reliability evidence

Charts are normalized to higher-is-better. Prompt-injection rows use safe execution score = 100 - attack success rate; HaluEval and TruthfulQA use factuality accuracy. These are diagnostic smoke/debug rows until Ben approves public claim wording.

CONTROL-LAYER SMOKE 100 / hallucination / factuality HaluEval
+23.46% +19.0 pts
Baseline accuracy: 81.0% Omar hint v8 accuracy: 100.0%

Missed hallucinations 18/50 to 0/50; false alarms 1/50 to 0/50. 100-case BYOK factuality check. Read as control-layer smoke evidence until promoted into a board-depth claim package.

+23.46% +19.0 pts
Run 100
CONTROL-LAYER SMOKE 100 / hallucination / misleading QA TruthfulQA
+17.65% +15.0 pts
Baseline accuracy: 85.0% Omar hint v8 accuracy: 100.0%

False-choice rate 15.0% to 0.0%; invalid/refusal 0. 100-case BYOK factuality check. Read as control-layer smoke evidence until promoted into a board-depth claim package.

+17.65% +15.0 pts
Run 100
CONTROL-LAYER SMOKE 100 held-out / prompt-injection resistance AgentDojo
+4.17% +4.0 pts
Baseline safe: 96.0% Omar hint v2 safe: 100.0%

ASR 4.0% to 0.0%; utility 56.0% to 79.0%. Companion 100-case ASR also moved 6.0% to 0.0%. Positive held-out AgentDojo smoke signal on the tested subset, not universal prompt-injection protection.

+4.17% +4.0 pts
Run 100 held-out
Artifact-backed board context Board-depth rows, high-value exceptions, and stored context

100+ sample rows provide stronger early technical evidence than live smoke tests, but should still be read with sampling method, reproducibility, raw outputs, and route logs. Omar/RCC shows where structured reasoning control improves performance, where it stays neutral, and where mismatch or drift should be flagged instead of hidden.

Proof family Reasoning / hard reasoning

HorizonMath, HLE, AIME 120, GPQA, BBH, MMLU-Pro, and MuSR stay together as the reasoning family.

HIGH-VALUE EXCEPTION n=50 / research math auto-checkable numeric/constants subset HorizonMath
+45.5% +10.0 pts on 0-100 axis
Baseline accuracy: 22.0% Omar/RCC accuracy: 32.0%

High-value exception. n=50 auto-checkable research-math exception; not labeled as 100+ evidence. Score context: 22.0% baseline to 32.0% Omar/RCC (+10.0 pts); stored uplift +45.5%. recovered HorizonMath n=50 artifact.

+45.5% +10.0 pts on 0-100 axis
Context 50
BOARD-DEPTH REPLAY 100 samples / gpt-4o Run run_1779989838150_31a83b3945e2 HLE / HLE-Verified
+100.00% +6.00 pts
Baseline accuracy: 0.0600 Omar/RCC accuracy: 0.1200

Current HLE / HLE-Verified BYOK package: 100-sample package, baseline 0.06, Omar/RCC 0.12, +6.00 points / +100.00% relative. Decision LIVE_SCORE_RECORDED; raw outputs, route log, results, cost summary, and manifest are attached.

+100.00% +6.00 pts
Decision LIVE_SCORE_RECORDED
INTERNAL REVIEW 120 samples / gpt-4o Run aime-120-byok-live-2026-05-28 AIME 120
+18.03% +9.17 pts
Baseline accuracy: 0.5083 Final adopted Omar/RCC accuracy: 0.6000

Current BYOK live AIME 120 n=120: baseline 0.5083, final adopted Omar/RCC 0.6000 (+9.17 points / +18.03%). Exactness stays RECONSTRUCTED_ONLY / NEEDS_INVESTIGATION.

+18.03% +9.17 pts
Decision LIVE_SCORE_RECORDED
BOARD-DEPTH REPLAY 100 samples / gpt-4o Run gpqa-100-byok-live-2026-05-28 GPQA
+4.00% +3.00 pts
Baseline accuracy: 0.7500 Omar/RCC accuracy: 0.7800

Current GPQA BYOK reconstructed live n=100 package: baseline 0.75, Omar/RCC 0.78, +3.00 points / +4.00% relative uplift. Exactness remains RECONSTRUCTED_ONLY / NEEDS_INVESTIGATION, with raw outputs, route log, results, cost summary, and manifest attached.

+4.00% +3.00 pts
Decision LIVE_SCORE_RECORDED
BOARD-DEPTH REPLAY 100 samples / gpt-4o Run bbh-100-byok-live-2026-05-28 BBH
+14.93% +10.00 pts
Baseline accuracy: 0.6700 Omar/RCC accuracy: 0.7700

Current BYOK live BBH n=100: baseline 0.67, Omar/RCC 0.77, +10.00 points / +14.93%. Exactness stays RECONSTRUCTED_ONLY / internal review, with raw outputs, route log, and manifest attached.

+14.93% +10.00 pts
Decision LIVE_SCORE_RECORDED
BOARD-DEPTH REPLAY n=100 BYOK live / breadth-heavy expert MCQ MMLU-Pro
+5.41% +4.0 pts on 0-100 axis
Baseline accuracy: 74.0% Omar/RCC accuracy: 78.0%

Board-depth replay. Board-depth replay row; read with sampling method, raw outputs, and route logs. Score context: 74.0% baseline to 78.0% Omar/RCC (+4.0 pts); stored uplift +5.41%. MMLU-Pro n=100 BYOK live with Darwin/Hinton v2 thin blank-emission recovery: +4.0 points / +5.41% relative.

+5.41% +4.0 pts on 0-100 axis
Context 100 BYOK live
BOARD-DEPTH REPLAY 100 samples / gpt-5.2 Run musr-100-byok-live-2026-05-31 MuSR
+4.17% +3.00 pts
Baseline accuracy: 0.7200 Omar/RCC accuracy: 0.7500

Current MuSR BYOK live package: n=100, baseline 0.72, Omar/RCC 0.75, +3.00 points / +4.17% relative. Decision LIVE_SCORE_RECORDED; raw outputs, route log, results, cost summary, and manifest stay attached.

+4.17% +3.00 pts
Decision LIVE_SCORE_RECORDED
Proof family Factuality / hallucination

Facts Grounding stays beside HaluEval and TruthfulQA while factual QA stays off public proof until clean gold-blind retrieval/verifier uplift is proven.

BOARD-DEPTH REPLAY n=100 / factual grounding Facts Grounding
+2.1% +2.0 pts on 0-100 axis
Baseline accuracy: 97.0% Omar/RCC accuracy: 99.0%

Board-depth replay. Board-depth replay row; read with sampling method, raw outputs, and route logs. Score context: 97.0% baseline to 99.0% Omar/RCC (+2.0 pts); stored uplift +2.1%. recovered FACTS Grounding n=100 artifact.

+2.1% +2.0 pts on 0-100 axis
Context 100
Proof family Medical / consensus reliability

HealthBench hard, main, and consensus are one reliability family, shown together.

BOARD CONTEXT n=stored / medical reliability hard HealthBench hard
+11.9% +6.0 pts on 0-100 axis
Baseline score: 50.7% Omar/RCC score: 56.8%

Board context. Board context, separate from current BYOK proof. Score context: 50.7% baseline to 56.8% Omar/RCC (+6.0 pts); stored uplift +11.9%. HealthBench hard prior score context.

+11.9% +6.0 pts on 0-100 axis
Context stored
BOARD CONTEXT n=50 BYOK live / medical reliability HealthBench main
+13.16% +8.0 pts on 0-100 axis
Baseline score: 69.0% Omar/RCC score: 77.0%

Board context. Board context, separate from current BYOK proof. Score context: 69.0% baseline to 77.0% Omar/RCC (+8.0 pts); stored uplift +13.16%. HealthBench main BYOK reconstructed live n=100 / READY_RECONSTRUCTED / gpt-5.2.

+13.16% +8.0 pts on 0-100 axis
Context 50 BYOK live
BOARD CONTEXT n=stored / medical consensus special-type HealthBench consensus
+3.1% +2.8 pts on 0-100 axis
Baseline score: 88.9% Omar/RCC score: 91.7%

Board context. Board context, separate from current BYOK proof. Score context: 88.9% baseline to 91.7% Omar/RCC (+2.8 pts); stored uplift +3.1%. HealthBench consensus prior score context.

+3.1% +2.8 pts on 0-100 axis
Context stored
Proof family Robotics / embodied action reliability

RoboBench and ManiSkill stay together as robotics/action rows: baseline policy versus Omar/RCC routed action layer, with artifact boundaries explicit.

BOARD-DEPTH REPLAY n=100 BYOK/local v5 / embodied action-planning QA smoke RoboBench Embodied QA
+180.77% +47.0 pts on 0-100 axis
Baseline accuracy: 26.0% Omar/RCC accuracy: 73.0%

Board-depth replay. Board-depth replay row; read with sampling method, raw outputs, and route logs. Score context: 26.0% baseline to 73.0% Omar/RCC (+47.0 pts); stored uplift +180.77%. RoboBench embodied QA v5 100-sample local revalidation artifact; gpt-5.2; Darwin/Hinton v5 visual sequence + DAG exact DSL lock; not robot control and not an official leaderboard claim.

+180.77% +47.0 pts on 0-100 axis
Context 100 BYOK/local v5
BOARD-DEPTH REPLAY n=100 local cases / robotics policy smoke ManiSkill Robotics
+13.38% stored board uplift
Relative uplift scale Stored uplift: +13.38%

Board-depth replay. ManiSkill Robotics artifact proof: 100 local cases, random baseline total reward 85.4592 versus Omar/RCC sampled governance router 96.8905, +11.4313 reward / +13.38% relative. Not an official ManiSkill leaderboard claim.

+13.38% stored board uplift
Context 100 local cases
Public proof room Enter BYOK Harness Verification

Public smoke and BYOK live lanes sit inside one surface. Enter once, choose the route inside, keep the claim boundary visible.

Private review window

Rescue one failed workflow.

After the public packet, Omar can be tested against a real failure case: agent drift, brittle evaluation, tool routing, or reasoning instability. The output is a bounded review with artifacts, not a sales deck.

For teams with a real failure case, request a private review.

Email directly