Omar OS-1

AI reliability infrastructure for action control.On verified routes: higher accuracy, fewer tokens.Pareto improvement before outputs become actions.

Open OS-1 Claudex

Benchmark replay · 20 executable lanesBYOK Harness Verification

Reviewer key. Same row, same model, same settings. Temperature, token budget, retry budget, scorer, and route package stay fixed. Raw outputs, route log, results, cost summary, and manifest stay attached, so a reviewer can inspect the route instead of trusting a screenshot.

Score Summary

20 executable BYOK lanes are available. Separately, 19 scored evidence rows produce an unweighted +38.40% mean relative uplift, with each displayed row counted exactly once.

ABCD is both an executable matched-live lane and one stored historical evidence row; it is counted once, not twice. AgentDojo remains executable for verification but is withdrawn from the scored mean. Accuracy, safe-execution, and normalized reward remain separate metrics; no universal uplift claim.

Executable BYOK lanes20 Scored evidence rows19 Disclosed rows / turns9,043 Rows with undisclosed n2 Harness max configured9,343 Mean evidence-row uplift+38.40%

9,043 = 6,273 ABCD turns + 2,770 disclosed evaluation rows. 2 HealthBench context rows do not disclose n. Separately, the 20-lane harness exposes up to 9,343 configured rows / turns; neither number is a count of completed BYOK executions.

Benchmark board
Proof family Conversation control / efficiency

Baseline → OS-1. Same reading order on every row.

ABCD
−19.27%tokens 2.39→1.93Mresult
−19.27%tokens 2.39→1.93Mresult
ZebraLogicBench
+29.6 ptsgain 55.2→84.8%result
+29.6 ptsgain 55.2→84.8%result
BBEH
+35.4 ptsgain 18.2→53.6%result
+35.4 ptsgain 18.2→53.6%result
PPR/BIPIA
+17.8 ptsgain 80.2→98%result
+17.8 ptsgain 80.2→98%result
HLE / HLE-Verified
+6.0 ptsgain 6→12%result
+6.0 ptsgain 6→12%result
BBH
+10.0 ptsgain 67→77%result
+10.0 ptsgain 67→77%result
MuSR
+3.0 ptsgain 72→75%result
+3.0 ptsgain 72→75%result
Control-layer verification Prompt-injection and factuality checks sit beside reliability evidence

Baseline → OS-1. Same reading order on every row.

HaluEval
+19.0 ptsgain 81→100%result
+19.0 ptsgain 81→100%result
TruthfulQA
+15.0 ptsgain 85→100%result
+15.0 ptsgain 85→100%result
Artifact-backed board context Board-depth rows, high-value exceptions, and stored context

Baseline → OS-1. Same reading order on every row.

Proof family Reasoning / hard reasoning

Baseline → OS-1. Same reading order on every row.

HorizonMath
+10.0 ptsgain 22→32%result
+10.0 ptsgain 22→32%result
HLE / HLE-Verified
+6.0 ptsgain 6→12%result
+6.0 ptsgain 6→12%result
AIME 120
+9.2 ptsgain 50.8→60%result
+9.2 ptsgain 50.8→60%result
GPQA
+3.0 ptsgain 75→78%result
+3.0 ptsgain 75→78%result
BBH
+10.0 ptsgain 67→77%result
+10.0 ptsgain 67→77%result
MMLU-Pro
+4.0 ptsgain 74→78%result
+4.0 ptsgain 74→78%result
MuSR
+3.0 ptsgain 72→75%result
+3.0 ptsgain 72→75%result
Proof family Factuality / hallucination

Baseline → OS-1. Same reading order on every row.

Facts Grounding
+2.0 ptsgain 97→99%result
+2.0 ptsgain 97→99%result
Proof family Medical / consensus reliability

Baseline → OS-1. Same reading order on every row.

HealthBench hard
+6.0 ptsgain 50.7→56.8%result
+6.0 ptsgain 50.7→56.8%result
HealthBench main
+8.0 ptsgain 69→77%result
+8.0 ptsgain 69→77%result
HealthBench consensus
+2.8 ptsgain 88.9→91.7%result
+2.8 ptsgain 88.9→91.7%result
Proof family Robotics / embodied action reliability

Baseline → OS-1. Same reading order on every row.

RoboBench Embodied QA
+47.0 ptsgain 26→73%result
+47.0 ptsgain 26→73%result
ManiSkill Robotics
+11.4 ptsgain 85.5→96.9result
+11.4 ptsgain 85.5→96.9result
Public proof room Enter BYOK Harness Verification

Public previews and BYOK live lanes sit inside one surface. Enter once, choose the route, and inspect the evidence boundary.