Omar OS-1
AI reliability infrastructure for action control.On verified routes: higher accuracy, fewer tokens.Pareto improvement before outputs become actions.
Open OS-1 ClaudexReviewer key. Same row, same model, same settings. Temperature, token budget, retry budget, scorer, and route package stay fixed. Raw outputs, route log, results, cost summary, and manifest stay attached, so a reviewer can inspect the route instead of trusting a screenshot.
20 executable BYOK lanes are available. Separately, 19 scored evidence rows produce an unweighted +38.40% mean relative uplift, with each displayed row counted exactly once.
ABCD is both an executable matched-live lane and one stored historical evidence row; it is counted once, not twice. AgentDojo remains executable for verification but is withdrawn from the scored mean. Accuracy, safe-execution, and normalized reward remain separate metrics; no universal uplift claim.
9,043 = 6,273 ABCD turns + 2,770 disclosed evaluation rows. 2 HealthBench context rows do not disclose n. Separately, the 20-lane harness exposes up to 9,343 configured rows / turns; neither number is a count of completed BYOK executions.
Baseline → OS-1. Same reading order on every row.
Baseline → OS-1. Same reading order on every row.
Baseline → OS-1. Same reading order on every row.
Baseline → OS-1. Same reading order on every row.
Baseline → OS-1. Same reading order on every row.
Baseline → OS-1. Same reading order on every row.
Baseline → OS-1. Same reading order on every row.