decision models for the computer-use loop

Pick the right action in milliseconds. Know when to abstain.

system1 is the open decision layer for computer-use agents. Your planner builds a bounded table of candidate actions from the accessibility tree; system1 picks one - fast, calibrated, locally hosted - and escalates to the planner when confidence is low.

Apache-2.0 · macOS · Windows · Linux · single GPU or a MacBook

Grounding is solved. Deciding is not.

Not another agent framework

system1 composes with any planner - GLM, Claude, Hermes, a shell script. It makes one decision type: given this state and these pre-validated candidates, which one?

Not a grounding model

UI-Venus, OS-Atlas and Qwen-UI-Agent solve coordinates. system1 solves which-action-next: selection over a candidate table with opaque ids, so it never invents params.

Calibrated or silent

Every decision carries confidence. Below the floor, system1 abstains and hands back to the planner. Fail-closed by construction - the executor refuses unattended destructive actions.

// any planner, any language:
const decision = await fetch("http://localhost:8100/v1/decide", {
  method: "POST",
  headers: { "content-type": "application/json" },
  body: JSON.stringify({
    state: { app: "Finder", context: { goal: "delete the old folder" } },
    candidate_actions: [
      { id: "a1", kind: "click", description: "Delete button" },
      { id: "a2", kind: "click", description: "Cancel" },
    ],
  }),
}).then((r) => r.json());

// -> { "action_id": "a1", "confidence": 0.75,
//      "abstained": false, "latency_ms": 0.4 }
# MCP: system1_decide / system1_verify in any MCP planner
{
  "mcpServers": {
    "system1": {
      "command": "/opt/system1/.venv/bin/system1",
      "args": ["mcp"]
    }
  }
}

# or speak the SystemOne API everyone already ships:
POST /v1/systemone/choice   # noul / choice / score

Built for the edge of the loop

The decision model runs where the mouse is: a single 4090, a MacBook’s unified memory, or plain CPU. Targets from the project plan: p50 < 50 ms on 4090, p50 < 25 ms on a high-end MacBook, sub-10 ms on the tiny tier.

Single GPU
RTX 4090 24GB
Clef-flash FP8 (~9 GiB)
p50 < 50 ms
MacBook
16-36GB unified
distilled 1-2B head
p50 < 25 ms
CPU
any machine
Laya-class tiny head
p50 < 10 ms

Quickstart

git clone https://github.com/speer-ai/system1
cd system1
pip install -e ".[dev]"

# zero-GPU baseline policy, HTTP + MCP + bench:
system1 serve --port 8100  # SystemOne-compatible API
system1 mcp                # MCP stdio server
system1 bench points.jsonl # score decision points
# score a benchmark file: one JSONL record = one decision point
{"state": {"app": "Finder", "context": {"goal": "delete folder"}},
 "candidate_actions": [
   {"id": "a1", "kind": "click", "description": "Delete button"},
   {"id": "a2", "kind": "click", "description": "Cancel"}],
 "gold_action_id": "a1"}

# -> top-1 accuracy, risk-coverage curves, ECE,
#    p50/p95 latency per hardware tier

The heuristic policy that ships in-repo is a deliberate floor: deterministic, sub-millisecond, zero dependencies - so the whole stack is testable without a GPU and every learned policy has something to beat.

The field, verified

s1-bench measures decision models specifically for the computer-use loop: top-1 accuracy over candidate tables, abstain quality, and latency on the hardware that matters. Landscape from our October 2026 survey:

ModelSizeMedian latencyMultimodalLicenseTier
Clef-flash (Cloudflare)9B38.8 msyesApache-2.04090
Kev-9B9B + LoRA51.4 msnoApache-2.04090
Clef (Cloudflare)27B209.3 msyesApache-2.04090 / DGX
Strands Decider1.9B~30 msnoopenMacBook
Laya (Convai)421M5.8 msnoApache-2.0MacBook / CPU

Nobody has benchmarked a decision model specifically for selecting from pre-validated candidate action tables derived from an accessibility tree. That is the wedge. s1-bench lands as an open HuggingFace dataset + scoring harness, grown from real recorded sessions.

Safe by default, not by promise

Fail-closed executor

The executor adapter is dry-run unless explicitly authorized, and refuses to dispatch at all without a deliberate opt-in token. Destructive actions (deletes, payments, sign-ins) escalate to a human gate instead of firing.

Escalation is a return, not a retry storm

Low confidence or failed verification returns control to the System-2 planner with a compact context bundle. The planner owns the next candidate table; system1 never invents actions.

The System-1 layer for computer use

Open, reproducible, citable. Bring your planner; keep your data local.

system1 · Apache-2.0 · built by Nick Speer & agents