SafeCoBrowser

Agent & CLI Reference

How an external AI agent (Claude Code, Codex, or any MCP client) drives SafeCoBrowser — the tools, the permission model, and the CLI. You act only within the mode the user grants, per tab, and they can revoke instantly.

Overview

SafeCoBrowser is a real, logged-in browser the human controls. Your agent is a guest that can see and act only inside the permission they grant — per tab, right now — and they can cut access instantly. You are the “brain”; SafeCoBrowser is the gated “body.”

  • Every tab is Off by default — until a mode is granted, every call returns permission_denied.
  • You cannot change the permission mode — only the user’s UI can.
  • You act on the active tab; to work on another, bring it to the front with switch_tab first (and you can never open or close tabs).
  • You never get cookies, passwords, or the profile — only brokered tools.
  • Every call (allowed or denied) is written to a tamper-evident audit log.

Connect your agent (MCP)

On launch SafeCoBrowser starts a localhost server (bound to 127.0.0.1 only) and writes its endpoint + bearer token to ~/.safecobrowser/endpoint.json. The token is required even on localhost; the server also validates Host/Origin.

~/.safecobrowser/endpoint.json
{ "url": "http://127.0.0.1:8676", "token": "<bearer>" }

The MCP endpoint is POST {url}/mcp (Streamable HTTP). Register it with Claude Code:

TOKEN=$(node -pe "JSON.parse(require('fs').readFileSync(require('os').homedir()+'/.safecobrowser/endpoint.json')).token")
claude mcp add --transport http safecobrowser http://127.0.0.1:8676/mcp \
  --header "Authorization: Bearer $TOKEN"

The server name is safecobrowser; it exposes the tools below. Calls are gated by the tab’s mode, and effectful / run_js calls require approval.

Driving it from another computer. The server is loopback-only by default. To connect a remote agent, either tunnel the port (ssh -N -L 8676:127.0.0.1:8676 you@host, or Tailscale — recommended, encrypted) and point at 127.0.0.1 as usual, or enable Settings → Agent connection → “Allow LAN connections” (default off). That rebinds the server to your LAN and shows a LAN MCP URL to use in place of 127.0.0.1. LAN mode is plaintext HTTP — use it only on a trusted network; for the internet keep it off and use the tunnel. The bearer token, the Host/Origin allowlist, and the per-tab AI block all still apply.

CLI

The bundled safecobrowser CLI is a thin client over the same broker — handy for scripting and testing.

safecobrowser status                 # session status (+ the active tab id)
safecobrowser tools                  # list tools with their mode / risk / approval
safecobrowser invoke <tool> [json]   # invoke a tool through the broker
safecobrowser audit [n]              # last n audit records (default 20)
safecobrowser audit verify           # verify the audit log's hash chain

status/tools/invoke need SafeCoBrowser running; audit reads the local log and works offline.invoke exits 0 on success, 2 on a denial. Note: the CLI cannot raise the mode either.

Permission modes

An ordered ladder, set per tab by the user. A tool runs only if the tab’s current mode is at least the tool’s minimum.

ModeUnlocks (cumulative)
blockedNothing — AI is off (default)
readread_page, screenshot, locate, read_screen_text, list_recipes, get_recipe
inspect+ inspect_element, read_console, read_network, read_network_body
act+ navigate, click, fill, scroll_to, and the coordinate tools move_to / click_at / scroll / press_key / type_text (all behind approval)
develop+ run_js (full page control)

Read the active tab’s current mode with get_mode (or safecobrowser status) instead of guessing from denials — it’s read-only, you still can’t change it. A few tools sit off this ladder and work at any mode (including Off): get_mode, list_tabs, switch_tab, and submit_feedback — they aren’t page operations.

Tools

Inputs are validated by the broker; a bad shape returns invalid_input.

ToolMode · risk · approvalInput → Output
get_modeany · low · nonone → { tab, mode } (read-only; off the ladder)
list_tabsany · low · nonone → { tabs[{ tab, active, mode, title, url }] } (titles/URLs privacy-filtered; off the ladder)
switch_tabany · low · no{ tab } → { switched, tab?, reason? } — brings a tab to the front; later calls target it. Never changes a tab’s mode.
read_pageread · low · nonone → { url, title, text, links[] }
screenshotread · low · nonone → { mimeType, base64 } (PNG)
locateread · low · no{ text?, selector? } → { count, matches[{ matched, tag, x, y, rect, inViewport, obscured }] } — coordinates from the DOM (no screenshot/vision); feed x,y to click_at
read_screen_textread · low · nonone → { count, words[{ text, x, y, rect, confidence }], note? } — offline OCR of the visible page into words + their CSS-viewport coordinates; feed x,y to click_at on canvas/no-DOM pages where locate finds nothing. Not audited.
list_recipesread · low · nonone → { domain, recipes[{ name, description?, steps }] }
get_reciperead · low · no{ name } → { name, domain, description?, steps[…] } (sensitive values withheld)
inspect_elementinspect · low · no{ selector } → { matched, tagName?, attributes?, text?, outerHTML?, rect? }
read_consoleinspect · low · no{ limit? } (≤1000) → [{ level, text, ts }] — level ∈ log|info|warning|error; filter level:"error" to check a page for JS errors / uncaught exceptions
read_networkinspect · low · no{ limit? } (≤1000) → [{ method, url, status?, ts }]
read_network_bodyinspect · low · no{ limit? } (≤1000; last 50/page) → [{ method, url, status, contentType, body, truncated, ts }] — XHR/fetch response bodies, any origin, text/JSON only, ~64KB cap; only captures after the current grant (cleared on grant/Stop AI), never audited
navigateact · medium · YES{ url } → { ok, url, title?, note? } — loads a URL (http/https only); returns the COMMITTED url after redirects so you can tell when a site bounced you to a login. Without it, reaching a new page needed run_js (the Developer grant).
clickact · medium · YES{ selector } → { clicked, matched, realInput?, note? }
fillact · medium · YES{ selector, value } → { filled, matched, note?, realInput? } — honest: reads the field back, returns filled:false (+note) if it didn’t land
scroll_toact · low · YES{ text?, selector? } → { found, matched?, x?, y?, inViewport?, obscured? } — scrolls a match into view; returns its settled centre for click_at
move_to / click_atact · medium · YES{ x, y[, button] } → { done, realInput, note? } — real trusted input at a viewport coordinate (computer use)
scroll (coord)act · medium · YES{ x, y, dy, dx? } → { done, realInput } — wheel-scroll at a point; positive dy scrolls down
press_keyact · medium · YES{ key } → { done, realInput } — one allowlisted key (Enter/Tab/Esc/arrows/…); no modifier chords
type_textact · medium · YES{ text } → { done, realInput } — types at the focused field; audited as a char count, never the text
run_jsdevelop · HIGH · YES{ script } (≤100k chars) → script return value
submit_feedbackany · medium · YES{ message, email? } → { ok } — off the ladder; approval-gated; never sends page content
  • fill never logs the value (it may be a password) — only the selector (#email (value hidden)). Filling does not submit; request the submit as a separate click.
  • run_js is full page control — its own grant, the script is shown in the approval card, and every script is logged. Prefer the narrower tools where they suffice.
  • screenshot base64 is large — save it to a file, don’t echo it.
  • list_recipes / get_recipe read the user’s saved recipes (recorded how-to tutorials) for the current site — read one, then reproduce it with click/fill. Scoped to the active tab’s domain; sensitive field values are withheld.
  • Prefer locate over screenshots for finding things. On a real DOM, locate returns an element’s coordinates straight from the page layout by text/selector (fast, no vision pass), with inViewport + obscured flags. Pattern: locate → pick a match with inViewport:true, obscured:false → click_at. If the match is inViewport:false, call scroll_to first (it scrolls the element in and returns its settled coordinates). It’s read-only — it never scrolls or changes the page.
  • Coordinate ("computer use") tools — move_to, click_at, scroll, press_key, type_text — deliver real trusted input at a viewport point, for canvas/WebGL apps and pages with no usable DOM. Coordinates are CSS viewport pixels; the approval card shows a screenshot with a crosshair on your target. Prefer the selector tools (click/fill) or locate+click_at when the DOM is usable — a coordinate is a poor approval prompt. Not a stealth feature: real input only, no humanized motion/timing.
  • read_screen_text is locate for pixels. When the page has no usable DOM — a canvas/WebGL app, or a page that’s a single image — locate finds nothing. read_screen_text OCRs the visible page (offline, on-device) into words with their CSS-viewport coordinates, so you get a target for click_at/type_text without a screenshot + vision pass. It’s not AI vision (you already see the screenshot) — it’s fast, precise text→coordinate extraction. Read-only, no approval, and the recognized text is never written to the audit. On a real DOM, prefer locate.
  • list_tabs / switch_tab let you see the open tabs and bring one to the front (your next calls then target it). Switching is not escalation — it never changes a tab’s mode, so a tab that’s Off still exposes nothing; call get_mode after switching. Both are governed by the user’s Allow agent tab control setting (default on); if it’s off, list_tabs shows only the active tab and switch_tab is denied.

Approvals & auto-approve

Every effectful tool requires approval — click, fill, scroll_to, the coordinate tools (move_to/click_at/scroll/press_key/type_text), and run_js. You invoke the tool, SafeCoBrowser shows the user an approval card with the concrete effect (the selector, the query for scroll_to, a screenshot + crosshair for the coordinate tools, or the full run_js script), and your call blocks until they answer.

  • Approve → the tool runs, you get { ok: true, … }.
  • Reject → approval_rejected. No answer in ~120s → approval_required.
  • The user may enable per-tab auto-approve (one toggle for click/fill, a separate one for run_js) — approved calls then run without a card, but are still logged. Don’t rely on it being on.
  • Real input (you don’t control this): each tab has a user-set Real input toggle. When it’s on, your click/fill are delivered as real, trusted input (the cursor actually moves and clicks; text is typed key by key) instead of synthesized events — which some sites require. You can’t read or set it and it never changes your mode; it only changes how the action is delivered. Your result carries realInput: true|false so you know which path ran, and the audit log tags those actions [real input]. If click reports clicked:false with a note like target obscured, an overlay was covering the point — it will not click the wrong thing; clear it and retry. It is not a stealth/anti-detection feature (no humanized motion or timing) — if a site still blocks real, trusted input, ask the user rather than trying to evade.

Results & errors

Every invocation returns an InvokeResult:

{ ok: true,  output: <tool output> }            // success
{ ok: false, reason: <DenyReason>, message? }   // denial / failure

All refusals fail closed. The reasons:

reasonMeaning / what to do
permission_deniedTab mode below the tool’s minimum (incl. AI off). State the mode you need; don’t retry blindly.
invalid_inputInput failed the schema. Fix the shape and retry.
approval_requiredNeeded approval, none granted (incl. ~120s timeout).
approval_rejectedThe user said no. Stop — don’t retry the same action.
revokedGrant pulled mid-flight (Stop AI / mode change / tab closed). Re-confirm before continuing.
timeoutHandler exceeded its budget. Retry once.
handler_errorThe tool threw. Report the message.
unknown_toolNo such tool — check the name against safecobrowser tools.

Privacy filter

The user can turn on a privacy filter — their own list of sensitive strings (name, address, account numbers) that SafeCoBrowser replaces with labels like [Name] or[Account]. There is no tool for it and the agent can’t see or change it.

  • What you read may be redacted. read_page, inspect_element, screenshot, and even run_js read the rendered DOM, which is where redaction happens — so you’ll see [Name] where the real value was. That’s intentional.
  • Don’t try to defeat it. Treat a label as a value you’re not meant to have; if you genuinely need it, ask the user.
  • Best-effort, not total. It redacts visible text nodes only — not form-field values, attributes, or network metadata — so a real value may still surface; treat anything sensitive accordingly.

Audit log

Every brokered decision is appended to ~/.safecobrowser/audit-log.jsonl, hash-chained for tamper-evidence, surviving restart, and viewable live in the in-app Activity panel.

safecobrowser audit 20        # recent records
safecobrowser audit verify    # re-walk the hash chain (detects edits/deletes/reorder)

click/inspect_element log their selector; run_js logs the full script; fill logs the selector but never the value; switch_tab logs the target tab id.

Agent tips

  • Look before you leap: use read_page / inspect_element to find real selectors before click/fill.
  • Diagnose with Inspect: combine read_console + read_network (look for non-2xx) before proposing a fix; use read_network_body to read the actual XHR/fetch payload (e.g. the JSON behind a feed) — it only captures responses received after AI was granted.
  • Fill then submit, deliberately: fill never submits — request the submit as its own approved click.
  • Re-read after navigation: a click that navigates makes your prior DOM knowledge stale.
  • Respect denials: approval_rejected means “no” — pick a different approach. permission_denied means “raise the mode,” which only the user can do.
  • Work across tabs deliberately: list_tabs to find the tab → switch_tab to it → get_mode (a freshly-switched tab is often Off — ask the user to raise it) → then read/act. Don’t hammer a tab you just switched to.
  • Minimize run_js: prefer the narrower tools; keep any script tight and readable (the user reads it in the card).
  • Follow a recipe’s intent, not just its rows: if get_recipe’s description states a pattern (“click each row’s download button, incrementing the index until none match”), read the live page and apply it to all items — there may be more than were recorded.
# typical loop (CLI form; MCP tool calls are 1:1)
safecobrowser status
safecobrowser invoke read_page
safecobrowser invoke inspect_element '{"selector":"form button[type=\"submit\"]"}'
safecobrowser invoke fill  '{"selector":"#email","value":"[email protected]"}'
safecobrowser invoke click '{"selector":"button[type=\"submit\"]"}'