Agent & CLI Reference
How an external AI agent (Claude Code, Codex, or any MCP client) drives SafeCoBrowser — the tools, the permission model, and the CLI. You act only within the mode the user grants, per tab, and they can revoke instantly.
Overview
SafeCoBrowser is a real, logged-in browser the human controls. Your agent is a guest that can see and act only inside the permission they grant — per tab, right now — and they can cut access instantly. You are the “brain”; SafeCoBrowser is the gated “body.”
- Every tab is Off by default — until a mode is granted, every call returns
permission_denied. - You cannot change the permission mode — only the user’s UI can.
- You act on the active tab; to work on another, bring it to the front with
switch_tabfirst (and you can never open or close tabs). - You never get cookies, passwords, or the profile — only brokered tools.
- Every call (allowed or denied) is written to a tamper-evident audit log.
Connect your agent (MCP)
On launch SafeCoBrowser starts a localhost server (bound to 127.0.0.1 only) and writes its endpoint + bearer token to ~/.safecobrowser/endpoint.json. The token is required even on localhost; the server also validates Host/Origin.
~/.safecobrowser/endpoint.json
{ "url": "http://127.0.0.1:8676", "token": "<bearer>" }The MCP endpoint is POST {url}/mcp (Streamable HTTP). Register it with Claude Code:
TOKEN=$(node -pe "JSON.parse(require('fs').readFileSync(require('os').homedir()+'/.safecobrowser/endpoint.json')).token")
claude mcp add --transport http safecobrowser http://127.0.0.1:8676/mcp \
--header "Authorization: Bearer $TOKEN"The server name is safecobrowser; it exposes the tools below. Calls are gated by the tab’s mode, and effectful / run_js calls require approval.
Driving it from another computer. The server is loopback-only by default. To connect a remote agent, either tunnel the port (ssh -N -L 8676:127.0.0.1:8676 you@host, or Tailscale — recommended, encrypted) and point at 127.0.0.1 as usual, or enable Settings → Agent connection → “Allow LAN connections” (default off). That rebinds the server to your LAN and shows a LAN MCP URL to use in place of 127.0.0.1. LAN mode is plaintext HTTP — use it only on a trusted network; for the internet keep it off and use the tunnel. The bearer token, the Host/Origin allowlist, and the per-tab AI block all still apply.
CLI
The bundled safecobrowser CLI is a thin client over the same broker — handy for scripting and testing.
safecobrowser status # session status (+ the active tab id)
safecobrowser tools # list tools with their mode / risk / approval
safecobrowser invoke <tool> [json] # invoke a tool through the broker
safecobrowser audit [n] # last n audit records (default 20)
safecobrowser audit verify # verify the audit log's hash chainstatus/tools/invoke need SafeCoBrowser running; audit reads the local log and works offline.invoke exits 0 on success, 2 on a denial. Note: the CLI cannot raise the mode either.
Permission modes
An ordered ladder, set per tab by the user. A tool runs only if the tab’s current mode is at least the tool’s minimum.
| Mode | Unlocks (cumulative) |
|---|---|
blocked | Nothing — AI is off (default) |
read | read_page, screenshot, locate, read_screen_text, list_recipes, get_recipe |
inspect | + inspect_element, read_console, read_network, read_network_body |
act | + navigate, click, fill, scroll_to, and the coordinate tools move_to / click_at / scroll / press_key / type_text (all behind approval) |
develop | + run_js (full page control) |
Read the active tab’s current mode with get_mode (or safecobrowser status) instead of guessing from denials — it’s read-only, you still can’t change it. A few tools sit off this ladder and work at any mode (including Off): get_mode, list_tabs, switch_tab, and submit_feedback — they aren’t page operations.
Tools
Inputs are validated by the broker; a bad shape returns invalid_input.
| Tool | Mode · risk · approval | Input → Output |
|---|---|---|
get_mode | any · low · no | none → { tab, mode } (read-only; off the ladder) |
list_tabs | any · low · no | none → { tabs[{ tab, active, mode, title, url }] } (titles/URLs privacy-filtered; off the ladder) |
switch_tab | any · low · no | { tab } → { switched, tab?, reason? } — brings a tab to the front; later calls target it. Never changes a tab’s mode. |
read_page | read · low · no | none → { url, title, text, links[] } |
screenshot | read · low · no | none → { mimeType, base64 } (PNG) |
locate | read · low · no | { text?, selector? } → { count, matches[{ matched, tag, x, y, rect, inViewport, obscured }] } — coordinates from the DOM (no screenshot/vision); feed x,y to click_at |
read_screen_text | read · low · no | none → { count, words[{ text, x, y, rect, confidence }], note? } — offline OCR of the visible page into words + their CSS-viewport coordinates; feed x,y to click_at on canvas/no-DOM pages where locate finds nothing. Not audited. |
list_recipes | read · low · no | none → { domain, recipes[{ name, description?, steps }] } |
get_recipe | read · low · no | { name } → { name, domain, description?, steps[…] } (sensitive values withheld) |
inspect_element | inspect · low · no | { selector } → { matched, tagName?, attributes?, text?, outerHTML?, rect? } |
read_console | inspect · low · no | { limit? } (≤1000) → [{ level, text, ts }] — level ∈ log|info|warning|error; filter level:"error" to check a page for JS errors / uncaught exceptions |
read_network | inspect · low · no | { limit? } (≤1000) → [{ method, url, status?, ts }] |
read_network_body | inspect · low · no | { limit? } (≤1000; last 50/page) → [{ method, url, status, contentType, body, truncated, ts }] — XHR/fetch response bodies, any origin, text/JSON only, ~64KB cap; only captures after the current grant (cleared on grant/Stop AI), never audited |
navigate | act · medium · YES | { url } → { ok, url, title?, note? } — loads a URL (http/https only); returns the COMMITTED url after redirects so you can tell when a site bounced you to a login. Without it, reaching a new page needed run_js (the Developer grant). |
click | act · medium · YES | { selector } → { clicked, matched, realInput?, note? } |
fill | act · medium · YES | { selector, value } → { filled, matched, note?, realInput? } — honest: reads the field back, returns filled:false (+note) if it didn’t land |
scroll_to | act · low · YES | { text?, selector? } → { found, matched?, x?, y?, inViewport?, obscured? } — scrolls a match into view; returns its settled centre for click_at |
move_to / click_at | act · medium · YES | { x, y[, button] } → { done, realInput, note? } — real trusted input at a viewport coordinate (computer use) |
scroll (coord) | act · medium · YES | { x, y, dy, dx? } → { done, realInput } — wheel-scroll at a point; positive dy scrolls down |
press_key | act · medium · YES | { key } → { done, realInput } — one allowlisted key (Enter/Tab/Esc/arrows/…); no modifier chords |
type_text | act · medium · YES | { text } → { done, realInput } — types at the focused field; audited as a char count, never the text |
run_js | develop · HIGH · YES | { script } (≤100k chars) → script return value |
submit_feedback | any · medium · YES | { message, email? } → { ok } — off the ladder; approval-gated; never sends page content |
fillnever logs the value (it may be a password) — only the selector (#email (value hidden)). Filling does not submit; request the submit as a separateclick.run_jsis full page control — its own grant, the script is shown in the approval card, and every script is logged. Prefer the narrower tools where they suffice.screenshotbase64 is large — save it to a file, don’t echo it.list_recipes/get_reciperead the user’s saved recipes (recorded how-to tutorials) for the current site — read one, then reproduce it withclick/fill. Scoped to the active tab’s domain; sensitive field values are withheld.- Prefer
locateover screenshots for finding things. On a real DOM,locatereturns an element’s coordinates straight from the page layout by text/selector (fast, no vision pass), withinViewport+obscuredflags. Pattern:locate→ pick a match withinViewport:true, obscured:false→click_at. If the match isinViewport:false, callscroll_tofirst (it scrolls the element in and returns its settled coordinates). It’s read-only — it never scrolls or changes the page. - Coordinate ("computer use") tools —
move_to,click_at,scroll,press_key,type_text— deliver real trusted input at a viewport point, for canvas/WebGL apps and pages with no usable DOM. Coordinates are CSS viewport pixels; the approval card shows a screenshot with a crosshair on your target. Prefer the selector tools (click/fill) orlocate+click_atwhen the DOM is usable — a coordinate is a poor approval prompt. Not a stealth feature: real input only, no humanized motion/timing. read_screen_textislocatefor pixels. When the page has no usable DOM — a canvas/WebGL app, or a page that’s a single image —locatefinds nothing.read_screen_textOCRs the visible page (offline, on-device) into words with their CSS-viewport coordinates, so you get a target forclick_at/type_textwithout a screenshot + vision pass. It’s not AI vision (you already see thescreenshot) — it’s fast, precise text→coordinate extraction. Read-only, no approval, and the recognized text is never written to the audit. On a real DOM, preferlocate.list_tabs/switch_tablet you see the open tabs and bring one to the front (your next calls then target it). Switching is not escalation — it never changes a tab’s mode, so a tab that’s Off still exposes nothing; callget_modeafter switching. Both are governed by the user’s Allow agent tab control setting (default on); if it’s off,list_tabsshows only the active tab andswitch_tabis denied.
Approvals & auto-approve
Every effectful tool requires approval — click, fill, scroll_to, the coordinate tools (move_to/click_at/scroll/press_key/type_text), and run_js. You invoke the tool, SafeCoBrowser shows the user an approval card with the concrete effect (the selector, the query for scroll_to, a screenshot + crosshair for the coordinate tools, or the full run_js script), and your call blocks until they answer.
- Approve → the tool runs, you get
{ ok: true, … }. - Reject →
approval_rejected. No answer in ~120s →approval_required. - The user may enable per-tab auto-approve (one toggle for click/fill, a separate one for run_js) — approved calls then run without a card, but are still logged. Don’t rely on it being on.
- Real input (you don’t control this): each tab has a user-set Real input toggle. When it’s on, your
click/fillare delivered as real, trusted input (the cursor actually moves and clicks; text is typed key by key) instead of synthesized events — which some sites require. You can’t read or set it and it never changes your mode; it only changes how the action is delivered. Your result carriesrealInput: true|falseso you know which path ran, and the audit log tags those actions[real input]. Ifclickreportsclicked:falsewith anotelike target obscured, an overlay was covering the point — it will not click the wrong thing; clear it and retry. It is not a stealth/anti-detection feature (no humanized motion or timing) — if a site still blocks real, trusted input, ask the user rather than trying to evade.
Results & errors
Every invocation returns an InvokeResult:
{ ok: true, output: <tool output> } // success
{ ok: false, reason: <DenyReason>, message? } // denial / failureAll refusals fail closed. The reasons:
| reason | Meaning / what to do |
|---|---|
permission_denied | Tab mode below the tool’s minimum (incl. AI off). State the mode you need; don’t retry blindly. |
invalid_input | Input failed the schema. Fix the shape and retry. |
approval_required | Needed approval, none granted (incl. ~120s timeout). |
approval_rejected | The user said no. Stop — don’t retry the same action. |
revoked | Grant pulled mid-flight (Stop AI / mode change / tab closed). Re-confirm before continuing. |
timeout | Handler exceeded its budget. Retry once. |
handler_error | The tool threw. Report the message. |
unknown_tool | No such tool — check the name against safecobrowser tools. |
Privacy filter
The user can turn on a privacy filter — their own list of sensitive strings (name, address, account numbers) that SafeCoBrowser replaces with labels like [Name] or[Account]. There is no tool for it and the agent can’t see or change it.
- What you read may be redacted.
read_page,inspect_element,screenshot, and evenrun_jsread the rendered DOM, which is where redaction happens — so you’ll see[Name]where the real value was. That’s intentional. - Don’t try to defeat it. Treat a label as a value you’re not meant to have; if you genuinely need it, ask the user.
- Best-effort, not total. It redacts visible text nodes only — not form-field values, attributes, or network metadata — so a real value may still surface; treat anything sensitive accordingly.
Audit log
Every brokered decision is appended to ~/.safecobrowser/audit-log.jsonl, hash-chained for tamper-evidence, surviving restart, and viewable live in the in-app Activity panel.
safecobrowser audit 20 # recent records
safecobrowser audit verify # re-walk the hash chain (detects edits/deletes/reorder)click/inspect_element log their selector; run_js logs the full script; fill logs the selector but never the value; switch_tab logs the target tab id.
Agent tips
- Look before you leap: use
read_page/inspect_elementto find real selectors beforeclick/fill. - Diagnose with Inspect: combine
read_console+read_network(look for non-2xx) before proposing a fix; useread_network_bodyto read the actual XHR/fetch payload (e.g. the JSON behind a feed) — it only captures responses received after AI was granted. - Fill then submit, deliberately:
fillnever submits — request the submit as its own approvedclick. - Re-read after navigation: a click that navigates makes your prior DOM knowledge stale.
- Respect denials:
approval_rejectedmeans “no” — pick a different approach.permission_deniedmeans “raise the mode,” which only the user can do. - Work across tabs deliberately:
list_tabsto find the tab →switch_tabto it →get_mode(a freshly-switched tab is often Off — ask the user to raise it) → then read/act. Don’t hammer a tab you just switched to. - Minimize
run_js: prefer the narrower tools; keep any script tight and readable (the user reads it in the card). - Follow a recipe’s intent, not just its rows: if
get_recipe’s description states a pattern (“click each row’s download button, incrementing the index until none match”), read the live page and apply it to all items — there may be more than were recorded.
# typical loop (CLI form; MCP tool calls are 1:1)
safecobrowser status
safecobrowser invoke read_page
safecobrowser invoke inspect_element '{"selector":"form button[type=\"submit\"]"}'
safecobrowser invoke fill '{"selector":"#email","value":"[email protected]"}'
safecobrowser invoke click '{"selector":"button[type=\"submit\"]"}'