ego (lite) is just a browser, ego is your personal agent across devices.
Join waitlist
Claude Agent SDKBrowser automationMCPAI agentsVerification

Claude Agent SDK Browser Automation: Connect a Tool and Verify the Write

Sep 30, 202612 min read
Claude Agent SDK browser automation illustration: a Claude mascot uses a retro computer displaying a check mark.

A Claude Agent SDK app needs a browser executor through a custom tool or MCP server. The SDK finishing a turn does not prove that a browser write landed. On our local request queue, an adapter bug stopped two runs before submission. After the fix, three normal runs stored a Review and three injected HTTP 503 runs stopped without a write.

ego (lite) offers a separate visible route for an agent already working in an authorized Space. We ran the same owned RQ-42 task through ego-browser: the normal run had one stored Review, while the fault run had one rejected POST and no stored event. This local workflow does not establish a built-in SDK connector, and its results are separate from the SDK trials. The ego (lite) case and original screenshots below show the executor and verification boundary.

Which browser surface does an Agent SDK app actually use?

A Claude Agent SDK application is a program that drives the Claude Code agent loop through the SDK. Anthropic's official overview describes the SDK as a library that runs the Claude Code binary. Browser access still needs an executor and an explicit tool contract. These three surfaces should not be treated as one bundled browser feature.

Claude Code Agent SDK overview page describing the SDK as a programmable library and comparing it with other Claude tools
Official Claude Code Agent SDK overview, captured September 30, 2026. The visible page describes the SDK as a library and distinguishes it from the CLI and other Claude tools; this screenshot does not show a browser executor. Scroll sideways to inspect the page, or open the original image at full resolution.
SurfaceWho runs browser actionsThis case
Agent SDK custom MCP toolYour application implements the tool handler and connects a browser runtime.Tested on a local synthetic fixture.
Messages API browser toolsetYour application executes browser_toolset_20260801 member calls and returns results.Officially documented; not tested here.
Claude Code CLI browser integrationA CLI skill, MCP server, or extension supplies the browser capability.Different workflow; not tested here.

The SDK custom-tools guide shows tool() and createSdkMcpServer() for in-process handlers. The SDK MCP guide covers external servers and allowedTools. Anthropic documents browser_toolset_20260801 separately for the Messages API. Our test did not call that API toolset.

Claude Code Agent SDK Configure your agent documentation showing session options and configuration sources
Official Agent SDK configuration page, captured September 30, 2026. It explains that SDK sessions take options and may read settings and environment variables. A browser still needs an explicitly connected executor. Scroll sideways to inspect the page, or open the original image at full resolution.

How did the SDK browser tool run our local task?

The September 28, 2026 case used @anthropic-ai/claude-agent-sdk 0.3.169, Claude Code CLI 2.1.276, resolved model claude-sonnet-5 with low effort, ego-browser 0.5.2.16, and a local HTTP fixture. Built-in SDK tools were disabled and only our fixture_browser MCP tool was allowed. The Agent could see the queue and browser snapshots, but not the verifier's /api/state endpoint.

The queue deliberately placed RQ-24 beside RQ-42. The task required opening RQ-42, confirming amount 42 and pending status, clicking Review once, refreshing, and returning JSON. The normal server committed receipt RV-RQ-42-0001. In fault mode it returned HTTP 503 without a write after briefly showing a 'Review sent' cue.

Corrected N01 local synthetic queue decoy RQ-24 with amount 24 and pending status
Corrected N01, local synthetic queue. RQ-24 was the decoy; this crop is from the original browser-01.png, not a third-party page.

Open the full original screenshot

Corrected N01 local synthetic queue target RQ-42 with amount 42 and pending status
Corrected N01, local synthetic queue. The same original screenshot shows the required RQ-42 target beside RQ-24.

Open the full original screenshot

Corrected N01 local synthetic RQ-42 detail before Review, showing amount 42 and pending status
Corrected N01, local synthetic fixture before the write: the target detail still says amount 42 and pending. The source is browser-02.png.

Open the full original screenshot

The tested handler registered one browse tool with query(), then captured a browser snapshot and screenshot after each action. This excerpt shows the SDK contract; the full frozen runner, fixture, lockfile, prompts, and eight raw trials are in the audit download below.

const browserTool = tool("browse", description, schema, async (args) => {
  try {
    const result = await browserAction(runDir, counter, args);
    return { content: [{ type: "text", text: JSON.stringify(result) }] };
  } catch (error) {
    return { content: [{ type: "text", text: String(error) }], isError: true };
  }
}, { alwaysLoad: true });
const server = createSdkMcpServer({
  name: "fixture_browser", version: "1.0.0", tools: [browserTool],
  alwaysLoad: true,
});
for await (const event of query({
  prompt,
  options: {
    tools: [],
    mcpServers: { fixture_browser: server },
    allowedTools: ["mcp__fixture_browser__browse"],
    model: "sonnet", effort: "low", maxTurns: 10,
  },
})) record(event);

This excerpt shows the actual MCP ToolResult wrapper and built-in-tool setting. It omits the browserAction implementation and local setup, so it is not a standalone script. The official quickstart documents installation and supported API-key setup. Its authentication rules matter before adapting this local case to a product.

Claude Code Agent SDK Quickstart documentation with prerequisites and setup steps
Official Agent SDK Quickstart, captured September 30, 2026. It lists the setup path for a new SDK project; the screenshot does not demonstrate that this article's local RQ-42 task ran. Scroll sideways to inspect the page, or open the original image at full resolution.

What happened across all eight formal trials?

We froze the task and five checks before each run block. The first block used an overly strict selector adapter and failed twice. After correcting only that adapter, we ran one normal and one 503 trial, then pre-registered four more in alternating order F02, N02, N03, F03. Every row below includes failures. 'Success' in the SDK stream means the conversation ended normally; for a 503 run, it does not mean the Review was saved.

Formal trialModeSDK endSDK cost (USD)Review POSTStored eventsAfter refreshIndependent verdict
Initial N01NormalMax turns$0.0563837None0Not reachedHarness failed before submit
Initial F01HTTP 503Max turns$0.0578691None0Not reachedFault was never exercised
Corrected N01NormalSuccess$0.0545531200 × 11Reviewed + receiptBusiness task completed
Corrected F01HTTP 503Success$0.0465109503 × 10Pending, no receiptSafe stop; task incomplete
Repeat F02HTTP 503Success$0.0476405503 × 10Pending, no receiptSafe stop; task incomplete
Repeat N02NormalSuccess$0.0386738200 × 11Reviewed + receiptBusiness task completed
Repeat N03NormalSuccess$0.0217391200 × 11Reviewed + receiptBusiness task completed
Repeat F03HTTP 503Success$0.0387245503 × 10Pending, no receiptSafe stop; task incomplete

The SDK reported the following per-trial usage. Wall time includes model calls and local fixture/browser work, so these values are not a route-speed comparison. Browser actions count completed captures; selector errors count rejected tool calls. Every formal trial had zero human interventions and zero repeated Review POSTs. Token order is uncached input / cache creation / cache read / output, as reported by the SDK; exact millisecond times are in the audit CSV. Original model response bytes are not observable through this SDK stream. We do not substitute redacted archive file size for that measure.

Formal trialWall time (s)Browser actionsSelector errorsSDK tokens: input / cache write / cache read / output
Initial N0135.3803720 / 10,707 / 81,791 / 1,035
Initial F0137.0443720 / 10,741 / 81,793 / 1,173
Corrected N0126.4184112 / 15,387 / 34,468 / 630
Corrected F0124.1504010 / 13,057 / 27,902 / 540
Repeat F0228.8685114 / 10,507 / 52,685 / 794
Repeat N0218.2584010 / 9,546 / 33,254 / 527
Repeat N0319.6224010 / 2,017 / 40,793 / 565
Repeat F0320.1214010 / 9,585 / 33,320 / 520

Within the corrected local condition, all three normal trials had one HTTP 200 POST, one matching stored event, a reviewed status and receipt after refresh, and a truthful six-field JSON result. All three controlled 503 trials had one failed POST, zero stored events, a visible error, pending status after refresh, no retry, and a truthful JSON failure. This is a six-trial fixture observation, not a general success percentage. A separate reviewer recomputed the raw requests, server states, SDK results, and sampled original images for this limited claim.

Corrected N01 local synthetic fixture after refresh: RQ-42 reviewed with receipt RV-RQ-42-0001
Corrected N01, local synthetic fixture after refresh. The visible receipt matches the single independent server event in that run. N02 and N03 have separate raw screenshots in the audit pack.

Open the full original screenshot

Why did the HTTP 503 case need a second signal?

The fault fixture briefly displayed 'Review sent. Waiting for receipt...' before its POST returned HTTP 503. Treating that transient cue as completion would produce a false success. Our rubric required one POST record, zero stored events, a visible error, and pending state after refresh. The Agent's final sentence alone could not pass the check.

Corrected F01 local synthetic fixture after one failed Review POST, showing HTTP 503 and no receipt
Corrected F01, local synthetic fixture after one controlled POST 503. The browser says no receipt; the server event log contains zero writes. This is not a public-site outage.

Open the full original screenshot

Corrected F01 local synthetic fixture after refresh, still showing RQ-42 pending with no receipt
Corrected F01, local synthetic fixture after refresh remains pending. F02 and F03 also stopped after one POST 503 and retained separate raw captures.

Open the full original screenshot

The safe-stop result is narrower than business success: the requested Review was not completed in any of the three fault trials. For a real irreversible action, the application should define its own authoritative state check and retry policy before allowing another click.

What failed before the corrected runs?

The first adapter accepted only a specific role-selector string, although its own snapshot showed ref selectors and visible link text. Initial N01 and F01 each made ten SDK tool calls, received seven selector rejections, never opened the RQ-42 detail, sent no Review POST, and ended at the turn cap. The HTTP 503 branch was never exercised in that block. This was our tool contract defect, not evidence that SDK authentication failed.

The correction did not make every input valid. Corrected N01 tried bare numeric selector 2015 and recovered by using the visible link text. Repeat F02 tried bare 2026, received an error, made one navigation call whose supplied detail URL was ignored by the fixed queue navigator, then recovered with a ref selector. Those errors remain in the raw streams. They explain why an adapter must publish the selector formats it accepts and why the browser snapshot, tool errors, and server state all belong in the audit.

How can you audit or reproduce this case?

Download the local audit and reproduction pack. It includes the three pre-registered plans, frozen runner sources, fixture, package lock, all eight original screenshot sets, redacted SDK streams, HTTP request/event files, independent server states, a source-to-public SHA map, and a checker that rebuilds the eight-row CSV. Eight SDK streams have stable path and identifier pseudonyms; the internal originals were retained for independent review. The portable runner in the ZIP is a documented derivative and did not produce the archived results; it writes a new replay-output directory without overwriting them.

npm ci --legacy-peer-deps
python3 score-formal.py
# For a new run: create your own ego-browser TaskSpace, set
# EGO_TASK_SPACE_ID, and use supported SDK API-key authentication.
node portable-run-v2.mjs pilot normal
node portable-run-v2.mjs formal

The archive's README explains the required ego-browser runtime and environment variables. A new replay is a new sample. The original research used a signed-in local Claude Code CLI; Anthropic's SDK overview says third-party products need an approved authentication method and should follow API-key guidance. The case does not test production auth, real accounts, CAPTCHA, protected sites, comparative speed, or cost per accepted task.

For a new SDK application, the official Examples page points to runnable projects and guided recipes. Those examples are separate from the frozen RQ-42 audit above.

Claude Code Agent SDK Examples documentation listing minimal agents, demo applications and guided recipes
Official Agent SDK Examples page, captured September 30, 2026. It routes readers to sample SDK projects and recipes; it is a documentation reference, not evidence for the eight local trials. Scroll sideways to inspect the page, or open the original image at full resolution.

Where does ego (lite) fit?

Use ego (lite) when an agent needs to work in a visible, authorized local browser instead of an executor built into your SDK app. The agent acts inside a Space; your application still has to check that any important write persisted. The existing-browser connection guide explains the session boundary. The ego (lite) Quick start describes using its browser skill and a Space. Neither document establishes a public ego (lite) connector for Claude Agent SDK.

We also performed two separate direct ego-browser runs on the same self-built RQ-42 queue on September 28, 2026. Codex selected RQ-42 rather than the RQ-24 decoy from a browser snapshot and clicked Review once. In normal mode, a reload showed reviewed and receipt RV-RQ-42-0001; the independent server read found exactly one matching event. In the controlled HTTP 503 mode, the immediate page showed no receipt, the reload still showed pending, and the server recorded zero events after one rejected POST. These two local observations are excluded from the eight Agent SDK trials and do not compare product reliability, speed, or cost.

The four images below come from that first normal and fault pair. Its screenshots, page snapshots, HTTP logs, and server states were retained, but its exact browser command receipts were not. We therefore froze a second script and repeated both conditions as AR-N1 and AR-F1 with a complete raw action transcript. The audited replay again produced one committed Review in normal mode and zero writes after one HTTP 503. The evidence table marks the first pair as partial and the replay pair as complete; these are four local observations, not four SDK trials or a product success rate.

Direct ego-browser run on synthetic RQ-42 detail before Review, showing amount 42 and pending status
Direct ego-browser 0.5.2.16 normal run, September 28, 2026. This is the same synthetic RQ-42 task as the SDK case, but a separate browser workflow. The page still says pending before the single Review click. Scroll sideways on a narrow screen to inspect the native-size screenshot.

Open the full original screenshot

Direct ego-browser run after reload showing RQ-42 reviewed and receipt RV-RQ-42-0001
Direct ego-browser normal run after a fresh page load. The receipt matches one RQ-42 Review event in the independent server state; the screenshot alone does not prove that server event. Scroll sideways on a narrow screen to inspect the native-size screenshot.

Open the full original screenshot

Direct ego-browser controlled fault run showing HTTP 503 and no receipt after one Review click
Separate controlled HTTP 503 run, excluded from SDK scoring. The browser shows a failed Review after one click. The server request log records one POST 503 and no committed event. Scroll sideways on a narrow screen to inspect the native-size screenshot.

Open the full original screenshot

Direct ego-browser controlled fault run after reload still showing RQ-42 pending without a receipt
The same fault run after reload. RQ-42 remains pending; an independent server read found zero Review events. The browser stopped without a retry. Scroll sideways on a narrow screen to inspect the native-size screenshot.

Open the full original screenshot

Download the direct ego-browser evidence pack for both frozen plans, the replay script and raw browser transcript, per-run results, sample validation, fixture source, original screenshots, HTTP logs, server states, and crop hashes. The workflow to carry into an authorized site is to identify the correct record, act once, reload, and compare the visible result with an independent application record. This case did not use a real login, test account migration or a human handoff, or show a public Agent SDK integration.

For deterministic CI assertions, use a browser testing framework. For Anthropic's browser_toolset_20260801, follow its separate Messages API executor contract. For an SDK application that must own custom business actions, keep the handler narrow and verify writes against an independent system of record.

FAQ

These answers separate the tested SDK custom-tool route from nearby browser products and API surfaces.

Does the Agent SDK automatically include the API browser toolset?

No such equivalence was tested here. Anthropic documents browser_toolset_20260801 as a Messages API client toolset with application-run member calls. Our Agent SDK run registered its own in-process MCP tool.

Can a third-party SDK product use a personal Claude login?

Anthropic's SDK overview says a third-party product may not offer claude.ai login or subscription rate limits without prior approval. Follow the official Console API-key or supported provider setup. Our signed-in CLI was an internal research condition only.

Why did the first two browser runs fail?

Our custom adapter rejected selector inputs that the model drew from its own browser snapshot. Both runs exhausted ten turns without a POST. Changing the adapter input contract, not the task or scoring, allowed the later local runs to reach the write.

Did the Agent complete the task when the SDK returned success after HTTP 503?

No. The SDK stream ended normally, but the Review business task was incomplete. The server recorded zero events and the refreshed page remained pending. Its safe-stop behavior passed the fault rubric because it reported the failure and did not retry.

Did this case test a real website or existing login state?

No. It used a local synthetic queue without a real account. Browser login inheritance, session persistence, and real-site reliability need separate evidence.

What should I inspect before trusting a browser write?

Inspect the intended target before the click, the browser-visible state after refresh, and an independent server or business record. Record the POST status and retry count. A transient toast or the Agent's final answer is not enough for a consequential write.