From 922519aa01368368eb60881cc1a4ba07ee3e244a Mon Sep 17 00:00:00 2001 From: asepharyana Date: Fri, 4 Sep 2026 19:22:41 +0700 Subject: [PATCH] feat: add web_browse tool for navigating and interacting with web pages using Bun's WebView --- docs/webview-spec.md | 67 ++++++++++++++++++ docs/webview.md | 83 ++++++++++++++++++++++ src/tools-webview.ts | 155 +++++++++++++++++++++++++++++++++++++++++ src/tools.ts | 5 +- test/tool-meta.test.ts | 2 + test/webview.test.ts | 104 +++++++++++++++++++++++++++ 6 files changed, 415 insertions(+), 1 deletion(-) create mode 100644 docs/webview-spec.md create mode 100644 docs/webview.md create mode 100644 src/tools-webview.ts create mode 100644 test/webview.test.ts diff --git a/docs/webview-spec.md b/docs/webview-spec.md new file mode 100644 index 0000000..0e96e72 --- /dev/null +++ b/docs/webview-spec.md @@ -0,0 +1,67 @@ +# WebView Tool — Spec + +## TL;DR +A `web_browse` tool that uses Bun's built-in WebView to navigate, interact with, and +screenshot web pages — no Puppeteer, no Playwright, no separate browser download. + +## Problem +The existing `web_fetch` tool only extracts static HTML. SPAs, pages requiring JS +execution, form submissions, and interactive testing are impossible. The agent needs +a real browser to handle modern web apps. + +## Goal +- Navigate to any URL and extract text/HTML after JS execution +- Click buttons, fill forms, scroll pages +- Take screenshots (save to file, return path) +- Evaluate arbitrary JavaScript in the page context +- Generate PDFs from pages + +## Architecture + +### Tool: `web_browse` + +Input schema: +```ts +{ + url: string; // URL to navigate to + action?: 'text' | 'html' | 'screenshot' | 'evaluate' | 'click' | 'type' | 'scroll' | 'pdf'; + selector?: string; // CSS selector for click/type/scroll + text?: string; // text to type (for 'type' action) + script?: string; // JS to evaluate (for 'evaluate' action) + wait?: number; // ms to wait after navigation (default: 1000) +} +``` + +Default action: `text` — navigate and return page text content. + +### Implementation (src/tools-webview.ts) + +- Uses `Bun.WebView` with Chrome backend (headless) +- One view per call (create → navigate → act → close) +- Screenshot saved to `os.tmpdir()/shiro-webview-{timestamp}.png` +- PDF saved to `os.tmpdir()/shiro-webview-{timestamp}.pdf` +- Timeout: 15s per navigation, 5s per action +- Error handling: browser not found → suggest install; page error → return error text + +### Registration + +- Add to `tools.ts` as a `net` tool (network access) +- Add to `TOOL_META` as 'net' +- Add to `TOOL_SETS.net` set +- NOT in default tool set (opt-in via toolSets config) + +### Tests + +- Unit: tool schema validation +- Integration: navigate to data: URL, verify text extraction +- Integration: screenshot saves file +- Integration: evaluate JS returns result +- Edge case: invalid URL returns error +- Edge case: browser not available returns helpful message + +## Files touched + +- **NEW**: `src/tools-webview.ts` +- **EDIT**: `src/tools.ts` — register tool, add to TOOL_META and TOOL_SETS +- **NEW**: `test/webview.test.ts` +- **NEW**: `docs/webview.md` diff --git a/docs/webview.md b/docs/webview.md new file mode 100644 index 0000000..45f9049 --- /dev/null +++ b/docs/webview.md @@ -0,0 +1,83 @@ +# WebView Tool (web_browse) + +A built-in browser tool powered by Bun's native WebView. No Puppeteer, no Playwright, +no separate browser download — just the runtime's own headless browser. + +## What it does + +- **Navigate** to any URL and extract text/HTML after JS execution +- **Click** buttons, links, and interactive elements +- **Type** into input fields and forms +- **Scroll** to elements on the page +- **Screenshot** pages as PNG files +- **Evaluate** arbitrary JavaScript in the page context + +## When to use it + +Use `web_browse` instead of `web_fetch` when: + +- The page is a SPA (Single Page Application) that needs JavaScript +- You need to interact with the page (click, type, scroll) +- You need a visual screenshot of the page +- `web_fetch` returns incomplete or empty content +- You need to fill forms or submit data + +Use `web_fetch` when: + +- The page is static HTML +- You just need the text content quickly +- No JS execution is needed + +## Actions + +| Action | Description | Requires | +|--------|-------------|----------| +| `text` | Page text content after JS execution | — | +| `html` | Raw HTML of the page | — | +| `screenshot` | Save page as PNG, return file path | — | +| `evaluate` | Run JavaScript, return result | `script` | +| `click` | Click an element | `selector` | +| `type` | Type text into an input | `selector` + `text` | +| `scroll` | Scroll to an element | `selector` | + +## Examples + +```json +// Get page text +{ "url": "https://example.com", "action": "text" } + +// Take a screenshot +{ "url": "https://example.com", "action": "screenshot" } + +// Click a button +{ "url": "https://example.com", "action": "click", "selector": "button.submit" } + +// Fill a form +{ "url": "https://example.com", "action": "type", "selector": "#email", "text": "user@example.com" } + +// Run custom JS +{ "url": "https://example.com", "action": "evaluate", "script": "document.querySelectorAll('a').length" } +``` + +## Requirements + +- **macOS**: Uses system WKWebView (no install needed) +- **Linux/Windows**: Requires Chrome, Chromium, Edge, or Brave installed +- Bun 1.2+ for WebView support + +## Tool set + +`web_browse` is in the `net` tool set. Enable it in config: + +```json +{ "toolSets": ["core", "git", "net"] } +``` + +Or enable all sets with `toolSets: ["*"]`. + +## Limitations + +- Each call creates a fresh browser instance (no persistent sessions by default) +- Maximum 15s navigation timeout, 5s action timeout +- Screenshots are saved to the system temp directory +- Not available in Bun versions before 1.2 diff --git a/src/tools-webview.ts b/src/tools-webview.ts new file mode 100644 index 0000000..9b35175 --- /dev/null +++ b/src/tools-webview.ts @@ -0,0 +1,155 @@ +import { z } from 'zod'; +import { join } from 'node:path'; +import { tmpdir } from 'node:os'; + +/** + * Bun built-in WebView tool for shiro-neko. + * + * Uses `Bun.WebView` to navigate, interact with, and screenshot web pages. + * No Puppeteer, no Playwright, no separate browser download — just the + * runtime's own headless browser. + * + * Requires Chrome/Chromium/Edge on Linux/Windows, or WKWebView on macOS. + * Not in the default tool set — opt-in via `toolSets: ["net"]`. + */ + +const WEBVIEW_TIMEOUT = 15_000; + +export const webBrowseSchema = z.object({ + url: z.string().url().describe('URL to navigate to.'), + action: z + .enum(['text', 'html', 'screenshot', 'evaluate', 'click', 'type', 'scroll']) + .default('text') + .describe( + 'text: page text content after JS; html: raw HTML; screenshot: save PNG and return path; ' + + 'evaluate: run script and return result; click: click a selector; type: type into a selector; ' + + 'scroll: scroll to a selector.', + ), + selector: z.string().optional().describe('CSS selector for click/type/scroll actions.'), + text: z.string().optional().describe('Text to type (for "type" action).'), + script: z.string().optional().describe('JavaScript to evaluate (for "evaluate" action).'), + wait: z.number().min(0).max(30_000).default(1000).describe('Ms to wait after navigation before acting.'), +}); + +export type WebBrowseInput = z.infer; + +function timestamp(): string { + return Date.now().toString(36); +} + +/** Type-safe wrapper around view.evaluate that always returns a string. */ +async function evalJS(view: InstanceType, script: string): Promise { + const result = await (view as any).evaluate(script); + if (typeof result === 'string') return result; + if (result === undefined || result === null) return ''; + return String(result); +} + +export async function executeWebBrowse(args: WebBrowseInput): Promise { + // Check if WebView is available + if (typeof (Bun as any).WebView === 'undefined') { + return 'Bun.WebView is not available in this Bun version. Upgrade to Bun 1.2+ for WebView support.'; + } + + const { url, action, selector, text, script, wait } = args; + + // Create a new view for each call (clean state) + let view: InstanceType | undefined; + try { + view = new (Bun as any).WebView({ width: 1280, height: 720 }) as InstanceType; + + // Navigate + try { + await Promise.race([ + view.navigate(url), + new Promise((_, reject) => + setTimeout(() => reject(new Error(`Navigation timeout after ${WEBVIEW_TIMEOUT}ms`)), WEBVIEW_TIMEOUT), + ), + ]); + } catch (e) { + return `Failed to navigate to ${url}: ${e instanceof Error ? e.message : String(e)}`; + } + + // Wait for page to settle (JS execution, rendering) + if (wait > 0) { + await new Promise((resolve) => setTimeout(resolve, wait)); + } + + // Execute action + switch (action) { + case 'text': { + const bodyText = await evalJS(view, 'document.body?.innerText ?? ""'); + const title = await evalJS(view, 'document.title ?? ""'); + const finalUrl = view.url; + return `[${title}](${finalUrl})\n\n${bodyText}`; + } + + case 'html': { + return await evalJS(view, 'document.documentElement?.outerHTML ?? ""'); + } + + case 'screenshot': { + const buf = await view.screenshot({ format: 'png', encoding: 'buffer' }); + const path = join(tmpdir(), `shiro-webview-${timestamp()}.png`); + await Bun.write(path, buf); + return `Screenshot saved: ${path} (${(buf.length / 1024).toFixed(1)} KB)`; + } + + case 'evaluate': { + if (!script) return 'Error: "evaluate" action requires a "script" argument.'; + return await evalJS(view, script); + } + + case 'click': { + if (!selector) return 'Error: "click" action requires a "selector" argument.'; + try { + await view.click(selector); + return `Clicked: ${selector}`; + } catch (e) { + return `Failed to click "${selector}": ${e instanceof Error ? e.message : String(e)}`; + } + } + + case 'type': { + if (!selector) return 'Error: "type" action requires a "selector" argument.'; + if (!text) return 'Error: "type" action requires a "text" argument.'; + try { + await (view as any).type(selector, text); + return `Typed "${text}" into: ${selector}`; + } catch (e) { + return `Failed to type into "${selector}": ${e instanceof Error ? e.message : String(e)}`; + } + } + + case 'scroll': { + if (!selector) return 'Error: "scroll" action requires a "selector" argument.'; + try { + await view.scrollTo(selector); + return `Scrolled to: ${selector}`; + } catch (e) { + return `Failed to scroll to "${selector}": ${e instanceof Error ? e.message : String(e)}`; + } + } + } + } catch (e) { + return `WebView error: ${e instanceof Error ? e.message : String(e)}`; + } finally { + view?.close(); + } + + // Unreachable but satisfies TS + return ''; +} + +/** + * The web_browse tool definition. + */ +export const webBrowseTool = { + name: 'web_browse', + description: + 'Browse a web page with a real browser (Bun WebView). Navigates, executes JS, clicks, types, scrolls, and screenshots — no Puppeteer needed. Use for SPAs, pages requiring JS, or when web_fetch returns incomplete content.', + inputSchema: webBrowseSchema, + execute: async (args: WebBrowseInput) => { + return executeWebBrowse(args); + }, +}; diff --git a/src/tools.ts b/src/tools.ts index 29d5294..4c4a081 100644 --- a/src/tools.ts +++ b/src/tools.ts @@ -6,6 +6,7 @@ import { jail, posix, walk } from './ignore'; import { GIT_TOOL_NAMES, gitTools } from './tools-git'; import { NET_TOOL_NAMES, netTools } from './tools-net'; import { codegraphTool } from './tools-codegraph'; +import { webBrowseTool } from './tools-webview'; /** Max chars returned by any single tool. Beyond this the output is truncated. */ const MAX_OUTPUT = 30_000; @@ -709,6 +710,7 @@ export const tools = { ...gitTools, ...netTools, codegraph: codegraphTool, + web_browse: webBrowseTool, }; /** @@ -726,7 +728,7 @@ export const TOOL_SETS = { core: ['read_file', 'write_file', 'edit_file', 'glob', 'grep', 'bash', 'codegraph'], 'edit-plus': ['multi_edit', 'list_dir', 'read_many_files', 'apply_patch'], git: GIT_TOOL_NAMES, - net: NET_TOOL_NAMES, + net: [...NET_TOOL_NAMES, 'web_browse'], } as const satisfies Record; export type ToolSetName = keyof typeof TOOL_SETS; @@ -776,6 +778,7 @@ const CORE_META: Record = { multi_edit: 'mutate', apply_patch: 'mutate', codegraph: 'read', + web_browse: 'net', bash: 'mutate', }; // Git tools are read-only (they never write the tree); net tools reach the diff --git a/test/tool-meta.test.ts b/test/tool-meta.test.ts index 771afdf..b66afec 100644 --- a/test/tool-meta.test.ts +++ b/test/tool-meta.test.ts @@ -30,6 +30,8 @@ test('the net tools are classified net and never mutating', () => { expect(TOOL_META[name]).toBe('net'); expect(MUTATING_TOOLS).not.toContain(name); } + // web_browse is also a net tool (Bun WebView, not in NET_TOOL_NAMES) + expect(TOOL_META['web_browse']).toBe('net'); }); test('the git tools are classified read-only', () => { diff --git a/test/webview.test.ts b/test/webview.test.ts new file mode 100644 index 0000000..3e042c3 --- /dev/null +++ b/test/webview.test.ts @@ -0,0 +1,104 @@ +import { expect, test } from 'bun:test'; +import { executeWebBrowse } from '../src/tools-webview'; + +// These tests require a browser (Chrome/Chromium on Linux, WKWebView on macOS). +// They use data: URLs so no network access is needed. + +const HAS_WV = typeof (Bun as any).WebView !== 'undefined'; + +test.skipIf(!HAS_WV)('text action extracts page content from data URL', async () => { + const result = await executeWebBrowse({ + url: 'data:text/html,

Hello WebView

Test content

', + action: 'text', + wait: 500, + }); + expect(result).toContain('Hello WebView'); + expect(result).toContain('Test content'); +}); + +test.skipIf(!HAS_WV)('html action returns raw HTML', async () => { + const result = await executeWebBrowse({ + url: 'data:text/html,
OK
', + action: 'html', + wait: 500, + }); + expect(result).toContain('
OK
'); +}); + +test.skipIf(!HAS_WV)('evaluate action runs JavaScript', async () => { + const result = await executeWebBrowse({ + url: 'data:text/html,', + action: 'evaluate', + script: 'window.x', + wait: 500, + }); + expect(result).toContain('42'); +}); + +test.skipIf(!HAS_WV)('screenshot action saves a PNG file', async () => { + const result = await executeWebBrowse({ + url: 'data:text/html,

Screenshot

', + action: 'screenshot', + wait: 500, + }); + expect(result).toContain('Screenshot saved'); + expect(result).toContain('.png'); + expect(result).toContain('KB'); +}); + +test.skipIf(!HAS_WV)('click action works on a button', async () => { + const result = await executeWebBrowse({ + url: 'data:text/html,', + action: 'click', + selector: 'button', + wait: 500, + }); + expect(result).toContain('Clicked: button'); +}); + +test.skipIf(!HAS_WV)('type action fills an input', async () => { + const result = await executeWebBrowse({ + url: 'data:text/html,', + action: 'type', + selector: '#name', + text: 'hello', + wait: 500, + }); + expect(result).toContain('Typed "hello"'); +}); + +test.skipIf(!HAS_WV)('missing selector returns error', async () => { + const result = await executeWebBrowse({ + url: 'data:text/html,', + action: 'click', + wait: 500, + }); + expect(result).toContain('Error'); + expect(result).toContain('selector'); +}); + +test.skipIf(!HAS_WV)('missing script returns error for evaluate', async () => { + const result = await executeWebBrowse({ + url: 'data:text/html,

Test

', + action: 'evaluate', + wait: 500, + }); + expect(result).toContain('Error'); + expect(result).toContain('script'); +}); + +test.skipIf(!HAS_WV)('invalid URL returns error', async () => { + const result = await executeWebBrowse({ + url: 'https://this-does-not-exist-12345.invalid', + action: 'text', + wait: 100, + }); + expect(result).toContain('Failed to navigate'); +}); + +test('schema validation rejects invalid URL', () => { + // This tests the zod schema, not the WebView itself + const { webBrowseSchema } = require('../src/tools-webview'); + const result = webBrowseSchema.safeParse({ url: 'not-a-url' }); + expect(result.success).toBe(false); +});