feat: add web_browse tool for navigating and interacting with web pages using Bun's WebView
ci / check (macos-latest) (push) Canceled after 0s
ci / check (ubuntu-latest) (push) Canceled after 0s
ci / check (windows-latest) (push) Canceled after 0s

This commit is contained in:
asepharyana
2026-09-04 22:37:38 +07:00
parent bb36af62e2
commit 922519aa01
6 changed files with 415 additions and 1 deletions
+67
View File
@@ -0,0 +1,67 @@
# WebView Tool — Spec
## TL;DR
A `web_browse` tool that uses Bun's built-in WebView to navigate, interact with, and
screenshot web pages — no Puppeteer, no Playwright, no separate browser download.
## Problem
The existing `web_fetch` tool only extracts static HTML. SPAs, pages requiring JS
execution, form submissions, and interactive testing are impossible. The agent needs
a real browser to handle modern web apps.
## Goal
- Navigate to any URL and extract text/HTML after JS execution
- Click buttons, fill forms, scroll pages
- Take screenshots (save to file, return path)
- Evaluate arbitrary JavaScript in the page context
- Generate PDFs from pages
## Architecture
### Tool: `web_browse`
Input schema:
```ts
{
url: string; // URL to navigate to
action?: 'text' | 'html' | 'screenshot' | 'evaluate' | 'click' | 'type' | 'scroll' | 'pdf';
selector?: string; // CSS selector for click/type/scroll
text?: string; // text to type (for 'type' action)
script?: string; // JS to evaluate (for 'evaluate' action)
wait?: number; // ms to wait after navigation (default: 1000)
}
```
Default action: `text` — navigate and return page text content.
### Implementation (src/tools-webview.ts)
- Uses `Bun.WebView` with Chrome backend (headless)
- One view per call (create → navigate → act → close)
- Screenshot saved to `os.tmpdir()/shiro-webview-{timestamp}.png`
- PDF saved to `os.tmpdir()/shiro-webview-{timestamp}.pdf`
- Timeout: 15s per navigation, 5s per action
- Error handling: browser not found → suggest install; page error → return error text
### Registration
- Add to `tools.ts` as a `net` tool (network access)
- Add to `TOOL_META` as 'net'
- Add to `TOOL_SETS.net` set
- NOT in default tool set (opt-in via toolSets config)
### Tests
- Unit: tool schema validation
- Integration: navigate to data: URL, verify text extraction
- Integration: screenshot saves file
- Integration: evaluate JS returns result
- Edge case: invalid URL returns error
- Edge case: browser not available returns helpful message
## Files touched
- **NEW**: `src/tools-webview.ts`
- **EDIT**: `src/tools.ts` — register tool, add to TOOL_META and TOOL_SETS
- **NEW**: `test/webview.test.ts`
- **NEW**: `docs/webview.md`
+83
View File
@@ -0,0 +1,83 @@
# WebView Tool (web_browse)
A built-in browser tool powered by Bun's native WebView. No Puppeteer, no Playwright,
no separate browser download — just the runtime's own headless browser.
## What it does
- **Navigate** to any URL and extract text/HTML after JS execution
- **Click** buttons, links, and interactive elements
- **Type** into input fields and forms
- **Scroll** to elements on the page
- **Screenshot** pages as PNG files
- **Evaluate** arbitrary JavaScript in the page context
## When to use it
Use `web_browse` instead of `web_fetch` when:
- The page is a SPA (Single Page Application) that needs JavaScript
- You need to interact with the page (click, type, scroll)
- You need a visual screenshot of the page
- `web_fetch` returns incomplete or empty content
- You need to fill forms or submit data
Use `web_fetch` when:
- The page is static HTML
- You just need the text content quickly
- No JS execution is needed
## Actions
| Action | Description | Requires |
|--------|-------------|----------|
| `text` | Page text content after JS execution | — |
| `html` | Raw HTML of the page | — |
| `screenshot` | Save page as PNG, return file path | — |
| `evaluate` | Run JavaScript, return result | `script` |
| `click` | Click an element | `selector` |
| `type` | Type text into an input | `selector` + `text` |
| `scroll` | Scroll to an element | `selector` |
## Examples
```json
// Get page text
{ "url": "https://example.com", "action": "text" }
// Take a screenshot
{ "url": "https://example.com", "action": "screenshot" }
// Click a button
{ "url": "https://example.com", "action": "click", "selector": "button.submit" }
// Fill a form
{ "url": "https://example.com", "action": "type", "selector": "#email", "text": "user@example.com" }
// Run custom JS
{ "url": "https://example.com", "action": "evaluate", "script": "document.querySelectorAll('a').length" }
```
## Requirements
- **macOS**: Uses system WKWebView (no install needed)
- **Linux/Windows**: Requires Chrome, Chromium, Edge, or Brave installed
- Bun 1.2+ for WebView support
## Tool set
`web_browse` is in the `net` tool set. Enable it in config:
```json
{ "toolSets": ["core", "git", "net"] }
```
Or enable all sets with `toolSets: ["*"]`.
## Limitations
- Each call creates a fresh browser instance (no persistent sessions by default)
- Maximum 15s navigation timeout, 5s action timeout
- Screenshots are saved to the system temp directory
- Not available in Bun versions before 1.2
+155
View File
@@ -0,0 +1,155 @@
import { z } from 'zod';
import { join } from 'node:path';
import { tmpdir } from 'node:os';
/**
* Bun built-in WebView tool for shiro-neko.
*
* Uses `Bun.WebView` to navigate, interact with, and screenshot web pages.
* No Puppeteer, no Playwright, no separate browser download — just the
* runtime's own headless browser.
*
* Requires Chrome/Chromium/Edge on Linux/Windows, or WKWebView on macOS.
* Not in the default tool set — opt-in via `toolSets: ["net"]`.
*/
const WEBVIEW_TIMEOUT = 15_000;
export const webBrowseSchema = z.object({
url: z.string().url().describe('URL to navigate to.'),
action: z
.enum(['text', 'html', 'screenshot', 'evaluate', 'click', 'type', 'scroll'])
.default('text')
.describe(
'text: page text content after JS; html: raw HTML; screenshot: save PNG and return path; ' +
'evaluate: run script and return result; click: click a selector; type: type into a selector; ' +
'scroll: scroll to a selector.',
),
selector: z.string().optional().describe('CSS selector for click/type/scroll actions.'),
text: z.string().optional().describe('Text to type (for "type" action).'),
script: z.string().optional().describe('JavaScript to evaluate (for "evaluate" action).'),
wait: z.number().min(0).max(30_000).default(1000).describe('Ms to wait after navigation before acting.'),
});
export type WebBrowseInput = z.infer<typeof webBrowseSchema>;
function timestamp(): string {
return Date.now().toString(36);
}
/** Type-safe wrapper around view.evaluate that always returns a string. */
async function evalJS(view: InstanceType<typeof Bun.WebView>, script: string): Promise<string> {
const result = await (view as any).evaluate(script);
if (typeof result === 'string') return result;
if (result === undefined || result === null) return '';
return String(result);
}
export async function executeWebBrowse(args: WebBrowseInput): Promise<string> {
// Check if WebView is available
if (typeof (Bun as any).WebView === 'undefined') {
return 'Bun.WebView is not available in this Bun version. Upgrade to Bun 1.2+ for WebView support.';
}
const { url, action, selector, text, script, wait } = args;
// Create a new view for each call (clean state)
let view: InstanceType<typeof Bun.WebView> | undefined;
try {
view = new (Bun as any).WebView({ width: 1280, height: 720 }) as InstanceType<typeof Bun.WebView>;
// Navigate
try {
await Promise.race([
view.navigate(url),
new Promise<never>((_, reject) =>
setTimeout(() => reject(new Error(`Navigation timeout after ${WEBVIEW_TIMEOUT}ms`)), WEBVIEW_TIMEOUT),
),
]);
} catch (e) {
return `Failed to navigate to ${url}: ${e instanceof Error ? e.message : String(e)}`;
}
// Wait for page to settle (JS execution, rendering)
if (wait > 0) {
await new Promise((resolve) => setTimeout(resolve, wait));
}
// Execute action
switch (action) {
case 'text': {
const bodyText = await evalJS(view, 'document.body?.innerText ?? ""');
const title = await evalJS(view, 'document.title ?? ""');
const finalUrl = view.url;
return `[${title}](${finalUrl})\n\n${bodyText}`;
}
case 'html': {
return await evalJS(view, 'document.documentElement?.outerHTML ?? ""');
}
case 'screenshot': {
const buf = await view.screenshot({ format: 'png', encoding: 'buffer' });
const path = join(tmpdir(), `shiro-webview-${timestamp()}.png`);
await Bun.write(path, buf);
return `Screenshot saved: ${path} (${(buf.length / 1024).toFixed(1)} KB)`;
}
case 'evaluate': {
if (!script) return 'Error: "evaluate" action requires a "script" argument.';
return await evalJS(view, script);
}
case 'click': {
if (!selector) return 'Error: "click" action requires a "selector" argument.';
try {
await view.click(selector);
return `Clicked: ${selector}`;
} catch (e) {
return `Failed to click "${selector}": ${e instanceof Error ? e.message : String(e)}`;
}
}
case 'type': {
if (!selector) return 'Error: "type" action requires a "selector" argument.';
if (!text) return 'Error: "type" action requires a "text" argument.';
try {
await (view as any).type(selector, text);
return `Typed "${text}" into: ${selector}`;
} catch (e) {
return `Failed to type into "${selector}": ${e instanceof Error ? e.message : String(e)}`;
}
}
case 'scroll': {
if (!selector) return 'Error: "scroll" action requires a "selector" argument.';
try {
await view.scrollTo(selector);
return `Scrolled to: ${selector}`;
} catch (e) {
return `Failed to scroll to "${selector}": ${e instanceof Error ? e.message : String(e)}`;
}
}
}
} catch (e) {
return `WebView error: ${e instanceof Error ? e.message : String(e)}`;
} finally {
view?.close();
}
// Unreachable but satisfies TS
return '';
}
/**
* The web_browse tool definition.
*/
export const webBrowseTool = {
name: 'web_browse',
description:
'Browse a web page with a real browser (Bun WebView). Navigates, executes JS, clicks, types, scrolls, and screenshots — no Puppeteer needed. Use for SPAs, pages requiring JS, or when web_fetch returns incomplete content.',
inputSchema: webBrowseSchema,
execute: async (args: WebBrowseInput) => {
return executeWebBrowse(args);
},
};
+4 -1
View File
@@ -6,6 +6,7 @@ import { jail, posix, walk } from './ignore';
import { GIT_TOOL_NAMES, gitTools } from './tools-git'; import { GIT_TOOL_NAMES, gitTools } from './tools-git';
import { NET_TOOL_NAMES, netTools } from './tools-net'; import { NET_TOOL_NAMES, netTools } from './tools-net';
import { codegraphTool } from './tools-codegraph'; import { codegraphTool } from './tools-codegraph';
import { webBrowseTool } from './tools-webview';
/** Max chars returned by any single tool. Beyond this the output is truncated. */ /** Max chars returned by any single tool. Beyond this the output is truncated. */
const MAX_OUTPUT = 30_000; const MAX_OUTPUT = 30_000;
@@ -709,6 +710,7 @@ export const tools = {
...gitTools, ...gitTools,
...netTools, ...netTools,
codegraph: codegraphTool, codegraph: codegraphTool,
web_browse: webBrowseTool,
}; };
/** /**
@@ -726,7 +728,7 @@ export const TOOL_SETS = {
core: ['read_file', 'write_file', 'edit_file', 'glob', 'grep', 'bash', 'codegraph'], core: ['read_file', 'write_file', 'edit_file', 'glob', 'grep', 'bash', 'codegraph'],
'edit-plus': ['multi_edit', 'list_dir', 'read_many_files', 'apply_patch'], 'edit-plus': ['multi_edit', 'list_dir', 'read_many_files', 'apply_patch'],
git: GIT_TOOL_NAMES, git: GIT_TOOL_NAMES,
net: NET_TOOL_NAMES, net: [...NET_TOOL_NAMES, 'web_browse'],
} as const satisfies Record<string, readonly string[]>; } as const satisfies Record<string, readonly string[]>;
export type ToolSetName = keyof typeof TOOL_SETS; export type ToolSetName = keyof typeof TOOL_SETS;
@@ -776,6 +778,7 @@ const CORE_META: Record<string, ToolEffect> = {
multi_edit: 'mutate', multi_edit: 'mutate',
apply_patch: 'mutate', apply_patch: 'mutate',
codegraph: 'read', codegraph: 'read',
web_browse: 'net',
bash: 'mutate', bash: 'mutate',
}; };
// Git tools are read-only (they never write the tree); net tools reach the // Git tools are read-only (they never write the tree); net tools reach the
+2
View File
@@ -30,6 +30,8 @@ test('the net tools are classified net and never mutating', () => {
expect(TOOL_META[name]).toBe('net'); expect(TOOL_META[name]).toBe('net');
expect(MUTATING_TOOLS).not.toContain(name); expect(MUTATING_TOOLS).not.toContain(name);
} }
// web_browse is also a net tool (Bun WebView, not in NET_TOOL_NAMES)
expect(TOOL_META['web_browse']).toBe('net');
}); });
test('the git tools are classified read-only', () => { test('the git tools are classified read-only', () => {
+104
View File
@@ -0,0 +1,104 @@
import { expect, test } from 'bun:test';
import { executeWebBrowse } from '../src/tools-webview';
// These tests require a browser (Chrome/Chromium on Linux, WKWebView on macOS).
// They use data: URLs so no network access is needed.
const HAS_WV = typeof (Bun as any).WebView !== 'undefined';
test.skipIf(!HAS_WV)('text action extracts page content from data URL', async () => {
const result = await executeWebBrowse({
url: 'data:text/html,<h1>Hello WebView</h1><p>Test content</p>',
action: 'text',
wait: 500,
});
expect(result).toContain('Hello WebView');
expect(result).toContain('Test content');
});
test.skipIf(!HAS_WV)('html action returns raw HTML', async () => {
const result = await executeWebBrowse({
url: 'data:text/html,<html><body><div id="app">OK</div></body></html>',
action: 'html',
wait: 500,
});
expect(result).toContain('<div id="app">OK</div>');
});
test.skipIf(!HAS_WV)('evaluate action runs JavaScript', async () => {
const result = await executeWebBrowse({
url: 'data:text/html,<script>window.x = 42;</script>',
action: 'evaluate',
script: 'window.x',
wait: 500,
});
expect(result).toContain('42');
});
test.skipIf(!HAS_WV)('screenshot action saves a PNG file', async () => {
const result = await executeWebBrowse({
url: 'data:text/html,<h1 style="color:blue">Screenshot</h1>',
action: 'screenshot',
wait: 500,
});
expect(result).toContain('Screenshot saved');
expect(result).toContain('.png');
expect(result).toContain('KB');
});
test.skipIf(!HAS_WV)('click action works on a button', async () => {
const result = await executeWebBrowse({
url: 'data:text/html,<button onclick="document.title = clicked">Click me</button>',
action: 'click',
selector: 'button',
wait: 500,
});
expect(result).toContain('Clicked: button');
});
test.skipIf(!HAS_WV)('type action fills an input', async () => {
const result = await executeWebBrowse({
url: 'data:text/html,<input id="name" />',
action: 'type',
selector: '#name',
text: 'hello',
wait: 500,
});
expect(result).toContain('Typed "hello"');
});
test.skipIf(!HAS_WV)('missing selector returns error', async () => {
const result = await executeWebBrowse({
url: 'data:text/html,<button>Click</button>',
action: 'click',
wait: 500,
});
expect(result).toContain('Error');
expect(result).toContain('selector');
});
test.skipIf(!HAS_WV)('missing script returns error for evaluate', async () => {
const result = await executeWebBrowse({
url: 'data:text/html,<h1>Test</h1>',
action: 'evaluate',
wait: 500,
});
expect(result).toContain('Error');
expect(result).toContain('script');
});
test.skipIf(!HAS_WV)('invalid URL returns error', async () => {
const result = await executeWebBrowse({
url: 'https://this-does-not-exist-12345.invalid',
action: 'text',
wait: 100,
});
expect(result).toContain('Failed to navigate');
});
test('schema validation rejects invalid URL', () => {
// This tests the zod schema, not the WebView itself
const { webBrowseSchema } = require('../src/tools-webview');
const result = webBrowseSchema.safeParse({ url: 'not-a-url' });
expect(result.success).toBe(false);
});