What If Your AI Agent Never Had to Leave the Browser?

작성자

카테고리:

← 피드로
DEV Community · heavy shield · 2026-09-29 개발(SW)

heavy shield

Introduction

Most AI agents today run in a Python process on a server or your laptop. They call APIs, maybe execute shell commands, and return text. But what if the agent’s entire runtime lived inside a browser tab? No backend, no container, no SSH. Just JavaScript and WebAssembly, with a Python kernel compiled to WASM.

This post walks through a minimal agent loop that runs entirely client-side using Pyodide — Python in the browser via WebAssembly. We’ll build a tool-using agent that can do arithmetic, read from a virtual filesystem, and stop under explicit conditions. All code is runnable in a modern browser.

Problem

Server-side agents have friction:

  • Deployment: You need a host, secrets management, and network egress.
  • Latency: Every tool call is a round trip.
  • Security surface: Arbitrary code execution on a server is a serious risk.
  • Demo friction: Sharing a working agent means sharing infrastructure.

A browser-native agent flips this. The sandbox is the browser tab. The runtime is WebAssembly. The only network call is loading the Python runtime itself.

Solution

We’ll use Pyodide to run CPython in the browser. The agent loop is a simple ReAct-style loop: the model proposes a tool call, we execute it in Python, append the result, and repeat until a termination condition is met.

We won’t call a real LLM here — instead we use a deterministic policy function so the demo is reproducible and offline. Swap the policy for a fetch to an LLM API and the loop is unchanged. We’ll call this a “ReAct-style” loop because it follows the observe-think-act pattern, not because it implements the exact paper.

Termination conditions (explicit):

  1. The policy returns a final action.
  2. The step count exceeds MAX_STEPS (default 8).
  3. A tool raises an exception that the policy cannot recover from.

Implementation

1. HTML shell

<!doctype html>
<html>
<head><meta charset="utf-8"><title>Browser Agent</title></head>
<body>
  <pre id="log"></pre>
  <script src="https://cdn.jsdelivr.net/pyodide/v0.26.2/full/pyodide.js"></script>
  <script type="module">
    const log = (m) => document.getElementById('log').textContent += m + '\n';
    const pyodide = await loadPyodide();
    await pyodide.runPythonAsync(await (await fetch('agent.py')).text());
    const result = await pyodide.runPythonAsync('run_agent("What is 21 * 2 plus 8?")');
    log(result);
  </script>
</body>
</html>

Enter fullscreen mode Exit fullscreen mode

2. The agent module (agent.py)

import json
import re

MAX_STEPS = 8

# --- Tools -------------------------------------------------------------

def tool_calc(expr: str) -> str:
    """Evaluate a pure arithmetic expression.

    SECURITY WARNING: eval() executes arbitrary Python. This implementation
    restricts input to digits and operators via a regex whitelist. Do not
    remove the whitelist. Do not pass user-controlled strings from an
    untrusted source. For production, use ast.literal_eval or a parser.
    """
    if not re.fullmatch(r"[0-9+\-*/(). ]+", expr):
        raise ValueError(f"unsafe expression: {expr!r}")
    return str(eval(expr, {"__builtins__": {}}, {}))


_FS = {"notes.txt": "remember: 42"}

def tool_read_file(path: str) -> str:
    if path not in _FS:
        raise FileNotFoundError(path)
    return _FS[path]


TOOLS = {
    "calc": tool_calc,
    "read_file": tool_read_file,
}

# --- Policy (replace with an LLM call) ---------------------------------

def policy(question: str, history: list) -> dict:
    """Deterministic stand-in for an LLM. Returns a dict action.

    Swap this for a fetch() to your LLM of choice. The loop below does not
    care where the action came from.
    """
    if not history:
        # First turn: extract a math expression from the question.
        m = re.search(r"([0-9+\-*/(). ]+)", question)
        if m:
            return {"type": "tool", "name": "calc", "args": {"expr": m.group(1).strip()}}
        return {"type": "final", "content": "no expression found"}

    last = history[-1]
    if last["role"] == "tool" and last["name"] == "calc":
        # Second turn: add 8 as required by the question.
        return {"type": "tool", "name": "calc", "args": {"expr": f"{last['result']} + 8"}}
    if last["role"] == "tool" and last["name"] == "calc":
        return {"type": "final", "content": last["result"]}
    return {"type": "final", "content": "done"}

# --- Agent loop --------------------------------------------------------

def run_agent(question: str) -> str:
    history = []
    for step in range(MAX_STEPS):
        action = policy(question, history)
        history.append({"role": "assistant", "action": action})

        if action["type"] == "final":
            return f"[step {step}] {action['content']}"

        if action["type"] == "tool":
            fn = TOOLS.get(action["name"])
            if fn is None:
                history.append({"role": "tool", "name": action["name"],
                                "error": "unknown tool"})
                continue
            try:
                result = fn(**action["args"])
                history.append({"role": "tool", "name": action["name"],
                                "result": result})
            except Exception as e:
                history.append({"role": "tool", "name": action["name"],
                                "error": str(e)})

    return f"[halted: exceeded MAX_STEPS={MAX_STEPS}]"

Enter fullscreen mode Exit fullscreen mode

3. What actually happens

  1. The browser loads Pyodide (about 6 MB gzipped).
  2. agent.py is fetched and executed in the WASM interpreter.
  3. run_agent runs the loop. Step 0 calls calc("21 * 2") → 42. Step 1 calls calc("42 + 8") → 50. Step 2 returns final.
  4. The result is printed to the page.

No server. No API key. The entire agent state lives in the tab and is discarded when you close it.

4. Swapping in a real LLM

Replace policy with something like:

import json
from js import fetch  # Pyodide exposes the browser fetch API

async def policy_llm(question, history):
    resp = await fetch(
        "https://api.example.com/v1/chat",
        {"method": "POST",
         "headers": {"Content-Type": "application/json"},
         "body": json.dumps({"q": question, "history": history})}
    )
    data = await resp.json()
    return json.loads(data.action)

Enter fullscreen mode Exit fullscreen mode

You’ll need to make run_agent async and await the policy. The rest of the loop is identical. Note that calling an LLM from the browser exposes your API key to the user — use a short-lived token or a proxy you control.

Key Takeaways

  • Browser-native agents are real. Pyodide gives you CPython in a tab. That’s enough to run a tool-using loop, a virtual filesystem, and a policy function.
  • Explicit termination matters. Define MAX_STEPS, a final action, and error handling. Without them, loops run forever and burn tokens.
  • The sandbox is the feature. No server means no remote code execution surface, but it also means no secrets. Design accordingly.
  • Never eval untrusted input. The tool_calc example uses a regex whitelist and an empty __builtins__ map. That is a mitigation, not a guarantee. For anything real, use ast.literal_eval or a dedicated parser.
  • The policy is swappable. The loop is agnostic to whether the action came from a regex, a local model, or a remote LLM. That separation is what makes the pattern portable.

Try it: drop the two files in a folder, serve with python -m http.server, and open the page. You’ll have an agent that never leaves the browser.

원문에서 계속 ↗