Browser agents that click through the web.
Agents that open a browser, read the page, click, type and move on: filling in forms, testing your own app, collecting data from sites without an API. Each step sends the page or a screenshot back to the model.
Pre-launch. Accounts are open; the API and checkout open soon.
How it runs.
01
Give it a task and a browser.
Your browser framework drives the browser; the model decides each next step.
02
It looks, then acts.
Each step sends the page or a screenshot and gets back a click, some typing or an answer.
03
Many sessions at once.
One session per lane, for testing, collecting data or chores.
Why lanes fit it.
Screenshots add up.
A single task can take dozens of steps, each with a fresh page or screenshot. On lanes that isn’t a line item.
Retries don’t cost more.
When a page changes and the agent tries again, it holds its lane longer. Nothing else.
Models that see.
MiMo V2.5 and Kimi K3 read screenshots; the one-lane MiMo handles most steps.
Works with what you use.
OpenAI TypeScript
import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.voidstone.net/v1", apiKey: process.env.VOIDSTONE_API_KEY,}); const stream = await client.chat.completions.create({ model: "glm-5.3", messages: [{ role: "user", content: "Refactor this function." }], stream: true,});for await (const chunk of stream) { process.stdout.write(chunk.choices[0]?.delta?.content ?? "");}