Deep research, as deep as it needs to go.
Ask a question and let an agent search, read dozens or hundreds of pages, follow the citations and write a report. Almost all of it is reading, and reading is input tokens.
Pre-launch. Accounts are open; the API and checkout open soon.
How it runs.
01
A lead agent plans the search.
It breaks the question into threads and decides what to look for.
02
Readers fan out.
Quick models read pages in parallel and pull out what matters, with sources.
03
One model writes it up.
A report with citations, which you can send back for another pass.
Why lanes fit it.
Input-heavy by nature.
A single report can read millions of tokens of pages before it writes a paragraph. None of them are billed per token.
Go another round.
Asking for a deeper second pass costs time, not money.
Parallel readers.
More lanes means more pages read at once, so reports come back sooner.
Works with what you use.
LangChain
python
import osfrom langchain_openai import ChatOpenAI llm = ChatOpenAI( model="kimi-k3", base_url="https://api.voidstone.net/v1", api_key=os.environ["VOIDSTONE_API_KEY"],)print(llm.invoke("Summarise this pull request.").content)