Evals and synthetic data at full size.
Test your own agent against thousands of simulated users and scenarios, grade the transcripts with another model, and generate the data to do it. Run the full suite on every change instead of a sample.
Pre-launch. Accounts are open; the API and checkout open soon.
How it runs.
01
Generate the scenarios.
A model writes personas, tasks and edge cases for your product.
02
Run them.
Simulated users talk to your agent in parallel, one conversation per lane.
03
Grade them.
A stronger model scores each transcript against your rubric.
Why lanes fit it.
Run the whole suite.
No reason to sample down to save money: the full set costs the same.
On every change.
Rerun after each prompt tweak or model swap.
Parallel by design.
Lanes are how many conversations run at once, so a bigger plan finishes sooner.
Works with what you use.
Anthropic Python
python
import osimport anthropic client = anthropic.Anthropic( base_url="https://api.voidstone.net/anthropic", api_key=os.environ["VOIDSTONE_API_KEY"],) message = client.messages.create( model="kimi-k3", max_tokens=1024, messages=[{"role": "user", "content": "Review this diff."}],)print(message.content[0].text)