Models

Frontier open models, more on the way.

Every model runs on the same lanes and every token is unmetered on paid plans. Use the id as the model field in any OpenAI- or Anthropic-compatible tool.

Pre-launch. Accounts are open; the API and checkout open soon.

Models
6
Makers
5
Context
Up to 1M
Speed
up to 150 tok/s

6 of 6 models

  • Kimi K3

    Moonshot AI

    Strongest all-rounder for long agentic sessions. Reads screenshots.

    Model id
    kimi-k3
    Context
    1M tokens
    Lanes per request
    6 lanesx2 at Turbo
    Runs on
    ProScale
  • GLM 5.3

    Z.ai

    Deliberate, always reasons first. Good at long-horizon refactors.

    Model id
    glm-5.3
    Context
    1M tokens
    Lanes per request
    2 lanesx2 at Turbo
    Runs on
    SoloProScale
  • MiMo V2.5

    Xiaomi

    Omnimodal: reads images, video and audio alongside code. Quick for its size.

    Model id
    mimo-v2.5
    Context
    1M tokens
    Lanes per request
    1 lanex2 at Turbo
    Runs on
    SoloProScale
  • GLM 5.3 Flash

    Z.ai

    The quick GLM. Cheap on lanes, good for edits, tests and subagents.

    Model id
    glm-5.3-flash
    Context
    1M tokens
    Lanes per request
    1 lanex2 at Turbo
    Runs on
    SoloProScale
  • DeepSeek V4.1 Flash

    DeepSeek

    Fastest in the lineup. Ideal for subagents and tight edit loops.

    Model id
    deepseek-v4.1-flash
    Context
    1M tokens
    Lanes per request
    1 lanex2 at Turbo
    Runs on
    FreeSoloProScale
  • gpt-oss-120b

    OpenAI

    OpenAI's open-weight reasoning model. Low, medium or high effort; quick, with solid tool calling.

    Model id
    gpt-oss-120b
    Context
    128K tokens
    Lanes per request
    1 lanex2 at Turbo
    Runs on
    SoloProScale

Speed per request: Normal up to 150 tok/s, Turbo 150–400 tok/s, depending on model. Tokens are unmetered on every paid plan.

Coming soon

More models, the day they drop.

We’re adding GPU capacity right now. Until it’s online the lineup is deliberately short: every model we list, we can serve at full speed on every lane. New frontier open models join as capacity allows, aiming for release day.

Estimates per request, depending on model. We're adding GPU capacity right now to raise them.

See plans
  • Next frontier release

    Added as soon as it can run on every plan.

    Coming soon

  • Larger context

    Longer windows as memory comes online.

    Coming soon

  • Faster lanes

    Normal speed past 150 tok/s as more GPUs come online.

    Coming soon

  • More makers

    A wider lineup of open models for agent work.

    Coming soon