OpenHands Model Router on your own endpoint: what works, what isn't wired yet, what it costs

Model Router is the headline feature of OpenHands v1.25.0: a classifier model reads your task and routes it to the right LLM profile. We wired it to our own OpenAI-compatible endpoint on release day and ran the full loop. The short version: the routing mechanism itself works end to end on a custom endpoint, and a routing decision costs about $0.0025 in classifier tokens. Two gaps are worth knowing before you plan around it, one in how you get the new build at all, one in the marquee "Run at conversation start" switch. Both look like early-build gaps in a project shipping eight releases in 25 days, and we have reported what we found to the project.

Disclosure: Wallaby Token sells API access to kimi-k3, the endpoint we tested against, so read this as a field report from our own setup, not a neutral review.

Getting v1.25.0 at all

If you installed OpenHands the documented way, uv tool install openhands followed by openhands serve, you are not running v1.25.0. Even if you installed today. The launcher pulls docker.openhands.dev/openhands/openhands:latest, and both :latest and :main currently resolve to builds from July 21–24, 2026, with the image label reading main-arm64. The v1.25.0 feature set lives in a different image, ghcr.io/openhands/agent-canvas, an all-in-one build that bundles the agent server, the frontend, and the automations service. The local setup docs don't mention it yet.

One more wrinkle: version numbers now run on three independent tracks. The app release is v1.25.0, our agent-canvas image reports git ref v1.53.0, and the agent server has its own numbering. When you report or compare issues, note which track you are on.

Wiring it up on a BYO endpoint

The clean path is four API calls: create a provider connection with your base URL and key, create LLM profiles that inherit credentials from the connection, create the router as a meta-profile with a classifier model and routing classes, then bind both to an agent profile and launch conversations against it.

Two things to budget for. First, the default router templates assume OpenHands' own model catalog: the prefilled classifier is minimax-m3, and the sample model table quotes benchmark scores for models you don't have. On your own endpoint you rewrite all of it. Second, each routing class's model field must match a saved profile name, case-insensitively, rather than a raw model ID. Neither is documented anywhere we could find; we learned both from the API schema and the tracker.

The switch that isn't wired yet

v1.25.0's settings UI includes a "Run at conversation start" toggle, added in PR #17797, that is supposed to classify each new conversation's first message automatically. PR #17840 goes further and auto-enables it when you create your first router.

On the current self-hosted build, that toggle doesn't do anything yet, though the reason turned out to be subtler than we first reported. Our initial pass concluded the setting never saved: we had written run_router_at_conversation_start into the meta-profile itself, where the API quietly drops unknown fields. After we filed #18068, the project's automated triage pointed us at the correct path — the toggle persists through PATCH /api/settings as misc_settings.app_preferences.run_router_at_conversation_start, an opaque frontend-owned container. Retested that way, the value does save and survives on disk.

Two genuine gaps remain. The settings page doesn't read the value back: with true stored, a freshly reloaded UI still renders the toggle off. And nothing consumes the preference: we started a new conversation with it on and an active router bound, and our gateway's billing log shows zero classifier calls — the agent simply answered. The string run_router_at_conversation_start appears nowhere in the installed backend source, so in this build the preference is stored but inert. We've added the full retest to the issue thread.

OpenHands Model Router settings after a reload: the preference is stored as true on disk, yet the "Run on first message" toggle still renders off

What actually routes today

The routing loop works when the agent invokes the route_task_to_model tool, which an active router injects into the system prompt automatically. We asked the agent to route a task explicitly and watched the full cycle: a classifier call of 709 tokens in and 44 out, about $0.0025 at our endpoint's prices; an observation reading "Classified task as 'trivial…' and switched to LLM profile 'kimi-k3-fast'"; a live swap of the agent's LLM configuration mid-conversation; and the task continuing on the routed profile.

The routed conversation: "Task routed (classified as trivial → kimi-k3-fast) and done", with the finished task below

The caveat: left to itself, the agent never calls the tool. We handed it a genuinely harder task, an LRU cache with TTL expiry, thread-safety, and a pytest suite, and it ground through the whole thing on the default profile without routing. Until the start-toggle lands in the backend, routing only happens when you ask for it.

One honest limitation of our test: our gateway serves a single model, so both routing targets pointed at kimi-k3 under different profile names. That proves the mechanism of classification, selection, and live switch, but says nothing about routing quality across genuinely different models.

The default-profile trap

One more for the error map: creating a conversation over the API with an empty agent_settings object silently ignores both your active profile and your router. The built-in default agent profile carries a soft reference to a default LLM profile that doesn't exist, and resolution falls through to a hardcoded gpt-5.6 pointed at api.openai.com. On a network where that host is unreachable, the run burns four-plus retry rounds over several minutes instead of failing fast. The fix is to always pass an explicit agent_profile_id bound to your own profiles.

Smaller notes

  • Conversation Metrics still shows $0.0000 for custom models, since LiteLLM's price map has no entry for them; your provider's console is the source of truth.
  • Conversation title generation now follows the running profile's LLM, per #16886. We verified titles billing to our own endpoint. A welcome one.
  • The settings blob carries a meta_profile_llms field that stays empty no matter what we did; possibly vestigial.

Verdict

Model Router is a good idea with a working core: classification is cheap, the switch is clean, and the profile-and-connection credential model held up throughout. But on release day, self-hosted users get a routing tool the agent won't use on its own and a start-toggle that saves fine but routes nothing, behind an install path that doesn't deliver the release. We would wait for the backend to catch up before designing workflows around automatic routing, and we will update this page when it does. Our v1.25.0 quick look covers what else shipped, and the setup guide covers getting OpenHands running against your own endpoint in the first place.