Max
The hardest and highest-stakes work. Tokaroo selects economically within the strongest measured task-specific quality band.
Install one Tokaroo key in your backend or harness. Tokaroo connects the agent and application ecosystem to the live model network, selects the compatible provider and model, handles fallback, and meters the result.
Server-side and trusted harness use only. Never ship a permanent key in browser or mobile code.
Keep provider selection, fallback, billing, and routing operations out of your application code.
curl https://api.tokaroo.com/v1/chat/completions \
-H "Authorization: Bearer $TOKAROO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "auto",
"messages": [
{"role": "user", "content": "Ship the fix"}
]
}'Tokaroo uses fresh task-specific evaluations internally, then selects the route that best matches your chosen operating objective.
The hardest and highest-stakes work. Tokaroo selects economically within the strongest measured task-specific quality band.
The best all-around mode for important technical, business, production, and customer-facing work. Most serious work belongs here.
Automatically applies the capability each step needs by balancing quality, cost, speed, reliability, tools, and escalation risk.
Rapid answers, edits, and back-and-forth. It keeps an acceptable task-specific quality bar, then prioritizes latency.
Everyday questions and routine tasks where you need an answer, not a project. Coding and website builds are intentionally excluded.
Choose the outcome; Tokaroo manages model selection, provider health, and fallback.
Aug 28, 2026 | Fresh Tokaroo evaluation evidence
Tokaroo is the router, not the harness host. Your app or agent runtime keeps control of tools, files, and execution.
Ten active Gwen engines plus OpenCode under compatibility evaluation. Each public profile is promoted only after its Tokaroo wire passes conformance checks.
Responses native | OpenAI-first
Messages native | Anthropic-first
OpenAI-compatible | optimized routing
OpenAI-compatible | conditional
OpenAI-compatible | Moonshot-first
OpenAI-compatible | DeepSeek-first
OpenAI-compatible | Qwen-first
GenerateContent native | Google-first
OpenAI-compatible | xAI-first
OpenAI-compatible | optimized routing
OpenAI-compatible | in evaluation
Harness-aware routing preserves each tool's native wire and favors model families proven to handle its tool loop well. The preference applies only inside the selected quality band; verified cross-family fallback remains available unless you explicitly pin a model.
Gwen chooses and runs the harness. Tokaroo supplies scoped runtime credentials, makes the optimized routing decision, enforces budgets, and records usage. That same pattern is available to other AI products.
Tokaroo routes metered TTS, transcription, realtime voice, and avatar requests to Raspy, with compatible managed fallback where available. Customers keep the same account, key, budgets, and usage ledger.
Use environment variables or a secrets manager in your backend or harness. Browser and mobile clients should call your backend or use short-lived, scoped credentials.
Start with Auto. Choose Basic, Fast, Pro, or Max when the workload has a clear cost, latency, or capability requirement.
Aug 28, 2026 | Fresh Tokaroo evaluation evidence
Aug 29, 2026 | Fresh Tokaroo evaluation evidence
Evaluation evidence is limited to successful Tokaroo tests in the latest 30-day window; imported zero-sample signals are excluded. The concrete serving model, serving provider, and private settlement cost remain internal.
Harness and model family are separate concerns: Hermes, Kimi, DeepSeek, Qwen, and the other profiles can all use Tokaroo's five optimized routing modes.