Model Router

Router playground

Run a task and see what the Router would pick, side by side with the default model the request would otherwise have gone to.

baseline for comparison
sent with each request
IT Support
Triage and next action
The hero scenario. Most tickets resolve on the fast tier in well under a second.
Fast~153 tok
Load
IT Support
Route a new ticket
A short structured task that should never reach the frontier tier. Category, priority, and queue in one pass.
Fast~93 tok
Load
IT Support
Unlock an account
High-volume L1 request and a clean self-service deflection candidate.
Fast~65 tok
Load
IT Support
VPN will not connect
Multi-step troubleshooting the balanced tier handles without escalating.
Balanced~30 tok
Load
IT Support
New-hire device setup
Checklist generation. A leader sees onboarding time drop.
Balanced~22 tok
Load
IT Support
Summarize a P1 incident
High-stakes synthesis. The cascade escalates to the frontier tier on purpose.
Frontier~25 tok
Load
IT Support
Detect a major incident
Pattern detection across a spike of tickets. Catch outages early.
Balanced~30 tok
Load
IT Support
Draft a KB article
Turn a resolved ticket into reusable knowledge. Feeds deflection.
Balanced~24 tok
Load
IT Support
Build an escalation package
Helpdesk work at its most valuable. Pattern analysis across many tickets, handled well below the frontier rate.
BalancedLong prompt~251 tok
Load
Prepare an escalation package for a recurring issue. Eleven users across two sites report that a managed ThinkPad drops the corporate wireless network roughly every forty minutes and reconnects on its own. You have the ticket history for all eleven users, the device inventory, and the wireless controller logs for the same period. Produce: 1. What the eleven cases have in common and what they do not, compared across device model, driver version, site, floor, and access point. 2. The hypotheses this pattern supports, ranked, with the evidence supporting and contradicting each. 3. What has already been tried across these tickets and what happened, so tier 2 does not repeat it. 4. The single next diagnostic step, and what its result would rule in or out. 5. A four-sentence summary a tier 2 engineer can read in under a minute. Do not recommend a fix the data cannot support. If the pattern points at the network rather than the devices, say so plainly and name what the network team should check.
MyAI
HR policy answer
Grounded answer from internal policy. Retrieval carries it, so the balanced tier is enough.
Balanced~25 tok
Load
MyAI
Summarize a benefits notice
Grounded summarization. The balanced tier handles it well.
Balanced~25 tok
Load
MyAI
Expense policy question
Policy lookup with a clear yes or no per item. The balanced tier answers it at a fraction of the frontier rate.
Balanced~25 tok
Load
MyAI
Answer from policy, with citations
The single largest workload in the savings model. Retrieval does the heavy lifting, so the balanced tier clears the bar.
BalancedLong prompt~221 tok
Load
An employee asks whether unused parental leave carries over after they transferred between two regions mid-year. Answer using the internal sources provided, which include the global leave policy, two regional addenda, and a benefits FAQ that is partly out of date. Requirements: 1. Give the direct answer first, in two sentences. 2. Cite the specific policy section behind each part of the answer. 3. Where the sources disagree, name the conflict explicitly, say which source governs, and say why. Do not quietly pick one. 4. Where the answer depends on facts you were not given, list exactly what the employee needs to confirm. 5. Flag any source that appears out of date and do not rely on it without saying so. Do not answer from general knowledge of employment law. If the provided sources do not settle the question, say so and route it to HR with the precise question to ask.
Sales
Draft an account summary
Everyday drafting where a frontier model is overkill.
Balanced~28 tok
Load
Sales
Competitive battlecard
Structured bullets a rep uses in the field.
Balanced~32 tok
Load
Sales
Executive proposal paragraph
Quality bar where the Router picks the cheapest frontier model that still clears it.
Frontier~29 tok
Load
xClaw
Explain and fix code
A tool-calling task. Shows the Router carrying agent workflows, not just chat.
FrontierTools~25 tok
Load
xClaw
Generate unit tests
Structured output. Shows agent-grade primitives beyond chat.
FrontierTools~24 tok
Load
xClaw
Explain a SQL query
A reasoning task the balanced tier handles cleanly and cheaply.
Balanced~21 tok
Load
xClaw
Review a pull request
The highest-volume frontier task in engineering. Long context in, structured judgement out, and no tools needed.
FrontierLong prompt~250 tok
Load
You are reviewing a pull request on a payment service before it merges to main. The change touches three files: the checkout handler, a shared validation library, and the token verification middleware. Review it for: 1. Security. Injection, authentication or authorization gaps, secrets committed in code, unsafe deserialization. 2. Regression risk. Behavior other services depend on, changed function signatures, altered error handling. 3. Correctness. Off-by-one errors, null and empty cases, race conditions, unhandled rejections. 4. Test coverage. Which changed paths have no test, and what the missing test should assert. For each finding give the file and approximate line, a severity of blocker, major, or minor, what breaks and under what condition, and the smallest safe fix. Do not restate what the code does. If a finding is speculative, mark it as needs verification and say exactly what to check. End with one merge recommendation: approve, approve with comments, or request changes.
xClaw
Plan a legacy modernization
Long-horizon reasoning over a codebase with hidden dependencies. The clearest case for a frontier model.
FrontierLong prompt~241 tok
Load
We are planning the modernization of a four-year-old payment form that three downstream services depend on. It has no test coverage and emits events that the checkout, fraud-check, and analytics services listen for. Produce a modernization plan that: 1. Maps the current behavior that must be preserved, including every event emitted, its payload shape, and which service consumes it. 2. Identifies the shared validation logic and says whether it can be extracted without changing behavior for the other consumers. 3. Proposes an incremental sequence of changes, smallest first, where every step is independently shippable and reversible. 4. Names the characterization tests to write before any refactor, and states what each one pins down. 5. Flags the changes that cannot be made safely without a coordinated release, and names which teams must be involved. Do not propose a full rewrite. Assume feature work cannot pause. Give the rollback path for every step.
xClaw
Generate tests from a change
Structured, bounded output. The balanced tier clears the bar here, which is where the saving comes from.
BalancedLong prompt~265 tok
Load
A change to an order-processing module introduces partial-batch processing. A batch that fails part way through now commits the orders that already succeeded and returns the failures for retry, where previously the whole batch was rolled back. Generate the test suite that should ship with this change. Cover: 1. The new behavior the change introduces, including the happy path and every branch. 2. Boundary conditions. Empty input, a single element, maximum page size, exactly at the limit, and one over the limit. 3. Error paths. Malformed payload, downstream timeout, and partial failure part way through a batch. 4. Regressions. Behavior that existed before the change and must still hold. For each test give a name that states the expected behavior, the arrange and act steps, and the specific assertion. Use table-driven tests where cases differ only by input. Flag any behavior in the change that cannot be tested without refactoring, and say what refactor would make it testable. Do not write tests that only assert the implementation back to itself.
xClaw
Find the root cause of an incident
Evidence weighing under time pressure. Escalates on purpose, and the answer is worth the frontier rate.
FrontierLong prompt~250 tok
Load
You are the on-call engineer. Checkout error rate went from 0.2 percent to 11 percent over eight minutes, then partially recovered. You have the application logs, the request traces, and the deployment timeline for the previous six hours. Work through this in order: 1. Establish the timeline. First anomaly, peak, partial recovery, and everything that changed in the six hours before it started. 2. Separate correlation from cause. List every candidate cause with the evidence for and against each one. 3. Identify the most likely root cause, say how certain you are, and say what would confirm it. 4. Give the immediate mitigation and the exact command or configuration change it needs. 5. Give the durable fix, stated separately from the mitigation. 6. Name the detection gap. What should have alerted sooner, and at what threshold. Do not speculate beyond the evidence provided. If the data cannot support a root cause, say so and name the specific query or log line that would close the gap.
xClaw
Document an undocumented service
Comprehension and synthesis rather than invention. The balanced tier holds the quality bar at a fraction of the cost.
BalancedLong prompt~256 tok
Load
Produce an onboarding document for a service that has no documentation. You have the repository structure, the public interface definitions, the database schema, and six months of commit history. The document must contain: 1. What the service does, in three sentences a new engineer can repeat back. 2. The domain model. Core entities, their relationships, and the invariants the code enforces. 3. The request lifecycle for the two highest-traffic endpoints, from entry to response, naming each component touched. 4. Every external dependency, what happens when it is unavailable, and whether that failure is handled or propagates. 5. The parts of the code that are load-bearing and poorly understood, inferred from commit churn and revert history. 6. The five questions a new engineer will ask in week one that this document does not answer. Write for an engineer who is competent but has never seen this codebase. Do not walk through the code line by line. Where you are inferring intent rather than reading it, say so.
Qira
Quick device question
The high-volume, low-value question. Answering it on the cheapest model in the pool is where routing pays for itself.
Fast~13 tok
Load
Qira
Summarize meeting notes
Short structured extraction from a short document. Nothing here earns a frontier rate.
Fast~17 tok
Load
General BU
Rewrite policy text
Everyday content task that rarely needs a frontier model.
Balanced~21 tok
Load
General BU
Translate an announcement
Many languages across regions. The balanced tier covers them.
Balanced~19 tok
Load
General BU
Exec briefing from notes
Synthesis at a quality bar that escalates to the frontier tier.
Frontier~24 tok
Load
General BU
Extract obligations from a contract
Document processing, the second largest workload. Precision matters more than price, so the cascade escalates.
FrontierLong prompt~214 tok
Load
Extract the obligations from this master services agreement and its two amendments so the delivery team can work from them without reading the contract. Produce a table with one row per obligation containing: who owes it, us or the customer; what is owed; the trigger that starts the clock; the deadline; the consequence of missing it; and the source clause. Then produce three short lists: 1. Obligations with a hard date in the next ninety days. 2. Obligations whose terms were changed by an amendment, showing the original and the current version. 3. Terms ambiguous enough that a reasonable person could read them two ways, with both readings stated. Do not summarize the contract and do not give an opinion on enforceability. Where an amendment supersedes the original, the table must show the current state and the change must appear in list two.
Loading approved models…

Picks the cheapest model that still clears the quality bar.

Cost is usually inversely related to quality percentage.

Third column shows the best alternative automatically.

Every run is scored against .

Compared model responses

Nothing run yetLoad a use case from the library or pick a couple of models, then run the task to see them side by side against the default.

Prototype note: model outputs are simulated for demonstration and do not call live providers. Each answer is an illustrative example of what that particular model produces for that kind of work. A model does not answer at one level across every task, so a specialised model can sit below its price tier on work outside its strength and at the top of it on work inside, and models that would answer alike on a given task share a sample. Note that the answer quality figure in the catalog is a single average for each model and does not vary by task, so it will not always track the difference between the answers shown here. Cost figures follow the published rate for each model in the catalog. Token counts, latency, and answer quality are representative and are replaced by measured values once the latency sample and the judge run land. The routing strategy, quality band, and default-model comparison follow the behavior specified in the Lenovo Model Router PRD. The Lenovo logo is the official mark, embedded for offline use.