option
Home
News
New $1.5B Router Model Hits 93% Accuracy, Eliminating Expensive Retraining Costs

New $1.5B Router Model Hits 93% Accuracy, Eliminating Expensive Retraining Costs

November 24, 2025
149

Researchers at Katanemo Labs have unveiled Arch-Router, an advanced routing model and framework engineered to intelligently direct user queries to the most appropriate large language model (LLM).

For companies developing products that leverage multiple LLMs, Arch-Rolver tackles a central dilemma: how to automatically route each request to the ideal model for the task, without depending on inflexible logic or expensive retraining whenever updates are needed.

The challenges of LLM routing

As the variety of available LLMs expands, developers are shifting from single-model configurations to multi-model architectures that utilize the distinct capabilities of different models for specialized functions—such as generating code, summarizing text, or editing images.

LLM routing has become an essential technique for constructing and running such systems, serving as an intelligent traffic director that guides each user query to the model best suited to handle it.

Current routing approaches generally fit into two main groups: task-based routing, which assigns queries according to predefined task categories, and performance-based routing, which seeks the best trade-off between expense and output quality.

However, task-based systems often falter when user intent is ambiguous or changes over the course of a conversation—especially in multi-turn dialogues. Performance-based routing, meanwhile, tends to prioritize static benchmark results, frequently overlooking actual user preferences and adapting slowly to new models without costly retraining.

As the researchers at Katanemo Labs state in their paper, a deeper issue is that “existing routing methods have practical limitations in real-world applications. Most are optimized for benchmark performance but ignore human preferences, which are guided by subjective evaluation criteria.”

The team emphasizes the importance of routing systems that “reflect subjective human judgments, provide greater transparency, and remain easily adjustable as both models and applications evolve.”

A new framework for preference-aligned routing

To overcome these issues, the researchers developed a “preference-aligned routing” framework that matches incoming queries to routing rules based on custom user preferences.

In this system, users define their routing policies using natural language through a two-tier “Domain-Action Taxonomy.” This structure reflects how people naturally describe tasks: starting with a broad category—the Domain, such as “legal” or “finance”—and drilling down to a specific task—the Action, like “summarization” or “coding.”

Each policy is then mapped to a preferred model, empowering developers to base routing choices on practical requirements rather than benchmark metrics alone. According to the paper, “This taxonomy acts as a mental model to help users create well-defined, structured routing policies.”

The routing procedure operates in two phases. First, a preference-aligned router model evaluates the user’s query alongside all available policies and picks the best-fitting one. Second, a mapping function connects the selected policy to its assigned LLM.

Because the logic for selecting a model is separated from the policy definition, developers can add, remove, or update models just by editing the routing rules—without retraining or changing the router. This separation enables the necessary flexibility for production environments, where models and applications are constantly changing.

Preference-aligned routing framework (source: arXiv)
Preference-aligned routing framework Source: arXiv

The policy selection is powered by Arch-Router, a compact 1.5-billion-parameter language model optimized for preference-aware routing. Arch-Router takes the user query and the full list of policy descriptions as input, then outputs the identifier of the most suitable policy.

Since policies are included in the input, the system can adjust to new or updated routes during inference through in-context learning—no retraining required. This generative strategy enables Arch-Router to leverage its pre-trained understanding to interpret the meaning of both the query and the policies, and to analyze complete conversation histories in one go.

One common worry with including lengthy policy lists in a prompt is the risk of higher latency. However, the team built Arch-Router for high efficiency. “Even with extensive routing policies, we can expand Arch-Router’s context window with very little effect on latency,” says Salman Paracha, co-author of the paper and Founder/CEO of Katanemo Labs. He points out that latency is mainly determined by output length, and Arch-Router only outputs a short policy name—such as “image_editing” or “document_creation.”

Arch-Router in action

To create Arch-Router, the team fine-tuned a 1.5B parameter variant of the Qwen 2.5 model using a carefully assembled dataset of 43,000 examples. They then benchmarked it against leading proprietary models from OpenAI, Anthropic, and Google across four public datasets designed to test conversational AI systems.

The findings indicate that Arch-Router achieved the top overall routing score of 93.17%, outperforming all other models—including top-tier proprietary ones—by an average of 7.71%. The model’s edge became more apparent in longer conversations, showcasing its superior ability to maintain context across multiple exchanges.

Arch-Router vs other models (source: arXiv)
Arch-Router vs other models Source: arXiv

In real-world use, this methodology is already being applied in multiple settings, notes Paracha. For instance, in open-source coding platforms, developers rely on Arch-Router to guide different parts of their workflow—like “code design,” “code understanding,” and “code generation”—to the LLMs most effective for each step. Similarly, organizations can route document creation tasks to a model such as Claude 3.7 Sonnet while sending image editing requests to Gemini 2.5 Pro.

The system is also well-suited “for personal assistants across various fields, where users perform a range of activities from summarizing text to answering factual queries,” Paracha explained, adding that “in such situations, Arch-Router helps product teams consolidate and improve the user’s overall experience.”

This framework is built into Arch, Katanemo Labs’ AI-native proxy server for agents, which supports the implementation of granular traffic management rules. For example, when adding a new LLM, a team can route a small percentage of traffic under a certain policy to the new model, validate its performance using internal analytics, and then confidently shift all traffic over. The company is also working to integrate its tools with evaluation platforms to make this workflow even smoother for corporate developers.

At its core, the objective is to help organizations move beyond disconnected AI implementations. “Arch-Router—and the Arch platform overall—enables developers and businesses to evolve from fragmented LLM usage to a unified, policy-governed system,” Paracha states. “When users perform a wide range of tasks, our platform converts that diversity of tasks and models into a cohesive experience, making the final product feel seamless and intuitive.”

Related article
Sam Altman Sparks Debate Over AI's Deceleration Sam Altman Sparks Debate Over AI's Deceleration Listen onApple PodcastsListen onSpotifyOpenAI CEO Sam Altman recently suggested that it may be time to “pace the rate of AI development” to allow society to “harden around some of these new capability levels.”On the latest episode of TechCrunch’s Equ
OpenAI fights Apple trade secret lawsuit OpenAI fights Apple trade secret lawsuit OpenAI rebutted Apple’s trade secret allegations on Tuesday, arguing the lawsuit is unfounded.“We take these claims seriously but see no evidence supporting them,” OpenAI stated, as reported by Bloomberg’s Ed Ludlow on X. “We support fair competition
OpenAI robotics head Caitlin Kalinowski resigns over Pentagon partnership OpenAI robotics head Caitlin Kalinowski resigns over Pentagon partnership OpenAI robotics leader Caitlin Kalinowski has stepped down following the company’s controversial partnership with the Department of Defense.“This wasn’t an easy call,” Kalinowski explained in a social media statement. “While AI plays a vital role in
Related Special Topic Recommendations
SEO Best AI SERP Analysis Tools for Search Strategy
Best AI SERP Analysis Tools for Search Strategy

2026 Latest Best Top-rated AI SERP Analysis Tools for Search Strategy are curated by XIX.AI through rigorous real-world tests and weekly updated rankings. These powerful tools help you unlock hidden optimization opportunities, boost content performance, and gain a competitive edge in search. Discover your perfect tool to elevate your search strategy today. Explore now!

13 tools
xix.ai
Data Analysis Best AI Anomaly Detection Tools for KPI Monitoring across SaaS and Ecommerce Teams
Best AI Anomaly Detection Tools for KPI Monitoring across SaaS and Ecommerce Teams

2026 Latest Best Top-rated AI Anomaly Detection Tools for KPI Monitoring in SaaS and Ecommerce teams! XIX.AI has curated a powerful, game-changing collection based on rigorous real-world tests and weekly updated rankings. You’ll find detailed free vs paid comparison insights to help you identify the must-try solution that boosts productivity and unlocks your AI edge. Explore now!

10 tools
xix.ai
writing Best AI Outline Generators for Long-Form SEO Articles
Best AI Outline Generators for Long-Form SEO Articles

2026 Latest Best Top-Rated AI Outline Generators for Long-Form SEO Articles, meticulously curated by XIX.AI. These powerful tools offer game-changing assistance in creating high-quality content quickly, boosting writing efficiency significantly. Get a free vs paid comparison along with real-world tests and detailed rankings to help you find the must-try option that suits your needs. Explore now to unlock your AI edge.

8 tools
xix.ai
Education and Learning AI Study Tools for Homework and Exam Prep
AI Study Tools for Homework and Exam Prep

2026 Latest Best AI Study Tools for Homework and Exam Prep! XIX.AI curates a top-rated list of powerful, game-changing tools that help students boost productivity, streamline homework completion, and ace exams through real-world tests. Get a free vs paid comparison, detailed rankings, and must-try options to unlock your AI edge. Explore now!

10 tools
xix.ai
Music composition AI Vocal Demo Tools for Songwriters, Hooks, Toplines, and Multilingual Draft Sessions
AI Vocal Demo Tools for Songwriters, Hooks, Toplines, and Multilingual Draft Sessions

2026 Latest Best AI Vocal Demo Tools for Songwriters, Hook Creators, and Multi-Language Content Teams! XIX.AI has curated a top-rated list of powerful game-changing tools that go through rigorous real-world tests. You’ll find detailed free vs paid comparison data, comprehensive rankings, and must-try options to help you boost writing efficiency and unlock your creative potential. Explore now to discover your perfect tool for all your content needs!

9 tools
xix.ai
Business Best AI Competitive Research Tools for Small Businesses
Best AI Competitive Research Tools for Small Businesses

2026 Latest Best Top-rated AI Competitive Research Tools for Small Businesses! XIX.AI has curated a highly powerful game-changing collection, updated weekly with rigorous real-world tests and detailed rankings. You can find a comprehensive free vs paid comparison to help you identify the must-try tools that boost your productivity and give you a competitive edge. Explore now to discover your perfect tool!

9 tools
xix.ai
Comments (1)
0/500
WillGarcía
WillGarcía April 5, 2026 at 10:00:35 PM EDT

Arch-Routerの構想は面白いね。社内でどのLLMを使うか毎回悩んでたから、これがあれば効率化に繋がりそう。ただ、精度93%って、結局残りの7%で重大なミスルーティングが起きたりしない? 医療や法務のようなクリティカルな分野への適用は少し不安かな。😅 開発元のKatanemo Labs、これでインフラ市場に本格参戦するつもり?

OR