VILNIUS, Lithuania – September 17, 2026 – In 2026, the industry shift toward autonomous coding agents has triggered a “token paradox.” The per unit cost of intelligence has never been cheaper. However, the massive volume of background AI traffic has made AI spend the fastest-growing expense in engineering budgets. Addressing this, nexos.ai has launched its smart router, an intelligent routing capability that automatically matches engineering tasks to the most cost-effective AI models. It routes complex tasks to frontier models and routine work to low-cost alternatives. The result is a dramatic reduction in AI spend with zero disruption to developer workflows.
The growing costs of coding agents
Coding agents are currently the fastest-growing and least controllable category of AI spending inside engineering organizations. Tools like Claude Code, Codex, and Cursor autonomously decide which model to call and when. This makes cost optimization a major challenge that simply swapping to a cheaper model cannot solve.
"The instinct is to just point your coding agent at a cheaper model, but that approach fails in reality," says Žilvinas Girėnas, head of product at nexos.ai. "Point it at one cheap model and quality goes down on hard tasks. Point it at one expensive model and you end up paying Opus prices for requests the agent itself considered Haiku-grade work. The answer is knowing what the agent is actually trying to do at any given moment."
Mirror benchmarking: Routing based on live session data
Through the nexos.ai platform, this challenge can be observed firsthand. The platform continuously benchmarks performance, evaluates new models, and tests workflows to identify the most efficient cost-to-quality ratios. This visibility led directly to the development of the smart router.
There is a unique structural difference to how the smart router functions. While public benchmarks are useful for shortlisting models, they test static tasks. The smart router relies on a novel benchmarking method, Mirror benchmarking. With it, it continuously evaluates live production traffic at the session level. By tracking exactly how sessions grow, where costs land, and how real-world failures occur, the platform captures the unique mix of actual customer demands. This visibility into real-time cost-to-quality ratios directly drove the development of the smart router, ensuring it holds up against the unpredictable reality of production.
84% of coding tasks don't require expensive models
In production testing, nexos.ai’s smart router preserved the tiered structure that coding agents already use internally. The results revealed a stark division of labor:
- 16% of requests (Planning): Routed to frontier reasoning models (like Claude Opus) to determine what to build and how.
- 84% of requests (Editing): Routed to cost-efficient, open-weight models (like Kimi or GLM) to carry out the existing plan, apply diffs, and write files.
"In our production testing, this approach cut costs by 59.2% on one workload – saving over $5,400 on traffic that would have cost more than $9,200 at frontier-model prices," adds Girėnas. "On another test, the cut was 60.4%. The quality stayed intact because the frontier model still made every decision that actually mattered. The goal was never to eliminate the expensive model; it was to ensure it only runs where it changes the outcome."
Reducing frontier vendor dependency
The industry is shifting from finding the "best" model to finding the best price-to-quality ratio. With more specialized and open-weight models entering the market, token consumption has become a strategic diversification play.
Beyond immediate cost savings, routing 84% of requests to open-weight models removes enterprise dependency on any single AI provider. The router reads each request without altering it and switches models only at natural breakpoints in the session, keeping the cache largely intact to prevent costly rebuilds.
By utilizing the router as the decision layer, frontier vendors maintain their influence only where they earn their premium price, allowing the rest of the workload to move to whichever model performs best at market rates.
About nexos.ai
nexos.ai is a cutting-edge AI infrastructure company providing a centralized platform for enterprises to seamlessly integrate and manage multiple AI models. Founded in 2024 by Tomas Okmanas and Eimantas Sabaliauskas – who also co-founded bootstrapped global ventures including the $3B cybersecurity unicorn Nord Security and Oxylabs – nexos.ai addresses the urgent enterprise need to efficiently deploy, manage, and optimize AI models within organizations. Originating in the ecosystem of Lithuania-based tech accelerator Tesonet, the company attracted its first investment of $8M in early 2025 from Index Ventures, Creandum, Dig Ventures, and a number of prominent angel investors.
Media contact
Rasa Daunoravičienė
CMO
M.: 37067095388