Nvidia SoL-Pi: How to Cut Coding Agent Token Costs in Half

Alex da Cruz
Alex da Cruz is a full-stack developer based in São Paulo, Brazil. He works with React, TypeScript and automation, and uses AI daily to solve real problems in code and operations — not as a demo. He has run an e-commerce operation end to end, and now builds and maintains the automation pipeline behind this blog. He writes about what he actually tests.
Running autonomous coding agents for long stretches causes token usage to balloon quickly as single predictions turn into massive reasoning chains. According to Nvidia researchers, the primary cost drivers often lie not in the underlying model, but in the harness—the control layer managing how an agent interacts with its environment.
What mechanisms does SoL-Pi use to reduce token waste?
The Nvidia research team built a system that tests agent traces and deploys four specific optimization mechanisms to eliminate redundant computational steps:
- Action Fusion: Merges sequential steps, such as a code edit followed immediately by a test run, removing an extra language model call.
- Online Context Compact: Trims accumulated context after planning steps without sacrificing crucial details.
- ObservationPack: Archives lengthy tool outputs, replacing them with brief summaries on subsequent steps.
- Evidence-Preserving Reducer: Routes bulky error logs to a cheaper model for distillation while keeping critical clues intact.
Combined, these four mechanisms yield an efficiency variant that uses 49 percent fewer tokens while retaining 93.7 percent of standard performance on EdgeBench.
What are the practical cost savings for development teams?
In financial terms, these reductions translate to spending $894 instead of $1,339 on benchmark runs. The authors estimate direct savings of $8.75 to $13.50 per hour compared to native Codex and Claude Code harnesses, dropping expenses by $4.36 to $5.71 per hour against the baseline Pi harness.
However, engineering leads should note trade-offs. Shorter contexts can reduce prompt cache reuse, and performance gains vary across different tests—such as Terminal-Bench 4, where the system completed fewer tasks overall despite consuming a quarter less budget.
Sources
Frequently asked questions
- What does the Nvidia SoL-Pi system do?
- SoL-Pi automatically optimizes the control layer, or harness, between a coding model and its environment, reducing token usage by up to 49 percent.
- How much money does SoL-Pi save on development?
- The system saves between $8.75 and $13.50 per hour compared to native Codex and Claude Code harnesses, dropping test run costs from $1,339 to $894.
- Does context reduction impact code quality?
- Yes, while token usage drops significantly, the efficiency variant retains about 93.7 percent of the original benchmark performance and can occasionally affect prompt cache reuse.
Comments
0 comments
Be the first to comment.
Continue Lendo

Manus 2.0: mobile control, video editing and AI agents
Manus 2.0 adds mobile remote control, video editing, multiplayer game hosting, and cuts agent running costs by 32 percent.

Shopify WebMCP Checkout: AI Agents Complete Purchases
Shopify's new WebMCP checkout support allows browser-based AI agents to complete purchases natively using structured APIs.

Meta Muse: Free Ubuntu cloud computer for every user
Meta's Muse agent equips every user with a free Ubuntu Linux cloud computer and isolated runtime cells for complex task execution.