
A coding agent can consume a lot of money without learning something new at every step. It rereads files, processes long logs, and repeatedly sends old results back to the language model. NVIDIA’s SoL-Pi research project targets that orchestration: Recorded token traffic on EdgeBench falls by 44.7 to 49 percent compared with Pi. The tradeoff is a slightly lower average score, and outside that test the balance becomes more pronounced.
Key takeaways
- SoL-Pi optimizes a coding agent’s working environment without retraining the main model.
- The efficiency configuration retains approximately 94 percent of Pi’s EdgeBench score at roughly one-third lower API cost.
- Four mechanisms combine actions, archive outputs, compact context, and selectively process long diagnostic logs.
- The open extension is designed for Pi. Its features are disabled by default; it is not a savings switch for every agent application.
Costs also accumulate between reasoning tasks
A language model answers a request using the context supplied with it. For an agent, that context includes not just the user’s question but also tool results and previous work. A long error message may therefore travel with later requests even though only a particular line matters to the next decision. The model is not simply doing more work on the code; some of the overhead comes from how its information is organized.
That organization is handled by the harness, the orchestration layer connecting the model, tools, and environment. Among other things, it determines which results remain visible and when another model call is needed. SoL-Pi changes those workflows. The approach complements improvements in the capability and cost of new models: Even an unchanged model can operate more economically in a differently organized working environment.
The research was developed with GPT-5.6 Sol and subsequently tested with Claude Opus 5. These are the models used in this experiment, not automatically the best choices today for every project. The fixed model baseline is especially useful to the study’s claim. It helps distinguish orchestration savings from improvements caused by switching models.
Four changes to the workflow
Action Fusion combines a file change with the subsequent validation command in one tool call. When a clearly specified edit should be followed by the appropriate test anyway, the local environment can run both actions and report them together. Our interpretation is that an extra reasoning turn here may be organizational overhead rather than additional problem-solving.
ObservationPack stores large tool outputs locally and replaces repeated transmission with a reference and a short excerpt. If the agent needs more later, it can retrieve exact portions of the original. The practical difference from a simple summary matters: When an error message is unclear, the original observation remains accessible. Less text in the active context does not necessarily mean losing the full log.
Evidence-Preserving Reducer uses a cheaper model to condense long diagnostic outputs into a compact evidence receipt. Retained quotations are checked against the archived original; if verification fails, the original result is preserved. That check establishes that quoted passages match their source. It does not by itself prove that a summary understood every relevant cause. Developers therefore still need a way to return to the original when necessary.
Online Context Compact treats completed subtasks as opportunities to compact context. Whether that pays off also depends on the overhead of rewriting and the reuse of previously processed inputs. This explains why “summarize earlier” alone is not a universal savings rule. Compaction can shorten future requests while also invalidating previously useful caching.
Fewer tokens do not automatically buy the same utility
The paper distinguishes the complete efficiency configuration from a variant using only the strongest individual mechanism for each model. The largest token reduction belongs to the efficiency stack. Combining its lower consumption with the higher score of another configuration into one promise would be misleading. Anyone deciding whether to adopt it must compare the quality and cost of the same variant.
On 63 CPU tasks from Terminal-Bench 4, SoL-Pi solves 15 tasks, while Pi and Codex each solve 18. Lower overall spending therefore also buys fewer completed tasks in this case. That does not make the extension useless, but it shows why “nearly identical performance” is too broad without a testing context. Savings per run and savings per successfully completed assignment are different metrics.
The token percentage is not an invoice total, either. The study uses API prices from August 17, 2026; different token categories and cache usage have different costs. Those figures therefore cannot be presented as guaranteed savings for a current subscription. Nor do they show that a Codex or Claude Code usage allowance will last proportionally longer after installation. The service actually used would need the same workflows and the same accounting.
Who could benefit from an independent comparison
Developers already using Pi for long, recurring development tasks can consider the extension as a candidate for a controlled comparison. The public repository describes a standalone extension for an unmodified Pi installation. Every mechanism requires explicit activation. The documentation offers a cautious starting point with the local Action Fusion and ObservationPack features before adding extra model calls or automatic context compaction.
In our view, a meaningful comparison should run the same tasks with the same model, permissions, and success criteria. It should measure not just tokens and cost but also passing checks, necessary rework, and time to a usable result. Local log archives also need a deliberate retention policy; the documentation notes that a reducer can send diagnostic material to its configured model.
SoL-Pi offers a concrete new possibility: Making agents more efficient by letting their environment identify and reduce repeated overhead. The convincing everyday evidence would be a repeatable reduction in cost per assignment actually completed. What matters is not the most striking percentage but whether the selected combination preserves enough context in a developer’s own project and continues to perform the work the agent was hired to do.

