Choosing by model name alone leaves out costs that appear as soon as an agent starts using tools. At the Opus 5 launch on July 24, 2026, Anthropic announced a beta for changing tools without invalidating the cache. That infrastructure detail deserves attention precisely because it affects a decision made during execution.
Three days earlier, on July 21, I questioned the choice of Sol in Cursor when Codex was available. I wanted to understand the combination of model and product chosen for the work. Anthropic's announcement adds a question to that concern: what happens to reused context when the configuration changes partway through a task?
Caching often appears in conversation with a generic promise of savings. I want to know which part of the workflow is being reused and which changes invalidate that reuse. Without that, any cost comparison becomes hard to explain. The beta proposes making certain changes less expensive; its concrete benefit depends on the announced conditions and how the task uses tools.
In a hypothetical example, an agent starts investigating with enough information and later needs an additional source. Adding a tool may be the correct next step. I want to examine the consultation's result and the cache's behavior separately. The former helps assess the information obtained. The latter shows the conditions under which previous work was reused.
A configuration that is cheaper to change still needs to respect scope. Discovering an available tool does not make its use necessary. If it helps resolve a question within the task, good. If it appears merely because it can be added, there is another possibility to maintain without a clear reason. Easier changes still need a purpose.
Understanding this beta requires recording the configuration and varying one thing at a time. Changing the model, tools, and request together leaves little basis for attributing the difference. After understanding a concrete change, the whole workflow can be assessed with less guessing about which part influenced the result and which part stayed the same.
My next step is to read the invalidation rules and choose a task that actually needs to change its tool set. That transition is where I want to check the promise. My question about Cursor and Codex still applies: the choice needs to explain the whole execution experience. Caching belongs in that explanation when it affects the work, with an observed benefit before I claim any savings of my own.