root@construct:~/logs/quando-a-ferramenta-troca-o-modelo$
<-- back to /logs
2026-08-08//LOG

If the tool switched models, show the switch

I want to know which model answered before I praise it or complain about it. That seems basic, but a name in the selector can conceal an execution that took another route. Comparing models without seeing that switch creates confidence the evidence cannot support.

On August 7, 2026, Anthropic announced new biology safeguards for Fable 5. According to its tests, the change reduced biology-related fallbacks by approximately 85%. Requests in certain areas would still go to Opus 5. The percentage describes the company's internal result; the detail I care about as a developer is identifying which model produced each answer.

In July, I questioned why Sol was being chosen in Cursor when Codex was available. I also asked for a comparison to establish which model would suit a particular role. That was a concrete tool choice, and the names were already getting tangled: the model on one side, the application and its capabilities on the other. Adding invisible fallback makes that conversation harder.

A weak answer can come from the selected model or an alternative invoked along the way. It can also reflect a product restriction. If I record everything under the first name on the screen, my analysis starts with a mistake. Build a routing rule from that impression and the mistake gets its own configuration file. Wonderful, now it is reusable.

The user still has a real complaint about a poor experience. Establishing where the answer came from helps resolve it. I want to distinguish an unclear request from an access restriction or a model switch. Each diagnosis calls for a different change. Habitually asking people to rewrite their prompt wastes their time.

An interface can show the switch beside the result and make further details available when needed. For a task with several calls, I need to reconstruct the sequence without hunting for identifiers scattered across another dashboard. Cost follows the same execution history: combining one model's prices with another model's answers ruins a comparison, however tidy the table looks.

My rule for evaluating this is straightforward: record the selected model, the model that actually answered, and the restrictions in effect. If the tool withholds that information, my conclusion applies to the whole product. I will not assign individual credit or blame to a model that may never have performed that part of the work.

Retrospective written in October 2026. The post date identifies the week revisited; the opinions draw on later experience.

Sources: Anthropic

The Broad Way | Kinho.dev