root@construct:~/logs/agente-botoes-aplicativo$
<-- back to /logs
2026-03-25//LOG

What good is a model that cannot reach the app?

A better model needs a way to reach the work, including the application's buttons. That is why Anthropic's computer use announcement interests me. In March 2026, the company introduced computer control in Cowork and Claude Code as a preview. Its approach preferred connectors and used the screen when a more direct route was unavailable.

Looking back in October 2026, I remember a frustration I expressed in September: Claude could do things Codex was still struggling with. At the same time, I considered the Codex model better. I wanted that capability to show up in the experience. Preferring one model and another product was a pretty annoying position to be in.

That comparison does not identify computer use as the cause. It does make my expectation clear: the application needs to give the model a way to execute. Understanding a request perfectly helps little when the information sits behind access the tool cannot provide. I can spend an afternoon refining instructions and still face the same closed door.

A connector that returns the right data handles part of that journey. Screen interaction extends the reach to applications without a dedicated integration, but adds work: finding the location and interpreting what appeared. The announcement itself acknowledges that this route can be slower and require another attempt. I prefer the direct path when it exists.

To compare products, I want a short task that crosses this boundary. Locating a setting and reporting its state is a test I propose without needing to change anything. The expected result is clear. I can examine how the agent got there and check the evidence behind its answer.

I also want to count human preparation. If I have to open every screen and arrange everything before asking for help, that belongs in the task's cost. A demonstration can look great after somebody prepares the entire scene. During work, that somebody is me, with other things to do.

My decision is to examine access before blaming reasoning. When an agent stalls, I want to know which tool it had and where its route ended. The next comparison needs the same intended result and an account of how much work came back to me. After that, changing the model becomes a useful experiment instead of another hopeful click in a menu.

Retrospective written in October 2026. The post date identifies the week revisited; the opinions draw on later experience.

Sources: Anthropic

The Broad Way | Kinho.dev