root@construct:~/rants/modelo-melhor-trabalho-pronto$
<-- back to /rants
2026-09-03//OPINIAO

Better model, same delivery: where is the progress?

A better model needs to improve the work that reaches me. I like seeing capability advance, but I get impatient when the tool leaves that improvement stuck behind a poor integration. The promise of completing whole tasks raises the expectation because handing over the entire objective is exactly what I want.

On September 3, 2026, OpenAI introduced GPT-6 Astra for coding, computer use, and professional work. Distribution began with a limited group of organizations. I am revisiting the announcement with a later impression: on September 15, 2026, I said the models had become much better.

That same day, the conversation about moving from Claude to Codex exposed the concrete frustration. I considered the Codex model better, yet Claude could already do things Codex still struggled with in that experience. I asked for an investigation. When capability seems stronger and the workflow feels worse, I want to find where the improvement is getting lost.

An agent can understand an assignment while receiving insufficient context. It can know which hypothesis to verify without reaching the necessary environment. From my chair, those differences arrive as unfinished work. Fixing them requires locating the obstacle. Replacing the model without identifying which part failed can leave the same difficulty waiting in the next execution.

A limitation should become visible before it contaminates the conclusion. If the agent could not open a necessary file, I want to know while there is still a useful decision to make about proceeding. Thinking longer cannot recover an instruction the tool never sent. The product needs to expose that absence without forcing me to infer it from the answer's tone.

The comparison I care about follows an ordinary task from its initial context to a reviewable deliverable. I want to see what happens when new information arrives and whether the agent preserves the parts of the request that remain valid. A broad demonstration becomes less useful when it hides those transitions. I prefer a smaller set of tasks whose operation I can inspect.

My positive impression of the models still stands. Now I expect that capability to reach everyday work. I will judge the next tool by the task it can carry through, including the point where it actually needs me. My intervention should resolve a specific decision and let the rest proceed. Coordinating every stage myself would mean paying for autonomy while keeping the dispatcher's job.

Retrospective written in October 2026. The post date identifies the week revisited; the opinions draw on later experience.

Sources: OpenAI

The Broad Way | Kinho.dev