root@construct:~/rants/agentes-demais-tokens-rapidos$
<-- back to /rants
2026-05-10//OPINIAO

Spare agents can waste my time too

I want parallel work that finishes tasks; spare agents give me more things to follow. In June 2026, I asked for persistent agents to gain speed. The next day, I suspected there were too many and suggested reducing the number. My desire to accelerate quickly met the need to understand what was running.

That sequence is what I bring, in October 2026, to this look back at the Gemma 4 drafters announced in May. Google promised faster generation through speculative decoding. I like waiting less. Organizing requests remains my responsibility, and extra capacity lets me distribute useful work or multiply a poorly formulated question.

Opening several investigations before deciding what I expect from each is easy. They all look busy. If they cover the same question, answers can arrive quickly and still require me to reconcile similar explanations. I have gained speed in receiving material and created work in reading it. The activity report looks lively; my task remains open.

Persistent agents interest me because of continuity. I want work advancing without rebuilding the organization around every request. Each task needs a boundary and somewhere to return its result. Distributing a vague demand across several conversations makes the coordinator responsible for discovering afterward what each one should have done.

My attention belongs in the calculation too. I may be reading one analysis when another finishes. Even with good answers, I have to decide what deserves attention now and what can wait. The number of executions should fit prepared work and my ability to use the results. Filling every available slot is no obligation I accepted by opening the tool.

My June suspicion establishes no ideal number. It requires me to ask what another execution contributes. Independent research with its own question has a clear reason to run in parallel. When I am still formulating the main problem, spreading that uncertainty merely distributes the mess more efficiently. A faster response can still leave the original decision unresolved.

I will leave room to increase or reduce the number as demand changes. The next agent needs a different question or an independent responsibility, with a destination for its answer. If I cannot describe that, I will let the current task finish and use its result first. I prefer spending new speed on closing a pending item before manufacturing another reading queue.

Retrospective written in October 2026. The post date identifies the week revisited; the opinions draw on later experience.

Sources: Google

The Broad Way | Kinho.dev