A model that cannot fit the task's environment is not an execution option yet. It can be interesting and worth following, but available memory does not increase because I liked the demo. That constraint belongs in the decision before I start arguing about answer quality.
On June 19, 2026, during a conversation about models, I was direct: I wanted the best one available. Looking back at Google's June 5 announcement gives that preference a concrete constraint. Google introduced Gemma 4 checkpoints using quantization-aware training, or QAT, aiming to reduce memory requirements for local execution while preserving quality during compression.
What interests me is expanding the set of models I can consider in an environment. Fitting changes the conversation. Before comparing answers, there is now a possible execution to examine. Preserving quality remains the manufacturer's proposition, something to check on the chosen work; reducing size alone does not establish whether the useful capability survived.
I also refuse to hide a machine replacement inside the word available. A candidate that runs in the current environment and one that requires different equipment demand different decisions. Comparing only their demonstration answers erases the effort needed to reach those answers. The surprise then arrives during setup, when the task should already be moving.
For a local evaluation, I will start by recording a configuration that fits. From there, I want to observe the whole task, including the compromises I accept in its output. A model name is insufficient to repeat the comparison. Changing the configuration halfway through and crediting every difference to the model produces a conclusion that is difficult to check.
Maintenance belongs in the calculation too. Somebody needs to understand how to update the environment and notice when a change alters test conditions. I have little interest in building a permanent laboratory merely to sustain a preference for a name. I want enough work to make the choice repeatable, with information that helps when the decision comes around again.
My decision is to put memory and environment in the same question as quality. QAT deserves evaluation because it opens a concrete possibility for local execution. First choose a viable configuration, then examine whether the answer is useful. If a candidate requires replacement equipment, that requirement goes into the comparison explicitly. The machine gets a say in the decision, however excited I am about the release.