EvalSignal 008 / Independent model evaluation
A newsletter issue is not supposed to age this quickly.
Five days ago, we published our guide to inexpensive Qwen models for large vision workloads. The results were useful and accurate. They are also already becoming obsolete.
Two new models arrived within hours of each other:
Both are more modern and substantially more capable than the models in our previous shortlist. They combine strong visual understanding with useful text and tool-planning intelligence, support text, images, and video, and offer roughly million-token context windows.
They remain firmly in the budget API tier. During its launch promotion, GLM-5.3 Flash is even cheaper than every model in our previous comparison.
We ran each model through two frozen 100-case suites:
Every model received the same prompts, screenshots, tool definitions, temperatures, and scoring rules.
That produced 400 final case results with zero errors.
The planning suite does not measure every form of intelligence or execute the complete task. It measures whether a model can interpret an instruction and choose a sensible first action. The vision suite requires important visual facts to be identified without critical contradictions.
| Model | Pass | Rubric | Vision | Text |
|---|---|---|---|---|
| Qwen3.8 Flash | 79 | 94.8% | 8.42s | 2.55s |
| GLM-5.3 Flash | 76 | 93.8% | 5.49s | 7.24s |
Qwen3.8 Flash won vision quality. It recorded three more strict passes, performed better on the most challenging cases, and was especially strong on authentication screens.
GLM-5.3 Flash won vision latency. It processed screenshots roughly three seconds faster at the median while finishing only one rubric point behind Qwen.
The text-planning result was more nuanced:
Qwen is therefore the faster planner. GLM is slightly more consistent about emitting structured tools and following the expected initial route.
| Model | Pass | 100-case cost |
|---|---|---|
| Qwen3.8 Flash | 79 | $0.041 |
| GLM-5.3 Flash | 76 | $0.021 promo |
| Qwen3-VL-32B Instruct | 69 | $0.022 |
| Qwen3-VL-30B-A3B Instruct | 68 | $0.024 |
Qwen3.8 Flash costs roughly two cents more per 100 screenshots than the previous Qwen leaders, but delivers ten additional strict passes and much stronger general-purpose capabilities.
GLM-5.3 Flash currently costs less than the old models while beating the previous best result by seven passes. However, its price is temporarily discounted by 50% through September 9. Its normal pricing is almost identical to Qwen3.8 Flash.
Neither model is genuinely small.
Qwen3.8 Flash has a 125B-parameter main model but activates only 6B parameters per token. GLM-5.3 Flash stores 320B parameters while activating 18B per token.
This sparse architecture lets the models retain far more capacity than yesterday's 8B and 30B budget models without performing dense computation across every parameter for every generated token.
The result is a new class of API model: enormous total capacity, relatively inexpensive inference, million-token context, and credible performance across both vision and text.
Best overall default: Qwen3.8 Flash
Choose Qwen when vision accuracy and fast text planning matter most. It has the stronger difficult-case floor and the best overall result in this comparison.
Best value during the promotion: GLM-5.3 Flash
Choose GLM when screenshot latency and immediate cost matter most. Its vision quality is close to Qwen, and its temporary price is exceptional.
After the promotion
Treat their token prices as approximately equal. Choose Qwen for vision hit rate and text speed. Choose GLM for screenshot speed and more consistent structured tool output.
Keep an older Qwen3-VL model only when one narrow metric dominates
Qwen3-VL-30B-A3B remains faster on screenshots, while Qwen3-VL-32B remains an inexpensive specialist. For a new mixed text-and-vision deployment, however, the new Flash generation is the more capable starting point.
Our previous budget shortlist did not become wrong. The market simply moved faster than expected.
Qwen3.8 Flash and GLM-5.3 Flash offer better vision, much broader intelligence, modern sparse architectures, and roughly million-token context without leaving the inexpensive hosted tier.
For most new deployments, start with Qwen3.8 Flash.
While its launch discount remains active, GLM-5.3 Flash may be the best value available.