Qwen3.8 Flash and GLM-5.3 Flash reset the budget multimodal tier.

EvalSignal 008 / Independent model evaluation

Two new Flash models just reset the budget multimodal tier

A newsletter issue is not supposed to age this quickly.

Five days ago, we published our guide to inexpensive Qwen models for large vision workloads. The results were useful and accurate. They are also already becoming obsolete.

Two new models arrived within hours of each other:

Both are more modern and substantially more capable than the models in our previous shortlist. They combine strong visual understanding with useful text and tool-planning intelligence, support text, images, and video, and offer roughly million-token context windows.

They remain firmly in the budget API tier. During its launch promotion, GLM-5.3 Flash is even cheaper than every model in our previous comparison.

What we tested

We ran each model through two frozen 100-case suites:

Every model received the same prompts, screenshots, tool definitions, temperatures, and scoring rules.

That produced 400 final case results with zero errors.

The planning suite does not measure every form of intelligence or execute the complete task. It measures whether a model can interpret an instruction and choose a sensible first action. The vision suite requires important visual facts to be identified without critical contradictions.

The results

Model Pass Rubric Vision Text
Qwen3.8 Flash 79 94.8% 8.42s 2.55s
GLM-5.3 Flash 76 93.8% 5.49s 7.24s

Qwen3.8 Flash won vision quality. It recorded three more strict passes, performed better on the most challenging cases, and was especially strong on authentication screens.

GLM-5.3 Flash won vision latency. It processed screenshots roughly three seconds faster at the median while finishing only one rubric point behind Qwen.

The text-planning result was more nuanced:

Qwen is therefore the faster planner. GLM is slightly more consistent about emitting structured tools and following the expected initial route.

Why the previous shortlist is already outdated

Model Pass 100-case cost
Qwen3.8 Flash 79 $0.041
GLM-5.3 Flash 76 $0.021 promo
Qwen3-VL-32B Instruct 69 $0.022
Qwen3-VL-30B-A3B Instruct 68 $0.024

Qwen3.8 Flash costs roughly two cents more per 100 screenshots than the previous Qwen leaders, but delivers ten additional strict passes and much stronger general-purpose capabilities.

GLM-5.3 Flash currently costs less than the old models while beating the previous best result by seven passes. However, its price is temporarily discounted by 50% through September 9. Its normal pricing is almost identical to Qwen3.8 Flash.

Large models with small active paths

Neither model is genuinely small.

Qwen3.8 Flash has a 125B-parameter main model but activates only 6B parameters per token. GLM-5.3 Flash stores 320B parameters while activating 18B per token.

This sparse architecture lets the models retain far more capacity than yesterday's 8B and 30B budget models without performing dense computation across every parameter for every generated token.

The result is a new class of API model: enormous total capacity, relatively inexpensive inference, million-token context, and credible performance across both vision and text.

Which one should you choose?

Best overall default: Qwen3.8 Flash
Choose Qwen when vision accuracy and fast text planning matter most. It has the stronger difficult-case floor and the best overall result in this comparison.

Best value during the promotion: GLM-5.3 Flash
Choose GLM when screenshot latency and immediate cost matter most. Its vision quality is close to Qwen, and its temporary price is exceptional.

After the promotion
Treat their token prices as approximately equal. Choose Qwen for vision hit rate and text speed. Choose GLM for screenshot speed and more consistent structured tool output.

Keep an older Qwen3-VL model only when one narrow metric dominates
Qwen3-VL-30B-A3B remains faster on screenshots, while Qwen3-VL-32B remains an inexpensive specialist. For a new mixed text-and-vision deployment, however, the new Flash generation is the more capable starting point.

Bottom line

Our previous budget shortlist did not become wrong. The market simply moved faster than expected.

Qwen3.8 Flash and GLM-5.3 Flash offer better vision, much broader intelligence, modern sparse architectures, and roughly million-token context without leaving the inexpensive hosted tier.

For most new deployments, start with Qwen3.8 Flash.

While its launch discount remains active, GLM-5.3 Flash may be the best value available.

Sources

Previous EvalSignal comparison

Full benchmark, methodology, and architecture analysis