Exploratory benchmark · Qwen2.5 vs Qwen3 · Local Ollama run

Qwen2.5 7B vs Qwen3 4B & 8B: Local Writing Coach Benchmark.

We gave Qwen2.5 7B, Qwen3 4B and Qwen3 8B the same 20 paired writing-coach cases. Qwen2.5 7B and Qwen3 4B produced the same complete-case outcome on all 20 cases, while Qwen3 8B completed one additional case. The study also compares error localization, output discipline and cold-start speed on Windows.

Published August 15, 2026 · Reproducible reference-based evaluation · Local Windows and Ollama run

3local Qwen models
2Qwen generations
20paired writing cases
60completed responses

Why compare Qwen2.5 7B with Qwen3 4B and 8B?

The earlier LinguaPilot benchmark compared three sizes inside the Qwen3 family. This second study asks a different practical question: how does the previous-generation Qwen2.5 7B compare with a smaller Qwen3 4B model and the larger Qwen3 8B model in the same writing-coach workflow?

For local writing tools, model generation and parameter count do not automatically translate into a better practical result. Correction success, error localization, explanation-language compliance, structured output and local execution speed are therefore kept as separate dimensions.

Study scope: this is an exploratory, reference-based comparison of three local Qwen models on the same 20 English writing cases with explanations requested in French. It describes the differences observed in this run rather than presenting a universal model ranking.

How the benchmark was run

  1. Same 20 cases for every modelQwen2.5 7B, Qwen3 4B and Qwen3 8B received 16 sentences with expected errors and 4 already-correct controls.
  2. Same writing-coach taskEach model had to provide a minimally corrected text, an optional improved version, detected errors and explanations in French.
  3. Frozen reference comparisonOutputs were evaluated offline against versioned expected corrections and acceptable local operations. The models did not see the hidden reference.
  4. Paired case analysisThe primary outcome was complete-case success: every expected error corrected without an incorrect replacement, false error or configured fact loss.
  5. Local cold-start timingLoad, generation and total cold-start time were measured on the same Windows benchmark setup so responsiveness could be compared alongside correction behavior.

Main Qwen2.5 vs Qwen3 benchmark results

The most distinctive result is the match between Qwen2.5 7B and Qwen3 4B: both completed 18 of 20 cases, applied 19 of 21 expected corrections and reached 97.6% error-localization F1. Qwen3 8B completed 19 of 20 cases, applied 20 of 21 expected corrections and localized all 21 expected error regions in this run.

Reference-based results from the paired 20-case run
ModelComplete casesExpected correctionsError localization F1Mean cold-start time
Qwen3 4BSame paired outcome as 7B18/20 · 90%19/21 · 90.5%97.6%23.99 s
Qwen2.5 7B18/20 · 90%19/21 · 90.5%97.6%54.37 s
Qwen3 8B19/20 · 95%20/21 · 95.2%100%60.24 s
Complete-case correction comparison for Qwen2.5 7B, Qwen3 4B and Qwen3 8B
Complete-case correction: Qwen2.5 7B and Qwen3 4B both completed 18 of 20 cases. Qwen3 8B completed 19 of 20, a one-case observed advantage on the primary outcome in this run.

The result that makes this comparison different

Qwen2.5 7B and Qwen3 4B did not merely finish with the same 18/20 score. Their paired complete-case outcomes matched on every tested case: both succeeded on the same 18 cases and both failed on the same 2 cases. There were no discordant complete-case outcomes between these two models in this run.

Paired complete-case evidence behind the visible scores
Model pairBoth succeededBoth failedDifferent outcomesInterpretation
Qwen2.5 7B vs Qwen3 4B18/202/200/20Same complete-case outcome on all 20 paired cases
Qwen2.5 7B vs Qwen3 8B18/201/201/20One-case observed advantage for Qwen3 8B
Qwen3 4B vs Qwen3 8B18/201/201/20One-case observed advantage for Qwen3 8B

Practical interpretation: in this 20-case writing benchmark, moving from Qwen2.5 7B to Qwen3 4B preserved the same primary correction outcome case by case while reducing local cold-start time substantially. Qwen3 8B added one complete case and reached full error localization in the tested set.

What each model did well

Qwen3 4B

Fastest route to the same primary outcome as 7B

Matched Qwen2.5 7B case by case on complete-case success while running much faster locally.

  • 18/20 complete cases
  • 19/21 expected corrections
  • 23.99 s average cold-start time

Qwen2.5 7B

Strong previous-generation reference point

Matched Qwen3 4B on correction coverage and error localization, with clean explanation-language compliance in this run.

  • 18/20 complete cases
  • 97.6% error-localization F1
  • 54.37 s average cold-start time

Qwen3 8B

Highest correction coverage in this run

Completed one additional case and localized all expected error regions, with a cold-start time close to Qwen2.5 7B.

  • 19/20 complete cases
  • 20/21 expected corrections
  • 100% error-localization F1

Try the same local writing workflow in real Windows apps.

LinguaPilot lets you select your own text, use an Ollama model locally, and receive a corrected version, an improved version and an explanation without moving every draft into a browser chatbot.

Download the 14-Day Free Trial

Windows 10/11 · Local Ollama mode · No LinguaPilot account required

Local speed changes the practical picture

Qwen3 4B was the fastest model in this run at 23.99 seconds average total cold-start time. Qwen2.5 7B averaged 54.37 seconds, while Qwen3 8B averaged 60.24 seconds. That puts Qwen2.5 7B about 9.7% below Qwen3 8B on mean cold-start time, while Qwen3 4B was about 56% below Qwen2.5 7B.

Cold-start load generation and total response time for Qwen3 4B, Qwen2.5 7B and Qwen3 8B
Cold-start performance on the tested Windows setup. Qwen2.5 7B sat much closer to Qwen3 8B than to Qwen3 4B on local response time.

Error localization separates the 8B result

Qwen3 8B localized all 21 expected error regions with no false-positive region in this run. Qwen2.5 7B and Qwen3 4B each localized 20 of 21 expected regions, also with no false-positive region. Their localization F1 scores were therefore identical at 97.6%, compared with 100% for Qwen3 8B.

Error localization precision recall and F1 for Qwen3 4B, Qwen2.5 7B and Qwen3 8B
Error localization is shown separately from complete-case correction so detection discipline is not collapsed into one overall score.

Output discipline and explanation checks

All three models produced parseable JSON, usable core output and full-contract-compliant output in all 20 cases. The configured fact-preservation check also passed all 20 cases for every model.

For explanation-language compliance, Qwen2.5 7B passed all 40 evaluated fields and Qwen3 8B passed all 41. Qwen3 4B recorded 37 passes and 3 warnings. Among the 16 error-containing cases, the action–explanation consistency check recorded 16 consistent results for each model, with no contradictory, ambiguous or missing entries.

How to read these checks: they are useful supporting signals for an indicative writing-coach benchmark. They measure configured rules such as language compliance, structured output, contradictions and selected pedagogical patterns; they are not presented as a single pedagogical-quality score.

What this Qwen2.5 vs Qwen3 run suggests

  • Qwen3 4B is the standout efficiency result: it matched Qwen2.5 7B on every paired complete-case outcome while averaging less than half its cold-start time.
  • Qwen2.5 7B remains a credible writing baseline: it matched Qwen3 4B on complete cases, expected-correction coverage and error-localization F1.
  • Qwen3 8B provided the strongest correction coverage in this set: one additional complete case, one additional expected correction and 100% error-localization F1.
  • The speed gap between 7B and 8B was relatively modest: Qwen2.5 7B averaged about 9.7% lower cold-start time than Qwen3 8B in this run.
  • Structured output was not a differentiator here: all three models completed the full output contract in all 20 cases.

Methodology and scope

  • Exploratory design: 20 paired cases provide a focused comparison of this specific writing-coach scenario rather than a universal model ranking.
  • One language pair: English writing with explanations requested in French.
  • Two Qwen generations: the comparison is intentionally centered on Qwen2.5 7B, Qwen3 4B and Qwen3 8B rather than unrelated model families.
  • Model labels: the benchmark records qwen2.5:7b, qwen3:4b-q4_K_M and qwen3:8b-q4_K_M as the tested model identifiers.
  • One machine: local timing depends on hardware, background workload, Ollama configuration and model format; the values describe the tested setup.
  • Pedagogical checks: automatic checks are used as supporting indicators and are kept separate from the primary correction outcome.
  • Evaluation reference: responses were compared with the same frozen, versioned correction reference and acceptance rules across the three models.

From benchmark evidence to everyday writing

A local writing assistant has to balance more than correction accuracy. It also has to respond at a practical speed, preserve the user’s text, return predictable output and explain changes clearly enough to be useful.

LinguaPilot lets users choose an Ollama model that fits their own Windows hardware and writing needs, while keeping correction, improvement and explanation in the same desktop workflow.

Your next correction should teach you something.

Use LinguaPilot free for 14 days with the Ollama model already installed on your PC.

Download Free Trial