Research & benchmarks

Research for AI writing tools that should help people learn.

A growing collection of transparent, reference-based studies on local AI writing models, multilingual feedback and real Windows performance. The goal is to provide practical evidence about the strengths, trade-offs and differences observed between models in controlled writing-coach scenarios.

Latest study

Qwen2.5 7B vs Qwen3 4B & 8B writing coach benchmark

Three local Qwen models from two generations received the same 20 writing-coach cases. The most distinctive result: Qwen2.5 7B and Qwen3 4B produced the same complete-case outcome on every paired case, while Qwen3 8B completed one additional case.

Exploratory local benchmark

What changes when Qwen2.5 7B meets Qwen3 4B and 8B in the same writing workflow?

Compare paired correction outcomes, expected-correction coverage, error localization, explanation-language compliance, structured output and cold-start performance on Windows.

Qwen2.5 vs Qwen320 paired cases60 completed responsesOllama on Windows

Previous study

Qwen3 writing coach benchmark: 4B vs 8B vs 14B

The first LinguaPilot benchmark compared three Qwen3 sizes on the same 20 reference-based writing cases and showed how model size affected correction coverage and local cold-start performance.

Qwen3 model-size benchmark

Which Qwen3 size offered the best practical balance for writing feedback in that run?

See complete-case correction, expected-correction coverage, error localization, explanation-language compliance and cold-start performance across Qwen3 4B, 8B and 14B.

3 Qwen3 models20 paired cases60 completed responsesOllama on Windows

Research principles

Practical evidence for informed local-model choices.

Reference based

Models are compared against a frozen correction reference

The evaluated models do not see the expected corrections. Their outputs are compared offline with the same versioned cases and acceptance rules.

Multidimensional

Correction, explanation and execution are separated

A fast model is not automatically a better coach. Correction success, language compliance, structured output and local speed are shown as different dimensions.

Paired cases

The same cases are used for every model

Pairing helps reveal whether a visible score difference comes from many cases or from one isolated example.

Scope visible

Results are interpreted within the tested setup

Each study identifies its dataset, language pair and execution context so readers can understand where the findings are most useful.

From research to daily writing

Use the local model that fits your own Windows setup.

LinguaPilot lets you select text in almost any Windows app, receive a corrected version, an improved version and an explanation, and choose Ollama for local processing or an optional cloud provider for other tasks.