Reference based
Models are compared against a frozen correction reference
The evaluated models do not see the expected corrections. Their outputs are compared offline with the same versioned cases and acceptance rules.
Research & benchmarks
A growing collection of transparent, reference-based studies on local AI writing models, multilingual feedback and real Windows performance. The goal is to provide practical evidence about the strengths, trade-offs and differences observed between models in controlled writing-coach scenarios.
Latest study
Three local Qwen3 models received the same 20 writing-coach cases: correct English text, identify the real errors and explain the changes in French. The study also measured cold-start execution with Ollama on Windows.
Exploratory local benchmark
See complete-case correction, expected-correction coverage, error localization, explanation-language compliance and cold-start performance—with the limits shown beside the results.
Research principles
Reference based
The evaluated models do not see the expected corrections. Their outputs are compared offline with the same versioned cases and acceptance rules.
Multidimensional
A fast model is not automatically a better coach. Correction success, language compliance, structured output and local speed are shown as different dimensions.
Paired cases
Pairing helps reveal whether a visible score difference comes from many cases or from one isolated example.
Limits visible
Each study identifies its dataset, language pair, hardware and execution conditions so readers can understand where the findings are most useful.
From research to daily writing
LinguaPilot lets you select text in almost any Windows app, receive a corrected version, an improved version and an explanation, and choose Ollama for local processing or an optional cloud provider for other tasks.