Most people learn about readability scores in school and promptly forget them. Then they start worrying about AI detection and someone mentions Flesch-Kincaid and the memory surfaces, vague and unhelpful.
Here's the clear version: what the main readability scores actually measure, what the numbers mean in practice, and why they matter more than ever now that AI detection tools have made readability a proxy for "did a human write this?"
The Main Readability Scores
Flesch Reading Ease
The oldest and most widely used formula. Developed by Rudolf Flesch in 1948, based on two variables: average sentence length and average number of syllables per word.
The formula: 206.835 – (1.015 × avg sentence length) – (84.6 × avg syllables per word)
The scale runs from 0 to 100, but inverted from what you might expect:
| Score | Description | Typical Audience |
|---|---|---|
| 90–100 | Very easy | 5th grade |
| 70–80 | Easy | 6th grade |
| 60–70 | Standard | 7th grade |
| 50–60 | Fairly difficult | High school |
| 30–50 | Difficult | College level |
| 0–30 | Very confusing | Professional/academic |
Most popular magazines aim for 60–70. Academic journals often land in the 20–40 range. A score of 45 isn't bad — it just means the writing is complex.
What moves this score: Long sentences drag it down. Long words drag it down. Short sentences and short words push it up. That's the whole formula.
Flesch-Kincaid Grade Level
Same underlying variables as Flesch Reading Ease, different formula, different output. This version gives you a US school grade level rather than a 0–100 score.
The formula: (0.39 × avg sentence length) + (11.8 × avg syllables per word) – 15.59
A score of 8 means an 8th grader should be able to read it. A score of 12 means high school senior. A score of 16 means college level.
This is the one embedded in Microsoft Word, Google Docs' readability statistics, and most writing tools. It's the one educators reference most often.
For blog content: aim for Grade 7–9. For academic writing: Grade 10–14 depending on the discipline. For legal or technical writing: Grade 12–16.
Dale-Chall Readability Formula
The Dale-Chall formula takes a different approach. Instead of counting syllables, it compares your words against a list of 3,000 "familiar words" — the words that 80% of 4th graders know.
Words not on the list count as "difficult." The formula penalizes text that uses lots of unfamiliar vocabulary, regardless of syllable count.
The output is a grade-level equivalent:
| Score | Grade Level |
|---|---|
| 4.9 or below | Grade 4 or below |
| 5.0–5.9 | Grades 5–6 |
| 6.0–6.9 | Grades 7–8 |
| 7.0–7.9 | Grades 9–10 |
| 8.0–8.9 | Grades 11–12 |
| 9.0–9.9 | College level |
| 10+ | Professional |
Dale-Chall is considered more accurate for texts using specialized or technical vocabulary — it distinguishes between "long words that educated people know" and "genuinely unfamiliar words." A chemistry paper about molecular spectroscopy will score differently on Dale-Chall vs Flesch-Kincaid, even with similar sentence lengths.
Gunning Fog Index
Developed by Robert Gunning in 1952 for business writing. The formula: 0.4 × (avg sentence length + percentage of complex words), where "complex words" means words with three or more syllables, excluding proper nouns and easy compound words.
Output is grade level. A Fog Index of 12 means a 12th grader can comfortably read it.
The name comes from Gunning's idea that business prose was unnecessarily "foggy" — complex sentences and polysyllabic words obscuring simple ideas. He developed the index to help writers cut through that fog.
Target zones: Below 12 for general audiences. Below 8 for broad popular appeal. Academic and professional writing routinely runs 15–18 without this being a problem.
SMOG Index
SMOG stands for Simple Measure of Gobbledygook. Developed in 1969 by G. Harry McLaughlin, it's one of the most accurate formulas for predicting reading comprehension on health and medical materials.
The SMOG calculation requires a 30-sentence sample and counts polysyllabic words (3+ syllables). The formula: 3 + √(count of polysyllabic words × (30 / sentence count))
SMOG tends to give slightly higher grade level estimates than Flesch-Kincaid for the same text. Researchers in health literacy consider it more reliable — studies show it predicts comprehension difficulty better than most alternatives for medical patient education materials.
For general use, Flesch-Kincaid is more practical. For healthcare, government communications, or anything where comprehension at a specific reading level is critical, SMOG is worth adding.
The AI Detection Connection
Here's where readability scores become something more than a writing quality metric.
AI-generated text has a characteristic readability profile. Specifically: AI writing shows very low variance in readability scores across paragraphs.
Run a Flesch-Kincaid Grade Level analysis on each paragraph of a GPT-5 essay. You'll typically see scores clustering around 10–12, with few paragraphs varying more than 2–3 grade levels from the average. The model produces consistently complex, formal prose.
Human writing doesn't work that way. Real writers vary dramatically. A researcher might explain a concept at Grade 7 and then immediately dive into technical detail at Grade 14. An essayist might open a paragraph with a one-sentence punch (Grade 4) and follow it with a long, subordinate-clause-heavy sentence (Grade 16). That variance is natural. It's hard to fake.
This means readability variance is a signal that AI detection tools use, even if they don't always name it explicitly. A document with a consistent Grade 11 readability across all paragraphs triggers detection algorithms differently than a document swinging between Grade 6 and Grade 15.
TextSight's readability analysis is part of the Humanization Score precisely because readability patterns are informative. A very flat readability profile — everything at the same grade level, paragraph after paragraph — is one of the statistical signatures that lowers your score.
What AI Writing Looks Like on These Scales
Based on analysis of GPT-4o and GPT-5 outputs across a range of writing tasks:
Flesch Reading Ease: Typically 35–50. AI writes formal, moderately complex prose by default. It rarely writes very easy (above 70) unless explicitly prompted to simplify.
Flesch-Kincaid Grade Level: Typically 9–13. The model tends toward high school to early college register.
Variance across paragraphs: Low. The standard deviation of grade level scores across paragraphs in AI text is usually 1.5–3. In human academic writing, it's typically 3–6.
Gunning Fog: Typically 11–14.
None of these numbers by themselves identify AI writing. A human writing a formal essay might land in the same ranges. The flag is the combination of where the scores cluster and how little they vary.
How to Use Readability Scores Practically
For writers concerned about AI detection
Intentionally vary your sentence structure. Write a long, complex sentence with multiple clauses. Then write a short one. Three words. The contrast is natural for human writers and unusual for AI output. It also lowers your readability variance in ways that detection algorithms register.
Mix technical vocabulary with plain language. Don't sustain formal register for entire documents. Even in academic writing, the clearest writers drop into plain language for key ideas before returning to technical vocabulary. That register-switching is a human tell.
Let some paragraphs be simple. Not every paragraph needs to be Grade 12. A paragraph that opens with a clear, short claim at Grade 7 and then develops at Grade 12 is both better writing and harder to flag.
Check TextSight's readability feedback alongside the Humanization Score. The readability analysis tells you something specific about your writing that the score alone doesn't — it identifies whether your prose is varying enough to read as human.
For educators evaluating student work
A student whose readability scores are suspiciously consistent — every paragraph within one grade level of every other — has a document worth examining more closely. It's not proof of anything. It's a pattern.
Compare the out-of-class submission to any in-class writing samples. Readability variance should be similar. A student who writes at Grade 7 in class and submits a Grade 11 paper with low variance is worth a conversation.
For content teams and SEO writers
Different platforms have different optimal readability targets. Blog content: Grade 7–9. White papers: Grade 11–13. Email newsletters: Grade 6–8. Product pages: Grade 6–8.
If your AI-drafted content is consistently landing at Grade 12 when your target is Grade 8, the readability scores are telling you something directly useful — edit toward simpler sentences and more common vocabulary, which also happens to look more human.
The Practical Checklist
Before you submit or publish anything:
-
Run Flesch-Kincaid Grade Level — aim for your target audience level.
-
Check paragraph-level variance — if you're using a tool that shows paragraph-by-paragraph scores, look for the range. A range of less than 3 grade levels across all paragraphs is a signal. Most editing tools, including Hemingway App, will give you an overall score; for paragraph-level analysis, paste sections individually.
-
Look at the Hemingway App sentence highlighting — the red sentences (very hard to read) and yellow sentences (hard to read) indicate Grade 12+ complexity. Having some of these is fine. Having only these is a readability variance problem.
-
Use TextSight's readability analysis alongside the Humanization Score — the two signals together are more informative than either alone. Try it at textsight.ai.
-
Rewrite for variety, not just simplicity — the goal isn't to lower your grade level across the board. It's to create variance. Mix simple and complex. Mix short and long. That's what human writing looks like.
One More Thing
Readability scores are tools, not verdicts. A Flesch-Kincaid Grade Level of 14 doesn't mean your writing is bad — it might mean your audience is advanced and your subject is complex. A score of 6 doesn't mean your writing is good — it might just mean you're using short words.
The goal is always writing that serves the reader. Readability metrics help you check whether it does. The AI detection angle is real and worth knowing, but the core value of readability analysis is simpler: it tells you whether your writing is appropriately calibrated for the people who are going to read it.
Get that right, and the AI detection scores tend to follow.
Related reading:
- How to Write a Research Paper That Doesn't Read Like AI
- How to Humanize ChatGPT Text
- 5 Free Tools Every Student Needs in 2026
- An Honest Review of the AI Detector Market in 2026
Try it on your own writing