Guides Development

How Text Analysis and Editing Tools Work

How words and reading time are counted, Flesch readability, naming-case styles, line diffs, sorting, de-duplication and escaping for JSON, regex and SQL.

Last reviewed:

AdSense Placeholder
Slot: header_reference_page
On this page

Most text tools do simple things, but a few details decide whether they give you the answer you expected: how a "word" is counted, how readability is scored, what makes two lines "the same", and which characters need escaping. This guide explains those, with every example run through the site's own tools.

Counting words and time

A word is a run of letters or digits, optionally joined by an apostrophe or hyphen (don't, well-known are one word each). Sentences end at ., !, ? or a blank line. From a word count you get time: silent reading averages about 238 words a minute and speaking about 150, so 1,000 words take roughly 4.2 minutes to read. Use the Word Counter, which also counts characters, and the Word Frequency Counter to see which terms repeat: in the cat and the hat and the bat, "the" is 37.5% of all words and "and" 25%.

Scoring readability

The Flesch Reading Ease score uses only sentence length and syllables per word:

206.835 − 1.015 × (words / sentences) − 84.6 × (syllables / words)

Higher is easier: 90 to 100 is very easy, 60 to 70 plain English, 30 to 50 difficult and below 30 very difficult. The Flesch-Kincaid grade turns the same two numbers into a US school grade: 0.39 × (words / sentences) + 11.8 × (syllables / words) − 15.59. Two real runs:

Text Words Sentences Syllables Reading Ease Grade
"The cat sat on the mat. It was a sunny day, and the cat was happy." 16 2 18 103.5 (very easy) 0.8
"The committee reviewed the proposal and approved the revised budget." 10 1 18 44.4 (difficult) 9.6

Short words and short sentences raise the score. Syllables are counted by an approximation, so treat scores as guidance. The Readability Score Checker reports several indexes at once. Aim for plain language: shorter sentences and everyday words.

Changing case

The same phrase in the common naming styles, handy for code and URLs:

Style Result Typical use
camelCase wordFrequencyCounter JavaScript variables
PascalCase WordFrequencyCounter Class names
snake_case word_frequency_counter Python and database columns
kebab-case word-frequency-counter URLs and CSS classes
SCREAMING_SNAKE_CASE WORD_FREQUENCY_COUNTER Constants
Title Case Word Frequency Counter Headlines
Sentence case Word frequency counter Ordinary text

The Case Converter converts between them, and for line breaks use the Remove Line Breaks (note that Windows ends lines with two characters, \r\n, while Linux and macOS use \n).

Comparing texts

A line-based diff finds the longest common subsequence of lines, then reports the rest as removed or added. Comparing three lines with a revised four-line version:

Old text New text Diff
red green blue red yellow blue black red
− green
+ yellow
blue
+ black

That is 2 added, 1 removed and 2 unchanged. A changed line counts as one removal plus one addition. The Text Diff Checker shows this side by side; for structured data use the JSON Diff Checker.

Sorting and removing duplicates

  • Sort order matters. Plain alphabetical order puts 10 before 2: the result is 1, 10, 2. Natural order compares numbers by value and gives 1, 2, 10. Use the Sort Lines with the mode you need.
  • Duplicates depend on case. Of a, b, a, A, the Remove Duplicate Lines keeps 3 lines, a, b, A, but when case is ignored only 2 remain, a, b.
  • Pull out what you need. The Text Extractor lists the emails, URLs, phone numbers or numbers in a block of text.

Escaping special characters

Characters like quotes and backslashes mean something in code, so they must be escaped when they are data. The String Escape & Unescape handles each language:

Language Input Escaped
JSON He said "hi" plus a newline He said \"hi\"\n
Regex a.b*c a\.b\*c
SQL O'Brien O''Brien

Escaping SQL by hand is a last resort: use parameterised queries, which prevent SQL injection entirely. Escape for the context you are in; HTML needs entities instead, see HTML Entities Cheat Sheet.

Placeholder text

Lorem ipsum is dummy text derived from a work by Cicero, used to fill layouts without distracting with meaning. Generate any amount with the Lorem Ipsum Generator, and replace it with real copy before launch.

Common mistakes

  • Trusting a readability score for non-English text. The formulas were tuned for English.
  • Sorting numbers as text. Use natural or numeric sort.
  • Ignoring invisible differences. Trailing spaces and different line endings make "identical" lines differ.
  • Escaping for the wrong language. A JSON-escaped string is not safe for a regex or SQL.

Try these tools

See also

  • Guide How JSON Works
    The six JSON value types, why parsing fails, traps with big integers and key order, JSONPath queries and converting JSON to TypeScript.
  • Cheat sheet HTML Entities Cheat Sheet
    The HTML entities you actually use, with name, decimal code and meaning, which five characters to escape and how entities are decoded.
  • Cheat sheet Regular Expressions (Regex) Cheat Sheet
    Regex syntax in one page: character classes, anchors, quantifiers, groups, lookaround, flags and ready-to-use patterns.
  • Glossary Regular expression
    A regular expression (regex) is a pattern written in a compact syntax that describes a set of strings, used to search.
  • Glossary SQL injection
    SQL injection is an attack in which input is crafted so that part of it is executed as database commands.
  • Glossary Lorem ipsum
    Lorem ipsum is dummy Latin-like text used to fill page layouts and designs so that reviewers focus on the design rather than the words.

Frequently Asked Questions

For general audiences, 60 to 70 (plain English) is a good target. Technical or legal writing often scores lower; the score is a guide, not a rule.

AdSense Placeholder
Slot: footer_leaderboard