High-Frequency Words (HFWs)

Comprehensive Definition

High-Frequency Words are the words that occur most often in written and spoken English. They are identified through corpus analysis, in which large bodies of text are counted, ranked, and aggregated to produce ordered lists of the most common words. The concentration is striking. The single most common word in English, "the", accounts for around six to seven per cent of all words in typical written text. The first hundred High-Frequency Words make up roughly half of everything a reader will encounter, and the first three hundred make up around sixty-five per cent. This skewed distribution follows a broader pattern in language known as Zipf's law, in which a small number of words appear extremely often and the vast majority appear rarely.

This frequency profile has direct instructional consequences. Because so few words carry so much of the reading load, the return on securing automatic recognition of these words is exceptionally high. A student who can recognise the first hundred High-Frequency Words instantly has, in effect, made half of every page in front of them effortless, freeing cognitive resources for the content words that carry meaning. A student who cannot recognise these words automatically will encounter friction on every line of every text.

For instructional purposes, High-Frequency Words are grouped according to how their spellings relate to phonic knowledge rather than by frequency alone. Decodable High-Frequency Words, such as "and", "it", "had", "can", and "but", follow the regular phoneme-grapheme correspondences taught in early Systematic Synthetic Phonics and are taught through phonics like any other decodable word. Irregular High-Frequency Words, such as "said", "was", "the", and "of", contain one or more graphemes that do not follow the most common pattern or that have not yet been taught at the student's stage. These are taught using the Heart Word method, in which the regular parts of the word are decoded as usual and the irregular grapheme is identified and held in memory by heart. A central principle is that even irregular words usually contain regular parts, so the visual memory load is far smaller than treating the word as a whole shape implies.

Practical Example

The Dolch list, compiled in 1936, and the Fry list, developed from the 1950s onwards, are two of the most widely used High-Frequency Word lists internationally. In Australia, the Oxford Wordlist is commonly used and is derived from the writing of Australian children in the early years of school. Different lists rank words slightly differently because they draw on different source corpora, including children's reading material, adult writing, or spoken language, but the most frequent words are remarkably consistent across them. PLD's Systematic Synthetic Phonics Kits sequence High-Frequency Words deliberately, introducing decodable High-Frequency Words through phonics as their corresponding patterns are taught, and introducing irregular High-Frequency Words as Heart Words in small sets at appropriate stages.

Common Misconception

Misconception: All High-Frequency Words are irregular and must be memorised visually.

Correction: Frequency and regularity are independent properties. Many High-Frequency Words are entirely decodable using early phonic knowledge, including "and", "it", "had", "can", "in", "on", and "but". These should be taught through phonics, not flashcards. Even the words that are genuinely irregular almost always contain regular elements. In "said", "s" and "d" follow the expected correspondences, leaving only the "ai" to be learned by heart. Treating High-Frequency Words as a single block of words to be visually memorised ignores the structure within them and wastes the phonic information already available to the student.

Frequently Asked Questions

Q: What is the goal of High-Frequency Word instruction? A: Automatic recognition of the words that carry the largest share of reading volume. Because the first few hundred High-Frequency Words make up the majority of any text, instant recognition of these words removes the most pervasive source of friction in reading. The goal is not to add another category of word to the curriculum but to ensure that the words students will see thousands of times become effortless as early and as durably as possible.

Q: What is the difference between a High-Frequency Word and a sight word? A: High frequency is a property of language, established by counting how often a word appears across large samples of text. A sight word is a property of an individual reader, referring to any word that reader recognises instantly without conscious decoding. The terms are often used interchangeably in classrooms, but they refer to different things. The instructional goal is for High-Frequency Words to become sight words for every reader, achieved through phonic decoding and orthographic mapping rather than through memorisation of word shapes.

Q: Should High-Frequency Words be taught instead of phonics? A: No, and the question itself reflects a misunderstanding that has caused real harm in early literacy. High-Frequency Words are taught alongside phonics, integrated into the same scope and sequence. Front-loading a large bank of memorised whole words before students have phonic knowledge encourages guessing strategies and bypasses the very process that builds durable reading. Within a structured literacy framework, decodable High-Frequency Words are taught through their phonic structure, and irregular ones are taught as Heart Words in small sets as they are needed for connected text.

Q: Where do High-Frequency Word lists come from, and which list should a school use? A: All major High-Frequency Word lists are derived from corpus analyses, but they differ in which corpus they draw on. The Dolch list is based on children's reading materials from the early twentieth century. The Fry list is based on a broader sample of written English and is updated periodically. The Oxford Wordlist is based on Australian children's writing in the early years of school and is therefore particularly well suited to the Australian context. The choice of list matters less than choosing one and sequencing it deliberately within the program's broader scope and sequence. PLD's resources draw on these established lists while organising the words to align with the program's phonic stages.

Q: Why do frequency counts vary between lists? A: Because they draw on different corpora. A list based on children's storybooks will rank words such as "said", "little", and "look" higher than a list based on adult news writing, which will rank words such as "government", "however", and "company" higher. The most common function words, including "the", "of", "and", "to", and "in", appear at the top of every list regardless of corpus, but the rankings further down diverge. This is why the choice of corpus matters, and why an Australian list based on Australian children's writing has practical advantages for Australian classrooms.

Q: Why are these particular words so frequent? A: The words at the top of every High-Frequency Word list are overwhelmingly function words: articles ("the", "a"), prepositions ("of", "in", "to"), pronouns ("I", "you", "it"), conjunctions ("and", "but"), and auxiliary verbs ("is", "was", "have"). Function words appear constantly because they hold sentences together, regardless of the topic of the text. Content words, by contrast, vary with subject matter and are spread thinly across a much larger vocabulary. This is also why so many High-Frequency Words appear irregular: many function words derive from Old English and have retained spelling conventions that no longer match contemporary phonic patterns.

Related Terms
Related Products