How to Create Reading Comprehension Assessments: 2026 Guide

TL;DR
Creating reading comprehension assessments requires understanding the different assessment types (screening, diagnostic, formative, summative, progress monitoring), choosing the right format for what you actually want to measure, and calibrating question difficulty using frameworks like Bloom’s Taxonomy and Depth of Knowledge. This glossary defines every term teachers need, explains when and why each matters, and offers practical guidance for building assessments that improve instruction rather than just generating grades.
Two out of five American fourth graders scored below Basic proficiency on the 2024 NAEP reading assessment. That number isn’t a rounding error. It means roughly 40% of nine-year-olds in the United States cannot demonstrate even the most fundamental reading comprehension skills expected for their grade.
At the twelfth-grade level, the picture doesn’t improve much. Only 35% of seniors performed at or above the NAEP Proficient level in 2024, down five percentage points from 1992.
These numbers make one thing clear: the way we assess reading comprehension isn’t just an academic exercise. Assessment drives instruction. When teachers create reading comprehension assessments that accurately reveal what students can and cannot do, they gain the data needed to actually close gaps. When they rely on mismatched or shallow assessments, gaps widen.
This glossary breaks down every term K-12 educators need to create, select, and interpret reading comprehension assessments with confidence. It’s organized thematically so each section builds on the last.
Create assessments in minutes with TeachTools’ AI quiz generator, built for K-12 teachers.
Types of Reading Comprehension Assessments
Before picking a format or writing a single question, teachers need to know which type of assessment fits their purpose. These five types aren’t mutually exclusive. Effective reading programs use all of them at different points in the school year.
Screening Assessment
A brief, standardized test given to all students, typically three times per year, to identify who may be at risk for reading difficulty. Think of it as a net that catches students who need closer attention before they fall further behind.
Why it matters: Screening catches problems early. Beginning in the 2025-2026 school year, states like Maryland and Missouri are mandating approved universal screeners grounded in the science of reading. Missouri law now requires all public and charter schools to use a state-approved screening assessment.
Common confusion: Screening doesn’t diagnose the problem. It flags that a problem likely exists. A student who fails a screener needs follow-up diagnostic assessment, not an immediate intervention plan.
Diagnostic Assessment
Delivered before or early in instruction to pinpoint specific gaps in prerequisite knowledge. While a screener tells you a student is struggling with reading, a diagnostic assessment tells you why: is it decoding, vocabulary, fluency, or comprehension itself?
Why it matters: Without diagnosis, intervention is guesswork. A student who struggles to decode multisyllabic words needs a different instructional response than one who decodes fluently but can’t make inferences.
Related terms: Screening assessment, progress monitoring
Formative Assessment
Frequent, low-stakes assessments embedded within instruction. Formative assessments in reading comprehension take many forms: asking students to retell what they read, engaging in text discussions, responding to short written prompts, or completing quick multiple-choice checks.
Why it matters: Formative assessment creates a feedback loop. It’s not about grading; it’s about adjusting instruction in real time. If half the class can’t identify the main idea of Tuesday’s passage, Wednesday’s lesson should address that directly.
Teachers looking to create formative assessments quickly can save significant prep time by using templates and AI-assisted tools rather than building everything from scratch.
Common confusion: Formative assessment is often confused with “practice.” The distinction is that formative assessment requires the teacher to analyze responses and use them to inform next steps, not just assign a completion grade.
Summative Assessment
An evaluation given at the end of a unit, grading period, or course to measure what students have learned. Summative reading comprehension assessments cover more content than formative checks and carry more weight in grading.
Why it matters: Summative data feeds into report cards, school accountability, and program evaluation. But summative assessments are backward-looking. They tell you what happened, not what to do next.
Related terms: Formative assessment, NAEP achievement levels
Progress Monitoring
Repeated measurement, often biweekly, tracking whether a student is responding to an intervention or instructional change. Progress monitoring typically uses standardized tools like maze passages or oral reading fluency measures.
Why it matters: Progress monitoring answers the question every teacher and interventionist needs answered: “Is what I’m doing working?” Without it, students can spend months in an ineffective program before anyone notices.
Understanding these types is foundational. When teachers create reading comprehension assessments, the first decision is always purpose: are you screening, diagnosing, checking understanding, measuring outcomes, or tracking growth?
Assessment Formats and Methods
The format you choose determines what depth of comprehension you actually test. This is one of the most overlooked aspects of assessment design. Each method below has strengths and blind spots.
Multiple-Choice Questions (MCQs)
The most common format on standardized reading tests. Students read a passage and select the best answer from several options.
Why it matters: MCQs are efficient to administer and score, which makes them practical for large groups. When constructed well, they can assess deep comprehension, including inferential and evaluative thinking, not just literal recall.
Common confusion: Many teachers assume MCQs only test surface-level understanding. That’s true of poorly written MCQs, but a well-crafted question with plausible distractors can require genuine reasoning. The quality of the question matters more than the format.
For guidance on building effective quiz questions at different difficulty levels, see this guide on customizing quiz difficulty across grade bands.
Constructed Response / Short Answer
Students write a brief answer to a question about a passage. This format requires students to generate rather than recognize the correct response.
Why it matters: Constructed responses reveal thinking in ways MCQs cannot. A student might select the right multiple-choice answer through elimination, but a written response shows whether they truly understood the text.
Important caveat: Unlike multiple-choice questions, short open-ended questions require a verbal response, which may put certain groups of students at a disadvantage, including students with disabilities and English learners. Teachers should consider whether they’re assessing reading comprehension or writing ability.
Cloze Test
A passage with regularly deleted words where students supply the missing word based on context.
Why it matters: Cloze tests are quick to create and administer. They measure a student’s ability to use context clues and syntactic knowledge.
Critical limitation: Research from literacy experts like Timothy Shanahan shows that cloze tests correlate best with comprehension questions based on single sentences. Correlations drop when students need to synthesize information across an entire passage. In other words, cloze tasks test sentence-level processing more than passage-level comprehension. Teachers should not treat cloze results as a complete picture of reading comprehension.
Maze Task
A multiple-choice variant of the cloze test. The first sentence remains intact, then every seventh word is replaced with three options (one correct, two distractors). Students typically have three minutes to complete the task, and their score is the number of correct choices.
Why it matters: Maze tasks are a staple of progress monitoring programs like Acadience Reading. They’re fast, standardized, and easy to score across an entire class.
Same limitation as cloze: Maze tasks provide reasonable predictions of reading comprehension, but they do so based on sentence-level interpretation. They shouldn’t be the only comprehension measure in a teacher’s toolkit.
Retelling / Oral Retell
A student reads or listens to a passage, then retells the story or information in their own words.
Why it matters: Retelling assesses whether a student can organize and recall key information without the scaffolding of answer choices. It reveals sequence understanding, main idea grasp, and the ability to distinguish important details from minor ones.
Best for: K-3 students and students whose writing skills don’t yet match their comprehension ability.
Running Record
A teacher listens to a student read aloud and marks errors, self-corrections, and patterns in real time. Common in guided reading, primarily used in K-3.
Why it matters: Running records capture not just accuracy but reading behaviors, revealing whether a student self-monitors, uses phonics strategies, or relies too heavily on guessing.
Common confusion: Running records assess reading accuracy and fluency, not comprehension directly. Pair them with a retell or comprehension questions for a fuller picture.
The format you select when you create reading comprehension assessments should match the skill you’re trying to measure. Use maze tasks when you need to screen 30 students in five minutes. Use constructed responses when you need to assess inferential thinking in depth.
Reading Frameworks That Guide Assessment
Why do these theoretical models matter for assessment? Because they tell teachers what to test for. Without a framework, assessment becomes a grab bag of random questions. These three models shape modern reading instruction and, by extension, modern reading assessment.
Simple View of Reading (SVR)
Proposed by Gough and Tunmer in 1986, the Simple View of Reading is expressed as a formula: Reading Comprehension = Decoding × Language Comprehension. Both components are necessary. If either is zero, reading comprehension is zero.
Why it matters for assessment: SVR tells teachers that a student struggling with comprehension might have a decoding problem, a language comprehension problem, or both. Assessment needs to distinguish between these causes. A student who can’t decode the words on the page will fail a reading comprehension test regardless of how strong their reasoning skills are.
Scarborough’s Reading Rope
Developed by Hollis Scarborough in 2001, the Reading Rope is a visual model showing how skilled reading weaves together two strands: word recognition (phonological awareness, decoding, sight recognition) and language comprehension (background knowledge, vocabulary, language structure, verbal reasoning, literacy knowledge). When all strands are tightly woven, proficient reading emerges. When even one strand is weak, the whole rope is compromised.
Why it matters for assessment: The Reading Rope gives teachers a checklist of sub-skills to assess. It moves beyond “can they answer questions about a passage” to “which specific strand is breaking down?” Teachers who align assessments to learning objectives grounded in these sub-skills get far more actionable data.
Science of Reading
The science of reading is a body of research, built over decades using rigorous methodologies, showing that learning to read requires explicit, systematic, and cumulative instruction. It is not a curriculum or a program. It’s the evidence base that informs which curricula and programs work.
Why it matters for assessment: The science of reading movement is reshaping assessment requirements at the state level. States are mandating approved screeners. The Reading League’s 2026 Summit theme, “From Confusion to Clarity: Turning Data Into Instructional Impact,” reflects the field’s current focus on using assessment data effectively, not just collecting it.
Teachers creating comprehension assessments should ensure they’re measuring skills the research says matter, not just whatever happens to be easy to test.
Cognitive Complexity Frameworks for Writing Questions
Explore 23+ free AI tools for teachers
Browse All Tools →Here’s a problem most teachers recognize intuitively but may not have quantified: studies of textbook comprehension questions found that questions clustered heavily at lower cognitive levels, with 26% at the comprehension level and 17% at the knowledge/remember level. That means a large share of pre-packaged questions test only recall and basic understanding.
When you create reading comprehension assessments from scratch, you have the opportunity to do better. Two frameworks help.
Bloom’s Taxonomy (Revised)
One of the most widely known cognitive frameworks in education. The revised version organizes thinking into six levels of increasing complexity:
- Remember — Recall facts or basic concepts (Who is the main character?)
- Understand — Explain ideas or concepts (What does the author mean by…?)
- Apply — Use information in new situations (How would this character respond if…?)
- Analyze — Draw connections among ideas (What evidence supports the author’s claim?)
- Evaluate — Justify a stand or decision (Is the narrator reliable? Why or why not?)
- Create — Produce new or original work (Write an alternative ending that stays consistent with the character’s motivations)
Practical guidance: When building a reading comprehension quiz or test, aim for a deliberate distribution across levels. A good rule of thumb: no more than 20-30% of questions at the Remember and Understand levels. Push students toward Analysis and Evaluation, especially in grades 4-12 where passage-level reasoning is the goal.
Depth of Knowledge (DOK)
Derived from Norman Webb’s taxonomy, DOK levels describe the cognitive demand of a task:
- Level 1 (Recall): Students receive or recite facts. Only a shallow understanding of the text is needed.
- Level 2 (Skill/Concept): Requires comprehension and processing of text beyond simple recall. Students might compare two characters or identify cause-and-effect relationships.
- Level 3 (Strategic Thinking): Deeper knowledge. Students must reason, plan, and use evidence. Example: analyzing how an author’s word choice shapes the reader’s perception of a character.
Key difference from Bloom’s: Bloom’s describes types of thinking. DOK describes the depth required. A question can involve “analysis” (Bloom’s Level 4) but still be DOK Level 1 if the answer is explicitly stated in the text.
Teachers who want to build assessments that genuinely push comprehension, not just test memory, should map their questions against both frameworks. For a deeper look at writing effective assessment rubrics, see this guide on creating meaningful rubrics.
Measurement Scales and Benchmarks
When selecting passages for custom assessments, teachers need to consider text difficulty. Two common measurement systems help.
Lexile Level / Lexile Measure
Created by MetaMetrics, the Lexile Framework for Reading assigns quantitative measures to both readers and texts. Measures range from below 200L for beginning readers to above 1600L for advanced readers. The system matches students to appropriately challenging text based on reading ability, not grade level.
Why it matters for assessment creation: If you’re building a reading comprehension assessment, the passage needs to be at the right difficulty for your students. A passage that’s too easy won’t reveal comprehension gaps. A passage that’s too hard turns a comprehension assessment into a decoding test.
NAEP Achievement Levels
The National Assessment of Educational Progress reports student performance across three levels: NAEP Basic, NAEP Proficient, and NAEP Advanced. These are performance standards describing what students should know and be able to do at each benchmark.
Why it matters: NAEP data provides the national context for reading comprehension. When only 31% of fourth graders perform at or above Proficient, it signals systemic gaps that classroom-level assessment can help address.
Comprehension Sub-Skills Worth Assessing
Reading comprehension is not a single skill. Strong assessments target specific sub-skills so teachers know exactly where a student needs support.
Literal Comprehension
Recalling facts and details explicitly stated in the text. This is the most basic level, the “what happened” questions. Necessary but insufficient on its own.
Inferential Comprehension
Answering questions about information implied but not directly stated. This requires students to connect clues in the text, draw conclusions, and read between the lines. Inferential comprehension is where many struggling readers break down, even those who can answer literal questions accurately.
Evaluative Comprehension
Judging the quality, value, or credibility of a text. Example: “Is the author’s argument convincing? What evidence is missing?” This level is critical for upper elementary and secondary students developing critical literacy.
Main Idea / Central Theme
Identifying the overarching message or argument. Seems simple, but it requires the ability to distinguish key ideas from supporting details, a skill that develops gradually.
Text Structure
Recognizing how a text is organized: cause-effect, compare-contrast, sequence, problem-solution, description. Students who understand text structure comprehend more efficiently because they can predict how information will unfold.
Vocabulary in Context
Understanding word meaning based on surrounding text rather than memorized definitions. This sub-skill sits at the intersection of vocabulary knowledge and comprehension, making it an efficient assessment target.
When you create reading comprehension assessments, label each question by the sub-skill it targets. This practice, called tagging, turns a quiz from a blunt score into diagnostic information. To explore reading comprehension activities that target these sub-skills, check out this collection of comprehension activities for classroom use.
Well-Known Reading Assessment Programs
Teachers don’t always build assessments from the ground up. These widely used programs provide standardized tools for screening, progress monitoring, and benchmarking.
DIBELS (Dynamic Indicators of Basic Early Literacy Skills)
A set of short fluency measures used K-6 for screening and progress monitoring. DIBELS assesses phonemic awareness, alphabetic principle, accuracy and fluency, and comprehension.
Acadience Reading
Formerly DIBELS Next, Acadience Reading includes oral reading fluency, retell, and maze components. Together, these three measures provide a more complete picture of reading proficiency than any single measure alone.
AIMSweb
A web-based progress monitoring and screening system used by schools and districts to track student growth and identify at-risk readers.
NAEP
The National Assessment of Educational Progress tests reading comprehension in grades 4, 8, and 12. It’s the largest nationally representative assessment and the source of the proficiency statistics cited throughout this glossary.
Important note: These formal programs serve critical purposes, but they don’t cover everything. Teachers frequently need to create supplemental classroom-level assessments tailored to specific texts, units, and skills. Pre-packaged activities in workbooks don’t always match the text being taught, the grade-level difficulty needed, or the specific comprehension sub-skill under focus.
That gap between formal assessment programs and daily instructional needs is exactly where custom assessment creation matters most.
The K-3 vs. 4-12 Assessment Sequence Difference
This distinction catches many teachers off guard, especially those transitioning between grade bands.
For K-3, the assessment sequence begins with the most discrete reading skills: phonological awareness, phonics, and letter recognition. It gradually expands to include fluency, vocabulary, and comprehension. This makes sense because younger students are still learning to decode. You can’t assess comprehension of a passage a child can’t read.
For grades 4-12, the sequence flips. Assessment begins with the most global skill, reading comprehension, and works backward to identify underlying weaknesses when comprehension breaks down. By this point, most students can decode grade-level text. The question shifts from “can they read the words” to “can they understand what the words mean together.”
Teachers creating assessments for early elementary need to ensure foundational skills are solid before layering on comprehension questions. Teachers in upper grades should start with passage-level comprehension and drill down only when results indicate a need.
How AI Tools Speed Up Assessment Creation
Practitioners on Reddit and in teaching forums consistently report one pain point: creating custom reading comprehension assessments takes too long. Teachers know that pre-packaged assessments often miss the mark, but writing original questions, selecting appropriately leveled passages, and aligning everything to standards is a hours-long process, often done at night or on weekends.
This is where AI-assisted assessment creation fills a genuine need. Instead of starting from a blank page, teachers can input a topic, grade level, and desired skill focus, then generate a draft assessment in minutes. The teacher still reviews, edits, and adjusts, but the heavy lifting of initial creation is handled.
The key when evaluating any AI tool for assessment creation is privacy. Teachers working with student data need to ensure FERPA compliance. For a practical checklist, read this guide on using AI in the classroom without violating FERPA.
Build reading assessments in minutes with TeachTools’ quiz generator, designed for K-12 standards alignment and print-ready output.
Putting It All Together
Creating reading comprehension assessments that actually improve instruction requires more than choosing a passage and writing five questions. It requires understanding:
- What type of assessment fits the moment (screening, diagnostic, formative, summative, or progress monitoring)
- Which format matches the comprehension depth you want to measure (MCQ for efficiency, constructed response for depth, maze for progress monitoring)
- Which framework guides your question design (Bloom’s and DOK to ensure you’re testing more than recall)
- Which sub-skills you’re targeting (literal, inferential, evaluative, main idea, text structure, vocabulary)
- What text level is appropriate for your students (Lexile matching)
When teachers understand these terms and concepts, assessment creation becomes faster, more intentional, and far more useful. The data you collect stops being just a gradebook entry and starts driving the kind of targeted instruction that moves students from “below basic” to proficient.
Explore all 23 teacher tools on TeachTools to streamline assessment creation, lesson planning, and classroom materials.
Frequently Asked Questions
What is the difference between formative and summative reading comprehension assessments?
Formative assessments happen during instruction and are designed to give teachers real-time feedback on student understanding. They’re low-stakes: think quick retells, exit tickets, or short quizzes. Summative assessments happen at the end of a unit or grading period and measure what students learned overall. Both are necessary, but they serve different purposes. Formative tells you what to teach next; summative tells you what was learned.
How do I choose between multiple-choice and constructed-response questions?
It depends on what you’re trying to measure. Multiple-choice questions are efficient for assessing a range of skills quickly, and well-written MCQs can test deep comprehension. Constructed responses are better when you need to see a student’s thinking process, especially for inferential and evaluative comprehension. Many effective assessments use both formats together.
Why do cloze and maze tasks have limitations for assessing comprehension?
Research shows that cloze and maze tasks correlate most strongly with sentence-level comprehension. They test how well students interpret individual sentences using context clues, but they don’t measure the ability to synthesize information across an entire passage. They’re useful for screening and progress monitoring, not as standalone comprehension assessments.
How often should I create reading comprehension assessments for my classroom?
Screening typically happens three times per year. Progress monitoring can be biweekly for students receiving intervention. Formative assessment should be ongoing, embedded in daily or weekly instruction. Summative assessments align with the end of units or grading periods. The frequency depends on the type and purpose.
What is the science of reading, and how does it affect assessment?
The science of reading is a body of research showing that reading requires explicit, systematic instruction. It’s not a program or curriculum but an evidence base. It affects assessment because states are increasingly mandating that schools use approved screeners and assessments aligned to this research. Teachers creating their own assessments should ensure they’re measuring skills the evidence says matter, like decoding, vocabulary, fluency, and comprehension, rather than relying solely on tradition.
How do Bloom’s Taxonomy and Depth of Knowledge differ?
Bloom’s Taxonomy categorizes types of thinking (remember, understand, apply, analyze, evaluate, create). Depth of Knowledge describes how deeply a student must engage with content. A question might involve “analysis” on Bloom’s but still be DOK Level 1 if the answer is stated directly in the text. Using both frameworks together gives the clearest picture of question quality.
What role do Lexile levels play when I create reading comprehension assessments?
Lexile levels help you match passage difficulty to student reading ability. If a passage is far above a student’s Lexile range, the assessment measures decoding struggle rather than comprehension. If it’s too easy, it won’t reveal gaps. When creating custom assessments, selecting passages within or slightly above students’ Lexile ranges produces the most accurate comprehension data.
Can AI tools help teachers create reading comprehension assessments effectively?
Yes, when used as a starting point rather than a final product. AI tools can generate draft questions, suggest passage-level activities, and align items to grade-level standards in minutes. The teacher’s role is to review, adjust for their specific students, and ensure the assessment targets the right sub-skills and cognitive levels. The time savings are significant, especially for teachers who need to create assessments tailored to texts not covered by pre-packaged materials.