Designing Assessments That Reveal Thinking
Once the intended cognition has been identified, the next question is how to design an assessment that provides evidence of it.
Over the past two decades, the Stanford History Education Group (now the Digital Inquiry Group) has provided some of the most influential practical responses to this problem. Their History Assessments of Thinking (HATs) represent one of the most sustained attempts to translate theories of historical thinking into classroom assessments.
Rather than treating historical thinking or working with evidence as a single, holistic construct, the designers of HATs isolated disciplinary processes such as sourcing, contextualization, corroboration, and close reading into focused assessment tasks. By narrowing the targeted cognition, they created observations that offered teachers clearer evidence of students’ historical reasoning.
The same basic logic guides the assessments developed through To The Past. Once we have identified the cognition we want students to demonstrate, we need to ask what students will do that will allow us to observe it. The task, source, prompt, and response format all become part of that design.
This does not require recreating the full complexity of historical inquiry in every assessment. In fact, a more focused task can sometimes provide clearer evidence of a particular form of historical reasoning. The question is not whether an assessment captures everything a student can do as a historian. It is whether it provides meaningful evidence of the particular thinking we intended to observe.
The Hager Letter
Perhaps the best way to illustrate the challenges of designing a historical thinking assessment is by walking one through the process.
While exploring the Canadian Letters and Images Project, I came across a 1915 letter written by Canadian soldier Allan Hager to his sister. Hager himself is not historically significant, but what caught my attention was the letter’s vivid portrayal of a soldier grappling with the realities of war. What kinds of historical reasoning might this source invite? More importantly, how could I design a task that would make that reasoning visible?
The letter was nearly two pages long. It ranged from humorous stories about sightseeing in Canterbury and a bicycle accident to reflections on training recruits before culminating in a vivid description of a Zeppelin attack that had killed many of Hager’s comrades.
Although reading the entire letter would likely be worthwhile, I had other instructional priorities. I wanted a task that could be completed quickly while still isolating a single historical thinking skill. I decided to highlight the final two paragraphs of his letter.
“We had a heavy loss last night, somewhere about seventeen men killed and about thirty wounded, out of our brigade. A German Zeppelin went over our lines. I just missed being one of the unfortunates by the skin of my teeth. All our guard was blown to atoms. It was my turn to be on guard, but as I was down at the ranges coaching for the 7th Brigade, I missed my turn. I saw the whole thing five minutes after it happened. It was all too terrible to mention. The men were shattered to a thousand pieces. It is just a taste of what it will be soon. You people over in Canada have not the slightest idea of how bad the war really is. Even here in England a person never knows when he is going to be killed. I have barely escaped two or three times now. I don’t think I had better tell you any more this time.”
“The Canadians are making a great name for themselves over here. We all swore last night when we saw so many of our comrades killed that the Germans would get no quarter from the 5th Artillery Brigade. I will send you some souvenirs sometime soon.”
— Your brother Allan.
From Open Question to Forced Choice
My first instinct was simply to ask students to infer what the letter revealed about Hager’s emotional response to his wartime experience. I quickly abandoned the idea. The problem was not that the question was uninteresting. It was that it created too much ambiguity about the cognition I would be observing. Students could reasonably produce a wide range of interpretations, and I would then have to determine whether differences in their responses reflected differences in historical reasoning, differences in what they noticed in the source, or simply different interpretations of an inherently open question.
My second idea was to present a single statement suggesting that Hager had been emotionally devastated by the war and ask students whether they agreed or disagreed, supporting their answer with evidence.
This was better, but it introduced another problem. Students could search the letter for evidence confirming the interpretation they had been given without seriously considering whether another interpretation might be better supported. The cognitive task became one of finding supporting evidence rather than evaluating competing inferences.
Eventually I settled on presenting two competing interpretations, asking students to determine which was better supported by the available evidence. Students therefore had to discriminate between competing interpretations by evaluating the relative strength of the evidence supporting each. This is closer to an important feature of historical reasoning: historians do not simply ask whether an interpretation is possible or plausible. They ask how well it is supported by the available evidence, or which is most plausible.
“The letter suggests that Allan Hager has become detached and indifferent to the loss of his fellow soldiers and to the devastation of war.”
“The letter shows that the devastation of war has taken a significant emotional toll on Allan Hager and has left him psychologically worn down.”
Writing two plausible interpretations proved surprisingly difficult. The distractor needed to represent a genuine reading of the source while remaining, in my judgment, less well supported than the alternative. If one interpretation was obviously weaker, students would have little reason to weigh the evidence carefully.
In fact, I hoped at least a handful of students would select the less-supported interpretation. A task that produces unanimous agreement may provide little opportunity for students to discriminate between competing interpretations or for teachers to observe differences in the quality of their reasoning.
When students challenged the distinction between the two interpretations, the disagreement often led us back to the source itself. That was precisely what I wanted.
The purpose of the assessment was not to discover whether students could guess the teacher’s preferred answer. It was to reveal how they evaluated evidence, distinguished stronger from weaker interpretations, and justified historical claims.
Where the Evidence Actually Lives
I chose to include a selected-response item because it provides useful scaffolding. Asking students to commit to the interpretation they believe is best supported gives them a starting point for explaining their reasoning.
But the selected response itself provides limited evidence of the underlying cognition. The evidence of thinking lies primarily in the justification that follows.
“Because he sounds sad.”
Cites Hager’s reluctance to describe the carnage, his observation that civilians in Canada “have not the slightest idea how bad the war really is,” and his decision to withhold further details from his sister.
Both students have selected the same interpretation. Their responses, however, provide very different evidence about how they arrived there.
A response is evidence of cognition only to the extent that the observation makes the relevant thinking visible.
The task does not become useful because it produces an answer. It becomes useful because the response allows us to examine the reasoning behind that answer.
Why the Hager Task Worked
The Hager assessment succeeded not because it captured every dimension of historical thinking, but because its design decisions were deliberately connected to the cognition it was intended to reveal — in this case, evaluating competing inferences.
Selected because it supported more than one plausible interpretation while containing sufficient evidence for students to weigh competing inferences.
The selected-response item scaffolded students’ thinking by asking them to commit to an interpretation before explaining it.
The written explanation revealed the reasoning behind their judgment.
Rewarded the quality of the evidence and justification rather than simply the conclusion.
Each element contributed to the observation. This is the central insight of the Observation vertex of Pellegrino, Chudowsky, and Glaser’s (2001) Assessment Triangle: the source, prompt, response format, and scoring criteria should work together to produce evidence of the particular cognition an assessment is intended to observe.
I also worked to ensure that potential barriers did not unnecessarily interfere with the thinking I intended to observe. In the Hager assessment, for example, I excerpted the letter to focus students’ attention on the passages most relevant to the task, limited cognitive load, provided support for potentially unfamiliar vocabulary, and limited the writing required to a brief justification. These decisions reduced unnecessary reading and writing demands without removing the historical reasoning students were expected to perform.
In designing the Hager assessment, a simple question guided most decisions:
If a student struggles with this task, what do I want that struggle to tell me?
If the answer is that the student is struggling to evaluate competing interpretations, the observation is providing useful evidence. If the struggle instead reflects unnecessary reading, writing, or task demands, the assessment is revealing something other than the thinking I intended to observe.
Designed for the Realities of Teaching
Finally, these decisions have the added benefit of increasing the practicality of the assessment. Teachers make dozens of instructional decisions every day while balancing limited planning time, large classes, and competing curricular demands. An assessment may be theoretically elegant, but if it requires an hour to administer, another hour to mark, and several more hours to interpret, it is unlikely to become a regular part of classroom practice.
Effective assessment must therefore be designed with the realities of teaching in mind. The brevity and narrow focus of the example above — and others in the To The Past library — are not a compromise. They are a design feature.
By lowering the practical costs of assessment, they make it more likely that teachers will use assessments regularly and respond to the evidence they generate. In many cases, a single historical source, one carefully designed prompt, and a brief written justification proved sufficient to reveal the targeted cognition.
Frequent, brief assessments allow teachers to observe growth over time rather than relying on isolated performances. A five-minute assessment that teachers administer every week will often do more to improve learning than a beautifully designed assessment that is used only once a semester.
This does not mean that longer, richer historical inquiries should be abandoned. Essays, document-based questions, debates, and research projects remain essential components of history education because they allow students to integrate multiple aspects of historical thinking in authentic contexts. However, these tasks are often too complex and time-consuming to serve as frequent formative assessments — which is why brief, focused tasks became my priority.
Shorter, more focused assessments complement rather than replace more complex inquiry by providing teachers with regular opportunities to monitor students’ developing historical reasoning throughout a unit of study.
Toward an Assessment Design Philosophy
In developing formative assessments for the Canadian Big Six framework, one lesson emerged repeatedly.
Do not let perfection become the enemy of good.
No assessment can capture every dimension of historical thinking, nor should it try. The goal is not to replicate the full complexity of historical inquiry in every assessment, but to make particular components of historical reasoning visible.
By narrowing the cognitive focus of a task, teachers can more clearly observe specific aspects of students’ reasoning — whether they attend to source metadata, identify relevant contextual information, draw defensible inferences, or corroborate competing accounts. In this way, formative assessments function less as culminating performances than as diagnostic windows into students’ historical thinking.
This focused approach is particularly valuable as students develop their historical reasoning. Specific disciplinary processes can be practiced deliberately before students are expected to coordinate them in more complex forms of historical inquiry (Breakstone, Shemilt, Lee). Short assessments can therefore help students develop individual disciplinary moves — selecting relevant evidence, drawing inferences, explaining significance, corroborating sources, and connecting evidence to historical claims — before integrating them into larger inquiries.
Taken together, these assessments provide something that large summative tasks often cannot: a clearer signal of student thinking. Individually, each offers only a narrow glimpse of students’ reasoning. Collectively, however, they reveal patterns that allow teachers to provide more targeted feedback and make better instructional decisions.
The power of these assessments also lies in their frequency.
One short assessment provides a snapshot. Ten provide a trajectory.
When students repeatedly engage in historical reasoning, they begin to internalize its expectations (Monte-Sano & La Paz, 2012). Historical thinking becomes less something they perform for an assessment and more a habitual way of approaching historical questions. Assessments that repeatedly make students’ reasoning visible can therefore do more than measure historical thinking — when the evidence is used to shape subsequent instruction and feedback, they can also help cultivate it (Black & Wiliam, 1998).
In my own classroom, these tasks became far more than assessment tools. They became teaching tools. Rather than asking students to construct sweeping interpretations from large collections of documents, they invited them to grapple with manageable historical problems:
Which interpretation is better supported?
How reliable is this source?
What evidence justifies this claim?
These are questions students can genuinely debate. Through repeated encounters with such problems, students develop the habits of mind that larger inquiries ultimately demand.
An assessment is not formative simply because it is brief or administered during instruction. It becomes formative when the evidence it produces informs what happens next. A ten-minute CHAT that never influences teaching is no more formative than a unit test returned weeks later. Conversely, a carefully designed discussion, written response, or selected-response question can serve formative purposes if it helps teachers decide what students need next.
Observation provides the evidence. Interpretation determines what that evidence means and how it is used.
