Interpretation — Reading the Evidence of Student Thinking — ToThePast
Interpretation

Reading the Evidence of Student Thinking

We have now arrived at the third corner of the Assessment Triangle—and, in many ways, the most demanding. Once a clear model of historical thinking has been developed, and tasks developed to elicit evidence of that thinking, we find ourselves faced with student work and must ask: “What does this student work actually reveal about their thinking? Answering that question allows us to provide meaningful feedback, make informed instructional decisions, and, ultimatley help them grow as thinkers.

A Worked Example

What Would You Infer?

Two students read a speech delivered by John A. Macdonald in the Legislative Assembly of the Province of Canada during the Confederation debates in 1865 and are asked to evaluate it as evidence of how Canadians felt about Confederation.

Primary SourceJohn A. Macdonald, 1865

“I have had the honor of being charged, on behalf of the Government, to submit a scheme for the Confederation of all the British North American Provinces — a scheme which has been received, I am glad to say, with general, if not universal, approbation in Canada. The scheme, as propounded through the press, has received almost no opposition. While there may be occasionally, here and there, expressions of dissent from some of the details, yet the scheme as a whole has met with almost universal approval, and the Government has the greatest satisfaction in presenting it to this House.”

Student A

“Macdonald wanted Confederation to happen, so in his speech he mostly talks about why it would be good and probably won’t even mention the bad stuff because he wanted the other government people to vote for it, so I don’t think this source gives both sides and I’d use it carefully as evidence of how people felt about Confederation.”

Student B

“This speech reflects Macdonald’s perspective and bias. His intended audience and purpose influence the information he presents, which may affect the reliability of the source. Historians must consider the author’s perspective and bias before deciding how useful the source is.”

What can we confidently infer about each student’s thinking?

— Take a moment before reading on. —

Student B’s response is polished, concise, and filled with the academic vocabulary we often associate with historical thinking: bias, perspective, audience, and reliability. Student A, by comparison, writes awkwardly. The language is informal, the sentence structure is clumsy, and the explanation lacks precision.

Yet Student B’s response remains largely generic. If we replaced Macdonald with almost any historical figure, the response would still make sense.

Student A, despite weaker writing, identifies Macdonald’s position on Confederation and uses that specific knowledge to infer how it shapes the evidence the source provides. The student recognizes that Macdonald was trying to persuade other politicians to support Confederation and uses that understanding to evaluate the source’s usefulness as evidence of how Canadians felt about Confederation. The writing is rough, but the historical reasoning is more developed.

Is Student B wrong? Not necessarily. The student demonstrates an emerging understanding that authors have perspectives, write for particular audiences, and pursue specific purposes. What is missing is the application of those ideas to this particular source. The response identifies perspective as something that exists but never explains how Macdonald’s support for Confederation shaped the arguments he chose to make.

We have now arrived at the third corner of the Assessment Triangle — and, in many ways, the most demanding. We began by defining the cognitive processes that constitute historical thinking, then designed assessments that made those processes visible. This page asks a different question:

What does this student work actually reveal about the student’s thinking?

Answering that question allows us to provide meaningful feedback, make informed instructional decisions, and, over time, develop increasingly confident judgments about students’ historical reasoning.

First Principle

Assessment Is Fundamentally an Act of Interpretation

Teachers, like historians, never observe their subject directly. Historians infer the past from the traces it leaves behind. Teachers likewise infer students’ thinking from the traces it leaves behind — exit slips, classroom discussions, margin annotations, HATs, CHATs, DBQs, and countless other classroom performances.

These responses are not historical thinking itself. They are evidence from which teachers infer the thinking that produced them. Student work never speaks for itself; every conclusion about what a student understands, misunderstands, or is ready to learn next depends upon how we interpret that evidence.

Teachers do possess one important advantage. Historians cannot generate new evidence from the past. Teachers can. Through carefully designed assessment tasks, teachers create opportunities for students to reveal their thinking. The purpose of these observations is not simply to assign scores or measure achievement, but to generate evidence from which defensible inferences about student thinking can be drawn.

The quality of assessment therefore depends not only on the quality of the evidence collected, but on the quality of the inferences we are able to draw from it.

Second Principle

Interpret From a Cognitive Model

How should teachers read the evidence of student thinking that assessments produce?

The Assessment Triangle offers a useful corrective here: interpretation, cognition, and observation are not sequential steps but three points of a single system, each shaping the others. The cognitive model names the thinking we hope to see. That naming shapes the tasks we design to draw it out. Student responses, in turn, only make sense when read against that same model — the one that helped build the task in the first place.

Take that model away, and interpretation drifts. It has nothing to anchor to. This is precisely why building a cognitive model for evidence mattered so much in the first place. It was not simply a preliminary exercise completed before getting to the “real” work of assessment. It provides the foundation for interpreting the evidence that follows.

Consider a DBQ. Students may be sourcing documents, contextualizing them, corroborating evidence, constructing an argument, and communicating their conclusions simultaneously. That complexity can provide valuable evidence of students’ ability to coordinate historical reasoning, but it can also make interpretation difficult. If a response is weak, what exactly has the student struggled to do?

A cognitive model gives the teacher a way to make that question more precise. Rather than simply evaluating the finished product, the teacher can examine what the response reveals about the particular cognitive processes involved.

Narrower tasks, designed to target a smaller number of cognitive processes at a time, offer a different advantage. They reduce the ambiguity surrounding interpretation. When teachers know exactly what form of reasoning a task was designed to elicit, their inferences about student thinking become more defensible and the feedback that follows becomes more precise and actionable.

None of this means narrow tasks are better than complex ones. They serve different purposes. A DBQ asks students to coordinate reasoning under conditions that resemble the complex intellectual work of the discipline. A narrowly targeted task allows us to focus more closely on a particular process so that we can observe and interpret it with greater precision.

Narrowly focused tasks and richer inquiries therefore serve complementary purposes. One isolates particular forms of reasoning so they can be observed, interpreted, and developed with precision. The other requires students to coordinate those forms of reasoning in ways that more closely resemble authentic historical inquiry. Effective assessment needs both.

Whatever assessment we use, one thing remains constant: walking into the interpretation of student work with a clear cognitive model changes everything downstream. It provides the basis for deciding what evidence matters and what that evidence might reveal about the student’s thinking.

Third Principle

Ground Interpretation in the Specific Content of the Assessment

A mathematics teacher assessing the solution of a linear equation can expect broadly similar reasoning even when the numbers change. Historical thinking does not work this way. A sourcing task may consistently ask students to consider the evidentiary value of a source, but what constitutes a strong response depends upon the source itself. One document may demand close attention to authorship. Another may hinge on context or purpose. The historical content determines which considerations become most important.

This relationship between cognition and content may help explain why the “skills versus content” debate has proven so persistent in history education.

Historical thinking cannot be practiced in the abstract. Historians do not contextualize in general — they contextualize particular events. They do not corroborate abstract propositions — they corroborate specific accounts. What distinguishes historical reasoning, then, is precisely how much the content shapes what successful reasoning looks like.

This has important implications for assessment. A generic descriptor such as connects the author’s perspective to the evidentiary value of the source identifies the targeted cognition, but it does not tell teachers what that reasoning should look like in response to a particular source. When interpreting student work, we must therefore have a clear sense of what evidence would demonstrate the targeted reasoning within this particular task. Whenever possible, assessment-specific rubrics or interpretive guides can help establish what that evidence might look like.

I have been teaching for some time, but I still find myself uncertain, particularly when I encounter a new assessment. My first inclination is to sit down and complete it myself, the way a student would. Then I’ll send it to a trusted colleague: What do you think — is this what we’re going for? Is this what you hoped students would notice? Is this what you’d count as convincing evidence? If not, what would an exemplary response look like, and why?

A cognitive model tells teachers what kind of reasoning they are trying to observe. An interpretive guide, or a set of exemplars, helps them determine how that reasoning is likely to surface within the particular source and question. It can identify the features of the source worth attending to, the kinds of explanations that would count as convincing evidence, and the misconceptions likely to appear along the way.

This kind of guidance does not replace professional judgment. It sharpens it. The clearer teachers are about what counts as convincing evidence, the more consistent their interpretations become and the better positioned they are to diagnose student work.

Fourth Principle

Interpret to Diagnose

The distinction between judgment and diagnosis matters because explanation, not evaluation, is what gives assessment its instructional value.

Judgment

How good is this performance?

Diagnosis

What does this performance reveal about the student’s thinking?

It is easy to move almost habitually from observation to evaluation. A student demonstrates weak analysis. Their sourcing is limited. The argument is persuasive. These judgments may be accurate, but they do not explain the thinking that produced them. Before making a judgment, it is often more useful to ask a different set of questions:

What did the student do well?

Where does their reasoning begin to break down?

Did they explain why the source’s origin mattered, or simply identify it?

Did they recognize conflicting evidence without reconciling it?

These are diagnostic questions. They move beyond describing the quality of a performance toward explaining the thinking that produced it.

Diagnosis also changes what happens next. A verdict concludes the assessment process. A diagnosis points toward a response. Teachers seek to understand the reasoning that produced a student’s work before deciding how to respond instructionally.

For example, a student may recognize that an author’s purpose affects a source but fail to explain how. Another student may recognize conflicting evidence but treat the contradiction as a problem to be noted rather than something to be explained. These students may both receive a general judgment that their analysis needs improvement, but they require different instructional responses.

The distinction becomes especially important when providing feedback. When assessment is interpreted against a clearly specified cognitive model, teachers can identify which aspects of the targeted reasoning are well developed, where that reasoning begins to break down, and what students might do next to improve. Feedback shifts from evaluation toward explanation and prescription.

This is one reason narrower, more targeted assessment tasks are so valuable. When an assessment is designed to elicit evidence of a small number of clearly specified cognitive processes, teachers can identify more precisely where a student’s reasoning is succeeding and where it begins to break down. That diagnosis provides a stronger foundation for formative feedback than broad judgments about the quality of a finished product.

This is where diagnosis earns assessment its instructional value. It allows teachers — and increasingly students themselves — to see not merely whether a response is right or wrong, but what forms of historical reasoning have been demonstrated, where that reasoning begins to falter, and what students might do next.

Fifth Principle

The Goal Is Better Inference, Not Certainty

Teachers work with imperfect models, imperfect observational tools, and imperfect evidence. That’s simply the condition under which all assessment operates.

None of this makes assessment impossible. It simply means that certainty was never the goal to begin with. The goal is to make our inferences more valid, more consistent, more theoretically grounded, and more transparent — not perfect, but defensible, and better than the inferences we made last time.

The Digital Inquiry Group History Assessments of Thinking, the tasks described throughout this site, and whatever comes next are not perfect. That is not an admission of failure but an acknowledgement of the realities of assessment design. Every assessment involves trade-offs. No single task can simultaneously maximize authenticity, diagnostic precision, efficiency, reliability, and breadth. Each represents a series of deliberate design decisions about which qualities to prioritize and which limitations to accept.

The goal is not to eliminate those limitations, but to design assessments that nevertheless provide useful evidence about student thinking. Good assessment doesn’t pretend otherwise. It doesn’t chase certainty it can’t have. It manages uncertainty responsibly — narrowing it where possible, acknowledging it where it can’t be narrowed, and never mistaking a single performance for the whole of a student’s thinking.

Interpretation is the continual work of constructing the best explanation available from the evidence at hand — while remaining willing to revise it when better evidence arrives.

From Interpretation to Instruction

The Bridge to What Happens Next

The value of interpretation is ultimately practical. We do not assess students’ thinking simply to describe what they can currently do. We assess it because that evidence can help determine what should happen next.

The value of a To The Past assessment is not in the score it produces, but in the instructional information it provides. A short assessment that reveals a particular difficulty in historical reasoning can be more useful than a much larger assessment that tells us only that a student is struggling.

If a student can identify relevant source information but cannot explain how it affects the interpretation of the source, the teacher has something specific to respond to. The next lesson, discussion, model, or practice opportunity can address that particular difficulty. Interpretation therefore provides the bridge between assessment evidence and instructional action.

This is also why the distinction between observation and interpretation matters. Observation provides the evidence. Interpretation tells us what that evidence suggests about student thinking. Instruction responds to that interpretation. The process then begins again.

1 COGNITION 2 OBSERVATION 3 INTERPRETATION 4 INSTRUCTION

Cognition identifies the thinking we want students to develop. Observation provides an opportunity to see that thinking. Interpretation helps us understand what students’ responses reveal about their current reasoning. Instruction responds to that evidence, creating new opportunities to observe what students can now do.

Assessment is not a single event. It is a cycle of making thinking visible, interpreting the evidence, and using what we learn to help students develop.