Sound and Situated Assessment: The Choices We Make

I’ve given a few talks recently on assessment in an AI-infused world, and I’ve found myself returning to a question that has been sitting with me for some time: what kind of assessment do we need now, in an AI world full of challenges and big, pressing, complex problems? As I ran a session on assessment today, I thought it was a good time to consolidate a few thoughts.

For years, my own work has focussed on authentic assessment. In my own work, I’ve framed authentic assessment as something relevant to the self, the discipline, employability, and the wider world. But over the last couple of years, particularly with James Croxford, I’ve started to question whether the concept itself is still valuable or whether it has run its course. We have spent a long time debating what “authentic assessment” is, refining definitions, proposing frameworks, and critiquing each other’s interpretations. I regularly get papers to review on this topic – to expand or nuance the definition, or to test some aspect of authentic, but I think we have lost sight of the fact that this is a conceptual label that can’t hope to be a single answer to a multilayered set of issues. There is a great deal of thoughtful scholarship around authentic assessment, and I respect it. It has served a role in nudging forward what is possible in the assessment space, but there is also a lingering problem: if we are still struggling to pin the concept down, what is it helping us to do? That led us to ask: is it time to stop talking about authentic assessment? I’m increasingly persuaded that it might be.

Why do we need to move on from authentic assessment? Not because the ideas underpinning authentic assessment are wrong, far from it, but because the label itself may now be getting in the way. It risks becoming a catch-all term that tries to do too much: relevance, real-world application, creativity, employability, social justice, performance, compliance, and more. When everything sits under one label, clarity is lost. So rather than asking whether assessment is “authentic” or not, I’ve started to think more simply about effective assessment.

The starting point, for me, is to ask: what is this assessment trying to do? Assessment can serve multiple purposes. It can drive learning. It can test the application of knowledge under specific conditions. It can develop creativity and problem-solving. It can ensure compliance with professional standards. It can test procedural competence. It can support students to engage with complex, real-world issues. These purposes often require different designs. Trying to achieve all of them at the same time, within a single assessment, is unrealistic.

I’m interested in an approach that focuses on doing the right thing at the right time.

I’m interested in making sound, situated assessment judgments. And yes, I did cross out that sentence about the right thing; as my workshop today reminded me – assessment doesn’t have a single correct response – many answers may be possible. A sense of the right answer can, I think, lead to positioning and tension rather than negotiation in this space.

Constructive alignment and validity are critical concepts in grounding judgments, but I think alone they are not enough here. It is not just about aligning learning outcomes, teaching activities, and assessment tasks. It is about aligning assessment with a broader system: institutional priorities and strategy, disciplinary norms, graduate attributes, resource constraints, staff expertise, and the values and requirements of the programme (including any professional body elements).

In my doctoral work, I conceptualised an ecosystem behind feedback, recognising the act of giving feedback is shaped by external, institutional, cultural, and individual factors brought together through actors’ internal conversations and outward socialisation. Those forces still feel highly relevant eleven years on, and in the space of assessment design as well as for feedback generation. The interaction of these- and maybe other factors – shape what is possible, valued, rewarded and effective.

I am not suggesting that we need to analyse every one of these factors every time we design an assessment. But at the level of a programme or course, there is real value in stepping back and asking: What are our governing variables? Where do we have flexibility?; Where are the constraints? What is our assessment philosophy (i.e. what matters)? Howdoes that fit with our institution and context? From there, assessment design becomes a series of informed choices. In my logical brain – I am seeing a decision tree of if this, then that … but it’s never that simple, of course.

And what choices do we have? We might choose higher or lower fidelity to professional contexts (i.e. assessments that emulate professional environments or not); we might prioritise security or openness; we might design tasks that emphasise compliance and standardisation, or ones that invite creativity and agency; we might include elements of public exhibition, collaboration, or individual performance, or not at all. None of these choices are inherently better than others – they are only more or less appropriate in context. This is why I find it increasingly difficult to make general claims about “good” assessment across the sector. What works in one context may be entirely inappropriate in another – assessment must be brokered and it sits within a localised context.

How is AI situated in a multilayered environment underpinned by agency and choice.?

Add then, thoughts about AI: AI does not simply “disrupt” assessment; it exposes the assumptions that underpin it. It forces us to confront questions about what we value: process or product, individual performance or team effort, originality or effective use of tools etc. AI makes a contextual approach to assessment even more important. It pushes us to be explicit about our purposes and deliberate in our design choices, rather than relying on inherited formats. It should force us to ask why we work this way or that way.

So perhaps we should stop asking whether assessment is authentic and start asking whether it is effective for this purpose, in this context, with these constraints, for these students and these staff and systems. That requires us to understand our own assessment ecosystem (including and especially resources and the operational reality), articulate a clear vision, and make intentional choices. It also requires us to accept that there is no single right answer to effective assessment in an AI world – we can only consider, navigate, decide, try, reflect, iterate, listen, evaluate and evolve – just as generations of reflective educators have always done.

AI Declaration & reflection: This article was dictated roughly on my walk I then pasted the text from voice note in to Perplexity and required a ‘tidy up’ of the transcript before manually editing it through.  The diagram was iterated with perplexity – AI created a starter based on my article, then I fused elements of my own ecosystem diagram into the base diagram and iterated to balance simplicity with structure. As a footnote sometimes this interaction with AI gave me a perfectionist complex and made me question whether I should bother, because AI never says – that’s great! This is a risk I am noticing – AI is not a saviour, nor is it neutral – it’s a mixed picture. So Perplexity …“assessment must be brokered and it sits within a localised context” → slightly clunky; consider “assessment must be brokered within a localised context” – I’m going with clunky 🙂

2 responses to “Sound and Situated Assessment: The Choices We Make”

  1. Great article Lydia. I was musing on your point that ‘perhaps we should stop asking whether assessment is authentic and start asking whether it is effective for this purpose, in this context, with these constraints, for these students and these staff and systems’. This reminds me of the discussions I have about establishing the validity and reliability of qualitative data (often wrongly confused with the dominant narrative associated with quant around objectivity and consistency). In the qual sphere this is best established by considering the appropriateness of the qualitative approaches and tools on offer in the context under study, reflections on the multiple ways evidence might be collected, and decisions why some are more appropriate, and others less so.

    1. Thanks Liz, I see the parallels with quals and quants for sure. Sometimes it feels like we want assessment to have a foot in both camps – objective, neat processes, replicable etc. but then situated and negotiated – I guess that is what makes this a permanent wicked problem.

Leave a Reply

Discover more from Home

Subscribe now to keep reading and get access to the full archive.

Continue reading