A field of identical pale architectural forms, with one cut away to reveal illuminated layers, pathways and structures beneath the surface.

When Answers Become Abundant, Learning Evidence Must Change

Generative AI is reducing the cost of producing a plausible answer. It is not reducing the importance of learning, judgement or responsibility.

That distinction is becoming difficult to ignore.

A learner can now produce a polished explanation, plan, calculation or recommendation without doing all the cognitive work that the finished artefact appears to represent. Another learner may use the same tools to question, test, revise and deepen their understanding. The visible products can look remarkably similar. The learning underneath them may be entirely different.

This does not make the answer meaningless. It makes the answer insufficient on its own.

The distinction became concrete during a recent workshop with tutors at Y Education in Christchurch. The group held different positions on AI: enthusiastic, cautious, neutral and resistant. We did not need to resolve those differences. The more durable question was the same with AI present or absent:

What thinking did the learner still need to do?

The workshop used five conditions for keeping that thinking visible: Attempt, Question, Check, Explain Judgement and Apply. The model worked in the room. But a workshop is a data point, not proof. When the model was tested against research, its central premise survived while several of its claims had to become narrower. Assistance needed to be calibrated. Questions needed guidance. Checking needed to belong to the discipline. Explanation needed adaptive follow-up. Application needed changed conditions. Feedback, present throughout the workshop but absent from the diagram, needed to become explicit.

The larger signal is not the emergence of one framework. It is the weakening of an old educational shortcut:

A finished answer can no longer be treated as a complete account of the capability that produced it.

Five conditions for visible thinking—Attempt, Question, Check, Explain Judgement and Apply—arranged as a cycle on a foundation of context, learner readiness and scaffolding, with feedback looping through the process.
Revised Five Conditions for Visible Thinking. The five conditions sit on a foundation of context, learner readiness and scaffolding, while feedback closes the loop by changing what happens next. This is a research-informed design heuristic, not a validated assessment model.

1. Purpose of This Briefing

Primary Goal

To examine how abundant AI-generated answers are changing the evidence required to support defensible judgements about learning and capability.

Core Strategic Premise

As assistance becomes more capable and more readily available, education systems need to distinguish assisted performance from developed capability. Confidence should increasingly rest on evidence gathered across a learning process: what the learner attempts, asks, checks, decides, explains, applies, revises and can do when conditions change.

What This Briefing Is Not

This is not an argument for banning AI, forcing every learner to use it, replacing written work, recording every learning interaction, or treating professional conversation as proof of authorship.

It is also not a claim that the Five Conditions for Visible Thinking framework has been validated. The framework is a research-informed design heuristic. Its usefulness depends on how it is enacted, the learner’s readiness, the domain, the quality of feedback and the judgement of the educator.

Desired Reader Outcome

Readers should leave with:

  • a sharper distinction between performance and learning;
  • a clearer account of why final products have reduced standalone evidential value;
  • a practical way to design for visible cognitive work without creating a surveillance system;
  • a set of implications for teaching, assessment and learning assurance;
  • clear limits on what visible evidence can and cannot establish.

2. Executive Signal Summary

Signal 1 — Assisted performance can conceal weaker learning

A 2025 randomised field experiment involving nearly 1,000 secondary mathematics students found that access to a GPT-4-based tutor improved performance during practice. Students with access to an unrestricted ChatGPT-like tutor later performed 17 per cent worse than the no-AI control group on an unassisted examination. A tutor constrained to provide teacher-designed hints largely removed that negative effect, although it did not produce a positive examination effect (Bastani et al., 2025).

The study is bounded to one school, mathematics review sessions and particular tutor designs. It does not show that AI generally harms learning. It does show that better assisted performance can coexist with weaker independent learning—and that the design of assistance matters.

Signal 2 — A finished answer is becoming weaker as standalone evidence

The answer still matters. Accuracy, quality and completion remain legitimate parts of assessment. But when different combinations of human and machine work can produce similar outputs, the product cannot carry the full evidential burden once placed upon it.

The relevant question shifts from “Was an answer produced?” to “What combination of evidence justifies confidence that the learner can understand, judge, adapt and take responsibility for it?”

Signal 3 — Visible thinking must be designed, guided and interpreted

Making thinking visible does not mean asking learners to expose every internal step. It means designing selected moments where consequential cognitive work becomes available for feedback and professional judgement.

Research on inquiry learning reinforces the importance of guidance. A meta-analysis of 72 studies found that guidance improved inquiry activities, successful performance and learning outcomes (Lazonder and Harmsen, 2016). Learner agency and instructional support are not opposites. Strong design increases agency by giving learners the models and scaffolds needed to conduct better inquiry.

Signal 4 — Feedback completes the evidence loop

Visible thinking has limited educational value if nothing meaningful responds to it. Feedback is not uniformly beneficial, but a large meta-analysis of 435 studies found a mean effect of d = 0.48, with substantial variation according to the information conveyed and the context in which it was used (Wisniewski, Zierer and Hattie, 2020).

The design question is therefore not only “What thinking will become visible?” It is also “What will talk back, and what must the learner do next?”

Signal 5 — Capability evidence needs to accumulate across a process

No single observation, explanation, transcript or final product proves learning. Confidence becomes stronger when different forms of evidence converge: an initial attempt, guided inquiry, discipline-specific checking, an adaptive explanation of judgement, application under changed conditions, and a visible response to feedback.

The strategic movement is from product inspection toward an evidence ecology.

3. Why This Matters

A. Learning Assurance

AI is creating a performance–learning gap. Systems can observe a high-quality result while remaining uncertain about what was learned, retained or transferable.

This is now a learning-assurance question, not merely an academic-integrity question. TEQSA’s 2026 work on quality learning in an AI-integrated future places adaptive capability, evaluative judgement, critical thinking, ethical reasoning and learning assurance at the centre of institutional response (TEQSA, 2026).

The enduring obligation is not to preserve unaided production for its own sake. It is to maintain defensible confidence that learning outcomes and human capability are genuinely present.

B. Assessment and Credential Trust

Assessment often asks a finished artefact to perform several functions at once: demonstrate knowledge, indicate authorship, evidence reasoning, show application and justify a credential decision.

Generative AI exposes how much was previously inferred rather than directly observed.

The appropriate response is not to abandon products. It is to stop treating them as self-interpreting. A report, design, calculation or answer should sit inside a wider pattern of evidence that helps an educator decide what the learner understands and can do.

This extends the argument made in SIGNAL 001. The collapse of familiar assessment assumptions creates a need for better capability verification. SIGNAL 004 moves that argument into the learning process itself.

C. Equity, Access and Learner Agency

AI assistance can remove barriers, support language production and give learners access to explanations they might not otherwise receive. The same assistance can also hide a capability gap or displace work the learner needed to do.

A simplistic “attempt without help” rule would reproduce its own inequities. Learners enter with different prior knowledge, language resources, confidence, experience and support. Productive effort for one learner can be empty failure for another.

The design challenge is calibrated assistance: enough support to make meaningful participation possible, while preserving the cognitive work through which capability develops.

D. Workload and Professional Practice

Educators reasonably resist models that appear to add another layer of documentation. The Christchurch workshop suggested a more useful starting point: much visible-thinking practice already exists.

Tutors ask follow-up questions. They watch learners perform. They notice hesitation. They ask why one option was chosen over another. They demonstrate, prompt, correct, revisit and vary the task. The problem is not always an absence of evidence. It is that these interactions may be treated as incidental rather than recognised as part of learning and assessment design.

The opportunity is to become more deliberate about selected evidence, not to capture everything.

4. Current System Reality

Observation 1 — Product quality and learner capability are separating

AI can improve the visible quality of work faster than many learning and assessment systems can adapt their interpretation of that work.

The Bastani study offers a particularly clear example: students could look more successful during assisted practice while learning less for an unassisted test. The problem is not unique to AI, but AI makes the separation easier to produce and harder to see.

Observation 2 — Institutional attention still concentrates on tool use

Many responses continue to ask whether AI was permitted, declared or detected. Those questions have administrative value, but they do not answer the educational question.

Two learners can both declare AI use and demonstrate very different levels of understanding. Two learners can both avoid AI and still demonstrate different quality of reasoning. Tool status does not by itself establish capability.

Observation 3 — Useful evidence already exists outside the final artefact

Vocational and work-based learning often contains multiple natural feedback channels: materials resist, measurements fail, customers respond, safety standards apply, equipment behaves unexpectedly and other practitioners ask questions.

This does not make vocational assessment automatically trustworthy or AI-proof. It does mean the environment can provide evidence that is difficult to reproduce through a single polished submission.

Observation 4 — Visible evidence remains open to misinterpretation

A learner can produce a convincing explanation without understanding it. A live answer can be rehearsed. A successful performance can depend on an unusually familiar context. An educator can over-read confidence or fluency.

Self-explanation is supported as a learning activity; a meta-analysis by Bisra and colleagues found an overall benefit from prompting learners to explain material to themselves (Bisra et al., 2018). But learning from explanation is not the same claim as proving authorship or understanding through explanation.

Visible evidence still requires professional interpretation.

Observation 5 — Transfer remains difficult

Capability becomes more credible when a learner can adapt it beyond the conditions in which it was first learned. But transfer is not a single phenomenon and cannot be assumed merely because an activity appears authentic. Barnett and Ceci’s taxonomy shows that transfer varies across multiple dimensions, including the knowledge domain, physical and social context, temporal context and functional demands (Barnett and Ceci, 2002).

Application should therefore include changed conditions, not merely repetition with realistic furniture.

5. What the System May Be Misunderstanding

Misunderstanding 1 — “The issue is whether learners use AI”

AI use is only one variable. The deeper issue is whether assistance supports or displaces the work through which learning occurs.

The more useful distinction is not AI versus no AI. It is productive assistance versus cognitive substitution.

Misunderstanding 2 — “A correct answer demonstrates learning”

A correct answer demonstrates that a correct answer was produced. Depending on the conditions, it may also contribute evidence of learning. It does not automatically establish how the result was produced, what was understood, whether the learner can recognise an error, or whether the capability will transfer.

The answer remains evidence. It is no longer enough evidence by itself.

Misunderstanding 3 — “Visible thinking means collecting more process data”

Exhaustive capture can turn learning into surveillance and increase workload without strengthening judgement. Prompt histories and process logs may create volume rather than insight.

Visibility should be selective and consequential. The aim is to reveal moments that matter: an initial interpretation, a choice between alternatives, a check against evidence, a response to challenge, a revision after feedback or an application under changed conditions.

Misunderstanding 4 — “Learners should struggle before receiving help”

Attempt is valuable only when the learner has enough knowledge and support to enter the task. Withholding help is not a pedagogy.

The design question is: What can this learner do from their current knowledge before substantial assistance, and what scaffold will make that effort productive?

Misunderstanding 5 — “If learners explain their judgement, we know the work is theirs”

An explanation can strengthen evidence. It cannot authenticate learning on its own.

Adaptive dialogue is stronger than a fixed written rationale because the learner must respond to an unanticipated follow-up, relate a decision to evidence and revise when a condition changes. Even then, the result is one part of a broader judgement, not proof.

Misunderstanding 6 — “Authentic tasks solve the problem”

Authenticity helps when the work introduces real constraints, consequences, standards and relationships. It does not remove the need to interpret evidence.

The important question is not whether the task looks like work. It is whether the learner can adapt capability when the work talks back.

6. Emerging Adaptation Patterns

Pattern 1 — From answers to guarded assistance

The contrast between unrestricted answers and teacher-designed hints in the Bastani study points toward a practical design principle: preserve the cognitive work while providing support around it.

This may include staged hints, questions before solutions, partial examples, worked examples with missing steps, or access to stronger assistance only after a meaningful attempt.

The goal is not friction for its own sake. It is assistance that leaves the learner something important to do.

Pattern 2 — From open questioning to guided inquiry

Learners do not automatically know which question will reveal an assumption, expose uncertainty or test a claim. Questioning itself requires modelling and practice.

Useful scaffolds include asking learners to identify what is uncertain, generate competing explanations, predict where an answer could fail, or name the evidence that would change their mind.

Pattern 3 — From generic checking to discipline-specific verification

“Check the AI” is too vague to guide action.

A builder, nurse, historian, engineer, barista and community practitioner work with different standards, tolerances, risks and evidence. Checking becomes meaningful when it belongs to the work.

Research on AI-assisted decision-making provides a related caution. Buçinca and colleagues found that cognitive forcing interventions reduced over-reliance on AI recommendations more effectively than explanations alone, although participants often preferred the less demanding designs (Buçinca et al., 2021).

Convenience and confidence are not substitutes for verification.

Pattern 4 — From fixed rationales to adaptive professional conversation

Written explanations can support reflection and contribute evidence, but they are increasingly easy to generate after the fact.

Professional conversation introduces responsiveness. The educator can ask why one option was rejected, what evidence mattered, what risk was accepted, or what would change the decision. The value lies less in oral performance than in the learner’s capacity to connect, defend and revise judgement.

Pattern 5 — From repetition to changed-condition application

Application becomes more informative when something changes: the client, material, constraint, standard, time pressure, evidence or consequence.

The learner is not merely asked to reproduce a known response. They must recognise what remains relevant and what requires adaptation.

Pattern 6 — From feedback received to feedback used

Feedback should not end with comments delivered. The evidentially important movement is what the learner does in response.

Did they correct the error? Revise the plan? Ask a better question? Change the technique? Explain why they retained the original decision? Apply the learning in a new context?

Feedback closes one loop and opens the next attempt.

7. Strategic Implications

For Educators

The core design question becomes:

What must the learner still do for this activity to develop and reveal capability?

This question works whether AI is encouraged, constrained, peripheral or absent. It directs attention away from ideological agreement about tools and toward the quality of learning.

For Assessment Designers

Assessment should use multiple evidence points proportionate to the stakes. A final product can remain central, but higher-confidence decisions may also require evidence of process, judgement, feedback response and application.

Not every task needs every condition. The framework is a design language, not a compliance checklist.

For Providers

Learning assurance will increasingly depend on programme-level evidence design rather than isolated changes to individual tasks. Providers need a coherent account of where capability becomes visible, how evidence is interpreted and where stronger verification is required.

Professional learning should therefore address the judgement of evidence, not only the operation of AI tools.

For Vocational and Work-Based Learning

Vocational education may have an advantage because work already contains material, social and professional feedback. But that advantage needs to be recognised and designed deliberately.

The conversation on the workshop floor, the correction during practice and the adaptation when conditions change may be educationally important evidence. They should not automatically become bureaucratic records, but neither should they remain invisible simply because the assessment system privileges a final written product.

For Policy and Quality Assurance

Policy needs to move beyond permitted-use categories toward clearer expectations for learning assurance. The policy question is not only what tools are allowed. It is how institutions justify confidence that learners have achieved the intended capability under AI-integrated conditions.

8. Intervention Options

Option 1 — Retain product-centred assessment with minor disclosure controls

This is the lowest-disruption option. It preserves familiar tasks and asks learners to declare assistance.

Its limitation is evidential: disclosure does not establish what was learned, and undeclared use remains difficult to interpret.

Option 2 — Add process capture and checkpoints

Drafts, prompt histories, reflections and staged submissions can make parts of the process visible.

This can strengthen evidence when checkpoints are purposeful. It can also create performative compliance, excessive data and additional workload. Process capture should be selective rather than exhaustive.

Option 3 — Build visible-thinking conditions into task design

Tasks can preserve a meaningful attempt, scaffold useful questions, require discipline-specific checks, elicit adaptive judgement, vary application conditions and require a response to feedback.

This is the strongest near-term option because it improves learning design without depending on a new platform or universal surveillance.

Option 4 — Develop programme-level evidence ecologies

Programmes can map where different forms of evidence accumulate across courses, placements, demonstrations, conversations and assessments.

This reduces the pressure on any single artefact and supports risk-proportionate verification. It requires stronger moderation, staff capability and evidence governance.

Option 5 — Build detection and surveillance infrastructure first

This option should be treated cautiously. It may create the appearance of control while leaving the learning-design problem unresolved.

Detection can support investigation in limited contexts. It is not a substitute for evidence that capability developed.

9. Emerging Directional Principles

Principle 1 — Learning before performance

Improved output is valuable only when the system can still protect the development of capability underneath it.

Principle 2 — Preserve consequential cognitive work

Assistance should support participation without removing every meaningful decision, check or attempt.

Principle 3 — Guidance enables agency

Scaffolding is not the opposite of learner agency. Appropriate guidance gives learners better tools for questioning, checking and acting independently.

Principle 4 — Evidence belongs to a domain

Checking and judgement should use the standards, risks, methods and consequences of the relevant discipline or workplace.

Principle 5 — Visibility without surveillance

Make selected consequential moments visible. Do not collect everything merely because it can be collected.

Principle 6 — Conversation strengthens evidence; it does not prove authorship

Adaptive follow-up can reveal connections and responsiveness that fixed artefacts cannot. It remains one evidential source among several.

Principle 7 — Feedback must change what happens next

Feedback becomes educationally significant when the learner interprets and acts upon it.

Principle 8 — Transfer requires changed conditions

Application should test adaptation, not only repetition.

Principle 9 — Confidence should rest on converging evidence

No single artefact, explanation or observation should carry more certainty than it can support.

10. Watchlist

Evidence Worth Monitoring

  • Replications of AI tutoring studies across ages, subjects, cultures and levels of prior knowledge.
  • Research that distinguishes assisted performance from retained learning and later transfer.
  • Evidence about which AI guardrails preserve productive cognitive work without excluding learners.
  • The validity, reliability and workload implications of professional conversation and adaptive questioning.
  • Feedback designs that improve learner action rather than merely increasing comment volume.

System Signals Worth Monitoring

  • TEQSA and other quality-assurance movement from integrity guidance toward programme-level learning assurance.
  • Changes to NZQA guidance on AI, assessment, moderation and evidence.
  • Provider adoption of multi-point evidence strategies.
  • Learner responses to additional friction, checkpoints and professional conversation.
  • Whether visible-thinking approaches strengthen agency or drift into surveillance.
  • Whether vocational providers begin recognising naturally occurring evidence without over-bureaucratising practice.

Framework Tests Worth Running

  • Where do the Five Conditions fail to distinguish strong from weak learning design?
  • When do Check and Explain Judgement overlap too heavily to remain useful as separate prompts?
  • What counts as feedback from the work itself, and how should learners respond?
  • Which conditions are essential at different levels of learner expertise?
  • Can the model improve educator judgement without becoming another rubric?

11. Conclusion

The central change is not that learners can now get answers from machines.

Learners have always received assistance—from teachers, peers, books, worked examples, workplaces and one another. The important change is the speed, fluency and availability with which assistance can now produce the visible result.

That weakens the old shortcut between answer and capability.

Education does not need to make every thought visible. It does need enough well-chosen evidence to support confidence that the learner can understand, judge, adapt and take responsibility.

The Five Conditions for Visible Thinking offer one practical language for that work:

  • Attempt, calibrated to readiness;
  • Question, supported by guidance;
  • Check, grounded in the discipline;
  • Explain Judgement, tested through adaptive follow-up;
  • Apply, under changed conditions;
  • all resting on context and scaffolding;
  • all strengthened by feedback that changes what happens next.

This is not a validated solution. It is a design proposition that can now be tested more honestly.

The durable question is not whether the learner used AI.

It is not even whether the answer was correct.

It is:

What thinking remained with the learner—and what can they now do because of it?

References

Key Questions

Why is a correct answer no longer sufficient evidence of learning?

Because AI and other forms of assistance can produce similar visible results while preserving very different amounts of learner reasoning, judgement and understanding. The answer remains relevant, but it needs to be interpreted alongside other evidence.

Does visible thinking mean recording every step?

No. Exhaustive capture can increase workload and surveillance without improving judgement. Visible thinking should focus on selected consequential moments.

Should learners always attempt a task before receiving help?

No. Attempt should be calibrated to prior knowledge, task complexity and learner readiness. The aim is productive effort, not unsupported struggle.

Does professional conversation prove that learning occurred?

No. Adaptive conversation can strengthen evidence by testing connections, responsiveness and judgement. It remains one source within a wider evidence ecology.

What role should AI play?

The framework does not prescribe one position. AI can be encouraged, constrained, peripheral or absent. The important question is which cognitive work the learner must still perform.

Continue Exploring

Subscribe

Subscribe to receive future Signal Intelligence briefings on AI, education, capability, assessment, trust and the infrastructure of human learning.


Discover more from THISISGRAEME

Subscribe to get the latest posts sent to your email.


Comments

Kia ora! Hey, I'd love to know what you think.

Discover more from THISISGRAEME

Subscribe now to keep reading and get access to the full archive.

Continue reading