How can we stop or detect AI use?
This approach often adds restrictions, surveillance, or detection measures to an existing task without first asking whether the task still measures the intended learning.
SAGE is a validated pedagogical framework that guides students through structured AI interaction — and verifies they own the result.
The Structured AI-Guided Education (SAGE) framework is a validated pedagogy for developing students' capacity to orchestrate generative AI outputs rather than passively accept them. It treats GenAI as an assistive tool whose outputs require explicit, evidence-based decisions to accept, modify, or reject; it is neither an authority to mimic nor a tool to ban.
SAGE was developed through three years of sustained empirical research involving more than 1,500 students across 30+ studies, spanning cybersecurity management, data analytics, Professional Communication, and systems analysis and design, delivered across campuses in Australia, China, and the United Kingdom. The framework has been referenced by researchers at institutions including the University of Warwick and the National University of Singapore.
Version 3.0 of this guide defines the canonical SAGE Student Orchestration Levels and retains the sixth step — Defend — introduced in Version 2, which places assurance of individual competence within the cycle itself rather than relying on an additional supervised assessment. The complete evidence base, discipline-specific examples, implementation templates, and detailed rubrics are available in Version 1 of this guide.
Assessment redesign in the age of generative AI should not begin only with the question, “How can we reduce the risk of students using AI?” It should begin with the learning outcome, the capability students must demonstrate, and the evidence required to show that learning has occurred.
This approach often adds restrictions, surveillance, or detection measures to an existing task without first asking whether the task still measures the intended learning.
Define the capability, decide whether it should be demonstrated with or without GenAI, and select evidence that validly demonstrates student learning.
Begin with the learning outcome rather than the existing assessment format.
Identify foundational knowledge, reasoning, or procedural competence that must remain personally internalised.
Consider whether graduates will be expected to use, test, challenge, and take responsibility for AI-supported work.
Select evidence that matches the learning: performance, explanation, application, artefacts, decisions, or supervised defence.
A written submission is not always the most valid evidence of these capabilities. Similarly, conducting an assessment during a scheduled class does not automatically make it secure. Evidence should be selected because it matches the capability, not merely because it is familiar or convenient.
SAGE operationalises this approach. Steps 1–5 structure critical and responsible human–AI collaboration. Step 6, Defend, provides direct verification where the intended capability requires assurance of individual understanding, judgement, or performance.
At the core of SAGE is a six-step cycle that moves students from passive AI consumption to critical AI orchestration, culminating in supervised demonstration of the competency developed through the preceding five steps. The cycle is designed to produce graduates who can deploy generative AI as a professional tool within the constraints of their discipline. They will be taught skills to evaluate AI outputs against industry standards, regulatory requirements, and domain-specific evidence rather than accepting them uncritically. This is the competency that employers now expect. It is not the ability to prompt an AI tool, but rather the ability to judge, correct, and take professional responsibility for what AI produces.
The cycle is sequenced to move students progressively through Bloom's revised taxonomy. Steps 1–3 operate at the Apply and Analyse levels: students generate output, compare it against domain standards, and modify it with evidence. Step 4, AI Critic, inverts the AI relationship: the student assigns AI a critical role and must audit the critic by evaluating the evaluator, operating at Bloom's Evaluate level. Step 5 requires metacognitive synthesis. This is the highest cognitive operation in the taxonomy, in which students articulate patterns of AI strength and failure across the task. Step 6 closes the cycle by requiring the student to demonstrate, under supervised conditions, that the competency developed through the preceding steps has been individually internalised rather than merely documented.
SAGE operates through two stages that move students from guided practice to independent application with supervised assurance. This progression mirrors the cognitive trajectory defined by Bloom's revised taxonomy: Stage 1 scaffolds the lower-order operations of Apply and Analyse under instructor guidance, while Stage 2 requires students to operate independently at the Evaluate and Create levels before demonstrating competency under supervised conditions. The two-stage design ensures that open and assurance tasks are both addressed within a single integrated teaching sequence rather than treated as separate assessment instruments.
Students practise the SAGE cycle under instructor guidance with structured prompts, worked examples, and immediate feedback. Scaffolding is heavy. The goal is to build the evaluation, refinement, and reflection skills that students will apply independently in Stage 2.
Students apply the SAGE cycle autonomously in summative tasks with reduced scaffolding. Process logging remains embedded in the open task for formative purposes. Assurance of individual attainment is placed in the Defend step under supervised conditions.
Students use a generative AI tool to produce an initial output for the task at hand. In Stage 1, standardised prompts are provided to ensure comparable outputs across the cohort. In Stage 2, students design their own prompts but must first produce a human baseline (outline, notes, or preliminary analysis) before engaging the AI tool. This baseline demonstrates independent understanding of the problem and prevents complete delegation.
Students compare the AI output against authoritative sources relevant to their discipline: industry standards, regulatory frameworks, peer-reviewed research, clinical guidelines, or case-study constraints. They identify what is present, what is missing, and what is incorrect. In Stage 1, instructors provide checklists and evaluation templates. In Stage 2, students identify evaluation criteria independently.
Students modify the AI output on the basis of their evaluation. Each modification is documented as an accept, modify, or reject decision with a brief evidence-based justification citing specific standards, sources, or contextual constraints. The refinement produces a professional-standard artefact that the student can defend as their own intellectual product.
Students ask GenAI to act as a critic, examiner, reviewer, auditor, or devil's advocate. They must then evaluate the AI's criticism rather than automatically accepting it, deciding which recommendations to accept, modify, or reject and providing reasons. This step develops metacognitive sophistication and strengthens the student's authority over the AI.
Students produce a metacognitive analysis documenting the AI's strengths and weaknesses as observed during the task, the specific domain errors identified, the corrective reasoning applied, and the broader implications for AI reliability in their discipline. Functional reflections — those that critique domain content directly — are distinguished from procedural reflections that merely describe workflow benefits.
Students demonstrate, under supervised conditions, that they can reproduce or explain the reasoning documented in their open assessment without reliance on the artefacts that produced it. The format is determined by disciplinary context (see Section 6). The defining criterion is that the student shows they own the competency — they can identify risks, explain trade-offs, justify decisions, and respond to challenge questions — not merely that they submitted a document containing these elements. This step carries the institutional assurance function: it is the evidence that the intended learning outcomes have been individually achieved.
The SAGE Student Orchestration Levels describe the observable quality of a student's collaboration with generative AI. They are different from an AI guidance or permission level. A guidance level states what use of GenAI is permitted in an assessment; an orchestration level describes the capability that the learning design seeks to develop and that assessment criteria may evaluate.
The levels represent a developmental progression. Passive Acceptor is a baseline condition that SAGE is designed to remediate, rather than a desired learning outcome.
Accepts AI output with minimal critical evaluation, disciplinary verification, or justification.
Filters AI output using disciplinary knowledge and evidence, accepting, modifying, or rejecting selected elements.
Systematically combines human and AI contributions through explicit trade-off reasoning, contextual judgement, and iterative refinement.
Proactively identifies gaps, assumptions, and bias and develops a traceable synthesis, reframing, solution, or contextual adaptation that extends beyond both the original human position and the AI output.
Steps 1–5 provide the process through which stronger orchestration capability can develop. Generate creates material for examination; Evaluate introduces authoritative anchors; Refine makes human contribution visible; AI Critic tests the developing work; and Reflect requires students to account for their decisions. Step 6, Defend, verifies whether the student has personally internalised the reasoning represented in the final work.
The Defend step is format-agnostic. The principle remains constant: the student demonstrates competency under conditions where the process cannot be simulated. The following table illustrates how Defend may be implemented across representative disciplines.
| Discipline | Defend Format | What the Student Demonstrates |
|---|---|---|
| Cybersecurity | Timed risk-scoring exercise or incident response viva | Classifies and prioritises threats for an unseen scenario using the same frameworks applied in the open task. Explains why specific controls were selected over alternatives. Responds to challenge questions on regulatory alignment. |
| Programming | Supervised code walkthrough or live debugging session | Explains design decisions in submitted code. Debugs a seeded error in a related module under observation. Demonstrates understanding of logic, structure, and trade-offs rather than surface familiarity with the output. |
| Health Sciences | Structured clinical reasoning viva or case interpretation | Applies clinical decision-making to a new patient scenario. Justifies care plan modifications against clinical guidelines. Identifies where AI-generated recommendations require human override based on patient-specific factors. |
| Business | Oral defence of strategic recommendation with examiner challenge | Presents and defends a strategic position. Responds to examiner objections with evidence-based reasoning. Demonstrates understanding of stakeholder constraints, market assumptions, and ethical implications. |
| Education | Teaching demonstration with reflective justification | Delivers a short teaching segment based on a lesson plan developed with AI assistance. Explains pedagogical choices, adaptation for specific learner needs, and why certain AI-suggested approaches were modified or rejected. |
| Engineering | Design defence with constraint interrogation | Defends design decisions under examiner questioning. Explains trade-offs between competing requirements. Demonstrates that safety, sustainability, and regulatory considerations were understood, not merely included in the document. |
| Law | Moot argument or case analysis viva | Argues a legal position with reference to relevant authorities. Responds to judicial challenge on the strength of the analysis. Distinguishes AI-generated legal summaries from independently reasoned legal arguments. |
The inclusion of a supervised assurance step is not a theoretical preference. It is an empirically grounded response to a documented structural limitation of unsupervised assessment under GenAI-rich conditions.
In a structured audit of 25 group submissions across two cybersecurity management cohorts, assessments designed using the full SAGE protocol, incorporating base prompts, structured decision tables, mandatory AI interaction logs, and reflective commentary, were examined for process fidelity using a five-check protocol. Full traceability between documented AI outputs and human evaluation claims was not achieved in most student submissions. Only 3 of 25 submissions (12%) produced evidence chains that were substantially auditable. The remaining submissions exhibited logical inconsistencies between reported AI positions and appended outputs, compliance-pattern text in evaluation cells, and structural indicators consistent with audit trail simulation (Elkhodr & Gide, 2026, STEM Education).
These findings are consistent with sector-wide guidance. TEQSA's 2025 resource Enacting Assessment Reform in a Time of Artificial Intelligence identifies structural assessment redesign, not detection, as the sustainable response. SAGE with Defend responds by embedding AI-supported learning and supervised assurance within one continuous six-step pedagogical process.