How First-Year Students Report Using AI in Permitted Assessments

Dashboard of post-activity metacognitive reflections following a structured SAGE assessment

Research Context

The Challenge: Whilst institutional policies on generative AI proliferate across higher education, task-proximal evidence of how students report navigating these tools in authentic assessment contexts remains limited. Assumptions about student behaviour often circulate without evidence collected close to an actual assessment event.

What We Studied: This research examined first-year ICT students at Central Queensland University during a supervised assessment in which generative AI use was explicitly permitted under institutional AI Collaboration guidelines. The assessment was structured around the five SAGE learning-process steps used in the activity: Generate, Evaluate, Refine, AI Critic and Reflect. Students completed a sequence of SAGE tasks. In selected tasks, they analysed a peer-reviewed article and used it, together with additional lecturer-provided sources, as authoritative anchors for evaluating and orchestrating AI-generated output. They then completed an embedded 12-item reflection instrument.

How We Collected Data: During Term 2, 2025, 167 student responses were selected for analysis from a cohort of approximately 300 students. The analysed responses addressed prompt frequency, source checking, revision and interaction strategies through structured metacognitive questions; 163 of the 167 also contained usable open-ended reflections. The reflection instrument was embedded directly after the structured SAGE activity.

Why This Matters: This dashboard presents exploratory descriptive evidence of how students understood and reported their orchestration of generative AI immediately after completing a complex, source-anchored SAGE activity. It examines metacognition following structured practice rather than general perceptions of AI in the abstract.

Unit: COIT11239 (Professional Communication for ICT) Dataset: 167 analysed responses from a cohort of approximately 300 Context: First-year, first-term students Framework: SAGE (Structured AI-Guided Education)

Key Findings Summary

The dashboard presents descriptive findings across the reflection items. The 4--8 prompt band was the most frequently selected category (55.1%), higher-frequency verification-oriented options totalled 73.0%, and verification accuracy was the most frequently selected challenge (77.8%). English writing confidence was selected as an influence by 46.7%. These aggregate findings do not establish respondent-level co-occurrence, performance effects or stable learner profiles.

73%
Reported Source Checking
Higher-frequency verification options
55.1%
4--8 Prompt Band
Most frequently selected category
77.8%
Verification Challenge
Most frequently selected difficulty
46.7%
English Confidence
Selected as an AI-use influence
Reader note applying to all questions: The questions below followed a complex, structured SAGE activity with defined Generate, Evaluate, Refine, AI Critic and Reflect stages. Students were taught and required to practise generative AI orchestration using authoritative sources. The instrument was therefore a post-activity metacognitive reflection, not a decontextualised perception survey. Its responses provide indirect evidence of how students understood and transferred orchestration skills into the assessment process.
1. Initial Approach Strategies (Q1)
What we asked: How did you begin interacting with the AI tool: conversation style, upload everything at once, or something different for each section?
Meaning Behind the Data: Conversational interaction (39.5%) and article-first sequential processing (36.5%) were the two most frequently selected starting approaches. The monolithic upload-and-extract option was selected by 7.2%. These selections describe declared starting strategies and do not establish interaction quality or the reasons underlying them. The study did not use statistical clustering or assign respondents to mutually exclusive profiles, so variation across items is reported by item and conceptual domain rather than as empirically identified student types.

Initial Approach Distribution (Q1)

How students began their AI interaction
2. Prompt Frequency Patterns (Q2)
What we asked: How many separate times did you interact with the AI during this assessment—just once, a few times, or many iterations?
Meaning Behind the Data: The 4--8 prompt band was the most frequently selected category (55.1%). Prompt frequency was not linked to grades, prompt logs or output-quality measures, so the bands are descriptive rather than indicators of effective or ineffective interaction.

Prompt Frequency Distribution (Q2)

Reported number of prompts used
3. Depth of Engagement: Revision Strategies (Q3)
What we asked: What did you do with AI outputs—use them as-is with minor tweaks, rewrite completely in your own words, or ask follow-up questions for refinement?
Meaning Behind the Data: Revision-oriented options were selected frequently: 37.7% reported rewriting after checking the article and 43.1% reported follow-up questioning or iterative refinement. Surface-level options were selected less frequently. These are declared practices rather than independently observed revision traces.

Revision Strategy Breakdown (Q3)

Are students "Copy-Pasting"?

Detailed Breakdown

Four distinct revision approaches
Follow-up Questions (43.1%, n=72)
Iterative refinement through continued dialogue
Verification-Rewrite (37.7%, n=63)
Checked article then rewrote AI output
Minor Formatting (8.4%, n=14)
Surface-level adjustments only
Combined Responses (6.0%, n=10)
Merged multiple AI outputs
4. Reported Source-Checking in a SAGE-Aligned Context (Q4)
What we asked: How often did you check AI responses against the original article—regularly, occasionally, or rarely?
Meaning Behind the Data: The assessment included SAGE-aligned verification prompts. In this context, 73.0% selected higher-frequency source-checking options, while 5.4% selected minimal verification. These values describe declared practice and do not demonstrate that SAGE caused the distribution.

Verification Behavior (Q4)

Did they check the article?
5. Assessment Components Requiring Rewriting (Q5)
What we asked: Which part of the assessment required the most rewriting after receiving AI help—summary, analysis, recommendations, or something else?
Meaning Behind the Data: Specific recommendations was the most frequently selected component requiring rewriting (31.1%), followed by article summary (24.6%). The item records students' perceptions and does not identify whether rewriting arose from AI limitations, task difficulty, prompt quality, domain knowledge or assessment expectations.

Reported Rewriting Requirements (Q5)

Components students identified as requiring rewriting
6. Reported Influences on AI Use (Q7)
What we asked: What factors influenced your decision to use AI—time pressure, English confidence, maintaining your voice, or something else? (Select up to 2)
Meaning Behind the Data: Time available was the most frequently selected influence (73.1%), followed by English writing confidence (46.7%). The latter suggests that some respondents may perceive AI as writing support; language background and outcomes were not analysed, so no equity effect is established.

Top Influencing Factors (Q7)

Why do students turn to AI? (Beyond just "Time")
7. Engagement Patterns: Interaction Strategies (Q8)
What we asked: Which best describes your overall interaction pattern—provided all info at once, built responses step-by-step, or engaged multiple times with verification checks?
Meaning Behind the Data: Multiple interactions interspersed with source checking was the most frequently selected pattern (58.1%), while comprehensive single-input use was selected by 3.0%. These self-reports are compatible with checkpoint-based engagement, but the distribution cannot be attributed causally to SAGE and does not establish naturally occurring behaviour.

Engagement Pattern Distribution (Q8)

Interaction strategies employed during the assessment
8. Reported Challenges (Q9)
What we asked: What was most difficult about using AI effectively—creating useful prompts, verifying accuracy, maintaining your authentic style, or something else? (Select up to 3)
Meaning Behind the Data: Verifying AI accuracy was the most frequently selected challenge (77.8%), followed by maintaining authentic style (64.7%) and producing useful prompts (50.9%). These selections identify perceived difficulties; they do not directly measure competence.

Challenge Hierarchy: The Friction Points (Q9)

What was hardest for students?
9. Intended Strategy Changes (Q10)
What we asked: If you could redo the assessment, would you change your AI approach—engage more iteratively, verify more thoroughly, or keep the same strategy?
Meaning Behind the Data: Engaging more iteratively was the most frequently selected intended change (46.7%), followed by more strategic use (28.1%) and more thorough verification (25.7%). Because the item was coded at option level, the selections cannot be combined into a unique percentage of students who would change strategy.

Retrospective Pivot (Q10)

"If you could do it again, what would you change?"
10. Requested Institutional Support (Q11)
What we asked: What kind of institutional support would most help you use AI correctly and ethically—verification guidance, prompt examples, practice opportunities, or integrity frameworks?
Meaning Behind the Data: Verification guidance was the most frequently selected support option (51.5%), followed by prompt examples (33.5%) and practice with feedback (29.3%). Because this item was analysed as option-level selections, percentages should not be added to estimate a unique respondent total.

Prioritized Support Needs (Q11)

Specific interventions requested by students
11. Thematic Analysis of Concerns and Needs (Q12)
What we asked: In your own words (2-3 bullet points), what could CQUniversity do more to help you use AI correctly and ethically?
Meaning Behind the Data: The 163 usable open-ended responses were analysed thematically, with computational lexical outputs used as audit and sensitising aids. Policy clarity was the largest dominant response-level category (44.8%), followed by practical training (35.0%) and support infrastructure (20.2%).

Primary Concern Distribution

What topics dominated reflections?

Desired Training Curriculum

Specific topics requested for workshops
12. Provisional Institutional Response Options
What this section provides: Provisional response options informed by reported needs; these actions were not tested in the study.
Interpretive boundary: The proposed options translate reported concerns and support requests into actions that institutions could evaluate. The present study did not test whether these interventions improve learning, integrity or confidence.
Immediate Term (0-3 months)
Policy Clarification & Legitimation
  • Publish unified AI policy: Define acceptable use with concrete examples differentiating "idea generation" vs "drafting" vs "editing assistance"
  • Address three ambiguity layers: Definitional (what counts as AI assistance?), procedural (disclosure requirements), evaluative (how will it be assessed?)
  • Establish transparency protocols: Create AI declaration templates for assessments specifying usage levels
  • Communicate institutional stance: Move from "detection and punishment" to "partnership and development" messaging
Mid Term (3-12 months)
Competency Development via Exemplars
  • Develop verification frameworks: Create discipline-specific heuristics for fact-checking AI outputs (51.5% requested this)
  • Build prompt engineering library: Annotated case studies showing effective vs ineffective prompts across assessment types (33.5% demand)
  • Establish practice infrastructure: Consequence-free "sandbox" environments with formative feedback on AI orchestration quality (29.3% requested)
  • Create comparison exemplars: Side-by-side demonstrations of legitimate AI collaboration vs problematic over-reliance
  • Integrate into curriculum: Embed AI literacy explicitly within core units rather than treating as supplementary skill
Long Term (12-24 months)
Assessment Transformation & Infrastructure
  • Transition to process-based assessment: Evaluate AI interaction quality through prompt sequences, verification protocols, and revision strategies rather than output alone
  • Legitimate hybrid authorship models: Develop contribution frameworks explicitly articulating human value-addition in AI-mediated work
  • Emphasize application over extraction: Design assessments requiring context-specific reasoning where AI struggles (recommendations, contextual analysis) rather than information retrieval where AI excels (summarisation)
  • Consult student representatives: Include student perspectives when developing and evaluating AI policies and support resources
  • Establish AI Support Hub: Centralized resource providing workshops, drop-in consultations, and discipline-specific guidance
13. Key Interpretive Implications
From data to interpretation: The aggregate patterns suggest three cautious implications for AI literacy, assessment design and institutional support:

Verification-Confidence Tension

At the cohort level, 73.0% selected higher-frequency verification-oriented options and 77.8% selected verification accuracy as the most frequently reported challenge. This construct-level convergence may suggest:

  • Reported source checking was not accompanied by strong confidence in verification adequacy
  • Uncertainty may reflect limited domain knowledge, metacognitive calibration or unclear assessment expectations
  • Students may benefit from explicit verification instruction, worked examples and calibration feedback
  • Respondent-level and trace-based research is needed before individual co-occurrence or competence can be established

Temporal Reallocation, Not Reduction

Although time was the most frequently selected influence (73.1%), students also reported iterative workflows and revision-oriented practices. This cohort-level pattern suggests:

  • AI adoption transforms rather than eliminates intellectual labour requirements
  • Students invest substantial cognitive effort in verification and revision activities
  • Time gains from AI-assisted information gathering are reinvested in critical evaluation and contextual application
  • The "efficiency" narrative obscures the reality of shifted rather than reduced cognitive load

Equity-Relevant Consideration

English writing confidence was the second-most frequently selected influence (46.7%), suggesting a possible writing-support role for some students:

  • The finding does not establish improved language learning, equalised outcomes or reduced disadvantage
  • Future research should examine language background, access, dependency and learning outcomes at respondent level
  • Institutional policy should consider possible writing-support benefits alongside authorship, access and learning risks
  • The present result is an equity-relevant signal rather than evidence of an equity effect