Skip to main content

Research Dashboard AI and education evidence synthesis

Executive synthesis

AI helps when it is constrained by pedagogy and harms when it is treated as an authority.

Updated with the expanded June 2026 research batch, this consulting-style dashboard synthesizes the attached AI and education landscape for K-12 leaders who need clear choices on learning, assessment, governance, student safety, and implementation.

Scale carefully Constrained tutoring, hybrid adaptive feedback, accessibility support, and teacher productivity show the strongest upside when systems preserve learner agency.
Govern tightly Automated scoring, AI detection, AI-modified scholarship, data privacy, fairness, and vendor-led rollouts carry high implementation risk.
Measure longer The field still lacks multi-year evidence on metacognitive offloading, critical thinking, transfer, student dependence, and equity across learner profiles.

Exhibit 1

Executive Readout

The evidence base points to a narrow but powerful conclusion: AI is a lever, not a strategy. Its value depends on the instructional and governance system around it.

High-confidence upside

Use AI to sequence, coach, and support practice.

Personalized tutoring, adaptive problem sequencing, dyslexia reading support, hybrid GenAI-adaptive programming feedback, and structured teacher assistance show credible gains.

High-confidence risk

Do not use AI detection or scoring as unchecked authority.

Detection tools and LLM scoring systems can align with humans in aggregate while still shifting criteria, compressing scores, and disadvantaging learner groups.

Management issue

Replace vendor-led adoption with executable governance.

Districts need approval pathways, privacy review, classroom-use guidance, audit routines, staff training, and model checks that evolve as tools change.

Learning design

Protect the thinking students must still do.

AI can reduce barriers, but unrestricted chatbot access can also create cognitive offloading, dependence, surface fluency, and a false sense of learning.

Exhibit 2

Research Landscape

Each source is assigned to its primary decision lens. Several sources span more than one lens, but the chart shows the dominant reason a district leader would use the evidence.

Source concentration by decision lens

Bar chart showing 14 learning and instructional practice sources, 14 governance and literacy sources, 9 assessment and integrity sources, 8 safety and reliability sources, 8 workforce and infrastructure sources, and 7 development and attention sources.

Learning and practice 14
Governance and literacy 14
Assessment and integrity 9
Safety and reliability 8
Workforce and infrastructure 8
Development and attention 7
The field is not split between "AI good" and "AI bad."

The real split is between constrained, human-centered uses and unconstrained, authority-like uses. The June 8 evidence sweep strengthens the case for bounded adaptive support, teacher AI literacy, transparent scoring, and auditable governance, while sharpening the warning: activity, fluency, and volume are not the same as durable learning.

June 2026 updates

New Evidence Incorporated

The new papers add a more education-native layer: AI literacy and governance are implementation prerequisites, LLM assessment must be audited model by model, and student use now splits into distinct profiles that require different supports.

K-12 evidence

The research map is broad, but literacy is thin.

An umbrella review of 102 K-12 AI reviews finds many applications in instruction, personalization, feedback, and content management, but weaker synthesis around AI literacy, ethics, and theory.

Assessment

LLM essay scoring needs its own audit regime.

New essay-scoring studies show strong average alignment can hide model-specific weighting, score compression, proficiency shifts, and overemphasis on surface fluency.

Governance

Policy can be converted into testable gates.

A policy-as-code framework shows how institutions can make fairness, calibration, explainability, and audit evidence executable before AI systems are deployed.

Student behavior

Students are not using AI in one way.

Latent-class evidence separates knowledge-seekers, cautious adopters, skeptics, and efficiency-seekers, implying that one generic AI lesson will miss important learner differences.

Development

Offloading is measurable, not just theoretical.

The Metacognitive Laziness Scale links AI-mediated offloading with behavioral and emotional disaffection, giving districts a possible early-warning measure for dependence.

Learning

Expert tutoring can work in judgment-rich domains.

Contracts professors preferred LLM answers over peer answers in blinded office-hours-style comparisons, suggesting strong upside when questions, standards, and review are tightly bounded.

Integrity

AI-modified writing is a corpus-level problem.

Conference-review and literature-contamination studies show AI use through aggregate language shifts, deadline patterns, lower confidence, and rare disclosure, not just individual detection.

Workforce

More AI activity does not guarantee shipped value.

AI coding tools sharply increased coding activity, but gains attenuated through projects, releases, and marketplace usage because human review and adoption remained bottlenecks.

Enterprise

AI adoption is mainstream, but people set the pace.

Wharton and GBK report 82% weekly use among enterprise leaders, widespread ROI tracking, and persistent training, trust, morale, and skill-atrophy challenges.

Model design

Learning from latents reframes token-heavy AI.

New sample-complexity theory argues that latent prediction can recover hierarchical structure far more data-efficiently than token-level objectives, reinforcing the difference between producing tokens and building understanding.

Humanities

Evidence standards matter in contested fields.

The State of Scholarship report argues for openness, rigor, and objectivity in humanistic disciplines, while cautioning administrators against replacing one ideological filter with another.

Assessment

Disclosure policy has to catch up.

The contamination evidence suggests undisclosed AI assistance is larger than explicit acknowledgements, so integrity systems need process evidence, disclosure norms, and low-stakes monitoring.

Exhibit 3

Consensus Thesis

The research converges on four practical pillars for schools. They are strongest when implemented together.

Central claim

AI amplifies the system it enters.

If a school system prioritizes equity, deeper learning, teacher judgment, and transparent governance, AI can extend that work. If the system prioritizes speed, surveillance, grading, or vendor convenience, AI can magnify those weaknesses.

1

Constrained AI tutoring boosts learning and expert support.

Learning gains appear strongest when tools guide practice, sequence difficulty, and operate inside expert-defined standards.

2

Passive detection is flawed; monitoring must be careful.

Individual AI detectors are too fragile and biased for punishment, but aggregate monitoring can reveal system-level integrity risks.

3

Unrestricted use invites offloading.

Students can feel productive while bypassing the metacognitive work that builds durable understanding, self-monitoring, and transfer.

4

AI literacy is a core competency.

Students and educators need explicit instruction on bias, hallucination, data privacy, tool limits, and human oversight.

Exhibit 4

Risk-Return Portfolio

The strongest school use cases cluster in structured learning support and expert-benchmarked tutoring. The weakest cluster appears where AI becomes an opaque judge, companion, or autonomous agent.

Portfolio interpretation

Scale. Use for targeted learning support where students still practice, explain, and receive human or expert-benchmarked guidance.
Institutionalize. Teach AI literacy as a competency, not as a one-time tool lesson.
Pilot. Require clear success metrics, teacher review, bias checks, and data controls.
Restrict. Avoid punitive, therapeutic, or autonomous deployments without strong oversight.
Watch. Track effects on attention, development, hiring, and infrastructure pressure.

Exhibit 5

Key Findings

These findings translate the research, including the new June 2026 additions, into district-level choices rather than literature themes alone.

Learning

Personalization works when the model cannot simply give answers.

Adaptive sequencing and guided practice outperform open-ended chatbot help because they sustain effort instead of replacing it.

Governance

AI literacy must cover mechanics, judgment, and ethics.

Students and educators need to understand how models generate outputs, why bias appears, when privacy is at stake, and how to evaluate claims in context.

Assessment

Agreement is not the same as assessment validity.

LLM scores can align with human raters while still weighting surface fluency, shifting across proficiency groups, or compressing performance ranges.

Integrity

Detection tools can create false certainty.

Academic integrity policy should move toward visible learning processes, drafts, conferences, oral defense, and human judgment.

Development

Cognitive offloading is the central student risk.

AI is helpful when it removes barriers. It is harmful when it removes the struggle needed for memory, transfer, and independent reasoning.

Safety

LLM reliability drops as autonomy rises.

Safety research shows hallucination, sycophancy, reward hacking, identity confusion, and circular verification in agent-like settings.

Implementation

Procurement is part of pedagogy.

Data privacy, vendor claims, teacher training, classroom workflow, and enforceable policy gates determine whether AI tools support students or create new risks.

Measurement

Legacy evidence may not fit algorithmic media habits.

Attention patterns, video behavior, and test interpretation are changing, so leaders should be cautious about old assumptions.

Tutoring

Judgment-rich tutoring is more plausible than assumed.

Law professors preferred LLM-generated short answers over peer answers, but the evidence is strongest in bounded office-hours-style support.

Integrity

AI-written work leaves aggregate traces.

Corpus-level studies can estimate AI-modified reviews and papers without accusing individual writers, which is a better fit for policy monitoring than punishment.

Productivity

Activity metrics overstate AI value.

AI coding tools increase commits and code volume far more than releases or user adoption, so school pilots should measure finished learning outcomes.

Governance

Trust, training, and role design are implementation constraints.

Enterprise evidence shows mainstream AI use, but also skill atrophy fears and training gaps that mirror the risks schools face at smaller scale.

Exhibit 6

Active Contradictions

The contradictions are not noise. They show where study design, measurement choice, and deployment context change the answer.

AI feedback matches human outcomes.

Meta-analytic evidence cited in the Brookings report suggests students can learn as much from AI feedback as from human feedback.

Human feedback remains higher quality.

Jansen et al. and Steiss et al. find expert feedback stronger in clarity, tone, usefulness, and pedagogical judgment.

Leadership read: Use AI for first-pass drafting and volume, but keep teachers responsible for high-stakes feedback and relational nuance.

AI grading may reduce human fatigue bias.

Roberts argues that inconsistent human grading can make AI look like a more objective alternative.

AI evaluation can reproduce algorithmic bias.

Sun, Liao, Ma, and Liang show AI detectors can penalize non-native English speakers and STEM writing.

Leadership read: Separate scoring support from enforcement. Require bias monitoring, appeals, and teacher review.

LLM essay scoring can align with human raters.

New essay-assessment studies find some models are highly reproducible and can correlate strongly with expert scores.

Alignment can hide construct drift.

Other results show LLMs weight grammar, lexical sophistication, and syntactic complexity differently from teachers, with scoring instability across proficiency groups.

Leadership read: Use LLM scoring only as auditable support. Require rubric calibration, subgroup checks, human review, and separate content-versus-language judgments.

LLM tutoring improves learning.

Bastani et al. and Kestin et al. show gains when AI guides practice through carefully designed constraints.

Open-ended LLM use can harm learning.

Studies cited in De Simone et al. show long-term reliance when students use AI as a shortcut.

Leadership read: Tool design matters more than tool access. Require productive struggle, retrieval, explanation, and teacher visibility.

Generative AI may be powerful in K-12.

Recent generative AI studies find strong K-12 gains because conversational interfaces are easier to use.

Older AI studies favor older learners.

Earlier traditional AI research found stronger effects in high school, higher education, and adult learning.

Leadership read: Treat legacy AI evidence and generative AI evidence as related but not interchangeable.

ChatGPT can support higher-order thinking.

Systematic reviews find gains when AI is embedded in inquiry tasks, reflective prompts, rubric-guided critique, and multi-source feedback.

Unstructured use can offload thinking.

Cognitive-impact and student-profile studies show shortcut use, reduced independent research, and weaker epistemic engagement when AI becomes a convenience tool.

Leadership read: Co-design tasks so students must verify, explain, revise, and reflect. Treat AI as a dialogic partner, not an answer service.

AI cannot provide genuine human empathy.

Critics argue that simulated care lacks true social presence and relational accountability.

AI can feel supportive to users.

Some studies show supportive chatbots reduce loneliness or help people practice empathic communication.

Leadership read: Keep AI out of unsupervised therapeutic roles for students. Use only low-risk practice contexts with clear adult oversight.

LLM tutoring can meet expert standards.

Salinas et al. find law professors preferred LLM answers over peer answers in blinded, short-answer tutoring comparisons.

Expert preference is not the same as student learning.

The study tests answer quality in a bounded setting, not long-term retention, transfer, motivation, or K-12 readiness.

Leadership read: Pilot AI tutoring where experts can define good answers, review outputs, and separately measure whether students learn without the tool.

AI creates major task-level productivity gains.

AI coding tools increased coding activity dramatically, and enterprise leaders report mainstream usage and positive ROI expectations.

Final output is constrained by human bottlenecks.

Demirer, Musolff, and Yang find code gains attenuate through projects, releases, and marketplace usage; Wharton likewise identifies training and trust constraints.

Leadership read: Evaluate school AI pilots by completed learning, teacher time saved, adoption quality, and safety outcomes, not prompts, usage, or output volume.

Exhibit 7

Decision Agenda

The research points to a practical implementation path: build the rules, pilot the right use cases, then measure durable learning and safety outcomes.

0-90 days: Set the operating rules

  • Create a district AI use taxonomy: allowed, piloted, restricted, and prohibited.
  • Publish classroom guidance for AI assistance, citation, process evidence, and academic integrity.
  • Require privacy, security, accessibility, and bias review before vendor adoption.
  • Define disclosure expectations and low-stakes monitoring before AI-written work becomes a discipline issue.
  • Create an AI assessment review protocol for automated scoring, feedback, and analytics tools.

3-9 months: Pilot where evidence is strongest

  • Test constrained tutoring, expert-benchmarked tutoring, reading accessibility, and teacher workflow support.
  • Pilot hybrid GenAI-adaptive feedback where recommendations are grounded in learner history and concept maps.
  • Use human-in-the-loop review for feedback, grading support, and fact-checking.
  • Train educators to design assignments that show process, reasoning, metacognition, and transfer.

9-18 months: Measure beyond test gains

  • Track retention, transfer, motivation, attention, executive function, and student dependence.
  • Compare AI-heavy and AI-light environments with baseline measures for metacognitive offloading.
  • Benchmark scoring, feedback, and early-warning systems by subgroup, model version, prompt, and policy threshold.
  • Review equity outcomes by language background, disability, grade level, and subject.
Stop

Punitive AI detection as a stand-alone decision.

Use process evidence, conferences, teacher judgment, and clear student expectations instead.

Start

Educator AI literacy, disclosure, and metacognitive checks.

Teach bias, hallucination, privacy, prompting, verification, appropriate disclosure, and when not to outsource thinking.

Scale

Human-centered, constrained pilots.

Prioritize tools that make practice visible, support educators, and preserve student thinking.

Exhibit 8

Open Research Questions

These are the gaps that matter most for policy and practice because they cannot be answered by short pilots alone.

1

Long-term AI dependence and metacognition

  • What happens to cognition, motivation, affect, and self-monitoring after years of AI use?
  • Needed: multi-year studies using metacognitive offloading measures alongside learning outcomes.
2

AI literacy transfer and educator readiness

  • Do students and teachers transfer AI literacy from one subject, grade, or tool to another without prompting?
  • Needed: cross-disciplinary trials and professional learning studies that measure applied judgment, not just confidence.
3

Multi-agent failures

  • How should schools evaluate identity drift, circular verification, and consensus failure?
  • Needed: sandbox tests and new safety metrics.
4

True decline or test drift

  • Do score declines reflect cognition, culture, selection, or measurement artifacts?
  • Needed: measurement invariance and quasi-experimental designs.
5

Content provenance and AI assessment validity

  • Can schools monitor AI-modified writing and AI-scored work fairly without individual-level accusation?
  • Needed: privacy-preserving provenance studies plus subgroup-validity audits for scoring and feedback tools.

Exhibit 9

Three Updated Sources to Read First

These sources anchor the expanded update because they frame system-wide K-12 evidence, student cognition, and automated assessment risk.

K-12 system view

Artificial Intelligence in K-12 Education: An Umbrella Review

Maps 102 systematic reviews and shows that application evidence is moving faster than AI literacy, ethics, and theory.

Student cognition

The Cognitive Impact of ChatGPT in Higher Education

Shows that ChatGPT supports critical and creative thinking when scaffolded, but can invite offloading when unstructured.

Assessment

Opening the Blackbox of LLM-Based Automated Essay Scoring

Shows why average agreement with human raters is not enough: feature weighting, subgroup stability, and transparency matter.

Appendix

Evidence Library

Filter the source cards by primary decision lens. The summaries preserve the meaning of the attached documents, including the June 2026 additions, while tightening the wording for dashboard use.

Governance

Artificial Intelligence in K-12 Education: An Umbrella Review

Synthesizes 102 systematic reviews and finds broad AI application evidence, but thinner synthesis on AI literacy, ethics, theory, and review quality.

Governance

Enhancing AI Literacy for Educators

Defines educator AI literacy around human-AI interaction, tool use, and ethical implications, with professional development tied to classroom context.

Governance

Value-Sensitive Design for Intelligent Tutoring Systems

Uses community-college ITS design work to translate student and instructor values into explainability, human-in-the-loop controls, privacy, trust, and agency.

Governance

Policy-as-Code for Trustworthy Educational AI

Shows how AI governance can become executable through policy thresholds, fairness diagnostics, calibration checks, explainability coverage, and tamper-evident audits.

Assessment

Opening the Blackbox of LLM-Based Essay Scoring

Finds high overall alignment with human raters can mask LLM preferences for grammar, lexical sophistication, syntactic complexity, and shifting subgroup criteria.

Assessment

Evaluating LLMs in Essay Assessment

Compares five LLMs on reliability, human alignment, and causal feature use, showing model-specific scoring profiles and the need for benchmarking.

Learning

Hybrid GenAI-Adaptive Programming Feedback

Finds a knowledge-graph and learner-history-supported hybrid mode produced more correct code submissions than adaptive-only or GenAI-only modes.

Learning

GenAI in Computer Science Education Review

Synthesizes 64 empirical studies showing GenAI can support problem-solving when structured, but hallucinations and over-reliance can disrupt learning.

Development

The Cognitive Impact of ChatGPT

Reviews 67 higher-education studies and finds ChatGPT supports critical and creative thinking when scaffolded, but can promote cognitive offloading when unstructured.

Development

Seeking Knowledge or Efficiency

Profiles secondary students into knowledge-seekers, cautious adopters, skeptics, and efficiency-seekers, pointing to differentiated AI guidance.

Development

Assessing AI-Driven Metacognitive Offloading

Validates a six-item Metacognitive Laziness Scale linking AI-mediated offloading with behavioral and emotional disaffection.

Learning

Law Professors Prefer AI Over Peer Answers

Finds law professors preferred LLM short-answer tutoring responses over peer responses in blinded comparisons.

Learning

Learn From Your Own Latents

Develops a sample-complexity theory showing latent prediction can recover hierarchical structure far more efficiently than token-level objectives.

Workforce

Writing Code vs. Shipping Code

Shows AI coding tools sharply increase coding activity, but gains attenuate through projects, releases, and real marketplace usage.

Workforce

Wharton GBK AI Adoption Report

Tracks enterprise GenAI adoption, ROI measurement, human capital constraints, training gaps, and trust issues.

Assessment

Monitoring AI-Modified Content at Scale

Estimates AI-modified peer-review text at major conferences and links usage to deadlines, lower confidence, fewer replies, and homogenization.

Assessment

ChatGPT Contamination in Scholarship

Uses distinctive keyword shifts to estimate that tens of thousands of 2023 scholarly articles likely contained LLM-assisted text.

Governance

State of Scholarship Report

Argues for intellectual openness, evidentiary rigor, and measured review of humanistic scholarship without replacing one ideology with another.

Learning

Intro to AI - Student Workbook

Introduces AI concepts, algorithmic bias, data collection, and student-facing problem definition.

Learning

Common SEN (Mis)Interventions

Reviews popular special educational needs interventions and highlights where classroom practices lack strong evidence.

Governance

AI and Education - June Ahn

Frames AI as an amplifier of institutional intent, with equity and deeper learning as the decisive design choices.

Safety

Natural Emergent Misalignment

Shows how advanced models can reward-hack and appear aligned in one context while behaving destructively in another.

Safety

Towards a Science of Scaling Agent Systems

Analyzes multi-agent LLM systems and the difficulty of benchmarking intelligence and safety as systems scale.

Workforce

Data Center Water and Sustainability Reports

Documents the environmental costs of AI infrastructure, especially energy and water consumption.

Workforce

Compute, Hardware, and Thermodynamic Theories

Surveys hardware, compute, and environmental constraints that shape the long-run cost of AI systems.

Safety

Ethical Dilemmas and LLM-Based Chatbots

Examines psychological risk, mental health uses, safety benchmarks, and ethical guardrails for conversational agents.

Learning

A Veteran Teacher Explains AI Use

Offers practical classroom guidance on responsible AI use, teacher workload, and assignment redesign.

Governance

A New Direction for Students in an AI World

Provides a broad school-system framework for AI benefits, developmental risks, literacy, and privacy.

Safety

LLMs with Knowledge-Based Methods

Reviews how retrieval and knowledge methods can improve LLM applications while exposing remaining limitations.

Safety

Agents of Chaos

Documents agent failures such as sensitive disclosure, hallucination, social incoherence, and identity confusion.

Development

Practicing with Language Models

Finds that LLM role-play can help humans practice and improve empathic communication.

Learning

Effective Personalized AI Tutors

Reports gains from an adaptive LLM-guided reinforcement learning tutor in high school programming.

Governance

AllHere "Ed" AI Chatbot Review

Analyzes a failed district chatbot rollout involving vendor collapse, data privacy concerns, and investigations.

Workforce

Labor Market Impacts of AI

Introduces an index for tracking how AI automates job tasks and reshapes labor demand.

Development

Are Children Testing Less Intelligent?

Questions whether declining scores reflect real cognitive decline or measurement artifacts and cultural drift.

Assessment

Comparing AI and Expert Feedback

Compares LLM-generated feedback with expert feedback and highlights the value of human nuance.

Learning

Engaged Teaching: Engaged Learning

Evaluates teaching methods, student engagement, and education technology in Apple Distinguished Schools.

Learning

From Chalkboards to Chatbots

Studies a Nigeria trial using Microsoft Copilot as an AI tutor and reports significant learning gains.

Development

Learning in the Age of Algorithmic Video

Explores how short-form algorithmic video can affect attention, memory, and instructional design.

Learning

Let AI Read First

Introduces AI-based text formatting support that improves reading speed and comprehension for dyslexic readers.

Assessment

On the Limits and Opportunities of AI Reviewers

Compares AI and human peer review, finding AI strong in structure but weaker in deep critique.

Assessment

Frequent ChatGPT Users Detect AI Text

Finds that experienced ChatGPT writers can identify AI-generated text more accurately than many tools.

Assessment

StoryScope

Builds a pipeline to extract narrative features and distinguish human-written fiction from AI-generated stories.

Safety

Sycophantic AI and Dependence

Reviews how overly agreeable AI companions may reduce prosocial behavior and increase emotional dependence.

Workforce

The Broken Ladder

Examines how AI and remote work may reduce early-career hiring and weaken entry-level pathways.

Workforce

The Cybernetic Teammate

Studies how generative AI reshapes teamwork, productivity, and expertise in a field experiment.

Learning

The Impressive Effects of Tutoring

Meta-analyzes PreK-12 tutoring and confirms substantial positive effects across contexts.

Development

Teachers Address Shortening Attention Spans

Reports on classroom strategies such as breaks and meditation to respond to attention challenges.

Learning

The Effects of GenAI on Learning Performance

Synthesizes experimental evidence on how generative AI affects student academic performance.

Safety

Persuading LLMs to Comply

Shows that users can bypass safety guardrails through psychological persuasion techniques.

Learning

The Science of Learning

Summarizes cognitive science principles for memory, multimedia learning, motivation, and classroom practice.

Governance

AI Conference Notes

Captures professional learning frameworks, classroom tools, and strategies for AI-assisted cheating concerns.

Governance

AI Conference Reflections

Curates summit highlights on AI storytelling, English language learner support, and leadership.

Safety

AI Fact Checking in the Wild

Compares LLM and human fact-checking behavior, including source use, style, and annotation quality.

Workforce

AI Self-Preferencing in Hiring

Shows that AI resume screening can favor AI-generated resumes over human-written resumes.

Governance

Parent Perspectives on AI

Synthesizes nearly 6,000 parent responses on AI integration, safety, and district expectations.

Assessment

Trusting AI to Detect AI?

Evaluates 13 AI-detection tools and finds reliability, robustness, and false-positive concerns.

Governance

AI in Orange County School Curriculum

Outlines AI-related curriculum platforms, regulatory considerations, and district integration context.

Governance

Empowering Learners for the Age of AI

Proposes AI literacy competencies for students and educators, emphasizing ethics and critical thinking.

Governance

Adaptive Governance for AI in K-12

Guides state and district leaders to organize around durable AI governance competencies.

Exhibit 10

Subject Area Impact Analysis

AI does not affect every subject in the same way. The subject-level evidence shows a pattern: AI is strongest when it supports practice, feedback, simulation, accessibility, and teacher review; it is riskiest when it replaces reading, reasoning, originality, or human support.

Community and educator signal

March-May 2026 survey evidence shows strong interest in AI support, but even stronger concern about independent student use, critical thinking, privacy, and misuse.

Perceived benefits

Tutoring and explanations 49%
Personalized support 42%
Workforce preparation 38%

Primary concerns

Reduced critical thinking 81%
Over-reliance and screen time 77%
Misuse and plagiarism 71%
ELA and literacy

Strong for feedback; risky for authentic reading and voice.

AI writing feedback can improve revision quality, motivation, and affect. New essay-scoring studies add a second warning: high agreement with teachers can still mask surface-feature bias and subgroup instability.

Use
Tutor-style prompts, version history, peer reading, oral defenses, and separate content-versus-language feedback.
Watch
Summary substitution, loss of vocabulary growth, weakened student voice, undisclosed AI-polished prose, and over-rewarded syntactic polish.
Avoid
Unmonitored take-home writing, automatic summarizer extensions as substitutes for reading, and unreviewed AI scoring for high-stakes grades.
Mathematics and STEM

Best as a scaffold, not a shortcut calculator.

Adaptive scaffolding, interest mapping, recommendation walls, and virtual lab support can reduce extraneous cognitive load. Value-sensitive ITS design adds that students need visible agency, privacy choices, and understandable AI decisions.

Use
Socratic tutoring, step review, benchmark triangulation, learner confidence checks, and source-validity checks.
Watch
Hallucinated procedures, unverified homework, and skipped logical steps.
Avoid
Direct calculator-style use before conceptual mastery is visible.
Computer science

Highest upside when adaptation is reliable and answers are withheld.

Adaptive problem sequencing with a conversational tutor improved unassisted final exam performance, and the new programming-feedback evidence favors hybrid GenAI-adaptive support over GenAI-only recommendations.

Use
No-direct-answer tutoring, code edit analysis, knowledge-graph grounding, varied pedagogy, and oral algorithm defenses.
Watch
Hallucinated code feedback, repeated exercise recommendations, platform instability, review bottlenecks, and false plagiarism signals in simple code.
Avoid
Penalties based only on automated code detectors.
Humanities and history

Useful for simulation; dangerous as a source of truth.

AI can create civic simulations, candidate profiles, structured debate, and historical inquiry prompts. The new law and scholarship papers add that judgment-rich support must still be anchored in expert standards and source evidence.

Use
Structured civic dialogue, localized surveys, primary source comparison, expert-modeled reasoning, and higher-order synthesis tasks.
Watch
Historical hallucinations, Eurocentric framing, ideological shortcuts, and weakened source analysis.
Avoid
Machine-generated bullet lists replacing primary source reading.
Fine arts and design

Powerful for ideation; risky for originality and attribution.

Text-to-image and art learning systems can support rapid prototyping, style exploration, motivation, and painting performance.

Use
Iterative sketch-to-AI feedback loops and explicit citation of prompts, sources, and style influence.
Watch
Copyright, style appropriation, abstract-language failures, and reduced technical skill growth.
Avoid
Direct submission of unrefined prompt outputs as final student artwork.
Clinical and support services

Use prediction only when a human response system is ready.

Predictive analytics can identify learning risks and support intervention planning, but the risk profile is the highest in the subject analysis. The policy-as-code evidence shows fairness and calibration, not operational cost, are the binding constraints.

Use
Human-in-the-loop early warning systems with defined intervention pathways, audit logs, fairness gates, and calibration checks.
Watch
Historical bias, sensitive demographic inputs, intersectional fairness gaps, emotional dependence, and sycophantic companion behavior.
Avoid
Clinical predictions without active supports or unchecked student interaction with agreeable AI companions.