For schools, colleges and learning products
The education AI agent I would pilot first
Start with a low-stakes, measurable workflow that helps staff or students without giving a model authority over grades, discipline, admissions or safeguarding.
Education teams often begin with the word agent and then search for a task. I reverse that order. Find a repetitive queue with clear source material, a named owner and a safe escalation path. A student-services FAQ grounded in approved policies is easier to evaluate than an autonomous tutor expected to know every learner and subject.
The design must respect age, privacy, accessibility and institutional policy from the beginning. FERPA, COPPA and local requirements are not badges a vendor can solve in a feature list. The institution needs to decide what data enters the system, who can see it, how errors are corrected and where a human remains responsible.
Choose assistance before authority
Let the agent retrieve a policy, draft a routine response or prepare a lesson option. Do not let it make consequential decisions about students or silently change records.
Ground answers in approved material
A school calendar, handbook, course catalogue and support directory should be versioned sources. The agent should cite the source and admit when the answer is not present.
Design the handoff first
Define urgent, sensitive and ambiguous cases before launch. Students must know when they are speaking with AI and how to reach a person.
Minimise student data
Do not collect a full record to answer a general question. Separate content generation from identifiable learner data wherever possible and document retention.
Keep teachers in pedagogical control
An agent can create variations, but the teacher sets objectives, selects material, checks accuracy and adapts for the class.
Evaluate different learners
Test reading levels, disability access, multilingual use, edge cases and unequal error rates. Average satisfaction can hide a poor experience for the students most dependent on support.
A responsible eight-week pilot
- 1
Define one queue
Do: Choose a bounded workload and record its current volume, handling time and error types.
Example: Answering registrar questions from published policy is clearer than 'personalise education.'
Checkpoint: The task has an owner and baseline.
- 2
Classify risk
Do: List data used, people affected and consequences of a wrong answer.
Example: A wrong cafeteria time differs from incorrect financial-aid guidance.
Checkpoint: High-risk topics route to staff.
- 3
Build the source set
Do: Approve, date and version every document the agent may use.
Example: Archive the superseded attendance policy instead of leaving both searchable.
Checkpoint: Answers can cite current sources.
- 4
Write refusal and escalation rules
Do: Specify what the agent will not decide and how a person takes over.
Example: Safeguarding language triggers the institution's existing protocol, not a generated counselling response.
Checkpoint: Staff rehearse the handoff.
- 5
Test with real scenarios
Do: Use anonymised common, ambiguous, adversarial and sensitive questions.
Example: Include a student who omits key context or uses informal language.
Checkpoint: Failures are categorised, not averaged away.
- 6
Launch narrowly
Do: Release to a small group with visible feedback and manual review.
Example: Begin with one programme or staff team before campus-wide access.
Checkpoint: There is a stop condition and rollback.
Education agent proposals I would reject
Automated grading as the first pilot
The stakes, validity and appeal requirements are too high for an immature deployment.
Uploading everything
A larger knowledge base can contain conflicts, restricted information and obsolete policy.
A generic tutor for every subject
Subject accuracy, pedagogy and learner needs require narrower design and expert review.
No disclosed handoff
A student should not have to discover by frustration that no person is monitoring the interaction.
Vendor claims as evaluation
Test the institution's workflows and population. A benchmark is not local evidence.
Time saved as the only metric
Measure correctness, resolution, equity, escalation quality and staff burden created by review.
Planning prompts for an education team
Use-case risk screen
For this proposed education workflow, map users, data, decisions, harms, required sources, human owner, escalation route and evidence needed before launch. Do not assume the agent may make consequential decisions.
Use in a cross-functional workshop.
Evaluation set builder
Create test-case categories from the approved policy and anonymised ticket patterns I provide: routine, ambiguous, conflicting-source, out-of-scope, accessibility, multilingual and urgent. Do not create or infer student records.
A test set should predate launch.
Citation-first answer rule
Answer only from the supplied approved sources. Cite the document and section. If sources conflict or do not answer the question, say so and route to [TEAM]. Never infer an individual student's eligibility or status.
A useful system instruction for policy Q&A.
Pilot review
Compare baseline and pilot results for accuracy, resolution, handling time, escalation, user satisfaction and subgroup differences. Identify where apparent time savings moved work to reviewers.
Use real measured data, not estimates.
Reader workbook
Working notes for the education ai agent i would pilot first
Reading the guide is only the first pass. The notes below turn its recommendations into evidence you can inspect, discuss with another person, and revise. Complete them with real material from your situation. Do not let an AI assistant fill gaps with plausible facts. When a policy, price, specification, source, system state, or personal experience matters, open the authoritative record and put the verified detail in your working document.
Working note 1: Choose assistance before authority
Begin with a concrete example from the last thirty days. Record what happened, what information was available at the time, who made the decision, and what the result was. Then apply the principle above to that example. The useful output is not a general agreement that the principle sounds sensible; it is one changed action, one piece of evidence you will collect, and one condition that would make you choose a different approach. Write those three items in language another person could audit.
Evidence to leave behind
A dated note that connects βChoose assistance before authorityβ to one actual decision, names the evidence used, records uncertainty, and identifies the next person or check required before the decision becomes final.
Working note 2: Ground answers in approved material
Test this principle against a difficult case rather than the easiest one. List the constraint most likely to be ignored, the person who carries the downside if the advice is wrong, and the source that can settle a factual disagreement. Next, describe a small trial that is reversible and produces a visible result. Decide in advance what would count as improvement, no change, or harm. This turns a broad recommendation into a decision with limits instead of another optimistic intention.
Evidence to leave behind
A dated note that connects βGround answers in approved materialβ to one actual decision, names the evidence used, records uncertainty, and identifies the next person or check required before the decision becomes final.
Working note 3: Design the handoff first
Explain this idea to a colleague, teacher, adviser, reviewer, or teammate without using jargon. Ask them where the explanation assumes knowledge that has not been demonstrated. Add the missing source, example, calculation, test, or observation. Finally, write the strongest reasonable objection and a response that acknowledges the trade-off. If the response depends on a vendor claim or an AI answer, mark it unverified until it has been checked against a primary source or real result.
Evidence to leave behind
A dated note that connects βDesign the handoff firstβ to one actual decision, names the evidence used, records uncertainty, and identifies the next person or check required before the decision becomes final.
Working note 4: Minimise student data
Begin with a concrete example from the last thirty days. Record what happened, what information was available at the time, who made the decision, and what the result was. Then apply the principle above to that example. The useful output is not a general agreement that the principle sounds sensible; it is one changed action, one piece of evidence you will collect, and one condition that would make you choose a different approach. Write those three items in language another person could audit.
Evidence to leave behind
A dated note that connects βMinimise student dataβ to one actual decision, names the evidence used, records uncertainty, and identifies the next person or check required before the decision becomes final.
Working note 5: Keep teachers in pedagogical control
Test this principle against a difficult case rather than the easiest one. List the constraint most likely to be ignored, the person who carries the downside if the advice is wrong, and the source that can settle a factual disagreement. Next, describe a small trial that is reversible and produces a visible result. Decide in advance what would count as improvement, no change, or harm. This turns a broad recommendation into a decision with limits instead of another optimistic intention.
Evidence to leave behind
A dated note that connects βKeep teachers in pedagogical controlβ to one actual decision, names the evidence used, records uncertainty, and identifies the next person or check required before the decision becomes final.
Working note 6: Evaluate different learners
Explain this idea to a colleague, teacher, adviser, reviewer, or teammate without using jargon. Ask them where the explanation assumes knowledge that has not been demonstrated. Add the missing source, example, calculation, test, or observation. Finally, write the strongest reasonable objection and a response that acknowledges the trade-off. If the response depends on a vendor claim or an AI answer, mark it unverified until it has been checked against a primary source or real result.
Evidence to leave behind
A dated note that connects βEvaluate different learnersβ to one actual decision, names the evidence used, records uncertainty, and identifies the next person or check required before the decision becomes final.
A review record for a responsible eight-week pilot
Keep one row for every pass through the workflow. The record should make progress and failure equally easy to see. A polished output with no trace of its sources, assumptions, checks, or human decisions is difficult to improve and dangerous to trust. The following review questions are deliberately tied to the steps above.
After step 1
Review βDefine one queueβ
Record what you actually did, not what the plan said you would do. Attach the relevant output or source. Use this checkpoint as the acceptance test: The task has an owner and baseline. If it is not met, note whether the cause was missing information, weak skill, unclear ownership, insufficient time, a faulty assumption, or an external constraint. Choose one correction and repeat the smallest affected step rather than restarting the entire workflow.
After step 2
Review βClassify riskβ
Record what you actually did, not what the plan said you would do. Attach the relevant output or source. Use this checkpoint as the acceptance test: High-risk topics route to staff. If it is not met, note whether the cause was missing information, weak skill, unclear ownership, insufficient time, a faulty assumption, or an external constraint. Choose one correction and repeat the smallest affected step rather than restarting the entire workflow.
After step 3
Review βBuild the source setβ
Record what you actually did, not what the plan said you would do. Attach the relevant output or source. Use this checkpoint as the acceptance test: Answers can cite current sources. If it is not met, note whether the cause was missing information, weak skill, unclear ownership, insufficient time, a faulty assumption, or an external constraint. Choose one correction and repeat the smallest affected step rather than restarting the entire workflow.
After step 4
Review βWrite refusal and escalation rulesβ
Record what you actually did, not what the plan said you would do. Attach the relevant output or source. Use this checkpoint as the acceptance test: Staff rehearse the handoff. If it is not met, note whether the cause was missing information, weak skill, unclear ownership, insufficient time, a faulty assumption, or an external constraint. Choose one correction and repeat the smallest affected step rather than restarting the entire workflow.
After step 5
Review βTest with real scenariosβ
Record what you actually did, not what the plan said you would do. Attach the relevant output or source. Use this checkpoint as the acceptance test: Failures are categorised, not averaged away. If it is not met, note whether the cause was missing information, weak skill, unclear ownership, insufficient time, a faulty assumption, or an external constraint. Choose one correction and repeat the smallest affected step rather than restarting the entire workflow.
After step 6
Review βLaunch narrowlyβ
Record what you actually did, not what the plan said you would do. Attach the relevant output or source. Use this checkpoint as the acceptance test: There is a stop condition and rollback. If it is not met, note whether the cause was missing information, weak skill, unclear ownership, insufficient time, a faulty assumption, or an external constraint. Choose one correction and repeat the smallest affected step rather than restarting the entire workflow.
A red-team pass before you rely on the result
Use the failure modes from this guide as a final challenge, not as a warning box you read and forget. Assign each one to a reviewer, or take them one at a time yourself. The reviewer should point to evidence in the work and should be allowed to say that the evidence is insufficient.
Could βAutomated grading as the first pilotβ be happening here?
Find the strongest sign that it is, then the strongest sign that it is not. Do not accept confidence, fluent wording, a high score, or a successful first attempt as proof. Write the additional check that would change the decision and name who owns that check.
Could βUploading everythingβ be happening here?
Find the strongest sign that it is, then the strongest sign that it is not. Do not accept confidence, fluent wording, a high score, or a successful first attempt as proof. Write the additional check that would change the decision and name who owns that check.
Could βA generic tutor for every subjectβ be happening here?
Find the strongest sign that it is, then the strongest sign that it is not. Do not accept confidence, fluent wording, a high score, or a successful first attempt as proof. Write the additional check that would change the decision and name who owns that check.
Could βNo disclosed handoffβ be happening here?
Find the strongest sign that it is, then the strongest sign that it is not. Do not accept confidence, fluent wording, a high score, or a successful first attempt as proof. Write the additional check that would change the decision and name who owns that check.
Could βVendor claims as evaluationβ be happening here?
Find the strongest sign that it is, then the strongest sign that it is not. Do not accept confidence, fluent wording, a high score, or a successful first attempt as proof. Write the additional check that would change the decision and name who owns that check.
Could βTime saved as the only metricβ be happening here?
Find the strongest sign that it is, then the strongest sign that it is not. Do not accept confidence, fluent wording, a high score, or a successful first attempt as proof. Write the additional check that would change the decision and name who owns that check.
How to use the prompts without outsourcing judgment
Before running any prompt, replace every placeholder, remove private information that is not required, and state which supplied sources the assistant may use. Save the initial input, output, corrections, and final human decision. This creates a record of how the tool contributed and makes it easier to spot when a later answer contradicts an earlier assumption.
- Use-case risk screen: define the expected output before sending it, verify every material claim afterwards, and use this practical boundary: Use in a cross-functional workshop.
- Evaluation set builder: define the expected output before sending it, verify every material claim afterwards, and use this practical boundary: A test set should predate launch.
- Citation-first answer rule: define the expected output before sending it, verify every material claim afterwards, and use this practical boundary: A useful system instruction for policy Q&A.
- Pilot review: define the expected output before sending it, verify every material claim afterwards, and use this practical boundary: Use real measured data, not estimates.
End by writing a short decision note in your own words: what you learned, what remains uncertain, which source or test carries the most weight, what you decided, and when the decision should be reviewed. That note is often more valuable than the original AI output because it captures accountable judgment rather than a temporary answer.
What success looks like in education
A successful agent makes a narrow service more consistent and accessible while leaving responsibility visible. Staff know its limits, students can reach a person, and every answer can be traced to an approved source.
Only after that evidence should the institution expand to a second workflow. Scale is a result of trust earned in operation, not the starting objective.