The useful boundary: evidence versus judgment
A performance review combines several things that should not be blended carelessly: what happened, what impact it had, how the employee experienced the work, how the behavior compares with the role expectation, and what the organization decides to do next. AI can help sort those layers. It should not silently collapse them into a score.
“The project shipped on March 18 after the agreed date” is an observation. “The delay created two weeks of downstream rework” is an impact claim that needs evidence. “The employee does not care about deadlines” is an interpretation. “Meets expectations” is a judgment under the company’s criteria. “No promotion” is an employment decision. A responsible workflow keeps those statements visible as different kinds of information.
I ask AI to label the difference. That makes the draft less dramatic but more useful. If a manager feels strongly about someone’s performance but cannot supply examples, the correct output is an evidence gap and a conversation, not a more persuasive paragraph.
The review cycle I would use
| Stage | AI may help with | Humans must own |
|---|---|---|
| Prepare | Collect themes and identify missing examples. | What evidence is relevant and how it was obtained. |
| Draft | Behavior-based wording and clear structure. | Accuracy, context, fairness, and employee voice. |
| Calibrate | Evidence table and questions for discussion. | Standards, comparisons, ratings, and decision authority. |
| Discuss | Conversation questions and action summary. | Listening, explanation, disagreement, and commitments. |
| Record | Format a confirmed summary. | Approved system, access, retention, and final record. |
Start with a review packet, not a chat box
Before I ask AI for help, I create a review packet with only the material needed for the task. It may include the role expectations, agreed goals, project outcomes, dated examples, customer or stakeholder feedback that the process permits, the employee’s self-review, and notes about support or blockers. Each item should have a source and a date where possible.
I remove data that is not necessary. Medical details, protected characteristics, compensation information, private complaints, credentials, and disciplinary records should not be pasted into an unapproved tool. Even in an approved workspace, minimization is useful: the model does not need an employee’s full personal history to help organize a project outcome.
I also label perspective. An employee’s self-report is important, but it is still a report to discuss. A stakeholder comment is evidence to examine, not automatically truth. A manager note written months ago may contain interpretation. Ask AI to preserve those distinctions instead of blending everything into one narrative.
The review packet should stay connected to the organization’s approved HR system and access controls. A generated draft in a casual chat is not automatically the official personnel record.
Self-reviews: let employees show the work
A self-review can be a better starting point than a manager’s memory because the employee knows the work that was invisible from the outside. ChatGPT can help an employee organize accomplishments into outcomes, collaboration, challenges, lessons, and next-period goals. The employee should verify every claim and edit the voice.
I would ask the tool to preserve uncertainty. If an employee says a launch increased retention, the draft should leave a placeholder for the source and period rather than silently upgrade the statement into a fact. If a contribution was shared across a team, the language should not claim sole credit.
Managers should read a self-review for questions, not just for evidence that supports a predetermined view. What work is the manager not seeing? Which blocker needs action? Where does the employee want development? What goal was unclear or changed mid-cycle? AI can make those questions easier to find, but the manager must ask them.
Manager reviews: write about behavior and impact
Useful feedback is specific enough that the employee can act on it. I use a simple structure: situation, behavior, impact, and next step. The situation identifies the work. The behavior says what the person did or did not do. The impact explains why it mattered. The next step says what to repeat, change, or discuss.
Weak feedback says, “Communication needs improvement.” A stronger version says, “During the March launch, two dependency changes were not included in the weekly status update, so the design review moved by two days. For the next launch, send a written risk update by Wednesday and flag changes in the project channel the same day.” The second version still needs the manager to verify the facts, but it gives the employee a fairer starting point.
AI is useful for finding vague verbs and labels. Ask it to highlight words such as “attitude,” “professionalism,” “ownership,” or “not collaborative” and require an example for each. Sometimes the label is accurate shorthand between experienced managers; it is not sufficient evidence for a formal review.
Check opportunity and resources too. Did the employee know the expectation? Have access to the system? Receive the necessary training? Carry a reasonable workload? Raise a blocker? The impact may be real even when the cause is not simple. The review should not let a model choose the explanation.
Do not let AI become a hidden employment decision-maker
The EEOC has warned that AI and other software used in employment decisions can create disability discrimination risks, including through screening and assessment. The concern is broader than a chatbot writing a paragraph: a tool may influence evaluation, ranking, monitoring, promotion, pay, or termination.
For performance reviews, I would prohibit asking AI to infer personality, motivation, emotional state, leadership potential, honesty, or culture fit from writing, speech, attendance patterns, or communications. I would also prohibit using it to rank employees or choose a rating without a documented, human-controlled process.
When a review involves an accommodation, medical information, protected leave, harassment, discrimination, retaliation, discipline, termination, or a complaint, follow the organization’s HR process and involve the qualified owner. AI can help prepare questions or organize approved facts, but it cannot replace that process.
Calibration: organize disagreement without hiding it
Calibration works best when managers compare evidence against shared criteria, not when they compare personalities. AI can prepare a neutral evidence sheet: goal, outcome, scope, source, context, evidence gap, and question for the panel. It can also flag when one manager’s language is much more certain than another’s.
I would never ask it to produce a leaderboard. A ranking can hide the criteria, reproduce historical bias, and make an uncertain input look precise. Instead, ask what evidence supports each conclusion, whether the same standard is being applied, and what information would change the view.
Calibration participants should be able to challenge the evidence and add context. If the employee faced a changed goal, a staffing gap, an approved accommodation, or a dependency outside their control, that context may matter. The record should distinguish what is known from what the panel is deciding.
Use AI after the meeting to organize confirmed actions and unresolved questions, not to manufacture consensus. The final rating and any related decision follow the company’s approved process.
Six prompts for evidence-first reviews
1. Organize evidence by goal
Using only these documented work examples and goals, group the evidence by goal or competency. For each example, state the action, result, scope, and source. Separate confirmed evidence, employee-reported context, manager interpretation, and missing evidence. Do not assign a rating or infer personality, intent, health, or protected traits.
2. Find evidence gaps
Review these performance notes for unsupported conclusions. List each judgment, the evidence that would support it, the missing context, and one neutral question the manager could ask. Check whether the employee had the required resources, authority, time, and clear expectations. Do not write final feedback yet.
3. Draft behavior-based feedback
Draft feedback from the confirmed examples using situation, behavior, impact, and next step. Use specific, respectful language. Do not exaggerate, diagnose, infer motives, mention protected traits, invent a consequence, or turn one event into a pattern. Mark any sentence that requires manager or HR review.
4. Prepare self-review themes
Organize this employee’s self-review into accomplishments, measurable outcomes, collaboration, challenges, lessons, and next-period goals. Preserve the employee’s claims as claims until verified. Flag missing metrics and suggest questions, but do not inflate the contribution or write in a voice the employee would not use.
5. Prepare calibration questions
Create a calibration evidence sheet from these review notes. Show the standard being applied, confirmed outcomes, missing evidence, comparable examples, context that may affect interpretation, and questions for the human panel. Do not rank people, recommend a rating, or recommend pay or promotion.
6. Audit the final draft
Audit this performance review draft for factual support, behavior-based wording, consistency with the stated criteria, unnecessary personal information, protected-trait assumptions, unequal standards, accidental legal conclusions, unsupported rating language, and tone. Return MUST FIX, HR REVIEW, and READY TO EDIT. Do not rewrite until the issues are listed.
The final review pass before the meeting
Before a review draft reaches an employee, I run five checks. First, evidence: can I point to the source for each important claim? Second, consistency: did I apply the same criteria used for comparable work? Third, context: did I check resources, expectations, changes, and the employee’s perspective? Fourth, privacy: does the document contain only what the audience needs? Fifth, consequence: does the language imply a rating, promise, or employment action that has not been approved?
| Red flag | What I do next |
|---|---|
| A conclusion without an example | Find the evidence or rewrite it as an open question. |
| A label about personality or intent | Replace it with observable behavior and impact. |
| A single dramatic incident | Check pattern, context, response, and policy before formal use. |
| Sensitive personal information | Minimize, remove, or use the approved HR process. |
| A model-suggested rating | Discard it and return to the organization’s human criteria. |
Privacy and account controls
The tool and account matter. OpenAI documents different controls for business offerings and personal workspaces, including business data not being used for model training by default. That is relevant to a manager choosing a workflow, but it is not a blanket authorization to process employee records. The organization still needs a purpose, access boundary, retention rule, approved account, and accountable reviewer.
I ask managers to use data minimization before the prompt: remove names where they are not needed, replace exact identifiers with roles, keep only the evidence relevant to the question, and avoid uploading an entire personnel file to solve a small drafting task. I also make sure the final review record is stored in the approved HR system rather than in a personal chat.
For self-reviews, explain what AI assistance is permitted and what must remain the employee’s own work. For manager reviews, explain who can access drafts, how corrections are handled, and how employees can raise a concern. Transparency does not solve every risk, but secrecy makes trust harder.
A review-cycle pilot
Week one: select a low-risk test set from completed work and remove restricted information. Write the review criteria and evidence standard before asking AI for language. Test organization and gap-finding, not ratings.
Week two: have two managers use the same prompt on comparable examples, then compare the edits. Look for different standards, missing context, labels, and unsupported confidence. Ask HR to review the workflow rather than only the final sentence.
Week three: use the tool for a self-review or manager draft with consent and approved controls. Keep the employee’s source material, the draft, and the human edits distinguishable. Measure correction time and questions raised.
Week four: decide which use cases are approved, restricted, or prohibited. Document the reviewer, storage location, retention, and escalation. Do not expand because the first draft sounded good; expand only when the review standard works.
My recommendation
Use AI to make evidence easier to see and feedback easier to understand. Keep the employee’s work, context, and voice visible. Ask for gaps and questions before asking for polished prose.
Do not use an AI system as the final judge of a person. Ratings, rankings, compensation, promotion, discipline, termination, and other employment decisions require the organization’s accountable human process. When the situation is sensitive, the correct AI workflow may be a redacted preparation step or no AI at all.
A fair review is not the one with the most elegant language. It is the one an employee can understand, challenge, and act on because the evidence, standard, context, and next step are clear.