Customer research
Use AI to analyze customer feedback without inventing demand
Ask AI to label individual comments before summarizing themes. Count unique customers, separate requests from underlying problems, preserve conflicting feedback, and describe small samples as signals to investigate rather than proof of market demand.
In this article
A convincing summary can exaggerate a small signal
A founder pastes ten customer messages into an AI assistant and asks what to build next. The response identifies strong demand for integrations, better reporting, and a simpler onboarding flow. It sounds like a product strategy. Yet perhaps three messages came from the same customer, two requests were hypothetical, and the reporting complaint described a bug rather than a missing feature.
I use AI to make feedback easier to inspect before I ask it to summarize anything. The first task is to preserve who said what and what evidence the comment actually contains. A small, labeled dataset can support useful follow-up questions. It cannot tell you the priorities of a whole market simply because the model groups the comments into neat themes.
This article uses a fictional appointment-booking app. The examples are invented to demonstrate the method. They do not represent a customer study or measured demand. The process works with support tickets, interview notes, cancellation comments, and product reviews, provided you understand how the sample was collected.
The outcome I want is a short evidence map: repeated problems, differences between users, questions that need another conversation, and one sensible experiment. That is enough to guide the next week of research without pretending that ten messages have settled the roadmap.
Record how the feedback reached you
Before classification, label the source of each comment. A cancellation survey, a public review, a sales call, and an unsolicited support ticket come from different situations. People who contact support are not necessarily representative of satisfied customers. Prospects describing an imagined workflow are not reporting the same experience as active users.
Create a record ID and a pseudonymous customer ID. The record ID lets you cite the exact comment. The customer ID helps prevent repeated messages from becoming multiple independent votes. You do not need personal names in the analysis when a consistent identifier will do.
Add the date, product version if known, and customer context that changes interpretation. For the booking app, a solo tutor and a multi-location studio may need different reporting. Preserve that distinction if it is actually stated. Do not let the model infer company size or profession from writing style.
Keep the original text in a separate column and avoid rewriting it before classification. A sentence that sounds contradictory may contain the most important insight. The user says scheduling is easy but managing changes is hard; simplifying that into onboarding problem loses the point. Clean formatting while preserving meaning, uncertainty, and the user's actual example.
Separate the request from the problem underneath it
Imagine four fictional comments. A tutor wants a weekly email showing canceled lessons. A studio owner asks for an Excel export to reconcile attendance. Another tutor says the dashboard total is wrong after rescheduling. A prospect asks whether reports can use the company's colors. These all mention reporting, but they do not describe one interchangeable need.
The first is about staying informed. The second is about moving data into another process. The third may be a correctness problem. The fourth is a presentation preference from someone who has not used the product. Grouping all four as demand for an advanced reporting dashboard would erase the differences that should guide your response.
Ask the assistant to label the expressed request and the underlying problem separately. The underlying problem should stay close to the evidence. It can be a cautious interpretation, but it should not expand into a psychological theory about the customer. If the comment does not explain why the request matters, mark that question as unanswered.
This distinction often changes the cheapest next step. A weekly notification might solve one problem without building a dashboard. Fixing a calculation could solve another without adding a feature. The model can suggest these possibilities, but a conversation with the customer is still needed to confirm that the proposed solution addresses the actual situation.
R01 / C01: I miss canceled lessons unless I check every day. Could you email a weekly list? R02 / C02: I need attendance in Excel for our monthly reconciliation. R03 / C03: After rescheduling, the total still shows the original booking. R04 / C04: We are considering the app. Can reports use our colors?
Label one comment at a time with a small codebook
A codebook is simply a list of labels and what they mean. Start with a few categories relevant to the product: scheduling, notifications, reporting, billing, reliability, and other. Define them in one sentence each. Add a separate field for whether the comment is a bug report, feature request, question, or positive observation.
Do not force every comment into one theme if it clearly addresses two issues. Allow a primary and secondary label, with a short explanation. At the same time, avoid an unlimited number of invented categories. If each comment gets its own label, grouping provides little value. The codebook should help you compare records without flattening them.
Ask the assistant to show the supporting phrase for each label. That makes review faster: you can see whether notifications was assigned because the customer requested an email or because the model guessed that reminders would help. The explanation should be brief enough to inspect alongside the original comment.
Review the first small batch manually. Correct inconsistent labels and revise definitions before processing more records. If rescheduling errors appear under both scheduling and reliability, decide how you want to handle overlap. Keep the decision in the codebook so the next batch follows the same logic.
Label each feedback record using this codebook: [LABELS AND DEFINITIONS]. Return record ID, customer ID, primary theme, optional secondary theme, comment type, stated request, underlying problem if supported, and supporting phrase. Use 'unclear' when the comment does not establish the problem. Do not infer customer attributes. Do not recommend roadmap priorities yet. [RECORDS]
Count customers and comments separately
A vocal customer can send five tickets about the same problem. Those tickets matter because they may show repeated disruption, but they are not five independent customers requesting a feature. Report both record count and unique customer count when the distinction is relevant.
For a small sample, plain counts are often more honest than percentages. Three of ten comments mentioning exports sounds appropriately limited. Thirty percent of customers demand exports sounds like a market estimate, especially if the sample was selected from a support inbox. The wording should reflect the collection method and the denominator.
If a customer mentions two themes, the theme counts may add up to more than the number of customers. Explain that rather than forcing a misleading total. Categories are not always mutually exclusive. A useful table can show theme, comment count, unique customers, source types, and a short evidence note.
Watch for duplicate content copied across systems. One interview may appear in a transcript, a research note, and a support ticket. If the same observation enters the dataset three times, label the common origin. AI can flag similar wording, but you should confirm duplicates before removing records that may represent separate experiences.
Keep the disagreement that a summary wants to hide
Suppose two studio managers want a detailed dashboard while three solo tutors want fewer screens. The summary should not average those preferences into users want a moderately detailed dashboard. The disagreement suggests different jobs, contexts, or levels of complexity. Preserving it is more useful than finding a smooth middle sentence.
Ask for negative cases: records that contradict or limit each proposed theme. If the model says onboarding is confusing, look for comments from users who completed setup easily and note what differed. That does not disprove the problem. It helps identify the conditions under which the problem occurs.
Also distinguish absence from opposition. A customer who did not mention exports has not necessarily rejected exports. Feedback usually covers whatever was salient in the moment. Treating silence as a vote against a feature can be as misleading as treating one request as universal demand.
A good synthesis might say that export requests came from organizations reconciling attendance outside the app, while solo tutors mainly wanted a digest. That sentence offers a segmentation hypothesis grounded in the comments. The next step is to test it with more users, not announce that the market has two definitive segments.
Turn each theme into a better follow-up question
The most useful output from a small feedback analysis is often a question. For the export request, ask the customer to walk through their last reconciliation. What information did they move, where did it go, and what happened when a record was missing? The answer may reveal that a simple attendance file is enough.
Avoid asking whether the customer would like your proposed feature. Many people will agree to an appealing hypothetical. Ask about a recent event instead: when did the problem last occur, what did you do, how long did it take, and what consequence followed? You are looking for the actual job and workaround.
AI can draft an interview guide from the theme table. Tell it to produce neutral questions and label the assumption each question tests. Review for leading wording. Would automatic reports save you hours assumes both a solution and a benefit; how did you prepare the last report invites evidence.
Keep questions short enough for a natural conversation. A five-part research question may look thorough on paper but be difficult to answer aloud. Ask one thing, listen, and follow the detail. The guide should help you notice important gaps without turning the interview into a checklist recital.
For each theme, propose two neutral follow-up questions about a recent real event. State the assumption each question tests. Do not ask whether the customer likes a proposed feature or suggest a time saving. Prioritize the current workaround, frequency, consequence, and information needed to finish the task.
Choose a small experiment before a large build
For the booking app, a reasonable first experiment might be a manually prepared attendance export for a few customers who already reconcile data elsewhere. Define what you want to learn: which columns they use, whether the file fits the workflow, and what still requires manual repair. The experiment is about usefulness, not proving the original idea right.
Set a concrete observation. Can the customer complete their usual reconciliation with the file? Which fields do they edit? Do they request it again? These questions are more informative than asking whether the export looks good. A positive reaction is encouraging, but repeated use in the intended task is stronger evidence of fit.
Keep the experiment separate from a public promise. Explain what is being tried and what is not yet a permanent product feature. If you are testing a manual service behind a proposed interface, describe the arrangement honestly to the participants. A learning exercise should not create commitments your team cannot sustain.
Write a stopping condition too. If customers need different data definitions or the source is unreliable, pause the interface work and fix the information problem first. AI can help compare experiment notes with the original hypothesis, but the decision belongs to the team responsible for the product and customer relationship.
Write a research note that a colleague can challenge
A useful research note starts with what you reviewed: ten support messages from six customers during a specific period, for example. Then state the main signals, show representative record IDs, preserve the contradictory cases, and name the questions that remain. The sample description belongs near the conclusion, not hidden in a footnote.
Use wording that matches the evidence. In this sample, several studio users needed attendance data outside the app is better than customers demand integrations. The first sentence identifies a context and a behavior. The second skips from a few messages to a broad product direction.
Include a short section on what the analysis cannot establish. You may not know how common the problem is among all active users, whether people would pay for a solution, or whether a different workflow would solve it. These are useful boundaries because they tell the team what to investigate next.
Google's prompting guidance supports providing clear instructions and examples. For this workflow, the examples belong in the codebook and the review process. They help the model label comments consistently; they do not turn a small dataset into a representative survey. Keep that distinction visible when sharing the AI-assisted synthesis.
Keep a living evidence map instead of a permanent verdict
Save the codebook, labeled records, and research note together. When new feedback arrives, apply the same definitions first, then revise the codebook if the product has changed. Record revisions so a change in theme counts is not mistaken for a change in customer behavior when it actually came from different labels.
Revisit earlier conclusions when new evidence conflicts with them. If later interviews show that the export request was a workaround for an incorrect dashboard, the priority may shift toward reliability. That is a successful research update, not a failure of the first analysis. The initial note should have been provisional enough to accommodate learning.
Keep the original comments available to the people making the decision. A summary is convenient, but direct evidence helps a team challenge interpretation. The assistant's role is to make that evidence easier to navigate. It should not become the only voice through which the customer is heard.
For your next batch, start with a handful of records and ask for labels with supporting phrases. Correct those labels, count unique customers, and write one follow-up question per theme. That small routine can improve product discussions immediately, even before you have enough feedback for stronger conclusions.
The practical payoff is a roadmap conversation with fewer invented certainties. You can explain what customers said, what you think it might mean, and what you will do to find out. Those are three different things, and a useful AI workflow keeps them distinct.
Ask whose experience is absent from the file
A support inbox describes people who contacted support. It may tell you little about people who left silently, never completed signup, or use the product without difficulty. Before treating a theme as a roadmap priority, write down which groups produced the records and which groups are missing. This is a sampling limitation, not something the chatbot can repair by generating more confident prose.
For the fictional booking app, complaints from established customers might highlight calendar management while omitting the confusion that stops new customers during setup. Both issues can matter, but the inbox alone cannot compare their prevalence. A small set of onboarding conversations or an appropriately collected funnel report would answer a different question.
I would add a coverage note beside the findings: These records came from recent support contacts; they do not represent all customers or prospective buyers. Then specify the next evidence needed before making a larger decision. This keeps the useful observations intact while preventing a convenient dataset from becoming an unsupported claim about the entire market.