Back to blog

Should we entrust AI with analysing user feedback? (artificial intelligence)

We compared the results of four AI models and three human analysts while labelling 100 user feedback entries. (artificial intelligence)

2026-03-10
6 min read

We compared the results of four language models with human analyses

Users constantly talk about services, apps and products: in Google reviews, app stores, forums, on social media and even through customer service channels. These are not responses to questionnaires, but voluntary, spontaneous opinions. That makes them highly valuable data, but also noisy. To use them for product or service development decisions, it is worth structuring this organic feedback in some way.

The challenge, then, is not a lack of data, but interpreting it.

What are people actually talking about? What themes, complaints and development needs emerge from the texts? The clearest answers come from identifying patterns in the feedback, then labelling and categorising the texts accordingly. Small volumes of data can be processed effectively by people, but when volumes are large and analysis is ongoing, an appealing question arises: can we entrust the work to a language model? And if so, which one, how should we use it, and what pitfalls should we expect? We explored these questions through user reviews of a banking mobile app.

Method

We asked four language models (GPT/ChatGPT, Sonnet/Claude, RoBERTa and LLaMA) and three people to label 100 user feedback entries according to predefined categories. We had the language models perform the same task three times. This allowed us to examine not only how closely AI solutions resembled human reasoning, but also the models’ consistency (in other words, whether individual models could reproduce their own results). (artificial intelligence)

ChatGPT and Claude performed the labelling task based on a prompt; we used the RoBERTa and LLaMA models via the NLP Cloud (natural language processing) categorisation platform, while the human participants received written instructions.

NLP Cloud (natural language processing)

NLP Cloud is an online platform that makes a range of artificial intelligence models available both through its website and via an API. It offers tools for several types of task, including a module for text categorisation. Through this, we controlled the RoBERTa and LLaMA models (unlike Claude and ChatGPT) not with textual instructions, but programmatically via an API. (natural language processing) (application programming interface)

We iterated once on the prompts for ChatGPT and Claude, so these models carried out the labelling in two rounds, 2*3 times in total (see Table 1).

1. táblázat: Használt modellek és instrukciók típusa
Table 1: Models used and types of instruction

Lesson 1: the importance of how instructions are phrased

In the first round, ChatGPT achieved 100% consistency, meaning it produced exactly the same result in all three runs. At first glance, this seemed like an excellent result, but it was actually so stable for the wrong reason: the model fell into anchoring bias, assigning virtually everything to the same single category, regardless of the content of the texts.

With the same prompt, Claude performed as we expected. It used several categories, and its results matched across the three runs 83% of the time.

Anchoring bias

The phenomenon in which the first piece of information received (the “anchor”) disproportionately influences subsequent judgements and decisions, even when the anchor is partly or entirely irrelevant. (For further examples and research on anchoring bias in AI, see this link.) (artificial intelligence)

We then refined the prompt and ran it three times with both models. In this second round, ChatGPT became anchored in only one run, while the other two produced more varied labelling. As a result, however, the three runs differed substantially, leaving ChatGPT with 10% consistency in this round. This is unacceptable for structured, repeatable analytical tasks.

For Claude, the refined prompt reduced stability: consistency fell from 83% to 63% (see Table 2). This also shows that different models may respond differently to the same prompt and the same modification.

We did not run LLaMA and RoBERTa using textual prompts, but through a solution designed specifically for text categorisation, using Python code. Here, LLaMA showed 95% consistency and RoBERTa 100% consistency.

2.táblázat: Modellek konzisztenciája három futtatás alapján
Table 2: Model consistency based on three runs

The consistency of results is greatly influenced by how we run the task and the constraints we set around it. Chat-based generative models (ChatGPT, Claude) are more sensitive to small differences in wording, so the same task can produce different labels across multiple runs. Although LLaMA is also a generative model, like the models behind ChatGPT and Claude, we ran it in a structured way (via a categorisation API, using code and a fixed set of labels), which gave us much more stable results. RoBERTa works on a different principle from the other three models and is specifically optimised for classification. As such, 100% consistency is to be expected from this model. (application programming interface)

Lesson 2: consistency ≠ quality

For ChatGPT, perfect consistency concealed false results. By contrast, RoBERTa was 100% stable while using a range of labels, as expected. However, when we compared RoBERTa’s solution with the results of human labelling (using the consensus labels of the three participants), the two matched only 25% of the time.

So RoBERTa was stable, but its solution differed substantially from human reasoning.

Lesson 3: people do not always agree with one another either

Before expecting AI models to provide consistent solutions that also make sense to us, it is worth looking at how people perform on the same labelling task. Three assessors worked independently on the same 100 user feedback entries, using the same 10 predefined labels. (artificial intelligence)

We compared the three participants’ solutions pairwise and found that they matched 44%, 46% and 58% of the time, respectively (see Figure 1). There was complete agreement (where all three participants chose the same label) for one-third of the texts, while there was complete disagreement (where three different labels were assigned) in 20% of cases.

1. ábra: A három értékelő személy megoldása közötti átfedések
Figure 1: Overlap between the solutions of the three assessors

This is not a sign of carelessness, but a question of the precision and clarity of the labels. It is not only the names of the labels that matter, but also how many labels apply to a given text. If several labels fit a text but only one can be selected, that does not support a clear decision. The precision and clarity of the category system is a key issue for both people and AI. (artificial intelligence)

Where people’s decisions differ, we cannot expect AI to give a uniform answer either. (artificial intelligence)

Summary of results

2. ábra: A vizsgált modellek megoldásainak konzisztenciája és minősége (az emberi megoldáshoz viszonyítva)
Figure 2: Consistency and quality of the solutions produced by the models studied (relative to the human solution)

Of the four models, LLaMA offered the best balance of consistency and quality: it was 95% consistent, and its solution matched human labelling 63% of the time (see Figure 2). Claude’s solution also matched that of the human participants 63% of the time, but this model was less consistent (63%).

It may seem low that the AI solutions matched those of humans at most 63% of the time. However, the human participants did not think alike either: during labelling, they reached the same decision at most 58% of the time. (artificial intelligence)

In this light, AI–human agreement of around 60% is not weak performance, but an approximation of the level of agreement between different ways of human thinking. (artificial intelligence)

What does this mean in practice?

It does not mean that AI is unreliable, nor does it mean that it replaces people. Rather, it means we need to ask the question differently: not “Which AI is best?”, but “Which model is suitable for which task, under what conditions?” (artificial intelligence)

  1. If large volumes need to be processed automatically, LLaMA in structured mode is the most balanced choice.
  2. If we want a quick, exploratory analysis and the results will be reviewed by a person, Claude or ChatGPT can also do the job. In this case, careful prompt development and testing are essential.
  3. If the category system is well defined, RoBERTa may be the best solution after fine-tuning.

Content labelling is not merely a technological question; it is also a task of interpretation. Model performance can only be properly understood in the context of the quality of the category system, the task definition and human context.

Even at its best, AI does not deliver absolute truth; it approximates human consensus. (artificial intelligence)

Author: Zsuzsa Székely, Experience Designer at Works.

Next post

Where is the design profession heading? Customer-centred thinking in an AI-accelerated world – Part 2 (artificial intelligence)

Read more