Feedback mechanisms are terrible data sources
I swear, the most painful part of kids going back to school is me having to wake up on like 5 hours of sleep to get them packed up and ready for school... 😵
If you pay attention to product experiences on the modern internet, it's not hard to find various forms of "feedback" mechanisms. Whether it's star ratings, Likert scales of 5 happy faces, satisfaction mini surveys, the persistent cancer that is Net Promoter Score, and even simple thumbs up/down systems are plastered everywhere these days. Most product organizations adopt one (and often many) of these mechanisms because they ostensibly want to collect information to improve their product over time.
As someone who has sat on the collecting side of much of these product feedback efforts over the past two decades... the regrettable truth is that the VAST majority of these mechanisms aren't being used to make product improvements. There's a ton of complicated reasons why that is the case, and this week I just feel like making a big ol' list of why such efforts are so much less effective than we want them to be.
First, users don't respond often and choose when they respond
As a researcher, I'd love it if every single user would give clear feedback on their experience every single time they touched the product I'm working on, without fail. That way I would have a full population data instead of a sample. I'd also be collecting data at every interaction whether it went good or bad. That gets rid of so much self selection bias inherent in such measures...
But as a user of a product, even I find myself getting annoyed at interruptions trying to collect feedback from me. I'm using a product to accomplish some task, like ordering lunch, and even a single click feedback question on a single screen can be annoyance-inducing enough that I will spend extra time searching for a way to get OUT of giving a response. In my head, I avoid responding because "why encourage the product team by giving them the information they're asking for? That'd only encourage them!" – which isn't particularly true and I literally know better about this whole process than a very large percent of the population! But that's the thought that goes through my head nevertheless because I'm hungry and not in the mood to see an NPS question or whatever.
I'm pretty sure I'm not alone in this line of thinking because it's well known in the industry that response rates to these kinds of prompts for feedback tend to have pathetic response rates. Often it's fractions of a single percentage point of impressions. Of the people who do provide feedback, it's usually some highly self-selected group – often someone who is upset enough to provide feedback, or something excited to give positive feedback. There's rarely any middle ground because those individuals are satisfied enough that they're just going to go on with their life.
The collected data resists interpretation
Since we know people dislike providing feedback already, researchers have little choice but to annoy the user with as few questions as possible. The most common choice is to ask just one single question. When you do that, and the response rate is also terrible, the natural instinct is to ask a question that is pretty broad. "How satisfied are you with our service?" "Is this recommendation useful to you?" and so on. They're also usually multiple choice questions because, again, we know that the vast majority of users can't be bothered to fill out an empty text box with a detailed response. It's our one shot at getting information from the user, so let's please try to get something – goes the thinking process.
The end result is that we collect a data point that is extremely hard to interpret. The most common complaint is that we collect a "thumbs down/negative" feedback and we have no idea what the user means. If we pull these users into usability studies and directly ask them what they meant by giving the negative feedback, we find a menagerie of responses ranging from "I didn't like the response [for purely aesthetic, sometimes even irrational] reasons" to "I was confused what product this was and was angry at another product, not yours". Sometimes we even get fun answers like "I read the scale the wrong way" or "I clicked that one by accident and didn't mean it".
To put it succinctly, if we assume that the collected data has any signal in it at all, it is usually buried under a ton of noise. What's our most robust way to deal with noisy signals – getting a big sample size. Given our typical response rate, how much sample size can we get? ... That's the problem.
Since we usually don't have the sample size that allows the mean/median to give particularly interesting results, researchers wind up doing the other thing that can potentially boost signal – try to filter out "bad" data. That's its own process fraught with issues, but what choice do you have when you need to find results. Very often they try to triangulate the feedback to more objective metrics signals in an effort to identify user feedback that "legitimately associates with a poor experience."
The feedback can't be tied directly into system feedback, so it's very "laggy"
Because all this data is so noisy and hard to interpret, everyone is rightly hesitant to hook this up into a closed feedback/training loop. Add in the possibility that some bad actor may attempt to manipulate the system by flooding the survey with spurious data, and what happens is very little of these feedback mechanisms directly affect how the system operates on a day-to-day level.
Instead, some sort of process is created to take the collected data, analyze it (that's usually us!) and provide recommendations as to what needs improvement. It all sounds great on paper, but everyone involved in the process is super busy and the review process can easily slip through the cracks. Weeks, months, quarters, can go by before an update pass happens. This is especially bad when there's not enough data being collected to do meaningful analysis and there's no choice but to wait. That means that many users will give feedback to a system and then realize that it doesn't seem to actually do anything to their experience. That in turn slowly convinces them that they shouldn't bother providing feedback, which hurts response rates more, and makes everything even worse.
Even if there's signal, it's not really actionable
Let's say you, like many of us now, are tasked with helping evaluate the chatbot du jour of your company. By a miracle, you overcome a lot of the previous hurdles I mentioned and a ton of users helpfully flag a bunch of responses as bad, that's great! But what do you DO with that information?
This is like how my kid will tell me that dinner tastes bad, or that they "don't like chicken". It's information, but just as I know that the flavor tastes fine to my adult tastes, and I've literally seen that kid eat all sorts of chicken foods just the past week, the feedback doesn't help me figure out what I need to improve. I still have to draw on other information to decide what to try doing next. In UX, there's always more ways to do things wrong than right, so getting feedback that merely eliminates one bad version out of nearly infinite bad versions is surprisingly unhelpful.
Most developer teams (rightfully) don't like it when people come up to them with blanket statements of "software not working" that don't highlight a clear direction, and UXR and analyst teams have to spend a lot of time getting to that point.
And so the data languishes...
Given all the issues, as teams have to wait to "collect more data to get signal", priorities shift, re-orgs happen, and before anyone realizes it, there's no one actually paying attention to the data that is actually coming in. The whole operation gets largely forgotten before it has a chance to become useful. Across my career, I've lost count of the times where I've heard this exact conversation:
Exec: "We're collecting [this feedback signal] what are we doing with it?"
Lead: "I'm not sure, let's ask around"
[Asking around happens]
Lead: "So the answer is 'nothing', we're not doing anything with it."
Then, guess who gets to field the "data question" of whether that data is useful for anything.
Oh, and the answer is no. The data isn't useful for anything.
What's to be done?
This is still an open question.
Nowadays I try much harder to target the question in a way to get a useful response when one rolls in, but given the nature of the beast it's a lot of "it depends". The tension between all the factors like how and when you show a survey, what's being asked, whether the user even understands what you're asking is always pulling the data collection in all sorts of directions. Plus, you won't know what works until you test it out live. Each situation is different from the last.
As I mentioned before, I think it's a bit more valuable to get a positive response than a negative one since they're often rarer and is associated with a direction (the current thing vs infinite alternatives). It still doesn't absolve the positive responses from all the potential biases and sampling issues, but at least once the noise settles out you have a hint that something in the design was working and you should be trying to isolate it.
In other products in the wild, I've seen signs that other UX Researchers are coming to similar conclusions because there's been more creativity in eliciting feedback lately. For example, I've seen AI systems present two competing responses that force the user to pick one. I can imagine on the back end how that data is significantly more rich because you can compare the generating parameters on the back end for clearer signal. I've also see follow-up surveys with fine grained questions and attention checks to help weed out bad responses. I'm not sure how much signal those teams are deriving from such more advanced methods, but I sure hope it's more than nothing.
Standing offer: If you created something and would like me to review or share it w/ the data community — just email me by replying to the newsletter emails.
Guest posts: If you’re interested in writing something, a data-related post to either show off work, share an experience, or want help coming up with a topic, please contact me. You don’t need any special credentials or credibility to do so.
"Data People Writing Stuff" webring: Welcomes anyone with a personal site/blog/newsletter/book/etc that is relevant to the data community.
Counting Stuff Official Forums: Discuss posts, or other data topics with the community.
About this newsletter
I’m Randy Au, Quantitative UX researcher, former data analyst, and general-purpose data and tech nerd. Counting Stuff is a weekly newsletter about the less-than-sexy aspects of data science, UX research and tech. With some excursions into other fun topics.
All photos/drawings used are taken/created by Randy unless otherwise credited.
Supporting the newsletter
All Tuesday posts to Counting Stuff are always free. The newsletter is self hosted. Support from subscribers is what makes everything possible. If you love the content, consider doing any of the following ways to support the newsletter:
- Consider a paid subscription – the self-hosted server/email infra is 100% funded via subscriptions, get access to the subscriber's area in the top nav of the site too
- Send a one time tip (feel free to change the amount)
- Join the Approaching Significance Discord — where data folk hang out and can talk a bit about data, and a bit about everything else. Randy moderates the discord. We keep a chill vibe.
- Get merch! If shirts and stickers are more your style — There’s a survivorship bias shirt!