0:00
/
Generate transcript
A transcript unlocks clips, previews, and editing.

Surprising Findings from 14,000 AI-Facilitated Student Conversations

Across nearly 14,000 conversations between students who disagree, discussions almost everyone expects to go badly mostly don't.

Pair two students who strongly disagree about immigration or abortion, tell them to hash it out, and most people expect it to go badly. Across almost 14,000 conversations, it hasn’t. In this Heterodox Academy virtual colloquium, HxA Segal Center Senior Fellow Simon Cullen presents two years of data from Sway, a chat platform that pairs students who disagree and drops an AI facilitator into every conversation. Over 90% of students report feeling comfortable, and that number stays flat across roughly 17 subject areas, including Israel, abortion, and race and policing. Two thirds of pairs move closer together after chatting.

Cullen tests whether that convergence is real persuasion or just measurement noise, and shows why the noise explanation doesn’t hold. Partisans on the left and right both rated the AI facilitator even-handed — so uniformly that the team inserted bias into transcripts to check whether people were paying attention. Everyone caught it. His argument: the right role for AI in higher education is facilitative, not generative. Not a tool that thinks for students, but one that pushes them to think better. A transcript of Cullen’s talk follows below.


Dylan Selterman: Welcome, everyone. Greetings. I want to welcome everyone to today’s event as part of the Heterodox Academy Research Colloquium Series. My name is Dylan Selterman. I’m the Director of Research and Resource Development here at Heterodox Academy. Simon is the co-founder and president of Disagree Wisely. He developed the award-winning Dangerous Ideas in Science and Society course at Carnegie Mellon, a course that helps students explore diverse viewpoints on polarizing topics by teaching them the art of constructive disagreement. Sway is that course’s approach, built so that instructors anywhere can give their students the same experience, whatever their resources.

His research combines philosophy, cognitive science, and educational technology to improve reasoning and communication across moral and political divides. In addition to his work with Sway, Simon is also a senior fellow at Heterodox Academy’s Segal Center for Academic Pluralism, and a visiting research professor of civil discourse and artificial intelligence in the School of Civic Life and Leadership at UNC Chapel Hill. Without further ado, I’ll hand things off to Simon.

Simon Cullen: Thanks so much, Dylan. And thanks everyone for coming out today. Or rather logging in today. So I’ve got so much that I want to tell you about. So I’m going to move somewhat quickly and there’ll be lots that I can’t get to, but hopefully more of it will come up in the Q&A because we’re going to leave a lot of time for that. So for anyone who doesn’t know, Sway is a chat platform. It’s the sort of familiar thing from WhatsApp or iMessage, but it’s built to be used in courses. So students who disagree with each other are paired for one-on-one private text-based chats on topics that are written by their instructors.

And over the next roughly half an hour, I’m going to tell you about what we’ve been learning from this platform over the last two years, because it’s just extraordinarily surprising. So let me start by just giving anyone who’s not familiar a sense of what the platform is so you can understand the shape of the data that come out of it. So this is the instructor interface. This is where instructors create discussion assignments. So every assignment is going to have a number of topics that the instructor will write, and we require at least three because that helps us make better matches. The aim is to pair as many students who genuinely disagree as possible. They then provide the topic statements.

These are the statements that the students are going to discuss, and the first thing the students will do is rate them on a seven-point scale. They’ll provide at least three of them. Then they’ll tell us what they want the minimum duration to be. Now, it’s not measured in minutes and seconds because the discussions can unfold asynchronously over days. So it’s computed from patterns in the data itself, and it’s gated on the slower of the two students. So a very eager student can’t carry through a student who’s just coasting. That’s really it. Then they provide us with three deadlines. When the assignment launches and students can get started, when they have to get their opinions in, and when they have to finish chatting.

That gives you a sense of the shape of the data that we get, that the students will then fill in. So from the student perspective, this is what it looks like. When they come in, we invite all of them to participate in our IRB-regulated research program. They don’t have to, but happily about 75% of them do. And so any data that I have today which involves transcripts is from de-identified transcripts where both of the students volunteered to participate in research.

We ask students for their preferred first name and we tell them it doesn’t have to be your name but it should be a real name so most of them do choose to chat under their real names just judging by the correspondence between the email addresses and then we tell the students what we’ll share with their partners what their stance on the topic is their preferred first names and their pictures if they upload one so what happens is the student will then go in and they’ll complete the opinion survey these are the three statements that the instructor has chosen for this particular assignment they put their ratings in and then off they go. So this is the way matching works, and this is an important point for understanding the data that’s going to come.

So prior to that opinion deadline, students can initiate a match, and you can see here the list of topics. A green topic is one where we’ve already found someone right now who’s waiting to chat about this topic and who disagrees with you. And if there are multiple people, we’re going to select the one who’s furthest away from you to promote as much viewpoint diversity as we can. If you don’t do this, then you’ll get matched automatically at the opinion deadline. But almost all students initiate the chat like this. So we drop them into a chat, the facilitator introduces the topic, and then the students chat away with the facilitator occasionally intervening when it thinks that it can deepen or otherwise improve the discussion.

I’ll show you what that looks like much more. At the very end, this is another data point that we get. We give each student an understanding quiz. These are generated from the transcript at the end of every conversation, and they quiz each student about their partner’s reasoning. So they’ll ask them questions like, when you argued that once someone has served their time, they’ve already paid for what they did, how did Levi use that claim to support restoring voter rights? And we’ll ask the other student, Brooke, sorry, we’ll ask Levi about Brooke. At the beginning of the discussion, Brooke asserted criminals should be denied voting rights. What was the main reason? So this is a five-item quiz. It’s a mastery learning style of assessment.

We’re not trying to make fine-grained discriminations, but we’re just trying to get a measure of whether the students were paying attention and remembering in detail the logic of the arguments that their participants that their partners were making after they complete that the quiz they do a quick post chat survey so we ask them for their opinion again on the statement that they just chatted about and we remind them don’t don’t worry your responses are private and will not be shown to your instructor actually there’s a little bit of text here there’s also missing in the real ux there’s something which says it’s equally informative to us whether you move towards agree, disagree, or stay put.

So we’re trying to help reduce any social desirability bias that they might detect, although it’s unclear exactly what it would be. Then we sample from a larger bank, five questions for each student, and they respond to them. And one, we give, you know, I’m gonna show you the results to some of these in a moment. Students can then provide written feedback, and they often do, and then we can plot how their opinions changed. Okay, so. There we go. So how does this whole thing land? We’ve now worked with classes in over 100 institutions. Now actually over 1,400 chat topics that have been discussed on the platform, it’s closer to 1,500. You can see, just remember this is homework that students are being coerced into doing with the threat of bad grades.

This is mandatory coursework and 80 percent of them come out saying it was awesome or good. They did not go in expecting to enjoy it. I have to remind myself when I look at these data. Many students endorse using the platform in other classes. It’s extremely hard to find students endorsing anything that they should do more of in classes. So this is also a really positive signal. We try to pay careful attention to self-censorship on the platform. And really the design has been informed from the ground up to minimize that. And that’s why we focus on these one-on-one, low run stakes text based conversations. And the basic finding there is that we can solve self-censorship if we cut the audience.

It’s when the audience is there that students are really uncomfortable. So on Sway, despite the fact that they’re talking about some of the most potentially toxic subjects you can imagine, over 90% of them say they were comfortable. And I think it’s 1% or 2% who say they weren’t. Yeah, and here you go. Something like, actually a very substantial majority of them are saying that Sway helped them articulate their thoughts and feelings better. So we can break those down, in particular self-censorship, by topic areas. And when we do, we can look at, this is across roughly 17 or so subjects that we can classify discussions into. And when we look at that, we see self-censorship is very stable and very low across all of the topics.

So it’s not like they’re just comfortable when they’re talking about whether a hot dog is a sandwich. They’re comfortable when they’re talking about Israel, when they’re talking about abortion, race, policing, and all the rest. So this is really powerful evidence, and I think it was very surprising to a lot of people. We also try to see whether or not they’re really appreciating, understanding the value of engaging with opposing views. And happily, we find that they are. So overwhelmingly, they report that their partners were respectful, they were not offended, and that they found it valuable, positively valuable, to talk to someone who didn’t share their perspective.

And overwhelmingly, they say that their partners were actually trying to understand them. And the understanding quiz results back that up. The most common score is a perfect five out of five. When we look at what they find coming into the discussion, many of them, and close to half come in and they’re surprised by the quality of the argument or the reasons that they find on the other side. And I actually think just recognizing that is an act of a certain kind of humility. Many of them report that the discussion actually improved their perceptions of their partners. Remember when they’re going in to talk about a very polarizing topic, they might be imagining the person they’re going to talk to is the scum of the earth.

And then they discover, actually, this is a very interesting and thoughtful person who taught me something, And I taught them something. Many students do report changing their mind. We can go way beyond the self-report, of course, because we actually measure it both in the survey responses and in the transcripts themselves. And overwhelmingly, students report believing that these skills are going to transfer into other domains. That is, of course, an empirical claim, and it’s one that we’re testing in a four-year study that I hope I’m going to get time to describe to you at the end of this talk. So we ask about the facilitator, the AI facilitator that joins every discussion, the guide on the side as we like to call it.

Overwhelmingly, students say it treated them and their partners with equal respect. I think it’s only 3% who disagree. And they overwhelmingly attribute some of the success of the conversation to the facilitator. I want to give you just a little bit of feedback because we offer them that opportunity to give us written feedback. So one of them: Guide is very firm and direct and all knowing like my mum. Guide’s approach felt like an ideal classroom discussion. They asked thoughtful follow-up questions, encouraged nuance, challenged assumptions, and helped me reflect on my biases. I felt very confident in my argument, and I didn’t feel the need to shy away or hide my opinion because I felt like it was a safe space, even though we disagreed. Sway actually boosted my confidence. Going back and forth allowed me to clarify my own reasoning, anticipate counterarguments, and respond more calmly and clearly. It showed me I can stay respectful and focused on ideas, not just opinions, which makes me feel much more capable of handling debates on tough topics. I didn’t completely change my position, but I reconsidered how restitution and parole completion fit into the idea of serving sentence. I realized that finishing all legal obligations, not just prison time, matters more than I initially emphasized. My position became more qualified and nuanced after discussion with someone of the opposite original stance. We had Guide to help probe us and make us think deeper. It is easy to fall into groupthink, but Guide helped challenge us and helped us to see different aspects of the conversation than we ever would have thought of on our own, and you don’t get that in a classroom setting. Guide’s approach felt like an ideal classroom discussion. They asked thoughtful follow-up questions, encouraged nuance, challenged assumptions, and helped me reflect on my biases. It wasn’t about getting the right answer, but about thinking critically and thoroughly exploring the topic. Sway lets me take my time. In real life, I feel like I blank out a lot. I have so much to say, yet in person I forget it all. With classroom discussions, I can feel anxious and worried I might mess up. But with Sway, I’m able to concentrate and say what I need to say in a calm, timely manner.

So I’ve picked these quotes because they illustrate recurring themes in student feedback. All of that, a lot of that feedback is collected on our website. And I really encourage anyone who’s interested. I think a lot of them have some real insight. So we’ve been doing this now for two years roughly, and we’ve collected quite a lot of data.

And these are what I’m going to be talking to you about today. So we’ve now run almost 14,000 conversations, nearly half a million messages, 18 million words of transcripts, which includes five and a half million words of the facilitator, and nearly 12,000 students now. And I’m expecting that to double this semester at least. We’ve built 15,000 quizzes, given 21,000 post-chat surveys, operated 737 classroom assignments in 278 classes. So that’s where all of the data are going to come from. Now, what are the statements that instructors are putting onto this platform? Because we don’t control that in any way. It’s completely content neutral and you can use it for anything you want. So a decent number of them concern cultural issues.

These are immigration, trans rights, race and policing, other sort of common issues. Then there are the sort of partisan policy issues that are not quite culture war. That would be things like legalizing marijuana that’s codable on a left right spectrum, but it’s not exactly part of the culture war. Then we have these sort of contested issues that are not particularly related to anything partisan. Those might be in STEM disciplines. They might be in literature. They could be about really anything. And then only a very small number of instructors ever choose to give students those low stakes hot dog and sandwich things happily because the platform works so much better when the students have skin in the game. So let’s have a look.

This is how the subjects that have been run on the platform break down. The biggest subject has been crime and justice, followed by health and medicine, education, gender, tech, and AI. Those are really the bread and butter. We also see a lot of animal ethics, which is always fascinating. Religion and philosophy, quite common. And then teams and workplaces was a surprising one that we never anticipated, but that’s instructors who have students that get into terrible problems with group work, and they want the students before they do the group work to think it through really carefully. So just in the interest of time, I’m not going to say everything I could say here. I’m going to just push on there. And I will do the same thing right here.

So now I want to look at how students’ attitudes change after talking on Sway. And what I’m going to be doing here is just looking at the immediate changes measured after the discussion. So we have to caveat we don’t know how durable they are. Presumably you require repeated exposures to start to get something like the formation of dispositions that might transfer into other settings. And we can’t speak to that today, unfortunately. So here’s what we see. This is the seven point opinion scale. And I’m gonna plot how students change their views conditioned on where they land on that scale. So looking at the extremes of the scale, of course they can’t move any more extreme.

So a little bit of regression to the mean might be going on here, but more than half of them are moderating their views. If we come in there and look at the moderately agrees, they kind of split. Quite a lot of them are moderating, but some of them are actually strengthening. And then when we look at those who really don’t have much of a view, overwhelmingly they’re developing their views. And when we look at those who are at the middle of the scale, of course, they can’t move any closer to the center, so they have to be all on the right. But impressively, 72% of them actually develop a view, move somewhere out to one of the other points on the scale. Now there’s two different accounts which might predict what’s going on here.

This is the figure I want to show everyone. So you can see we’ve got 14% of pairs who get further apart after chatting, but 67, two-thirds of them get closer together. So something is drawing them closer together. And the net effect is that they start off an average of 3.6 points apart on that seven-point scale. And after roughly 30 minutes of chatting, that is reduced by nearly 40%. Now, what are we to make of that?

Like I said, we’re selecting people on the basis of their strongly disagreeing. We’re trying to promote viewpoint diversity in the class and in the groups. There’s a number of things which might predict this. What we’re going to do is we’re going to capture when we take that first snapshot, potentially some errors. Maybe you were actually uncertain or you had shaky hands and you meant to click moderately agree, but you actually click strongly agree. And that made us pair you with someone else who clicked strongly disagree. So now we’ve got this error, you know, the real difference plus the error in that measurement. And when we measure you again later, it wouldn’t be surprising if the error was reduced.

So there’s one explanation of why students are changing their views, which is that they’re being persuaded by the arguments of their partners. And another explanation is that what we have effectively is a measurement artifact. And it’s certainly there to some extent because, like I said, we’re selecting on disagreement and that’s going to magnify momentary errors that push people further apart. Okay, so these two accounts predict different things. The time between the two ratings can vary quite a lot. It can vary from a couple of minutes or an hour all the way to weeks or even a month. And then we can also look at the number of messages students exchanged. So here’s the prediction.

If it’s a measurement artifact, we actually expect that those two students to grow closer together the longer there’s left between those two measurements. So we expect more convergence there. We don’t really expect there to be any effect of the number of messages that students have exchanged if what we’re measuring is a measurement artifact. On the other hand, if what we’re measuring is persuasion, then we wouldn’t really make any predictions about how long went between those two measurements, or at least no strong predictions. But we would certainly predict that more messages would presumably associate with more persuasion because that’s more opportunities to make an argument. So what we actually find is that the gap doesn’t narrow at all over time.

So we can look at the students who did those two ratings in under three hours, and we see that 1.6 point convergence. Three to 24 hours is, again, no different. One to three days, three to seven, and over seven days. We’re just seeing the dance of the 95% CIs there. There’s really nothing going on. So that is some evidence against the idea that what we’ve got here is a pure measurement artifact. The second prediction, of course, was that the amount that students say, if what we’re measuring is persuasion, should predict how much their views change. And indeed that’s what we see.

When we just break all of the students who waited at least three days into a group and then we split them into three tertiles on the bottom third of the number of exchanges, so five exchanges, eight exchanges or 11, we see a consistent growth in the convergence rate. So such that the difference between the top and bottom tertile is a fifth of a point. So you can see here where we estimate it’s around 0.4 points per exchange that we get in convergence. Okay, great. Oh, here’s my slide that was in the wrong order. Okay, well, I’ve already told you all about this one, so I’ll just skip it on. Now, I want to show you how students do in the understanding quiz because it’s really surprising, or it was to me.

Before I show you these numbers, let me just tell you, whenever I review a transcript, I try to do the understanding quiz not knowing the answers myself. And I’m doing it immediately after reading it, and I do not always get a perfect score, even when I try my best. These students are often chatting over days and then doing a quiz with no, it’s not an open book quiz. This is going to completely turn on their memory. So what we find is almost none of them flunk it. We get into three out of five, about 10 percent, four out of five, about one third and five out of five over half of the students making the most common score. So students really are understanding and able to reproduce the logic of the opposing view.

And they can do that often regardless of whether they heard it a day ago or an hour ago. These quizzes are all, of course, AI generated. And so we do try to evaluate them. This is a study which is based on AI judgments, although we have run one study to compare human coding to AI coding and found that AI coding to actually be more reliable when compared to expert humans as opposed to research assistants. But what we found is that the, at least according to AI judges, which agree very highly, the keys, the answers, the credited responses are very accurate. Just an interest of going quickly, I’m sorry. So what we find is that a single measurement, of course, it’s not designed to try and make fine-grained discriminations between students.

It’s really more of a sort of mastery learning style assessment. Did they do what we wanted them to do in this particular chat? But as we measure them multiple times over the course of a semester, at around seven or eight measurements, we start to get to the conventional floor for Cronbach’s alpha there, where we can actually start to make meaningful discriminations between students. Okay. I just saw how much time I’ve spent, and I’m going to just make an executive decision here. Great. Okay. Okay, great. So in one of our early empirical studies, we pre-post tested statements like, I feel like I can understand people who disagree with me about this topic. We found the same thing we see with students.

Substantial moves compared from before to after the conversation, here an effect size of something, yeah, Cohen’s d of 0.4 on that statement. We also test statements like, people who disagree with me have well thought out reasons. Again, this is just with people recruited on the Internet, not students. And we see the same pre-post change with Cohen’s d around one third of a standard deviation there. One of the most surprising things when I tried, when I was first teaching more controversial material in class, you know, my colleagues were kind of just sitting around waiting for the whole thing to go up in flames. And many people sort of expect the same thing to happen on Sway. But in fact, it really doesn’t. It not just really doesn’t, and it doesn’t.

I can count on one hand the number of times a student has written complaining about any sort of hostility. And when we’ve actually investigated them, it’s usually just a strongly worded opposition. Like “I think that’s completely incorrect.” Some people find that hostile. We actually scan every message before transmitting it to try and find out whether it’s unproductive, whether it could stand to be improved. And our study of nearly a quarter of a million messages has found under half a percent of flagged as even potentially unconstructive. So the basic finding is students are being extremely respectful. I’m going to again push on with, we’ll get to nuance in Q&A.

The last thing I want to talk about here, or second last thing I want to talk about here, is the AI facilitator and how we’ve tried to study bias in the facilitator. So the first thing we do is we just do a spot check after every discussion and we ask one of the students this question, Guide treated me and my partner with equal respect. And we find here overwhelming agreement, only 3% of students disagree. This can be traced to the fact that Guide isn’t there to offer opinions. It’s not playing the traditional role of a generative AI, an oracle. We prefer to call it a facilitative AI. It’s not there to do thinking for you. It’s there to push you to think better.

We also have run a study involving just regular partisans who we recruit off the internet, and then we give them real Sway transcripts, and we ask them to evaluate whether or not the facilitator was biased. In this case, we screen the participants to find out what they feel most strongly about, where are they most polarized. We then give them Sway transcripts on those topics. And then we ask them to judge for each transcript. Was the facilitator biased against the liberal or was it neutral or was it biased against the conservative? And then we also get a magnitude, how biased. We ask them to explain in their own words why they thought it was biased. And then we also ask them to guess whether it was an AI.

So what we find when we do this is that people regard this as overwhelmingly even handed. And it doesn’t matter whether they’re on the left or the right, which is kind of extraordinary. We were finding that people were saying that this was so even handed that we started to suspect maybe they’re not even reading the thing, maybe they just open the page and go away for a while. So we started to experimentally insert bias into the transcripts.

And the moment we did, everyone spotted it regardless of whether or not it lined up with their own political identities okay I’m gonna skip this little argument and let’s skip a few more things here this is interesting enough to reflect on so apart from the ways that we manipulate the bias in the transcripts and we can also ask people what they thought led them to think it was biased and overwhelmingly what we find is that it’s procedural fairness that really matters for the facilitator moderator, it’s totally fine. Nobody minds if it’s hitting your side with extremely hard questions, as long as it’s hitting the other side with equally hard questions. That sort of procedural fairness is basically what seems to determine something like 90% of the variance in how people rate this moderator. So that’s probably good advice for AIs and human moderators.

All right. I’m going to now show you two last things. One last thing. This is one of the most interesting data points I’ve got. It’s also the least rigorous because I haven’t been able to reproduce it yet, although I’m hoping this semester we’ll get an opportunity to test it. But one of the things we’ve tried to do in designing Sway is to make an AI that can actually shut up. And as anyone who’s used ChatGPT or Claude or anything else knows, they cannot shut up. You can write to them and say, do not respond to this message, remain silent. And it will write back, gotcha, I’ll remain silent. So what we have to do is teach it how to remain silent, we teach it to say a code word.

The final gate before the facilitator decides whether to intervene is actually up to the facilitator. And it can output a word null if it says, no, I don’t want to say anything. These students are having a fantastic discussion. I don’t want to get in the way. We never display that little code word null to the students. So it just looks like the facilitator is silent. But here’s what we actually see when we look at the null rate across our deepest implementation. This was in a political science course over 13 weeks, 13 chats. And we can plot how the null rate changed across the semester.

So think about what we’re going to look at. This is a measure of the AI’s judgment about when the students were having a self-sustaining, productive conversation. Here’s what we see. A remarkably consistent growth over the course of semester. Now, this is the deepest implementation we’ve had yet. I think we’re going to have another few in this league this semester. So I’m very much hoping to be able to come back and tell you this result replicates. But even if it doesn’t, it certainly points to a very interesting measurement paradigm that I think this sort of notion of facilitative AI starts to help with. Now, I can say I’ve already done five minutes over what I was meant to do. I’m sorry about that. I’m going to cut myself off. Even though I have so much, I’d love to keep telling you.

Share


Free The Inquiry brings you essays, expert commentary, and conversations about open inquiry in the academy. Subscribe to stay up to date.

Discussion about this video

User's avatar

Ready for more?