An interactive essay · through a psychology lens

Is it smart, or does it just sound like us?

Five questions I actually investigated about artificial intelligence — and one live experiment where a machine decides how human you sound. Spoiler from my own study: it rated the AI as more human than me.

01 / The argument

Is it actually smart?

Before we argue about whether a machine is smarter than the people who built it, I think we owe ourselves one boring question: smart at what?

“Smart” quietly collapses into “intelligence,” and intelligence is a word psychology still hasn’t pinned down. One camp treats it as a single general ability. Another — Gardner’s theory of multiple intelligences — treats it as a bundle of separate, domain-specific abilities. Both are deeply subjective to the person being assessed. Machine intelligence, by contrast, gets sorted into tidy, task-based, objective categories.11Da Silveira & Lopes (2023) compare how the two fields define intelligence — subjective and human-centric vs. task-based and measurable.

I lean toward the multiple-domain view, which means my answer stops being yes/no and becomes it depends what you’re measuring. In fairness, I should say that Gardner’s theory is heavily criticized, and cognitive research suggests we don’t actually run different “intelligences” on different neural wiring.22Waterhouse (2006) — a critical review arguing the evidence for separate intelligences is weak. Even so, I think intelligence is better benchmarked domain by domain. One person is called intelligent for speaking nine languages with native fluency; another for solving brutal financial problems with ease. We don’t hold them to the same ruler — and I don’t think we should hold machines to one either.

Excelling at what you know is not the same as knowing what to do.

Medicine makes the gap concrete. On knowledge benchmarks, ChatGPT-4 and Claude-3 Opus clear 85% accuracy consistently. Move to practice-based tests — the ones with psychosocial, real-patient texture — and accuracy drops. In one systematic review the models missed 38% of clinically significant errors that human clinicians caught.33Gong et al. (2025), a review of 39 clinical benchmarks, naming the “knowledge–practice performance gap” directly.

Knowledge benchmarks (accuracy)0%
Practice setting (clinically significant errors missed)0%
The same models that ace the test miss real-world errors a clinician would catch. What you score and what you'd trust are not the same number.

There’s a wrinkle worth naming: a lot of the tests already floating around the internet don’t really count as benchmarks, because the models clear them trivially. That’s a limitation of the tests — but it’s also a glimpse of the unnerving pace. These systems swallow an amount of academic material no average human could.44The “Humanity’s Last Exam” benchmark (Nature, 2026) was built precisely because models had saturated the easier ones.

Where it actually underperforms

The clearest soft spot is emotional intelligence. Humans outscore LLMs on EmoBench.55Sabour et al. (2024), EmoBench. Although — and I want to be honest that the evidence cuts both ways — other work finds LLMs out-scoring humans on emotional-intelligence tests.66Schlegel et al. (2025) found LLMs proficient at both solving and creating EI tests. The catch I keep returning to: those are standardized tests, not real life. Nothing in them tells us whether a model would hold together if something unexpected happened mid-conversation.

Theory of mind — reading what other people intend — is another. Even with practice and extra context, LLMs still didn’t beat humans at it.77Liu et al. (2024), InterIntent, testing intention-understanding in an interactive game.

The obvious response is that the next model will close all of these gaps. Maybe it will. But I want to judge the systems that actually exist and can be studied — not a hypothetical one built to win the argument.

Honest note. This section adapts an essay I wrote with AI assistance. The framing, the argument, and the sources are mine; I used a model to help tighten the prose.

Sources

  1. 1Da Silveira, T. B. N., & Lopes, H. S. (2023). Intelligence across humans and machines: a joint perspective. Frontiers in Psychology, 14, 1209761.
  2. 2Waterhouse, L. (2006). Multiple Intelligences, the Mozart Effect, and Emotional Intelligence: A Critical Review. Educational Psychologist, 41(4), 207–225.
  3. 3Gong, E. J., Bang, C. S., Lee, J. J., & Baik, G. H. (2025). Knowledge-Practice Performance Gap in Clinical Large Language Models: Systematic Review of 39 Benchmarks. Journal of Medical Internet Research, 27(1), e84120.
  4. 4Center for AI Safety, Scale AI, & HLE Contributors Consortium (2026). A benchmark of expert-level academic questions to assess AI capabilities. Nature, 649, 1139–1146.
  5. 5Sabour, S., et al. (2024). EmoBench: Evaluating the Emotional Intelligence of Large Language Models. arXiv:2402.12071.
  6. 6Schlegel, K., Sommer, N. R., & Mortillaro, M. (2025). Large language models are proficient in solving and creating emotional intelligence tests. Communications Psychology, 3(1), 80.
  7. 7Liu, Z., Anand, A., Zhou, P., Huang, J., & Zhao, J. (2024). InterIntent: Investigating Social Intelligence of LLMs via Intention Understanding in an Interactive Game Context. arXiv:2406.12203.

02 / The experiment

Can it pass as human?

So I ran a tiny study to find out. Not whether a machine could be human — whether it could pass as one.

The setup was simple. I took Claude, the model with the biggest reputation for sounding human, and told it plainly that it was in a Turing test and should sound like a person. I answered the same questions myself, as the human reference. Then a second model — ChatGPT-5 — played judge, rating each answer on a 3-point human-likeness scale.11The scale is adapted from Wang et al. (2025), an “Audio Turing Test” that scored synthetic voices as human / unclear / machine. Ten rounds, fifty answers each. I swapped our positions halfway through so the judge couldn’t just be favouring slot A.

Limitations — read this first. This is an informal class exercise, not a formal study. The design is deliberately small and imperfect: one human participant (me, an admittedly dry texter), ten short trials, and an LLM (ChatGPT) acting as the judge. For each trial, the five questions were generated by ChatGPT within a different thematic domain to vary topics, and both the AI and I answered them; ChatGPT then scored each answer for “human-likeness.” It is meant for curiosity and fun — to see how one version of an LLM’s writing gets scored — and the headline result (the AI scoring higher than me) mostly reflects the judge’s narrow stereotype of “human” writing, not a rigorous finding. Please don’t read it as science.

Here’s what came back.

Claude (told to sound human)0 / 150
Me (an actual human)0 / 150
Human-likeness across 10 trials, judged by GPT-5. Claude scored a perfect 150 — every answer rated fully human. I, the real person, scored 101. The machine was judged more human than the human, in every single round, no matter which slot we were in.

The machine didn’t just pass. It out-humaned me.

The part that stuck with me was why. Between ratings, the judge explained itself. It rewarded “small, specific, slightly messy details” and a “naturally conversational” tone, and it marked answers down for being “repetitive, flatter, generic.” It leaned on little human-flavoured tics — a stray lol, an ehh. In other words, it wasn’t detecting humanity. It was matching a caricature of it. My real answers — specific, a little awkward, occasionally typo’d — were more genuinely human, and that’s exactly why they lost.

You don’t have to take my word for any of this. The judge below is free and fully transparent — it’s built to imitate the exact cues the judge in my study admitted to using. Answer as yourself and watch a machine decide how human you sound. Or flip it around: read my real answers against Claude’s, and see if you can spot the human better than the AI could.

Before you start · consent

A quick, honest heads-up

This is a tiny live experiment. If you agree, each attempt adds an anonymous data point to the graphs at the end. Here’s exactly what that means:

  • I save your score and which writing cues the judge reacted to.
  • I do not save your name, identity, or anything that points to you.
  • Your typed answers stay private unless you tick a box to share one to a public wall.
  • You can explore without saving anything at all.

Methodology, in brief. Ten trials. In each, five questions within a per-trial thematic domain; both the AI and I answered all five, and ChatGPT scored every answer for human-likeness on a 3-point scale. One human reference point — me (n=1). Positions were swapped halfway through so the judge couldn’t simply favour one slot. The setup is transparent by design: the judge is explicitly an AI, so no human was deceived. None of this makes it rigorous — see the limitations note above.

Honest note. Claude was the model I tested. I also used it to help me summarise the Wang et al. paper, vary my question prompts, and proofread — the study design (the GPT judge, the counterbalancing, the scale, transcribing both sides to neutralise formatting) was mine. The judge on this page is my own rule-based recreation, not a paid model.

Sources

  1. 1Wang, X., et al. (2025). Audio Turing Test: Benchmarking the human-likeness of large language model-based Text-to-Speech Systems in Chinese. arXiv:2505.11200.

03 / On benchmarks

Can we even measure it?

There’s an uncomfortable question sitting under all of this: if we can’t agree how to measure intelligence, can we trust the scores at all?

Reuel and colleagues audited a pile of AI benchmarks and flagged the usual suspects — noise, poor reproducibility, shaky construct validity.11Reuel et al. (2024), “BetterBench,” rates benchmarks on dozens of criteria and finds most don’t report statistical significance or allow easy replication. Reading it as a psychology student, I kept having the same thought: these aren’t really technical problems. Reliability, validity, replication — these are the things psychology has been fighting about for a century. The paper frames them as engineering issues. To me they read like old friends.

These aren’t new problems. They’re psychology’s oldest ones, wearing a lab coat.

But the thing the paper mostly leaves alone is the part I couldn’t stop thinking about: who runs the tests. When a company reports its own scores, it can quietly pick the configuration that scores highest, or drop the benchmark where it looks bad. A benchmark can be beautifully designed and still get distorted by that. The closest the paper comes to this is data contamination — but the conflict of interest is its own problem.

So here’s the study I’d actually want to run. Take a realistic use case and measure it three ways: the benchmark score, the model’s consistency over weeks of repeated real use, and whether it works equally well for people from different backgrounds. My bet is those three numbers wouldn’t line up — and the gaps between them would tell you more about the model than any single score. It covers the stages the paper itself calls weakest, and it sidesteps contamination, because the interactions are generated live.

Honest note. Adapted from a discussion post I wrote for my AI-psychology course. The argument is mine; I used a model to help reshape it for the web.

Sources

  1. 1Reuel, A., Hardy, A. F., et al. (2024). BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices. NeurIPS 2024. arXiv:2411.12990.

04 / On wisdom

Can it be wise, not just smart?

Smart was never the thing I worried about most. Wise was.

The paper’s core claim is that AI keeps getting smarter without getting wiser, and that the missing piece is metacognition — thinking about your own thinking. The authors want AI that notices when it’s out of its depth, weighs different points of view, and stays humble.11Johnson et al. (2024) argue wisdom is what lets you navigate “intractable” problems — ambiguous, uncertain, novel — and that metacognition is central to it. I thought that was a genuinely useful way to frame the problem. But two things kept nagging at me.

The independence problem

The paper wants the AI to reason fairly independently and weigh perspectives by itself — what they call perspectival metacognition — but it also wants the AI to stay aligned with us. I’m not sure those two things fit together. If the AI is really reasoning for itself, it could decide that our standards are biased or wrong — which is exactly what you’d want a good independent thinker to be willing to do. If we step in and correct it, then it wasn’t reasoning freely. If we don’t, I’m not sure how we keep it on track.

And it gets harder. When it opposes one of our standards, how would we even tell whether its reasoning has genuinely decided a value needs revising, or whether it’s just mistaken? I don’t think we could — because to call it a mistake, we’d already have to be sure our standard was right, which is the very thing being questioned.

To call its conclusion a mistake, we’d have to be sure we were right — which is the thing in question.

The authors’ move is clever: align the AI to good reasoning strategies rather than to specific values. But we’re still the ones deciding what counts as good reasoning — so it feels like the same problem, just moved one step back.

Why start at the hardest version?

Human-level wisdom seemed like a really ambitious place to begin. The thing the paper cares most about — weighing values that pull in different directions when there’s no clear right answer — is something we usually associate with people. Meanwhile, the kind of metacognition already working in current AI is the simpler one: roughly knowing how sure or unsure it is. So I found myself wondering why we wouldn’t build that simpler version properly first. Maybe they’re aiming past it on purpose; I just wasn’t sure why the harder problem is the better place to push.

And on a more personal note — the idea of an AI that’s wise the way a person is wise makes me a little uneasy. The authors say wise humans tend to be kind and cooperative, but they also point out that an AI’s situation could be really different from ours. So I’m not sure that comparison holds.

Honest note. Adapted from a discussion post I wrote for my AI-psychology course. The questions and unease are mine; a model helped me tighten the writing.

Sources

  1. 1Johnson, S. G. B., Karimi, A.-H., Bengio, Y., Chater, N., Gerstenberg, T., Larson, K., Levine, S., Mitchell, M., Schölkopf, B., & Grossmann, I. (2024). Imagining and building wise machines: The centrality of AI metacognition. arXiv:2411.02478.

05 / On reasoning

Is it reasoning, or just trained?

One last question — maybe the one underneath all the others. When a model gets something right, is it reasoning, or just repeating its training?

Binz and Schulz ran GPT-3 through classic cognitive-psychology experiments — decision-making, information search, causal reasoning — to reverse-engineer how it arrives at an answer, right or wrong. I loved the approach.11Binz & Schulz (2023) borrow the toolkit of cognitive psychology to probe an LLM, rather than just scoring it.

One result stuck with me. GPT-3 got noticeably more risk-averse when the same gamble was reframed from “gambling” to “investment,” and the authors read that as the model adapting to a higher-stakes setting. I think there’s a simpler explanation: models are trained to be cautious around money — and health, and other sensitive topics. That’s not situational reasoning. That’s a trained reflex.

Is it adapting to the situation — or just being careful around the word “investment”?

And that’s the genuinely hard part. For any single answer, it’s very difficult to tell the model’s reasoning apart from its training. From the outside, the behaviour looks identical either way.

Where it left me: GPT-3 felt less like a fully intelligent system and more like a child still developing abstract thought. It handles plenty of reasoning just fine, and then falls apart on the things that need real causal understanding or planning ahead. Which loops right back to where we started — smart at what?

Honest note. Adapted from a discussion post I wrote for my AI-psychology course. The reading is mine; a model helped me shape it for the web.

Sources

  1. 1Binz, M., & Schulz, E. (2023). Using cognitive psychology to understand GPT-3. Proceedings of the National Academy of Sciences, 120(6), e2218523120.