I went back to the primary papers for this piece — not the press releases, not the LinkedIn summaries, the actual PDFs — because the first draft I’d assembled from secondary coverage contained a claim that didn’t survive contact with the source. It said that eye-tracking data showed people looking at AI-generated art differently, implying an unconscious bias that shows up in attention patterns. The study it was citing found the opposite: gaze behavior did not differ based on believed authorship. Only people’s stated opinions changed. That distinction matters, and I’ve flagged it below in a dedicated correction box rather than quietly fixing it and pretending the error never existed.

That’s the standard for the rest of this article. Every statistic below traces to a named study with a DOI. Where a widely cited figure turned out to be softer than reported — sample sizes that shrank after exclusions, single-author studies presented as consensus, working papers that later became peer-reviewed journal articles — I’ve noted it rather than smoothing it over.

Quick answer — read this first
  • AI beats the average human on standardized divergent-thinking tests (the Alternate Uses Task and the Divergent Association Task) — this now holds across multiple independent studies, including one with over 100,000 human participants
  • The most creative humans still win, and the gap is widening, not closing. The 100,000-person study found the top half of human participants outperforms every AI model tested, and the top 10% pull further ahead still
  • AI helps weak performers more than strong ones. A Science Advances experiment on 293 analyzed writers found AI lifted below-average writers substantially; the benefit for already-strong writers was much smaller
  • AI collapses collective diversity even as it raises individual quality. The Wharton-authored study — now published in Nature Human Behaviour — found only 6% of AI-generated ideas were unique across sessions, versus 100% for humans
  • The “AI art bias is unconscious” claim is not well supported. The one eye-tracking study on this found explicit ratings dropped for AI-labeled art, but gaze and pupil measures showed no implicit difference at all
  • Structured human-AI collaboration outperforms either working alone — but the benefit differs by experience level: it mainly speeds up idea generation for novices and mainly improves refinement for experts

01 — OriginsWhere the “AI is creative” headlines come from

Three separate research programs, running independently between 2023 and January 2026, converged on a similar headline: AI chatbots now outperform the average human on standardized creativity benchmarks. That’s not spin — it’s a real, replicated finding. But “outperforms the average human on a benchmark” and “is more creative than humans” are different claims, and the gap between them is where most of the bad takes live.

256 humans vs. 3 chatbots, Scientific Reports (Koivisto & Grassini, 2024)
100K+ humans vs. multiple LLMs, Scientific Reports (Bellemare-Pépin et al., Jan 2026)
293 writers analyzed (of 500 recruited), Science Advances (Doshi & Hauser, 2024)
6% of AI ideas stayed unique across sessions, Nature Human Behaviour (Meincke, Nave & Terwiesch, 2025)

The earliest and most-cited of these is Mika Koivisto and Simone Grassini’s 2024 study in Scientific Reports, which compared 256 human participants against three chatbots (ChatGPT-3.5, ChatGPT-4, and Copy.Ai) on the Alternate Uses Task — the standard test where you generate as many uses as possible for an everyday object like a brick or a paperclip. n=256 humans + 3 chatbots, 11 sessions each; DOI: 10.1038/s41598-023-40858-3 On average, chatbot responses scored higher on both semantic distance (0.95 vs. 0.91) and human-rated creativity (2.91 vs. 2.47). But — and the paper is explicit about this — the human sample’s range was far wider: more low-quality ideas, but also higher peaks. The best human ideas still matched or beat the best AI ideas.

The much larger and more recent confirmation is the January 2026 study led by Professor Karim Jerbi at Université de Montréal, with co-first authors Antoine Bellemare-Pépin and François Lespinasse, in collaboration with researchers from Concordia, University of Toronto Mississauga, Mila, and Google DeepMind — including Yoshua Bengio as a co-author. Published in Scientific Reports, Jan 21, 2026; n=100,000 humans across five English-speaking countries, benchmarked against GPT-4, Gemini, Claude and others using the Divergent Association Task and creative-writing tasks Their finding: some LLMs, GPT-4 among them, now exceed average human performance on divergent linguistic creativity. But the average performance of just the top half of human participants exceeds every AI model tested — and the top 10% of humans open an even wider gap. As Jerbi put it, the result is “surprising — even unsettling” but not a story of AI overtaking human creativity outright.

I checked whether this reflects a moving goalpost — are these tasks getting easier for AI as models improve, or is human performance the constant? The tasks (AUT, DAT) haven’t changed since the 1960s. What’s changed is which models clear the average-human bar. That’s a real capability shift, not test inflation.


02 — The catchWhat AI is actually good at — and the catch nobody mentions

Speed and volume are the obvious wins, and they’re real enough that they don’t need a citation to justify — anyone who has used an AI tool for a first draft already knows this. The more interesting finding is who benefits and by how much.

Anil Doshi (UCL School of Management) and Oliver Hauser (University of Exeter Business School) ran a preregistered experiment recruiting 500 writers through Prolific; after dropouts, non-consent, and participants who used AI outside their assigned condition, 293 stories were analyzed. Science Advances, DOI: 10.1126/sciadv.adn5290, published 12 July 2024 Writers with access to generative AI ideas produced stories judged more novel and more useful — and the effect was strongest for writers who scored lower on an independent creativity measure going in. Access to five AI ideas raised “usefulness” scores by 9.0% relative to no-AI writers; access to just one AI idea still produced a measurable, statistically significant lift.

Here’s the part that gets dropped in most summaries: AI-assisted stories were more similar to each other than human-only stories were. Doshi and Hauser call this a social dilemma — individually better off, collectively narrower. That’s the same shape of finding as the Wharton diversity study covered in Section 5, discovered independently, in a completely different creative domain (fiction writing vs. product brainstorming). Two unrelated research teams landing on the same structural trade-off is stronger evidence than either result alone.

Second-order read — not stated directly in either paper

If AI tools lift weak performers more than strong ones, the market-level effect isn’t “AI democratizes creativity” in the way it’s usually framed. It’s closer to: the floor of average creative output rises, the ceiling doesn’t move, and the tools that produced that lift are the same tools narrowing the range of ideas being lifted. You get more competent-to-good work, converging on a smaller set of concepts, while the top tier — who use AI more sparingly and as a springboard rather than a source — becomes comparatively easier to distinguish, not harder.

There’s also a quieter finding worth flagging: AI is fluent but not self-critical. A single-author 2025 study by Joy Desdevises (OCTO Technology / Accenture) had ChatGPT-4o perform the classic “egg task” — a paradigm specifically designed to measure fixation bias, i.e., the tendency to stay anchored to conventional uses of an object rather than exploring novel ones. Frontiers in Psychology, DOI: 10.3389/fpsyg.2025.1628486; ChatGPT-4o compared against n=47 human participants plus aggregated data from 8 prior studies using the same task ChatGPT-4o was more prolific than the humans, but showed a comparable fixation bias — most of its ideas stayed inside conventional categories — and, notably, it struggled to tell its own original ideas apart from its conventional ones. Humans, imperfect as they are at generating novel ideas, are reliably better at recognizing which of their own ideas actually are novel. This is one study with a single named author and a modest human comparison sample, so I’m treating it as suggestive rather than settled — but it lines up with the fixation patterns other researchers have found using different methods.


03 — The ceilingWhere humans still win, decisively

Every study in this piece agrees on one point: peak human creativity has not been matched. Not “not yet, probably soon” — the 100,000-person study, the newest and largest dataset available, found the gap between top-tier humans and top-tier AI models is wider than the gap between average humans and average AI performance. If anything, the ceiling is separating rather than converging.

The mechanistic explanation researchers point to is that creativity isn’t produced by one system. It draws on the interplay between the brain’s executive control network, the default mode network associated with unstructured mind-wandering, and bursts of activity in regions linked to sudden insight. That’s a genuinely different computational process than sampling the next most probable token from a learned distribution — which is what an LLM is doing, however sophisticated the sampling has become. An LLM doesn’t wander toward an unrelated idea and then recognize it as useful; it doesn’t have the underlying process that produces the “aha” in the first place.

Human creativity
Divergent thinking (average) Below leading AI
Divergent thinking (top 10%) Ahead of all AI tested
Idea uniqueness across sessions 100%
Recognizes own original ideas Reliable
Audience-perceived authenticity Higher, explicitly
Output volume per hour Limited
vs
AI creativity
Divergent thinking (average) Beats avg. human
Divergent thinking (top 10%) Falls short
Idea uniqueness across sessions 6%
Recognizes own original ideas Weak, one study
Audience-perceived authenticity Lower, explicitly
Output volume per hour Extraordinary

04 — CorrectionThe bias claim that doesn’t hold up

Fact-check note — corrected during research for this piece

A claim circulates widely that people don’t just say they prefer human art over AI art — their eye movements supposedly prove an unconscious bias too. I traced this to a real, well-designed study and it does not say that.

Caitlin Cunningham, Gabriel Radvansky, and James Brockmole (University of Notre Dame) recorded eye movements from 92 participants viewing 24 pieces of DALL-E 2-generated religious art — a domain chosen specifically to maximize any bias against AI. Frontiers in Psychology, DOI: 10.3389/fpsyg.2025.1509974 Half the group was told the art was made by art students; the other half was told it was AI-generated. Result: participants’ stated opinions of quality and artistic merit were significantly more positive when they believed a human made the piece. But their gaze patterns — fixation counts, fixation duration, saccade amplitude, pupil size — showed no measurable difference between the two groups. The authors’ own conclusion is that they found no evidence the bias operates at an implicit, perceptual level; it appears to be an explicit, conscious judgment.

Widely repeated claim“Eye-tracking proves people unconsciously perceive AI art as different, even when they can’t say why.”
What the study foundGaze behavior was statistically indistinguishable between groups. The bias showed up only in people’s conscious, stated ratings — not in how they looked at the art.

This doesn’t erase the underlying finding — people really do rate AI-labeled art lower once they know the label, and that’s a real, replicated, practically important effect for anyone building a brand around authenticity. But it’s a conscious, belief-driven judgment, not evidence of some deeper perceptual “tell.” That distinction changes what you’d actually do about it: disclosure and framing decisions matter more than trying to make AI output “pass” some instinctive human detector, because the research so far hasn’t found one.


05 — The real findingThe diversity collapse: the finding that matters most

This is the study I keep coming back to, and it’s had an unusual path to credibility: it started as Wharton-affiliated working-paper coverage in 2025 and was subsequently published as a peer-reviewed article in Nature Human Behaviour — a substantially higher evidentiary bar than the “Wharton study” framing most coverage still uses.

Lennart Meincke (Mack Institute research fellow), Gideon Nave, and Christian Terwiesch (both Wharton professors) reanalyzed data from earlier ChatGPT brainstorming experiments, this time measuring not just idea quality but idea diversity across participants. Nature Human Behaviour, DOI: 10.1038/s41562-025-02173-x; reanalysis across five prior experiments, diversity measured via semantic similarity In 37 of 45 comparisons, ChatGPT-assisted ideas were significantly less diverse than ideas produced by unaided humans or humans using web search — a pattern that held regardless of which similarity-measurement method the researchers used.

The illustrative example: participants were asked to invent a toy using a fan and a brick. Among those using ChatGPT, 94% of ideas shared overlapping concepts — nine separate participants, working independently, named their invention “Build-a-Breeze Castle.” The human-only group produced no overlapping names at all.

“If you rely on ChatGPT as your only creative advisor, you’ll soon run out of ideas, because they’re too similar to each other.”

Christian Terwiesch, Co-Director, Wharton’s Mack Institute for Innovation Management

Terwiesch has been explicit that this isn’t a claim about representational diversity — it’s structural, closer to a biology metaphor: an ecosystem that loses genetic variety becomes fragile, not because any one organism is worse, but because there’s less range for the system to draw on when conditions change. Applied to brainstorming: individual AI-assisted ideas can be excellent, and the pool they come from can still be dangerously narrow.

Cross-source synthesis — this argument isn’t in any single cited paper

Put three independent findings together and a structural pattern emerges that none of the individual papers claims on its own. The Wharton/Nature Human Behaviour reanalysis shows AI-assisted ideation converges across sessions and across people. The Science Advances writing experiment shows the same convergence effect in a completely different creative domain — fiction — discovered by a different team using different methods. And the Notre Dame eye-tracking study shows audiences form real, if purely conscious, negative judgments once they know AI was involved. Combine these and the practical risk isn’t “AI-generated content is bad” — plenty of it tests well individually. It’s that markets, teams, and creative pipelines that lean on AI as the primary idea source are quietly converging toward a smaller set of concepts, at the exact moment audiences are primed to notice and discount output once they suspect that’s what happened.

Editorial synthesis — sources: Meincke, Nave & Terwiesch, Nat. Hum. Behav. 10.1038/s41562-025-02173-x · Doshi & Hauser, Sci. Adv. 10.1126/sciadv.adn5290 · Cunningham et al., Front. Psychol. 10.3389/fpsyg.2025.1509974

06 — Original frameworkThe Creative Substitution Matrix

Most “will AI replace creative work” arguments treat creativity as one thing. The research above suggests it isn’t — AI substitutes cleanly for some creative tasks and badly for others, and the difference isn’t about the subject matter, it’s about two independent variables: how much a task rewards recombining known patterns versus generating something genuinely outside them, and how much the audience’s judgment depends on knowing who or what made it. Mapping those two axes against each other gives a usable decision tool — one I haven’t seen published elsewhere in this exact form.

High conventionality · Low authorship-sensitivity

AI Substitution Zone

Ad copy variations, listicle drafts, product description templates, internal brainstorm volume. Audience doesn’t care who made it; task rewards fast recombination. This is where AI’s AUT/DAT advantage translates directly into value.

Low conventionality · Low authorship-sensitivity

AI-as-Springboard Zone

Early-stage concepting, “unstuck me” ideation, cross-domain pattern recombination. Nobody’s judging authorship yet — but the fixation-bias research suggests AI needs a human to select which of its outputs is genuinely novel, since it can’t reliably tell the difference itself.

High conventionality · High authorship-sensitivity

Draft-and-Sign Zone

Ghostwritten essays, routine brand copy under a personal byline, template-driven design work presented as bespoke. AI can do the heavy lifting, but the Notre Dame findings say disclosure and framing decisions carry real reputational weight here.

Low conventionality · High authorship-sensitivity

Human-Exclusive Zone

Signature art, memoir, brand-defining campaigns, anything sold partly on the story of who made it and why. This is where peak-human data (Section 3) and audience-authenticity data (Section 4) compound rather than cancel out — AI’s structural disadvantages stack here.

Horizontal axis: authorship-sensitivity (low → high, left to right) · Vertical axis: task conventionality (high → low, top to bottom)

The practical use of this matrix isn’t to sort your entire creative output once — it’s to sort individual briefs before you decide how much AI involvement is appropriate. A brief that starts in the AI Substitution Zone can drift into the Draft-and-Sign Zone the moment you decide to publish under a named author’s voice. That drift is usually where the trouble starts.


07 — Original toolThe Diversity Risk Checklist

The Wharton/Nature Human Behaviour finding is only actionable if you can tell whether it applies to your own pipeline. Here’s a simple three-factor scoring tool built directly from what that study — and the Science Advances replication of the same effect in a different domain — identified as the actual mechanism of convergence: shared prompts, shared tools, and unreviewed output.

  1. Prompt homogeneity — Score 0 if every contributor writes their own prompt from a unique angle; 1 if most start from a shared template with light edits; 2 if the whole team uses one prompt or one brief verbatim.
  2. Tool overlap — Score 0 if the team uses varied tools/models for ideation; 1 if most use the same model but different accounts/sessions; 2 if everyone shares one tool, one account, one conversation thread.
  3. Review diversity — Score 0 if a human actively curates and discards overlapping AI ideas before use; 1 if a single person skims for obvious duplicates; 2 if AI output goes to production with no dedicated novelty review.
0–2: Low risk 3–4: Moderate — add prompt variation 5–6: High — pipeline is likely converging without anyone noticing

A score of 5 or 6 doesn’t mean the individual outputs are bad — the research is clear that AI-assisted output often tests well piece by piece. It means that, in aggregate, your team is producing “Build-a-Breeze Castle” nine times over without realizing it, because nothing in the process is designed to catch that.


08 — Collaboration dataWhat actually works: human-AI collaboration

The most methodologically careful collaboration study I found is a 2025 two-phase experiment out of Hongik University — Nan Wang, Hyunsuk Kim, Junfeng Peng, and Jiayi Wang — which proposed and tested a structured Human-AI Co-Creative Design Process (HAI-CDP) against a traditional design workflow. Frontiers in Computer Science, DOI: 10.3389/fcomp.2025.1672735; IRB-approved, Hongik University No. 7002340-202506-HR-009-01; 2×2 factorial design (process type × experience level), plus semi-structured interviews

The structured collaboration process outperformed the traditional process on creative performance overall. But the value split by experience level in a way that matters for how you’d actually deploy this: for novice designers, the process’s main benefit was generating a wider spread of ideas faster — the same volume advantage this article has already documented. For experienced designers, the benefit showed up differently — in refinement and quality of already-strong concepts, not in idea generation itself. Experienced practitioners, in other words, don’t need AI to have ideas. They use it to make good ideas better, faster.

The workflow implied by the data isn’t complicated, and it maps directly onto the Substitution Matrix above: a human defines the brief and the constraint — the part requiring judgment about audience and purpose that AI has no reliable substitute for; AI generates volume and variation within that constraint; a human curates, using the kind of active novelty-discrimination the fixation-bias research found AI itself lacks; AI assists with refinement and iteration on the selected direction; a human does the final check, particularly for the emotional and cultural fit that the Notre Dame authenticity data says audiences notice.


09 — ApplicationWhat this means for your work

For: individual creative professionals

Compete on the ceiling, not the floor

The data is consistent across every study cited here: AI wins on average-case speed and volume; it has not closed the gap at the top, and the 100,000-person study suggests that gap is widening. What this means practically: stop trying to out-produce AI on raw output — that’s a contest you structurally can’t win. The differentiators the research keeps surfacing are judgment (recognizing which idea is actually novel, which AI is measurably bad at) and authorship — audiences rate work higher when they believe a human made it, even though that bias is conscious rather than instinctive, which means disclosure and story matter more than trying to disguise AI involvement.

For: teams and creative leads

Run the diversity checklist before you scale AI adoption

The Wharton/Nature Human Behaviour reanalysis and the Science Advances writing study found the same convergence risk independently, in different domains. What this means practically: if your team scores 5 or 6 on the checklist in Section 7, your creative pipeline is probably narrower than it looks from the outside, even if each individual output tests fine. Force divergence deliberately — different prompts, different tools, deliberately odd constraints — rather than assuming variety will emerge on its own from a shared tool.

For: brand and marketing teams

Assume your competitors are converging with you

If your category is briefing similar prompts into similar tools, the diversity-collapse research says you’re likely landing in the same conceptual neighborhood as competitors who did the same thing. What this means practically: proprietary context — brand history, specific customer language, constraints your AI tool wasn’t optimized around — is now a competitive input, not a nice-to-have. Treat prompt and constraint design as a defensible asset, because “we used AI” is no longer a differentiator when everyone else did too.


10 — Evidence tableFull evidence table

Finding Source & scope What it means ⚠ Limitation
AI beats average human on AUT divergent thinking Koivisto & Grassini, Scientific Reports 2024; n=256 humans + 3 chatbots; DOI 10.1038/s41598-023-40858-3 AI is a capable brainstorm partner on this benchmark AUT measures one narrow facet of creativity; three chatbots tested, all now outdated models
Top-10% humans exceed all AI models on divergent creativity Bellemare-Pépin, Lespinasse, Jerbi et al., Scientific Reports, Jan 2026; n=100,000 humans; multiple LLMs Peak human creativity gap is widening, not closing Human sample skews English-speaking, five countries only; LLM sample is a snapshot of models available at test time
AI lifts weak writers more than strong writers Doshi & Hauser, Science Advances 2024; 500 recruited, 293 analyzed after exclusions; DOI 10.1126/sciadv.adn5290 AI narrows the individual-quality gap between writers UK-based Prolific sample, not professional writers; single fiction-writing task
Only 6% of AI ideas unique vs. 100% human, across sessions Meincke, Nave & Terwiesch, Nature Human Behaviour 2025; DOI 10.1038/s41562-025-02173-x; reanalysis of 5 prior experiments AI use reduces collective idea diversity even as individual quality rises Specific to ChatGPT with repeated/independent prompting; reanalysis of existing datasets, not a fresh RCT
ChatGPT shows fixation bias, can’t reliably self-assess originality Desdevises, Frontiers in Psychology 2025; ChatGPT-4o vs. n=47 humans + 8 aggregated prior studies; DOI 10.3389/fpsyg.2025.1628486 AI needs human curation to separate novel from conventional output Single-author study; single task (egg task); aggregated benchmark data, not a matched control group
AI art rated lower — but only in explicit judgment, not gaze behavior Cunningham, Radvansky & Brockmole, Frontiers in Psychology 2025; n=92; eye-tracking; DOI 10.3389/fpsyg.2025.1509974 Authorship bias against AI art is real but conscious, not perceptual Religious-art domain specifically, chosen to maximize bias; may not generalize to all art categories
Structured human-AI collaboration beats either alone; benefit differs by experience Wang, Kim, Peng & Wang, Frontiers in Computer Science 2025; IRB-approved 2×2 design; DOI 10.3389/fcomp.2025.1672735 Collaboration model works, but deploy it differently for novices vs. experts Design domain specifically; findings may not transfer directly to writing, music, or fine art
All entries verified against primary sources with DOIs, checked directly rather than via secondary coverage, August 2026. Evidence strength varies: multi-thousand-participant peer-reviewed studies are stronger than single-author or small-n findings, which are marked accordingly. This is an active research field — expect revision as replications accumulate.

FAQQuestions people actually ask about this

Is AI more creative than humans now?

On standardized divergent-thinking benchmarks, AI now beats the average human — that’s a real, replicated finding across multiple independent studies through January 2026. But the most creative humans still outperform every AI model tested, and that gap appears to be widening rather than closing as models improve.

Does using ChatGPT for brainstorming actually hurt creativity?

Not at the individual level — individual AI-assisted ideas tend to score as good or better than unaided human ideas. The documented harm is at the group level: teams and individuals repeatedly prompting the same tool tend to converge on the same small set of concepts, which is a diversity problem, not a quality problem.

Can people actually tell if art was made by AI?

Not reliably through gaze or attention, according to the one eye-tracking study that tested this directly. But once people are told a piece was AI-generated, they rate it lower — a conscious, belief-driven effect rather than an instinctive detection ability.

Who benefits most from AI creative tools — beginners or experts?

Both, but differently. Weaker or less experienced creators see the largest raw quality lift. More experienced creators see less benefit to raw idea generation but more benefit to refining and polishing ideas they’ve already had, according to the Hongik University collaboration study.

How do I know if my team’s AI use is creating a diversity problem?

Run the three-factor Diversity Risk Checklist in Section 7 — score prompt homogeneity, tool overlap, and review diversity from 0–2 each. A combined score of 5 or 6 signals a pipeline likely converging on repeated concepts without anyone noticing, based on the mechanism identified in the Wharton/Nature Human Behaviour research.


GlossaryTerms used in this article

Alternate Uses Task (AUT)
A standard divergent-thinking test asking participants to generate as many uncommon uses as possible for an everyday object (e.g., a brick), scored on fluency, flexibility, originality, and elaboration.
Divergent Association Task (DAT)
A creativity benchmark that asks participants to produce ten words as semantically unrelated to each other as possible; wider semantic distance scores higher.
Fixation bias
The tendency to stay anchored to a conventional or first-available use of something rather than exploring less obvious, more original possibilities.
Collective diversity (idea diversity)
How different a pool of ideas from multiple contributors is from each other — distinct from individual idea quality, which measures how good any single idea is on its own.
Divergent vs. convergent thinking
Divergent thinking generates many possible ideas or solutions; convergent thinking narrows a set of options down to the single best one. Creativity research typically treats both as necessary but separate skills.

The honest summary: for average-case creative output, AI has already changed the baseline, and pretending otherwise doesn’t help anyone. For peak creative work — the kind built on genuine novelty, cultural fluency, and a story of authorship that audiences actually respond to — the current evidence, including the largest and newest study available, says the gap hasn’t closed. What should worry practitioners isn’t AI getting “too creative.” It’s the opposite: entire teams and markets converging on the same handful of AI-shaped ideas without anyone designing a way to notice it happening.