Why Generative AI Often Deepens the Innovation Bottlenecks It Promises to Remove

Innovation teams now share near-identical foundation models and prompt libraries, but results remain uneven. Some report more creative output; others see homogenised ideas that could have come from a single writer. New research from academics at Harvard, Wharton, Northwestern and Columbia, in a working paper invited by the International Journal of Research in Marketing, argues the difference is not the model but how AI interacts with long-standing human bottlenecks.

The authors examine four stages of the innovation pipeline: ideation, screening, consumer insight, and post-launch market learning. For each stage, they ask what human constraint slows the work, how generative AI affects that constraint when used without deliberate redesign, and what a more careful process would look like.

The central finding is that AI's default effect is often to deepen the bottleneck, not dissolve it. Models trained on aggregate human output reproduce the same anchoring, bias and rationalisation that caused the problem, only faster and at scale. The paper identifies one stage—post-launch feedback synthesis—where AI genuinely shrinks a bottleneck, but even that creates a new risk of selective, more persuasive confirmation bias.

The implication for leaders is less about choosing a better model and more about diagnosing the mechanism behind each constraint before adopting a tool.

Advertisement

How AI Deepens Ideation, Screening, Consumer Insight and Market Learning

Ideation: Convergence on the statistically typical

When an LLM is used without structure, it produces the most statistically probable answer and then anchors the human reader to that option. Because many companies use similar training data, independently creative teams can converge on the same design directions, reducing differentiation even as individual output rises. The paper's specific, somewhat counterintuitive fix is to aim the intervention at the model rather than the person: chain-of-thought prompting that asks the model to revise toward bolder, more distinct territory works better than telling people to think more broadly, which can actually increase fixation.

Screening: Fluency gets mistaken for quality

Stage-gate panels already underweight novel ideas because they feel riskier. Generative AI worsens this because AI-generated pitches are fluent and polished by design, and evaluators mistake that fluency for substance. The structural fix is to strip submissions to a common format and conceal whether an idea came from a human or a model. The paper also cites a field experiment in which evaluators given an LLM recommendation plus a written rationale became more likely to rubber-stamp rejections without becoming better judges. Removing the rationale improved decisions, especially in borderline cases—the opposite of what many companies are building.

Consumer insight: Simulated customers understate the frictions that kill adoption

AI-simulated consumer segments can produce demand curves that are not just noisy but systematically wrong. A model treats a higher price as a signal of premium quality rather than as the same product, just more expensive, so it can generate a flat or upward-sloping curve where real shoppers would buy less. Digital twins built from rich behavioural profiles behave more rationally than actual humans and therefore underweight switching costs, sunk costs and framing effects. The research recommends using simulated personas only as a filter for low-stakes questions, and keeping lead-user panels and ethnographic fieldwork for decisions where those frictions decide the launch.

Diffusion and market learning: A more convincing echo chamber

After launch, AI's ability to ingest and cluster reviews, tickets, forums and social chatter is a genuine bottleneck reduction: no analyst team can match the scale. The risk is that frequency-weighted summaries bury the weak signals that matter most, such as an emerging use case or a segment silently failing to adopt. Worse, an AI can be queried in ways that telegraph the answer a team prefers, and its fluent, confident synthesis becomes a citable voice for decisions already made. The open question is whether AI nets out to more honest reporting or merely a more persuasive rationalisation.

Advertisement

A Three-Question Diagnostic for Innovation Leaders Adopting AI

For innovation leaders, the paper's value is a diagnostic question rather than a checklist of use cases: what is actually slowing this stage—an information problem, a judgment problem, or an incentive problem?

  • For ideation: Use chain-of-thought prompting that directs the model to revise toward bolder, more distinct concepts; do not rely on exhorting teams to think more broadly after they have seen the AI's first answer.
  • For screening: Strip all submissions to a common format and hide human versus AI origin before scoring. In the field experiment cited, a black box AI recommendation without a written rationale improved evaluator decisions; consider withholding LLM rationales, especially for borderline cases.
  • For consumer insight: Treat AI-simulated consumers as a fast, low-cost filter, not as a demand forecast. Keep lead-user panels and ethnographic observation for anything that depends on switching costs, sunk costs or how choices are framed.
  • For post-launch review: Decide which feedback themes warrant action before AI synthesis is shared, and phrase AI queries neutrally so the system does not merely confirm a preferred conclusion.

Risk & Opportunity Assessment

Commercial RiskMediumTeams that treat simulated customer responses as demand forecasts risk launching products that ignore switching costs and sunk-cost behavior, the frictions the paper says kill adoption.
Competitive RiskHighCompanies using similar foundation models for ideation can converge on the same handful of design directions, eroding differentiation across a market.
Regulatory RiskLowThe paper does not identify a direct regulatory or compliance constraint; the bottlenecks described are cognitive, social and organizational.
Reputation RiskMediumAI-generated fluent pitches and favorable synthetic feedback can make flawed concepts look validated, increasing the reputational cost when launches underperform.
Technology DisruptionHighThe research argues that bottlenecks rooted in lived experience, irrational behavior or organizational self-interest cannot be fixed by model scaling, challenging overconfident AI adoption.
Commercial OpportunityMediumDeliberate redesigns—chain-of-thought prompting, blind screening, removing LLM rationales, and using AI for high-volume feedback synthesis—can reduce genuine innovation frictions.