Large Language Models are producing increasingly similar content. A study published on arXiv on August 19, 2026 documents a statistically significant decline in output diversity over three years. The investigation by Nirav Patel, Josiah Crossman, Eva Aggarwal and Emily Wenger analyzed real user queries and standardized creativity tests using sentence-embedding similarity.
The research focuses on open-ended, creative tasks where originality, diversity and creativity are equally important as quality – an area not covered by existing benchmarks. The authors justify the relevance of their research question with the massive proliferation of LLMs: ChatGPT recorded over 700 million weekly active users in July 2025, sending more than 2.5 billion messages daily – approximately 29,000 messages per second. Anthropic and DeepSeek also have millions of users.
Creative tasks dominate usage
According to the study, approximately 70 percent of non-professional ChatGPT usage falls into practical instruction (including creative ideation) and writing. These two categories account for over half of all ChatGPT messages. Anthropic reports that nearly 20 percent of Claude messages fall into categories such as content creation or multidisciplinary academic research and writing.
The researchers analyzed LLM outputs with two datasets: Infinity-Chat100, a real collection of open user queries, and the Alternate Uses Task, an established psychometric creativity assessment instrument. As their analysis method, they used sentence-embedding similarity to examine trends in LLM responses across different model versions.
Output diversity declines statistically significantly
The study shows a statistically significant decline in output diversity over time. This suggests that LLM outputs are converging in their creative substance – different models are increasingly producing similar content. The research confirms earlier findings: scientific papers show that LLM outputs can appear individually creative but are collectively homogeneous – resembling either other creative content from the same LLM or content from other LLMs.
A study from January 2025 titled "We're Different, We're the Same: Creative Homogeneity Across LLMs" addressed similar issues. An analysis from September 2025 discussed the regression-to-the-mean phenomenon, in which LLMs tend toward safer, more generic phrasing. An empirical comparative study between human and ChatGPT writing showed in September 2025 the homogenizing effect of LLMs on creative diversity. Another analysis from November 2025 questioned whether LLM creativity has reached its peak.
Warning of constrained human creativity
The authors warn of potential negative consequences: LLM-driven homogenization could progressively reduce human agency in creative collaboration between humans and AI. The convergence of LLM outputs could restrict the diversity of ideas users encounter. This could influence downstream human creativity.
The authors call for careful review of the role of LLMs in the human creative process. The study is referred to as a "preliminary analysis," suggesting limited scope. Further research is needed to clarify whether this trend continues, how different usage scenarios are affected by LLM homogenization, and how users are impacted in their own creativity by more limited output diversity.
