What is Reddit After All? And Why Does It Matter for Search Visibility?
Ask ChatGPT which running shoe to buy, ask Perplexity which CRM is worth the money, ask Google’s AI Overview whether a piece of software is any good, and there is a strong chance the answer you receive was assembled, at least in part, from anonymous strangers arguing on a forum.
However, Reddit is not an encyclopaedia nor a publisher. It is a pile of pseudonymous threads, moderated by unpaid volunteers, governed by norms rather than editorial standards. Yet, Google paid roughly sixty million dollars a year for gaining access to it, OpenAI signed its own agreement, and Reddit’s chief legal officer now describes the platform as the most commonly cited source in AI-generated answers.
How did this happen? To answer this question and understand how Reddit grew to become this ultra-important source for LLMs, we got to step far, far back (no, not the Stone Age!)
How it all started
Reddit was founded in June 2005 by Steve Huffman and Alexis Ohanian, both freshly graduated from the University of Virginia. They had pitched Paul Graham a mobile food-ordering app, been rejected, and been invited back with a different idea.
The founding premise was simple: build a site where readers decide what appears on the front page. Submissions are voted up or down by the community. Comments arrived in December 2005. A merger with Aaron Swartz’s Infogami followed shortly after, bringing a rewritten codebase and a third co-founder into the mix.
Two structural choices from that period explain almost everything that came later.
Pseudonymity brings honesty
Reddit accounts are detached from real identity. That has obvious costs, and Reddit has spent much of its history paying them. It also has a specific benefit: people say things under a throwaway username that they would never put on LinkedIn. Honest reports of side effects, salary numbers, product failures, regret over a purchase. This is exactly the class of information that marketing content systematically suppresses.
Subreddits as delegated governance
One of the best things Reddit did was to not try to moderate its use centrally. It handed communities their own rules, their own moderators and their own culture. Over 100,000 active communities now operate this way. The result is a corpus that is not one undifferentiated forum but thousands of small, self-policing knowledge bases, each with its own tolerance for self-promotion, low-effort posts and undisclosed advertising.
Ownership changed repeatedly. Condé Nast acquired Reddit on 31 October 2006 for a reported ten million dollars, then spun it out in September 2011 as an independent subsidiary of parent company Advance Publications (Seattle Times, 2024). Reddit listed on the New York Stock Exchange in March 2024. Through all of it, the underlying mechanics stayed intact.
Data that became the product
At some point, Reddit realised that it was sitting on top of a goldmine.
On 18 April 2023, Reddit announced it would begin charging for API access, a resource that had been free since 2008. The stated rate was $0.24 per 1,000 API calls, effective 1 July. Christian Selig, developer of the popular third-party client Apollo, calculated that this would cost him roughly twenty million dollars a year, and announced Apollo would shut down (Variety, 2023).
The response was the largest coordinated protest in the platform’s history. Between 12 and 14 June 2023, more than 8,000 subreddits went private or read-only, including some of the largest communities on the site. Tens of thousands of moderators participated. Reddit itself went down for around three hours under the load (TechCrunch, 2023; Ars Technica, 2023).
The justification Huffman offered was simple. Reddit needed to be self-sustaining and could no longer subsidise commercial entities requiring large-scale data use. The reference to language models trained on Reddit data was explicit. The API fight was not really about third-party apps. It was about establishing that Reddit’s corpus was a licensable asset rather than a public utility.
Two events in 2024 confirmed the thesis. In February, on the same day it filed for its IPO, Reddit announced a content licensing agreement with Google reported at roughly sixty million dollars annually, giving Google real-time structured access to Reddit’s Data API and allowing Reddit content to be displayed more prominently across Google products (NBC News/AP, 2024). In May, Reddit and OpenAI announced a comparable partnership bringing Reddit content into ChatGPT, with OpenAI also becoming an advertising partner (OpenAI, 2024). Reddit’s shares jumped sharply on the news.
Why AI companies need Reddit
A model with the entire indexed web available to it still has to decide what to retrieve and what to cite, and the pattern is consistent enough across independent studies that commercial arrangements alone cannot account for it. These are the four properties that influence it the most:
Format matches retrieval
A Reddit thread is natively a question followed by ranked answers. Retrieval-augmented systems chunk documents and pull passages. A comment that reads “we moved from Tool A to Tool B eighteen months ago, here is what broke and what improved” is a self-contained, extractable passage. A brand landing page rarely is.
Experience is scarce
Language models are trained on an internet in which most commercial content is written by people who have never used the product they are describing. First-hand experience with very specific accounts of using product has become immensely valuable and it’s what most users seek when using Reddit
Community validation is a quality signal
Upvote counts, comment depth, awards and the presence of corrections give a retrieval system something to weight beyond raw text. Where a thread contains an error, there is usually a reply saying so.
Recency
Licensed API access means Google and OpenAI receive Reddit content close to real time. In late 2023, as part of its core updates, Google rolled out a series of ranking changes the industry came to call “hidden gems”, designed to surface first-hand knowledge and personal insight from forums, comments and small blogs (Search Engine Roundtable, 2023).
The effect was dramatic. Glenn Gabe’s analysis found Reddit’s search visibility up 378 percent by Semrush data and 978 percent by Sistrix data following the change, and, crucially, that this was not a Reddit-specific favour: of 97 forums he tracked across verticals, 88 percent saw year-on-year visibility growth above 100 percent (GSQi, 2024).
That is the important point for anyone reading the current AI citation data. Reddit’s prominence in AI answers is the continuation of a search-quality trend that predates AI answers, not a break from it.
What the citation data shows
The headline claim is easy to state and easy to overstate. Reddit is, by most measures, the single most-cited domain in AI-generated answers.
An analysis of 30 million directly cited sources by Peec AI, reported by Search Engine Land in March 2026, ranked Reddit first across ChatGPT, Google AI Mode, Gemini, Perplexity and AI Overviews, ahead of YouTube, LinkedIn, Wikipedia and Forbes. Semrush’s July 2025 study, covering 5,000 randomly selected keywords and more than 150,000 unique citations across four AI search platforms, found Reddit appeared as a leading citation source on every platform tested. Secondary coverage of that dataset put Reddit’s reference share at 40.1 percent, ahead of Wikipedia at 26.3 percent and YouTube at 23.5 percent.
Anyone using these numbers professionally should treat them with more care than they usually receive. Three caveats matter.
Denominators differ wildly
“Share of all citations across millions of domains” and “frequency of appearing at all in a given answer” produce numbers that differ by an order of magnitude from the same underlying data. A figure of 40 percent and a figure of 2 percent can both be accurate descriptions of the same platform.
Engine behaviour is not uniform
Ahrefs’ comparison of AI Mode and AI Overviews found Reddit cited at broadly similar rates in both, while Wikipedia appeared in 28.9 percent of AI Mode citations against 18.1 percent for AI Overviews, and Quora appeared 3.5 times more often in AI Mode (Ahrefs, 2025). Perplexity leans on Reddit far more heavily than ChatGPT does. Aggregate cross-engine figures obscure this.
Volatility is measured in weeks
Semrush’s 13-week study of more than 230,000 prompts found ChatGPT citing Reddit in close to 60 percent of prompt responses in early August 2025, collapsing to around 10 percent by mid-September, with no equivalent drop on AI Mode or Perplexity (Semrush, 2025). This means that the exact figures move constantly.
Why large-scale, low-value manipulation is going nowhere
Everything that makes Reddit valuable to AI systems is also what makes that value unstable, such as:
Manipulation risk
If community consensus is a quality signal, then it’s only natural that it becomes commercially attractive. Over four months in 2024 and 2025, researchers at the University of Zurich ran an experiment on a subreddit, deploying AI-generated personas including a claimed sexual assault survivor and a claimed trauma counsellor, posting 1,783 comments to test whether models could shift opinions. Moderators, who found out only after it had ended, described it as psychological manipulation. Reddit’s chief legal officer Ben Lee called it “deeply wrong on both a moral and legal level” and Reddit banned the associated accounts (Engadget, 2025; Retraction Watch, 2025).
If a small research team could do that undetected for four months, the implications for coordinated commercial use were obvious. Both Reddit’s enforcement systems and the FTC’s endorsement rules push against it, and the reputational downside for a brand caught doing it is severe, but the incentive is there.
The data itself
Reddit sued Anthropic in June 2025 over alleged unauthorised scraping, then sued Perplexity together with data intermediaries SerpApi, Oxylabs and AWM Proxy in October 2025, alleging an operation that accessed nearly three billion search results pages in a single two-week period. Lee described the surrounding market as an industrial-scale “data laundering” economy (Forbes, 2025; Reuters, 2025). Perplexity characterised the suit as a negotiating tactic. These cases will help define whether publicly visible user content is free to read but not free to take.
Reddit no longer wants to be just a supplier
This is the shift most SEO and GEO practitioners are underweighting. Reddit is building the destination itself. Weekly active users of Reddit’s search grew from 60 million to 80 million during 2025, while Reddit Answers, its AI-powered question interface, went from 1 million weekly users in Q1 2025 to 15 million by Q4. Huffman has said the company intends to be an end-to-end search destination (Search Engine Land, 2026; TechCrunch, 2026). Roughly 40 percent of conversations on the platform are commercial in nature, and Reddit says 84 percent of shoppers feel more confident about purchases after researching there.
The referral relationship is deteriorating
In Q2 2026 Reddit reported revenue of 805 million dollars, up 61 percent year on year, and 130.3 million daily active uniques. It also reported that US daily actives slipped sequentially, with Huffman describing search referrals as “choppy” and noting that the shift toward AI Overviews had not yet become a net positive (CNBC, 2026; Reddit Q2 2026 earnings call). In July 2026, the Wall Street Journal reported Reddit was weighing whether to renew the Google licensing agreement at all.
The relationship is therefore not stable. Reddit supplies the data that powers answer engines which reduce the traffic Reddit depends on to generate more data. Huffman’s framing on the Q1 2026 call was blunt: there is no artificial intelligence without actual intelligence, and it comes from Reddit.
What to do then?
What AI systems reward is specific, experiential and verifiable content with external signal of validation. Reddit happens to be the largest reservoir of it. The same properties can be built elsewhere: in expert-authored content with named authorship and real methodology, in earned coverage carrying specific quotable claims, in customer-facing documentation that answers a question directly rather than positioning around it.
Where Reddit itself is concerned, the durable approach is genuine participation in communities where the brand has something useful to contribute. Inauthentic mass tactics do not work, mostly because they carry escalating detection risk, violate FTC endorsement rules, and degrade the exact signal that makes the channel valuable in the first place.
Reddit is the clearest available evidence that what other people say about a brand, in places the brand does not control, now carries more weight in machine-generated answers than what the brand says about itself. It is a channel to be nurtured and influenced, but not gamed.
Sources
- Sequoia Capital, “Reddit ft. Steve Huffman: The Making (and Remaking) of the Front Page of the Internet”, Crucible Moments — https://sequoiacap.com/podcast/crucible-moments-reddit
- Y Combinator, Reddit company profile — https://ycombinator.com/companies/reddit
- The Seattle Times, “Condé Nast’s owners set to reap $1.4 billion windfall from Reddit”, March 2024 — https://www.seattletimes.com/business/conde-nasts-owners-set-to-reap-1-4-billion-windfall-from-reddit/
- Variety, “Reddit Protest: Subreddits Go Dark in Backlash Over API Pricing Move”, June 2023 — https://variety.com/2023/digital/news/reddit-blackout-dark-protest-api-charge-third-party-apps-1235640741/
- TechCrunch, “Thousands of subreddits go dark to protest Reddit’s API pricing”, June 2023 — https://techcrunch.com/2023/06/12/reddit-blackout-8000-subreddits-went-dark-protest-api/
- Trust and Safety Foundation, “The Reddit Blackout of 2023” — https://www.trustandsafetyfoundation.org/blog/blog/the-reddit-blackout-of-2023-moderators-lead-the-charge-for-a-site-wide-protest-of-api-changes
- NBC News / Associated Press, “Reddit strikes $60M deal allowing Google to train AI models on its posts, unveils IPO plans”, February 2024 — https://www.nbcnews.com/tech/tech-news/reddit-strikes-60m-deal-allowing-google-train-ai-models-posts-unveils-rcna140168
- OpenAI, “OpenAI and Reddit Partnership”, May 2024 — https://openai.com/index/openai-and-reddit-partnership/
- Search Engine Roundtable, “Google Search Hidden Gems Ranking Algorithm Rolling Out”, November 2023 — https://www.seroundtable.com/google-search-hidden-gems-ranking-algorithm-36388.html
- GSQi (Glenn Gabe), “Beyond Reddit and Quora: How Google’s Hidden Gems Update Yielded Explosive Growth In Search Visibility For Many Forums”, 2024 — https://www.gsqi.com/marketing-blog/beyond-reddit-and-quora-google-hidden-gems-update-forums-surge/
- Semrush, “How Google’s AI Mode Compares to Traditional Search and Other LLMs”, July 2025 — https://www.semrush.com/blog/ai-mode-comparison-study/
- Semrush, “The Most-Cited Domains in AI: A 3-Month Study”, November 2025 — https://www.semrush.com/blog/most-cited-domains-ai/
- Search Engine Land, “AI search engines cite Reddit, YouTube, and LinkedIn most: Study”, March 2026 — https://searchengineland.com/ai-search-engines-cite-reddit-youtube-and-linkedin-most-study-473138
- TechCrunch, “Reddit looks to AI search as its next big opportunity”, February 2026 — https://techcrunch.com/2026/02/05/reddit-looks-to-ai-search-as-its-next-big-opportunity/
- TechCrunch, “People are finally using Reddit’s search”, May 2026 — https://techcrunch.com/2026/05/01/people-are-finally-using-reddits-search/
- CNBC, “Reddit (RDDT) Q2 2026 earnings report”, July 2026 — https://www.cnbc.com/2026/07/30/reddit-rddt-q2-2026-earnings-report.html
- CNBC, “Reddit stock sinks on report it may not renew Google AI content deal”, July 2026 — https://www.cnbc.com/2026/07/22/reddit-stock-google-ai-content-deal.html














