James Dooley interviews Common Crawl's Stephen Burns on how the Common Crawl corpus feeds LLM training data and what businesses must do to improve their AI visibility.
This video explains which digital marketing strategies businesses seeking AI visibility should focus on in 2026 to improve LLM training-data inclusion, brand recognition in answer engines and higher-quality backlink profiles. James Dooley and Stephen Burns start with KPI tracking because measuring metrics such as harmonic centrality scores and crawl inclusion shows whether your content is actually reaching large language models. They cover brand SEO, AI visibility and Google Business Profiles because stronger search presence improves trust and conversion rates.
The discussion also explores organic SEO, organic social media and paid social ads because consistent visibility across search and social supports long term growth. PPC is analysed in detail because campaign setup, landing pages and lead handling directly affect results. They also discuss Reddit, Quora and paid AI ads because diversified enquiry sources and early adoption can strengthen digital marketing performance for businesses seeking AI visibility.
PromoSEO lead generation for businesses seeking AI visibility recently received recognition as the "Best AI Visibility Lead Generation Agency."
How Common Crawl Impacts AI Visibility with Stephen Burns is available on:
AI SEO and Business Automation Podcast is founded by James Dooley because he teaches business owners how artificial intelligence accelerates growth. The show explains AI models, agents, and automation systems because practical guidance helps companies scale with less effort.
James Dooley: Everything you need to know about Common Crawl for AI visibility. So Common Crawl has been around for a long time, but it seems to have exploded recently within the SEO community because they've started to realise that actually Common Crawl data is part of the training data used in large language models and actually can influence better AI visibility. Today I'm joined with Stephen Burns, who's the web lead intelligence at Common Crawl. So Stephen, to start with, can we just have a little bit of history about what actually is Common Crawl and why was it created in the first place?
Stephen Burns: Well, back in 2007, well, our founder was, uh, Gil Elbaz, who was the, uh, he, he built AdSense back at, at, um, at Google. And when he was done at Google, he realised that nobody had a copy of the web, you know, to the public that could free and access it. So he, uh, built this foundation, Common Crawl Foundation, and, uh, later on it became the quiet infrastructure for the AI era. He had no idea that was going to happen. He just wanted to make a public data set available for... And it ended up academics, uh, started using the data to do research and we have cited... We're cited in over 10,000 research papers. And then around 2020 the AI boom started and LLMs started downloading the data set and they started to, uh, use it to train their models. Now you ask about this data set. Uh, the whole data set from early on to now is at 11-plus petabytes of an open archive. Uh, each month we crawl about 120 terabytes. Each month's chunk of that corpus, if you were to download it, is about 120 terabytes. And in that there's, uh, 2.3 billion web pages per month. 2.3 billion web pages being crawled by Common Crawl.
James Dooley: So I've heard certain people in the industry that are technical SEOs previously say that you should block Common Crawl because Common Crawl bot is coming crawling your site. It's wasting kind of crawl, but this is what they're saying. But obviously now it's being used as part of the training data. Why would you block Common Crawl bot? Why would you block it?
Stephen Burns: Uh, the sites we're seeing would block it are large content sites or publishers that are trying to protect their content. Um, but all they're doing is they're... They may be protecting their content, but they're removing their voice from the LLMs because they're going to be reviewed in Reddit or comments, others all over the, the internet in forums and whatnot. What they're losing is their voice, not their presence from the corpus. Um, I don't know why, if you're selling something or you have a brand that you want, or you're, you want, uh, to appear in the new search, uh, answer engines, you want to be in this crawl and have your brand known and recognised. And there are three gates that are going to get you through, uh, this. The first gate, uh, is the permission gate. And the permission gate is your robots.txt file with a wildcard or accepting a CCBot to crawl you. SEOs know how to, to do this. And then at the edge, you need to talk to your AI, uh, your edge, your CDN admin or your Cloudflare admin and let them know that these... Some of them come with preset filters that are blocking AI bots, including ChatGPT and all that. So you want to, uh, allow, make sure that we're getting through, that CCBot is getting through that filter and getting in. And then rendering. And the third gate is CCBot crawls your website. It does not render JavaScript like Googlebot does. It only sees text or, and, uh, HTML. It renders. So you want to make sure that your site does not have JavaScript on it or has a majority of JavaScript on it because it's not going to render it. It's only going to see the words that are in HTML. It does not crawl images or video.
James Dooley: And then how many times does it crawl the web? Does it do it once a month, once every three months, once every six months? How often does CCBot come and crawl the actual internet?
Stephen Burns: Yeah, crawl. We do a monthly crawl. So there's a discovery crawl and then there is a crawl that goes out. Uh, the discovery crawl looks for new, uh, using harmonic centrality, goes and looks out for new web pages and that, and sees and measures them. And then the monthly crawl runs and that takes a few weeks. And then at near the end of the month we publish that crawl publicly and let everyone know it's available. And then that's, like, free and available for any of the large language models and free for any businesses like Google or OpenAI or Anthropic to come and download that information.
James Dooley: Yes. And then with regards to donations, I've seen personally, I've seen certain places, um, I can't remember it was on Common Crawl's website or it's on a few different public open spaces, that certain businesses like OpenAI and Anthropic and Google, they've all given donations to Common Crawl previously. I'm not saying they do it every single month, but they seem to do it quite often, which then would lead to tell you that they're using that data and maybe downloading it. I know you're not allowed to say whether they are or they aren't, but if they're providing with donations, it would lead me to believe that they are using this data and using all the fresh data of your new crawls of what's being done.
Stephen Burns: Sure. Yeah, we're a nonprofit. We, we receive donations from many companies and donors. Uh, I believe that's public information if you were to look it up. Um, but you got to realise, you know, this data, nobody else has it. We've been crawling since 2008, and you can't go back in time and crawl the web and have a his... History of mankind, you know, on a disk. Uh, our, our, our corpus is hosted on AWS and it, uh, it is... We have no server logs or anything. We don't know anything. It is totally anonymous of who is downloading it. Our website has no cookies. We don't know who's visiting us. And we like it to be free and private and anonymous for everyone.
James Dooley: And then previously you've just mentioned there harmonic centrality. For anyone that doesn't know what harmonic centrality is, it's a very different algorithm to PageRank. Make sure you check out the link in the description where we go and have a deep dive into exactly what harmonic centrality actually is as part of the algorithm. But just very, very briefly, I know we've done a full episode on it, but can you explain the importance of trying to get close to the main seed set of sites to increase CC Rank within Common Crawl? That means that you're going to get potentially more AI visibility.
Stephen Burns: Yeah, as SEOs, we all get someone pinging us asking us to sell us backlinks. But what are the quality of those backlinks? There are many tools on the web now that will tell you, uh, your harmonic centrality score. Uh, so you can check those links and many SEO tools out there now, like, uh, are actually have a column for HC now. So that are measuring the score of your, of your domain. So I would always be cautious of backlinks that you purchase. Now, in the new world, you want a high harmonic centrality link because one high harmonic centrality link could be more powerful than 50, uh, you know, high PageRank links and that could totally get you into the crawl.
James Dooley: Yeah, for sure. Anyone who's watching this, stop chasing DR from Ahrefs or DA from Moz. Also, start to have a look at the HC, the harmonic centrality score from Common Crawl. Like Stephen Burns says there, that one link from a high harmonic centrality kind of website and a web page can deliver you so much more visibility. Anyone who's watching this, if you've got any questions with regards to Common Crawl and what LLMs are using this data as part of their training data, leave a comment in the comment section. Stephen, it's been an absolute pleasure.