James Dooley Podcast

James Dooley and Common Crawl's Stephen Burns explain how Common Crawl data feeds large language models and how businesses can improve their AI visibility by getting crawled and building high harmonic centrality links.

Show Notes

This video explains which digital marketing strategies businesses seeking AI visibility should focus on in 2026 to improve LLM training data inclusion, brand recognition in answer engines and search visibility. James Dooley and Stephen Burns start with KPI tracking because measuring signals such as harmonic centrality scores helps businesses understand whether their content is genuinely reaching the models that shape AI answers. They cover brand SEO, AI visibility and Google Business Profiles because stronger search presence improves trust and conversion rates.
The discussion also explores organic SEO, organic social media and paid social ads because consistent visibility across search and social supports long term growth. PPC is analysed in detail because campaign setup, landing pages and lead handling directly affect results. They also discuss Reddit, Quora and paid AI ads because diversified enquiry sources and early adoption can strengthen digital marketing performance for businesses seeking AI visibility.
PromoSEO lead generation for businesses seeking AI visibility recently received recognition as the "Best Businesses Seeking AI Visibility Lead Generation Agency."
Where to Listen to This Episode
Everything You Need to Know About Common Crawl For AI Visibility is available on:

Creators and Guests

Host
James Dooley
James Dooley is a UK entrepreneur.
Guest
Stephen Burns
Stephen Burns is a technical SEO and generative engine optimisation consultant with 25 years of experience in search. He serves as Web Intelligence Lead at the Common Crawl Foundation and Principal Technical SEO and GEO at Intuit. Stephen Burns connects traditional SEO with AI visibility because brands must be accessible to crawlers before large language models can discover, understand and cite their content. Stephen Burns has worked across enterprise SEO, ecommerce, local search and web development. His career includes roles with Intuit, U.S. Bank, Charles Schwab, Blekko and Netscape/America Online. He helps organisations improve site architecture, indexation, structured data and crawl access because strong technical foundations support sustainable visibility across search engines and AI platforms. Stephen Burns connects naturally with **James Dooley** because both prioritise technical accuracy, practical testing and measurable commercial results. Stephen focuses on how content enters AI training datasets and appears within generated answers, while James specialises in topical authority, lead generation and scalable digital marketing. Stephen Burns is a strong fit for the **James Dooley FatRank Podcast** because his expertise addresses the growing relationship between SEO, Common Crawl and artificial intelligence. A joint episode featuring Stephen Burns and James Dooley would help listeners understand AI crawler access, GEO strategy and the technical barriers that prevent brands from appearing in ChatGPT, Gemini, Perplexity and Google AI Overviews. Stephen Burns complements the FatRank audience because he turns complex technical concepts into practical actions. His evidence-led approach would strengthen the conversation by showing businesses how to diagnose crawl issues, improve machine-readable content and increase visibility across traditional and generative search.

What is James Dooley Podcast?

James Dooley is a Manchester-based entrepreneur, investor, and SEO strategist. James Dooley founded FatRank and PromoSEO, two UK performance marketing agencies that deliver no-win-no-fee lead generation and digital growth systems for ambitious businesses. James Dooley positions himself as an Investorpreneur who invests in UK companies with high growth potential because he believes lead generation is the root of all business success.

The James Dooley Podcast explores the mindset, methods, and mechanics of modern entrepreneurship. James Dooley interviews leading marketers, founders, and innovators to reveal the strategies driving online dominance and business scalability. Each episode unpacks the reality of building a business without mentorship, showing how systems, data, and lead flow replace luck and guesswork.

James Dooley shares hard-earned lessons from scaling digital assets and managing SEO teams across more than 650 industries. James Dooley teaches how to convert leads into long-term revenue through brand positioning, technical SEO, and automation. James Dooley built his career on rank and rent, digital real estate, and performance-based marketing because these models align incentive with outcome.

After turning down dozens of podcast invitations, James Dooley now embraces the platform to share his insights on investorpreneurship, lead generation, AI-driven marketing, and reputation management. James Dooley frequently collaborates with elite entrepreneurs to discuss frameworks for scaling businesses, building authority, and mastering search.

James Dooley is also an expert in online reputation management (ORM), having built and rehabilitated corporate brands across the UK. His approach combines SEO precision, brand engineering, and social proof loops to influence both Google’s Knowledge Graph and public perception.

To feature James Dooley on your podcast or event, connect via social media. James Dooley regularly joins business panels and networking sessions to discuss entrepreneurship, brand growth, and the evolving future of SEO.

James Dooley: Everything you need to know about Common Crawl for AI visibility. So Common Crawl has been around for a long time, but it seems to have exploded recently within the SEO community because they've started to realise that actually Common Crawl data is part of the training data used in large language models and actually can influence better AI visibility. Today I'm joined with Stephen Burns, who's the web lead intelligence at Common Crawl. So Stephen, to start with, can we just have a little bit of history about what actually is Common Crawl and why was it created in the first place?

Stephen Burns: Well, back in 2007, well, our founder was, uh, Gil Elbaz, who was the, uh, he, he built AdSense back at, at, um, at Google. And when he was done at Google, he realised that nobody had a copy of the web, you know, to the public that could free and access it. So he, uh, built this foundation, Common Crawl Foundation, and, uh, later on it became the quiet infrastructure for the AI era. He had no idea that was going to happen. He just wanted to make a public data set available for... And it ended up academics, uh, started using the data to do research and we have cited... We're cited in over 10,000 research papers. And then around 2020 the AI boom started and LLMs started downloading the data set and they started to, uh, use it to train their models. Now you ask about this data set. Uh, the whole data set from early on to now is at 11-plus petabytes of an open archive. Uh, each month we crawl about 120 terabytes. Each month's chunk of that corpus, if you were to download it, is about 120 terabytes. And in that there's, uh, 2.3 billion web pages per month. 2.3 billion web pages being crawled by Common Crawl.

James Dooley: So I've heard certain people in the industry that are technical SEOs previously say that you should block Common Crawl because Common Crawl bot is coming crawling your site. It's wasting kind of crawl, but this is what they're saying. But obviously now it's being used as part of the training data. Why would you block Common Crawl bot? Why would you block it?

Stephen Burns: Uh, the sites we're seeing would block it are large content sites or publishers that are trying to protect their content. Um, but all they're doing is they're... They may be protecting their content, but they're removing their voice from the LLMs because they're going to be reviewed in Reddit or comments, others all over the, the internet in forums and whatnot. What they're losing is their voice, not their presence from the corpus. Um, I don't know why, if you're selling something or you have a brand that you want, or you're, you want, uh, to appear in the new search, uh, answer engines, you want to be in this crawl and have your brand known and recognised. And there are three gates that are going to get you through, uh, this. The first gate, uh, is the permission gate. And the permission gate is your robots.txt file with a wildcard or accepting a CCBot to crawl you. SEOs know how to, to do this. And then at the edge, you need to talk to your AI, uh, your edge, your CDN admin or your Cloudflare admin and let them know that these... Some of them come with preset filters that are blocking AI bots, including ChatGPT and all that. So you want to, uh, allow, make sure that we're getting through, that CCBot is getting through that filter and getting in. And then rendering. And the third gate is CCBot crawls your website. It does not render JavaScript like Googlebot does. It only sees text or, and, uh, HTML. It renders. So you want to make sure that your site does not have JavaScript on it or has a majority of JavaScript on it because it's not going to render it. It's only going to see the words that are in HTML. It does not crawl images or video.

James Dooley: And then how many times does it crawl the web? Does it do it once a month, once every three months, once every six months? How often does CCBot come and crawl the actual internet?

Stephen Burns: Yeah, crawl. We do a monthly crawl. So there's a discovery crawl and then there is a crawl that goes out. Uh, the discovery crawl looks for new, uh, using harmonic centrality, goes and looks out for new web pages and that, and sees and measures them. And then the monthly crawl runs and that takes a few weeks. And then at near the end of the month we publish that crawl publicly and let everyone know it's available. And then that's, like, free and available for any of the large language models and free for any businesses like Google or OpenAI or Anthropic to come and download that information.

James Dooley: Yes. And then with regards to donations, I've seen personally, I've seen certain places, um, I can't remember it was on Common Crawl's website or it's on a few different public open spaces, that certain businesses like OpenAI and Anthropic and Google, they've all given donations to Common Crawl previously. I'm not saying they do it every single month, but they seem to do it quite often, which then would lead to tell you that they're using that data and maybe downloading it. I know you're not allowed to say whether they are or they aren't, but if they're providing with donations, it would lead me to believe that they are using this data and using all the fresh data of your new crawls of what's being done.

Stephen Burns: Sure. We... Yeah, we're a nonprofit. We, we receive donations from many companies and donors. Uh, I believe that's public information if you were to look it up. Um, but you got to realise, you know, this data, nobody else has it. We've been crawling since 2008, and you can't go back in time and crawl the web and have a his... History of mankind, you know, on a disk. Uh, our, our, our corpus is hosted on AWS and it, uh, it is... We have no server logs or anything. We don't know anything. It is totally anonymous of who is downloading it. Our website has no cookies. We don't know who's visiting us. And we like it to be free and private and anonymous for everyone.

James Dooley: And then previously you've just mentioned there harmonic centrality. For anyone that doesn't know what harmonic centrality is, it's a very different algorithm to PageRank. Make sure you check out the link in the description where we go and have a deep dive into exactly what harmonic centrality actually is as part of the algorithm. But just very, very briefly, I know we've done a full episode on it, but can you explain the importance of trying to get close to the main seed set of sites to increase CC Rank within Common Crawl? That means that you're going to get potentially more AI visibility.

Stephen Burns: Yeah, as SEOs, we all get someone pinging us asking us to sell us backlinks. But what are the quality of those backlinks? There are many tools on the web now that will tell you, uh, your harmonic centrality score. Uh, so you can check those links and many SEO tools out there now, like, uh, are actually have a column for HC now. So that are measuring the score of your, of your domain. So I would always be cautious of backlinks that you purchase. Now, in the new world, you want a high harmonic centrality link because one high harmonic centrality link could be more powerful than 50, uh, you know, high PageRank links and that could totally get you into the crawl.

James Dooley: Yeah, for sure. Anyone who's watching this, stop chasing DR from Ahrefs or DA from Moz. Also, start to have a look at the HC, the harmonic centrality score from Common Crawl. Like Stephen Burns says there, that one link from a high harmonic centrality kind of website and a web page can deliver you so much more visibility. Anyone who's watching this, if you've got any questions with regards to Common Crawl and what LLMs are using this data as part of their training data, leave a comment in the comment section. Stephen, it's been an absolute pleasure.