Full transcript
Location pages dropping out of the index
0:00Google reached into the index, found
0:01pages that had been sitting there
0:02ranking for months, and pulled them back
0:04out again, flipped them to crawled,
0:07currently not indexed. About a third of
0:09the location pages on one client site
0:11gone in a couple of weeks, and nothing
0:13on those pages had changes. So, when I
0:15went looking at what sites this was
0:17hitting, there was a clear pattern to
0:19this. A site with two or three of these
0:21pages was fine. Nothing happened. A site
0:23with 30 or more, they were losing most
0:26of them. Now, one system generated all
0:29of them to the same spec at the same
0:30quality. The only variable was how many
0:33each site had, which means Google is
0:36grading these pages against each other
0:38on the same domain. And somewhere in
0:40there is a point where the next page you
0:42publish stops helping you and starts
0:45costing you. Now, I found where that
0:46point is and what it takes to get back
0:49across it. It took a human copywriter,
0:52yes they do still exist, three pages and
0:54about 11 hours of research. Now, first
0:57the timing because this got reported
0:59badly everywhere. There was no April
1:01core update. Google confirmed to this
1:03year, March 27th through April 8th and
1:06May 21st through June 2nd. What hit us
1:08landed in the gap between in early May.
1:11And it wasn't only my client's sites.
1:14Starting in late April, site owners
1:16everywhere were reporting the same
1:18thing. pages [snorts] that have been
1:19indexed for years. No manual action, no
Timing vs the March and May core updates
1:22crawl error, nothing broken, suddenly
1:24flipping into crawled, currently not
1:26indexed, and staying there. John Miller
1:29got asked about this directly on April
1:3130th. His answer was that super helpful.
1:35Some sites go up, some sites go down,
1:37and he didn't see anything exceptional.
1:39Now, if you were watching your pages
1:41disappear that week, that landed about
1:44as well as you'd expect it to. Now,
1:46here's the part that not a lot of people
1:48understand, but it changes how you read
1:51your search console. Crawled, currently
1:53not indexed, has nothing to do with
1:55ranking. It happens in a different part
1:58of Google's system before ranking is
2:01really in the picture at all. A core
2:03update moves around pages that are
2:05already in the index. This is Google
2:07looking at a page and deciding it isn't
2:11worth the storage. Storage costs money.
2:13Google crawls far more of the internet
2:15than it ever intends to keep. And every
2:18page it keeps is a page it pays to hold.
2:21So a page in that bucket got read and
2:25turned down later. This is a harsher
2:28judgment than a ranking drop. A ranking
2:30drop says you're not the best answer
2:32today. This it says the page doesn't
2:35need to exist. No user of Google
2:39anywhere on the planet ever is going to
2:41find value in the content on that page.
2:43Therefore, Google won't index it.
2:45Miller's been clear about what drives
2:47it. He said if their systems are
2:49seriously worried about the quality of
2:50the website, they crawl it less and
2:52index less. And this status is what that
2:55looks like from your side. Gary Illis
2:57said the same thing from the other
2:59direction. That general site quality
What “crawled, currently not indexed” actually means
3:01matters a lot for how many of these you
3:03see. So the status is a readout of
3:05something Google already decided. And
3:08this left us with location pages Google
3:10had accepted years ago and then rejected
3:13again. Nothing had changed on our end.
3:15So something changed in what Google was
3:17willing to hold on to. And the pages
3:19closest to the line were the ones that
3:22fell off it. The only way to find out
3:23which line was to take a few of them and
3:26change them. So we pulled three of the
3:28drop pages and I handed them to my human
3:30copywriter. We didn't touch anything
3:32else. We didn't source links. We didn't
3:34do schema. We didn't do anything
3:35technical. Same URL, same template.
3:36Everything was kept the same. The only
3:38variable was the content, the words on
3:41the page. He spent about three and a
3:43half hours on each one of them. And what
3:46he did with that time is the whole
3:48answer of what Google is looking for.
3:50So, let me walk through. Okay. He
3:52started in the search console and in the
3:54Google business profile performance
3:55insights looking for what Google already
3:58associated that page with. looking for
4:00what search queries that page was
4:02already indexed. He wasn't after the
4:05keywords that he liked or that he
4:06thought it should be indexed for. He
4:08wanted the queries the page was already
4:10surfacing for, even weekly, even ranked
4:13in position 95, because that tells you
4:15what Google is already thinking the page
4:18is about. You're adding onto something
4:20instead of arguing with Google's
4:21algorithm. Don't argue with machines. It
4:24rarely works out. Then he went through
4:26the client's call transcripts, actual
4:28recorded calls from actual customers,
4:30looking for what people say out loud
4:32when they call the business owner. This
4:34is different from how they type. And
4:36this exists nowhere else on the
4:38internet. Then he went to Reddit, not
4:40the topic in general, but that specific
4:43neighborhood, what people who live there
4:45complain about and warn each other
4:46about. This client is in Chicago, so
4:48there's a lot to work with. And then the
4:51piece that mattered the most. He took
4:53these pages and turned them into a
4:55story. An actual throughine connected to
4:57the business, its history, and why
4:59someone in that part of Chicago ends up
5:02calling their office. Most location
5:04pages are a pile of generic neighborhood
5:07facts with the service bolted on to the
5:08end. This was not what he wrote. Three
Rewriting three pages and what changed
5:11pages, 11 hours of work. All three got
5:15indexed within a week. Now, three pages
5:17is an anecdote. And yes, I'll say that
5:19plainly. We try to use real data in
5:22statistically significant samples before
5:24we make any conclusions. But it told us
5:26the failure was likely in the words and
5:28we finally had a description of what had
5:31been missing. So we took what the human
5:33did and rebuilt our content generation
5:36process around it. The research order,
5:38where the data comes from, how the
5:39outline gets built before anything gets
5:41ever written. Across every site we run
5:43it on, we now see location pages
5:45indexing at about 97%.
5:48They were somewhere between 60 and 70%
5:51before this change with Google's update.
5:53And that range is wide for a reason,
5:56which brings me back to what I said at
5:57the start of this about volume. The 60
6:00to 70 was never a flat rate. It depends
6:02almost entirely on how many of these
6:04pages a site already has. A site with
6:08two or three location pages, yeah, they
6:10indexed almost every time. A site with
6:1230, most of them get rejected because
6:15they're built the same way to the same
6:17spec. the count was the only difference.
6:19And once you understand what Google is
6:22actually deciding, it starts to make
6:23sense. Miller framed this in a way that
6:26I haven't been able to stop thinking
6:27about. He said, "The question isn't
6:29whether the page is good. Google's
6:31algorithm is very bad at deciding
6:34whether content is quality or not. The
6:37question is what is missing from the
6:39index if this page isn't in it." a
6:42solidly written article that already
6:44exists in a similar form 30 times over
6:47doesn't say anything new. And he took it
6:49further and said about AI content
6:51directly that you read some of these
6:53pages and you think, boy, anyone could
6:55have written that. That's how AI writers
6:57typically work. They produce the average
6:59of what already exists. So page one
7:02clears easily. Nothing else in the index
7:04covers that service in that neighborhood
7:06from that business. Something is missing
7:08without it. But by the time you get to
7:09page 30, you've answered that question
7:1129 other times, each page is competing
7:14against your own previous 29. And that
7:17they're all built the same way with the
7:19same depth. The honest answer to what's
Volume, sameness, and the missing-from-the-index test
7:20missing without this 30th article is not
7:24very much. And that's the threshold. It
7:26isn't a number Google publishes, and
7:28it's likely different for every site,
7:30for different niches and different
7:31areas. But the mechanism is the same
7:33everywhere. Every page you publish
7:35raises the bar for the next one because
7:37the next one has to be worth keeping in
7:39a world where your previous pages
7:41already exist. So volume on its own is
7:43fine. It turns on you when the pages
7:45stop getting more specific as you add
7:48them. All right, I want to show you what
7:49this looks like in the search console
7:51because I want to be careful about the
7:52difference between the number moving and
7:54the number meaning something. So this is
7:56the coverage report on the client site
7:58for crawled currently not indexed. This
8:00is the category where Google looked at
8:02the content on the page and decided it
8:04never needed to show it to anyone ever
8:06in the world. Okay. Now, this is a real
8:10business. So, 277 pages sitting in
8:13crawled not indexed. But we can see this
8:15peaked at 363 at the end of June. Now,
8:20watch how it moves through steps rather
8:21than slopes. One step down here,
8:25one step down here. This is Google
8:27reprocessing, rec crawling this website
8:29in batches, which is what you'd expect
8:31when a group of pages gets rewritten and
8:33recrolled all at once. Okay, now I want
8:35to be honest about these 277 pages. Most
8:38of them are not location pages. Most of
8:40this 277, something like 87% of them are
8:45machine translated Spanish pages, plus a
8:47bunch of duplicate junk, tracking
8:49parameters, HTTP versions of HTTPS,
8:51trailing/variants, stuff like this. None
8:54of this was ever really meant to be
8:55indexed. So, I don't really count it.
8:56The number that I actually watch is
8:58nine. Right now, there are nine location
9:01pages de-indexed. That's compared to 89
9:04that were de-indexed at the end of June.
9:06And two of those nine are duplicates,
9:08which we should fix rather than real
Reading the coverage report without lying to yourself
9:10failures. So, this is seven out of 89
9:13that are not indexed. That's 90%. This
9:16is below the 97% average, but I'm
9:18showing you this below average one on
9:20purpose because honestly, anyone can
9:21screenshot their best client. Now, those
9:23seven were last crawled between July 2nd
9:26and July 31st, and the rewrite that we
9:28published went live after that. This
9:30means Google hasn't seen the new version
9:32yet, and I expect them to be indexed
9:35when Google does. If they don't, I'll
9:37record another video and let you know
9:38why. But there's a second number on this
9:40report that almost matters more, and
9:43very few people check it. Let me show
9:45you. This one is discovered, currently
9:48not indexed. Okay, five pages on a site
9:51with hundreds of URLs, hundreds of
9:53pages. This is what tells me that Google
9:55sees this as a healthy domain because
9:58those two statuses answer different
9:59questions. Crawl, not indexed, means
10:01Google came, it read, and passed. A
10:05judgment about the page. A judgment
10:06about the quality. Discovered, not
10:08indexed, means Google knows the URL
10:10exists, but hasn't even bothered to go
10:13look at it. Okay, that one's about you.
10:16Google's crawling documentation explains
10:18why. They call it crawl demand, and it
10:21reflects how interested Google is in
10:22your pages, including how good it thinks
10:26they're going to be. Low demand means
10:28less crawling, no matter how fast your
10:30server is, and their own guidance.
10:32There's a lot of low-v value URLs will
10:35hurt a site's crawling and indexing. So,
10:38crawling first in that sentence, right?
10:40So, there's a progression available
10:42here. A template keeps producing pages
10:44that Google won't keep. that results in
10:46Google's interest in that pattern
10:48dropping and eventually the new one stop
10:50getting crawled at all. Now, Google
10:53hasn't laid this out as a formal
10:55sequence, so call it my read, my
10:56educated guess, but this lines up with
10:58everything they published about crawl
Discovered vs crawled, and crawl demand
11:00demand. So, on this site, the number is
11:02five. Google is still interested.
11:03They're crawling almost every new page
11:05we publish. It's just turning down these
11:07specific five pages. Okay, these pages
11:10have a problem. The site doesn't. So,
11:12what do you actually do with a page
11:14that's in this bucket? Because there's a
11:16way to drive this number to zero that
11:18really fixes nothing at all. And it's
11:20the first move a lot of people reach for
11:22the magical no index tag. Page shows up,
11:26it's on one of these reports, you don't
11:27want it, throw a no index tag on it, and
11:29it drops off the report. Number goes
11:31down, dashboard looks clean, and you
11:33have changed absolutely nothing. Those
11:35pages are already not indexed. That's
11:38what the status means. The tag moves
11:40them from one bucket in your report to a
11:42different bucket in your report while
11:44Google's opinion of your site sits
11:45exactly where it was. You've
11:47disconnected the gauge from the engine
11:49and then congratulated yourself on the
11:52reading. Okay. Now, forcing this also
11:54doesn't work. Resubmitting, pointing
11:56internal links at it, pinging the site
11:58map. Miller's been direct that this
12:00isn't a technical problem to be fixed.
12:03The question Google asked was, "What's
12:06missing from the index without this
12:08page?" and submitting it again doesn't
12:11change the answer. You have two real
12:13options here. Number one, improve it so
12:16it carries something that could only
12:18have come from this business that adds
Why noindex and resubmits don’t fix this
12:21meaningful information to the search
12:23results or delete the page. Deleting
12:26this took me a little while to accept.
12:28Google's own guidance on unhelpful
12:30content says removing it can help the
12:31rankings of your other content because
12:33this signal is sitewide. their words,
12:36content on a site with a lot of
12:38unhelpful content is less likely to do
12:40well, even the good content. So, a page
12:43that was never going to earn its place
12:44is costing you by existing, and taking
12:47it down is a repair of your site. Lily
12:50Ray did a study of 220 sites that scaled
12:53AI content, and it shows this from the
12:54other end. The brands that cut their
12:56content footprints are the ones where
12:58the traffic started recovering. And if a
13:00page still won't index after you've done
13:02the work, if after you've done the
13:03improvement, then the problem moved.
13:05It's sitting in whatever produced it.
13:07One page failing is a page problem. Six
13:10pages failing from the same generator is
13:13the generator's problem. So you should
13:15count rate per template instead of pages
13:17per site. Every page came out of some
Improve it, delete it, or fix the template
13:20process, a prompt, a build sheet, a
13:22writer following a brief. Maybe maybe
13:24you still use human writers like I do.
13:27Group your unindexed pages by what made
13:29them, how they were created, and watch
13:31that rate. The level barely matters. The
13:33slope does. A template that's steady at
13:368% is noise. But if that template jumps
13:38to 40% after you change something,
13:40that's Google telling you that change
13:42made it worse. And it's telling you
13:44before your rankings move. This is the
13:47whole value. It's early. Rankings lag.
13:49And by the time a position drops, you've
13:52already published another 100 pages off
13:54of a broken template. Let me close the
13:55loop on something because I called this
13:57a leading indicator and I want to tell
14:00you about that. Every site in Ray's
14:02study, as far as we know, they indexed
14:04fine. All 220 indexed rank, pulled
14:06traffic for 6 to 12 months, then got
14:08demoted. Indexation never caught it. So,
14:10if you walk out of here thinking that
14:12your coverage report is a safety net,
14:14you're you've got it wrong. And I'd
14:16rather say that than let you find out
14:18the hard way. What this catches is thin
Why indexation is not a safety net
14:21and redundant at the door before you've
14:23built anything on top of it. Cheap,
14:26real, more than most people are
14:28checking. It's a floor. What actually
14:30kills site is slower and it doesn't
14:33always show up in search console until
14:35the traffic's already gone. So, this is
14:37why the standard I hold my own clients
14:40to lives outside the report entirely.
14:42Could a competitor publish a near
14:44identical version of this page tomorrow
14:46using the same methodology? If they
14:49could, the page doesn't need to exist.
14:51And Google will work it out eventually
14:53whether or not it indexed today. Now,
14:55everything I've covered here happens
14:57after Google reads your page. There's a
15:00whole set of problems that can happen
15:02before that where Google can't read the
15:04site properly in the first place. And
15:07those are the ones that make an AI built
15:09website invisible from day one. That's
15:12this video on the screen. Go check that
15:14one out