TL;DR
- The JPL managed-fleet indexation census, run 11 September 2026, inspected all 575 URLs our seven Search Console properties submit for indexing, one URL Inspection call each.
- Of the 549 pages actually offered to Google, 77.6 percent are indexed. That is the number for every page type we submit.
- There is a figure circulating in SEO content that only about 60 percent of the blog posts you publish stay indexed, stated with no study behind it. We refused to repeat it and measured instead. On its own denominator it turns out to be close: 57.5 percent of our blog posts are indexed.
- The claim is still wrong in the way it gets used, because it is repeated as a fact about publishing in general. Everything that is not a blog post indexed at 81.4 percent.
- The gap is the finding. Blog posts were crawled and declined at 19.5 percent. Everything else, at 2.6 percent. Seven times the rate, same fleet, same day.
- The worst property in the fleet is ours, not a client’s: our AI and workflow automation site sits at 50 percent indexed, and 14 of its 15 declined pages are posts written by named AI agents under a human editor. We are publishing that because an agency that measures this should be willing to be measured by it.
Where does the “only 60 percent stays indexed” claim come from?
Nowhere you can check. It circulates in SEO content marketing stated flat, with no study attached, no date range, no sample, no method. We ran into it in a training video whose every on-screen number was the creator’s own unaudited screenshot.
It is directionally consistent with something real. Google has tightened index selection, and practitioners do see pages crawled and passed over. But a round number with nothing behind it is not a statistic, and repeating it to a client is not analysis. So we did not repeat it. We measured our own.
We can do that because we hold Search Console access across the sites we build and maintain. That is the only real advantage here: not a better opinion about indexation, an instrument pointed at a fleet.
How do you measure indexation across a whole fleet?
For each property, we took the URLs the site itself submits through its sitemap. That is the honest denominator: pages the owner has declared they want in the index. Then we asked Google, one URL at a time through the Search Console URL Inspection API, what it had done with each one. That returns a coverage state per URL, the same verdict you see when you paste a URL into Search Console by hand, just at fleet scale.
Two adjustments before any percentage, both in the direction of being harder on ourselves rather than easier:
- Intentional
noindexpages come out of the denominator. Twenty-three pages carry a deliberatenoindex. Counting them as indexation failures would inflate the problem with pages nobody wanted indexed. - Dead URLs come out too. Three sitemap entries return 404. Those are a sitemap hygiene defect, worth fixing and worth listing separately, but they are not Google declining a page. They are us advertising a page that no longer exists.
That leaves 549 pages genuinely offered for indexing. So it is reproducible: noindex status and every coverage state were read from what the API returned, not from our own view of the pages; page types were assigned by URL path; the submitted set came from each property’s sitemap index and its child sitemaps, deduplicated. Google publishes the quota, 2,000 inspection queries per property per day and 600 per minute, so this is a run anyone with Search Console access can repeat on their own sites. Every number below is against the 549, measured on a single day. No sampling, no extrapolation, no modelling.
What percentage of submitted pages does Google actually index?
| Coverage state | Pages | Share of offered |
|---|---|---|
| Submitted and indexed | 426 | 77.6 percent |
| Not indexed, all causes | 123 | 22.4 percent |
| Crawled - currently not indexed | 29 | 5.3 percent |
| Discovered - currently not indexed | 76 | 13.8 percent |
| URL is unknown to Google | 17 | 3.1 percent |
| Duplicate, Google chose different canonical than user | 1 | 0.2 percent |
77.6 percent indexed, across every page type we submit. That is the fleet number, and on its own it is the wrong number to compare against the claim we started with.
Was the 60 percent claim right?
On its own terms, close to it. The claim is about blog posts specifically, so we typed all 549 offered pages by URL path and ran the same split.
| Page type | Offered | Indexed | Indexed % | Crawled and declined |
|---|---|---|---|---|
| Blog posts | 87 | 50 | 57.5 percent | 17 (19.5 percent) |
| Everything else | 462 | 376 | 81.4 percent | 12 (2.6 percent) |
Our blog posts are 57.5 percent indexed. The unsourced figure said about 60 percent. We refused to repeat that number because nobody had measured it, and then our own measurement landed within three points of it. That is worth saying plainly rather than burying, and it is the part of this exercise we did not expect.
What is wrong with the claim is not the arithmetic, it is the scope it gets used at. It circulates as a fact about publishing, and it is repeated to owners as though every page they put up has a four-in-ten chance of vanishing. On our fleet that is false. Everything that is not a blog post indexed at 81.4 percent. The failure is concentrated in one page type, and the useful sentence is about the gap rather than the average.
Blog posts were crawled and declined at 19.5 percent. Everything else, at 2.6 percent. Same fleet, same crawler, same day, seven times the rate. That is the finding, and it is the one the borrowed number obscures by turning a page-type problem into a general dread.
Why is my page not indexed? Three states, three different fixes
Of the 549 pages our seven properties offer Google, 123 are not indexed. Search Console reports those as separate coverage states, and treating them as one number is how people end up fixing the wrong thing.
Crawled - currently not indexed. Google fetched the page, read it, and chose not to index it. This is the closest thing to a quality verdict Google will give you for free. It is 29 pages, 5.3 percent of the 549, and as the split above shows it is overwhelmingly a blog-post state.
Discovered - currently not indexed. Google knows the URL exists and has not fetched it. No verdict has been rendered. This is the largest slice at 76 pages, 13.8 percent of the 549. From practice rather than from this census, the lever on this state is internal linking, sitemap discipline, and giving the crawler a reason to come back, not prose. We have not measured that here either, and we will say so when we have.
URL is unknown to Google. Seventeen pages, 3.1 percent of the 549. The URL sits in a sitemap and Google has no record of it. On a new property that is mostly time. On an established one it is a discovery defect worth chasing. These states also move on their own: while we were writing this, three URLs the census had recorded as unknown had already become discovered by the same afternoon.
Collapse those three into one percentage and you get an audit finding that tells a client to rewrite 123 pages. Keep them apart and the actual instruction is: rewrite 29, fix the internal link graph for 76, and investigate why 17 are invisible.
What kind of pages did Google crawl and decline?
The 29 crawled-and-declined pages are not randomly distributed, and the pattern is uncomfortable if your content plan is built on volume.
- Seventeen are blog posts. Articles published to build topical coverage, read by Google, and left out of the index.
- Nine are near-duplicate service pages, mostly the same service rewritten per city on one client site. The city variants were declined while the parent service pages were indexed. Google does not say why, but the most economical reading is that it already has this page.
- One is an author archive that should never have been in a sitemap at all. A content-management default, submitted for indexing, declined. That is not a content problem, it is a configuration leak.
- Two are other pages: one vertical landing page and one transparency page.
Of the 29 pages Google crawled and declined, none were core service pages. Read that back as an instruction and it is unkind to the standard advice: the pages that pay indexed fine, and the pages published to support them are where index selection bit.
One practitioner note worth the space. Those nine near-duplicate city pages landed in crawled-and-declined, not in the duplicate-canonical state. Across the entire census exactly one page returned “Duplicate, Google chose different canonical than user.” If you are auditing your own city pages, do not go looking for them under the duplicate state. They will be sitting in the same bucket as your thin blog posts.
What does this look like on our own sites, including the bad one?
A fleet average hides the spread, and the spread is the interesting part. Here is every property. Client sites are unnamed; ours are named, because they are ours to expose.
| Property | Offered | Indexed | Indexed % | Crawled and declined |
|---|---|---|---|---|
| couvreursverifies.ca (our directory) | 155 | 134 | 86.5 percent | 0 |
| Client site A (manufacturer) | 20 | 18 | 90.0 percent | 0 |
| Client site B (services) | 144 | 116 | 80.6 percent | 0 |
| Client site C (services) | 75 | 52 | 69.3 percent | 10 |
| jpldigital.ca | 72 | 61 | 84.7 percent | 2 |
| Our FR sub-brand site | 23 | 15 | 65.2 percent | 2 |
| Our AI and workflow automation site | 60 | 30 | 50.0 percent | 15 |
Four of those seven are ours, including the best row and the worst one. Three are client sites and stay anonymous. The best is our roofing directory at 86.5 percent with nothing declined, and most of the explanation is that it is about seven weeks old with a clean, purpose-built URL set.
The worst is also ours, at 50 percent. Two things explain it and only one is comfortable. It is a young site publishing into a category where Google already holds abundant coverage, which is the condition under which index selection gets strict. It is also the property whose blog is written by named AI research agents under a human editor, disclosed on the site itself, and 14 of its 15 declined pages are those posts. We are not going to report the first reason and leave out the second.
Hold that finding at the size it actually is. It is one property, on one day, 15 pages, and it is confounded with site age and with category saturation. It is not a result about AI-written content in general and we are not going to dress it up as one. It is a real thing that happened on our own site, and it is the reason the next census matters more than this one.
Publishing the number is the point. If we ran this audit for you and showed you a 69 percent, you are entitled to ask what ours looks like.
What an indexation audit should check
Most audits you will be sold return a score. A score is not actionable. What you want out of this part of an audit is five things, and you can ask for them by name.
- Coverage state for every submitted URL, not a sample. Per URL, from Search Console, not inferred from a crawler. A third-party crawler tells you what your site says; only Google tells you what Google did.
- A denominator you agree with. Intentional
noindexout, dead URLs out and listed separately as a hygiene defect. If nobody shows you the denominator, the percentage means nothing. - The numbers split by page type. A single site-wide rate would have hidden the entire finding above. Blog posts, service pages and archives fail differently and a blended average tells you nothing about which.
- The three not-indexed states kept apart, because they route to three different pieces of work.
- The declined list itself, page by page. The pattern in that list is the finding. Ours said blog volume and city duplicates. Yours will say something else.
If you want to check one thing yourself this afternoon, open Search Console, paste your three most recently published pages into the URL inspection box, and read the coverage state. It takes two minutes, and it is the same instrument we just pointed at 575 URLs. When we run it across a whole property it is part of an audit sprint, and the declined list is the deliverable, not the score.
Does indexation affect AI visibility?
It is upstream of it. An answer engine that grounds its response in search cannot cite a page that is not in the index, so a page sitting in crawled-and-declined is invisible to both surfaces at once. We have written before about how AI systems choose what to cite and about the shared preference that concentrates those citations on a narrow set of sources. The layer above this one is what a model recalls when it never looks anything up, which no indexation work reaches. This is the floor under the case where it does look. Being in the index does not earn you a citation. Being out of it forecloses one.
What this measurement is not
Scope, plainly, because the whole point of the exercise was refusing to repeat an unsourced number:
- This is our fleet, not the web. Seven properties, 549 pages, Quebec SMB and JPL-owned sites. It is a real measurement of a real population, and that population is small and ours. Do not read 77.6 percent, or 57.5 percent, as an industry benchmark.
- Indexed is a floor, not an outcome. A page can be indexed and take no impressions at all. This census says nothing about traffic, rankings or revenue, and 100 percent indexed is not a goal worth chasing on its own.
- It is one day. Coverage states move, as the three that moved while we wrote demonstrate. A page in discovered-not-indexed today can be indexed next month without anyone touching it. We have not yet run this twice, so we cannot tell you the rate of change.
- Three sites are missing, and their absence probably flatters the number. Three client sites sit behind host-level bot protection that answers automated requests with a challenge instead of their sitemap, so we could not assemble their submitted URL list without going around a security control, which we will not do. Those are among the oldest libraries in the fleet. Our expectation from practice, not from this census, is that older libraries carry more declined pages, which would put the true figure below 77.6 percent rather than above it. We will know when we can measure them.
- Coverage state is a verdict, not an explanation. Google tells you it declined a page. It does not tell you why. The patterns above are our reading of the declined list, not Google’s stated reasons.
We will re-run this census and publish the movement. That second number is the one that will actually tell you whether rewriting declined pages works, and we would rather wait for it than assert it now.
