Fly Away Paul

HomeGuides

Site Health, Plainly: What a Walk of Your Site Really Checks — and the One Button Only You Can Press

"Site health" gets sold as a score out of a hundred, and the score usually means less than it appears to. This is what the phrase actually covers, what a proper walk of your site can and cannot see, what Google's own console knows that no outside tool ever will — and the single click in this whole trade that no software can do for you.

Every page you publish is read twice. A person reads it the way you wrote it, and a machine reads it the way a filing clerk would — following each link to see where it lands, checking the title against the titles of your other pages, looking for the one line that says what the page is about and the markup that says who and where your business is. Site health is the state of that second reading. When it goes wrong, nothing looks broken to you: the page renders, the photos load, the visitor you watch over their shoulder has a perfectly fine time. The clerk, meanwhile, walked into a door that wouldn't open, filed two of your pages under the same name, and left.

The practical questions behind the phrase come in four sizes, and it's worth having them straight before anyone shows you a score. Is each page working — does it answer when knocked on, does it have a title, does anything link into it? Is each page built properly — headings in order, images described, a stated address for the page that machines can rely on? Does the whole site hang together as a body of work rather than a pile of pages? And does it say, unambiguously, what kind of business this is and where it stands? A "health check" that only asks the first question can hand out perfect marks to a site failing the other three.

A health score is only as good as the list of things it checked — and that list ought to be thorough and clear if it is to be of any help.

Why the scores mislead

We learned this while building Fly Away Paul's own Site Health hub, which is why we're comfortable saying it bluntly. Two of the first sites our walk ever examined scored 100% — and the reason was that our early checks had never once looked at an image's alt text, a heading's order or a title's length. The sites weren't perfect; our eyesight was narrow. A perfect score meant "nothing we happened to look at was broken", which is a much smaller claim than it appears on a dashboard. When we widened the checks, the same score swung too far the other way — a page with one slightly long title was suddenly counted as dirty as a page answering 404 — and that reading was no more truthful than the first.

Both failures point at the same rule for reading any audit, ours included: ask what was checked, and ask what each finding costs. A dead link that stops the crawler cold and a description that runs a few characters over are different orders of problem, and a tool that counts them equally is producing a number rather than a reading. The other line worth keeping from the start: fixing site health removes drag — it stops you losing clicks and authority you'd already earned. It does not, by itself, move you up the rankings; that part is won by the pages themselves. Any tool implying that a tidy site outranks a useful one is overclaiming, and the rest of its report deserves reading in that light.

Included in the plan

Fly Away Paul walks your whole site the way Google does

Every page fetched and read, every internal link followed to see where it lands, titles, descriptions, headings, alt text and canonicals checked — then reported with each finding graded by what it costs, alongside a record of what was examined, so a clean bill means something. It re-walks monthly on the plan, speaks up when something breaks, and the whole list exports as a playbook your web team or an AI assistant can work through.

Walk your site →

The list Google shows you that no audit tool can see

Open Google Search Console — Google's free reporting desk for site owners, and worth connecting whoever you are — and sooner or later it will show you a wall of "not found" errors. This week, for one site we look after, that wall was 112 URLs long. Our walk of the same site reported none of them, and the walk was right to stay silent, because not one of those addresses is linked from anywhere on the site today.

What that wall actually is: Google's memory. A crawl — anyone's crawl — can only find what today's pages point at. Google, though, re-knocks on every address it has ever known: the site as it stood two redesigns ago, paths from an old sitemap, links other sites made years back, and a surprising number of addresses Google itself assembled wrongly and now dutifully reports as missing. When we checked that 112-row list, the majority were exactly this — the skeleton of a site that no longer exists, plus phantom addresses that never existed at all. No crawler can reproduce that report, and that's no failing of the crawler — the report describes the inside of Google, not the state of your site.

The important part, which the console's red styling doesn't tell you: most of it is fine. In a best-practice world, any removed page with a true successor would have been redirected at the moment it moved — but for a page that's simply gone, with no equivalent to point at, Google's own guidance is that "not found" is the correct answer, it costs your live pages nothing, and redirecting it anyway (usually to the homepage) creates the "soft 404" problem tools rightly complain about. So the wall isn't a to-do list. Buried inside it are the addresses that deserved a redirect and never got one, and telling those apart from the correctly dead is the actual work.

And be clear about how big that buried portion can be. The bulk minting of ghosts happens at a site rebuild: a redesign launches with a fresh structure, nobody writes the redirect map, and a hundred addresses that spent years earning links and visitors all start answering "not found" on the same day. The site we mentioned above had lived through exactly that — the majority of its wall was old pages with no home on the new site, which means the traffic they had built was lost with them, invisibly, at launch. Of everything in this guide, that is the failure with the most money attached. It is also more recoverable than it looks, because the record of what the old site had never went away — which is what the next section is about.

Reading your site's past without borrowing Google's

If you're wondering whether the answer is plugging your audit tool into Search Console, that was our first instinct too — and we decided against it, because we'd rather Fly Away Paul stand on evidence it can gather itself than ask for the keys to another product. It turns out Google's memory is not the only memory of your old site. The Internet Archive's Wayback Machine has been photographing the web for decades, and its public index holds the addresses your site used to have. So our walk now reads that inventory, knocks on every former address, and sorts the dead ones by the only question that matters — is anything still arriving here? Run after a rebuild, that read is the recovery map for the loss described above: it names which of the retired addresses were still earning, so the redirect map that never got written at launch can be written late. Below is the full field guide — every state a dead or dying address can be in, what the evidence means, and what each one deserves. Tap any row for the why.

Your own site still links to itFix the link

A door on your current site opens onto a wall — the worst state on this list, because you're sending your own visitors and your own crawl budget into it. Re-point the link to the right page, and add a redirect as well, so anyone holding the old address from outside lands well too.

Search still sends people to itRedirect it

Google is actively handing you visitors and they are landing on an error page — the strongest redirect case there is. One line of config points them at the successor page, and traffic you were losing today comes back today.

Other sites still link to itRedirect it

The vote of confidence a backlink carries keeps arriving at a dead address, and its value drains away there. These are the most expensive ghosts and the least visible — no visitor may have clicked in months, so nothing looks wrong. Redirect, and the link's worth flows to a live page again.

Arrivals, but none from searchRead the numbers

A handful of unreferred hits on a dead path is often a crawler probing common addresses rather than a person who needs rescuing. If you recognise the URL as a real former page, redirect it; if you don't, "not found" is the correct answer to give a robot guessing at doors.

A real old page with no successorRebuild or re-home

The rebuild casualty: the page earned links and visitors on the old site, and the new site has nothing equivalent to redirect it to. A redirect to its parent section stops the bleeding, but the fuller answer is a new home — content that proved it could earn traffic is the safest content to build again.

Answers 200 but shows "not found"Fix the status

The soft 404 — the page says "nothing here" to people while telling machines everything is fine. Google eventually notices the contradiction and trusts the site a little less for it. Make the address answer honestly: a real 404 if it's gone, a real redirect if it moved.

Redirects — straight to the homepagePoint it properly

The blanket redirect is the well-meant version of the soft 404: everything old gets funnelled to the front door, and both the visitor and Google arrive somewhere that has nothing to do with what they wanted. Each redirect should land on the page's actual successor, or its section — the homepage answer teaches Google to treat all your redirects as noise.

Nothing arrives at allLeave it be

No links, no visitors, no memory outside Google's. This is the healthy majority of any old site's ghosts — a retired page dying correctly. "Not found" is the right answer for it, and a tool that tells you to "fix" these is inventing work.

A console shows you the wall; a useful reading tells you which three of the hundred deserve a redirect and why, next to the reassurance that the rest are dying as they should.

What only Google can still tell you

Some of what Search Console shows, no outside tool can reconstruct — ours included. It reports Google's decisions: whether it chose a different canonical address for a page than the one you declared, whether a page was crawled but left out of the index, which searches your pages actually appeared for. Those are facts about Google's private ledger, not about your site, and no amount of crawling — ours or anyone's — can observe them from outside. The residue of its 404 memory that no archive holds sits in the same category, though as it happens that residue is made of the addresses nothing ever points at, which is precisely the kind that costs you nothing.

The practical stance we'd recommend to anyone: keep Search Console connected to your site and glance at it occasionally — it's free, it's yours, and it's the only window into those verdicts. Then make the verdicts boring. A site that declares a clean canonical on every page gives Google nothing to choose differently; a site whose sitemap is accurate and whose dead addresses are redirected or correctly buried gives the index nothing to puzzle over. Prevention reads the same as omniscience from the outside, and it's available to everyone.

The one button only you can press

There remains a single lever in this whole trade that no tool can pull on your behalf — not ours, not the expensive suites, none of them. When you've fixed something Google flagged, or published a page it hasn't picked up, Search Console has a button for it: "validate fix" on a flagged issue, "request indexing" on an inspected URL. It nudges Google to come and look sooner than its own schedule would. Google exposes no programmatic way to press it — the click lives in the console's interface and nowhere else, which is almost certainly deliberate, since a button robots could press would be pressed by robots all day long.

Two things follow from that. First, the button moves the inspector's schedule and nothing else — whatever you claimed to fix still has to hold up when Google arrives. Second, a well-kept site rarely needs it. Google finds changes on its own by re-reading your sitemap — one whose dates actually change when pages change is the standing "come and look" signal — and by following links from the pages it visits often, which is why new work linked prominently from your homepage gets found fast. The button is for the exceptions: a page that's clean, listed, linked and still sitting unindexed weeks later, or a fix you'd like off the console's red list before the next monthly sweep. When that's the situation, log in and press it. It takes ten seconds, and it is the one contribution to your site's health that will always be yours.

The button is a doorbell, not a repair — it brings the inspector round early, and the fix still has to be real when they arrive.

Where this lands

Site health, done properly, is less a score to chase than a standing answer to a short list of questions: is anything broken, is anything leaking value you already earned, is the site legible to the machines that decide when to show it — and is what remains on the list actually worth doing, or just red styling on housekeeping? You can hold any tool to that standard, including ours: it should say what it checked, grade what it found by cost, tell you the past is dying cleanly when it is, and admit which facts live inside Google where nobody else can read them.

Close the loop

The walk, the ghosts, and what to do about each

Fly Away Paul walks your live site, reads your site's past from the public archive, checks which dead addresses still carry visitors or links, and gives every finding a verdict — fix, redirect, or leave it be — with the reassurance said as plainly as the faults. It sits in the plan next to your rankings, AI visibility and backlink health, and everything exports as a playbook when the fixing isn't your job.

See your site's health →