Most technical SEO checklists are a list of everything that could be wrong with a website. That is useful once and overwhelming after that, because it gives you fifty items and no order.
This one is ordered, and the order comes from running the checks on this site. Where a number appears below, it is from technoxprt.com’s own Search Console rather than an example.
Work down the list. The early sections decide whether the later ones matter at all, because a page Google never fetches cannot be helped by its Core Web Vitals.
How to use this
- Do it in order. Crawling before indexing, indexing before performance.
- Stop at the first section where you find something broken and fix that before moving on.
- Most sites have two or three real problems, not fifty.
1. Crawling: Can Google Reach Your Pages?
Everything else is downstream of this. Start in Search Console under Settings → Crawl stats, a report most people never open.
Check your crawl budget is not being wasted
Crawl stats breaks every request Googlebot made in the last 90 days by response code, file type and purpose. Here is what this site’s looked like:
| Metric | This site | What it means |
| Total crawl requests | 4,780 in 90 days | About 53 a day |
| Returning 404 | 20% | Roughly 956 wasted fetches |
| HTML | 28% | JS, CSS and images take 42% |
| Purpose: discovery | 10% | 90% is re-checking known pages |
A fifth of the budget going to dead URLs is the finding that mattered. It is also the easiest thing on this entire list to fix.
Your checks:
- 404 share. Anything above a few per cent is budget being burned. Find what links to those URLs and either fix the link or redirect it.
- Discovery share. If it is very low and you have unindexed pages, crawl capacity is your bottleneck, not content.
- Average response time. Google crawls slow sites more cautiously. This site averages 556 ms, which is not a constraint. Two seconds or more would be.
- Host status. Any availability problems here outrank everything else on this page.
Check robots.txt is not blocking what you need
Load yoursite.com/robots.txt and read it. The classic failure is a Disallow: / left over from a staging site.
Do not block CSS or JavaScript. Google renders pages, and blocking the files it needs to render them means it sees a broken version of your site.
Remember that robots.txt controls crawling, not indexing. A blocked URL can still appear in results if other sites link to it. To keep a page out of the index you need a noindex tag, which means the page must stay crawlable for Google to see it.
Check your non-HTML content is reachable too
Only 28% of Googlebot’s requests on this site were HTML. The rest went to scripts, stylesheets and images, and images, video and PDFs are content that can rank in their own right.
- Images served from a crawlable path, not locked behind a script or a hotlink block that also refuses Googlebot.
- Real
<img src>with alt text, not CSS backgrounds, for anything you want found in image search. - Video with a real page around it and, ideally, VideoObject markup. An embed with no surrounding text gives Google nothing to index.
- PDFs get indexed and rank. If yours contain content that should live on a page, that is a content decision, not a technical one.
Check your sitemap is current and honest
- It exists, returns 200, and is submitted in Search Console.
- It contains only indexable URLs, no redirects, no 404s, no
noindexpages. lastmoddates are real. Setting every page to today teaches Google to ignore the field.- The URL count roughly matches the number of pages you actually want indexed.
2. Indexing: Is Google Keeping Your Pages?
Search Console → Pages splits your URLs into indexed and not indexed, with reasons. The reasons matter more than the totals.
Read the not-indexed reasons properly
Two of them are routinely misread, and they mean opposite things.
- Crawled – currently not indexed means Google fetched the page and declined it. On this site, 28 of the 51 URLs in that bucket were not articles at all, RSS feeds, paginated archives, author pages and two JavaScript files. No amount of rewriting indexes a JavaScript file.
- Discovered – currently not indexed means Google has never fetched the page. All 36 URLs in that bucket here showed “Last crawled: N/A”, and they were the site’s best current posts. Rewriting a page Google has not read cannot help.
Before you act on any URL, sort the list by what the URLs actually are. Feeds, pagination and archives get handled technically. Only what is left is a content question.
Check canonicals point where you think
Every indexable page should have a self-referencing canonical tag. Then verify Google agrees: the URL Inspection tool shows both your declared canonical and the one Google selected. When they differ, Google has decided two of your pages are the same thing.
Check for duplicate and near-duplicate pages
Parameter URLs, tag archives, printer versions and HTTP versions all create copies. Each one splits signals and consumes crawl budget.
Search your own site for your main topics. Two pages circling the same question is a merge, not an optimisation.
3. Site Structure and Internal Links
This section contains the single most useful thing found while auditing this site, and it contradicts what most checklists imply.
Sitewide links do almost nothing
Crawling all 123 posts on this site produced a link graph. Two of the never-crawled pages were linked from every single page on the site. 123 internal links each, from the theme’s related-posts widget.
Google had still never fetched them.
Meanwhile the pages with one or two genuine in-content links were in exactly the same state. The conclusion is blunt: theme-generated blocks (related posts, popular posts, sidebar lists), are discounted heavily. They are the same on every page, so they carry almost no information.
What counts is a link inside a sentence, in the body of a relevant page, at the point where the reader would want it.
Your structure checks
- Click depth. Important pages within three clicks of the homepage.
- Orphans. Any page with zero in-content links pointing at it. Crawl your own site to find them, Screaming Frog does this free up to 500 URLs.
- Link from crawled pages. A link from a page Google never visits is worthless. Check that your source pages are actually indexed.
- Descriptive anchors. The existing words in your sentence, not “click here” and not an exact-match keyword bolted on.
- Breadcrumbs present and marked up. They give Google a second, explicit signal about hierarchy, and they are what appears in place of the raw URL in results.
4. URLs, Redirects and Status Codes
Redirect checks
- Single hop. Chains waste crawl budget and lose a little signal at each step.
- 301, not 302, for anything permanent.
- Relevant destination. This is the one people get wrong.
On relevance: redirecting a deleted page to something only loosely related gets treated as a soft 404, which is worse than a clean 404. If genuinely nothing on your site answers what that page answered, let it 404. That is the correct signal, not a failure.
Also check where your redirects land. On this site, 52 of 70 redirects pointed at category archives that were set to noindex, so the destination could not rank at all. Five were repointed at real posts once that was spotted.
URL checks
- Readable, lowercase, hyphenated, no session IDs or dates you will regret.
- One protocol and one hostname. Pick HTTPS with or without www and redirect the other three variants to it.
- Consistent trailing slashes.
- No URLs left behind from an old permalink structure. This site still had a bare
/2102/sitting in Search Console. - Broken outbound links fixed. They do not carry a penalty, but a page full of dead references reads as unmaintained to a reader and to a reviewer.
5. Mobile
Google indexes the mobile version of your site. Not the desktop version with mobile as a consideration, but the mobile version is the one that counts.
The practical implication catches people out: if content, links or structured data exist on desktop but not on mobile, for indexing purposes they do not exist.
- Same content on both. Text hidden or stripped on mobile is text Google may not index.
- Same internal links on both. A desktop-only sidebar or mega-menu is a set of links Google does not see.
- Same structured data and metadata on the mobile template.
- Usable, not just responsive. Tap targets big enough, text readable without zooming, nothing overflowing horizontally.
- Interstitials restrained. A popup covering the content on arrival is a documented problem, not a grey area.
Check it the direct way: load the page in a mobile viewport and compare what is there against desktop. URL Inspection’s rendered screenshot shows you what Google actually got.
6. Performance and Core Web Vitals
Google’s three thresholds, from its own documentation, are:
| Metric | Measures | Good |
| Largest Contentful Paint (LCP) | Loading | 2.5 seconds or less |
| Interaction to Next Paint (INP) | Interactivity | 200 milliseconds or less |
| Cumulative Layout Shift (CLS) | Visual stability | 0.1 or less |
These are judged at the 75th percentile of real page loads, measured separately for mobile and desktop. A lab score in a testing tool is a diagnostic, not the thing being measured.
Where the wins usually are
- Images. Nearly always the biggest and easiest gain. Correct dimensions, modern formats, lazy loading below the fold, and explicit width and height to stop layout shift. Image SEO covers this properly.
- Third-party scripts. Chat widgets, analytics and ad tags are usually the biggest INP problem. Audit what is actually loading and remove what nobody uses.
- Server response. Caching and a CDN help here, and they also raise how much Google is willing to crawl.
- Fonts. Preload, and use
font-display: swapso text is visible while they load.
On WordPress specifically, plugin bloat is the usual cause; there are practical steps for speeding a WordPress site up that go beyond installing another caching plugin.
One caution worth stating plainly. Core Web Vitals is a real ranking signal and a small one. If your pages are not being indexed, performance work will not fix that, and it is the wrong thing to start with.
7. Rendering and JavaScript
Google renders JavaScript, but rendering is queued separately from crawling and costs it more. Anything critical should be in the initial HTML.
The test takes a minute: use URL Inspection, run “Test live URL”, then look at the rendered HTML and screenshot. If your main content, headings, links or canonical tag are missing there, Google is not reliably seeing them.
- Main content present without JavaScript executing, where possible.
- Internal links as real
<a href>elements, not click handlers on a div. - Content behind tabs and accordions present in the HTML, not fetched on click.
- Infinite scroll backed by paginated URLs that can be crawled directly.
8. Security
- HTTPS everywhere, with a valid certificate and every HTTP URL redirected. The HTTPS migration checklist covers the parts people miss.
- No mixed content. One image loaded over HTTP downgrades the whole page in the browser.
- Certificate expiry monitored. An expired certificate is an outage.
- Security headers set. These do not lift rankings directly, but they matter for how safe the site looks and for hardening it.
- Check Search Console → Security issues periodically. A hacked site loses traffic faster than any ranking factor gains it.
9. Structured Data
Structured data does not raise rankings. It changes how your result can appear, which changes clicks.
- Validate with Google’s Rich Results Test, not just a generic schema validator.
- Markup must match visible text. FAQ schema describing questions that are not on the page is a straightforward violation.
- Check Search Console’s enhancement reports for errors you have not noticed.
- Mark up what you actually are (Article, Product, LocalBusiness, FAQ), and stop there.
10. Beyond Google
This section exists because of a finding on this site that nothing in a standard checklist would have caught.
Bing had indexed zero pages. Searching the brand name returned only an Instagram profile, and Bing autocorrected the name to a different word.
Nothing was blocking it. robots.txt was open, bingbot got 200s, the sitemap worked. The cause was mundane: Bing Webmaster Tools held one sitemap, submitted in 2015, listing 7 URLs, last crawled in 2019. A current sitemap had simply never been submitted.
That matters beyond Bing, because Bing’s index also feeds DuckDuckGo and Microsoft Copilot.
- Bing Webmaster Tools set up, with a current sitemap submitted.
- IndexNow enabled so new URLs are pushed rather than waiting for a crawl.
- Consider an llms.txt file if you want to be explicit about what assistants should use.
Decide what AI crawlers may do, on purpose
These get blocked by accident more often than deliberately, usually by someone pasting a robots.txt snippet from a forum.
They are separate from Googlebot and they do different jobs, so blocking them has different consequences.
- GPTBot. OpenAI’s training crawler. Blocking it keeps your content out of future model training.
- OAI-SearchBot. OpenAI’s crawler for ChatGPT’s live search. Blocking this one removes you from ChatGPT’s search results, which is usually not what people intend when they block GPTBot.
- PerplexityBot, ClaudeBot, CCBot. The same trade-off in different places.
- Google-Extended. The one most often misunderstood.
Google’s own documentation is explicit about that last one: Google-Extended controls whether your content trains Gemini and grounds its answers, and it does not affect your inclusion in Google Search. It also has no user agent string of its own, crawling still happens under the normal Google agents, and the token exists only as a control in robots.txt.
So blocking Google-Extended costs you nothing in Search. Blocking OAI-SearchBot costs you visibility in ChatGPT. Those are different decisions and worth making separately.
Check what you are already blocking
- Read every
Disallowin your robots.txt and be able to say why each one is there. - Check your CDN and host separately. Cloudflare and several hosts now offer one-click AI bot blocking, and it works at the edge, above robots.txt. Your robots.txt can look completely open while the request is refused before it ever reaches WordPress.
- Check your SEO plugin. Some now write AI crawler rules into robots.txt for you.
Test it rather than assuming. Request a page with the bot’s user agent and see what comes back:
curl -s -o /dev/null -w "%{http_code}" -A "GPTBot" https://yoursite.com/
A 200 means it is getting through. A 403 means something is blocking it, and if that was not deliberate, this is the check that finds it.
The Tools You Actually Need
Everything in this checklist can be done on free tools. Paid tools mainly save time on large sites.
| Job | Tool | Cost |
| Crawl budget, indexing, CWV field data, structured data errors | Google Search Console | Free |
| Crawling your own site for orphans, broken links, redirect chains | Screaming Frog | Free to 500 URLs |
| Checking a single page as Google sees it | URL Inspection, “Test live URL” | Free |
| Lab performance diagnosis | PageSpeed Insights / Lighthouse | Free |
| Rich result eligibility | Google’s Rich Results Test | Free |
| Bing, DuckDuckGo and Copilot indexing | Bing Webmaster Tools | Free |
Two things worth knowing about that list. Search Console covers more of this checklist than any paid tool does, and it is the only one with Google’s own data rather than an estimate of it.
And every finding quoted in this guide came from those free tools. The crawl budget numbers, the indexing breakdown, the internal link graph and the Bing discovery all came out of Search Console, Bing Webmaster Tools and a crawl.
Two Checks Most Sites Can Skip
Both of these appear on nearly every technical checklist. Most sites reading one do not need either, and it is worth saying so rather than adding items nobody will action.
hreflang, if you publish in more than one language
Single-language site: skip this entirely. Multi-language or multi-region: hreflang tells Google which version to show which audience, and it is easy to break.
- Every version references every other version, including itself.
- The tags are reciprocal. A one-way reference is ignored.
- Language and region codes are valid, and
x-defaultis set for the fallback. - Only canonical URLs are referenced, never redirecting ones.
Log file analysis, if you are large enough to need it
Server logs show every request Googlebot actually made, which is more complete than Crawl stats. They tell you which sections get crawled, which get ignored, and how much budget goes to URLs you did not know existed.
For a site of a few hundred pages, Crawl stats answers the same questions well enough and takes a minute instead of an afternoon. The threshold where logs start earning their time is somewhere in the tens of thousands of URLs, or any site where crawl budget is a known constraint.
11. What to Actually Do First
A fifty-item checklist is only useful with an order attached, so here is the one this site’s data produced.
| Priority | Check | Why it is first |
| 1 | Host status and availability | Nothing else matters if the site is down for Googlebot |
| 2 | robots.txt not blocking anything important | One line can hide a whole site |
| 3 | 404 share in Crawl stats | Biggest, cheapest crawl-budget win |
| 4 | Indexing report, sorted by URL type | Tells you whether you have a technical or a content problem |
| 5 | In-content internal links to unindexed pages | The most direct lever you control |
| 6 | Canonicals and duplicates | Stops your own pages competing |
| 7 | Core Web Vitals | Real, but small, and useless on unindexed pages |
| 8 | Structured data | Affects appearance, not position |
Most sites stop finding real problems around item four or five.
How often to run it
Crawl stats and the indexing report are worth a monthly glance. The full list is a quarterly job, plus any time you change hosting, redesign, migrate, or move a large number of URLs.
Pair the technical checks with the SEO KPIs you already report on, so a fix can be tied to something that moved.
Frequently Asked Questions
How often should I do a technical SEO audit?
A full pass quarterly is enough for most sites, with a monthly look at Crawl stats and the indexing report. Run the full list immediately after a migration, redesign, hosting change, or any bulk URL change.
What is the most common technical SEO problem?
Wasted crawl budget, usually from 404s and redirect chains. On this site a fifth of all crawl requests returned 404, and it went unnoticed because nothing visibly broke.
Do I need paid tools for a technical audit?
No. Search Console covers crawling, indexing, Core Web Vitals and structured data for free, and Screaming Frog’s free tier crawls up to 500 URLs. Paid tools mainly save time on larger sites.
Does technical SEO still matter with AI search?
More, if anything. An assistant can only cite a page it can fetch, read and understand. Crawlability, clean HTML and accurate structured data are what make a page usable to one.
Are Core Web Vitals a ranking factor?
Yes, and a small one. They are worth fixing, but not before crawling and indexing problems. A page Google never fetches gains nothing from a fast load.
Should I worry about pages in “Crawled – currently not indexed”?
It depends entirely on what those URLs are. On this site, 28 of 51 were feeds, pagination, author archives and JavaScript files, which Google was right to skip. Sort the list by URL type before deciding anything.
Can technical SEO fix a site with no traffic?
It can remove the things stopping a site from ranking, which is not the same as making it rank. If pages are indexed and still invisible, the constraint is usually authority or content, not technical setup.
