Technical SEO Audit

FreeOfficial

Audit crawlability, rendering, metadata, and structured data with evidence-backed fixes. Use for indexing problems, migrations, or pre-launch SEO checks.

frontendseocrawlabilitystructured-datametadatarenderingaudit· v1· by SkillingMain
64
Usefulness score
999
Installs
No ratings yet
today
Last updated

Model requirements

Capability tier

basic

Recommended models
Any instruction-tuned LLM (7B+)Llama 3.1 8B (self-hosted)Phi-3.5 (self-hosted)Claude Haiku 3.5GPT-4o-miniGemini 2.0 Flash

Skill instructions

When to use

Use when asked to audit crawlability or indexation, diagnose "page not indexed" or ranking drops, review metadata or structured data, verify that a JavaScript-heavy site renders for bots, or run a pre-launch or pre-migration SEO check. Do not use for content strategy or keyword research — this skill covers the technical layer only.

Inputs to gather

  • Domain plus one representative URL per template (home, listing/category, detail, article).
  • Rendering mode per route (SSG, SSR, CSR, hybrid) and the framework.
  • Search Console access or exported coverage data, if available.
  • Target locales and whether hreflang is in scope.
  • Any recent migration or URL-structure change, with dates.
  • Locations of robots.txt and the sitemap(s).

Procedure

  1. Crawl fundamentals:
    • curl -sI https://<domain>/ — expect 200.
    • Test all four origin variants (http/https x www/non-www) plus a trailing-slash variant with curl -sIL, counting hops: every variant must 301 (not 302) directly to the canonical origin in ONE hop, no chains.
    • Fetch robots.txt: confirm it does not block CSS/JS/asset paths, has no accidental Disallow: /, and references the sitemap.
    • Fetch the sitemap: valid XML, only 200-status canonical indexable URLs, under 50,000 URLs and 50MB per file.
    • Verify lastmod reflects real content changes, not the deploy timestamp.
  2. Indexation controls per template:
    • Check BOTH the X-Robots-Tag response header and the meta robots tag; they can conflict, and the more restrictive one wins.
    • Canonical: absolute URL, self-referencing on canonical pages, appears exactly once, and agrees with the sitemap and internal links.
    • Decision point for faceted/parameter pages: use noindex or a canonical — never block them in robots.txt, because a blocked page's noindex can never be seen.
  3. Rendering audit — the highest-risk area on JS-heavy sites. For one URL per template, compare curl -s <url> raw HTML against the rendered DOM:
    • If the title, meta description, canonical, main content, or internal links exist only after JavaScript runs, flag the route as at-risk: search-engine rendering is delayed and inconsistent, and most non-Google crawlers (including LLM crawlers) never execute JS.
    • Recommend SSR, SSG, or prerendering for every indexable route.
    • Flag client-side injection that CHANGES server-sent metadata — conflicting signals are worse than missing ones.
  4. Metadata per template:
    • Unique title of roughly 50-60 characters with the differentiating keyword front-loaded.
    • Meta description of 120-160 characters written for click-through — it is not a ranking factor and search engines rewrite most of them; its job is CTR.
    • Open Graph and Twitter card tags on shareable templates.
    • Duplicate titles across URLs almost always indicate a template bug — report the root cause, not a URL list.
  5. Structured data:
    • Prefer JSON-LD delivered in the initial HTML, not injected post-load.
    • Match types to templates: Organization and WebSite sitewide, BreadcrumbList on hierarchical pages, Product with offers on product pages, Article with author and dates on editorial pages.
    • Validate with the schema.org validator or Rich Results Test — presence is not validity.
    • Confirm every marked-up value is visible on the page; invisible markup violates spam policy.
  6. Status-code hygiene:
    • Missing content returns a real 404/410, never a 200 "not found" page (soft 404).
    • Removed pages 301 to a RELEVANT equivalent, not the homepage.
    • Internal links point at final URLs so crawlers never pay for internal redirect hops.
  7. If multilingual:
    • hreflang must be bidirectional — every page in the cluster lists all alternates plus itself — and include x-default.
    • Each page's canonical must be self-referencing within its own locale; a canonical pointing across locales silently invalidates the cluster.
  8. Architecture and discovery:
    • Key pages reachable within 3 clicks of home.
    • Check for orphans: URLs in the sitemap that no internal link points to.
    • Pagination and "load more" use real <a href> links, not JS-only handlers, or deep items are undiscoverable.
  9. Page-experience layer: viewport meta present, HTTPS everywhere with no mixed content, no intrusive interstitials. Summarize Core Web Vitals status but delegate deep performance work to a dedicated performance audit.
  10. Prioritize findings: P0 blocks crawling or indexing entirely; P1 causes wrong-page indexing or duplication; P2 costs rich results or CTR; P3 is hygiene.

Output format

### Verdict
One paragraph: can the site be crawled, rendered, and indexed correctly today?
Name the single biggest risk.

### Findings by priority
[P0-P3] Title
- Affected: templates/URLs with counts
- Evidence: the actual command output, header, or tag observed
- Fix: the exact change (header value, tag, redirect rule, or code)
- Validation: the command or tool check that confirms the fix

### Template matrix
Table: template → title | description | canonical | robots | structured data (OK / issue ID)

### Post-fix monitoring
What to watch in Search Console (coverage states, enhancement reports) and for how long.

Quality checklist

  • Every claim is backed by a fetched response or a code reference — nothing asserted from assumption.
  • Raw vs. rendered HTML compared for at least one URL per template, with the differences listed.
  • All four protocol/host variants tested for redirect behavior and hop count.
  • Canonical tags, sitemap entries, and internal links cross-checked for agreement.
  • Structured data validated by a tool, not just observed to exist.
  • Findings reference templates and root causes, not enumerations of individual URLs.
  • No recommendation anywhere to robots.txt-block a page that needs deindexing.

Common pitfalls

  • Blocking URLs in robots.txt to remove them from the index. Blocked pages cannot be crawled, so their noindex is never seen — they linger as "indexed, no content". Deindex first, block later if ever.
  • A canonical that points to a URL which redirects or 404s — the signal is discarded and the duplicate cluster resolves unpredictably.
  • Sitemap lastmod auto-set to build time on every deploy: once it proves unreliable, crawlers ignore the field sitewide and recrawl prioritization suffers.
  • One-directional hreflang: a missing return tag invalidates that pair silently. No error is surfaced anywhere — it just stops working.
  • SPA route changes that update content but not title/canonical — every client-side route must own its metadata, or all routes compete as one page.
  • Soft 404s from empty states: a category page returning 200 with "0 results" gets classified as soft 404 and drags down crawl efficiency for the whole section.
  • Auditing only the homepage. Template-level bugs live on listing and detail pages, which is where the traffic is.
  • Treating a passing homepage curl as proof the site renders for bots when half the routes are CSR — always sample every template.
  • Recommending a blanket Disallow for staging cleanup on the production robots.txt — migrations have shipped this exact mistake; diff robots.txt as part of every deploy check.