astro-seo-enforcer: fail the build on SEO regressions
I added four gallery components to this blog last week. The build stayed green. The site still looked fine. And without noticing, I had pushed a page with two <h1> elements and a handful of <title> tags that ran well past 60 characters.
Nobody would have caught that for months, neither me nor a reviewer nor a test. astro-seo-enforcer exists to catch exactly that.
Where the idea came from
A few weeks earlier I had signed up for Ahrefs and pointed its Site Audit at this blog. It came back with a wall of warnings. Titles too long. Titles too short. Duplicate titles across the English and German versions. Meta descriptions missing or barely a sentence long. Pages with more than one H1. Nothing was broken, nothing was on fire, just dozens of small SEO papercuts that had been quietly shipping for months.
I worked through the list. Then it hit me that Ahrefs only tells you after the fact. You publish, it crawls a few days later, and then you find out. I wanted that feedback moved to the left. Same checks, but running before anything ships: at build time, in CI, on the exact HTML that would go live. That is the tool.
What astro-seo-enforcer does
astro-seo-enforcer is an Astro integration that hooks into astro:build:done, walks your output directory, parses every .html file, and runs a set of SEO rules against the real markup. Some rules look at one page in isolation. Others wait until every file has been parsed, because a duplicate title, a broken link or an orphaned page only exists relative to the rest of the site. If anything comes back a violation, it prints a grouped report and exits with a non-zero code, so your CI pipeline fails. Zero config to start, fully configurable when you need it.
Why check the built HTML
Your .astro files are not what ships. By the time a page is written to dist/, a layout has wrapped it, an i18n helper has swapped strings, a content collection has injected frontmatter, and any components on the page have added their own markup. A <title> that looked fine in the layout can be too long once the post title and the site suffix are joined. An <h1> can go missing because a template branch did not fire.
Linting the source cannot see any of that. astro-seo-enforcer runs after the build, on the same HTML a crawler downloads.
Setup
npx astro add astro-seo-enforcer
That installs the package and writes it into your integrations array. Manually it is a dev dependency plus four lines:
npm install -D astro-seo-enforcer
// astro.config.mjs
import { defineConfig } from 'astro/config';
import seoEnforcer from 'astro-seo-enforcer';
export default defineConfig({
site: 'https://example.com',
integrations: [seoEnforcer()],
});
Run astro build. The checks execute once the static files exist. It is a no-op during astro dev.
The per-page rules
Every rule is on by default. Set any to false to disable it, or pass an object to tune it.
| Rule | Default severity | What it checks |
|---|---|---|
title | error | <title> exists and is 30 to 60 characters. |
metaDescription | error | <meta name="description"> exists and is 50 to 160 characters. |
headingHierarchy | error | Exactly one <h1>, no skipped levels, optionally <h1> first. |
semanticHtml | error | At least one landmark element such as <main>, <header>, <footer>. |
imageAlt | error | Every <img> has an alt attribute. Empty alt is allowed, a missing attribute is not. |
canonical | error | Exactly one <link rel="canonical"> with a non-empty, absolute URL. |
anchorText | warning | No generic link text like “click here” or “read more”. No links without an accessible name. |
internalLinks | error | Every internal <a href> resolves to a file in the build output. Broken #fragment links are reported separately, at warning level. |
jsDependency | error | The <body> has real text content, not a near-empty shell waiting for client rendering. |
robots | warning | Warns when <meta name="robots"> contains noindex or nofollow. |
duplicateId | error | No id value is used twice in a document. |
imageSize | warning | Local images stay under maxBytes (200 KB by default), carry width/height so the browser can reserve space, and are not served more than 2x larger than they display. |
structuredData | warning | Every <script type="application/ld+json"> block parses. requireTypes asserts that a given @type is present; require flags pages that ship no JSON-LD at all. |
thinContent | warning | The main content region holds at least minWords words, 250 by default. |
The checks that need the whole build
These run once every file has been parsed, because they compare pages against each other or against the sitemap.
| Check | Default severity | What it checks |
|---|---|---|
duplicate title | error | The same <title> on two or more pages. |
duplicate metaDescription | warning | The same description on two or more pages. |
duplicate <h1> | warning | The same <h1> text on two or more pages, which is two of your own pages competing for one query. |
duplicateContent | warning | Pages whose visible text is near-identical to another page’s. Similarity is the Jaccard overlap of their five-word runs, and a pair at or above threshold (0.9) is reported on both pages. A second pass flags any page less than minUniqueRatio (0.2) of whose text is unique to it, which catches many-way templating that no single pair trips. |
orphanPages | warning | HTML pages that no other page links to and no sitemap lists. |
sitemapCoverage | warning | Every <loc> in sitemap*.xml resolves to a real file, every indexable page is listed, and nothing is both noindex and in the sitemap. |
The content comparisons look at the <main> or <article> region rather than the whole <body>, so shared nav and footer text does not count. Otherwise two pages would look alike purely because they share a layout, and a page with three sentences of text would clear the word count.
What happened when I ran it on this blog
The first run reported 843 errors and 2 warnings across 341 pages.
That number looks brutal until you break it down. The overwhelming majority came from auto-generated tag and category listing pages. This blog has around 200 of them, one per tag, and they all carry short, near-duplicate titles like #vlan | SlashGordon and meta descriptions of eight characters. Those pages are legitimately low value. I excluded them:
seoEnforcer({
exclude: ['tags/', 'de/tags/', 'categories/', 'de/categories/'],
});
That left about 67 real errors on actual content, and those were worth fixing.
Around 49 titles ran over 60 characters. This blog uses long descriptive titles on purpose, so instead of gutting them I added an optional seoTitle field to the post schema, letting the <title> tag differ from the on-page <h1>, and set a shorter tag on the 26 posts that needed one.
Ten short and four duplicate titles sat on the home, about, projects, all-posts, and tools index pages, things like Tools | SlashGordon used identically on the English and German version. I gave them unique, localised titles. Two thin meta descriptions on the projects pages got lengthened past 50 characters. And one post had two <h1> elements, because the markdown body still carried a # Heading on top of the layout’s <h1>; I removed the duplicate.
The whole cleanup took maybe an hour, most of it writing 26 short titles, and the report told me exactly which files and why.
Tuning the newer rules
The link, image and duplicate checks arrived in 1.1 through 1.5, and turning them on took some thresholds that are specific to this site.
imageSize defaults to a 200 KB ceiling per file. My galleries already emit WebP at quality 58 with widths capped at 1000px, and the heaviest a detailed 1000px HDR photo lands at after that is around 384 KB, so I raised maxBytes to exactly that. Anything above it is unoptimised, not just a big photo.
internalLinks treats a missing page as an error and a broken #fragment as a warning. That split suits this site: MapGallery and ImageTimeline render their anchor targets client-side, so an id can be absent from the built HTML and still work perfectly in the browser. A missing file is a real dead link; a missing fragment is a maybe.
duplicateContent needed the most thought. At the default minUniqueRatio of 0.2 the English and German home pages both tripped it at 15% unique, because they list the same recent post titles and excerpts and only the surrounding prose differs. That is not a mistake, and hreflang already tells Google the two are a pair, so I dropped the ratio to 0.1. The check stays live for templated page sets without shouting about a translation pair.
structuredData I set to require: true with requireTypes: ['WebSite']. The layout emits a WebSite JSON-LD block on every page, so requiring it turns the rule into a regression guard: if the block ever gets dropped, or JSON.stringify emits something invalid, the report says so. I deliberately did not require BlogPosting or BreadcrumbList, because requireTypes applies to every scanned page and this site does not emit them yet.
sitemapCoverage came with one ordering gotcha. This blog flattens Astro’s sitemap in its own astro:build:done hook, and the SEO checks read whatever sitemap*.xml is on disk. That works only because flatSitemap() is registered before seoEnforcer() in the integrations array and Astro runs the hooks in that order. Register them the other way round and the check reads a sitemap that is about to be rewritten.
Where it stands now
The latest build writes 370 pages, of which 77 are scanned after exclusions, and reports zero errors and eleven warnings.
The exclusion list has grown past the tag and category pages. Legal pages (impressum, privacy, datenschutz) are held to a template, not to article rules. The Game of Life route is an interactive app page rather than an article, so it was excluded too. Only the English one had been missing from the list, which is exactly the kind of asymmetry you find by reading a report instead of guessing.
Of the eleven warnings, nine are thinContent and two are duplicate <h1>. No article trips the word count. The pages that do are the short hub pages: about at 194 words in English and 183 in German, projects at 238 and 195, the tools index at 55 and 42, and the two FRITZ!Box decrypt tool pages at 149 and 141. The duplicate <h1> is the literal word “Tools” on both locale versions of the tools index.
None of that breaks the build, and it should not. Those are content decisions for the backlog. The warning tier exists for this: the report keeps a list of things worth writing without holding up a deploy.
The CI pattern I settled on
I do not want a failed local build every time I am mid-edit. I do want the pipeline to block a regression. So:
seoEnforcer({
failOn: process.env.CI ? 'error' : 'never',
rules: {
title: { minLength: 30, maxLength: 60, checkDuplicates: true },
metaDescription: { minLength: 50, maxLength: 160, checkDuplicates: true },
headingHierarchy: { requireSingleH1: true, enforceNoSkips: true, checkDuplicateH1: true },
imageSize: { severity: 'warning', maxBytes: 384 * 1024, maxScaleFactor: 2 },
internalLinks: { severity: 'error', checkFragments: true, fragmentSeverity: 'warning' },
duplicateContent: { severity: 'warning', threshold: 0.9, minUniqueRatio: 0.1 },
structuredData: { severity: 'warning', require: true, requireTypes: ['WebSite'] },
orphanPages: { severity: 'warning', entryPoints: ['index.html', 'de/index.html'] },
},
exclude: ['tags/', 'de/tags/', 'categories/', 'de/categories/'],
});
Locally, npm run build prints the report and carries on. In Gitea CI, where CI is set, a single error stops the build. New post, new component, template refactor: if any of them break a <title>, drop an <h1> or leave a link pointing at a page that no longer exists, I find out in the pipeline, not in the analytics three months later.
Every rule in that block is on by default. I spelled them out anyway, so the thresholds this site was tuned against are explicit instead of quietly inherited from whatever a future release decides the defaults should be.
What it is not
It is not a Lighthouse replacement. It does not measure performance and it does not run a full accessibility audit. It never fetches anything either: internal links are resolved against the files on disk, and external URLs are left alone. What it does is check the structural SEO basics that tend to rot silently, from the title and description on one page up to the link graph across the whole build.
FAQ
Does it slow the build down? Barely. It reads each HTML file once and parses it with a lightweight parser. On this blog the full run, image headers and cross-page checks included, takes about half a second.
Do the site-wide checks scale?
The duplicate-content pass compares every page against every other, so its cost grows with the square of the page count. There is a maxPages guard, 1500 by default, above which that pass is skipped with a notice instead. Everything else is linear.
Does it work with SSR or hybrid output?
It checks whatever HTML ends up in the output directory. For hybrid builds that means the pre-rendered pages. Pure on-demand routes are not in dist/, so they are not checked.
Can I use the checks outside a build?
Yes. runSeoChecks, resolveConfig and formatReport are exported for tests or a standalone script.
What Astro versions does it support? Astro 3 through 7.
Links
- npm:
astro-seo-enforcer - Source and full rule reference: github.com/SlashGordon/astro-seo-enforcer
- The post that triggered the cleanup: astro-gallery