Skip to content
AstroCraft Docs
On this theme

SEO

Indexa has no SEO package. Every meta, Open Graph and Twitter tag is emitted natively by src/layouts/BaseHead.astro, the JSON-LD is built by typed helpers in src/js/schema.ts, and robots.txt, llms.txt, the RSS feed and the sitemap are generated endpoints. Same stance as the icons and the motion catalog: owned, not vendored — which means every tag on the page is one you can read and change.

The one variable

site in astro.config.mjs, fed by SITE_URL, is the root of all of it. Canonical URLs, og:url, the JSON-LD @ids, the sitemap, the Sitemap: line in robots.txt and every link in llms.txt derive from it. Setting it once fixes them together; getting it wrong breaks six things silently, which is why a production deploy on the placeholder throws. Deployment covers the gate.

What every page emits

BaseHead resolves canonical and the social image to absolute URLs, then emits:

The basics — charset, viewport, generator, the font preload, favicons, a sitemap link and an RSS alternate link. Then <title>, the description, <link rel="canonical">, and a robots noindex, nofollow when the page asks for it.

The Open Graph set — og:type (article on a post, website elsewhere), title, description, og:url identical to canonical, site name, locale, and the image with og:image:alt, og:image:width and og:image:height. Real dimensions when the page passes a bundled image; the 1200×630 convention for the default, which is why public/og.jpg has to actually be that size.

On a post, the three article:* tags — published time, modified time when the post has one, and the author’s name.

The Twitter card — summary_large_image, title, description, image, image alt, and twitter:creator when siteData.author.twitter is set.

Two conversions happen here that are easy to get wrong elsewhere. og:locale wants language_TERRITORY, so en-US becomes en_US with a one-line replace. And canonical and og:url are the same value, not two independently built URLs — a pair that disagrees is a classic self-inflicted duplicate-content signal.

The JSON-LD graph

src/js/schema.ts is six builders and a serializer, each typed on its inputs and returning a plain node:

getOrganizationSchema and getWebSiteSchema build the two site-level nodes, linked by a stable @id from organizationId() — https://your-domain/#organization. getSiteSchema composes both, which is what BaseHead calls on every page, so the publisher wiring lives in one module rather than in the layout.

getArticleSchema builds a BlogPosting with its author, image and publisher reference. getBreadcrumbSchema turns an ordered trail into a BreadcrumbList with 1-indexed positions. Pages pass either through the schema prop, which BaseHead merges into the same graph:

const jsonLd = serializeJsonLd([...siteNodes, ...(schema ?? [])]);

A single node inlines directly; several are wrapped in @graph. The blog post page passes two — a BlogPosting and a breadcrumb — and every info page passes a breadcrumb, so the crumb trail in the markup and the one in the structured data always agree.

serializeJsonLd does one security-relevant thing: it escapes < to <, so a value containing </script> cannot break out of the tag it is inlined in. That is the reason to have a serializer at all rather than calling JSON.stringify at the call site, and schema.test.ts asserts it.

There is one honest placeholder. The Organization’s logo currently points at the default OG image, marked with a ponytail: note — swap it for a real brand logo before launch. And sameAs is empty, which is deliberate: those URLs are what disambiguate your organization from every other one with the same name, and a guessed profile URL is worse than none.

The four endpoints

robots.txt allows everything and appends the Sitemap: line built from site. It is an endpoint rather than a file in public/ for precisely that reason — a hard-coded domain in a static file drifts the day you move.

llms.txt is a curated map in the llmstxt.org shape: the core pages, then the index hubs, then the blog and its feed. It is hand-shaped and says so — which pages are worth an AI retrieval system’s attention is an editorial judgement, not a fact about the file tree, and this is a map rather than a second sitemap.

rss.xml is hand-rolled from the blog collection, escaping its own five XML entities. sitemap-index.xml comes from @astrojs/sitemap with the filter that drops the dev catalog, the 404 and the two account demos — so the sitemap and the noindex tags on those pages cannot contradict each other.

Per-page SEO

A page owns its own tags through BaseLayout’s props. The record pages are a good example of doing that from data rather than by hand:

<BaseLayout
  title={`${car.ref} — ${car.title}`}
  description={`${car.title} · ${car.mileage.toLocaleString("en-GB")} mi · ${car.fuel} · …`}
>

Sixty-nine unique titles and descriptions, none written twice, none of them able to drift from the record they describe. That is the pattern worth following for any new route family: derive the metadata from the same data the page renders.

NEXT STEPForms