---
title: Help agents find your markdown
description: Serve each page at its own .md URL, declare the markdown in your page head, add frontmatter, publish /sitemap.md and answer missing pages with a markdown 404.
canonical_url: https://rebilder.com/docs/discovery
last_updated: "2026-09-28"
---

# [Help agents find your markdown](https://rebilder.com/docs/discovery)

> Serve each page at its own `.md` URL, declare the markdown in your page head, add frontmatter, publish `/sitemap.md` and answer missing pages with a markdown 404.

- **Updated:** 2026-09-28
- **Author:** Rebilder
- **Section:** Concepts
- **Description:** Serve each page at its own .md URL, declare the markdown in your page head, add frontmatter, publish /sitemap.md and answer missing pages with a markdown 404.
- **Publisher:** Rebilder

## What each option adds

Content negotiation serves agents that ask for markdown. These options help the ones that do not: a person pasting a link into a chat, a tool with no `Accept` header, an agent looking for a list of your pages. All of them are off until you turn them on.

| Option | What agents get |
| --- | --- |
| `markdownUrls: true` | Each page’s markdown at its URL plus `.md`, for anyone who asks. |
| `markdownAlternateTypes` / `markdownAlternate` | A `<link rel="alternate" type="text/markdown">` in your page head. |
| `frontmatter: true` | YAML frontmatter at the top of the markdown: title, description, canonical URL and last-updated date. |
| `createSitemapMdRouteHandler` / `generateSitemapMd` | A `/sitemap.md` that lists your pages as markdown links. |
| `notFound` | A real 404 with a short markdown body and your discovery links, instead of your HTML error page. |

gateway-config.ts

```
import type { GatewayConfig } from '@rebilder/gateway'
import { sources } from './sources'

export const gatewayConfig: GatewayConfig = {
  storeId: 'my-site',
  sources,
  // /services/bike-fitting.md returns the markdown of /services/bike-fitting.
  markdownUrls: true,
  // title, description, canonical_url and last_updated from your source.
  frontmatter: true,
  // A markdown 404 for agents that ask for a page you do not have.
  notFound: {
    links: [
      { title: 'Sitemap', url: '/sitemap.md' },
      { title: 'llms.txt', url: '/llms.txt' },
    ],
    // Next.js and Node only: every page under /services/ has a source,
    // so a /services/ URL with no source is not a page.
    isMissing: (url) => url.pathname.startsWith('/services/'),
  },
}
```

Each one answers a check in Vercel’s agent-readability audit. See [Use Rebilder with @vercel/agent-readability](/docs/vercel-agent-readability) for the full list.

## .md URLs

With `markdownUrls: true`, `GET /services/bike-fitting.md` returns exactly what `/services/bike-fitting` returns to `Accept: text/markdown`, with the same `ETag`. The response carries `Link: <https://example.com/services/bike-fitting>; rel="canonical"`, so search engines credit the page, not the `.md` copy.

- `/index.md` is your homepage. A query string carries over to the page URL.
- The `.md` URL has one version, so browsers and crawlers get the markdown there too. Your page URL is unchanged: crawlers still always get its HTML.
- The gateway answers a `.md` URL only when a source answers the page. Other `.md` requests reach your site untouched, so a real file such as `/README.md` keeps serving. If a real file sits at the `.md` address of a sourced page, the source wins; narrow that source or your `match` router to hand the path back.
- GET and HEAD only. Access rules apply to agents as they do on the page URL.
- Each request is recorded as a `markdown` event for the `.md` URL.

## The markdown link in your page head

On Next.js, add the markdown type to your page metadata. `markdownAlternateTypes` returns it only when your config’s `match` router confirms a source for the URL, and `undefined` otherwise, so a page never advertises markdown it cannot serve.

app/services/[slug]/page.tsx

```
// app/services/[slug]/page.tsx
import type { Metadata } from 'next'
import { markdownAlternateTypes } from '@rebilder/gateway/next'
import { gatewayConfig } from '../../../lib/gateway-config'

export async function generateMetadata({ params }: { params: Promise<{ slug: string }> }): Promise<Metadata> {
  const { slug } = await params
  const url = `https://example.com/services/${slug}`
  return {
    alternates: { canonical: url, types: markdownAlternateTypes(gatewayConfig, url) },
  }
}
```

In other frameworks, call `markdownAlternate(gatewayConfig, url)` from `@rebilder/gateway`. It returns `{ rel, type, href }` or `null`. Render the tag in your page head:

HTML

```
<link rel="alternate" type="text/markdown" href="https://example.com/services/bike-fitting">
```

The `href` is the page URL, where an agent that sends `Accept: text/markdown` gets the markdown. That is the same URL the gateway’s `Link: rel="alternate"` header names on your HTML response, so the page makes one claim in both places.

## Frontmatter

With `frontmatter: true`, each markdown response starts with a YAML block. The values come from the source that answered and from the URL:

response

```
---
title: Bike fitting
description: A 90-minute fit on your own bike, with a written report.
canonical_url: https://example.com/services/bike-fitting
last_updated: "2026-09-28"
---

# [Bike fitting](https://example.com/services/bike-fitting)

> A 90-minute fit on your own bike, with a written report.
```

| Source | Fields |
| --- | --- |
| `document` | `title`, `description` from a text fact labelled `Description` (or else `summary`, or else a fact labelled `Summary`), `canonical_url`, `last_updated` from `updated` |
| `product` | `title`, `description`, `canonical_url`, `last_updated` from `updated` |
| `collection` | `title`, `canonical_url`, `last_updated` from `updated` |
| `policies` | `title` when the URL has one policy, `canonical_url` |
| `catalog` | `canonical_url` |

- A field your source does not have is left out. The gateway never fills in a date, so give pages an `updated` value to get `last_updated`.
- An invalid date or a URL that is not `http(s)` is dropped, not corrected.
- A value is written as it is when YAML reads it back unchanged, and quoted otherwise, so source text cannot end the block early.
- The block is capped at 2 KB. A longer description is dropped first.
- The block comes before the rest of the markdown, after the byte budget. It moves every fact down by its own length, so tools that measure where facts appear, including the Agent Readability Score, can report different numbers with it on.

## /sitemap.md

A markdown list of your pages that an agent can read without parsing XML. Like llms.txt, it lists what your `catalog`, `collection` and `policies` sources return for your base URL, then the sections you add:

app/sitemap.md/route.ts

```
// app/sitemap.md/route.ts
import { createSitemapMdRouteHandler } from '@rebilder/gateway/next'
import { gatewayConfig } from '../../lib/gateway-config'

export const GET = createSitemapMdRouteHandler(gatewayConfig, {
  baseUrl: 'https://example.com',
  sections: [
    { title: 'Services', links: [{ title: 'Bike fitting', url: 'https://example.com/services/bike-fitting' }] },
  ],
})
```

The response is `text/markdown; charset=utf-8`. List page URLs rather than `.md` URLs, so each page appears once. On other frameworks, serve the string from `generateSitemapMd(gatewayConfig, options)`.

## Markdown 404

With `notFound` set, an agent that asks for a page you do not have gets a 404 status with this body instead of your HTML error page:

response

```
# Page not found

There is no page at this address.

- [Sitemap](https://example.com/sitemap.md)
- [llms.txt](https://example.com/llms.txt)
```

| Adapter | How it knows the page is missing |
| --- | --- |
| `@rebilder/gateway/fetch`, `@rebilder/gateway/edge` | Your site returns 404 or 410. The gateway keeps your status and replaces the body for agents. |
| `@rebilder/gateway/next`, `@rebilder/gateway/node` | These run before your routes, so they ask `notFound.isMissing(url)`. Return `true` only for URLs your site does not serve. |

- Only markdown requests that no source answered are affected. People and search crawlers get your own error page.
- Relative link URLs resolve against the request’s origin. List only files you serve.
- The response carries `Vary: Accept` and `X-Robots-Tag: noindex`.
- The request is recorded as the same Agent Miss as before, so your Console numbers do not change.

## Dates in sitemap.xml

The gateway does not write your `sitemap.xml`. Add `<lastmod>` from the date each page states about itself, such as the `updated` value you give the gateway. Leave it off pages that state no date. A build time changes on every deploy whether or not the page did, and crawlers learn to ignore it.

## Check it

terminal

```
curl https://example.com/services/bike-fitting.md
curl -s https://example.com/services/bike-fitting | grep -o '<link rel="alternate" type="text/markdown"[^>]*>'
curl -H "Accept: text/markdown" https://example.com/services/no-such-page
npx @vercel/agent-readability audit https://example.com
```

The first command returns markdown. The second prints the link tag. The third returns the markdown 404. The audit checks all of them on a sample of pages from your sitemap.