I Serve a Markdown Twin of My Site to Robots
More of my first visits now come from AI agents than I used to think mattered. Recruiters run assistants to screen candidates before a human opens a tab. LLM crawlers index personal sites to answer natural language queries. If one of those systems hits a JavaScript-rendered portfolio full of animation wrappers, split-text spans, and sticky nav markup, it does not get my work. It gets soup.
I decided to fix that at the protocol level, not with a separate page no one maintains.
One URL, Two Representations
HTTP has had content negotiation since 1996. A client that wants a specific format sends an Accept header, and a server that understands the request can respond accordingly. Browsers ask for HTML. Agents can ask for anything.
I added a check in my site's middleware that fires before any page renders. If an incoming GET request carries text/markdown in its Accept header and is not already hitting an API route or a static asset, the request gets rewritten to a parallel /md/* route that returns clean Markdown generated from the same source files.
const accept = request.headers.get("accept") || "";
if (
request.method === "GET" &&
accept.includes("text/markdown") &&
!pathname.startsWith("/md") &&
!pathname.startsWith("/api") &&
!pathname.includes(".")
) {
const url = request.nextUrl.clone();
url.pathname = pathname === "/" ? "/md" : `/md${pathname.replace(/\/$/, "")}`;
return NextResponse.rewrite(url);
}
The rewrite is transparent to the caller. The URL never changes in the response. The agent asked for /blog/serving-markdown-to-robots and got exactly that, just in the format it can actually use.
Why Generated, Not Maintained
The twin is not a hand-written Markdown file I keep in sync. It is generated at request time from the same metadata and MDX that render the HTML page. The case study content, blog posts, and profile information all share a single source of truth.
This matters because the honesty property is structural, not disciplinary. I cannot let the two representations drift apart by forgetting to update one of them. There is nothing to forget. A change to the content source propagates to both the human page and the machine page automatically.
For case study pages, the generator strips JSX components down to their readable core: a DataCallout becomes a bullet, a PullQuote becomes a blockquote, an image component becomes an italicized caption with an absolute URL. The result is clean, structured text with no layout scaffolding in the way.
The Hygiene Details That Make It Safe
Serving a parallel representation creates two risks: search engines indexing the twin as a separate page (duplicate content), and the twin accidentally outranking the canonical page it mirrors. Both are handled at the response level, not in any CMS or meta tag that could drift.
Every Markdown twin response carries two headers. First, a Link header pointing at the canonical HTML URL:
Link: <https://www.lokeshsaini.com/blog/serving-markdown-to-robots>; rel="canonical"
Second, X-Robots-Tag: noindex, which tells crawlers not to index the twin at all. An agent reading the twin for content will ignore that header. A search crawler respects it.
There is also an llms.txt at the root of the site. Agents that check for a structured profile before crawling individual pages find a plain-text document with a summary, my availability, key links, and pointers to the full Markdown version and a JSON API endpoint. It is the equivalent of a robots.txt for language models: a known location, a known format, no JavaScript required.
The Easter Egg
I built a human-facing version of the same view. There is a lens on the site that renders any page as its Markdown twin in a terminal-style reader. You can flip to it on any case study or blog post and see exactly what the robots see: raw structure, no visual treatment, no animation.
It started as a debugging tool. I kept it because it is useful. If the machine-readable version of a page is confusing or thin, that is a signal the content itself has a problem, not a rendering problem.
What an Agent Can Do With This
An agent that hits the site with Accept: text/markdown can answer three questions from a single request: who is this person, what have they shipped, are they available. The home twin lists case studies with roles and tags, experience with metrics, and an explicit availability line in the header. A case study twin contains the full narrative, stripped of decoration but complete in substance.
That is enough context for a recruiting assistant to make a meaningful recommendation. It is enough for an LLM to give an accurate answer when someone asks it about me.
Treat Agents as an Audience
The standard advice for AI discoverability is semantic HTML and good meta descriptions. That advice assumes agents parse your HTML well. Many do not. Some execute JavaScript, some do not. Some respect robots.txt, some do not. Content negotiation sidesteps all of that: the agent declares what it wants, the server responds with exactly that.
The right frame is not "how do I make my page legible to bots." It is "agents are a first-class audience with their own representation." Give them one. Negotiate it at the protocol level. Generate it from the same source of truth as the human version.
The alternative is hoping your JavaScript renders fast enough and your DOM is clean enough. That is not a strategy. That is a wish.