robots.txt and sitemap.xml for a Small Developer Blog (That Search Engines Can Use)
Search features will not save thin content, but broken discovery setup can hide good content. For a small static developer blog, two files do most of the mechanical work: robots.txt and sitemap.xml. This guide keeps them honest and maintainable—especially on Cloudflare Pages where the publish directory is the whole product.
A sane robots.txt
For an open content site you want indexed, keep robots.txt simple:
User-agent: *
Allow: /
Sitemap: https://negency-lab-pilot.pages.dev/sitemap.xml
Only add Disallow rules for paths that truly should not be crawled (staging folders, private experiments). Do not use robots.txt as a security boundary—sensitive files should not be uploaded. If you later add a preview subdirectory you dislike indexing, disallow that prefix and exclude it from the sitemap.
Avoid huge stacks of bot-specific rules copied from forums. They rot, conflict, and rarely improve rankings. Prefer clarity.
sitemap.xml essentials
A sitemap lists canonical URLs you want discovered. For static sites, a hand-maintained or script-generated XML file is enough:
- Include homepage, About, Privacy, Contact, article index, and each article.
- Omit thank-you pages, duplicate sort orders, and parameter URLs you do not use.
lastmodshould change when content meaningfully updates—not on every deploy of an unchanged file.changefreqandpriorityare soft hints; do not obsess. Consistency beats fake “1.0 priority on everything.”
Validate that every sitemap URL returns 200 and matches the canonical link in the page head. Mixed trailing-slash policies cause duplicate URL noise—pick one and redirect if your host supports it.
Absolute URLs and host choice
Sitemaps need absolute URLs with scheme and host. During a pages.dev pilot, use the pages.dev host everywhere (canonical tags, sitemap, robots Sitemap line). When a custom domain becomes primary, update those strings in one pass and redeploy. Running two conflicting canonical hosts without a clear primary confuses reviewers and crawlers alike.
If you temporarily serve both hosts, pick one canonical and make the other redirect when possible. Pages custom domains usually serve the same content; still set canonical tags explicitly.
Updating when you publish
Every new article should:
- Add a file under
/articles/. - Link from
/articles/index.htmland optionally the homepage. - Append a
<url>entry to sitemap.xml. - Redeploy.
A tiny Python or Node script can regenerate the sitemap from a list of slugs to avoid XML typos. Commit the generator next to the site so the process is repeatable.
Search Console basics
After the site is live, add the property in Google Search Console (URL-prefix is fine for a pilot). Submit the sitemap URL. Inspect a few article URLs. Fix obvious 404s before you worry about impressions. Bing Webmaster Tools is optional but cheap insurance.
Do not expect instant indexing. Original long-form how-tos with clear titles earn crawl attention faster than doorway pages. Internal links from the homepage and article index matter more than fancy schema on day one—though basic Article structured data can come later.
Myths to ignore
- “robots.txt meta keywords boost SEO” — no.
- “Sitemap guarantees ranking” — it only helps discovery.
- “Disallow all then selectively allow” — easy to get wrong; not needed for open blogs.
- “Priority 1.0 makes homepage rank” — ignoring you is free.
Ship clean files, keep URLs stable, and spend remaining energy on the eighth article—not on XML trivia.
Complete example files
Here is a complete starter robots.txt for an open pilot:
User-agent: *
Allow: /
# Keep secrets off the server entirely; robots is not access control.
Sitemap: https://negency-lab-pilot.pages.dev/sitemap.xml
And a trimmed sitemap pattern (repeat url blocks for each article):
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://negency-lab-pilot.pages.dev/</loc>
<lastmod>2026-10-01</lastmod>
</url>
<url>
<loc>https://negency-lab-pilot.pages.dev/articles/robots-txt-sitemap-small-dev-blog.html</loc>
<lastmod>2026-10-01</lastmod>
</url>
</urlset>
Generate this from a list of paths in Python to avoid hand-editing XML errors. Keep the generator in-repo. When you switch primary host to a custom domain, pass the new base URL into the generator and redeploy in the same commit that updates canonical tags.
Where ads.txt fits
ads.txt is not a sitemap, but reviewers and ad systems expect it at the site root once you monetize. Before you have a publisher ID, ship a file that contains only a comment explaining that authorized sellers will be listed after AdSense approval. After approval, replace the comment with the official Google AdSense ads.txt lines from your account. Never invent seller IDs. Host the file over HTTPS at /ads.txt on the same hostname you use in AdSense.
If you operate both pages.dev and a custom domain, ensure ads.txt is reachable on the hostname you claim in the AdSense property. Some publishers keep both in sync during transition.
Internal linking beats fancy SEO plugins
Static sites do not need WordPress SEO plugins. They need: descriptive titles, unique meta descriptions, one H1, sensible H2s, and links between related tutorials. Add a “Related” paragraph at the end of each article with two internal links. Link from the homepage to your best pillar posts. That graph helps users and crawlers more than micro-optimizing priority fields in XML.
Track what you publish in a simple spreadsheet or Markdown table: slug, title, date, word count, sitemap yes/no. When something is missing from the sitemap, you will see it before Google does.
Also ensure ads.txt exists at the site root when you monetize: for AdSense it will list Google’s seller lines. Before approval, a short comment placeholder in the file (or a minimal stub) documents intent without inventing publisher IDs.