XML Sitemaps and Robots Directives: A Before-and-After Fix for Small Businesses

An XML sitemap helps search engines discover the URLs you want them to know about. robots.txt controls crawl access, while robots meta directives can control indexing and search-result presentation. Mixing up those jobs can leave an important WordPress page missing from Google even when it appears in your sitemap. This guide shows how to diagnose the conflict, fix it safely, and verify the result.

Written by Founder of ScanMySEO
Published Updated Reading time8 min read

The three controls are not interchangeable

The quickest way to troubleshoot this topic is to separate discovery, crawling and indexing. A sitemap is mainly a discovery signal: it tells search engines which URLs you consider important. Google says a sitemap can help it crawl a site more efficiently, but inclusion does not guarantee that a URL will be crawled or indexed. Google's sitemap documentation explains that distinction.

robots.txt has a different job. It tells compliant crawlers which URLs they may request. Google explicitly warns that it is not a reliable way to keep a web page out of Search: a blocked URL can still be discovered from links and may appear without Google having crawled its content. For deliberate index exclusion, Google recommends a noindex rule or access control such as authentication, depending on the goal. See Google's robots.txt guidance.

ControlWhat it is forWhat it does not guarantee
XML sitemapHelps search engines discover preferred URLs and understand useful metadata such as a meaningful lastmod.It does not force crawling, indexing or ranking.
robots.txtControls whether a compliant crawler may request a URL or path.It does not reliably remove a URL from search results.
Robots meta tag / X-Robots-TagControls indexing and some search-result presentation for content a crawler can access.A crawler cannot act on a page-level noindex rule it is prevented from fetching.

What a healthy small-business setup looks like

For a typical small WordPress site, the aim is simple: important pages should be internally linked, crawlable, indexable and represented by their preferred canonical URLs. Your sitemap should support that setup, not compensate for a site whose navigation or directives contradict it.

Google recommends including the URLs you want to see in search results and, where duplicate versions exist, listing the preferred canonical version rather than every variation. Google also notes that most content management systems can generate sitemaps automatically. WordPress core has included XML sitemap functionality since version 5.5, with the default sitemap index available at /wp-sitemap.xml, although an SEO plugin may provide or replace the sitemap you actually use. WordPress documents the core sitemap behaviour.

For a deeper explanation of what belongs in a sitemap, see ScanMySEO's XML sitemaps guide. The rest of this article focuses on the diagnostic problem: what to do when sitemap inclusion, crawl rules and indexing directives disagree.

Diagnose the problem in the right order

When an important page is missing from Google, do not start by repeatedly resubmitting the sitemap. Check the page itself and the controls around it in this order:

  1. Confirm the preferred URL. Open the live page and make sure you are inspecting the URL you actually want indexed, not a redirect, parameter version, duplicate or outdated address.
  2. Check whether the page is in the sitemap. If it is an important canonical page, inclusion is normally sensible. If it is missing, first confirm whether WordPress core or an SEO plugin is responsible for generating the sitemap before manually editing anything.
  3. Check crawl access in robots.txt. Visit https://yourdomain.example/robots.txt and look for rules that match the page's path. A broad directory rule can block an important page even if that page appears in the sitemap.
  4. Check page-level indexing directives. Inspect the HTML for a robots meta tag and, where relevant, the HTTP response headers for an X-Robots-Tag. A noindex rule is a direct reason for Google not to index the page once Google can crawl and process it.
  5. Check WordPress search visibility. On a production site, verify the setting under Settings → Reading → Search Engine Visibility. WordPress documents that the “Discourage search engines from indexing this site” setting can output a noindex,nofollow robots meta directive. That setting is a request to search engines, not a privacy or access-control feature. WordPress explains the setting and its effect.
  6. Use Google Search Console for the affected URL. URL Inspection shows what Google knows about a specific page and can test a live URL. For broader patterns, use the Page indexing report; the old “Coverage” wording is no longer the clearest name for the current report. Google's URL Inspection documentation explains the current workflow.

If you need a deeper robots-specific checklist, use ScanMySEO's robots.txt mistakes guide.

A before-and-after WordPress example

Consider a hypothetical small business that has launched a new page at https://www.example.com/services/boiler-repair/. The owner expects it to appear in Google because the page is in the XML sitemap, but URL Inspection reports that crawling is not allowed.

Before: the signals conflict

The sitemap contains the new service URL:

<url>
  <loc>https://www.example.com/services/boiler-repair/</loc>
</url>

But robots.txt contains a broad rule left over from development:

User-agent: *
Disallow: /services/

Sitemap: https://www.example.com/wp-sitemap.xml

Those instructions do not cancel each other out. The sitemap tells Google that the page exists; the Disallow rule tells Googlebot not to crawl URLs under /services/. Leaving the page in the sitemap does not override that block.

There is another common trap: adding <meta name="robots" content="noindex"> to a page while also blocking that page in robots.txt. Google can only read a page-level robots directive if it is allowed to fetch the page. Google therefore advises against relying on a combination where the crawler is blocked from seeing the noindex rule. Google's robots meta documentation states that these rules are followed only when crawlers can access the page.

After: discovery, crawling and indexing agree

After confirming that the service page is public, useful and intended for Google Search, the site owner removes the unintended block on the /services/ directory. The sitemap continues to list the preferred service URL. The page does not need an explicit index,follow tag: Google documents unrestricted indexing and serving as the default, so adding that tag usually does not solve an indexing problem.

A simplified robots.txt might now look like this:

User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php

Sitemap: https://www.example.com/wp-sitemap.xml

This is an example, not a universal WordPress template. Keep any site-specific rules that have a clear purpose, and check whether your hosting platform or SEO plugin generates a virtual robots.txt file before editing or replacing it.

The sitemap should then contain the canonical URLs the business actually wants indexed. Google says sitemap submission is only a hint and does not guarantee that Google will download the sitemap, crawl every listed URL or index those URLs. The useful reason to submit a sitemap in Search Console is that the Sitemaps report can show whether Google processed it and encountered errors. Google's build-and-submit guidance makes this explicit.

Use the control that matches your actual goal

Your goalSitemaprobots.txtRobots meta / access control
Have an important public page discovered and eligible for indexingInclude the preferred canonical URL.Allow crawling.Do not apply noindex. An explicit index,follow is normally unnecessary.
Keep a public page out of Google SearchNormally exclude it.Allow crawling so Google can see the indexing rule.Use noindex where appropriate.
Reduce crawling of a low-value URL patternDo not use the sitemap to promote those URLs.A targeted disallow may be appropriate after checking consequences.Do not depend on page-level noindex if crawling is blocked.
Keep private or sensitive content privateDo not list it.Do not treat robots rules as security.Use authentication or another genuine access-control method.

Five mistakes worth fixing first

  • Assuming “in the sitemap” means “will be indexed”. It does not. Indexing depends on crawl access, page quality, canonicalisation and other signals.
  • Using robots.txt to deindex a page. A disallowed URL can still be known to Google. Use the control that matches the goal.
  • Blocking a page that carries noindex. The crawler may be unable to read the directive you expect it to obey.
  • Listing duplicate, redirected or non-preferred URLs as if they were primary pages. Google recommends using preferred canonical URLs in sitemaps.
  • Trying to fix indexing with index,follow. The absence of restrictions is already the default for Google; the useful task is finding the actual conflicting rule or page-level problem.

Verify the fix instead of assuming it worked

After changing a sitemap or robots directive, verify each layer separately:

  1. Reload the live files. Confirm that the sitemap contains the intended canonical URL and that robots.txt no longer blocks it.
  2. Inspect the rendered page and response. Confirm that an unwanted robots meta tag or X-Robots-Tag is gone. If you intentionally use noindex, confirm Google can still crawl the page to see it.
  3. Run URL Inspection. Use the live test to check crawlability, then request indexing if this is an important changed page. A request is not a guarantee of indexing.
  4. Check the Sitemaps report. Make sure Google can fetch and process the submitted sitemap. For site-wide patterns, review the Page indexing report rather than treating every “not indexed” URL as an error; duplicate, redirected or intentionally excluded URLs can be valid outcomes.
  5. Recheck after Google recrawls. Search Console's indexed data reflects Google's last processing of the page, so a live fix and the indexed result may not update at exactly the same time.

For most small businesses, the practical rule is straightforward: keep important pages internally linked, list their preferred URLs in the sitemap, allow Google to crawl them, and avoid accidental noindex directives. Use robots.txt for crawl management, not as a substitute for indexing controls. When those signals agree, troubleshooting becomes much easier because each control is doing one clear job.

Hansel McKoy

Hansel McKoy is the founder of ScanMySEO and a technical SEO specialist with more than 10 years of experience across agency, in-house, public-sector, and founder-led roles.

Founder of ScanMySEO


Get More Out of ScanMySEO