Sitemaps and robots.txt sit next to each other in every SEO checklist, so people often treat them as two halves of the same tool. They are not. A sitemap helps Google discover your URLs; robots.txt tells crawlers where they may go. And the point that causes the most damage when it is misunderstood: robots.txt is not a way to keep a page out of Google.
Two files, two different jobs
An XML sitemap is a list of the URLs you want Google to know about. It helps discovery, especially for new pages or pages with few internal links. It does not force indexing, and it does not guarantee that Google will crawl every entry.
robots.txt is a plain text file at the root of your host, for example https://example.com/robots.txt, that tells crawlers which paths they may request. It manages crawling, not indexing. Google says it directly: robots.txt is not a mechanism for keeping a web page out of Google, and to keep a page out you should block indexing with noindex or password-protect the page.
Before you write either file, check what you already have. Google notes that if you use a CMS such as WordPress, Wix or Blogger, it has likely already made a sitemap available. In my experience, plenty of Arabic blogs run two sitemaps by accident: one from the CMS and one from a plugin. Pick one source.
Building both files step by step
- Find or create your sitemap. Open
/sitemap.xmlor check your CMS or SEO plugin settings. If nothing exists, generate one from your CMS or a script. - Respect the limits. Google limits a single sitemap to 50MB uncompressed or 50,000 URLs. Beyond that, split the list into several sitemaps and list them in a sitemap index file. Splitting by language (
sitemap-ar.xmlandsitemap-en.xml) also makes Search Console reporting easier to read. - Use absolute URLs. Google asks for fully qualified, absolute URLs, meaning
https://example.com/ar/..., never/ar/.... - Encode Arabic URLs properly. Save the file as UTF-8 and write Arabic paths in their percent-encoded form, the same form Google’s URL guidance recommends for non-ASCII characters in links.
- List only canonical, indexable pages. Leave out redirected URLs, pages marked noindex and duplicate versions.
- Set lastmod honestly, and skip the rest. Google ignores
priorityandchangefreq. It useslastmodonly if the value is consistently and verifiably accurate, so change it when the content really changes, not on every page every night. - Submit it. Add the sitemap in Search Console’s Sitemaps report, where you can see when Googlebot accessed it, and add a
Sitemap:line to robots.txt as well. - Write robots.txt carefully. Google supports four fields: user-agent, allow, disallow and sitemap. Other fields such as crawl-delay are not supported. Keep the file well under Google’s 500 KiB limit.
- Never block rendering resources. Do not disallow the CSS, JavaScript or image folders your pages need. Google renders pages using a recent version of Chrome, and blocked files can leave it looking at a broken layout.
- Remove pages the right way. For anything that must stay out of Google, use noindex or put it behind a password. If you disallow a page in robots.txt, Google cannot fetch it, so it would never see a noindex on it.
- Allow for caching. Google generally caches robots.txt for up to 24 hours, so an edit may take up to a day to be picked up.
Example: a bilingual site
Here is a sitemap entry for an Arabic article whose slug is تحسين-سرعة-الموقع, written in its encoded form:
<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<url>
<loc>https://example.com/ar/%D8%AA%D8%AD%D8%B3%D9%8A%D9%86-%D8%B3%D8%B1%D8%B9%D8%A9-%D8%A7%D9%84%D9%85%D9%88%D9%82%D8%B9</loc>
<lastmod>2026-02-01</lastmod>
</url>
<url>
<loc>https://example.com/en/site-speed</loc>
<lastmod>2026-01-28</lastmod>
</url>
</urlset>
There is no priority or changefreq, because Google ignores them. Each lastmod is the date the article’s text last changed.
And a simple robots.txt for a WordPress site with Arabic and English sections:
User-agent: * Disallow: /wp-admin/ Allow: /wp-admin/admin-ajax.php Sitemap: https://example.com/sitemap-ar.xml Sitemap: https://example.com/sitemap-en.xml
It blocks the admin area, keeps one file the front end commonly calls reachable, and points to both language sitemaps. It does not block themes, scripts or uploads. Short is good here: the fewer rules you write, the fewer you can get wrong.
Checklist
| Check | How | Tool |
|---|---|---|
| Only one sitemap source | Confirm the CMS and plugins are not both generating sitemaps | CMS settings |
| Within limits | Count URLs and file size per sitemap, split above 50,000 URLs or 50MB | Sitemap file and a text editor |
| URLs are absolute and encoded | Open the file and review several Arabic entries | Browser |
| Only canonical, indexable URLs | Spot check entries for redirects or noindex | Search Console URL Inspection |
| lastmod is accurate | Compare a few dates with real edit dates | CMS post history |
| Sitemap submitted and read | Check status and last read date | Search Console Sitemaps report |
| robots.txt uses supported fields | Remove crawl-delay and unsupported lines | Text editor |
| Rendering files are not blocked | Confirm no disallow rule covers CSS, JS or image folders | robots.txt review and URL Inspection |
| Private pages are protected | Use noindex or a password, not disallow | Page settings and server config |
Common mistakes
- Using Disallow to “deindex” a page. It stops crawling, not indexing, and it hides any noindex you add.
- Blocking CSS or JavaScript. It saves nothing and can leave Google rendering a broken page.
- Listing redirected, noindex or non-canonical URLs in the sitemap. That sends mixed signals about which URL you want.
- Updating lastmod on every page daily. Google uses lastmod only when it is consistently accurate, so a fake date makes it useless.
- Relying on crawl-delay for Googlebot. Google does not support that field.
- Expecting a sitemap to guarantee indexing. It helps discovery. Whether a page is indexed is still Google’s decision.
Sources
- Google Search Central, documentation on building and submitting a sitemap, consulted 4 February 2026, https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap
- Google Search Central, introduction to robots.txt, consulted 4 February 2026, https://developers.google.com/search/docs/crawling-indexing/robots/intro
- Google Search Central, documentation on how Google interprets the robots.txt specification, consulted 4 February 2026, https://developers.google.com/search/docs/crawling-indexing/robots/robots_txt
- Google Search Central, URL structure best practices (percent encoding of non-ASCII characters), consulted 4 February 2026, https://developers.google.com/search/docs/crawling-indexing/url-structure
- Google Search Central, documentation on how Google Search works (rendering with a recent version of Chrome), consulted 4 February 2026, https://developers.google.com/search/docs/fundamentals/how-search-works