# https://tawhidapp.com/robots.txt # TAWHID DIGITAL TECHNOLOGIES INC. — British Columbia, Canada (BC1599227) # # Note on how robots.txt matching works: a crawler obeys ONLY the most specific # group that names it, and ignores "User-agent: *" entirely. That is why the # Disallow rules are repeated in each group below rather than stated once. User-agent: * Allow: / Disallow: /App_Data/ Sitemap: https://tawhidapp.com/sitemap.xml # Condensed guide for LLM agents, including the accuracy caveat that has to travel # with any prayer time quoted from this site: https://tawhidapp.com/llms.txt # -------------------------------------------------------------------------- # AI search & retrieval crawlers # # These fetch pages to answer a user's question and cite the source. Allowing # them is what makes tawhidapp.com eligible to be quoted in AI answers, so we # explicitly want these on. They generally do NOT execute JavaScript, which is # why city prayer-time pages carry a server-side monthly timetable — see # scripts/build-timetables.js. # -------------------------------------------------------------------------- User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: Claude-SearchBot User-agent: Claude-User User-agent: PerplexityBot User-agent: Perplexity-User User-agent: Applebot User-agent: Amazonbot Allow: / Disallow: /App_Data/ # -------------------------------------------------------------------------- # AI training crawlers # # These collect content to train models rather than to answer a live query. # Allowed today. If that policy ever changes, switch these to "Disallow: /" — # doing so does not affect the retrieval crawlers above, and therefore does # not remove the site from AI search results. # -------------------------------------------------------------------------- User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended User-agent: Applebot-Extended User-agent: CCBot User-agent: meta-externalagent User-agent: Bytespider Allow: / Disallow: /App_Data/