User-agent: * Allow: / Disallow: /admin/ Disallow: /admin Disallow: /500 Disallow: /404 # Block authenticated app pages from crawling Disallow: /home Disallow: /dashboard Disallow: /messages/ Disallow: /user Disallow: /phases/ Disallow: /call/ Disallow: /billing_info Disallow: /payment/ Disallow: /payment_method Disallow: /phased_project/ Disallow: /unphased_project/ Disallow: /phased_project_call/ Disallow: /network Disallow: /questions Disallow: /search Disallow: /offer/ Disallow: /contact_me/ Disallow: /done Disallow: /forgot_password Disallow: /reset_password Disallow: /check_email Disallow: /verify_email Disallow: /signup/entrepreneur Disallow: /signup/consultant_mentor Disallow: /startupmap/admin Disallow: /badge/ Disallow: /widget/ Disallow: /*?search= # Full text search shards for /journal-officiel. 5,033 machine readable files # of official JO text, fetched by the search box only. Never a landing page, # never in a sitemap, and a duplicate surface if crawled. Disallow: /data/jo-fts/ # BULK MACHINE READABLE CORPORA. These exist so the site's own tools can load # them: /ecosystem and /startupmap read entities.json, the nomenclature search # reads nae-related.json, the JO hub reads jo-ref-index.json. A person using # those pages fetches each file once. Only a harvester wants all of them. # # THESE THREE RULES ARE DELIBERATELY *NOT* REPEATED IN THE NAMED GROUPS BELOW. # That is the exact opposite of the convention documented at the end of this # group, and it is intentional, so read this before "fixing" it: # # A named group REPLACES this one. By leaving these paths out of the # Googlebot, Bingbot, GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot and # every other named AI group, those crawlers keep FULL access to the data. # They are the high authority citers we want reading it. Only UNNAMED # crawlers are held out, and an unnamed crawler is what a scraper is. # # If you add a path here, do NOT copy it into the named groups. # # The five open CSV datasets are deliberately absent from this list. They are # what the site gives away on purpose and they stay crawlable by everyone: # cnrc-2026.csv, anae-2026.csv, journal-officiel-archive.csv, wilayas-2026.csv # and tarif-douanier-sections-chapitres.csv. # # NOTE THE LIMIT: robots.txt is advisory. It governs crawlers that choose to # obey it. It stops nothing that does not, so it is a signal, not a control. # The enforcement lives in nginx. Disallow: /data/entities.json Disallow: /data/jo-ref-index.json Disallow: /data/nae-related.json # /demo/ IS DELIBERATELY NOT DISALLOWED. DO NOT ADD A RULE FOR IT. # # /demo// holds eight standalone demonstration applications (public/demo/ # in the repo). Every one of them serves # # and that tag, not a rule in this file, is what keeps them out of the index. # # Disallowing them here would BREAK that, not reinforce it. These pages are # linked from /realisations, /tech and the service pages in all three language # editions. A crawler that is forbidden to fetch a linked URL never reads its # noindex, and Google will still list the address it found in those links, with # no title and no snippet, and no way to remove it while the block stands. The # only mechanism that actually drops a LINKED page from the index is letting the # crawler in so it can read the noindex. `follow` is kept on purpose so the # links each demo carries back to upgrowth.dz still count. # # Crawl cost of allowing them: 77 files, about 1.0 MB, against 13 594 # prerendered pages. Nothing here is a duplicate of a site page. # # Each demo folder also ships its own robots.txt in the source tree. Those are # NOT deployed and would be ignored anyway: only the origin root robots.txt is # ever read, so /demo//robots.txt has no effect at all. # A NAMED USER AGENT GROUP REPLACES THE WILDCARD GROUP, IT DOES NOT ADD TO IT. # These two groups used to be "Allow: /" plus one Disallow, which meant # Googlebot and Bingbot ignored every rule in the * group above: /admin, # /dashboard, /payment/, /user, /reset_password, /verify_email, all of it was # crawlable by the only two crawlers that matter. The * list is therefore # repeated verbatim in each named group. robots.txt has no include, so if a # path is added above it must be added here too. User-agent: Googlebot Allow: / Disallow: /admin/ Disallow: /admin Disallow: /500 Disallow: /404 Disallow: /home Disallow: /dashboard Disallow: /messages/ Disallow: /user Disallow: /phases/ Disallow: /call/ Disallow: /billing_info Disallow: /payment/ Disallow: /payment_method Disallow: /phased_project/ Disallow: /unphased_project/ Disallow: /phased_project_call/ Disallow: /network Disallow: /questions Disallow: /search Disallow: /offer/ Disallow: /contact_me/ Disallow: /done Disallow: /forgot_password Disallow: /reset_password Disallow: /check_email Disallow: /verify_email Disallow: /signup/entrepreneur Disallow: /signup/consultant_mentor Disallow: /startupmap/admin Disallow: /badge/ Disallow: /widget/ Disallow: /*?search= Disallow: /data/jo-fts/ User-agent: Bingbot Allow: / Disallow: /admin/ Disallow: /admin Disallow: /500 Disallow: /404 Disallow: /home Disallow: /dashboard Disallow: /messages/ Disallow: /user Disallow: /phases/ Disallow: /call/ Disallow: /billing_info Disallow: /payment/ Disallow: /payment_method Disallow: /phased_project/ Disallow: /unphased_project/ Disallow: /phased_project_call/ Disallow: /network Disallow: /questions Disallow: /search Disallow: /offer/ Disallow: /contact_me/ Disallow: /done Disallow: /forgot_password Disallow: /reset_password Disallow: /check_email Disallow: /verify_email Disallow: /signup/entrepreneur Disallow: /signup/consultant_mentor Disallow: /startupmap/admin Disallow: /badge/ Disallow: /widget/ Disallow: /*?search= Disallow: /data/jo-fts/ # Allow AI Bots (We want to be in the knowledge graph + AI citations) User-agent: GPTBot Allow: / User-agent: CCBot Allow: / User-agent: Google-Extended Allow: / User-agent: ClaudeBot Allow: / User-agent: anthropic-ai Allow: / User-agent: PerplexityBot Allow: / User-agent: Applebot Allow: / User-agent: Applebot-Extended Allow: / User-agent: FacebookBot Allow: / User-agent: Bytespider Allow: / # Newer AI search/grounding crawlers (distinct from the training bots above): # OpenAI search + user fetches, Anthropic Claude search + user fetches, # Perplexity user fetches, Meta AI, DuckDuckGo AI, Mistral user fetches. User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: Claude-SearchBot Allow: / User-agent: Claude-User Allow: / User-agent: Perplexity-User Allow: / User-agent: meta-externalagent Allow: / User-agent: DuckAssistBot Allow: / User-agent: MistralAI-User Allow: / # AI assistant context file (services, prices, canonical URLs) # https://www.upgrowth.dz/llms.txt # Sitemap Sitemap: https://www.upgrowth.dz/sitemap.xml