Search by

daikazu / robotstxt

daikazu

Dynamically generate robots.txt content based on the current Laravel environment.

Package info

github.com/daikazu/robotstxt

pkg:composer/daikazu/robotstxt

Fund package maintenance!

Daikazu

Statistics

Installs: 286

Dependents: 0

Suggesters: 0

Stars: 0

Open Issues: 0

v1.3.0 2026-09-23 15:37 UTC

This package is auto-updated.

Last update: 2026-09-23 15:43:17 UTC


README

Logo for ROBOTS.TXT

Latest Version on Packagist GitHub Tests Action Status GitHub Code Style Action Status Total Downloads

Dynamic robots.txt for your Laravel app.

A Laravel package for dynamically generating robots.txt files with environment-specific configurations. Control how search engines and AI crawlers interact with your site using modern features like Cloudflare's Content Signals Policy, per-environment rules, and flexible content signal directives.

Perfect for applications that need different robots.txt rules across environments (production, staging, local) or want granular control over AI training, search indexing, and content access.

Features

  • Environment-Specific Configuration - Different robots.txt rules for production, staging, local, etc.
  • Content Signals Support - Implement Cloudflare's Content Signals Policy to control AI training, search indexing, and AI content usage
  • Flexible User-Agent Rules - Define global rules or per-agent directives (disallow, allow, content signals)
  • Host Directive - Specify preferred domain for crawlers
  • Sitemap Management - Automatically include sitemap URLs
  • Custom Text - Add custom content to your robots.txt file
  • Human-Readable Policies - Optional policy comment blocks with custom or default text

Installation

You can install the package via composer:

composer require daikazu/robotstxt

You can publish the config file with:

php artisan vendor:publish --tag="robotstxt-config"

Remove the static public/robots.txt (Required)

New Laravel applications ship with a static public/robots.txt. Web servers serve files in public/ directly, so while that file exists, requests never reach Laravel and this package's output is never shown. Delete it:

rm public/robots.txt

Nginx Configuration (Required for Production)

If you're getting a 404 status (but still seeing content), you need to configure Nginx to pass robots.txt requests to Laravel:

For Laravel Herd: Add this to your site's Nginx config (via Herd UI or ~/.config/herd/Nginx/[site].conf):

location = /robots.txt {
      try_files $uri /index.php?$query_string;
      access_log off;
      log_not_found off;
  }

For Laravel Forge: Forge's default Nginx config includes a location = /robots.txt block that only serves the static file. Replace it with the block above.

For Laravel Vapor: No Nginx changes are needed. Just make sure public/robots.txt is deleted, otherwise it is uploaded as a static asset and served from the CDN.

For custom servers: Add to your server block in your Nginx config file.

Then restart Nginx/Herd.

Usage

After installation, the package automatically registers a route at /robots.txt that serves your dynamically generated robots.txt file.

Basic Configuration

The default configuration file (config/robotstxt.php) provides environment-specific settings. Here's a simple example:

return [
    'environments' => [
        'production' => [
            'paths' => [
                '*' => [
                    'disallow' => ['/admin', '/api'],
                    'allow' => ['/'],
                ],
            ],
            'sitemaps' => [
                'sitemap.xml',
            ],
        ],
    ],
];

This generates:

Sitemap: https://example.com/sitemap.xml

User-agent: *
Disallow: /admin
Disallow: /api
Allow: /

Content Signals (AI & Search Control)

Control how AI crawlers and search engines use your content with Cloudflare's Content Signals Policy:

'production' => [
    // Enable human-readable policy comment block
    'content_signals_policy' => [
        'enabled' => true,
        'custom_policy' => null, // or provide your own HEREDOC text
    ],

    // Global content signals (added to every User-agent group without its own signals)
    'content_signals' => [
        'search'   => true,   // Allow search indexing
        'ai_input' => false,  // Block AI input/RAG
        'ai_train' => false,  // Block AI training
    ],

    'paths' => [
        '*' => [
            'disallow' => [],
            'allow' => ['/'],
        ],
    ],
],

Generates:

# As a condition of accessing this website, you agree to abide by the following
# content signals:
# [Full policy text...]

User-agent: *
Content-Signal: search=yes, ai-input=no, ai-train=no
Allow: /

Per-Agent Content Signals

You can also define content signals for specific user agents. Per-agent signals replace the global signals for that agent; agents without their own signals inherit the global ones:

'paths' => [
    '*' => [
        'disallow' => [],
        'allow' => ['/'],
    ],
    'Googlebot' => [
        'content_signals' => [
            'search'   => true,
            'ai_input' => true,
            'ai_train' => false,
        ],
        'disallow' => ['/private'],
        'allow' => ['/'],
    ],
],

Generates (with the global signals from the previous example):

User-agent: *
Content-Signal: search=yes, ai-input=no, ai-train=no
Allow: /

User-agent: Googlebot
Content-Signal: search=yes, ai-input=yes, ai-train=no
Disallow: /private
Allow: /

To opt an agent out of the global signals without setting any of its own, give it a content_signals block with every value set to null:

'Bingbot' => [
    'content_signals' => ['search' => null, 'ai_input' => null, 'ai_train' => null],
    'allow' => ['/'],
],

Host Directive

Specify the preferred domain for crawlers:

'production' => [
    'host' => 'https://www.example.com',
    // ... other config
],

Generates:

Host: https://www.example.com

Custom Text

Add arbitrary custom content to the end of your robots.txt:

'production' => [
    'custom_text' => <<<'TEXT'
# Custom crawl-delay for specific bots
User-agent: Bingbot
Crawl-delay: 1
TEXT,
    // ... other config
],

Environment-Specific Rules

Define different rules for each environment. The current environment is read from APP_ENV. Any environment that isn't in the config (or has no paths) falls back to blocking all crawlers:

User-agent: *
Disallow: /

So staging, local and preview environments stay out of search engines unless you configure them otherwise.

return [
    'environments' => [
        'production' => [
            'paths' => [
                '*' => [
                    'disallow' => [],
                    'allow' => ['/'],
                ],
            ],
            'content_signals' => [
                'search' => true,
                'ai_input' => false,
                'ai_train' => false,
            ],
        ],
        'staging' => [
            'paths' => [
                '*' => [
                    'disallow' => ['/'],
                ],
            ],
        ],
        'local' => [
            'paths' => [
                '*' => [
                    'disallow' => ['/'],
                ],
            ],
        ],
    ],
];

Content Signal Values

  • true or 'yes' - Permission granted
  • false or 'no' - Permission denied
  • null - No preference specified (signal not included)

Content Signal Types

  • search - Building a search index and providing search results (excludes AI summaries)
  • ai_input - Inputting content into AI models (RAG, grounding, AI Overviews)
  • ai_train - Training or fine-tuning AI models

Complete Configuration Example

return [
    'environments' => [
        'production' => [
            // Content Signals Policy
            'content_signals_policy' => [
                'enabled' => true,
                'custom_policy' => null,
            ],

            // Global Content Signals
            'content_signals' => [
                'search'   => true,
                'ai_input' => false,
                'ai_train' => false,
            ],

            // User-Agent Rules
            'paths' => [
                '*' => [
                    'disallow' => ['/admin', '/api'],
                    'allow' => ['/'],
                ],
                'Googlebot' => [
                    'content_signals' => [
                        'search'   => true,
                        'ai_input' => true,
                        'ai_train' => false,
                    ],
                    'disallow' => [],
                    'allow' => ['/'],
                ],
            ],

            // Sitemaps
            'sitemaps' => [
                'sitemap.xml',
                'sitemap-news.xml',
            ],

            // Host
            'host' => 'https://www.example.com',

            // Custom Text
            'custom_text' => <<<'TEXT'
# Additional custom directives
User-agent: Bingbot
Crawl-delay: 1
TEXT,
        ],
    ],
];

Laravel Boost

If your application uses Laravel Boost, this package ships an AI guideline and a robotstxt-configuration skill, so your coding agent knows how to configure robots.txt correctly. They're picked up when you run:

php artisan boost:install

If Boost is already installed, run php artisan boost:update --discover to add them.

Testing

composer test

Roadmap

  • Crawl-delay Directive - Add native support for crawl-delay configuration per user-agent

Changelog

Please see CHANGELOG for more information on what has changed recently.

Contributing

Issues and pull requests are welcome on GitHub.

Security Vulnerabilities

Please review our security policy on how to report security vulnerabilities.

Credits

License

The MIT License (MIT). Please see License File for more information.