Review a source before collection
Document the source and access conditions.
Keep the right fields.
Workflow field map| Field / step | Input | Rule |
|---|---|---|
| Crawl rules | Source robots.txt | A crawler-control file is not an access license. |
| Terms and access | Applicable source terms and authorization | Review the actual intended collection. |
| Data use | Fields, retention and publication | Consider the relevant jurisdiction and data type. |
Read a crawl rule in its own scope.
Illustrative robots.txt file for a hypothetical store. The rule below asks matching crawlers to avoid /account/ and /checkout/. It does not grant a license to reuse the rest of the site.
User-agent: *
Disallow: /account/
Disallow: /checkout/Crawling, indexing and authorization are separate checks.
RFC 9309 defines crawler rules, not access authorization. Google also distinguishes crawl blocking from keeping a URL out of its index.
Scroll to compare →
| Keep | Record | Decision |
|---|---|---|
| Crawl rule | Rules for the relevant host, path and user agent | Check the actual source file and retain when it was reviewed. |
| Search indexing | A URL may be discovered even when crawling is blocked | Site owners should use an appropriate indexing control or authentication for their objective. |
| Access control | Login requirements, account permissions and technical restrictions | A robots.txt file is not a mechanism for securing private content. |
| Reuse decision | Terms, intended purpose, fields and applicable jurisdiction | Public accessibility alone does not settle copyright, contractual or privacy questions. |
Keep a review record with the collection request.
This is an operational checklist, not a legal conclusion about a particular collection. Escalate unresolved legal questions to someone qualified in the relevant jurisdiction.
Scroll to compare →
| Keep | Record | Decision |
|---|---|---|
| Source and scope | Exact hosts, public paths, required fields and intended use | Exclude private account data from a public product-data request. |
| Evidence checked | Dated crawl rules, applicable terms and any owner authorization | Keep enough context to revisit the decision when the source changes. |
| Collection limits | Request bounds, rate expectations and failure handling | Stop for unresolved access restrictions rather than treating a failure as permission. |
| Downstream handling | Storage period, recipients and publication rights | Review images, descriptions and personal data separately when included. |
Content reviewed 2026-09-15.
Before you use the output
Google describes robots.txt as a crawl-management mechanism, not a way to secure private information. Access and reuse questions need separate review for the specific source and jurisdiction.
Check one record first.
Start with the smallest useful scope. Compare the returned record with its source, confirm variant identity and inspect missing fields before applying the mapping to a larger dataset.
Keep extraction warnings and the retrieval timestamp with downstream output. When a required field is missing, leave it unresolved or return to the source; do not fill it from an assumption.