Extraction guide

Review a source before collection

Document the source and access conditions.

01 / INPUTSource URL + intended use
02 / YOUR WORKFLOWReview access → scope → retain decision
03 / OUTPUTA collection review record

Keep the right fields.

Workflow field map
Review a source before collection field map
Field / stepInputRule
Crawl rulesSource robots.txtA crawler-control file is not an access license.
Terms and accessApplicable source terms and authorizationReview the actual intended collection.
Data useFields, retention and publicationConsider the relevant jurisdiction and data type.

Read a crawl rule in its own scope.

Illustrative robots.txt file for a hypothetical store. The rule below asks matching crawlers to avoid /account/ and /checkout/. It does not grant a license to reuse the rest of the site.

Example
User-agent: *
Disallow: /account/
Disallow: /checkout/

Crawling, indexing and authorization are separate checks.

RFC 9309 defines crawler rules, not access authorization. Google also distinguishes crawl blocking from keeping a URL out of its index.

Scroll to compare →

Crawling, indexing and authorization are separate checks.
KeepRecordDecision
Crawl ruleRules for the relevant host, path and user agentCheck the actual source file and retain when it was reviewed.
Search indexingA URL may be discovered even when crawling is blockedSite owners should use an appropriate indexing control or authentication for their objective.
Access controlLogin requirements, account permissions and technical restrictionsA robots.txt file is not a mechanism for securing private content.
Reuse decisionTerms, intended purpose, fields and applicable jurisdictionPublic accessibility alone does not settle copyright, contractual or privacy questions.

Keep a review record with the collection request.

This is an operational checklist, not a legal conclusion about a particular collection. Escalate unresolved legal questions to someone qualified in the relevant jurisdiction.

Scroll to compare →

Keep a review record with the collection request.
KeepRecordDecision
Source and scopeExact hosts, public paths, required fields and intended useExclude private account data from a public product-data request.
Evidence checkedDated crawl rules, applicable terms and any owner authorizationKeep enough context to revisit the decision when the source changes.
Collection limitsRequest bounds, rate expectations and failure handlingStop for unresolved access restrictions rather than treating a failure as permission.
Downstream handlingStorage period, recipients and publication rightsReview images, descriptions and personal data separately when included.

Content reviewed 2026-09-15.

Before you use the output

Google describes robots.txt as a crawl-management mechanism, not a way to secure private information. Access and reuse questions need separate review for the specific source and jurisdiction.

Google: introduction to robots.txt ↗

Check one record first.

Start with the smallest useful scope. Compare the returned record with its source, confirm variant identity and inspect missing fields before applying the mapping to a larger dataset.

Keep extraction warnings and the retrieval timestamp with downstream output. When a required field is missing, leave it unresolved or return to the source; do not fill it from an assumption.

Continue this workflow