Understanding Website Crawling

Key concepts:

  • Seed URL / Website URL — starting point for the crawl
  • Host — site host shown in Website sources
  • Pages ingested — how much content was indexed
  • Expires — when verification/crawl entitlement expires
  • Status — lifecycle of the source

Common statuses (title-cased from the API): Pending Dns, Dns Failed, Crawling, Ingesting, Completed, Failed, Expired.