The workflow remains on this action while Firecrawl is scraping. The node
checks the crawl every two seconds and fails if Firecrawl reports a failed
crawl.
Crawl inputs
Per-page scrape inputs
These settings are applied to every page visited:Outputs
Expanding Pages gives each page:
Scope the crawl deliberately
Start with a modest Limit and Max Depth, inspect the returned URLs, and increase them only when needed. Query-heavy sites can expose many near-duplicate URLs; Ignore Query Parameters helps keep those from consuming the page limit.Troubleshooting
The crawl visits too many duplicate pages
The crawl visits too many duplicate pages
Enable Ignore Query Parameters and lower Max Depth. Tracking,
filtering, and pagination parameters often create many URL variants.
Pages above the starting path are missing
Pages above the starting path are missing
Enable Allow Backward Links. Leave it off when the crawl should stay
within a subsection of a site.
The action takes a long time
The action takes a long time
Reduce Limit, Max Depth, and Wait For. Full-page screenshots on
every page also add substantial work.