Skip to content
Proud to collaborate with Microsoft for Startups

firecrawl.crawling.crawl-urls ​

Crawl multiple URLs based on options

Overview ​

PropertyValue
Workflow typeAtomic
LibraryApp-firecrawl
Version1.0

Input Schema ​

FieldTypeRequiredDefaultDescription
organization_uuiduuidNo—Organization that owns the connection. Resolved from the authenticated context; callers never pass it
cloud_connection_uuiduuidNo—Firecrawl connection to use (API key and optional self-hosted base_url). Omit to use the organization's only active Firecrawl connection
urlstringYes—The base URL to start crawling from
excludepathslistNo—URL pathname regex patterns that exclude matching URLs from the crawl. For example, if you set "excludePaths": ["blog/.*"] for the base URL firecrawl.dev, any results matching that pattern will be excluded, such as https://www.firecrawl.dev/blog/firecrawl-launch-week-1-recap.
includepathslistNo—URL pathname regex patterns that include matching URLs in the crawl. Only the paths that match the specified patterns will be included in the response. For example, if you set "includePaths": ["blog/.*"] for the base URL firecrawl.dev, only results matching that pattern will be included, such as https://www.firecrawl.dev/blog/firecrawl-launch-week-1-recap.
maxdepthintegerNo—Maximum depth to crawl relative to the base URL. Basically, the max number of slashes the pathname of a scraped URL may contain.
maxdiscoverydepthintegerNo—Maximum depth to crawl based on discovery order. The root site and sitemapped pages has a discovery depth of 0. For example, if you set it to 1, and you set ignoreSitemap, you will only crawl the entered URL and all URLs that are linked on that page.
ignoresitemapbooleanNo—Ignore the website sitemap when crawling
ignorequeryparametersbooleanNo—Do not re-scrape the same path with different (or none) query parameters
limitintegerNo—Maximum number of pages to crawl. Default limit is 10000.
allowbackwardlinksbooleanNo—Allows the crawler to follow internal links to sibling or parent URLs, not just child paths. false: Only crawls deeper (child) URLs. → e.g. /features/feature-1 → /features/feature-1/tips ✅ → Won't follow /pricing or / ❌ true: Crawls any internal links, including siblings and parents. → e.g. /features/feature-1 → /pricing, /, etc. ✅ Use true for broader internal coverage beyond nested paths.
allowexternallinksbooleanNo—Allows the crawler to follow links to external websites.
delayfloatNo—Delay in seconds between scrapes. This helps respect website rate limits.
webhookjsonNo—A webhook specification object.
scrapeoptionsjsonNo——

Output Schema ​

FieldTypeRequiredDefaultDescription
organization_uuiduuidNo—Organization that owns the connection. Resolved from the authenticated context; callers never pass it
cloud_connection_uuiduuidNo—Firecrawl connection to use (API key and optional self-hosted base_url). Omit to use the organization's only active Firecrawl connection
urlstringYes—The base URL to start crawling from
excludepathslistNo—URL pathname regex patterns that exclude matching URLs from the crawl. For example, if you set "excludePaths": ["blog/.*"] for the base URL firecrawl.dev, any results matching that pattern will be excluded, such as https://www.firecrawl.dev/blog/firecrawl-launch-week-1-recap.
includepathslistNo—URL pathname regex patterns that include matching URLs in the crawl. Only the paths that match the specified patterns will be included in the response. For example, if you set "includePaths": ["blog/.*"] for the base URL firecrawl.dev, only results matching that pattern will be included, such as https://www.firecrawl.dev/blog/firecrawl-launch-week-1-recap.
maxdepthintegerNo—Maximum depth to crawl relative to the base URL. Basically, the max number of slashes the pathname of a scraped URL may contain.
maxdiscoverydepthintegerNo—Maximum depth to crawl based on discovery order. The root site and sitemapped pages has a discovery depth of 0. For example, if you set it to 1, and you set ignoreSitemap, you will only crawl the entered URL and all URLs that are linked on that page.
ignoresitemapbooleanNo—Ignore the website sitemap when crawling
ignorequeryparametersbooleanNo—Do not re-scrape the same path with different (or none) query parameters
limitintegerNo—Maximum number of pages to crawl. Default limit is 10000.
allowbackwardlinksbooleanNo—Allows the crawler to follow internal links to sibling or parent URLs, not just child paths. false: Only crawls deeper (child) URLs. → e.g. /features/feature-1 → /features/feature-1/tips ✅ → Won't follow /pricing or / ❌ true: Crawls any internal links, including siblings and parents. → e.g. /features/feature-1 → /pricing, /, etc. ✅ Use true for broader internal coverage beyond nested paths.
allowexternallinksbooleanNo—Allows the crawler to follow links to external websites.
delayfloatNo—Delay in seconds between scrapes. This helps respect website rate limits.
webhookjsonNo—A webhook specification object.
scrapeoptionsjsonNo——
status_codeintegerNo—HTTP status code of the completed call
responsejsonNo—Parsed JSON response body
failure_reasonstringNo——
failure_typestringNo——
failed_atstringNo——
failed_stepstringNo——
failed_layerstringNo——
failed_at_statestringNo——
errorstringNo——
error_typestringNo——

States ​

StateInitialTerminalSuccessAuto-advanceDescription
pendingYesNo—executeWaiting to call POST /crawl
completedNoYesYes—HTTP call succeeded
failedNoYesNo—HTTP call failed

State Diagram ​

Transitions ​

FromActionToDescription
pendingexecutecompletedPerform POST /crawl
* (any state)failfailedRecord the failure reason

Outcomes ​

OutcomeTypeDescriptionState Data Keys
completedSUCCESSThe provider accepted the requeststatus_code, response
failedFAILUREThe provider rejected the request or was unreachablefailure_reason, failure_type, status_code

Business Errors ​

CodeMessage Template
HTTP_AUTHENTICATION_FAILEDThe provider rejected the credential (HTTP {status_code})
HTTP_AUTHORIZATION_FAILEDThe credential is not permitted to perform this operation (HTTP {status_code})
HTTP_NOT_FOUNDThe provider has no resource at {path} (HTTP 404)
HTTP_RATE_LIMITEDThe provider rate-limited this request (HTTP 429)
HTTP_REQUEST_FAILEDThe provider returned HTTP {status_code} for {method}
HTTP_RESPONSE_SHAPE_UNEXPECTEDThe provider response for {method} {path} could not be parsed as JSON
FIRECRAWL_CONNECTION_MISSINGfirecrawl_connection_missing: this organization has no active {connection_type} connection. Create one with connection.setup (type {connection_type}).
FIRECRAWL_CONNECTION_AMBIGUOUSfirecrawl_connection_ambiguous: this organization has more than one active {connection_type} connection. Pass cloud_connection_uuid to choose one.
FIRECRAWL_CONNECTION_NOT_FOUNDfirecrawl_connection_not_found: no {connection_type} connection {connection_uuid} in this organization.
FIRECRAWL_CONNECTION_CREDENTIAL_MISSINGfirecrawl_connection_credential_missing: the {connection_type} connection has no api_token. Update the connection with a Firecrawl API key.

API Usage ​

bash
POST /api/workflows/start
Content-Type: application/json

{
  "workflow_type": "firecrawl.crawling.crawl-urls",
  "initial_data": {
    "url": "value"
  }
}