Skip to content
Proud to collaborate with Microsoft for Startups

firecrawl.extraction.extract-data ​

Extract structured data from pages using LLMs

Overview ​

PropertyValue
Workflow typeAtomic
LibraryApp-firecrawl
Version1.0

Input Schema ​

FieldTypeRequiredDefaultDescription
organization_uuiduuidNo—Organization that owns the connection. Resolved from the authenticated context; callers never pass it
cloud_connection_uuiduuidNo—Firecrawl connection to use (API key and optional self-hosted base_url). Omit to use the organization's only active Firecrawl connection
urlslistYes——
promptstringNo—Prompt to guide the extraction process
schemajsonNo—Schema to define the structure of the extracted data. Must conform to JSON Schema.
enablewebsearchbooleanNo—When true, the extraction will use web search to find additional data
ignoresitemapbooleanNo—When true, sitemap.xml files will be ignored during website scanning
includesubdomainsbooleanNo—When true, subdomains of the provided URLs will also be scanned
showsourcesbooleanNo—When true, the sources used to extract the data will be included in the response as sources key
scrapeoptionsjsonNo——
ignoreinvalidurlsbooleanNo—If invalid URLs are specified in the urls array, they will be ignored. Instead of them failing the entire request, an extract using the remaining valid URLs will be performed, and the invalid URLs will be returned in the invalidURLs field of the response.

Output Schema ​

FieldTypeRequiredDefaultDescription
organization_uuiduuidNo—Organization that owns the connection. Resolved from the authenticated context; callers never pass it
cloud_connection_uuiduuidNo—Firecrawl connection to use (API key and optional self-hosted base_url). Omit to use the organization's only active Firecrawl connection
urlslistYes——
promptstringNo—Prompt to guide the extraction process
schemajsonNo—Schema to define the structure of the extracted data. Must conform to JSON Schema.
enablewebsearchbooleanNo—When true, the extraction will use web search to find additional data
ignoresitemapbooleanNo—When true, sitemap.xml files will be ignored during website scanning
includesubdomainsbooleanNo—When true, subdomains of the provided URLs will also be scanned
showsourcesbooleanNo—When true, the sources used to extract the data will be included in the response as sources key
scrapeoptionsjsonNo——
ignoreinvalidurlsbooleanNo—If invalid URLs are specified in the urls array, they will be ignored. Instead of them failing the entire request, an extract using the remaining valid URLs will be performed, and the invalid URLs will be returned in the invalidURLs field of the response.
status_codeintegerNo—HTTP status code of the completed call
responsejsonNo—Parsed JSON response body
failure_reasonstringNo——
failure_typestringNo——
failed_atstringNo——
failed_stepstringNo——
failed_layerstringNo——
failed_at_statestringNo——
errorstringNo——
error_typestringNo——

States ​

StateInitialTerminalSuccessAuto-advanceDescription
pendingYesNo—executeWaiting to call POST /extract
completedNoYesYes—HTTP call succeeded
failedNoYesNo—HTTP call failed

State Diagram ​

Transitions ​

FromActionToDescription
pendingexecutecompletedPerform POST /extract
* (any state)failfailedRecord the failure reason

Outcomes ​

OutcomeTypeDescriptionState Data Keys
completedSUCCESSThe provider accepted the requeststatus_code, response
failedFAILUREThe provider rejected the request or was unreachablefailure_reason, failure_type, status_code

Business Errors ​

CodeMessage Template
HTTP_AUTHENTICATION_FAILEDThe provider rejected the credential (HTTP {status_code})
HTTP_AUTHORIZATION_FAILEDThe credential is not permitted to perform this operation (HTTP {status_code})
HTTP_NOT_FOUNDThe provider has no resource at {path} (HTTP 404)
HTTP_RATE_LIMITEDThe provider rate-limited this request (HTTP 429)
HTTP_REQUEST_FAILEDThe provider returned HTTP {status_code} for {method}
HTTP_RESPONSE_SHAPE_UNEXPECTEDThe provider response for {method} {path} could not be parsed as JSON
FIRECRAWL_CONNECTION_MISSINGfirecrawl_connection_missing: this organization has no active {connection_type} connection. Create one with connection.setup (type {connection_type}).
FIRECRAWL_CONNECTION_AMBIGUOUSfirecrawl_connection_ambiguous: this organization has more than one active {connection_type} connection. Pass cloud_connection_uuid to choose one.
FIRECRAWL_CONNECTION_NOT_FOUNDfirecrawl_connection_not_found: no {connection_type} connection {connection_uuid} in this organization.
FIRECRAWL_CONNECTION_CREDENTIAL_MISSINGfirecrawl_connection_credential_missing: the {connection_type} connection has no api_token. Update the connection with a Firecrawl API key.

API Usage ​

bash
POST /api/workflows/start
Content-Type: application/json

{
  "workflow_type": "firecrawl.extraction.extract-data",
  "initial_data": {
    "urls": "value"
  }
}