firecrawl.extraction.extract-data ​
Extract structured data from pages using LLMs
Overview ​
| Property | Value |
|---|---|
| Workflow type | Atomic |
| Library | App-firecrawl |
| Version | 1.0 |
Input Schema ​
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
organization_uuid | uuid | No | — | Organization that owns the connection. Resolved from the authenticated context; callers never pass it |
cloud_connection_uuid | uuid | No | — | Firecrawl connection to use (API key and optional self-hosted base_url). Omit to use the organization's only active Firecrawl connection |
urls | list | Yes | — | — |
prompt | string | No | — | Prompt to guide the extraction process |
schema | json | No | — | Schema to define the structure of the extracted data. Must conform to JSON Schema. |
enablewebsearch | boolean | No | — | When true, the extraction will use web search to find additional data |
ignoresitemap | boolean | No | — | When true, sitemap.xml files will be ignored during website scanning |
includesubdomains | boolean | No | — | When true, subdomains of the provided URLs will also be scanned |
showsources | boolean | No | — | When true, the sources used to extract the data will be included in the response as sources key |
scrapeoptions | json | No | — | — |
ignoreinvalidurls | boolean | No | — | If invalid URLs are specified in the urls array, they will be ignored. Instead of them failing the entire request, an extract using the remaining valid URLs will be performed, and the invalid URLs will be returned in the invalidURLs field of the response. |
Output Schema ​
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
organization_uuid | uuid | No | — | Organization that owns the connection. Resolved from the authenticated context; callers never pass it |
cloud_connection_uuid | uuid | No | — | Firecrawl connection to use (API key and optional self-hosted base_url). Omit to use the organization's only active Firecrawl connection |
urls | list | Yes | — | — |
prompt | string | No | — | Prompt to guide the extraction process |
schema | json | No | — | Schema to define the structure of the extracted data. Must conform to JSON Schema. |
enablewebsearch | boolean | No | — | When true, the extraction will use web search to find additional data |
ignoresitemap | boolean | No | — | When true, sitemap.xml files will be ignored during website scanning |
includesubdomains | boolean | No | — | When true, subdomains of the provided URLs will also be scanned |
showsources | boolean | No | — | When true, the sources used to extract the data will be included in the response as sources key |
scrapeoptions | json | No | — | — |
ignoreinvalidurls | boolean | No | — | If invalid URLs are specified in the urls array, they will be ignored. Instead of them failing the entire request, an extract using the remaining valid URLs will be performed, and the invalid URLs will be returned in the invalidURLs field of the response. |
status_code | integer | No | — | HTTP status code of the completed call |
response | json | No | — | Parsed JSON response body |
failure_reason | string | No | — | — |
failure_type | string | No | — | — |
failed_at | string | No | — | — |
failed_step | string | No | — | — |
failed_layer | string | No | — | — |
failed_at_state | string | No | — | — |
error | string | No | — | — |
error_type | string | No | — | — |
States ​
| State | Initial | Terminal | Success | Auto-advance | Description |
|---|---|---|---|---|---|
pending | Yes | No | — | execute | Waiting to call POST /extract |
completed | No | Yes | Yes | — | HTTP call succeeded |
failed | No | Yes | No | — | HTTP call failed |
State Diagram ​
Transitions ​
| From | Action | To | Description |
|---|---|---|---|
pending | execute | completed | Perform POST /extract |
* (any state) | fail | failed | Record the failure reason |
Outcomes ​
| Outcome | Type | Description | State Data Keys |
|---|---|---|---|
completed | SUCCESS | The provider accepted the request | status_code, response |
failed | FAILURE | The provider rejected the request or was unreachable | failure_reason, failure_type, status_code |
Business Errors ​
| Code | Message Template |
|---|---|
HTTP_AUTHENTICATION_FAILED | The provider rejected the credential (HTTP {status_code}) |
HTTP_AUTHORIZATION_FAILED | The credential is not permitted to perform this operation (HTTP {status_code}) |
HTTP_NOT_FOUND | The provider has no resource at {path} (HTTP 404) |
HTTP_RATE_LIMITED | The provider rate-limited this request (HTTP 429) |
HTTP_REQUEST_FAILED | The provider returned HTTP {status_code} for {method} |
HTTP_RESPONSE_SHAPE_UNEXPECTED | The provider response for {method} {path} could not be parsed as JSON |
FIRECRAWL_CONNECTION_MISSING | firecrawl_connection_missing: this organization has no active {connection_type} connection. Create one with connection.setup (type {connection_type}). |
FIRECRAWL_CONNECTION_AMBIGUOUS | firecrawl_connection_ambiguous: this organization has more than one active {connection_type} connection. Pass cloud_connection_uuid to choose one. |
FIRECRAWL_CONNECTION_NOT_FOUND | firecrawl_connection_not_found: no {connection_type} connection {connection_uuid} in this organization. |
FIRECRAWL_CONNECTION_CREDENTIAL_MISSING | firecrawl_connection_credential_missing: the {connection_type} connection has no api_token. Update the connection with a Firecrawl API key. |
API Usage ​
bash
POST /api/workflows/start
Content-Type: application/json
{
"workflow_type": "firecrawl.extraction.extract-data",
"initial_data": {
"urls": "value"
}
}