Skip to content
Proud to collaborate with Microsoft for Startups

media.audio.transcribe ​

Transcribe an audio StorageObject via the org AI connection (provider-blind), persist the transcript to storage, and index it as a copy GeneratedArtifact. Returns refs plus the transcript text; audio bytes never enter workflow state.

Overview ​

PropertyValue
Workflow typeAtomic
LibraryApp-media
Version1.0

Input Schema ​

FieldTypeRequiredDefaultDescription
workflow_run_idstringNo—Parent DAG workflow run id (engine-injected when used as a step)
organization_uuidstringYes——
storage_object_uuidstringYes—StorageObject UUID of the audio to transcribe. Produce one with media.ingest.from-url (remote URL) or storage.upload.presigned-url (browser/direct upload).
media_profile_uuidstringNo—Optional MediaProfile supplying storage connection + bucket + key_prefix. Unlike media.image.*, a draft profile is accepted: profile readiness is an image-smoke verdict and does not gate audio.
ai_provider_config_uuidstringNo—AIProviderConfig to source the AI connection from; defaults to the org default
connection_uuidstringNo—Explicit AI CloudConnection UUID. Any connection serving the OpenAI /audio/transcriptions contract works; a non-native type can opt in with supports_transcription=true on its config.
modelstringNo—Transcription model, or the Azure deployment name. Falls back to the resolved AIProviderConfig's default_model, then to whisper-1. Bind a dedicated STT provider config and this can be omitted.
storage_connection_uuidstringNo—Required unless resolved from media_profile_uuid / org default profile
bucketstringNo—Required unless resolved from media_profile_uuid / org default profile
keystringNo—Object key for the transcript; default <key_prefix>/transcript/<uuid>
languagestringNo—ISO-639-1 hint (e.g. 'pt'); improves accuracy and latency
promptstringNo—Optional biasing prompt — domain vocabulary, names, spellings
temperaturefloatNo——
persist_transcriptbooleanNo—Default true — write the transcript to storage as text/plain
register_artifactbooleanNo—Default true — index the transcript as a copy GeneratedArtifact
review_gatestringNo——
brandstringNo——
campaignstringNo——
purposestringNo——
session_uuidstringNo—When set, also stamp agents.cost-entry-record unit=audio_second
produced_by_kindstringNo——
produced_by_uuidstringNo——
created_by_uuidstringNo——

Output Schema ​

FieldTypeRequiredDefaultDescription
workflow_run_idstringNo—Parent DAG workflow run id (engine-injected when used as a step)
organization_uuidstringYes——
storage_object_uuidstringYes—StorageObject UUID of the audio to transcribe. Produce one with media.ingest.from-url (remote URL) or storage.upload.presigned-url (browser/direct upload).
media_profile_uuidstringNo——
ai_provider_config_uuidstringNo——
connection_uuidstringNo—Explicit AI CloudConnection UUID. Any connection serving the OpenAI /audio/transcriptions contract works; a non-native type can opt in with supports_transcription=true on its config.
modelstringNo——
storage_connection_uuidstringNo—Required unless resolved from media_profile_uuid / org default profile
bucketstringNo—Required unless resolved from media_profile_uuid / org default profile
keystringNo—Object key for the transcript; default <key_prefix>/transcript/<uuid>
languagestringNo——
promptstringNo—Optional biasing prompt — domain vocabulary, names, spellings
temperaturefloatNo——
persist_transcriptbooleanNo—Default true — write the transcript to storage as text/plain
register_artifactbooleanNo—Default true — index the transcript as a copy GeneratedArtifact
review_gatestringNo——
brandstringNo——
campaignstringNo——
purposestringNo——
session_uuidstringNo—When set, also stamp agents.cost-entry-record unit=audio_second
produced_by_kindstringNo——
produced_by_uuidstringNo——
created_by_uuidstringNo——
transcriptstringNo—Transcript text, truncated at 20000 chars
transcript_truncatedbooleanNo——
transcript_charsintegerNo——
transcript_storage_object_uuidstringNo—Full transcript persisted as text/plain; always complete
audio_storage_object_uuidstringNo——
audio_artifact_uuidstringNo—Existing GeneratedArtifact for the source audio, when indexed
artifact_uuidstringNo—GeneratedArtifact (kind=copy) indexing the transcript
duration_secondsfloatNo——
duration_sourcestringNo—provider
unitstringNo——
quantityfloatNo——
cost_usdstringNo——
providerstringNo——
verdictstringNo——
failure_reasonstringNo—Engine-stamped human-readable failure reason
failed_atstringNo—ISO timestamp when the workflow failed
failed_at_statestringNo—Engine-stamped state when the workflow failed
failed_stepstringNo—Engine-stamped step name (DAG path)
failed_layerintegerNo—Engine-stamped layer index (DAG path)
errorstringNo—Engine-stamped exception message
error_typestringNo—Engine-stamped exception class name
completed_atstringNo——

States ​

StateInitialTerminalSuccessAuto-advanceDescription
pendingYesNo—execute—
completedNoYesYes——
failedNoYesNo——

State Diagram ​

Transitions ​

FromActionToDescription
pendingexecutecompleted—
* (any state)failfailed—

Outcomes ​

OutcomeTypeDescriptionState Data Keys
completedSUCCESSMedia atom completed successfullycompleted_at
failedFAILUREMedia atom failedfailure_reason, error, error_type

API Usage ​

bash
POST /api/workflows/start
Content-Type: application/json

{
  "workflow_type": "media.audio.transcribe",
  "initial_data": {
    "organization_uuid": "value",
    "storage_object_uuid": "value"
  }
}