create_mode or agentic_thinking. The mode is inferred.
Spoken audio is the same task. See Speech.
POST /suno/v2/generate returns { "taskId" }. Poll Get Task Status. Schema: GET /suno/v2/generate/schema?task=create.
Parameters
audio_refs[]: clip_id required; style_prompt, lyrics, duration_s optional. Pick is ByModel — the clip does not have to belong to the pool account.
Images
image_urls is the official Simple + Image control. It is only for task=create with an empty prompt and no persona. Images alone are enough to enter Simple — you do not also need gpt_description_prompt. You can still send an idea, audio_refs, or attached_styles in the same body.
The server downloads each URL, uploads it on the account selected for this generate, and sends the resulting ids upstream. You only pass URLs. Omit the field to leave the request unchanged. Do not send an empty array.
Examples
Custom — you write the lyrics:prompt, no persona:
clip_id alone is enough:
prompt, no persona: