Skip to main content
Generation tools use the documented REST API and the same account credits and quotas. Viewing saved media and checking video status do not spend generation credits. Expand a tool below for its parameters. Parameters marked with * are required.

Text generation

Lists a page of templates plus the list of categories. Returns total, offset, and next_offset (null on the final page).
Returns one template with its input fields (name, label, type, options). Field names such as keywords and extra1…extra14 map directly onto generate_text parameters.
Runs a template and returns the generated text. Uses the account word quota and API token allowance.
Call get_template before generate_text so the agent knows which extra_fields a template expects.

Images

Generates or edits an image and waits for the result — this can take a few minutes. Uses image credits and requires a paid plan.

Video

Starts a generation and returns an id immediately. Defaults to MiniMax H3 Max, the same video model used by Marky Agent: 5 seconds at 768p with native audio, billed at 4 credits per second at 768p. With a first-frame image and no model, uses MiniMax H3 Max Image.
Returns { video: { id, url, status } }. status stays processing until the video is ready, then becomes completed with a download url.

Display images and videos

Shows saved images in a gallery or videos in a player in clients that support MCP Apps. Use IDs returned by generate_image, get_image, generate_video, or get_video.The viewer includes an original-file link. For processing videos, it checks the same ID every 15 seconds for up to five minutes, then offers Check status. It never starts a new generation or spends generation credits.Clients without embedded UI support still receive the media URLs. Refresh your connector’s tool list if display_media is missing.
Image and video tools also return structuredContent.media, an array with each result’s id, type, url, and status. Optional fields include prompt, model, and used_credits when available. Text results remain available for other MCP clients.

Speech

Returns a page of voices with id, name, language, and accent, including your custom and cloned voices. Returns total, offset, and next_offset (null on the final page).
Converts text to speech and returns the audio file URL. Billed against the monthly text-to-speech character quota.

Transcription

Starts a transcription of an audio or video file and returns a uuid immediately. Requires a paid plan.
The content field stays empty until the transcription finishes. Reuse the exact UUID and the same connected account. Report a missing or failed job; do not create a replacement unless the user asks for a new transcription.

Async tools

generate_video and transcribe_audio return an id immediately instead of waiting for the result. Use display_media for a video viewer that checks progress, or poll the matching get_video or get_transcription tool every 15–30 seconds until the result is ready. Always reuse the original ID. Videos typically take 1–5 minutes depending on model, duration, and resolution.

Account

Takes no parameters and returns { "connected": true }. Use it only when the user asks to verify their connection. It does not return an account ID, email, name, API key, or OAuth token, and it does not check quota or support account deletion.