minimax_voice-gen_speech-2-8-hd $0.001 per call

MiniMax: Speech 2.8 HD. Generate high-definition speech from up to 10,000 characters of text. (MiniMax · Voice generation)

For AI agents: call minimax_voice-gen_speech-2-8-hd
POST https://toll402.dev/v1/p/minimax.voice-gen.speech-2-8-hd
{
 "model": "speech-2.8-hd",
 "text": "Hello.",
 "stream": false,
 "output_format": "url",
 "voice_setting": {
  "voice_id": "English_expressive_narrator",
  "speed": 1,
  "vol": 1,
  "pitch": 0
 },
 "audio_setting": {
  "sample_rate": 32000,
  "bitrate": 128000,
  "format": "mp3",
  "channel": 1
 }
}
Answers 402 with x402 payment requirements; pay with any x402 client (@x402/fetch, x402 for Python, toll402-client, toll402-mcp) or send a prepaid-credits key as x-toll402-key (https://toll402.dev/credits). Paid per call via x402 (USDC on Base) or with a prepaid-credits key (x-toll402-key, card at /credits); charged only on success. Start at /llms.txt or connect the MCP server: npx -y toll402-mcp · remote https://toll402.dev/mcp.

Input schema

{
 "type": "object",
 "properties": {
  "model": {
   "type": "string",
   "enum": [
    "speech-2.8-hd"
   ],
   "examples": [
    "speech-2.8-hd"
   ]
  },
  "text": {
   "type": "string",
   "description": "Fewer than 10,000 characters. Use between speakable segments for a pause of x seconds.",
   "examples": [
    "A calm voice can make a complex idea feel simple."
   ]
  },
  "stream": {
   "type": "boolean",
   "description": "This catalog route is non-streaming so the relay returns one bounded JSON response.",
   "enum": [
    false
   ],
   "examples": [
    false
   ]
  },
  "output_format": {
   "type": "string",
   "description": "The returned audio URL is valid for 24 hours.",
   "enum": [
    "url"
   ],
   "examples": [
    "url"
   ]
  },
  "voice_setting": {
   "type": "object",
   "description": "voice_id is required. Optional controls include speed, vol, pitch, emotion, text_normalization and latex_read. speech-2.8 emotion is happy | sad | angry | fearful | disgusted | surprised | calm; whisper and fluent are 2.6-only (status 2013 ",
   "properties": {
    "voice_id": {
     "type": "string",
     "examples": [
      "English_expressive_narrator"
     ]
    },
    "speed": {
     "type": "number",
     "examples": [
      1
     ]
    },
    "vol": {
     "type": "number",
     "examples": [
      1
     ]
    },
    "pitch": {
     "type": "integer",
     "examples": [
      0
     ]
    },
    "emotion": {
     "type": "string",
     "description": "speech-2.8 allows happy | sad | angry | fearful | disgusted | surprised | calm. whisper and fluent are 2.6-only and return status 2013 on 2.8 even with a whispering-named voice_id; use a whispering-style voice_id, or speech-2.6 for emotion=",
     "enum": [
      "happy",
      "sad",
      "angry",
      "fearful",
      "disgusted",
      "surprised",
      "calm"
     ]
    }
   },
   "required": [
    "voice_id"
   ],
   "examples": [
    {
     "voice_id": "English_expressive_narrator",
     "speed": 1,
     "vol": 1,
     "pitch": 0
    }
   ]
  },
  "audio_setting": {
   "type": "object",
   "description": "Non-streaming output supports mp3, wav and flac. bitrate (mp3 only) is 32000 | 64000 | 128000 | 256000; 192000 is invalid (status 2013). sample_rate is 8000 | 16000 | 22050 | 24000 | 32000 | 44100.",
   "properties": {
    "sample_rate": {
     "type": "integer",
     "enum": [
      8000,
      16000,
      22050,
      24000,
      32000,
      44100
     ],
     "examples": [
      32000
     ]
    },
    "bitrate": {
     "type": "integer",
     "description": "mp3 only. 192000 is invalid (status 2013).",
     "enum": [
      32000,
      64000,
      128000,
      256000
     ],
     "examples": [
      128000
     ]
    },
    "format": {
     "type": "string",
     "description": "Non-streaming output supports mp3, wav and flac.",
     "examples": [
      "mp3"
     ]
    },
    "channel": {
     "type": "integer",
     "examples": [
      1
     ]
    }
   },
   "examples": [
    {
     "sample_rate": 32000,
     "bitrate": 128000,
     "format": "mp3",
     "channel": 1
    }
   ]
  },
  "language_boost": {
   "type": "string",
   "description": "Exact enum strings only. English(UK) and en-GB are invalid (status 2013); use English. auto detects the language.",
   "enum": [
    "Chinese",
    "Chinese,Yue",
    "English",
    "Arabic",
    "Russian",
    "Spanish",
    "French",
    "Portuguese",
    "German",
    "Turkish",
    "Dutch",
    "Ukrainian",
    "Vietnamese",
    "Indonesian",
    "Japanese",
    "Italian",
    "Korean",
    "Thai",
    "Polish",
    "Romanian",
    "Greek",
    "Czech",
    "Finnish",
    "Hindi",
    "Bulgarian",
    "Danish",
    "Hebrew",
    "Malay",
    "Persian",
    "Slovak",
    "Swedish",
    "Croatian",
    "Filipino",
    "Hungarian",
    "Norwegian",
    "Slovenian",
    "Catalan",
    "Nynorsk",
    "Tamil",
    "Afrikaans",
    "auto"
   ],
   "examples": [
    "auto"
   ]
  },
  "pronunciation_dict": {
   "type": "object",
   "description": "Custom pronunciation pairs in source/target form.",
   "examples": [
    {
     "tone": [
      "Treg/tree-g"
     ]
    }
   ]
  },
  "subtitle_enable": {
   "type": "boolean",
   "default": false
  },
  "subtitle_type": {
   "type": "string",
   "default": "sentence",
   "enum": [
    "sentence",
    "word"
   ],
   "examples": [
    "sentence"
   ]
  },
  "voice_modify": {
   "type": "object",
   "description": "Optional pitch, intensity, timbre and sound_effects controls."
  }
 },
 "additionalProperties": true,
 "required": [
  "model",
  "text",
  "stream",
  "output_format",
  "voice_setting"
 ]
}

MCP

{"mcpServers":{"toll402":{"command":"npx","args":["-y","toll402-mcp"],"env":{"TOLL402_WALLET_KEY":"0x..."}}}}

All tools · OpenAPI