メインコンテンツまでスキップ

If you're an agent

Use the company-hosted MCP connection. Your user does not need a Kitsch server, package download, or manually copied API key.

Release status

The hosted service is deployed at https://mcp.kitschlabs.com/mcp. HTTPS, readiness, OAuth metadata, and unauthenticated request rejection have been verified. Actual Kitsch OAuth connections in the clients below and paid audio requests have not yet been verified. Do not invent a package or download URL.

Connect first​

Follow the client-specific commands for OpenClaw, Hermes, Claude Code, Codex, or Cursor. Preserve unrelated MCP servers and existing settings.

  1. Register https://mcp.kitschlabs.com/mcp as a remote HTTP MCP server.
  2. Start the client's OAuth login flow and let the user finish Kitsch account sign-in and consent in their browser. Do not ask for passwords or API keys in chat.
  3. The user chooses an active workspace where they are an owner or admin. Synthesis and transcription consume that workspace's credits or incur usage charges on a postpaid plan.
  4. Refresh the client connection or start a new session. Discover tools and call list_voices to confirm access. Setup alone does not authorize a paid audio test.

The user can revoke access in the Kitsch app under Account settings → Connected agents. Revocation blocks new requests; it does not undo work already started.

Use the available tools​

Discover the actual tool schemas after connecting. Names may have a server prefix, and only tools covered by the approved permissions are available.

ToolPurpose
list_voicesGet available voices. Use a returned voice_id; do not invent one.
synthesize_speechCreate a complete MP3 or WAV file from user-authorized text.
prepare_transcriptionPrepare a direct file upload. This does not upload or transcribe the file.
get_audio_operationRecover the status or result of an existing operation without starting a new one.

For each new synthesis or transcription, generate and retain one operation ID:

const operation_id = Date.now() + "_" + crypto.randomUUID();

Use the current timestamp and a fresh UUIDv4. Keep the same ID for recovery. Do not reuse a documentation example or substitute a bare UUID.

Synthesize speech​

Call list_voices, then pass the chosen voice_id, text, and operation_id to synthesize_speech. MP3 is the default; output_format: "wav" selects WAV. Omit unspecified optional fields instead of sending empty strings.

The result contains a download link and MIME type. Links last up to 15 minutes; files are retained for 24 hours. Use get_audio_operation for a fresh link during retention. Treat the link as private. Optional return_audio: true also returns audio content for new results up to 2 MiB; recovery returns links.

Transcribe a file​

You need a shell or HTTP tool that can read the user's authorized audio file. A remote MCP server cannot read an arbitrary path on the user's computer.

  1. Read the filename and actual byte size. The maximum is 100 MiB.
  2. Call prepare_transcription with operation_id, filename, and size_bytes. Optional language and prompt must be omitted if unset.
  3. Use the returned method, upload_url, headers, and multipart_fields for the HTTP request, attaching the file under file_field. Send the file directly; do not stage it on another file-sharing service.
  4. Keep the transcript and operation ID. Do not print the upload credentials in chat or shell logs. The upload credential is single-use and expires after five minutes.

Do not pass audio_url or invent a transcribe_audio tool on the hosted server. If you have no file-capable shell or HTTP tool, explain that file transcription cannot be completed in the current client.

Recover without duplicate charges​

If the response is lost or times out, query get_audio_operation with the original ID. An unknown status does not establish that nothing ran or was billed. Do not automatically create a new operation or resend the file. Tell the user what is known and retain any request ID for support.

An uncertain synthesis failure can block new work in the same workspace for five minutes. Respect admission_blocked_until; it is not permission to retry a paid operation automatically. Disconnecting does not undo processing or charges.

Realtime audio is a separate integration​

These MCP tools return completed audio files and transcripts. Streamable HTTP does not mean live microphone capture or streaming playback, and connecting MCP does not change the agent's default voice provider. Realtime STT/TTS adapters for Hermes and OpenClaw are follow-up work.

Developers who intentionally need a local process can use the separate advanced stdio guide.