If you're an agent
Use the company-hosted MCP connection. Your user does not need a Kitsch server, package download, or manually copied API key.
The hosted service is deployed at https://mcp.kitschlabs.com/mcp.
HTTPS, readiness, OAuth metadata, and unauthenticated request rejection have
been verified. Actual Kitsch OAuth connections in the clients below and paid
audio requests have not yet been verified. Do not invent a package or download URL.
Connect first
Follow the client-specific commands for OpenClaw, Hermes, Claude Code, Codex, or Cursor. Preserve unrelated MCP servers and existing settings.
- Register
https://mcp.kitschlabs.com/mcpas a remote HTTP MCP server. - Start the client's OAuth login flow and let the user finish Kitsch account sign-in and consent in their browser. Do not ask for passwords or API keys in chat.
- The user chooses an active workspace where they are an owner or admin. Synthesis and transcription consume that workspace's credits or incur usage charges on a postpaid plan.
- Refresh the client connection or start a new session. Discover tools and call
list_voicesto confirm access. Setup alone does not authorize a paid audio test.
The user can revoke access in the Kitsch app under Account settings → Connected agents. Revocation blocks new requests; it does not undo work already started.
Use the available tools
Discover the actual tool schemas after connecting. Names may have a server prefix, and only tools covered by the approved permissions are available.
| Tool | Purpose |
|---|---|
list_voices | Get available voices. Use a returned voice_id; do not invent one. |
synthesize_speech | Create a complete MP3 or WAV file from user-authorized text. |
prepare_transcription | Prepare a direct file upload. This does not upload or transcribe the file. |
get_audio_operation | Recover the status or result of an existing operation without starting a new one. |
For each new synthesis or transcription, generate and retain one operation ID:
const operation_id = Date.now() + "_" + crypto.randomUUID();
Use the current timestamp and a fresh UUIDv4. Keep the same ID for recovery. Do not reuse a documentation example or substitute a bare UUID.
Synthesize speech
Call list_voices, then pass the chosen voice_id, text, and operation_id
to synthesize_speech. MP3 is the default; output_format: "wav" selects WAV.
Omit unspecified optional fields instead of sending empty strings.
The result contains a download link and MIME type. Links last up to 15 minutes;
files are retained for 24 hours. Use get_audio_operation for a fresh link during
retention. Treat the link as private. Optional return_audio: true also returns
audio content for new results up to 2 MiB; recovery returns links.
Transcribe a file
You need a shell or HTTP tool that can read the user's authorized audio file. A remote MCP server cannot read an arbitrary path on the user's computer.
- Read the filename and actual byte size. The maximum is 100 MiB.
- Call
prepare_transcriptionwithoperation_id,filename, andsize_bytes. Optionallanguageandpromptmust be omitted if unset. - Use the returned
method,upload_url,headers, andmultipart_fieldsfor the HTTP request, attaching the file underfile_field. Send the file directly; do not stage it on another file-sharing service. - Keep the transcript and operation ID. Do not print the upload credentials in chat or shell logs. The upload credential is single-use and expires after five minutes.
Do not pass audio_url or invent a transcribe_audio tool on the hosted server.
If you have no file-capable shell or HTTP tool, explain that file transcription
cannot be completed in the current client.
Recover without duplicate charges
If the response is lost or times out, query get_audio_operation with the original
ID. An unknown status does not establish that nothing ran or was billed.
Do not automatically create a new operation or resend the file. Tell the user
what is known and retain any request ID for support.
An uncertain synthesis failure can block new work in the same workspace for five
minutes. Respect admission_blocked_until; it is not permission to retry a paid
operation automatically. Disconnecting does not undo processing or charges.
Realtime audio is a separate integration
These MCP tools return completed audio files and transcripts. Streamable HTTP does not mean live microphone capture or streaming playback, and connecting MCP does not change the agent's default voice provider. Realtime STT/TTS adapters for Hermes and OpenClaw are follow-up work.
Developers who intentionally need a local process can use the separate advanced stdio guide.