Elevenlabs - Speech to Text

ElevenLabs

ElevenLabs

ElevenLabs on FuseAITools—five audio workflows: Multilingual v2 and Turbo 2.5 TTS, Speech-to-Text, Sound Effect v2, and AI Audio Isolation. Voice synthesis, transcription, SFX, and stem separation in the browser.

Multilingual v2
Turbo 2.5
Speech-to-Text
Sound Effect v2
AI Audio Isolation

Speech-to-Text Configuration

Click to upload audio fileSupports MP3, WAV, M4A formats, max 200MB
Upload audio file to recognize, supports multiple formats
Select the main language of the audio, or use auto detect
Identify and label different speakers
Mark audio events like music, noise, silence, etc.

Generation Result

No recognition result yet

Upload audio file and click "Start Recognition"
🎤Text-to-Speech: Supports multiple languages and voice styles, adjustable stability, similarity and style parameters
📝Speech-to-Text: High-precision speech recognition with speaker identification and audio event marking
🎵Sound Effect Generation: AI-driven sound effect generation with loop playback and duration control
✂️AI Audio Isolation: Intelligently isolate vocals and background music
🎙️ ElevenLabs · Voice & Audio · Five Workflows

📝 ElevenLabs Speech-to-Text

Upload an audio file (MP3, WAV, or M4A, max 200MB) and receive a transcript with optional speaker diarization and audio event tagging. Set language to auto detect or pick from supported ISO codes. Results include a word-level timeline you can click to seek playback.

What is ElevenLabs Speech-to-Text?

ElevenLabs Speech-to-Text on FuseAITools (elevenlabs_speech_to_text) transcribes uploaded audio. Required: audioUrl from upload (≤200MB). Optional: languageCode (auto or en/zh/ja/ko/es/fr/de/it/pt/ru), diarize for speaker identification, and tagAudioEvents for music/noise markers. Returns full text plus word timeline. Priced per minute of audio.

🎙️ ElevenLabs on FuseAITools

ElevenLabs on FuseAITools covers professional voice and audio in the browser—natural text-to-speech (Multilingual v2 or fast Turbo 2.5), accurate speech-to-text with optional diarization, AI sound effects, and audio isolation for stems. Pick a voice, paste or upload audio, and download results from history. Credits appear on the Generate button before you submit; new users receive 20 free credits on sign-up.

✨ ElevenLabs Core Features

Two TTS Models

Multilingual v2 for premium prosody; Turbo 2.5 for faster, lower-latency speech.

Speech-to-Text

Upload up to 200MB—auto language, diarization, event tags, and clickable word timelines.

Voice Fine-Tuning

Stability, similarity, style, and speed sliders plus optional context text for seamless multi-clip narration.

Cloud on FuseAITools

Generate, transcribe, and isolate in the browser—credits shown before submit; no local GPU required.

🎯 Built for These Scenarios

Video & podcast narrationDubbing & accessibilityMeeting transcriptsGame & film SFXVocal stem extraction

📊 ElevenLabs Workflow Quick Guide

WorkflowInputBest for
Multilingual v2 TTSVoice + text (≤5000 chars)High-quality narration, dubbing, audiobooks
Turbo 2.5 TTSVoice + text (≤5000 chars)Low-latency voice for assistants and batch runs
Speech-to-TextUploaded audio (≤200MB)Transcripts, subtitles, meeting notes
Sound Effect v2Text description (≤5000 chars)Game, video, and UI sound design
AI Audio IsolationUploaded audio (≤10MB)Vocal/instrument stems and remix prep

❓ FAQ (ElevenLabs)

MP3, WAV, and M4A uploads up to 200MB. Set language to auto detect or choose from English, Chinese, Japanese, Korean, Spanish, French, German, Italian, Portuguese, or Russian.

⚙️ ElevenLabs Technical Specs

Parameters below match the FuseAITools ElevenLabs form and API (elevenlabs_* model keys).

WorkflowmodelKeyRequiredPricing unitKey controls
Multilingual v2 TTSelevenlabs_text_to_speech_multilingualVoice + textcredits / 1K charsStability, similarity, style, speed (0.7–1.2); optional timestamps, context text, language code; MP3/PCM output
Turbo 2.5 TTSelevenlabs_text_to_speech_turboVoice + textcredits / 1K charsSame voice controls as Multilingual v2—optimized for faster generation
Speech-to-Textelevenlabs_speech_to_textUploaded audio URLcredits / minLanguage (auto or ISO code); speaker diarization; audio event tagging; word timeline in results
Sound Effect v2elevenlabs_sound_effectSound descriptioncredits / minDuration 0.5–22s; loop toggle; intensity (prompt influence); MP3/PCM output
AI Audio Isolationelevenlabs_audio_isolationUploaded audio URLcredits / minUpload MP3/WAV/M4A (≤10MB)—isolates vocals or instruments from mixed audio

TTS text and context fields: max 5000 characters each. STT uploads: max 200MB. Isolation uploads: max 10MB. Output formats include MP3 (128–320 kbps) and PCM (16–44.1 kHz).

ElevenLabs — Five Audio Workflows

Pick the mode that matches your starting material:

💳 New users get 20 free credits on sign-up. View pricing for subscription discounts and credit top-ups.
🔗 ElevenLabs pipeline tip Narrate with Multilingual v2 or Turbo 2.5 , transcribe meetings via Speech-to-Text , design SFX with Sound Effect v2 , and split stems with AI Audio Isolation . For full music tracks, pair with Suno Music Generation .