What is it
Miso One is the flagship product from Miso Labs built around the Miso TTS 8B open-weights model. This English text-to-speech system is designed for expressive, conversational speech, voice continuation, and low-latency voice-agent research. Featuring 8 billion parameters with open weights, Miso One enables researchers and developers to evaluate voice quality, latency, prompt-audio behavior, and local deployment tradeoffs in a realistic production context. The model focuses on English-language speech with rich emotion, natural pacing, and conversational prosody, and it supports conditioning on prompt audio to enable voice continuation and one-shot voice cloning. Open access via Hugging Face and official repos lets you inspect the code, download weights, and run inference in your own environment, making Miso One a practical tool for experimentation, prototyping, and evaluation before deciding on production deployment.
Features
- Expressive English speech: high-quality, emotionally varied conversational voice with natural rhythm and pacing.
- Audio-context conditioning: prompt-driven voice continuation and one-shot voice cloning capabilities to maintain identity and tone.
- Low-latency design: publicly claimed low latency suitable for interactive voice-agent workflows (reported around 110 ms in suitable environments).
- Open-weights and local deployment: 8B parameter TTS with open weights for local inference via official repos and Hugging Face pages.
- High-fidelity output: 48 kHz preview and streaming-ready audio workflows for creator-centric production.
- Voice design and cloning: support for private voice models and prompt-based voice design within safety and consent boundaries.
- Cross-tool workflow: integrated with live translation, streaming transcripts, and publish-ready audio for dynamic production pipelines.
- Safety and watermarking considerations: official notes on safety boundaries, consent, and watermark guidance to guide responsible use.
How to Use
- Start with understanding the model and licenses: review the official model card, safety notes, and setup requirements on the Hugging Face page and MisoTTS GitHub repository.
- Try the hosted demo: audition voice quality, emotional range, and English pronunciation in a quick, browser-based test before diving into local setup.
- Run local inference (recommended for full evaluation): install the repository, download the 8B weights, and benchmark latency and memory usage in your CUDA environment.
- Test prompt audio: evaluate voice continuation with consented audio prompts, including short prompts, noisy prompts, and longer generated continuations.
- Decide on deployment: choose between self-hosting, hosted access, or continuing with another model based on your hardware, latency requirements, and production needs.
Pricing and access: Miso One offers a Free plan at 120 characters per conversion, with paid tiers that unlock larger limits. Annual deals provide substantial savings (up to 50% off) and credits that cover TTS usage, Voice Design previews, and Voice Clone activities. Plans include Basic, Pro, and Enterprise, each with annual character limits, voice credits, clone allowances, and prioritized support. For teams and high-volume workflows, Enterprise provides the largest quotas and fastest support.
Pricing
- Free plan: 120 characters per conversion.
- Basic: $9.90/mo (annual price shown as $4.95/mo with annual saving); includes 960,000 TTS characters per year, 9,600 voice credits, up to 480 instant voice clones, and basic email support.
- Pro: $29.90/mo (annual price shown as $14.95/mo); includes 4,200,000 TTS characters per year, 42,000 voice credits, up to 2,100 instant voice clones, and priority support for voice workflows.
- Enterprise: $49.90/mo (annual price shown as $24.95/mo); includes 9,600,000 TTS characters per year, 96,000 voice credits, up to 4,800 instant voice clones, and priority support from the Miso team.
- Credits are shared across TTS, Voice Design, and Voice Clone; annual plans bundle extensive usage and private voice model creation.
Note: The published figures reflect annual-billed plans and are intended to support consistent production planning. Always consult the official pricing page for the latest terms and any promotions.
Tips
- Evaluate latency on your target hardware: Miso One’s 8B model is not a lightweight browser toy. Real-world latency depends on GPUs, batching, serving code, and prompt length. Use the hosted demo as a quick filter, then benchmark locally with your target GPU.
- Test consent and safety boundaries: when exploring voice cloning and continuation, ensure you have consented audio material and adhere to watermarking and consent guidelines before public deployment.
- Start with English-only expectations: Miso One’s public release focuses on English speech today; plan multilingual needs if required with alternative models or future updates.
- Leverage prompt audio strategy: experiment with audio-conditioned prompts to maintain consistent voice identity and expressiveness across longer interactions.
- Prepare for local deployment: ensure you have sufficient local hardware (GPU memory and compute) to run the 8B weights effectively, and review the repository’s setup requirements before production tests.
- Use live translation and captions for broader workflows: take advantage of features like live EN→ES translation and streaming transcripts to enrich voice-enabled experiences in multilingual or accessibility-focused apps.
- Consider end-to-end production readiness: while Miso One provides valuable evaluation data, assess watermarking, safety, and consent workflows in your production pipeline to meet compliance and user expectations.
Frequently Asked Questions
-
What is Miso One?
Miso One is the product-facing name for Miso Labs’ Miso TTS 8B open-weights English text-to-speech model, designed for expressive conversational speech, voice continuation, and low-latency voice-agent research. -
Is Miso One open source?
Yes. The model weights are openly available, with open-weights access via the official repository and Hugging Face pages to inspect code, download weights, and run inference locally. -
How many languages does Miso One support?
The current public release focuses on English. It is not a broad multilingual model at this time. -
Can I run Miso TTS 8B locally?
Yes. Open weights are provided for local inference, and developers can install the repository, download 8B weights, and benchmark latency in a CUDA environment. -
Does Miso One support voice cloning?
Yes. The model supports prompted voice cloning and voice continuation, with guidelines and safety notes to ensure consent and responsible use. -
Is Miso One production-ready?
Miso One is designed for evaluation and experimentation, with plans and pricing that support production testing. Production readiness depends on your hardware, latency requirements, safety practices, and compliance with watermarking and consent policies. -
Where can I access the model and documentation?
Access the official Hugging Face model page and the MisoTTS GitHub repository for model cards, licensing, safety notes, and setup instructions. -
What are the key advantages of Miso One?
The main advantages are open-weights access for local deployment, expressive English speech with low latency, audio-context conditioning for voice continuation, and structured pricing with scalable plans for creators and teams. These features facilitate thorough evaluation, local testing, and integration into voice-centric applications.
This comprehensive overview captures Miso One’s core value: an openly accessible, low-latency English TTS solution with expressive capabilities, designed for researchers and creators who need realistic voice agents, robust evaluation tools, and scalable deployment options.