首页全部工具分类
浏览工具

dia

A text-to-speech model for generating ultra-realistic dialogue in one pass.

free(有免费版)19.4k github stars

dia 简介

Dia is a 1.6B parameter text-to-speech model developed by Nari Labs. It is designed to generate natural-sounding spoken dialogue directly from text, supporting speaker differentiation and non-verbal cues.

The model supports special markup like [S1] and [S2] tags to distinguish speakers, and recognizes non-verbal expressions such as (laughs) or (coughs). Voice cloning functionality is also included, with an example script provided in the repository.

Dia2, an updated version, has been released on GitHub and Hugging Face, and the model is now accessible via Hugging Face Transformers.

核心功能

  • Generates multi-speaker dialogue using [S1] and [S2] speaker tags.
  • Supports non-verbal expressions like (laughs) and (coughs).
  • Includes voice cloning capability, demonstrated in voice_clone.py.
  • Available through Hugging Face Transformers for easy integration.
  • Released as open-weight, enabling local inference and customization.

优点 & 缺点

优点

  • • Open-weight model allows full transparency and local deployment.
  • • Supports expressive, multi-speaker dialogue generation in a single pass.
  • • Voice cloning feature enables personalized speech synthesis.

缺点

  • • Non-verbal tags may produce unexpected output, as noted in the README.
  • • No official web interface or hosted service — requires technical setup for use.
  • • Licensing details are not explicitly stated, creating uncertainty around commercial use.