ホームすべてのツールカテゴリー
ツールを見る

dia

A text-to-speech model for generating ultra-realistic dialogue in one pass.

free(無料プランあり)19.4k github stars

dia について

Dia is a 1.6B parameter text-to-speech model developed by Nari Labs. It is designed to generate natural-sounding spoken dialogue directly from text, supporting speaker differentiation and non-verbal cues.

The model supports special markup like [S1] and [S2] tags to distinguish speakers, and recognizes non-verbal expressions such as (laughs) or (coughs). Voice cloning functionality is also included, with an example script provided in the repository.

Dia2, an updated version, has been released on GitHub and Hugging Face, and the model is now accessible via Hugging Face Transformers.

主な機能

  • Generates multi-speaker dialogue using [S1] and [S2] speaker tags.
  • Supports non-verbal expressions like (laughs) and (coughs).
  • Includes voice cloning capability, demonstrated in voice_clone.py.
  • Available through Hugging Face Transformers for easy integration.
  • Released as open-weight, enabling local inference and customization.

メリット & デメリット

メリット

  • • Open-weight model allows full transparency and local deployment.
  • • Supports expressive, multi-speaker dialogue generation in a single pass.
  • • Voice cloning feature enables personalized speech synthesis.

デメリット

  • • Non-verbal tags may produce unexpected output, as noted in the README.
  • • No official web interface or hosted service — requires technical setup for use.
  • • Licensing details are not explicitly stated, creating uncertainty around commercial use.