ホームすべてのツールカテゴリー
ツールを見る

GPA

A unified, lightweight audio model for ASR, TTS, and voice conversion

free(無料プランあり)3.1k github stars

GPA について

GPA (General Purpose Audio) is a single compact model capable of performing automatic speech recognition (ASR), text-to-speech (TTS), and voice conversion (VC). It aims to unify multiple audio tasks without requiring separate large models.

The latest GPA-v1.5 release delivers near-state-of-the-art performance across ASR and TTS while remaining deployable via ONNX Runtime — supporting CLI tools, FastAPI services, and a browser-based UI. The model offers runtime-selectable decoder precision (INT8, FP16, FP32) to balance quality and compute requirements.

GPA is developed by AutoArk and published as an open-source project with documentation and runtime assets available on its GitHub and website.

主な機能

  • Single unified model supporting ASR, TTS, and voice conversion
  • ONNX Runtime support enabling cross-platform deployment (CLI, FastAPI, browser UI)
  • Runtime-selectable decoder precision (INT8, FP16, FP32) for quality/compute trade-offs
  • Lightweight architecture designed for efficiency without sacrificing near-SOTA performance
  • Pre-built runtime asset bundle for quick local or server-side setup

メリット & デメリット

メリット

  • • Unifies three major audio tasks in one model, reducing deployment complexity
  • • ONNX compatibility enables broad hardware and platform support including edge devices
  • • Flexible precision options allow tuning for quality or resource constraints

デメリット

  • • No explicit documentation on minimum hardware requirements or latency benchmarks
  • • Voice conversion capabilities are mentioned but lack detail on supported languages, speaker adaptation, or evaluation metrics
  • • No public information about commercial licensing terms despite likely open-source status