
GPA
A unified, lightweight audio model for ASR, TTS, and voice conversion
GPA 简介
GPA (General Purpose Audio) is a single compact model capable of performing automatic speech recognition (ASR), text-to-speech (TTS), and voice conversion (VC). It aims to unify multiple audio tasks without requiring separate large models.
The latest GPA-v1.5 release delivers near-state-of-the-art performance across ASR and TTS while remaining deployable via ONNX Runtime — supporting CLI tools, FastAPI services, and a browser-based UI. The model offers runtime-selectable decoder precision (INT8, FP16, FP32) to balance quality and compute requirements.
GPA is developed by AutoArk and published as an open-source project with documentation and runtime assets available on its GitHub and website.
核心功能
- Single unified model supporting ASR, TTS, and voice conversion
- ONNX Runtime support enabling cross-platform deployment (CLI, FastAPI, browser UI)
- Runtime-selectable decoder precision (INT8, FP16, FP32) for quality/compute trade-offs
- Lightweight architecture designed for efficiency without sacrificing near-SOTA performance
- Pre-built runtime asset bundle for quick local or server-side setup
优点 & 缺点
优点
- • Unifies three major audio tasks in one model, reducing deployment complexity
- • ONNX compatibility enables broad hardware and platform support including edge devices
- • Flexible precision options allow tuning for quality or resource constraints
缺点
- • No explicit documentation on minimum hardware requirements or latency benchmarks
- • Voice conversion capabilities are mentioned but lack detail on supported languages, speaker adaptation, or evaluation metrics
- • No public information about commercial licensing terms despite likely open-source status