Thư viện Đại học Duy Tân, Đà Nẵng, Việt Nam

CSDL Bài trích Báo - Tạp chí

Hiển thị Marc

ADAPT-TTS: high-quality zero-shot multi-speaker text-to-speech adaptive-based for Vietnamese

Tác giả: Phuong Pham Ngoc, Chung Tran Quang, Mai Luong Chi

Số trang: P. 159-173

Số phát hành: Tập 39 - Số 2

Kiểu tài liệu: Tạp chí trong nước

Nơi lưu trữ: 03 Quang Trung

Mã phân loại: 621

Ngôn ngữ: English

Từ khóa: Zero-shot TTS, multi-speaker, text-to-speech, diffusion models, mel-spectrogram denoiser, Extracting Mel-vector, EMV, Adapt-TTS

Chủ đề: Engineering

Tóm tắt:

In this paper introduce the Adapt-TTS model that allows high-quality audio synthesis from a small adaptive sample without training to solve these problems. Key recommendations: 1. The Extracting Mel-vector (EMV) architecture allows for a better representation of speaker characteristics and speech style; 2. An improved zero-shot model with a denoising diffusion model (Mel-spectrogram denoiser) component allows for new voice synthesis without training with better quality (less noise).

Tạp chí liên quan

Bài báo Giảng viên DTU

Thư mục chuyên đề

CSDL Bài trích Báo - Tạp chí

ADAPT-TTS: high-quality zero-shot multi-speaker text-to-speech adaptive-based for Vietnamese

Tóm tắt: