Fun-CosyVoice3-0.5B-2512 是阿里巴巴通义语音团队推出的开源文本转语音(TTS)模型,具备多语言、多方言合成、情感控制、音色克隆等实用功能,且部署门槛低,非常适合初学者上手。本教程将围绕最基础的环境搭建、核心功能使用及简单应用,带大家快速掌握模型的基本用法。
- 硬件要求:建议 GPU 显存 ≥ 12GB,无 GPU 环境可使用 CPU 推理,但速度较慢。
- 依赖工具:确保已安装 Git(用于克隆代码和下载模型)、Anaconda/Miniconda(用于创建独立环境)。
1. 模型加载
创建项目以及需要的环境:
# ====================== 操作步骤 ====================== # 1. 使用 conda 创建虚拟环境 cosyvoice-env # 2. 在 PyCharm 中创建 "TTS服务器" 项目,并配置环境为 cosyvoice-env # 3. 在 "TTS服务器" 项目目录中,下载模型 Fun-CosyVoice3-0.5B-2512 以及模型开发包 CosyVoice # 4. 进入 CosyVoice 目录,安装其依赖的包 # ===================================================== # ====================== 相关命令和链接 ====================== # 模型链接:https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512 # 下载开发包 git clone --recursive https://github.com/FunAudioLLM/CosyVoice.git # 下载模型 git clone https://www.modelscope.cn/FunAudioLLM/Fun-CosyVoice3-0.5B-2512.git # 创建虚拟环境,并安装开发包相关依赖 conda create -n cosyvoice-env python=3.10 conda activate cosyvoice-env pip install -r requirements.txt -i https://mirrors.aliyun.com/pypi/simple/ --trusted-host=mirrors.aliyun.com
在项目中创建 demo.py 文件,在该文件中编写模型使用相关代码。模型加载需要使用目录 CosyVoice/cosyvoice/cli/cosyvoice 下的 CosyVoice3 类,为了能够正确导入该类,需要将两个目录增加到包搜索路径:
import sys
sys.path.append('CosyVoice')
sys.path.append('CosyVoice/third_party/Matcha-TTS')
接下来,使用下面代码就可以正确加载模型:
from cosyvoice.cli.cosyvoice import CosyVoice3 # model_dir: 模型路径,必须 # load_trt=False: 是否使用 TensorRT 推理框架,我们场景简单,不使用 # load_vllm=False: 是否使用 vllm 推理框架,我们场景简单,不使用 # fp16=False: 是否在推理时,使用自动混合精度,有助于提高推理速度,使用 cosyvoice = CosyVoice3(model_dir='Fun-CosyVoice3-0.5B-2512', fp16=True)
2. 模型推理
CosyVoice3 模型提供了三种不同的推理函数,分别是:
- inference_zero_shot:参考音频的全部声学特征,追求原声复刻。例如:音色、语速、口音细节
- inference_cross_lingual:只迁移参考音频的音色,保持目标语言自然度
- inference_instruct2:用指令控制语音风格,可结合参考音色。例如:语速、情感、方言等
import sys
sys.path.append('CosyVoice')
sys.path.append('CosyVoice/third_party/Matcha-TTS')
from cosyvoice.cli.cosyvoice import CosyVoice3
import torchaudio
cosyvoice = CosyVoice3(model_dir='Fun-CosyVoice3-0.5B-2512', fp16=True)
# 1. inference_zero_shot
def demo01():
prompt_text = 'You are a helpful assistant.<|endofprompt|>Maintaining your ability to learn translates into increased marketability, improved career options and higher salaries.'
prompt_wav = 'audio/spk-3.wav'
tts_text = 'The device would work during the day as well, if you took steps to either block direct sunlightor point it away from the sun.'
generator = cosyvoice.inference_zero_shot(tts_text=tts_text, prompt_text=prompt_text, prompt_wav=prompt_wav)
for chunk in generator:
torchaudio.save('demo/demo01-1.wav', chunk['tts_speech'], cosyvoice.sample_rate)
tts_text = '空山新雨后,天气晚来秋。明月松间照,清泉石上流。竹喧归浣女,莲动下渔舟。随意春芳歇,王孙自可留。'
generator = cosyvoice.inference_zero_shot(tts_text=tts_text, prompt_text=prompt_text, prompt_wav=prompt_wav)
for chunk in generator:
torchaudio.save(f'demo/demo01-2.wav', chunk['tts_speech'], cosyvoice.sample_rate)
# 2. inference_cross_lingual
def demo02():
prompt_wav = 'audio/spk-3.wav'
tts_text = 'You are a helpful assistant.<|endofprompt|>The device would work during the day as well, if you took steps to either block direct sunlightor point it away from the sun.'
generator = cosyvoice.inference_cross_lingual(tts_text=tts_text, prompt_wav=prompt_wav)
for chunk in generator:
torchaudio.save('demo/demo02-1.wav', chunk['tts_speech'], cosyvoice.sample_rate)
tts_text = 'You are a helpful assistant.<|endofprompt|>空山新雨后,天气晚来秋。明月松间照,清泉石上流。竹喧归浣女,莲动下渔舟。随意春芳歇,王孙自可留。'
generator = cosyvoice.inference_cross_lingual(tts_text=tts_text, prompt_wav=prompt_wav)
for chunk in generator:
torchaudio.save(f'demo/demo02-2.wav', chunk['tts_speech'], cosyvoice.sample_rate)
# 注意:inference_cross_lingual 适合短文本
# 由于 inference_cross_lingual 要求输入的 tts_text 前面加上 You are a helpful assistant.<|endofprompt|>,
# 但是加上之后,长文本就不会自动切分,会把长文本一次性扔给模型,由于输入文本太长,模型生成的语音会存在发音不准确,有噪声,不完整的问题
# 如果不加 You are a helpful assistant.<|endofprompt|>,会生成多段语音,每段语音的语速、发音可能存在较大差异
# 3. inference_instruct2
def demo03():
prompt_wav = 'audio/spk-1.wav'
# =========== 1. 方言 ===========
tts_text = '那座古老的城堡笼罩在神秘的雾气中,吸引着冒险者前去探索。'
instruct_text = 'You are a helpful assistant. 请用粤语表达,并用合适的语速。<|endofprompt|>'
generator = cosyvoice.inference_instruct2(tts_text=tts_text, instruct_text=instruct_text, prompt_wav=prompt_wav)
torchaudio.save('demo/demo03-1.wav', next(generator)['tts_speech'], cosyvoice.sample_rate)
# =========== 2. 语速 ===========
tts_text = '那座古老的城堡笼罩在神秘的雾气中,吸引着冒险者前去探索。'
instruct_text = 'You are a helpful assistant. 请用尽可能快的语速。<|endofprompt|>'
generator = cosyvoice.inference_instruct2(tts_text=tts_text, instruct_text=instruct_text, prompt_wav=prompt_wav)
torchaudio.save('demo/demo03-2.wav', next(generator)['tts_speech'], cosyvoice.sample_rate)
# =========== 2. 情绪 ===========
tts_text = '这里一片荒凉,没有水,也没有生机,孤独感和无助让我心如刀割。'
instruct_text = 'You are a helpful assistant. 请使用愤怒的情绪。<|endofprompt|>'
generator = cosyvoice.inference_instruct2(tts_text=tts_text, instruct_text=instruct_text, prompt_wav=prompt_wav)
torchaudio.save('demo/demo03-3.wav', next(generator)['tts_speech'], cosyvoice.sample_rate)
# 4. 参考视频预处理
def demo04():
print('可用音色:', cosyvoice.list_available_spks())
prompt_text = 'You are a helpful assistant.<|endofprompt|>如果你对某件事情有强烈的感觉,你应该发声并采取行动。这是我生活的哲学。'
prompt_wav = 'audio/spk-1.wav'
cosyvoice.add_zero_shot_spk(prompt_text=prompt_text, prompt_wav=prompt_wav, zero_shot_spk_id='zh-spk')
prompt_text = 'You are a helpful assistant.<|endofprompt|>Maintaining your ability to learn translates into increased marketability, improved career optionsand higher salaries.'
prompt_wav = 'audio/spk-3.wav'
cosyvoice.add_zero_shot_spk(prompt_text=prompt_text, prompt_wav=prompt_wav, zero_shot_spk_id='en-spk')
# 保存参考音频信息
cosyvoice.save_spkinfo()
tts_text = '空山新雨后,天气晚来秋。明月松间照,清泉石上流。竹喧归浣女,莲动下渔舟。随意春芳歇,王孙自可留。'
generator = cosyvoice.inference_zero_shot(tts_text=tts_text, prompt_text='', prompt_wav='', zero_shot_spk_id='zh-spk')
for chunk in generator:
torchaudio.save(f'demo/demo04.wav', chunk['tts_speech'], cosyvoice.sample_rate)
# 注意:inference_zero_shot、inference_instruct2 虽然有 zero_shot_spk_id 参数,但是实现逻辑有 BUG, 长文本时会报错
if __name__ == '__main__':
demo01()
demo02()
demo03()
demo04()

冀公网安备13050302001966号