AI实现语音文字处理，PaddleSpeech项目安装使用 | 机器学习

2023年5月27日下午4:59 • 人工智能 • 阅读 130

前言

环境安装

1、conda安装Python3.9虚拟环境

2、安装Visual Studio 2019

3、安装requirements.txt

4、安装paddlepaddle和paddlespeech

前言

这段时间一直在研究飞浆平台，最近试了试PaddleSpeech项目，试着对文本语音做处理。整体的效果个人觉着不算特别优越，只能作为简单的学习使用。

项目github地址：github仓库

环境安装

首先，让我们来看看项目结构和安装文档。

[En]

First, let’s take a look at the project structure and installation documentation.

需要Python3.7以上、C++环境、requirements安装等等，下面按照我的顺序说一下。

1、conda安装Python3.9虚拟环境

使用conda安装python3.9环境，命令如下。

conda create -n py39 python=3.9

2、安装Visual Studio 2019

安装地址: Microsoft C++ 生成工具 – Visual Studio

注意安装的时候需要勾选C++桌面开发。

3、安装requirements.txt

使用命令安装requiremets.txt，命令如下：

pip install -r requirements.txt -i https://pypi.douban.com/simple

这里要注意一下，paddlespeech_ctcdecoders安装失败的话无所谓，可以略掉。

4、安装paddlepaddle和paddlespeech

命令如下：

pip install paddlepaddle -i https://mirror.baidu.com/pypi/simple
pip install paddlespeech -i https://pypi.tuna.tsinghua.edu.cn/simple

5、nltk_data下载

按照项目安装文档中的说明进行操作。

[En]

Follow the instructions in the project installation documentation.

我的本地目录地址如下

项目验证

我下面分别验证一下tts、asr以及标点恢复功能。

tts语音合成

使用命令如下：

paddlespeech tts –input “南京现在很冷，下次再去夫子庙吧。” –output C:\Users\xxx\Desktop\115.wav

执行过程

(dh_partner) D:\spyder\PaddleSpeech>paddlespeech tts –input “南京现在很冷，下次再去夫子庙吧。” –output C:\Users\xxx\Desktop\115.wav
phones_dict: None
[2022-01-05 17:23:43,642] [ INFO] [log.py] [L57] – File C:\Users\huyi.paddlespeech\models\fastspeech2_csmsc-zh\fastspeech2_nosil_baker_ckpt_0.4.zip md5 checking…

[2022-01-05 17:23:44,742] [ INFO] [log.py] [L57] – Use pretrained model stored in: C:\Users\huyi.paddlespeech\models\fastspeech2_csmsc-zh\fastspeech2_nosil_baker_ckpt_0.4
self.phones_dict: C:\Users\huyi.paddlespeech\models\fastspeech2_csmsc-zh\fastspeech2_nosil_baker_ckpt_0.4\phone_id_map.txt
[2022-01-05 17:23:44,743] [ INFO] [log.py] [L57] – C:\Users\huyi.paddlespeech\models\fastspeech2_csmsc-zh\fastspeech2_nosil_baker_ckpt_0.4
[2022-01-05 17:23:44,744] [ INFO] [log.py] [L57] – C:\Users\huyi.paddlespeech\models\fastspeech2_csmsc-zh\fastspeech2_nosil_baker_ckpt_0.4\default.yaml
[2022-01-05 17:23:44,744] [ INFO] [log.py] [L57] – C:\Users\huyi.paddlespeech\models\fastspeech2_csmsc-zh\fastspeech2_nosil_baker_ckpt_0.4\snapshot_iter_76000.pdz
self.phones_dict: C:\Users\huyi.paddlespeech\models\fastspeech2_csmsc-zh\fastspeech2_nosil_baker_ckpt_0.4\phone_id_map.txt
[2022-01-05 17:23:44,745] [ INFO] [log.py] [L57] – File C:\Users\huyi.paddlespeech\models\pwgan_csmsc-zh\pwg_baker_ckpt_0.4.zip md5 checking…

[2022-01-05 17:23:44,782] [ INFO] [log.py] [L57] – Use pretrained model stored in: C:\Users\huyi.paddlespeech\models\pwgan_csmsc-zh\pwg_baker_ckpt_0.4
[2022-01-05 17:23:44,783] [ INFO] [log.py] [L57] – C:\Users\huyi.paddlespeech\models\pwgan_csmsc-zh\pwg_baker_ckpt_0.4
[2022-01-05 17:23:44,783] [ INFO] [log.py] [L57] – C:\Users\huyi.paddlespeech\models\pwgan_csmsc-zh\pwg_baker_ckpt_0.4\pwg_default.yaml
[2022-01-05 17:23:44,785] [ INFO] [log.py] [L57] – C:\Users\huyi.paddlespeech\models\pwgan_csmsc-zh\pwg_baker_ckpt_0.4\pwg_snapshot_iter_400000.pdz
vocab_size: 268
frontend done!

encoder_type is transformer
decoder_type is transformer
C:\Users\huyi.conda\envs\dh_partner\lib\site-packages\paddle\framework\io.py:415: DeprecationWarning: Using or importing the ABCs from ‘collections’ instead of from ‘collections.abc’ i
s deprecated since Python 3.3, and in 3.10 it will stop working
if isinstance(obj, collections.Iterable) and not isinstance(obj, (
acoustic model done!

voc done!

Building prefix dict from the default dictionary …

[2022-01-05 17:23:51] [DEBUG] [init.py:113] Building prefix dict from the default dictionary …

Loading model from cache C:\Users\huyi\AppData\Local\Temp\jieba.cache
[2022-01-05 17:23:51] [DEBUG] [init.py:132] Loading model from cache C:\Users\huyi\AppData\Local\Temp\jieba.cache
Loading model cost 0.659 seconds.

[2022-01-05 17:23:52] [DEBUG] [init.py:164] Loading model cost 0.659 seconds.

Prefix dict has been built successfully.

[2022-01-05 17:23:52] [DEBUG] [init.py:166] Prefix dict has been built successfully.

C:\Users\huyi.conda\envs\dh_partner\lib\site-packages\paddle\fluid\dygraph\math_op_patch.py:251: UserWarning: The dtype of left and right variables are not the same, left dtype is padd
le.int64, but right dtype is paddle.int32, the right dtype will convert to paddle.int64
warnings.warn(
[2022-01-05 17:23:58,811] [ INFO] [log.py] [L57] – Wave file has been generated: C:\Users\xxx\Desktop\115.wav

生成的音频如下

asr语音识别

我就使用了tts生成的音频进行asr识别，看看效果，命令如下:

paddlespeech asr –lang zh –input C:\Users\xxx\Desktop\115.wav

执行结果如下

可以看到，最后打印的内容都是不加标点符号的文本输出，还是比较准确的。

[En]

You can see that the last printed content is unpunctuated text output, or relatively accurate.

标点恢复

试着用这个句子恢复标点符号。命令如下：

[En]

Try punctuation recovery with this sentence. The command is as follows:

paddlespeech text –task punc –input 南京现在很冷下次再去夫子庙吧

执行结果

在语义上似乎没有什么问题。

[En]

There seems to be nothing wrong with the semantics.

总结

我在前言中说效果不是很好的主要原因是因为速率比较慢，相比于类似阿里云提供的tts、asr接口来说，效率比较低。也可能和需要校验模型是否存在这些无关紧要的功能有关。可以考虑研究代码，自己重新封装一些服务，效果应该好的多。

还有补充一下，最近博主在参加评选 博客之星活动。如果你喜欢我的文章的话，不妨给我点个五星，投投票吧，谢谢大家的支持！！链接地址：https://bbs.csdn.net/topics/603956455

这个世界不会在意你的自尊，人们只会看到你的成就。在你有所成就之前，不要过分强调自尊。了不起的盖茨比

[En]

The world will not care about your self-esteem, people will only see your achievements. Don’t put too much emphasis on self-esteem until you have achieved something. The Great Gatsby

如果本文对你有用的话， 点个赞吧，谢谢！！！

Original: https://blog.csdn.net/zhiweihongyan1/article/details/122326644
Author: 剑客阿良_ALiang
Title: AI实现语音文字处理，PaddleSpeech项目安装使用 | 机器学习

原创文章受到原创版权保护。转载请注明出处：https://www.johngo689.com/526990/

转载文章受原作者版权保护。转载请注明原作者出处！

人工智能

【自取】最近整理的，有需要可以领取学习：

Linux核心资料大放送~

全栈面试题汇总（持续更新&可下载）

一个提高学习100%效率的工具！

【超详细】深度学习面试题目！

LeetCode Python刷题答案下载！

LeetCode Java版刷题答案下载！

LeetCode C++ 版本，抓紧保存！

LeetCode GO语言刷题答案下载！

准确率、召回率、F1值的思考

简述概念准确率（Accuracy）准确率（ACC）, 所有预测正确的占总样本的比重。精确率/查准率（Precision）精确率（P）：精确率/查准率，表示正确预测为正的占全…

人工智能 2023年5月30日
0059
【pytorch实战学习】第七篇：tensorboard可视化介绍

【pytorch学习实战】第一篇：线性回归【pytorch学习实战】第二篇：多项式回归【pytorch学习实战】第三篇：逻辑回归【pytorch学习实战】第四篇：MNIST数…

人工智能 2023年7月13日
0060
sklearn多分类求AUC，多分类report

目录 1.多分类求AUC 2.多分类sklearn report：注意：logits和score不同，二者关系是：score=softmax（logits）或sigmod（log…

人工智能 2023年6月30日
0083
语音合成论文优选：Parallel Tacotron 2: A Non-Autoregressive Neural TTS Model with Differentiable Duration Mod

免责声明：首选系列演讲合成论文以分享论文为主，分享论文不直接翻译，内容主要是我对论文内容的总结和个人观点。如果是转载，请注明出处。 [En] Disclaimer: the pre…

人工智能 2023年5月27日
0098
深蓝学院-视觉SLAM课程-第8讲笔记-回环检测

课程Github地址：https://github.com/wrk666/VSLAM-Course/tree/master最后一次课了，加油！内容 ; 1. 回环检测与词袋回顾…

人工智能 2023年6月2日
0085
python如何安装keras和tensorflow

目录一. 通过pip install kears 安装 keras 二. 安装tensorflow * 1.报错 No module named ‘tensorflo…

人工智能 2023年6月25日
0074
python error tokenizing data_python 问题杂烩

python 问题杂烩 python problem cookbook ParserError: Error tokenizing data. C error: Calling r…

人工智能 2023年7月8日
0083
机器学习深度神经网络——实验报告

机器学习实验报告〇、实验报告pdf可在该网址下载一、实验目的与要求二、实验内容与方法 * 2.1 深度神经网络的知识回顾 – 2.1.1 神经元模型 2.1.2 …

人工智能 2023年6月23日
0069
【OpenCV 例程300篇】03. 图像的显示（cv2.imshow）

专栏地址：『youcans 的 OpenCV 例程300篇 – 总目录』01. 图像的读取（cv2.imread）02. 图像的保存（cv2.imwrite）03. 图…

人工智能 2023年5月26日
00116
RAFT:使用深度学习的光流估计

在这篇文章中，我们将讨论两种基于深度学习的使用光流进行运动估计的方法。FlowNet是第一种用于计算光流的CNN方法，RAFT是目前最先进的估算光流的方法。我们还将看到如何使用作者…

人工智能 2023年7月13日
00128
[Pandas数据处理Debug记录]DataFrame.apply的使用

[Pandas数据处理Debug记录]DataFrame.apply的使用 Pandas数据处理Debug记录 * 问题1：apply函数通过args的传参抛出ValueError…

人工智能 2023年7月8日
0096
DDPM代码详细解读(2)：Unet结构、正向和逆向过程、IS和FID测试、EMA优化

以下是将 Unet_和门 _结构_结合的 _PyTorch 代码： import torch import torch.nn as nn import torch.nn.funct…

人工智能 2023年6月17日
0098
OpenCV学习（33）

一，相关OpenCV源码分析溯源在…lopencvlsourceslmodules\timgproclsrcl morph.cpp 路径中，我们可以发现erode（腐…

人工智能 2023年6月22日
0076
经济型EtherCAT运动控制器(八）：轴参数与运动指令

一、XPLC006E功能简介 XPLC006E是正运动运动控制器推出的一款多轴经济型EtherCAT总线运动控制器，XPLC系列运动控制器可应用于各种需要脱机或联机运行的场合。 X…

人工智能 2023年7月7日
0047
Sox(Sound eXchange)一款强大的音频处理工具格式转化、切割音频、合并音频等

Sox(Sound eXchange)是一款强大的音频处理工具，能够合并、拆分多通道；能播放能录音；可以截取音频的某一部分或删除开头结尾部分。能满足大部分音频处理的操作需求。安装…

人工智能 2023年5月27日
0075
（附源码）springboot社区疫情防控管理系统毕业设计 164621

系统实现 5.1 用户前台功能模块社区疫情防控管理系统，在系统首页通知公告、出入预约、隔离申请、打卡信息、社区疫情等内容，如图5-1所示。图5-1首页界面图登录，在登录页面通…

人工智能 2023年7月29日
0053

2024 年 5 月
一	二	三	四	五	六	日
		1	2	3	4	5
6	7	8	9	10	11	12
13	14	15	16	17	18	19
20	21	22	23	24	25	26
27	28	29	30	31

AI实现语音文字处理，PaddleSpeech项目安装使用 | 机器学习

1、conda安装Python3.9虚拟环境

2、安装Visual Studio 2019

3、安装requirements.txt

4、安装paddlepaddle和paddlespeech

5、nltk_data下载

tts语音合成

asr语音识别

标点恢复

大家都在看