【miscellaneous】使用Google语音识别引擎（Google Speech API）[3月5日修改]

wget -O "GoogleSpeechAPI.txt" --user-agent="Mozilla/5.0" --post-file=test.flac --header="Content-Type: audio/x-flac; rate=16000" "http://www.google.com/speech-api/v1/recognize?xjerr=1&client=chromium&lang=zh-CN&maxresults=1"

结果如下：

[javascript] view
plain copy

print ?

{
"status":0, /* 结果代码，详细见本文结尾 */
"id":"c421dee91abe31d9b8457f2a80ebca91-1", /* 识别编号 */
"hypotheses": /* 假设，即结果 */
[
{
"utterance":"下午好", /* 话语 */
"confidence":0.2507637 /* 信心，即准确度 */
}
]
}

注：注释后为手工添加的结果解释

返回结果太明了了！直接就能拿来用了不是~ 返回的编码是UTF-8。

对于编码格式，在测试中使用了FLAC编码，采样率为16kHz，经测试其他采样率同样可用，但一定要保证Header里的rate与实际数据相符。（关于其他格式的实验请看本文底部。）

总结：

1、基本流程：

一、从音频输入设备获取原始数据。

二、对原始数据进行包装、编码。

三、将编码后的音频POST至接口地址。

四、分析处理接口返回的JSON并得出结果。

2、请求接口

地址：http://www.google.com/speech-api/v1/recognize?xjerr=1&client=chromium&lang=zh-CN&maxresults=1

请求方式：HTTP POST

头部信息：Content-Type: audio/x-flac; rate=16000 （注：Content-Type根据所使用的编码格式不同而不同，详见文章底部。rate为音频采样率。）

请求数据：编码后的音频数据

3、音频编码格式：

FLAC或WAV或SPEEX

下面是我写的Qt(C++)中的请求：

[cpp] view
plain copy

print ?

void Protocol::Request_SPEECH(QByteArray & audioData)
{
if (!Nt_SPEECH)
{
QNetworkRequest request;
QString speechAPI = "http://www.google.com/speech-api/v1/recognize?xjerr=1&client=chromium&lang=zh-CN&maxresults=1";
request.setUrl(speechAPI);
request.setRawHeader("User-Agent", "Mozilla/5.0");
request.setRawHeader("Content-Type", "audio/x-flac; rate=16000");
Nt_SPEECH = NetworkMGR.post(request, audioData);
connect(Nt_SPEECH, SIGNAL(readyRead()), this, SLOT(Read_SPEECH()));
}
}

至于读取函数，就不贴在这里了，具体见：

Protocol: http://pastebin.com/6G6wggfF

AudioInput:

speechInput.h: http://pastebin.com/qdMPeWZD

speechInput.cpp: http://pastebin.com/567B47qF

main:

mainwidget: http://pastebin.com/c8bk7zd2

在翻阅Chromium源码的过程之中，还发现了其他有用的东西：

Speech Input API Specification http://www.w3.org/2005/Incubator/htmlspeech/2010/10/google-api-draft.html

到目前为止，Google好像还没有公开这个API，使用许可依旧不详，请求也没有用到任何认证。但它确实能用，而且十分方便，对于编写非商业程序的人来说，这个东西真的是再好不过了（因为它有着高的爆表的识别率）。

参考：

Chromium Repository http://src.chromium.org/viewvc/chrome/trunk/src/content/browser/speech/

Accessing Google Speech API / Chrome 11 http://mikepultz.com/2011/03/accessing-google-speech-api-chrome-11/

附：

1、SpeechInputError interface 错误信息

[cpp] view
plain copy

print ?

// This enumeration follows the values described here:
// http://www.w3.org/2005/Incubator/htmlspeech/2010/10/google-api-draft.html#speech-input-error
enum SpeechInputError {
// There was no error.
SPEECH_INPUT_ERROR_NONE = 0,
// The user or a script aborted speech input.
SPEECH_INPUT_ERROR_ABORTED,
// There was an error with recording audio.
SPEECH_INPUT_ERROR_AUDIO,
// There was a network error.
SPEECH_INPUT_ERROR_NETWORK,
// No speech heard before timeout.
SPEECH_INPUT_ERROR_NO_SPEECH,
// Speech was heard, but could not be interpreted.
SPEECH_INPUT_ERROR_NO_MATCH,
// There was an error in the speech recognition grammar.
SPEECH_INPUT_ERROR_BAD_GRAMMAR,
};

2、多种音频格式的测试

收到朋友的邮件说使用flac实在是很不方便，问我有没有更好的解决方法，于是我尝试将其他编码格式应用于Google Speech API。以下为结果：

1、WAV格式

请求Header：Content-Type: audio/L16; rate=16000

返回结果：识别成功

2、MP3格式

请求Header：Content-Type: audio/mpeg; rate=16000

返回结果：无法识别的编码

请求Header：Content-Type: audio/mpeg3; rate=16000

返回结果：无法识别的编码

请求Header：Content-Type: audio/x-mpeg; rate=16000

返回结果：无法识别的编码

请求Header：Content-Type: audio/x-mpeg-3; rate=16000

返回结果：无法识别的编码

请求Header：Content-Type: audio/mp3; rate=16000

返回结果：无法识别的编码

3、PCM格式

请求Header：Content-Type: audio/x-ogg-pcm; rate=16000

返回结果：无法识别的编码

请求Header：Content-Type: audio/pcm; rate=16000

返回结果：无法识别的编码

4、SPEEX格式

请求Header：Content-Type: audio/x-speex-with-header-byte; rate=16000

返回结果：识别成功

请求Header：Content-Type: audio/speex; rate=16000

返回结果：识别成功

由于识别接口并不开放，所以无法得知具体的支持格式，如果哪位朋友发现了新的支持格式，请一定要留言哦！

巴特西

【miscellaneous】使用Google语音识别引擎（Google Speech API）[3月5日修改]

最新文章

热门文章