Tech Blog
  • HOME
  • Blog
  • Comparing the speech recognition rates of OpenAI's Whisper and AmiVoice for "conference" audio

Comparing the speech recognition rates of OpenAI's Whisper and AmiVoice for "conference" audio

Published: 2023.06.26 Last updated: 2025.03.06

ando
Shogo Ando

Hello everyone.

In this article recently, we compared the recognition rates of OpenAI's Whisper and AmiVoice.

Looking back at this article, the results were quite predictable —"AmiVoice wins in areas where it excels, and Whisper wins in areas where it excels"— so this time we will compare them from the perspective of "which performs better with meeting audio?"

Verification method

The verification methods and conditions are as follows:

  • We used audio from a meeting provided by one of our clients for research and development purposes.
  • The conference audio was from four different industries, and the first 10 minutes of each was used, totaling approximately 40 minutes.
  • The above audio and its transcripts are not used to train the speech recognition engine.
  • Both AmiVoice and Whisper performed speech recognition processing in our local environment.
  • The speech recognition engines used were an equivalent of AmiVoice API's "Conversation_General Purpose" and Whisper's "large" and "large-v2".
  • Speech recognition processing was performed on AmiVoice around June 2022, on Whisper (large) around October 2022, and on Whisper (large-v2) in February 2023, and the latest versions at that time were used for each.
  • Speech recognition accuracy was measured character by character (not word by word).
  • We have checked and corrected any misrecognitions due to variations in spelling through automatic conversion and visual inspection. However, since we were checking visually, there may still be some oversights.
  • Filler (unnecessary words) are removed from both the correct sentences andspeech recognition results before calculation.

As a key point here, the conference audio in question comes from relatively large companies and organizations, where proceedings are properly moderated with individual microphones for each speaker, and speakers deliver their remarks at an appropriate volume to address all participants (not muttering or speaking in hushed tones). This audio is easy to understand even for human listeners, which means that the difficulty level of speech recognition for these conference recordings can be considered relatively low.

Measurement results

AmiVoice (Conversation_General-purpose)

Data Number of correct characters Number of insertion errors Number of deletion errors Number of substitution errors Speech recognition accuracy
2905 33 46 34 96.11 %
2616 23 21 27 97.29 %
3011 38 23 44 96.51 %
3047 36 21 39 96.85 %
Total 11579 130 111 144 96.68 %


*
Whisper (large)

Data Number of correct characters Number of insertion errors Number of deletion errors Number of substitution errors Speech recognition accuracy
2903 49 113 199 87.56 %
2630 37 190 60 89.09 %
3009 47 138 146 89.00 %
3046 65 223 90 87.59 %
Total 11588 198 664 495 88.29 %


*
Whisper (large-v2)

Data Number of correct characters Number of insertion errors Number of deletion errors Number of substitution errors Speech recognition accuracy
2901 60 132 161 87.83 %
2625 67 204 55 87.58 %
3011 40 117 122 90.73 %
3044 65 189 85 88.86 %
Total 11581 232 642 423 88.80 %

Consideration

AmiVoice produced the best results.In terms of character error rate (CER), AmiVoice achieved 100% - 96.68% = 3.32%, while Whisper (large) achieved 11.71% and Whisper (large-v2) achieved 11.20%, representing an extremely significant difference of more than three times.

AmiVoice's speech recognition rate of 96.68% is a relatively high level among our in-house experiments. The microphone was set up properly, the speaker spoke clearly to the participants, and the voice was easy for humans to hear, which is probably why it achieved such high accuracy.

What concerns me is that Whisper makes a fair number of false positives. When I checked the false positives, I found that there were two main types of false positives.

  • Whisper misrecognition pattern 1: Incorrect recognition of words with similar pronunciation, ignoring context

    • 競合→今日後
    • 好調に推移→好調に注意
    • どこも投資をしている→ドグも投資している
    • ◯◯か何かで→◯◯内科で
  • Whisper misrecognition pattern 2: Unusual Kanji conversions
    • 加重平均→過充平均
    • 四半期報告→市販機報告
    • 余資運用→吉運用
    • 残存期間→暫存期間
    • 外貨建て→外科だて
    • 一過性→一家性

Pattern 1 may sound a little like "そう聞こえなくもないかな?" to a human listener, but when you consider the context and other factors, it seems like a strange misrecognition. (This pattern also occurs with AmiVoice, but it was more frequent with Whisper in this audio.)

Pattern 2 uses kanji conversion that doesn't come up much in web searches, so it's a mystery why it produced such an output.

Furthermore, Whisper has a conspicuously high number of deletion errors. When I investigated the cause, I was surprised to find that Whisper has a tendency to intelligently delete sentences it deems unnecessary. For example, Whisper may output the following for utterances such as those shown below. *1

  • Speaker: "1月、いや2月の1日、あれ2月1日じゃなかったでしたっけすいません、え、いいんでしたっけ、その日の件ですが"
  • Whisper: "2月1日の件ですが" *2

While this processing is thought to improve the readability of speech recognition results, the omitted portions are treated as deletion errors in calculations, as these results cannot be considered correct. Deletion errors resulting from this behavior may not be problematic depending on the application, so it might be acceptable to discount them to some extent during evaluation.

Summary

This time, we compared AmiVoice and Whisper using conference audio.

The results showed a significant difference, with AmiVoice having an error rate (CER) of less than one-third.

Whisper seems to have a tendency to output similar-sounding words without considering the context, and it often misrecognizes words by converting them to kanji characters that are not commonly used in Japanese.

Whisper also tends to intelligently delete unnecessary text, which makes it prone to more deletion errors. These errors may not be a problem depending on your use case, so you might consider discounting them to some extent when comparing and evaluating.

Person who wrote this article

  • Shogo Ando

    While researching speech recognition, I found a local speech tech firm and decided to join the team, where I continue to work to this day.

    My hobbies include overseas travel, trying great food and visiting saunas.

    x : @anpyan

*1:Due to the sensitivity of the data, the actual speech has been dramatized.

*2:Just to be sure, I ran Whisper through the speech recognition of just this sentence, and the entire sentence was output in the speech recognition results without any omissions. Perhaps the system changes its behavior depending on the context and overall structure of the sentence.

Use API for Free