Comparing the speech recognition rates of OpenAI's Whisper and AmiVoice for "conference" audio


Hello everyone.
In this article recently, we compared the recognition rates of OpenAI's Whisper and AmiVoice.
Looking back at this article, the results were quite predictable —"AmiVoice wins in areas where it excels, and Whisper wins in areas where it excels"— so this time we will compare them from the perspective of "which performs better with meeting audio?"
Verification method
The verification methods and conditions are as follows:
- We used audio from a meeting provided by one of our clients for research and development purposes.
- The conference audio was from four different industries, and the first 10 minutes of each was used, totaling approximately 40 minutes.
- The above audio and its transcripts are not used to train the speech recognition engine.
- Both AmiVoice and Whisper performed speech recognition processing in our local environment.
- The speech recognition engines used were an equivalent of AmiVoice API's "Conversation_General Purpose" and Whisper's "large" and "large-v2".
- Speech recognition processing was performed on AmiVoice around June 2022, on Whisper (large) around October 2022, and on Whisper (large-v2) in February 2023, and the latest versions at that time were used for each.
- Speech recognition accuracy was measured character by character (not word by word).
- We have checked and corrected any misrecognitions due to variations in spelling through automatic conversion and visual inspection. However, since we were checking visually, there may still be some oversights.
- Filler (unnecessary words) are removed from both the correct sentences andspeech recognition results before calculation.
As a key point here, the conference audio in question comes from relatively large companies and organizations, where proceedings are properly moderated with individual microphones for each speaker, and speakers deliver their remarks at an appropriate volume to address all participants (not muttering or speaking in hushed tones). This audio is easy to understand even for human listeners, which means that the difficulty level of speech recognition for these conference recordings can be considered relatively low.
Measurement results
■ AmiVoice (Conversation_General-purpose)
| Data | Number of correct characters | Number of insertion errors | Number of deletion errors | Number of substitution errors | Speech recognition accuracy |
| ① | 2905 | 33 | 46 | 34 | 96.11 % |
| ② | 2616 | 23 | 21 | 27 | 97.29 % |
| ③ | 3011 | 38 | 23 | 44 | 96.51 % |
| ④ | 3047 | 36 | 21 | 39 | 96.85 % |
| Total | 11579 | 130 | 111 | 144 | 96.68 % |
*Whisper (large)
| Data | Number of correct characters | Number of insertion errors | Number of deletion errors | Number of substitution errors | Speech recognition accuracy |
| ① | 2903 | 49 | 113 | 199 | 87.56 % |
| ② | 2630 | 37 | 190 | 60 | 89.09 % |
| ③ | 3009 | 47 | 138 | 146 | 89.00 % |
| ④ | 3046 | 65 | 223 | 90 | 87.59 % |
| Total | 11588 | 198 | 664 | 495 | 88.29 % |
*Whisper (large-v2)
| Data | Number of correct characters | Number of insertion errors | Number of deletion errors | Number of substitution errors | Speech recognition accuracy |
| ① | 2901 | 60 | 132 | 161 | 87.83 % |
| ② | 2625 | 67 | 204 | 55 | 87.58 % |
| ③ | 3011 | 40 | 117 | 122 | 90.73 % |
| ④ | 3044 | 65 | 189 | 85 | 88.86 % |
| Total | 11581 | 232 | 642 | 423 | 88.80 % |
Consideration
AmiVoice produced the best results.In terms of character error rate (CER), AmiVoice achieved 100% - 96.68% = 3.32%, while Whisper (large) achieved 11.71% and Whisper (large-v2) achieved 11.20%, representing an extremely significant difference of more than three times.
AmiVoice's speech recognition rate of 96.68% is a relatively high level among our in-house experiments. The microphone was set up properly, the speaker spoke clearly to the participants, and the voice was easy for humans to hear, which is probably why it achieved such high accuracy.
What concerns me is that Whisper makes a fair number of false positives. When I checked the false positives, I found that there were two main types of false positives.
- Whisper misrecognition pattern 1: Incorrect recognition of words with similar pronunciation, ignoring context
- 競合→今日後
- 好調に推移→好調に注意
- どこも投資をしている→ドグも投資している
- ◯◯か何かで→◯◯内科で
- Whisper misrecognition pattern 2: Unusual Kanji conversions
- 加重平均→過充平均
- 四半期報告→市販機報告
- 余資運用→吉運用
- 残存期間→暫存期間
- 外貨建て→外科だて
- 一過性→一家性
Pattern 1 may sound a little like "そう聞こえなくもないかな?" to a human listener, but when you consider the context and other factors, it seems like a strange misrecognition. (This pattern also occurs with AmiVoice, but it was more frequent with Whisper in this audio.)
Pattern 2 uses kanji conversion that doesn't come up much in web searches, so it's a mystery why it produced such an output.
Furthermore, Whisper has a conspicuously high number of deletion errors. When I investigated the cause, I was surprised to find that Whisper has a tendency to intelligently delete sentences it deems unnecessary. For example, Whisper may output the following for utterances such as those shown below. *1
- Speaker: "1月、いや2月の1日、あれ2月1日じゃなかったでしたっけすいません、え、いいんでしたっけ、その日の件ですが"
- Whisper: "2月1日の件ですが" *2
While this processing is thought to improve the readability of speech recognition results, the omitted portions are treated as deletion errors in calculations, as these results cannot be considered correct. Deletion errors resulting from this behavior may not be problematic depending on the application, so it might be acceptable to discount them to some extent during evaluation.
Summary
This time, we compared AmiVoice and Whisper using conference audio.
The results showed a significant difference, with AmiVoice having an error rate (CER) of less than one-third.
Whisper seems to have a tendency to output similar-sounding words without considering the context, and it often misrecognizes words by converting them to kanji characters that are not commonly used in Japanese.
Whisper also tends to intelligently delete unnecessary text, which makes it prone to more deletion errors. These errors may not be a problem depending on your use case, so you might consider discounting them to some extent when comparing and evaluating.
Person who wrote this article

Shogo Ando
While researching speech recognition, I found a local speech tech firm and decided to join the team, where I continue to work to this day.
My hobbies include overseas travel, trying great food and visiting saunas.
: @anpyan
*1:Due to the sensitivity of the data, the actual speech has been dramatized.
*2:Just to be sure, I ran Whisper through the speech recognition of just this sentence, and the entire sentence was output in the speech recognition results without any omissions. Perhaps the system changes its behavior depending on the context and overall structure of the sentence.
Most Viewed Articles
- A quick explanation of how speech recognition works!
- Comparing the speech recognition rates of OpenAI's Whisper and AmiVoice for "conference" audio
- How to use the AmiVoice API free coupon
New Articles
- How to Choose a Speech Recognition API: 4 Comparison Points to Avoid Mistakes
- Things to know before implementation: the accuracy of Japanese speech recognition
- Thank you for all your submissions! Here's a summary of Zennfes Spring 2026.
Category List
- Tech (1)
- Introduction to Speech Recognition (16)
- How to Improve Speech Recognition Accuracy (12)
- Tried Building It (27)
- How to Use AmiVoice API (28)
- Comparison and Verification (7)
- Others (10)
