Telephony-based Speech Recognition
The limiting factor for telephony based speech recognition is the bandwidth at which speech can be transmitted. For example, a standard land-line telephone only has a bandwidth of 64 kbit/s at a sampling rate of 8 kHz and 8-bits per sample (8000 samples per second * 8-bits per sample = 64000 bit/s). Therefore, for telephony based speech recognition, acoustic models should be trained with 8 kHz/8-bit speech audio files.
In the case of Voice over IP, the codec determines the sampling rate/bits per sample of speech transmission. Codecs with a higher sampling rate/bits per sample for speech transmission (which improve the sound quality) necessitate acoustic models trained with audio data that matches that sampling rate/bits per sample.
Read more about this topic: Acoustic Model
Famous quotes containing the words speech and/or recognition:
“It is clear that not in one thing alone, but in many ways equality and freedom of speech are a good thing.”
—Herodotus (c. 484424 B.C.)
“The recognition of Russia on November 16, 1933, started forces which were to have considerable influence in the attempt to collectivize the United States.”
—Herbert Hoover (18741964)