Language models in automatic speech recognition
Keywords:
Language model, Speech recognition, Portuguese languageAbstract
In this paper we describe the work done with the updating and improvement of the language model component of a continuous speech recognition system for the Portuguese language. As a baseline system we used a large vocabulary speech recognition systern for the Portuguese language, developed for a Broadcast News (BN) recognition task. Two sources of performance improvement have been studied: the inclusion of more training data to better estimate the language model parameters, and the use of different discounting and pruning techniques. The results show that using more training data helped to achieve a small relative improvement in recognition accuracy (about 5%). Applying an entropy based pruning technique one can get up to more than 30% size reduction with a slightly increase on perplexity and WER.