Structures Used in Natural Language Processing
- Corpus – body of data, optionally tagged (for example, through part-of-speech tagging), providing real world samples for analysis and comparison.
- Text corpus – large and structured set of texts, nowadays usually electronically stored and processed. They are used to do statistical analysis and hypothesis testing, checking occurrences or validating linguistic rules within a specific subject (or domain).
- Speech corpus – database of speech audio files and text transcriptions. In Speech technology, speech corpora are used, among other things, to create acoustic models (which can then be used with a speech recognition engine). In Linguistics, spoken corpora are used to do research into phonetic, conversation analysis, dialectology and other fields.
Read more about this topic: Natural Language Processing Toolkits
Famous quotes containing the words structures, natural and/or language:
“If there are people who feel that God wants them to change the structures of society, that is something between them and their God. We must serve him in whatever way we are called. I am called to help the individual; to love each poor person. Not to deal with institutions. I am in no position to judge.”
—Mother Teresa (b. 1910)
“The great man knew not that he was great. It took a century or two for that fact to appear. What he did, he did because he must; it was the most natural thing in the world, and grew out of the circumstances of the moment. But now, every thing he did, even to the lifting of his finger or the eating of bread, looks large, all-related, and is called an institution.”
—Ralph Waldo Emerson (18031882)
“Please stop using the word Negro.... We are the only human beings in the world with fifty-seven variety of complexions who are classed together as a single racial unit. Therefore, we are really truly colored people, and that is the only name in the English language which accurately describes us.”
—Mary Church Terrell (18631954)