Structures Used in Natural Language Processing
- Corpus – body of data, optionally tagged (for example, through part-of-speech tagging), providing real world samples for analysis and comparison.
- Text corpus – large and structured set of texts, nowadays usually electronically stored and processed. They are used to do statistical analysis and hypothesis testing, checking occurrences or validating linguistic rules within a specific subject (or domain).
- Speech corpus – database of speech audio files and text transcriptions. In Speech technology, speech corpora are used, among other things, to create acoustic models (which can then be used with a speech recognition engine). In Linguistics, spoken corpora are used to do research into phonetic, conversation analysis, dialectology and other fields.
Read more about this topic: List Of Natural Language Processing Toolkits
Famous quotes containing the words natural and/or language:
“The enemy is like a woman, weak in face of opposition, but correspondingly strong when not opposed. In a quarrel with a man, it is natural for a woman to lose heart and run away when he faces up to her; on the other hand, if the man begins to be afraid and to give ground, her rage, vindictiveness and fury overflow and know no limit.”
—St. Ignatius Of Loyola (14911556)
“Theres language in her eye, her cheek, her lip,
Nay, her foot speaks; her wanton spirits look out
At every joint and motive of her body.”
—William Shakespeare (15641616)