IDF Information Theoretic Interpretation
Here is an interpretation from information theory. Suppose a query term appears in documents. Then a randomly picked document will contain the term with probability (where is again the cardinality of the set of documents in the collection). Therefore, the information content of the message " contains " is:
Now suppose we have two query terms and . If the two terms occur in documents entirely independently of each other, then the probability of seeing both and in a randomly picked document is:
and the information content of such an event is:
With a small variation, this is exactly what is expressed by the IDF component of BM25.
Read more about this topic: Okapi BM25
Famous quotes containing the word information:
“As information technology restructures the work situation, it abstracts thought from action.”
—Shoshana Zuboff (b. 1951)