Estimating the entropy of English (Problem 14.1-5 of textbook) Estimate the information per lette...
Estimating the entropy of English (Problem 14.1-5 of textbook) Estimate the information per letter in the English language by various method but is good enough to get a rough idea.) ter is independent of the others. (This is not true, (a) In the first method, assume that all 27 characters (26 letters and a space) are equiprobable. This is a gross approximation, but good for a (b) In the second method, use the table of probabilities of various characters (Attachment). (e) Use Zipf's law relating the word rank to its probability. In English prose, if we order words according to the frequency of usage so that the most frequently used word (the) is word number 1 (rank 1), the next most probable word (of) is number 2 (rank 2), and so on, then empirically it is found that P(r), the probability of the rth word (rank r) is very nearly P(r) = 01 Now use Zipf's law to compute the entropy per word. Assume that there are 8727 words. The reason for this number is that the probabilities P/r) sum to 1 for r from 1 to 8727 Zipf's law, surprisingly, gives reasonably good results. Assuming there are 5.5 letters (including space) per word on the average, determine the entropy or information per letter. Table P14.1-5.pdf