The joint probability distribution
|
So This could also be written: |
Notice that the probabilities in the table sum to 1
We can learn a joint probability distribution from data by counting occurrences and computing relative frequencies.
Example: Suppose we examine 10 patients and observe:
|
Data
|
Count each combination:
|
|
|
As we move from specific examples to general probability identities, we shift our notation:
Event-based notation:
Variable-based notation:
This shift allows us to write general identities like:
|
|
We can use Bayes rule to build a classifier:
Where
How to handle zeros for some attributes?
Solution: Laplace Smoothing (Add-one smoothing)
Example: Suppose our 10 patients were also checked for a rash, and none of the 3 flu patients had one:
Probability prologue, moved here from prob.md so the first probability meeting stays short enough to serve the bias/variance activity. About 22 minutes. Supplies the derivation the "Recall the naive Bayes classifier" slide below depends on.
P(x) = .2 P(~x) = .8 P(c | x) = .2 P(f | x) = .7 P(h | x) = .1 P(c | ~x) = .1 P(f | ~x) = .05 P(h | ~x) = .85