JMU CS 445: Probability and Naive Bayes

Disease and Fever Example

Consider a medical scenario where we observe a single patient:

  • Experiment: Examine one patient and record their disease status and fever status
  • Sample Space:

  • Disease Random Variable ():
    • is shorthand for the event
  • Fever Random Variable ():
    • is shorthand for the event

Joint Probability Distributions

The joint probability distribution assigns a probability to every combination of values of the random variables and :

X Y P(X,Y)
T cold .04
T flu .14
T healthy .02
F cold .08
F flu .04
F healthy
.68

So .

This could also be written:

Notice that the probabilities in the table sum to 1

"Learning" a Probability Distributions

We can learn a joint probability distribution from data by counting occurrences and computing relative frequencies.

Example: Suppose we examine 10 patients and observe:

Data

Patient Disease Fever
1coldyes
2healthyno
3fluyes
4healthyno
5coldno
6fluyes
7healthyyes
8fluno
9coldyes
10healthyno

Count each combination:

  • (cold, yes fever): 2 patients
  • (cold, no fever): 1 patient
  • (flu, yes fever): 2 patients
  • (flu, no fever): 1 patient
  • (healthy, yes fever): 1 patient
  • (healthy, no fever): 3 patients
  • Compute probabilities:

  • Estimated Joint Distribution
    Fever Disease P(X,Y)
    Tcold0.2
    Tflu0.2
    Thealthy0.1
    Fcold0.1
    Fflu0.1
    Fhealthy0.3

Notation: Events vs. Random Variables

As we move from specific examples to general probability identities, we shift our notation:

  • Event-based notation:

    • refers to the probability of a specific event — the random variables and taking on values and .
  • Variable-based notation:

    • refers to the joint distribution of the random variables and — a function that assigns probabilities to all combinations of values.
  • This shift allows us to write general identities like:

Marginalization

X Y P(X,Y)
T cold .04
T flu .14
T healthy .02
F cold .08
F flu .04
F healthy
.68
  • Given the joint probability distribution we can use marginalization to retrieve the probability distribution for any individual variable:
  • For example:
    • The probability that someone has the flu, regardless of fever:

    • What is the probability that someone has a fever regardless of health?


Conditional Probability

  • Definition:
  • For example:
  • It follows that:
    • (the chain rule)

Independence and Conditional Independence

  • Two random variables and are independent if and only if:
    • Equivalently:
  • and are conditionally independent given , if and only if:
    • Equivalently:

Bayes Theorem / Bayes Rule

  • Note that (by combining marginalization with the chain rule):

  • So Bayes rule can be expressed as:

Bayes Classifier

We can use Bayes rule to build a classifier:

Where corresponds to the class label and each is an attribute.

  • There is a serious problem with this! What is it?

Naive Bayes Classifier

  • We assume that the attributes are conditionally independent given class labels, so:
  • We can also recognize that is the same regardless of class so dividing with it won't change the class with the largest value.
  • These leads to the naive Bayes classifier:

Properties of Naive-Bayes

  • Pros:
    • Provides a meaningful class probability, not just a class label
    • Works in the face of missing attributes (just don't include them in the calculation)
    • Relatively easy to interpret: we can examine the class-conditional probabilities for individual attributes.
  • Cons:
    • Classification performance may be worse than other classifiers: Most real classification tasks will violate the independence assumption to some extent.

Implementation Issues - 1

  • Naive Bayes classifier:

  • Each is less then 1.
  • What is ? ?
  • Recall that
  • Also, the log function is monotonic: if then
  • So, practical implementations generally work with logs:

Implementation Issues - 2

  • How to handle zeros for some attributes?

    • If for some attribute value , then the entire product becomes 0
    • This means regardless of other evidence
    • Problem: A single zero probability can dominate the classification
  • Solution: Laplace Smoothing (Add-one smoothing)

    • Instead of:
    • Use:
    • Where is the number of possible values for attribute
  • Example: Suppose our 10 patients were also checked for a rash, and none of the 3 flu patients had one:

    • Without smoothing:
    • With Laplace:

Probability prologue, moved here from prob.md so the first probability meeting stays short enough to serve the bias/variance activity. About 22 minutes. Supplies the derivation the "Recall the naive Bayes classifier" slide below depends on.

P(x) = .2 P(~x) = .8 P(c | x) = .2 P(f | x) = .7 P(h | x) = .1 P(c | ~x) = .1 P(f | ~x) = .05 P(h | ~x) = .85