The following file contains a very simple implementation of a Naive Bayes text classifier with Laplace smoothing:
Here is some sample data derived from the Rotten Tomatoes Reviews Dataset:
These csv files contain two columns. The first is a class label: ‘1’ for “fresh” and ‘0’ for “rotten”. The second column contains short movie reviews. The goal here is to learn to classify movie reviews as either fresh or rotten.
Your goal in this assignment is to answer several questions by working with the provided data set and the provided classifier code. You are free to modify the provided code, but you should not make any changes to the algorithms that are currently in place. Make sure you comment any code changes to indicate how they are related to answering the questions.
Find the most “rotten” movie review among all of the reviews in the test set, as determined by the Naive Bayes classifier that has been trained on the training set. You’ll need to make some changes to the code since it currently only provides a class label without returning a score.
Provide the predicted class probability distribution associated with the following review:
Jurassic Park is a cautionary tale about science gone wrong and
filmmaking gone lazy. For all its groundbreaking effects, the plot is
held together with dino-sized leaps in logic and characters who make
decisions so dumb they deserve to be eaten. The kids are annoying, the
adults are incompetent, and Jeff Goldblum spends half the movie
shirtless and smirking like he’s in a cologne ad. It’s less a
thrilling adventure and more a theme park ride with a script written
on the back of a napkin.
Again, you will need to modify the existing classifier to convert the log-based scores to a normalized probability distribution.
Generate an ROC curve using the provided test data, where “fresh” is the positive class. This will require determining a score value for every test instance. You should make use of the sklearn ROC curve function in generating your figure.
Find the 10 words that, after training, are most indicative of rottenness, and the 10 words that are most indicative of freshness.
Answering this requires understanding how class scores are calculated by the classifier:
# Calculate scores for each class
for class_label in self.classes:
# Start with log of class prior
log_prob = math.log(self.class_priors[class_label])
# Add log probabilities of known words
for word in known_words:
word_prob = self.word_probs[class_label][word]
log_prob += math.log(word_prob) # <------
class_scores[class_label] = log_probThe indicated line can be seen as a weighted vote associated with one of the words in the review. If the word is more associated with one class than the other, it will have a relatively higher value. The most telling words are the words whose log class conditional probabilities differ the most between the two classes.
The classifier estimates word probabilities using Laplace smoothing. The formula from the slides,
\[P(X_i = x \mid Y = y) = \frac{\text{count}(X_i = x, Y = y) + 1}{\text{count}(Y = y) + |V_i|},\]
was written for attributes like Golf or Fedora, where each attribute has its own small set of possible values. Our text classifier instead has a single attribute, “which word,” whose possible values are the words in the vocabulary \(V\). Each word in a review is treated as an independent observation of that attribute, drawn from the same class-specific distribution no matter where in the review it appears. Adapted to that setting, the formula becomes
\[P(w \mid Y = y) = \frac{\text{count}(w, y) + \alpha}{N_y + \alpha |V|}\]
where \(\text{count}(w, y)\) is the number of times word \(w\) appears in
training reviews of class \(y\), \(N_y\) is the total number of words in
training reviews of class \(y\), and \(\alpha\) is the number of extra
times we pretend to have seen each word. Standard Laplace smoothing
uses \(\alpha = 1\), which is the classifier’s default. The raw counts
are stored in the classifier’s word_counts and total_words
attributes.
Find the word that appears most often in rotten training reviews but never appears in fresh ones. Report how many times it appears in rotten reviews, and the smoothed probability that the classifier assigns to it for each class.
Without smoothing, what would \(P(w \mid \text{fresh})\) be for your word from part (a), and what would that do to the fresh score of any review containing it? An earlier version of this classifier had no smoothing. To avoid taking the log of zero, it simply ignored every word that had a count of zero in either class. How many test reviews contain at least one word that it would have ignored? Explain why ignoring these words is a poor fix, using your word from part (a) as an example.
The alpha argument to the constructor controls the strength of
the smoothing. Plot test accuracy as a function of alpha for values
from 0.01 to 1,000,000, using a log scale for the x-axis. What happens
to the classifier’s predictions when alpha is very large, and why?
Write a positive movie review that will be classified as highly negative by our trained classifier. Provide your review as well as the probability distribution (using the same approach as Problem 2).
This assignment may be completed individually or in pairs. If you are working with a partner, you must notify me at the beginning of the project. My expectation for pairs is that both members are actively involved, and take full responsibility for all aspects of the project. In other words, I expect that you are either sitting (or virtually) together to work, and not that you are splitting up tasks to be completed separately. If both members of the group are not able to fully explain the code to me, then this does not meet this expectation.
Your answers should be provided in the following Jupyter Notebook file:
Submit your completed notebook, as well as your updated version of
naive_bayes.py through Gradescope. The submitted version of your
notebook should include the output of all code cells: I shouldn’t need
to run your notebook to see the answers.