{
 "cells": [
  {
   "cell_type": "markdown",
   "id": "c41ca1ff-7f90-4024-819b-3c0abfdc17b5",
   "metadata": {},
   "source": [
    "# PA2 - Naive Bayes Questions\n",
    "\n",
    "Modify this notebook so that the answers to the questions below are generated when it is executed.  You should also include markdown cells explaining or illustrating your answers where appropriate.\n",
    "\n",
    "Name: YOUR NAME HERE"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": 2,
   "id": "f0774a5a-1500-44fd-a6b6-3709fc5e63df",
   "metadata": {},
   "outputs": [],
   "source": [
    "# LOAD THE DATA AND TRAIN A CLASSIFIER\n",
    "from naive_bayes import NaiveBayesClassifier, load_and_process_data\n",
    "X_train, y_train = load_and_process_data(\"data_rt_train.csv\")\n",
    "X_test, y_test = load_and_process_data(\"data_rt_test.csv\")\n",
    "nb_classifier = NaiveBayesClassifier()\n",
    "nb_classifier.fit(X_train, y_train)"
   ]
  },
  {
   "cell_type": "markdown",
   "id": "33c76408",
   "metadata": {},
   "source": [
    "## Problem 1 - The Worst Movie\n",
    "\n",
    "Find the most \"rotten\" movie review among all of the reviews in the\n",
    "test set, as determined by the Naive Bayes classifier that has been\n",
    "trained on the training set. You'll need to make some changes to the\n",
    "code since it currently only provides a class label without returning\n",
    "a score."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "1989021e",
   "metadata": {},
   "outputs": [],
   "source": []
  },
  {
   "cell_type": "markdown",
   "id": "84a4cfee",
   "metadata": {},
   "source": [
    "## Problem 2 - Probability of Rottenness \n",
    "\n",
    "Provide the predicted class probability distribution associated with\n",
    "the following review:\n",
    "\n",
    "\n",
    "```\n",
    "Jurassic Park is a cautionary tale about science gone wrong and\n",
    "filmmaking gone lazy. For all its groundbreaking effects, the plot is\n",
    "held together with dino-sized leaps in logic and characters who make\n",
    "decisions so dumb they deserve to be eaten. The kids are annoying, the\n",
    "adults are incompetent, and Jeff Goldblum spends half the movie\n",
    "shirtless and smirking like he’s in a cologne ad. It’s less a\n",
    "thrilling adventure and more a theme park ride with a script written\n",
    "on the back of a napkin.\n",
    "```\n",
    "\n",
    "Again, you will need to modify the existing classifier to convert the\n",
    "log-based scores to a normalized probability distribution.\n",
    "\n",
    "*Hint:* the log scores get more negative with every word, and\n",
    "`math.exp` of anything below about -745 underflows to 0.  To avoid\n",
    "this, subtract the largest class score from every class score before\n",
    "exponentiating.  This rescales all of the scores by the same factor,\n",
    "which cancels out when you normalize."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "da41e400",
   "metadata": {},
   "outputs": [],
   "source": []
  },
  {
   "cell_type": "markdown",
   "id": "e23b81c2",
   "metadata": {},
   "source": [
    "## Problem 3 - ROC Curve\n",
    "\n",
    "Generate an ROC curve using the provided test data, where \"fresh\"\n",
    "is the positive class. This will require determining a score value for\n",
    "every test instance. You should make use of the [sklearn ROC curve\n",
    "function](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.roc_curve.html)\n",
    "in generating your figure."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "da5afc0c",
   "metadata": {},
   "outputs": [],
   "source": []
  },
  {
   "cell_type": "markdown",
   "id": "0efe9b18",
   "metadata": {},
   "source": [
    "## Problem 4 - Predictive Words\n",
    "\n",
    "Find the 10 words that, after training, are most indicative of\n",
    "rottenness, and the 10 words that are most indicative of freshness.\n",
    "\n",
    "Answering this requires understanding how class scores are calculated\n",
    "by the classifier:\n",
    "\n",
    "```python\n",
    "        # Calculate scores for each class\n",
    "        for class_label in self.classes:\n",
    "            # Start with log of class prior\n",
    "            log_prob = math.log(self.class_priors[class_label])\n",
    "\n",
    "            # Add log probabilities of known words\n",
    "            for word in known_words:\n",
    "                word_prob = self.word_probs[class_label][word]\n",
    "                log_prob += math.log(word_prob)                  # <------\n",
    "\n",
    "            class_scores[class_label] = log_prob\n",
    "```\n",
    "\n",
    "The indicated line can be seen as a weighted vote associated with one\n",
    "of the words in the review.  If the word is more associated with one\n",
    "class than the other, it will have a relatively higher value.  The\n",
    "most telling words are the words whose *log* class conditional\n",
    "probabilities differ the most between the two classes."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "72979a00",
   "metadata": {},
   "outputs": [],
   "source": []
  },
  {
   "cell_type": "markdown",
   "id": "f2ed0f5a",
   "metadata": {},
   "source": [
    "## Problem 5 - Laplace Smoothing\n",
    "\n",
    "The classifier estimates word probabilities using Laplace smoothing.\n",
    "The formula from the slides,\n",
    "\n",
    "$$P(X_i = x \\mid Y = y) = \\frac{\\text{count}(X_i = x, Y = y) + 1}{\\text{count}(Y = y) + |V_i|},$$\n",
    "\n",
    "was written for attributes like *Golf* or *Fedora*, where each\n",
    "attribute has its own small set of possible values.  Our text\n",
    "classifier instead has a single attribute, \"which word,\" whose\n",
    "possible values are the words in the vocabulary $V$.  Each word in a\n",
    "review is treated as an independent observation of that attribute,\n",
    "drawn from the same class-specific distribution no matter where in the\n",
    "review it appears (so shuffling a review's words would not change its\n",
    "score).  Adapted to that setting, the formula becomes\n",
    "\n",
    "$$P(w \\mid Y = y) = \\frac{\\text{count}(w, y) + \\alpha}{N_y + \\alpha |V|}$$\n",
    "\n",
    "where $\\text{count}(w, y)$ is the number of times word $w$ appears in\n",
    "training reviews of class $y$, $N_y$ is the total number of words in\n",
    "training reviews of class $y$, and $\\alpha$ is the number of extra\n",
    "times we pretend to have seen each word.  Standard Laplace smoothing\n",
    "uses $\\alpha = 1$, which is the classifier's default.  The raw counts\n",
    "are stored in the classifier's `word_counts` and `total_words`\n",
    "attributes.\n",
    "\n",
    "a) Find the word that appears most often in rotten training reviews\n",
    "but never appears in fresh ones.  Report how many times it appears in\n",
    "rotten reviews, and the smoothed probability that the classifier\n",
    "assigns to it for each class.\n",
    "\n",
    "b) Without smoothing, what would $P(w \\mid \\text{fresh})$ be for your\n",
    "word from part (a), and what would that do to the fresh score of any\n",
    "review containing it?  An earlier version of this classifier had no\n",
    "smoothing.  To avoid taking the log of zero, it simply ignored every\n",
    "word that had a count of zero in either class.  How many test reviews\n",
    "contain at least one word that it would have ignored?  Explain why\n",
    "ignoring these words is a poor fix, using your word from part (a) as\n",
    "an example.\n",
    "\n",
    "c) The `alpha` argument to the constructor controls the strength of\n",
    "the smoothing.  Plot test accuracy as a function of `alpha` for values\n",
    "from 0.01 to 1,000,000, using a log scale for the x-axis.  What happens\n",
    "to the classifier's predictions when `alpha` is very large, and why?"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "433f83df",
   "metadata": {},
   "outputs": [],
   "source": []
  },
  {
   "cell_type": "markdown",
   "id": "7cda0e34",
   "metadata": {},
   "source": [
    "## Problem 6 - Writing Deceptive Reviews\n",
    "\n",
    "Write a **positive** movie review that will be classified as **highly negative** by\n",
    "our trained classifier.  Provide your review as well as the\n",
    "probability distribution (using the same approach as Problem 2)."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "id": "bfd4f3ae",
   "metadata": {},
   "outputs": [],
   "source": []
  }
 ],
 "metadata": {
  "kernelspec": {
   "display_name": "Python 3 (ipykernel)",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "codemirror_mode": {
    "name": "ipython",
    "version": 3
   },
   "file_extension": ".py",
   "mimetype": "text/x-python",
   "name": "python",
   "nbconvert_exporter": "python",
   "pygments_lexer": "ipython3",
   "version": "3.12.11"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 5
}
