{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport string\nimport tensorflow as tf\nimport sklearn as sk # For AUC/ROC analysis\n\n\n# For filtering stopwords\nfrom nltk.corpus import stopwords\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-04-22T01:52:59.314995Z","iopub.execute_input":"2023-04-22T01:52:59.315485Z","iopub.status.idle":"2023-04-22T01:52:59.327006Z","shell.execute_reply.started":"2023-04-22T01:52:59.315444Z","shell.execute_reply":"2023-04-22T01:52:59.325791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Experiment Authors and Thanks\nThis project has been worked by Jann C. García Pagán & Sebastián A. Estrada López\nWith the help of Tech Exchange Instructor, Daniel Gillick and the TA's from the Google Tech Exchange Machine Learning course\n\nProject original description, data, information, code and rules can be found using \nthe following link: https://www.kaggle.com/competitions/jigsaw-multilingual-toxic-comment-classification/overview","metadata":{}},{"cell_type":"markdown","source":"### Split Data\n\nThis Kaggle competition provides us already with the data split for training, validation and test as 'csv' files. We can manipulate and use these data sets by converting them into pandas dataframes. The competition provided an approximate of 3% of the data as validation.","metadata":{}},{"cell_type":"code","source":"# File names\nTRAINING_DATA_PATH = \"/kaggle/input/jigsaw-multilingual-toxic-comment-classification/jigsaw-unintended-bias-train.csv\"\nVALIDATION_DATA_PATH = \"/kaggle/input/jigsaw-multilingual-toxic-comment-classification/jigsaw-toxic-comment-train.csv\"\n\n# Load the data\ntraining_dataframe = pd.read_csv(TRAINING_DATA_PATH)\nvalidation_dataframe = pd.read_csv(VALIDATION_DATA_PATH)\n\n# Verify Data is downloaded\ndisplay(training_dataframe)\ndisplay(validation_dataframe)\n\nprint(\"Train Shape:\", training_dataframe.shape)\nprint(\"Validation Shape:\", training_dataframe.shape)\nprint(\"Validation %:\", 100* len(validation_dataframe)/len(training_dataframe))\n","metadata":{"execution":{"iopub.status.busy":"2023-04-22T01:52:59.332439Z","iopub.execute_input":"2023-04-22T01:52:59.332917Z","iopub.status.idle":"2023-04-22T01:53:20.174287Z","shell.execute_reply.started":"2023-04-22T01:52:59.332879Z","shell.execute_reply":"2023-04-22T01:53:20.172786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From our data, we only care about the 'toxic' (what we are predicting) and the 'comment_text' (our input).","metadata":{}},{"cell_type":"code","source":"# Constants\nstop_words = set(stopwords.words('english'))\nOOV_TOKEN = \"<OOV>\"\nMAX_LENGTH = 20 # How long each input should be\nVOCABULARY_LIMIT = 100000 # How large our vocabulary should be\nTHRESHOLD = 0.5 # Threshold used in Data Anaylsis. probability > threshold = toxic\nTRAINING_COUNT = 500000 # How many training inputs to use.\n\n# NOTE: We do not convert to a numpy array just yet, as we need to process the data\nx_train = training_dataframe[\"comment_text\"][:TRAINING_COUNT]\nx_validation = validation_dataframe[\"comment_text\"]\n\ny_train = training_dataframe[\"toxic\"][:TRAINING_COUNT]\ny_validation = validation_dataframe[\"toxic\"]\n","metadata":{"execution":{"iopub.status.busy":"2023-04-22T01:53:20.177517Z","iopub.execute_input":"2023-04-22T01:53:20.178162Z","iopub.status.idle":"2023-04-22T01:53:20.185288Z","shell.execute_reply.started":"2023-04-22T01:53:20.178123Z","shell.execute_reply":"2023-04-22T01:53:20.184182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Cleansing the Data\n\nFor the data to be used as an input to the model, we need to do the following:\n* Limit all inputs to a certain length\n* Replace all words NOT in the vocabulary with a special OOV token\n* Have all comments be in numerical values (models do not work w/ strings)\n\nCurrently, **because we do not have a vocabulary, we cannot perform these preparations.** However, to ease the creation of our vocabulary, we will cleanse our data by performing the following:\n* Remove all stopwords\n    - This is because stopwords provide almost nothing to the outcome of the label. They are usually neutral, and are removed before reaching the model.\n    - To do this, we used the stopwords set from the NLTK library\n    \n    \n* Remove punctuations\n    - Punctuations bring a new level of complexity when it comes to word meaning. Removing them simplifies what data we have to work with. However, this comes as a cost as words such as \"engl!sh\" are now treated as \"englsh\" instead of \"english\". \n    \n    \n* Lowercase all words\n    - This is to ensure words like \"Hello\" and \"hello\" are counted as one.","metadata":{}},{"cell_type":"code","source":"def remove_stopwords(sequences):\n    return sequences.apply(lambda x: ' '.join([word for word in x.split() if word not in stop_words]))\n\n    \ndef remove_punctuations(sequences):\n    return sequences.str.replace('[{}]'.format(string.punctuation), '')\n\n\ndef lowercase_all(sequences):\n    return sequences.str.lower()\n\ndef cleanse_data(sequences):\n    sequences = remove_punctuations(sequences)\n    sequences = lowercase_all(sequences)\n    return remove_stopwords(sequences).to_numpy()\n\n    \nx_train_cleansed = cleanse_data(x_train)\nx_validation_cleansed = cleanse_data(x_validation)\n\nprint(x_train[0])\nprint(x_train_cleansed[0])","metadata":{"execution":{"iopub.status.busy":"2023-04-22T01:53:20.187116Z","iopub.execute_input":"2023-04-22T01:53:20.187449Z","iopub.status.idle":"2023-04-22T01:53:26.239099Z","shell.execute_reply.started":"2023-04-22T01:53:20.187419Z","shell.execute_reply":"2023-04-22T01:53:26.238222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Baseline Model\n\nThe initial baseline model used consisted on always predicting the most frequent classification label. For example, if there are more toxic comments (1) than non toxic comments (0) then the baseline predicts a 1 for all data. The opposite is true as well. After changing the data set from predicting a binary label to predicting a probability, the baseline model was changed into predicting the mean of all of the probabilities found in the data. ","metadata":{}},{"cell_type":"code","source":"# Predict the average of all probabilities as the baseline model for the experiment\npredicted_average = y_train.mean()\n# Make the same prediction for every label in the training set\nprediction = [predicted_average for i in range(len(y_validation))] \nprint(\"Baseline Prediction\", prediction[:10])\n\n# Generate the dataframe, and generate the .csv\nvalidation_predictions = {\n    \"id\": [],\n    \"toxic\": []\n}\n\n\nfor i in range(len(validation_dataframe)):\n    validation_predictions[\"id\"].append(i)\n    validation_predictions[\"toxic\"].append(prediction)\n    \n    \nvalidation_results = pd.DataFrame(validation_predictions)\n\ndisplay(validation_results)","metadata":{"execution":{"iopub.status.busy":"2023-04-22T01:53:26.240487Z","iopub.execute_input":"2023-04-22T01:53:26.240988Z","iopub.status.idle":"2023-04-22T01:53:26.496789Z","shell.execute_reply.started":"2023-04-22T01:53:26.240954Z","shell.execute_reply":"2023-04-22T01:53:26.495580Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Validation Results\n\nIntitially, the baseline was evaluated with F1 scoring. In order to calculate F1, first it is required to calculate the Accuracy and Recall. Accuracy can be defined as: what is actually true from the predicted truth labels. Recall can be defined as: How much of the truth was missed? \n\nAnother definitions to be learned for the validation phase are:\n\n* True Positive - The model correctly predicts the positive class and the actual label is true\n* True Negative - The model correctly predicts the negative class and the actual label is negative\n* False Positive - The model predicts incorrectly the positive class and the actual label is negative\n* False Negative - The model predicts incorrectly the negative class and the actual label is positive\n\nHowever, after deeper analysis, it was decided to change the database from predicting binary labels into predicting probabilities for a class. Analyzing the data presented in the validation set, it was observed that these contain binary labels, 0 and 1, while the training set contains probabilities. Therefore, the need of creating a new way to evaluate the prediction. It was found that in order to predict labels with predictions two scores can be used jointly: Receiver Operating Characteristic (ROC) and Area Under the Curve (AUC).\n\n* ROC is created by plotting the True Positive Rate (TPR) against the False Positive Rate (FPR) at various threshold settings. TPR is the proportion of true positives among all actual positives, and FPR is the proportion of false positives among all actual negatives.\n\n* The area under the ROC curve (AUC) is a metric that represents the overall performance of the classifier, with higher AUC values indicating better performance.\n\nThe performance of ROC & AUC score can be evaluated by the following table:\n\n|ROC & AUC Score| Interpretation |\n|-|-|\n|>0.8|Very Good Performance|\n|0.7-0.8|Good Performance|\n|0.5-0.7|Ok Performance|\n|<0.5|As good as random choice|\n\n","metadata":{}},{"cell_type":"code","source":"# Prediction Evaluation using ROC & AUC Score\n# y_train parameter should contain binary labels (0 & 1)\n# Prediction parameter expects the predicted probabilities of the positive class\nauc_roc_score = sk.metrics.roc_auc_score(y_validation,prediction)\nprint(\"ROC & AUC Score:\", auc_roc_score)\n","metadata":{"execution":{"iopub.status.busy":"2023-04-22T01:53:26.501132Z","iopub.execute_input":"2023-04-22T01:53:26.502058Z","iopub.status.idle":"2023-04-22T01:53:26.562246Z","shell.execute_reply.started":"2023-04-22T01:53:26.502017Z","shell.execute_reply":"2023-04-22T01:53:26.561066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Interprating Results\nInitially, with the original baseline and data set, It was observed from the evaluation that the baseline model gave a quantity of 0 True Positives and 0 False Positives. However, there is true negative and false negative predictions. This can mean that the baseline model is always predicting the negative class. Possible causes of this is that the data does not have many 'toxic' comments, so the majority of the classes are non-toxic (0). This could explain why there is no true positives (the baseline is not predicting anything with a 1) and why are there False negatives, since the model is predicting 0 for every case when it should predict a value of 1 for some comments. \n\nWith the new baseline and data set, The ROC & AUC score obtained from evaluation the baseline created gives 0.5, which is as good as random choice. This makes sense for the baseline. Eventhough the baseline is not exactly guessing, it is simply calculating the average of the probabilities and using that value as the prediction for all other labels, so at some degree it is guessing the same value for every label.","metadata":{}},{"cell_type":"markdown","source":"# Data Analysis\n\nTo understand the data a bit more, we will try to analyze if there is a correlation between:\n\n* the length of the comment and being toxic\n* the frequency of words and being toxic","metadata":{}},{"cell_type":"markdown","source":"We will make an histogram in order to observe if there is a relation between the length of a comment and if it is classified as a toxic comment. The question to be asked is: Is there a difference in length between toxic and non toxic comments? The expectation is for toxic comments to be shorter in length, since the tend observed by experience in looking at social media, toxic comments are usually short comments with emphasis in specific word or wor","metadata":{}},{"cell_type":"code","source":"# Create a list to store the lengths of the toxic comments found in the data set\ntoxic_comments_lengths = [len(x_train[i]) for i in range(len(x_train)) if y_train[i] >= THRESHOLD]\n# Create a list to store the lengths of the non toxic comments found in the data set\nnon_toxic_comments_lengths = [len(x_train[i]) for i in range(len(x_train)) if y_train[i] < THRESHOLD]\n\n\n# Plot a histogram graph to observe the lengths of each comment and how many comments have a specific length\nplt.hist([toxic_comments_lengths, non_toxic_comments_lengths],  density=True, bins=20, range=(0, 1000), label=['toxic','non toxic'])\nplt.legend()\nplt.show()\n\n# Verify the largest toxic comment for toxic and non toxic\nprint(\"Largest Toxic Comment: \", max(toxic_comments_lengths))\nprint(\"Largest Non Toxic Comment:\", max(non_toxic_comments_lengths))","metadata":{"execution":{"iopub.status.busy":"2023-04-22T01:53:26.563578Z","iopub.execute_input":"2023-04-22T01:53:26.563914Z","iopub.status.idle":"2023-04-22T01:53:27.497717Z","shell.execute_reply.started":"2023-04-22T01:53:26.563881Z","shell.execute_reply":"2023-04-22T01:53:27.496285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Histogram Analysis \n\nAt first, the counts of the lengths between 'toxic' and 'non toxic' comments did not gave much insight or information to be analysed. After gathering feedback from instructors, it was suggested to estimate the probability density of the distribution.This was recommended to gather a visual representation of the frequencies or probabilities of data points. From the histogram, it is observed that a high ammount of toxic comments do not have a long length of text when compared to non toxic comments. This is similar to the previous prediction that it was made above.","metadata":{}},{"cell_type":"markdown","source":"Next off, we will see what words occur more frequently in both classes. During analysis, we learned about \"stopwords\" \n> Stopwords are the words in a stop list (or stoplist or negative dictionary) which are filtered out (i.e. stopped) before or after processing of natural language data (text) because they are insignificant.\n\nIt was obvious that words such as \"the\" and \"a\" will be very frequent in our dataset. So when calculating frequency, we filtered out such words, using a stopword dataset provided by [nltk](https://www.nltk.org).\n\nOne other thing to consider was the presence of typos and punctuation. We decided to ignore typos overall, as some typos convey meaning to the context of the word (\"helloooooooooooooooooo\"). For punctuation, we decided to filter them out entirely, as they made up a majority of the \"frequent\" words. ","metadata":{}},{"cell_type":"code","source":"# Word to [positive, negative] count list\nword_count = dict()\n\nfor i, comment in enumerate(x_train_cleansed):\n    for word in comment.split():\n        \n        # Default value = [0, 0]\n        if word not in word_count:\n            word_count[word] = [0,0]\n\n        # Increase count\n        word_count[word][int(y_train[i] >= THRESHOLD)] += 1\n        \n    \n        \n# Turn dictionaries into lists, for easy sorting\nword_count_list = list(word_count.items())\n\n# Sort by frequency\np_count_list = sorted(word_count_list, key = lambda x: x[1][0], reverse = True)\nn_count_list = sorted(word_count_list, key = lambda x: x[1][1], reverse = True)\n\nWORDS_TO_DISPLAY = 10\n\n# Print top 10 words for both\nprint(\"Positive Comment word frequency\")\nprint(\"Word,\\tCount\")\nfor i in range(WORDS_TO_DISPLAY):\n    print(p_count_list[i][0], \"\\t\" , p_count_list[i][1][0])\n    \nprint(\"==========\")\nprint(\"Negative Comment word frequency\")\nprint(\"Word,\\tCount\")\nfor i in range(WORDS_TO_DISPLAY):\n    print(n_count_list[i][0], \"\\t\" , n_count_list[i][1][1])","metadata":{"execution":{"iopub.status.busy":"2023-04-22T01:53:27.499187Z","iopub.execute_input":"2023-04-22T01:53:27.499675Z","iopub.status.idle":"2023-04-22T01:53:36.148014Z","shell.execute_reply.started":"2023-04-22T01:53:27.499631Z","shell.execute_reply":"2023-04-22T01:53:36.146814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Positive Comments\n|Word|Count|\n|---|---|\n|article|75126|\n|page|57393|\n|wikipedia|42627|\n|would|39130|\n|please|38120|\n|one|36665|\n|talk|35021|\n|like|33540|\n|dont|30547|\n|see|28233|\n\n\n## Negative Comments\n|Word|Count|\n|---|---|\n|fuck|14143|\n|nigger|5537|\n|fucking|4955|\n|shit|4825|\n|like|4597|\n|dont|4232|\n|suck|3944|\n|wikipedia|3821|\n|hate|3683|\n|ass|3666|\n\n\nLooking at the data, most of these words make sense. All offensive words are mostly used in toxic comments, while positive words, such as \"please\", are used in non-toxic comments.\n\n\"Wikipedia\" appears in both types of comments equally, which makes sense, because all comments from the data set are obtained from Wikipedia. This gives the impression that this word will not contribute much in determining if a comment is toxic or not toxic.","metadata":{}},{"cell_type":"markdown","source":"## Determining the Vocabulary\n\nWe need a vocabulary to use as input to the model. To get it, we sort the word_count dictionary by total frequency of words, and limit it to just the top K words (constant is defined above).","metadata":{}},{"cell_type":"code","source":"# x[1][0] = positive count\n# x[1][1] = negative count\n# Sum them together for total frequency\n# TODO: There is an issue. Because negative comments are barely present in the data set\n# Negative words might be cut out first in favor of positive words\nvocabulary = sorted(word_count_list, key = lambda x: x[1][0] + x[1][1], reverse = True)\n\n# Limit vocabulary to K words\nvocabulary = vocabulary[:VOCABULARY_LIMIT]\n\n# Remove counts, we only care about the words\nvocabulary = [OOV_TOKEN] + [x[0] for x in vocabulary]\n\n# Create the vocabulary and reverse dictionaries\nid_to_word = dict([(key, value) for (key, value) in enumerate(vocabulary)])\nword_to_id = dict([(value, key) for (key, value) in enumerate(vocabulary)])","metadata":{"execution":{"iopub.status.busy":"2023-04-22T01:53:36.149203Z","iopub.execute_input":"2023-04-22T01:53:36.149536Z","iopub.status.idle":"2023-04-22T01:53:36.205831Z","shell.execute_reply.started":"2023-04-22T01:53:36.149504Z","shell.execute_reply":"2023-04-22T01:53:36.204599Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Processing the Data\n\nNow that we have a vocabulary, we can perform the 3 steps we mentioned before:\n* Limit all inputs to a certain length\n* Replace all words NOT in the vocabulary with a special OOV token\n* Have all comments be in numerical values (models do not work w/ strings)\n","metadata":{}},{"cell_type":"code","source":"def decode(sequence):\n    return \" \".join([id_to_word.get(i, OOV_TOKEN) for i in sequence])\n    \ndef encode(comment):\n    return [word_to_id.get(word, 0) for word in comment.split()]\n    \n\ndef pad_data(sequences, max_length = MAX_LENGTH):\n    # Pads each numerical array to a certain length\n    return np.array(list(\n        tf.keras.preprocessing.sequence.pad_sequences(\n          sequences, dtype=object, maxlen=max_length, padding='post', value=0)))\n\ndef text_to_sequences(comments):\n    # Encodes each comment into a numerical array\n    return np.array([encode(comment) for comment in comments])\n\ndef process_data(x_input):\n    return pad_data(text_to_sequences(x_input)).astype(int)\n\nx_train_processed = process_data(x_train_cleansed)\nx_validation_processed = process_data(x_validation_cleansed)\n\nprint(x_train_processed[0])","metadata":{"execution":{"iopub.status.busy":"2023-04-22T01:53:36.207385Z","iopub.execute_input":"2023-04-22T01:53:36.207741Z","iopub.status.idle":"2023-04-22T01:53:40.671544Z","shell.execute_reply.started":"2023-04-22T01:53:36.207708Z","shell.execute_reply":"2023-04-22T01:53:40.670285Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model\n\nNow that our data is processed, it's time to make our model. ","metadata":{}},{"cell_type":"code","source":"def plot_history(history):\n    plt.ylabel('Loss')\n    plt.xlabel('Epoch')\n    plt.xticks(range(0, len(history['loss'] + 1)))\n    plt.plot(history['loss'], label=\"training\", marker='o')\n    plt.plot(history['val_loss'], label=\"validation\", marker='o')\n    plt.legend()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-22T01:53:40.680585Z","iopub.execute_input":"2023-04-22T01:53:40.681000Z","iopub.status.idle":"2023-04-22T01:53:40.688781Z","shell.execute_reply.started":"2023-04-22T01:53:40.680965Z","shell.execute_reply":"2023-04-22T01:53:40.687384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def build_model(embedding_dim = 16, dropout_rate = 0.5):\n    # Clear session and remove randomness.\n    tf.keras.backend.clear_session()\n    tf.keras.utils.set_random_seed(0)\n    \n    model = tf.keras.Sequential()\n    model.add(tf.keras.layers.Embedding(\n        input_dim=VOCABULARY_LIMIT+1, # add 1 because of the OOV token\n        output_dim=embedding_dim,\n        input_length=MAX_LENGTH)\n    )\n    \n    model.add(tf.keras.layers.Conv1D(\n        filters=32, # number of filters (i.e., output channels)\n        kernel_size=4, # width of the filters\n        activation='relu' # activation function to use\n    ))\n    \n    model.add(tf.keras.layers.GlobalMaxPooling1D()) \n    model.add(tf.keras.layers.Dense(units=1, activation='sigmoid'))\n\n\n    model.compile(loss='binary_crossentropy',\n                    optimizer='adam',\n                    metrics=[tf.keras.metrics.AUC()])    \n    return model","metadata":{"execution":{"iopub.status.busy":"2023-04-22T01:53:40.690493Z","iopub.execute_input":"2023-04-22T01:53:40.691132Z","iopub.status.idle":"2023-04-22T01:53:40.700739Z","shell.execute_reply.started":"2023-04-22T01:53:40.691068Z","shell.execute_reply":"2023-04-22T01:53:40.699542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def build_embeddings_model(embedding_dim=2):\n    \"\"\"Build a tf.keras model using embeddings\"\"\"\n    tf.keras.backend.clear_session()\n    # Eliminate Randomness factor for consistent results\n    tf.keras.utils.set_random_seed(0)\n    \n    model = tf.keras.Sequential()\n    model.add(tf.keras.layers.Embedding(\n        input_dim=VOCABULARY_LIMIT+1, # add 1 because of the OOV token\n        output_dim=embedding_dim,\n        input_length=MAX_LENGTH)\n    )\n    \n    model.add(tf.keras.layers.Flatten())\n    \n    model.add(tf.keras.layers.Dense(\n        units=64,\n        activation='relu'))\n    \n    model.add(tf.keras.layers.Dense(\n        units=1,\n        activation='sigmoid'))\n    \n    model.compile(loss='binary_crossentropy',\n                  optimizer='adam',\n                  metrics=[tf.keras.metrics.AUC()])\n    \n    return model","metadata":{"execution":{"iopub.status.busy":"2023-04-22T01:53:40.702495Z","iopub.execute_input":"2023-04-22T01:53:40.703430Z","iopub.status.idle":"2023-04-22T01:53:40.715741Z","shell.execute_reply.started":"2023-04-22T01:53:40.703393Z","shell.execute_reply":"2023-04-22T01:53:40.714185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dimensions = [2, 4, 8, 16, 32, 64]\n\nmodels = [build_model, build_embeddings_model]\n\nfor build_m in models:\n    for d in dimensions:\n        print(\"Currently Evaluating Dimension:\", d)\n        model = build_m(d)\n        model.summary()\n\n        # Train the model\n        history = model.fit(\n            x = x_train_processed,             \n            y = y_train,                     \n            epochs=5,                    # number of passes through the training data\n            batch_size=64,                    # mini-batch size\n            validation_split=0.1,# use a fraction of the examples for validation\n            verbose=1                         # display some progress output during training\n        )\n\n        history = pd.DataFrame(history.history)\n        plot_history(history)\n\n        # Predict on the validation data\n        results = model.predict(x = x_validation_processed)\n\n        # Calculate Kaggle Score:\n        print(\"AUC/ROC Score:\", sk.metrics.roc_auc_score(y_validation, results))\n        print(\"=================\")\n    ","metadata":{"execution":{"iopub.status.busy":"2023-04-22T01:53:40.717198Z","iopub.execute_input":"2023-04-22T01:53:40.717596Z","iopub.status.idle":"2023-04-22T02:00:02.556412Z","shell.execute_reply.started":"2023-04-22T01:53:40.717562Z","shell.execute_reply":"2023-04-22T02:00:02.554823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# FNN Results after 5 Epochs\n|Embedding Size|Training Loss|Validation Loss|Training AUC/ROC|Validation AUC/ROC|Test AUC/ROC Score|\n|---|---|---|---|---|---|\n|2|0.2364|0.2633|0.8585|0.7942|0.9349|\n|4|0.2292|0.2689|0.8780|0.7826|0.9338|\n|8|0.2167|0.2791|0.9095|0.7678|0.9272|\n|16|0.2029|0.2920|0.9423|0.7493|0.9233|\n|32|0.1928|0.3040|0.9637|0.7418|0.9190|\n|64|0.1886|0.3018|0.9723|0.7426|0.9188|\n\n# CNN Results after 5 Epochs\n|Embedding Size|Training Loss|Validation Loss|Training AUC/ROC Score|Validation AUC/ROC Score|Test AUC/ROC Score|\n|---|---|---|---|---|---|\n|2|0.2398|0.2631|0.8526|0.7931|0.9334\n|4|0.2335|0.2669|0.8681|0.7874|0.9308\n|8|0.2248|0.2740|0.8896|0.7792|0.9269\n|16|0.2131|0.2804|0.9172|0.7691|0.9236\n|32|0.2032|0.2874|0.9414|0.7627|0.9220\n|64|0.1976|0.2881|0.9554|0.7625|0.9221\n\n\n","metadata":{}},{"cell_type":"code","source":"# Get predicted probabilities for the test data\nvalidation_predictions = model.predict(x_validation_processed)\n   \n    \nprint(\"========\")\n\n# Compute false positive rate and true positive rate\nfpr, tpr, thresholds = sk.metrics.roc_curve(y_validation, validation_predictions)\n\n# Compute the area under the curve (AUC)\nauc_score = sk.metrics.roc_auc_score(y_validation, validation_predictions)\n\n# Plot the ROC curve\nplt.plot(fpr, tpr, label='AUC = {:.2f}'.format(auc_score))\nplt.plot([0, 1], [0, 1], linestyle='--', color='gray')\nplt.xlabel('False Positive Rate')\nplt.ylabel('True Positive Rate')\nplt.title('Receiver Operating Characteristic (ROC) Curve')\nplt.legend()\nplt.show()\n\n# Calculate precision and recall for different thresholds\nprecision, recall, thresholds = sk.metrics.precision_recall_curve(y_validation, validation_predictions)\n\n# Plot the precision-recall curve\nplt.plot(recall, precision)\nplt.xlabel('Recall')\nplt.ylabel('Precision')\nplt.title('Precision-Recall Curve')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-04-22T02:00:02.559969Z","iopub.execute_input":"2023-04-22T02:00:02.560385Z","iopub.status.idle":"2023-04-22T02:00:16.287543Z","shell.execute_reply.started":"2023-04-22T02:00:02.560345Z","shell.execute_reply":"2023-04-22T02:00:16.286153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Evaluation History\n|Model|TP|TN|FP|FN|Accuracy|F1 Score|\n|---|---|---|---|---|---|---|\n|Baseline|0|6770|0|1230|0.84625|0|\n","metadata":{}},{"cell_type":"markdown","source":"# Submitting to Kaggle","metadata":{}},{"cell_type":"code","source":"# Unfortunately, the Kaggle Competition expects us to predict toxicity of non-English comments, something we were not aware of\n# when starting the competition. We decided to take a different approach in evaluating our Model\n\n# Please read the \"Important Notice\" section of our Final Report :)","metadata":{"execution":{"iopub.status.busy":"2023-04-22T02:00:16.289218Z","iopub.execute_input":"2023-04-22T02:00:16.289724Z","iopub.status.idle":"2023-04-22T02:00:16.295830Z","shell.execute_reply.started":"2023-04-22T02:00:16.289674Z","shell.execute_reply":"2023-04-22T02:00:16.294605Z"},"trusted":true},"execution_count":null,"outputs":[]}]}