{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"One of the more underappreciated aspects of Kaggle competitions is that we can \"repurpose\" them for all sorts of different tasks that go beyond the socope of the original context. Not all of the models that we build are equally generalizable, but some can be used for a wide variety of purposes. \n\nIn this kernel I'd like to see how do models built on [Toxic Comment Classification Challenge](https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/) perform on non-competition \"real world\" data. Here I will just use one model that was built inside of a [kernel](https://www.kaggle.com/tunguz/bi-gru-lstm-cnn-poolings-fasttext). The kernel scores in the 0.984x AUC range. It's a respectable score, but well below the top solutions that scored in the 0.988x range. I have used this approach to find out [how toxic are Hillary Clinton and Donald Trump tweets](https://www.kaggle.com/tunguz/how-toxic-are-hillary-and-trump-tweets/), with some interesting insights. \n\nFirst, let's load up the Python libraries."},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline  \n\nimport time\nstart_time = time.time()\nfrom sklearn.model_selection import train_test_split\nimport sys, os, re, csv, codecs, numpy as np, pandas as pd\nnp.random.seed(32)\nos.environ[\"OMP_NUM_THREADS\"] = \"4\"\nfrom keras.preprocessing.text import Tokenizer\nfrom keras.preprocessing.sequence import pad_sequences\nfrom keras.layers import Dense, Input, LSTM, Embedding, Dropout, Activation, Conv1D, GRU\nfrom keras.layers import Bidirectional, GlobalMaxPool1D, MaxPooling1D, Add, Flatten\nfrom keras.layers import GlobalAveragePooling1D, GlobalMaxPooling1D, concatenate, SpatialDropout1D\nfrom keras.models import Model, load_model\nfrom keras import initializers, regularizers, constraints, optimizers, layers, callbacks\nfrom keras import backend as K\nfrom keras.engine import InputSpec, Layer\nfrom keras.optimizers import Adam, RMSprop\nfrom keras.callbacks import EarlyStopping, ModelCheckpoint, LearningRateScheduler\nfrom keras.layers import GRU, BatchNormalization, Conv1D, MaxPooling1D\n\nimport logging\nfrom sklearn.metrics import roc_auc_score\nfrom keras.callbacks import Callback\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a2feb7cb56e4a6425fffbbd0f9627184962f8829"},"cell_type":"markdown","source":"Now we'll load the Quora datasets:"},{"metadata":{"trusted":true,"_uuid":"b0fbd1556270bfb4d99f0119ae4e8bcf88fad197"},"cell_type":"code","source":"train = pd.read_csv('../input/quora-insincere-questions-classification/train.csv').fillna(' ')\ntest = pd.read_csv('../input/quora-insincere-questions-classification/test.csv').fillna(' ')\ntest_qid = test['qid']\ntrain_qid = train['qid']\ntrain_target = train['target'].values\n\ntrain_text = train['question_text']\ntest_text = test['question_text']\n\nall_text = pd.concat([train_text, test_text])\n\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c4f1ed133e49e9ab28f054cfe779a0e79e4084a4"},"cell_type":"code","source":"embedding_path = \"../input/fasttext-crawl-300d-2m/crawl-300d-2M.vec\"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"799552635ee988fd76bfb82341b75f376aa0e69c"},"cell_type":"markdown","source":"We will embed words from these tweets into a word-vector space using one of the previously trained word embeddings. Here we use a 300-dimensional vector space that comes curtesy of FastText. Unfortunately, this embedding is not available for the Quora compatition, but as we are using this kernel just for the educational purposes, that will be fine. We will also limit the length of text to 220 words. This is an overkill for questions, but for general purpose it is rather small text length. The original was aimed at much longer text sizes, and this was a reasonable length for those purposes. The best embedding that we used in Toxic limited length to 900 words."},{"metadata":{"trusted":true,"_uuid":"6d67d7121715d290633c2486f6043c392996ff81"},"cell_type":"code","source":"embed_size = 300\nmax_features = 130000\nmax_len = 220\n\nlist_classes = [\"toxic\", \"severe_toxic\", \"obscene\", \"threat\", \"insult\", \"identity_hate\"]\ntrain_text = train_text.str.lower()\ntest_text = test_text.str.lower()\nall_text = all_text.str.lower()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a556d8cadd33e99ef3d358b2e96fa8cd52d5f01f"},"cell_type":"markdown","source":"In order for our pretrained models to work, we need to transform the text here into the appropriate vectorized format."},{"metadata":{"trusted":true,"_uuid":"26a7e8f5ad8a0e3b059d9f07aaf1700e1dae63a9"},"cell_type":"code","source":"tk = Tokenizer(num_words = max_features, lower = True)\ntk.fit_on_texts(all_text)\nall_text = tk.texts_to_sequences(all_text)\ntrain_text = tk.texts_to_sequences(train_text)\ntest_text = tk.texts_to_sequences(test_text)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b5f1678a1a3dcd275f2a51752d70b243ef6f3dda"},"cell_type":"markdown","source":"We also need to pad the tweets that are less than 220 words, which is essentially all of them."},{"metadata":{"trusted":true,"_uuid":"9424aaad0a211a24333d47fad7bb23133807a0d2"},"cell_type":"code","source":"train_pad_sequences = pad_sequences(train_text, maxlen = max_len)\ntest_pad_sequences = pad_sequences(test_text, maxlen = max_len)\nall_pad_sequences = pad_sequences(all_text, maxlen = max_len)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"afe93bb3386482e349d7b756f24562b6c4c00bb6"},"cell_type":"code","source":"def get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\nembedding_index = dict(get_coefs(*o.strip().split(\" \")) for o in open(embedding_path))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"070c7f41c93bfb0fdd6dfe30487e4cd548c15922"},"cell_type":"code","source":"word_index = tk.word_index\nnb_words = min(max_features, len(word_index))\nembedding_matrix = np.zeros((nb_words, embed_size))\nfor word, i in word_index.items():\n    if i >= max_features: continue\n    embedding_vector = embedding_index.get(word)\n    if embedding_vector is not None: embedding_matrix[i] = embedding_vector","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"513844c2abf71ede8040abdbb1e307810482a3a7"},"cell_type":"code","source":"model = load_model(\"../input/bi-gru-lstm-cnn-poolings-fasttext/best_model.hdf5\")\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bf893f5de8ac3b8a82212be83bd9e920fc0365d9"},"cell_type":"code","source":"train_pred = model.predict(train_pad_sequences, batch_size = 1024, verbose = 1)\ntest_pred = model.predict(test_pad_sequences, batch_size = 1024, verbose = 1)\n#all_pred = model.predict(all_pad_sequences, batch_size = 1024, verbose = 1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3bd4bb16c18e76a2bdc12384df8d12f3efef00cb"},"cell_type":"markdown","source":"Let's see what's the maximum probability for this model:"},{"metadata":{"trusted":true,"_uuid":"76aaefe824942985861f7aec9300ec18240f2971"},"cell_type":"code","source":"train_pred.max()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"066c3c4f61dffcfd0e3cbc01600201598212829a"},"cell_type":"code","source":"test_pred.max()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"eb396473227fedf12309fc8063c3e4f1e43c78fc"},"cell_type":"markdown","source":"In other words, at nearly 1.0 probability the model seems pretty confident about the \"toxicity\" of some of the tweets.\n\nNow let's put the predictions into a dataframe, so we can have a better view of them and how they relate to the actual tweets."},{"metadata":{"trusted":true,"_uuid":"94b4b00672991b9365ddfba91ca7cc94fcb04602"},"cell_type":"code","source":"toxic_predictions_train = pd.DataFrame(columns=list_classes, data=train_pred)\ntoxic_predictions_test = pd.DataFrame(columns=list_classes, data=test_pred)\ntoxic_predictions_train['question_text'] = train['question_text'].values\ntoxic_predictions_test['question_text'] = test['question_text'].values\ntoxic_predictions_train['qid'] = train_qid\ntoxic_predictions_test['qid'] = test_qid","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6f2fb32d4eaacb585f7bb97d830a17e3021ec574"},"cell_type":"code","source":"toxic_predictions_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c5247f6ea40c035286dfaec6564bc0c679a12264"},"cell_type":"code","source":"toxic_predictions_test.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4521e2c4ace223aa832096f9406df9404fc4f694"},"cell_type":"code","source":"toxic_predictions_train[list_classes].describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f134a4f871af071444973c38a915e68cc9134e39"},"cell_type":"code","source":"toxic_predictions_test[list_classes].describe()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"73f59074655bedeba58b98319d350fd3e5d6f363"},"cell_type":"markdown","source":"The worst 'toxic' train questions:"},{"metadata":{"trusted":true,"_uuid":"f3c411d4e09dc756f70c69396e9e103257f2cf03"},"cell_type":"code","source":"print(toxic_predictions_train.sort_values(by=['toxic'], ascending=False)['question_text'].head(10).values)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ba1ce487834a3e0c66ccc01432363e190c9cfd74"},"cell_type":"code","source":"print(toxic_predictions_train.sort_values(by=['severe_toxic'], ascending=False)['question_text'].head(10).values)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c505c9d0a4e4b34462a75a39a1cdbc0021ff8e27"},"cell_type":"code","source":"print(toxic_predictions_train.sort_values(by=['obscene'], ascending=False)['question_text'].head(10).values)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fea8ab575c67380f3d8208a91836ba261a387b88"},"cell_type":"code","source":"print(toxic_predictions_train.sort_values(by=['threat'], ascending=False)['question_text'].head(10).values)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8efc5670a7c9d9ee31fa45479a5ac24a6da9b07f"},"cell_type":"code","source":"print(toxic_predictions_train.sort_values(by=['insult'], ascending=False)['question_text'].head(10).values)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c64e30dc8857b20a2b16284fc0267d845a1f2c31"},"cell_type":"code","source":"print(toxic_predictions_train.sort_values(by=['identity_hate'], ascending=False)['question_text'].head(10).values)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"454a3a042f741285428bf412043a301f7d903f53"},"cell_type":"markdown","source":"In other words, the model seems not to work too well. The fact that soem of these are marked almost 1.0 on probablity scale shows the limitations of this model. It is very likely that Quora employs a very good set of toxicity-detection tools on their site, so the kinds of questions we get in this competition will most likley be already heavily vetted for inapropriate content. "}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}