{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 5GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# A couple of words to start with","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Hello, everyone! I'm Ivan Lipatov, 3rd year student of Higher School of Economics from Russia. I'm deeply interested in ML and DL topics and recently I've started my way up to the top on Kaggle. At time of posting this notebook I'm taking part in Jigsaw Multilingual Toxic Comment Classification and already have reached quite high results. That is why I decided to post a series of notebooks that show a sequential process of investigation of the competition problem. The aim of my work is not to teach Kaggle grandmasters how to win competitions, but to help those, who are only at the beginning of their Kaggle way, because as I've learned the most difficult part is to start :) \n\n(don't think that next stages of Ml research will be free lunch for you - there will be much work to do, but as far as you plunge into competition spirit it gets much easier to make your way through)","execution_count":null},{"metadata":{},"cell_type":"markdown","source":" # What data do we face","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"The goal of competition is to classify toxic comments in non-english languages having only english training data","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"#define global variables\nDATA_PATH = \"../input/jigsaw-multilingual-toxic-comment-classification\"\nsmall_ds_path = \"jigsaw-toxic-comment-train.csv\"\nlarge_ds_path = \"jigsaw-unintended-bias-train.csv\"\nval_ds_path = \"validation.csv\"\ntest_ds_path = \"test.csv\"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#download the data\nsmall_ds = pd.read_csv(os.path.join(DATA_PATH, small_ds_path), usecols=[\"id\", \"comment_text\", \"toxic\"])\nlarge_ds = pd.read_csv(os.path.join(DATA_PATH, large_ds_path), usecols=[\"id\", \"comment_text\", \"toxic\"])\nval_ds = pd.read_csv(os.path.join(DATA_PATH, val_ds_path))\ntest_ds = pd.read_csv(os.path.join(DATA_PATH, test_ds_path))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We have 4 datasets provided:\n\n1) ~224k examples train dataset with english comment and binary target label (either 0 or 1)\n\n2) ~1,87M examples train datset with english commentsand  target label varying berween 0 and 1 - estimation of probability to be target\n\n3) 8k examples validation dataset with data in several non-english languages\n\n4) ~63k examples test data with data in several non-english languages\n\n\nLet's take a more detailed look to each piece of the data","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"# Small train dataset","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"small_ds.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"vals = small_ds.toxic.value_counts()\nsns.barplot(vals.index, vals.values)\nplt.title(\"Non-toxic vs toxic occurence in data\")\nplt.ylabel(\"Number exmaples\")\nplt.xlabel(\"Target value\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So here is the first important insight - we face unbalanced classes problem - in following notebooks I will show to tackle it","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"toxic_examples = small_ds[small_ds[\"toxic\"] == 1].sample(5, random_state=42)[\"comment_text\"]\nfor comment in toxic_examples.values:\n    print(\"Next comment:\")\n    print(comment)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#generate wordcloud to get more intuition about toxicity\nfrom wordcloud import WordCloud\ntoxic_comments = \" \".join(small_ds[small_ds[\"toxic\"]==1][\"comment_text\"].values)\nwc = WordCloud().generate(toxic_comments)\nplt.imshow(wc, interpolation='bilinear')\nplt.axis(\"off\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Although it is not the most enjoyable actitivity, you can get more intuition about what means toxic by reading more comments. My own summary is that toxic means that a comment is written with offensive sense towards someone or something, often with profanity.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"non_toxic_examples = small_ds[small_ds[\"toxic\"] == 0].sample(5, random_state=42)[\"comment_text\"]\nfor comment in non_toxic_examples.values:\n    print(\"Next comment\")\n    print(comment)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"While non toxic comment is the ordinal speech about some topic - it seems quite understandable","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, (ax1, ax2) = plt.subplots(nrows=1, ncols=2)\nsmall_ds[\"num_words\"] = small_ds[\"comment_text\"].str.split().apply(len)\ntemp_ds = small_ds[small_ds[\"num_words\"] < 500]\nsns.violinplot(x=\"toxic\",y=\"num_words\", data=temp_ds, ax=ax1)\nax1.set_title(\"Distributions of number of words/sentences in toxic/nontoxic comments\")\n\nsmall_ds[\"num_sents\"] = small_ds[\"comment_text\"].str.split(\".\").apply(len)\ntemp2_ds = small_ds[small_ds[\"num_sents\"] < 100]\nsns.violinplot(x=\"toxic\",y=\"num_sents\", data=temp2_ds, ax=ax2)\n#ax2.set_title(\"Distribution of number of sentences in toxic/nontoxic comments\")\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(\"Number of words descriptive stats\")\nprint(small_ds[\"num_words\"].describe())\nprint()\nprint(\"Number of sentences descriptive stats\")\nprint(small_ds[\"num_sents\"].describe())","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Also for training we have larger dataset with comments, though toxicity there is distributed between 0 and 1.","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"large_ds.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"t = large_ds.toxic.round(1)\nt.value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sns.barplot(t.value_counts().index, t.value_counts().values)\n\nplt.ylabel(\"Num samples\")\nplt.xlabel(\"Probability of being toxic\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So we see that there are lots of totaly non-toxic comments and a few probably toxic - approximately the same we've seen in the small datasets - so the problem of unbalanced classes occurs again","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"From this data we need to decide which parts we can use as toxic, and which as non-toxic for training, as we need labels to be either 0 or 1","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"large_ds[\"rounded_toxic\"] = large_ds.toxic.round(1)\nmaybe_toxic = large_ds[(large_ds.rounded_toxic == 0.5) | (large_ds.rounded_toxic == 0.6)].comment_text\nprobably_toxic = large_ds[(large_ds.rounded_toxic == 0.7) | (large_ds.rounded_toxic == 0.8)].comment_text\nsurely_toxic = large_ds[(large_ds.rounded_toxic == 0.9) | (large_ds.rounded_toxic == 1.0)].comment_text","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# may be toxic examples\n\nfor comm in maybe_toxic.sample(3):\n    print(comm)\n    print()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"maybe_toxic_comments = \" \".join(maybe_toxic.values)\nwc = WordCloud().generate(maybe_toxic_comments)\nplt.imshow(wc, interpolation='bilinear')\nplt.axis(\"off\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#probably toxic examples\nfor comm in probably_toxic.sample(3):\n    print(comm)\n    print()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"probably_toxic_comments = \" \".join(probably_toxic.values)\nwc = WordCloud().generate(probably_toxic_comments)\nplt.imshow(wc, interpolation='bilinear')\nplt.axis(\"off\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"#surely toxic examples\nfor comm in surely_toxic.sample(3):\n    print(comm)\n    print()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"surely_toxic_comments = \" \".join(surely_toxic.values)\nwc = WordCloud().generate(surely_toxic_comments)\nplt.imshow(wc, interpolation='bilinear')\nplt.axis(\"off\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"We see that toxic == 0.5 is quite different from toxic == 1. Although, from examples it can be seen that toxic >= 0.5 are more likely toxic than non-toxic, from my point of view. Of course, it is still questionable which boarder is the best, but I decide to take 0.5. However, to give to the future model the information that originally not all the labels are straight zeros and ones we can apply technique called labels smoothing. The goal of this technique is to force the model to be less confident in its predictions - so to make resulting predicted probabilites closer to 0.5 than to 0 and 1. This can be reached by assigning downsizing labels == 1 and upsizing labels == 0 in training process due to some distribution for example uniform. \n\nHere is the nice guide to label smoothing https://towardsdatascience.com/label-smoothing-making-model-robust-to-incorrect-labels-2fae037ffbd0","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"Test and Validation datasets","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"val_ds.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"len(val_ds)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"vs = val_ds.lang.value_counts()\nsns.barplot(vs.index, vs.values)\nplt.xlabel(\"language\")\nplt.ylabel(\"Number of samples\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"ts = val_ds.toxic.value_counts()\nsns.barplot(ts.index, ts.values)\nplt.xlabel(\"Non-toxic vs Toxic\")\nplt.ylabel(\"Number of samples\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In validation dataset we have data in 3 languages: turkish, spanish and italian - approximately the same number of samples for each\n\nClass proportions are simillar to the training data we have, however train data differs from validation significantly because of the language. Nevertheless, it's okay as far as we want metrics on validation reflect the quality of the model on the test data. And as you can see from my following notebooks - results on validation and test are very close, which is very important for monitoring model quality while training","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"test_ds.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"len(test_ds)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"vc = test_ds.lang.value_counts()\nsns.barplot(vc.index, vc.values)\nplt.xlabel(\"language\")\nplt.ylabel(\"Number of samples\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Overall, there are three main insights from the data, which can help to create the better model in the future.\n\n1. Data is multilingual that is why it's rational to prefer those models that can handle multilingual input data and can take out the semantics not depending on language as we train on english , test on non-english\n2. Classes are unbalanced - which can cause a significant problem for model to converge. That is why, it is necessary to apply some methods to handle it\n3. Validation dataset can be a good reflection of the test dataset, that is why we can rely on metrics computed on it","execution_count":null},{"metadata":{},"cell_type":"markdown","source":"These are my following notebooks:\n\n1. https://www.kaggle.com/vgodie/first-baseline - there I'm considering different methods to approach the problem and show to set up zero baseline with BERT architecture\n2. https://www.kaggle.com/vgodie/class-balancing - there I share my ideas on how to fight unbalanced classes problem\n3. https://www.kaggle.com/vgodie/data-encoding - there I show how to make custom preprocessing data for more sophisticated models such as XLM-Roberta, however this notebook can be used as a template for the preprocessing for any transformer model\n4. https://www.kaggle.com/vgodie/xlm-roberta - and finally I build and train my best model with all data preparations and techniques discussed in the previous notebook","execution_count":null},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}