{"cells":[{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":false},"cell_type":"markdown","source":"# Simple EDA - Multilingual Toxic Comment\n\n## Description\n> This year, we're taking advantage of Kaggle's new TPU support and challenging you to build multilingual models with English-only training data.\n\n### Data description\n\n> **What should I expect the data format to be?**\n>\n> The primary data for the competition is, in each provided file, the comment_text column. This contains the text of a comment which has been classified as toxic or non-toxic (0...1 in the toxic column). The train set’s comments are entirely in english and come either from *Civil Comments* or *Wikipedia talk page* edits. The test data's `comment_text` columns are composed of multiple non-English languages.\n>\n> The `*-train.csv` files and `validation.csv` file also contain a toxic column that is the target to be trained on.\n>\n> The `jigsaw-toxic-comment-train.csv` and `jigsaw-unintended-bias-train.csv` contain training data (`comment_text` and `toxic`) from the two previous Jigsaw competitions, as well as additional columns that you may find useful.\n>\n> `*-seqlen128.csv` files contain training, validation, and test data that has been processed for input into BERT.\n\n### What am I predicting?\n> You are predicting the probability that a comment is `toxic`. A toxic comment would receive a `1.0`. A benign, non-toxic comment would receive a `0.0`. In the test set, all comments are classified as either a `1.0` or a `0.0`.\n\n### Columns\n- **id** - identifier within each file.\n- **comment_text** - the text of the comment to be classified.\n- **lang** - the language of the comment.\n- **toxic** - whether or not the comment is classified as toxic. (Does not exist in test.csv.)\n\n-----------------------------\n**I'll update this EDA notebook in the next days/weeks, stay tuned!**"},{"metadata":{"trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n\nimport plotly.express as px\nimport plotly.graph_objects as go\nimport plotly.figure_factory as ff\n\nfrom wordcloud import WordCloud, STOPWORDS","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"DIR_INPUT = '/kaggle/input/jigsaw-multilingual-toxic-comment-classification'","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Train dataset"},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df1 = pd.read_csv(DIR_INPUT + '/jigsaw-toxic-comment-train.csv')\ntrain_df1['src'] = 0\ntrain_df1.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df2 = pd.read_csv(DIR_INPUT + '/jigsaw-unintended-bias-train.csv')\ntrain_df2['src'] = 1\ntrain_df2.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Because in the test/validation data we only have `['id', 'comment_text', 'toxic']` columns, I drop anything else from train.\n\n*Note: The `toxic` ratio is not the same in the two source dataset!*"},{"metadata":{"trusted":true},"cell_type":"code","source":"keep_cols = ['id', 'comment_text', 'toxic', 'src']\ntrain_df = train_df1[keep_cols].append(train_df2[keep_cols])\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"del train_df1, train_df2","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['toxic'] = (train_df['toxic'] > 0.5).astype(np.uint)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"print(\"We have {} english comments in the train datasets.\".format(train_df.shape[0]))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['toxic'].value_counts(normalize=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df.groupby(by=['toxic', 'src']).count()[['id']]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = go.Figure([go.Bar(x=['Not-toxic', 'Toxic'], y=train_df.toxic.value_counts())])\nfig.update_layout(\n    title='Toxic/non-toxic comments distribution in the train dataset'\n)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"train_df['comment_text_len'] = train_df['comment_text'].apply(lambda x : len(x))\ntrain_df['comment_text_word_cnt'] = train_df['comment_text'].apply(lambda x : len(x.split(' ')))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.histogram(train_df, x='comment_text_len', color='toxic', nbins=200)\nfig.show(renderer=\"kaggle\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.histogram(train_df[train_df['src'] == 0],\n                   x='comment_text_len',\n                   color='toxic',\n                   nbins=200,\n                   title='Text length - Source: Jigsaw toxic comment (train)')\nfig.show(renderer=\"kaggle\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.histogram(train_df[train_df['src'] == 1],\n                   x='comment_text_len',\n                   color='toxic',\n                   nbins=200,\n                   title='Text length - Source: Jigsaw unintended bias (train)')\nfig.show(renderer=\"kaggle\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.histogram(train_df[train_df['src'] == 0],\n                   x='comment_text_word_cnt',\n                   color='toxic',\n                   nbins=200,\n                   title='Word count - Source: Jigsaw toxic comment (train)')\nfig.show(renderer=\"kaggle\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = px.histogram(train_df[train_df['src'] == 1],\n                   x='comment_text_word_cnt',\n                   color='toxic',\n                   nbins=200,\n                   title='Word count - Source: Jigsaw toxic comment (train)')\nfig.show(renderer=\"kaggle\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Test/Valid dataset"},{"metadata":{"trusted":true},"cell_type":"code","source":"valid_df = pd.read_csv(DIR_INPUT + '/validation.csv')\nvalid_df.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"In the validation set we have 15.3% toxic comments (in the train set the toxic comments ratio is only 6.3%)\n\n*Note: the train set is a combination of the previous two competitions' data*"},{"metadata":{"trusted":true},"cell_type":"code","source":"per_lang = valid_df['lang'].value_counts()\nfig = go.Figure([go.Bar(x=per_lang.index, y=per_lang.values)])\nfig.update_layout(\n    title='Language distribution in the validation dataset'\n)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"valid_df.toxic.value_counts(normalize=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"fig = go.Figure([go.Bar(x=['Not-toxic', 'Toxic'], y=valid_df.toxic.value_counts())])\nfig.update_layout(\n    title='Language distribution in the validation dataset'\n)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"per_lang = valid_df.groupby(by=['lang', 'toxic']).count()[['id']]\nper_lang","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data = []\n\nfor lang in valid_df['lang'].unique():\n    y = per_lang[per_lang.index.get_level_values('lang') == lang].values.flatten()\n    data.append(go.Bar(name=lang, x=['Non-toxic', 'Toxic'], y=y))\n\nfig = go.Figure(data=data)\nfig.update_layout(\n    title='Language distribution in the validation dataset',\n    barmode='group'\n)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df = pd.read_csv(DIR_INPUT + '/test.csv')\ntest_df.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"test_df['lang'].value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"per_lang = test_df['lang'].value_counts()\nfig = go.Figure([go.Bar(x=per_lang.index, y=per_lang.values)])\nfig.update_layout(\n    title='Language distribution in the test dataset',\n)\nfig.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"# Comments"},{"metadata":{"trusted":true},"cell_type":"code","source":"toxic_samples = train_df[train_df['toxic'] == 1].sample(n=5)['comment_text']\n\nfor toxic in toxic_samples.values:\n    print(\"\")\n    print(\"==============================\")\n    print(toxic)\n    print(\"==============================\")\n    print(\"\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Wordclouds - Frequent words:\n"},{"metadata":{"trusted":true},"cell_type":"code","source":"rnd_comments = train_df.sample(n=2500)['comment_text'].values\nwc = WordCloud(background_color=\"black\", max_words=2000, stopwords=STOPWORDS.update(['Trump', 'people', 'one', 'will']))\nwc.generate(\" \".join(rnd_comments))\n\nplt.figure(figsize=(20,10))\nplt.axis(\"off\")\nplt.title(\"Random words\", fontsize=20)\nplt.imshow(wc.recolor(colormap= 'viridis' , random_state=17), alpha=0.98)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"rnd_comments = train_df[train_df['toxic'] == 0].sample(n=10000)['comment_text'].values\nwc = WordCloud(background_color=\"black\", max_words=2000, stopwords=STOPWORDS.update(['Trump', 'people', 'one', 'will']))\nwc.generate(\" \".join(rnd_comments))\n\nplt.figure(figsize=(20,10))\nplt.axis(\"off\")\nplt.title(\"Frequent words in non-toxic comments\", fontsize=20)\nplt.imshow(wc.recolor(colormap= 'viridis' , random_state=17), alpha=0.98)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"rnd_comments = train_df[train_df['toxic'] == 1].sample(n=10000)['comment_text'].values\nwc = WordCloud(background_color=\"black\", max_words=2000, stopwords=STOPWORDS.update(['Trump', 'people', 'one', 'will']))\nwc.generate(\" \".join(rnd_comments))\n\nplt.figure(figsize=(20,10))\nplt.axis(\"off\")\nplt.title(\"Frequent words in toxic comments\", fontsize=20)\nplt.imshow(wc.recolor(colormap= 'viridis' , random_state=17), alpha=0.98)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}