{"cells":[{"metadata":{"_uuid":"7137f012df32b29a00905ff67526209e1f9bad0a"},"cell_type":"markdown","source":"## Load Librarys"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import os\nimport time\nimport numpy as np\nimport pandas as pd \nfrom collections import Counter\n\nfrom wordcloud import WordCloud\nimport matplotlib.pyplot as plt","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f6c026f971e4978391ecc9fba2eddde13b254ef2"},"cell_type":"markdown","source":"## Load Data"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"train_df = pd.read_csv(\"../input/train.csv\")\ntrain_df.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f2f6e54f75c745d7f64b8b023e697da6d59cbf21"},"cell_type":"code","source":"train_df.head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"14493cb23ad342df4d53be0c515cd35d7d53a0fe"},"cell_type":"code","source":"target_count = train_df.target.value_counts()\nprint('Class 0:', target_count[0])\nprint('Class 1:', target_count[1])\nprint('Proportion:', round(target_count[0] / target_count[1], 2), ': 1')\n\ntarget_count.plot(kind='bar', title='Count (target)');","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3502f21f18edd1357b58e441155820892f9e4f88"},"cell_type":"markdown","source":"Our sample is very unbalanced"},{"metadata":{"_uuid":"36aef49e03461fedc59f9432af5211cab379293e"},"cell_type":"markdown","source":"## Tokenize"},{"metadata":{"_uuid":"62779faf8f6a228730e419e9961150b4a90acf65"},"cell_type":"markdown","source":"We will drop stop words and tokenize text.\n\n**Stop words** usually refers to the most common words in a language. Text may contain stop words like ‘the’, ‘is’, ‘are’. Stop words can be filtered from the text to be processed. There is no universal list of stop words in nlp research, however the nltk module contains a list of stop words.\n\n**Tokenization** is the process of demarcating and possibly classifying sections of a string of input characters. The resulting tokens are then passed on to some other form of processing."},{"metadata":{"trusted":true,"_uuid":"504ec3c1ff4b253256814c148c1c242ced6f9ae4"},"cell_type":"code","source":"import nltk\nfrom nltk.corpus import stopwords\n\ndef tokenizer(file_text):\n    tokens = nltk.word_tokenize(file_text)\n\n    stop_words = stopwords.words('english')\n    tokens = [i for i in tokens if ( i not in stop_words )]\n    \n    return ' '.join(tokens)\n\ntrain_df.question_text = train_df.question_text.apply(lambda x: tokenizer(x))\n\ntrain_df.head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"45580981ead033d35d3fd372f3dcdffcd7d68ec9"},"cell_type":"markdown","source":"## Word Clouds"},{"metadata":{"_uuid":"2df01f210a4e8b702dd857d284b59b680c03f012"},"cell_type":"markdown","source":"Build our first  Word Clouds using all data."},{"metadata":{"trusted":true,"_uuid":"5282e3e1226e0441a482c6ad737599f4d07a8c0b"},"cell_type":"code","source":"text = ' '.join(train_df['question_text'].str.lower().values[-1000000:])\nwordcloud = WordCloud(max_font_size=None, background_color='black',\n                      width=1200, height=1000).generate(text)\nplt.figure(figsize=(12, 8))\nplt.imshow(wordcloud)\nplt.title('Top words in question text')\nplt.axis(\"off\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e486a7478136e31b33c6e25fc6153621035ec692"},"cell_type":"markdown","source":"**Сonclusion**\n\nThis cloud does not provide any useful information or insights.\n\nWe need to see the clouds for nontoxic toxic content in search of insight"},{"metadata":{"_uuid":"4c79c1c821e4d1cee940cb9efa986ed99f283fa8"},"cell_type":"markdown","source":"### Nontoxic content"},{"metadata":{"trusted":true,"_uuid":"fb89fe96ae1a782a2a42773378c1bbd0392cb6c9"},"cell_type":"code","source":"train_df[train_df['target']==0].question_text.head(10)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"63fdc1d968f21d4fef2e62ff60b917ac58e8076e"},"cell_type":"code","source":"text = ' '.join(train_df[train_df['target']==0].question_text.str.lower().values[-1000000:])\nwordcloud = WordCloud(max_font_size=None, background_color='black',\n                      width=1200, height=1000).generate(text)\nplt.figure(figsize=(12, 8))\nplt.imshow(wordcloud)\nplt.title('Top words in nontoxic question text')\nplt.axis(\"off\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5d1c6d988e1444ade0e707012440d6715f04c269"},"cell_type":"markdown","source":"**Сonclusion**\n\nThis cloud does not provide any useful information or insights same as previous.\n\nWe need to see the cloud of toxic content and compare them."},{"metadata":{"_uuid":"6abf4a6c7bf2c1a67a9e34194d48e947b234df49"},"cell_type":"markdown","source":"## Toxic content\n\nAttention, 18+ content\n\nThe text below may offend your feelings."},{"metadata":{"_uuid":"66fd0bf9c81b4d8d91e028b34d69923a195544ea"},"cell_type":"markdown","source":"![Toxic content](https://upload.wikimedia.org/wikipedia/commons/7/78/RARS_18%2B.svg)"},{"metadata":{"trusted":true,"_uuid":"e4946d0ec68e7e69ebcae7ead48dd273757695d6"},"cell_type":"code","source":"text = ' '.join(train_df[train_df['target']==1].question_text.str.lower().values[-1000000:])\nwordcloud = WordCloud(max_font_size=None, background_color='black',\n                      width=1200, height=1000).generate(text)\nplt.figure(figsize=(12, 8))\nplt.imshow(wordcloud)\nplt.title('Top words in toxic question text')\nplt.axis(\"off\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f18dfca3b537cfe6d968de02ff55f93c948764c9"},"cell_type":"markdown","source":"![trump](https://timedotcom.files.wordpress.com/2018/03/donald-trump-snl-baldwin-twitter.jpg)"},{"metadata":{"trusted":true,"_uuid":"8e1e012e7faa64cf695588752046e605491ad2d1"},"cell_type":"markdown","source":"**Сonclusion**\n\nDonald Trump on top of the world\n\nThere are several words that intersect in toxic and nontoxic question text. For example: India, People.\nWe need to understand what to do about it.\n\nThen you can pay attention to the typical rasism, sexist and political question text."},{"metadata":{"_uuid":"55cb10a358f4d4e43102fd6ba7d1414f59f180e3"},"cell_type":"markdown","source":"I was very surprised to find Donald Trump on top of toxic questions.\nOut of interest I decided to take a look, what kind of questions are these."},{"metadata":{"trusted":true,"_uuid":"69d441cdc544faefb6920a900f604631867be4e1"},"cell_type":"code","source":"pd.set_option('display.max_colwidth', -1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b9d91dc66e6f243b3a60e702c91dbdcc777da32c"},"cell_type":"code","source":"train_df[(train_df['target']==1) & (train_df['question_text'].str.contains(\"Trump\"))].head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"234bb011e317cbf7595a467d16e5eeb7d1f1ca4c"},"cell_type":"markdown","source":"And look at nontoxic questions about Donald Trump"},{"metadata":{"trusted":true,"_uuid":"d0fb6a53ebed9c0a87987000f7370e4bc420e9ae"},"cell_type":"code","source":"train_df[(train_df['target']==0) & (train_df['question_text'].str.contains(\"Trump\"))].head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f65714184b6904dfb1e141789f35ad384645d980"},"cell_type":"markdown","source":"The author is not English speaking. But I think that many questions with a target = 1 should be from 0 and questions from a target = 0 should be from 1.\n\nYes, I know that in the description of the competition was:\n>  The training data includes the question that was asked, and whether it was identified as insincere (target = 1). The ground-truth labels contain some amount of noise: they are not guaranteed to be perfect.\n\nI think this can seriously affect the generalizing ability of the model."},{"metadata":{"_uuid":"533542d03f9d5504fc539c6bf53d42fdb8679b6c"},"cell_type":"markdown","source":"## TODO\n1. Try n-grams\n1. Try Word2Vec\n1. Try sec2seq"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}