{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\nimport re\nfrom matplotlib import pyplot as plt\nimport seaborn as sns\n\nfrom nltk.util import ngrams\nfrom nltk.corpus import stopwords \nfrom nltk import FreqDist\nfrom collections import Counter\nfrom wordcloud import WordCloud\n\nplt.rcParams.update({'font.size': 22})\nstop_words = set(stopwords.words('english'))","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","collapsed":true,"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true,"_kg_hide-output":false,"_kg_hide-input":false},"cell_type":"code","source":"print(os.getcwd())\nfor path, dirs, files in os.walk(\"../\"):\n  print (path)\n  for f in files:\n    print (\"\\t{}\".format(f))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1eadcda24bfb410ec008957e2fdc339a7fc22a52"},"cell_type":"code","source":"df = pd.read_csv(\"../input/train.csv\")\ndf.info()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c08b31a25a885aa9b0b1e87d02be019c263d7577"},"cell_type":"markdown","source":"We have a healthy amount of data. There is over 1.3M sentences in the training set. For every question we have the identifier (useless), the text and the target. A target of 1 means the question is flagged as insincere, zero then means it is flagged as sincere.\n\nLet's see how do sincere and insincere questions look like"},{"metadata":{"trusted":true,"_uuid":"9aaa34b276b941088b1d74b1ee3e1c27237598c2"},"cell_type":"code","source":"print(df[(df[\"target\"] == 0)].iloc[:5,1].values)\nprint(\"\\n\")\nprint(df[(df[\"target\"] == 1)].iloc[:5,1].values)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e672d115f72368d7ab4131775f9885c31da61f65"},"cell_type":"markdown","source":"Insincere questions differ from sincere ones in that they include non accepted content such as incest. They are insincere as well if they take for granted a false premise, such as the US being a dictatorship. A full list of criteria for a question to be considered insincere can be found in the data section of the competition. It is expected to have some noise in the target values, so the data is not perfect.\n\nLet's start processing the data, in the next script we will remove punctuation marks and tokenize the data to get some insights into the structure of such sentences.\n\n"},{"metadata":{"trusted":true,"_uuid":"ac36a4d38d43a3ec986445dbffcd1e0e6bc8af5b"},"cell_type":"code","source":"strings = df['question_text'].values\nwords = []\nallCharsCount = 0\nremovedCount = 0\nremovedInfo = []\nfor s in strings:\n    s = s.lower()\n    \n    allCharsCount += len(s)\n    removed = re.sub(r'(?![a-zA-Z0-9\\s]).', \"\", s)\n    dif = len(s) - len(removed)\n    removedCount += dif\n    removedInfo.append([dif,dif/len(s)])\n    \n    tokenized = re.sub(r'(?![a-zA-Z0-9\\s]).', \" \", s).split(\" \")\n    words += [word for word in tokenized if len(word) != 0]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9c654dabf4c2af99d588bfb78f15a5702dc9a2b5"},"cell_type":"code","source":"print(\"{} chars removed ({:.3}%)\".format(removedCount, removedCount*100/allCharsCount))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f2dff356fcce4422351636ee5eb7c8a540299510"},"cell_type":"markdown","source":"2.52 percent of all the characters were considered punctuation marks and removed. It is a significant number of characters and we should dig deeper into the issue. Let's see which sentences were most affected by the character removal and how much of the sentence was removed in relation to the length of the sentence."},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"8df2b3fe5eece75dbc0f317ea323adbd6e8ce33c"},"cell_type":"code","source":"removedInfo = pd.DataFrame(removedInfo, columns=[\"sum\", \"percent\"])\nsortedRemovedInfo = removedInfo.sort_values(by=\"percent\")\nprint(\"\\tMost affected sentences:\\n\")\nfor x in np.arange(-1,-11,-1):\n    ind = sortedRemovedInfo.index[x]\n    print(\"with {:.2f}% of deletions:  {}\\n\".format(\n        sortedRemovedInfo.iloc[x,1]*100, strings[ind]))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b08016e243235ebf71868bf86dac2690dc1a8a95"},"cell_type":"markdown","source":"Looks like sentences with non latin characters took the worst part of it. If we stick to english to build the classifier then we should remove all the non english questions. For the moment let's keep studying how the character removal process went. At this point it would be useful to check the distribution of removed characters in the sentences in relation to the length of such sentences.\n\nIn the next script we plot the  distribution for percentage of characters removed and then we will do the same but ignoring the sentences most affected by the process since they are big outliers and they skew the distribution too much. We will plot the distribution for the 99.9% and 99.0% of least affected sentences"},{"metadata":{"trusted":true,"_uuid":"5ebff27b548129a7daef1c58c581cee42793cd17"},"cell_type":"code","source":"ind1Pct = int(sortedRemovedInfo.shape[0]*0.999)\nind2Pct = int(sortedRemovedInfo.shape[0]*0.99)\n\nfig, ax = plt.subplots(3,1,figsize=(15,15))\nfig.suptitle('Distribution for: Percent of characters removed in a sentence')\nax[0].set_title(\"Full distribution\")\nax[1].set_title(\"Lowest 99.9% distribution\")\nax[2].set_title(\"Lowest 99.0% distribution\")\nsns.distplot(sortedRemovedInfo.iloc[:,1], ax=ax[0])\nsns.distplot(sortedRemovedInfo.iloc[:ind1Pct,1], ax=ax[1])\nsns.distplot(sortedRemovedInfo.iloc[:ind2Pct,1], ax=ax[2])\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b60b193073ab7b0e5ed0c80bf7bb21a080f5cab6"},"cell_type":"markdown","source":"In the plot we can see that 99.9% of the data had less than 25% of their characters removed while 99% of the sentences had less than 9% of the characters removed. In general, most of the sentences had only around 2% of their characters removed.\n\nThere has to be a threshold of removed characters that, once overcame, the sentence loses too much of their meaning and becomes harmful for the model trying to learn from the data. We cannot know where that threshold is with the information we have right now. Perhaps a trial and error approach tuning the \"character removal filter\" hyperparameter can shed some light over the issue.\n\nIn the next section we will look at the length of the sentences. Perhaps we find that extreme lengths make for useless sentences"},{"metadata":{"trusted":true,"_uuid":"77635b24a1d05fc3dec31b09b3e5693bbf78b77d"},"cell_type":"code","source":"df[\"question_len\"] = df[\"question_text\"].apply(lambda x: len(x))\nlenSorted = df.sort_values(by=\"question_len\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b6e40cb0ab58c2f2ece49655bee085fd6faa3ce2"},"cell_type":"code","source":"print(\"Shortest questions:\")\nlenSorted[[\"question_text\", \"target\"]].iloc[:20]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"140cd51ac11664fad6592dc8d6611381050d24c9"},"cell_type":"markdown","source":"It looks like most of the shortest questions are flagged as insincere. If we keep these sentences, the model may learn that all short questions are insincere which is not the case. On the other hand if we remove the shortest questions, the model may perform better on longer questions which make up most of the dataset anyway, because it doesn't have to worry about this special case of sentence.\n\nIf the task of predicting insincere short questions is different enough from the process of predicting insincere longer questions, we will get better results training one model for each kind of sentence.\n\nLet's check the distribution for all lengths and then we will compare it to the distribution of short sentences"},{"metadata":{"trusted":true,"_uuid":"1e9a20ab44566222db6b5014e529b3e22f9b8aeb"},"cell_type":"code","source":"vals = lenSorted.iloc[:-10]\ninsincere = vals[(vals[\"target\"] == 1)][\"question_len\"]\nsincere = vals[(vals[\"target\"] == 0)][\"question_len\"]\nplt.figure(figsize=(15,5))\nplt.hist([insincere, sincere], stacked=True, bins = 50,\n         label=[\"insincere\", \"sincere\"])\nplt.legend()\nplt.title(\"Distribution of length of questions\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b6eb26d923972e1f54d69746f74f9617383e328c"},"cell_type":"code","source":"lowVals = lenSorted.iloc[:100]\ninsincere = lowVals[(lowVals[\"target\"] == 1)][\"question_len\"]\nsincere = lowVals[(lowVals[\"target\"] == 0)][\"question_len\"]\nplt.figure(figsize=(15,5))\nplt.hist([insincere, sincere], stacked=True, bins = 12,\n         label=[\"insincere\", \"sincere\"])\nplt.legend()\nplt.title(\"Distribution of the shortest 100 questions\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"22eb44a05b46b90587f3f3a12e3d8cd34e87a2e2"},"cell_type":"markdown","source":"From the plots we see that, the shortest questions have indeed a very different distribution than the rest of questions since most of them are flagged as insincere, however this trend only holds for some of the shortest of them all. The trend of having mostly sincere questions starts appearing as soon as length = 10. This means that the the distribution is different for a minute amount of questions and they will not have a big impact on the model performance. Anyway it may be beneficial to remove these sentences and let the model focus on the general case.\n\nNow let's check the longest questions"},{"metadata":{"trusted":true,"_uuid":"9880a57bfeb589b2d99bf747df29f0e5de67b90b"},"cell_type":"code","source":"lenSorted[[\"question_text\", \"target\"]].iloc[-20:]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"71ca1ed370f74486e3e3fbcae48e0782ee21b972"},"cell_type":"code","source":"highVals = lenSorted.iloc[-100:-10]\ninsincere = highVals[(highVals[\"target\"] == 1)][\"question_len\"]\nsincere = highVals[(highVals[\"target\"] == 0)][\"question_len\"]\nplt.figure(figsize=(15,5))\nplt.hist([insincere, sincere], stacked=True, bins = 12,\n         label=[\"insincere\", \"sincere\"])\nplt.legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"070c37eea27436426bbb8289dac74da8a24ca96b"},"cell_type":"markdown","source":"The distribution of sincere/insincere questions starts to difer from the general trend at around length = 255, but the amount of affected sentences is once again vary small and the distribution change is not too much different anyway. We may be safe by keeping the logest questions.\n\nTo get a feeling of the questions text content we can check out the most frequent words, bigrams and trigrams. That's what we will do in the next scripts, including a word cloud for the most common words as well."},{"metadata":{"trusted":true,"_uuid":"3479cca2843dbe17e92f97f20ecfa790982b41d1"},"cell_type":"code","source":"def word_analysis(tokens):\n    frequency_dict = Counter(tokens)\n    most_common = frequency_dict.most_common(200)\n    most_common = [entry for entry in most_common if (entry[0] not in stop_words)]\n    wc = WordCloud(background_color='white', width=2500, height=500)\n    wc.generate_from_frequencies(dict(most_common))\n    \n    most_common = dict(list(most_common[:20]))\n    \n    fig, axes = plt.subplots(2,1, figsize=(25,10))\n    axes[0].imshow(wc)\n    axes[0].axis('off')\n    axes[1].bar(most_common.keys(), most_common.values())\n    plt.xticks(rotation=45)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b790fd1da70211a92429ddfce375e9ceeda52dc9"},"cell_type":"code","source":"plt.rcParams.update({'font.size': 22})\nword_analysis(words)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"89e31b35d7b71cabfbf43365a008992071124378"},"cell_type":"markdown","source":"The most common words are quite generic, except for the word \"India\", Quora is probably very popular in India."},{"metadata":{"trusted":true,"_uuid":"4a05ff1838f909684676af03756c081b505a76bf"},"cell_type":"code","source":"bigrams = ngrams(words,2)\nbigram_dist = FreqDist()\nbigram_dist.update(bigrams)\n\nplt.figure(figsize=(20,8))\nbigram_dist.plot(25)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"662342ccdf5b37b33a94370254742d59ea8470a5"},"cell_type":"code","source":"trigrams = ngrams(words,3)\ntrigram_dist = FreqDist()\ntrigram_dist.update(trigrams)\n\nplt.figure(figsize=(20,8))\ntrigram_dist.plot(25)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"96e823f6229cde76a216a6be54a37940010be41a"},"cell_type":"markdown","source":"the most common bigrams and trigrams are \"question forming\" expressions, which make a lot of sense in a website for asking questions. Probably not one of these expressions will help the model at predicting whether the question is sincere of not, since they are very generic. they are also the most common expressions so there is a lot of not-so-useful data in the dataset."},{"metadata":{"trusted":true,"_uuid":"eb1c4309cdfcf0601fac06017477f3c1bee3e85a"},"cell_type":"code","source":"\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}