{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"**Hi!** \n\nI've been learning python and machine learning for 6 weeks now, and I wanted to try working on a Kaggle competition. So this is my very first kernel on Kaggle!\n\nI hope you will enjoy this kernel and find interesting highlights on this dataset from Quora.\n\nThis is still a work in progress, the next steps of the ML part will arrive soon!\n\nHere is the current architecture of this notebook :\n\n**Exploratory Data Analysis:**\n\n    - First insights\n    - Working on meta-features\n    - A little bit of topic modeling\n    - Insincere questions topic modeling with bi-grams\n    - Preprocessing with Spacy\n    - BONUS : questions id in train and test datasets\n\n**Machine Learning**\n\n    - Baseline model\n    "},{"metadata":{"_uuid":"90afdc069a17ccf2a039730044c269a5e5a47b75"},"cell_type":"markdown","source":"![ReadyURL](https://media.giphy.com/media/12WPxqBJAwOuIM/giphy.gif \"AreYouReady\")"},{"metadata":{"_kg_hide-input":false,"_kg_hide-output":false,"trusted":true,"_uuid":"3f40a9ebe18c1df6739e72110957400898a661d0"},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"import os\nimport string\nimport pickle\nimport random\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport spacy\nimport nltk\n\nimport warnings\nwarnings.filterwarnings('ignore')\n\nstop_words = set(nltk.corpus.stopwords.words('english')) \n\nsns.set()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"13a73ba4229fa22ac77e992e0d1846d8f064d360"},"cell_type":"markdown","source":"First I create the filepaths to the csv files, load the dataframes and check if everything is ok.**"},{"metadata":{"trusted":true,"_uuid":"94b0be36a83d7c0303982d19a500e3c096b713f8"},"cell_type":"code","source":"filepath_train = os.path.join('..', 'input', 'train.csv')\nfilepath_test = os.path.join('..', 'input', 'test.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ffeba9ca8bcfbb0102de09041db4ca9550ce09ca"},"cell_type":"code","source":"df_train = pd.read_csv(filepath_train)\ndf_test = pd.read_csv(filepath_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f5c6cd8a09574ab0d79df07317e9ea0e9168fd8b"},"cell_type":"code","source":"df_train.shape, df_test.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"942ded7d207a94c78da5fe1ad214adfec2cd0fbf"},"cell_type":"code","source":"df_train.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"994f8d114be2d31110d6e645484089977ff5110e"},"cell_type":"markdown","source":"![ExploURL](https://media.giphy.com/media/VEakclh6bmV4A/giphy.gif \"Exploration\")\n\nThe train dataframe is loaded, we can start our exploratory data analysis!"},{"metadata":{"_uuid":"0b3b9df679cdab9c07bae4d9f5ebbb230b1c65df"},"cell_type":"markdown","source":"# Exploratory Data Analysis"},{"metadata":{"_uuid":"7e8ca28b1453429608929ff131781e81edaed946"},"cell_type":"markdown","source":"## First insights"},{"metadata":{"trusted":true,"_uuid":"a7df01ec579ffd6afc6c05fbeb871b17a9458b4d"},"cell_type":"code","source":"df_train.info()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3fe61a7fe084450b648e3c8f4376b1cdc5f52e70"},"cell_type":"markdown","source":"We can see there are no missing values, which is great so far!\n\nLet's see how our targets are distributed."},{"metadata":{"trusted":true,"_uuid":"4a856c79aa603b33cdd43664d286d46577c1a0be"},"cell_type":"code","source":"# Separating the targets from the feature we will work on\nX = df_train.drop(['qid', 'target'], axis=1)\ny = df_train['target']\nX.shape, y.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6a2af06ce1226c6316bae1dc0f651e32c24e0a62"},"cell_type":"code","source":"n_0 = y.value_counts()[0]\nn_1 = y.value_counts()[1]\nprint('{}% of the questions in the train set are tagged as insincere.'.format((n_1*100/(n_1 + n_0)).round(2)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"55277528ea43860ad4c3f0b0ce490de1612a7aae","scrolled":false},"cell_type":"code","source":"# Visualizing some insincere questions randomly chosen\n\nnp.array(X[y==1])[np.random.choice(len(np.array(X[y==1])), size=15, replace=False)]\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"01ecdd8d8f34b0437fa47c977f3f6a4b94a8607b"},"cell_type":"markdown","source":"## Working on meta-features"},{"metadata":{"_uuid":"b9077b049b4370e7ba022cce652806d64bc2f8e7"},"cell_type":"markdown","source":"Here we'll try to create meta features to understand better the structure of the questions."},{"metadata":{"trusted":true,"_uuid":"8f5b8c680c66d403c28bf2d573130bd10a41fa10"},"cell_type":"code","source":"# Custom function to create the meta-features we want from X and add them in a new DataFrame\n\ndef add_metafeatures(dataframe):\n    new_dataframe = dataframe.copy()\n    questions = df_train['question_text']\n    n_charac = pd.Series([len(t) for t in questions])\n    n_punctuation = pd.Series([sum([1 for x in text if x in set(string.punctuation)]) for text in questions])\n    n_upper = pd.Series([sum([1 for c in text if c.isupper()]) for text in questions])\n    new_dataframe['n_charac'] = n_charac\n    new_dataframe['n_punctuation'] = n_punctuation\n    new_dataframe['n_upper'] = n_upper\n    return new_dataframe","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6d77fcd9649f5db774b165327ef3eda2d3a630c8"},"cell_type":"code","source":"X_meta = add_metafeatures(X)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d709b06b69ea836f262ab60be10e369534dc39e7"},"cell_type":"code","source":"X_meta.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"57fcf45edb116eb0e320124e5dc63187704f510f"},"cell_type":"code","source":"print('Number of characters description : \\n\\n {} \\n\\n Number of punctuations description : \\n\\n {} \\n\\n Number of uppercase characters description : \\n\\n {}'.format(\n    X_meta['n_charac'].describe(),\n    X_meta['n_punctuation'].describe(), \n    X_meta['n_upper'].describe()))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cc85ee4ad36d9db92fc88508738ff075a4264313"},"cell_type":"markdown","source":"Let's visualize our meta-features!"},{"metadata":{"trusted":true,"_uuid":"3319ef7ad5805c25fd8b12fa04422989eb7b201b"},"cell_type":"code","source":"# Separating X_meta with our targets in y\n\nX_meta_sincere = X_meta[y==0]\nX_meta_insincere = X_meta[y==1]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"67bbea0102a33c593e3a45c467d019aa3413f8a6"},"cell_type":"code","source":"_, axes = plt.subplots(2, 3, sharey=True, figsize=(18, 8))\nsns.boxplot(x=X_meta['n_charac'], y=y, orient='h', ax=axes.flat[0]);\nsns.boxplot(x=X_meta['n_punctuation'], y=y, orient='h', ax=axes.flat[1]);\nsns.boxplot(x=X_meta['n_upper'], y=y, orient='h', ax=axes.flat[2]);\n\nX_meta_charac = X_meta[X_meta['n_charac']<400]\nX_meta_punctuation = X_meta[X_meta['n_punctuation']<10]\nX_meta_upper = X_meta[X_meta['n_upper']<15]\n\nsns.boxplot(x=X_meta_charac['n_charac'], y=y, orient='h', ax=axes.flat[3]);\nsns.boxplot(x=X_meta_punctuation['n_punctuation'], y=y, orient='h', ax=axes.flat[4]);\nsns.boxplot(x=X_meta_upper['n_upper'], y=y, orient='h', ax=axes.flat[5]);","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"80a0e3f89d7a8755e9a7e2a0d44ddcc6d0021f75"},"cell_type":"markdown","source":"The second line of graphs is just a zoom in the interesting parts of the grpahs on the first line\n\nWe can see there is a slight difference between the distribution of the number of characters, of punctuations and of uppercase characters for sincere and insincere questions.\n\nLets take a look at the outliers for the number of characters (ie : n_charac > 400)"},{"metadata":{"trusted":true,"_uuid":"072bb3ca006ba68385ce738d9d3a69ae5f1dd2d5"},"cell_type":"code","source":"pd.concat([X_meta[X_meta['n_charac']>400], y], axis=1, join='inner')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fbccef1de16da180301cc13302baf0fb7b6c30f6"},"cell_type":"markdown","source":"As we can see, 3 of these questions are about math problems, and one about Star Trek."},{"metadata":{"_uuid":"c6cc8bc5bfa1f4ee9b876bdfebebf89b2167d529"},"cell_type":"markdown","source":"![SpockUrl](https://media.giphy.com/media/vp122eOzO0Hxm/giphy.gif \"fascinating\")"},{"metadata":{"_uuid":"7dfc6c7d1819ce6c08706ff23fac10785e8b10c7"},"cell_type":"markdown","source":"Over the three math questions, 1 has been classified as sincere, 2 as insincere. Lets take a closer look at the full text for these questions:"},{"metadata":{"trusted":true,"_uuid":"b9cdb4f411b3120946ae7725ac4f997a785159ff"},"cell_type":"code","source":"print(np.array(X_meta[X_meta['n_charac']>400]['question_text']))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5a2ff478d6eee378e5ef146efb81c6cdb7875ad8"},"cell_type":"markdown","source":"We can suppose that the math question classified as sincere might have been missclassified. \nMayber other questions have also been missclassified, leading to inaccuracy for our models."},{"metadata":{"_uuid":"5c4206053fd619b5d6c812b15dadc3fe1e75c62e"},"cell_type":"markdown","source":"Maybe it is interesting to calculate the punctuation ratio:"},{"metadata":{"trusted":true,"_uuid":"646b6cff39e7424ff0be4dde6b5bd655b89f47c5"},"cell_type":"code","source":"punctuation_ratio = 100*X_meta['n_punctuation'] / X_meta['n_charac']","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c5fa714a153fe3a8a822d3db483bd35dd62c7c5c"},"cell_type":"code","source":"plt.figure(figsize=(18, 8))\nsns.boxplot(punctuation_ratio, y, orient='h');","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ea46a285a7caaba159d3aac6e39237242f3cbabd"},"cell_type":"markdown","source":"We can see that the distribution here is slightly the same for sincere and insincere questions, so it will not be very usefull to keep this ratio as a feature.  Nevertheless, let's again take a look at these outliers over 50."},{"metadata":{"trusted":true,"_uuid":"ba891d01900f6e687b3cb5919357d470f57f51f6","scrolled":true},"cell_type":"code","source":"pd.concat([X_meta[punctuation_ratio>50], y, punctuation_ratio], axis=1, join='inner')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ab235b7aa2ceaf6b4205eb4addcea8f191efeacf"},"cell_type":"markdown","source":"Again, math formulas as sincere questions, nothing to see here, go next."},{"metadata":{"_uuid":"01b31205832acfb8a4938701d51ae41fb611e9e9"},"cell_type":"markdown","source":"![MoveURL](https://media.giphy.com/media/l0MYsTuL1N15t4FiM/giphy.gif \"MoveAlong\")"},{"metadata":{"_uuid":"86ec1da28be4dab53575dc99b1625b4b1f7fcbf9"},"cell_type":"markdown","source":"## A little bit of topic modeling"},{"metadata":{"_uuid":"68ac9dec85e2d402db2c832bdae6214009450ef9"},"cell_type":"markdown","source":"Here we'll try to find wich topics appear more often in sincere and insincere questions. To do so, I'll use a `CountVectorizer` and a `TruncatedSVD`in a pipeline (yes the pipeline is only here to show off)."},{"metadata":{"trusted":true,"_uuid":"d9ba4095d64261eff42e5537347f9eccd941264f"},"cell_type":"code","source":"from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer\nfrom sklearn.decomposition import TruncatedSVD\nfrom sklearn.pipeline import Pipeline","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"08306c6eb6db3aa3fb35930b7ac31af900e7e2fd"},"cell_type":"code","source":"vectorizer = CountVectorizer(stop_words='english')\nsvd = TruncatedSVD(n_components=1, random_state=42)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c48a72b52ba494f6c26cf704f6367f24d2c8366e"},"cell_type":"code","source":"preprocessing_pipe = Pipeline([('vectorizer', vectorizer), ('svd', svd)])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"156c63a5a8e6db95659a5542478d87b82d64f1de"},"cell_type":"code","source":"# Building the latent semantic analysis dataframe for sincere and insincere questions\n\nlsa_insincere = preprocessing_pipe.fit_transform(X[y==1]['question_text'])\ntopics_insincere = pd.DataFrame(svd.components_)\ntopics_insincere.columns = preprocessing_pipe.named_steps['vectorizer'].get_feature_names()\n\nlsa_sincere = preprocessing_pipe.fit_transform(X[y==0]['question_text'])\ntopics_sincere = pd.DataFrame(svd.components_)\ntopics_sincere.columns = preprocessing_pipe.named_steps['vectorizer'].get_feature_names()\n\ntopics_insincere.shape, topics_sincere.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"0e6ba6d00452f0188d7e029857bc2983e4517de1"},"cell_type":"code","source":"fig, axes = plt.subplots(1, 2, figsize=(22,10));\n\ntopics_sincere.iloc[0].sort_values(ascending=False)[:30].sort_values().plot.barh(ax=axes[0]);\ntopics_insincere.iloc[0].sort_values(ascending=False)[:30].sort_values().plot.barh(ax=axes[1]);","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d487d28f12c1d6ca154a6286279f37f572fda00c"},"cell_type":"markdown","source":"Soooo here we are. Some words are obviously more common in insincere questions, like 'white' and 'black', but other words of importance in our LSA are shared by both sincere and insincere questions at the same level, like 'people'. Maybe working on bi-grams or tri-grams will help us define more precislely what an insincere question looks like.\nBut before this, I wanted to try a Truncated SVD with 2 components, to see if I'll be able to link these two components to the results of my bi-grams study later."},{"metadata":{"_uuid":"f415a81e13971721280b58dc1b03b8b1d8125f4b"},"cell_type":"markdown","source":"![LetsgoURL](https://media.giphy.com/media/3o7TKUM3IgJBX2as9O/giphy.gif \"LetsGo\")"},{"metadata":{"trusted":true,"_uuid":"0b036c485382419317467676c062c6965c7610a9"},"cell_type":"code","source":"vectorizer = CountVectorizer(stop_words='english')\nsvd = TruncatedSVD(n_components=2, random_state=42)\n\npreprocessing_pipe = Pipeline([('vectorizer', vectorizer), ('svd', svd)])\n\n# Building the latent semantic analysis dataframe for sincere and insincere questions\n\nlsa_insincere_2 = preprocessing_pipe.fit_transform(X[y==1]['question_text'])\ntopics_insincere_2 = pd.DataFrame(svd.components_)\ntopics_insincere_2.columns = preprocessing_pipe.named_steps['vectorizer'].get_feature_names()\n\nlsa_sincere_2 = preprocessing_pipe.fit_transform(X[y==0]['question_text'])\ntopics_sincere_2 = pd.DataFrame(svd.components_)\ntopics_sincere_2.columns = preprocessing_pipe.named_steps['vectorizer'].get_feature_names()\n\n\nfig_1, axes_1 = plt.subplots(1, 2, figsize=(18, 8))\nfor i, ax in enumerate(axes_1.flat):\n    topics_insincere_2.iloc[i].sort_values(ascending=False)[:30].sort_values().plot.barh(ax=ax)\n    \nfig_2, axes_2 = plt.subplots(1, 2, figsize=(18, 8))\nfor i, ax in enumerate(axes_2.flat):\n    topics_sincere_2.iloc[i].sort_values(ascending=False)[:30].sort_values().plot.barh(ax=ax)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a967a4500aa79d21595e9d9bdead57f03c2ca9d8"},"cell_type":"markdown","source":"### Insincere questions topic modeling with bi-grams"},{"metadata":{"trusted":true,"_uuid":"391f912e17dcc8257fcf1a7c295c451fb09935d7"},"cell_type":"markdown","source":"Here I will also use a `CountVectorizer` and a `TruncatedSVD` with 9 components to identify the nine main topics of insincere questions, but with the parameter ngram_range set at (2, 2)  for the `CountVectorizer`"},{"metadata":{"trusted":true,"_uuid":"f87632172e94aabc7a2693187502a3e9cb420d0f","scrolled":false},"cell_type":"code","source":"vectorizer_22 = CountVectorizer(stop_words='english', ngram_range=(2, 2))\nsvd_10c = TruncatedSVD(n_components=9, random_state=42)\n\npreprocessing_pipe = Pipeline([('vectorizer_22', vectorizer_22), ('svd_10c', svd_10c)])\n\n# Building the latent semantic analysis dataframe for insincere questions\n\nlsa_insincere_10c = preprocessing_pipe.fit_transform(X[y==1]['question_text'])\ntopics_insincere_10c = pd.DataFrame(svd_10c.components_)\ntopics_insincere_10c.columns = preprocessing_pipe.named_steps['vectorizer_22'].get_feature_names()\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3e4ab070ee6a91d3f04e8a1b3346ebe00bc17b66"},"cell_type":"code","source":"fig, axes = plt.subplots(3, 3, figsize=(20, 12))\nfor i, ax in enumerate(axes.flat):\n    topics_insincere_10c.iloc[i].sort_values(ascending=False)[:10].sort_values().plot.barh(ax=ax)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e40de45e77289a3be6a506d4fd20c9e3a83e61c8"},"cell_type":"markdown","source":"We can see emerging topics, about Donald Trump or racism for instance. Let's do the same with bi-grams and tri-grams to see what happens."},{"metadata":{"trusted":true,"_uuid":"92488f02a596bd6d3b8e9eb64fde9fe6c28a5176"},"cell_type":"code","source":"vectorizer_23 = TfidfVectorizer(stop_words='english', ngram_range=(2, 3))\nsvd_9c = TruncatedSVD(n_components=9, random_state=42)\n\npreprocessing_pipe = Pipeline([('vectorizer_23', vectorizer_23), ('svd_9c', svd_9c)])\n\n# Building the latent semantic analysis dataframe for insincere questions\n\nlsa_insincere_9c = preprocessing_pipe.fit_transform(X[y==1]['question_text'])\ntopics_insincere_9c = pd.DataFrame(svd_9c.components_)\ntopics_insincere_9c.columns = preprocessing_pipe.named_steps['vectorizer_23'].get_feature_names()\n\nfig, axes = plt.subplots(3, 3, figsize=(20, 12))\nfor i, ax in enumerate(axes.flat):\n    topics_insincere_9c.iloc[i].sort_values(ascending=False)[:10].sort_values().plot.barh(ax=ax)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7a00b4061f64dc490f65be2ee7a03a15e88e84cf"},"cell_type":"markdown","source":"Here we have some issues due to our non-preprocessed data. Indeed, for instance, 'year old girl' and 'year old girls' are two different components for our Vectorizer.\nSo it is maybe time to preprocess ou raw data to make it more explicit!"},{"metadata":{"_uuid":"d06a3bbdd193690f005c98680c8dc9c8c5383703"},"cell_type":"markdown","source":"## Preprocessing with Spacy"},{"metadata":{"_uuid":"b97e083c65d4d044a993abb80e65f22f9101803b"},"cell_type":"markdown","source":"I'll use spacy to preprocess the questions."},{"metadata":{"trusted":true,"_uuid":"a642eb298806c3634c43ea446b50dadb796fcb4b"},"cell_type":"code","source":"nlp = spacy.load('en_core_web_sm', disable=['parser', 'ner'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9e5b657f79d8a487c288d1d47825b3a4b23ba627"},"cell_type":"code","source":"#Custom function to preprocess the questions\n\ndef preprocess(X):\n    docs = nlp.pipe(X)\n    lemmas_as_string = []\n    for doc in docs:\n        doc_of_lemmas = []\n        for t in doc:\n            if t.text.lower() not in stop_words and t.text.isalpha() == True:\n                if t.lemma_ !='-PRON-':\n                    doc_of_lemmas.append(t.lemma_)\n                else:\n                    doc_of_lemmas.append(t.text)\n        lemmas_as_string.append(' '.join(doc_of_lemmas))\n    return pd.DataFrame(lemmas_as_string)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"52b29fbf323e128f29f38ededeccd3a9503ec461"},"cell_type":"code","source":"%%time\nX_prep = preprocess(X['question_text'])\nX_prep.to_pickle('X_preprocessed.pkl')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"72b8f174795c76e109a167ff0133666e1af21e41"},"cell_type":"markdown","source":"After having preprocessed X one time, I saved the result as a pickle file and saved it, to avoid having to wait 15min each time I run this notebook. I therefore commented the cell code above and created the one below to load X preprocessed from the pickle file."},{"metadata":{"trusted":true,"_uuid":"998caf6e0707dd877addd414a59132365d5770e3"},"cell_type":"code","source":"X_prep = pd.read_pickle('X_preprocessed.pkl')\nX_prep.columns = ['question_text']\nX_prep.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a3eb2cbf67cfdef43e8f53511c594d2e99b60b40"},"cell_type":"markdown","source":"Now let's use our topic modeling code on this preprocessed `DataFrame` !"},{"metadata":{"trusted":true,"scrolled":false,"_uuid":"4292a03184f397f4517061192d4bdafb50399fba"},"cell_type":"code","source":"vectorizer_23 = TfidfVectorizer(stop_words='english', ngram_range=(2, 3))\nsvd_9c = TruncatedSVD(n_components=9, random_state=42)\n\npreprocessing_pipe = Pipeline([('vectorizer_23', vectorizer_23), ('svd_9c', svd_9c)])\n\n# Building the latent semantic analysis dataframe for insincere questions\n\nlsa_insincere_9c = preprocessing_pipe.fit_transform(X_prep[y==1]['question_text'])\ntopics_insincere_9c = pd.DataFrame(svd_9c.components_)\ntopics_insincere_9c.columns = preprocessing_pipe.named_steps['vectorizer_23'].get_feature_names()\n\nfig, axes = plt.subplots(3, 3, figsize=(20, 12))\nfor i, ax in enumerate(axes.flat):\n    topics_insincere_9c.iloc[i].sort_values(ascending=False)[:10].sort_values().plot.barh(ax=ax)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"50a9c08df7cf586f7d4f34e89952deb3feda0a05"},"cell_type":"markdown","source":"Based on this cleaner version of the histograms we plotted before, we can list a few main topics in insincere questions:\n\n- Donald Trump\n- Racism\n- Sex with a family member\n- People asking stupid questions on quora\n\nIt seems like we lost information, compared to the previous LSA on non-preprocessed data (less tdifferent topics in the 9 main components). I'll keep this in mind when i'll build machine learning models, to decide wether i'll train them on the raw or the preprocessed dataset.\n"},{"metadata":{"_uuid":"f3b52435e8d4e551877984bc191944b1a3e7e8d0"},"cell_type":"markdown","source":"To conclude this EDA : \n\n- Insincere questions main topics are really interesting from a social point of view. Someone with societal analysis skills would probably be interested by taking a look at these questions!\n\n- Some of the insincere questions might have been missclassified, which could lead do a decreased accuracy of our ML models.\n\n- Some people posting questions on Quora really are insane!"},{"metadata":{"_uuid":"cc778535d8e488da365869c7c4d3c64c11db682d"},"cell_type":"markdown","source":"![SocietyURL](https://media.giphy.com/media/10E3mQGzAxWFZm/giphy.gif \"Society\")"},{"metadata":{"_uuid":"64c94b15b606775dcd580a35366baef9e0f0a003"},"cell_type":"markdown","source":"## Bonus : questions id in train and test datasets"},{"metadata":{"_uuid":"002f4d0bb8a9bad0006984144e8c667cb0090fda"},"cell_type":"markdown","source":"I was wondering wether the test dataset was an extract from the train one, or completely different. To answer this existential question, I've been working on the 'qid' column.\n\nThe question id is an hexadecimal number. The first step here is to extract this value, and see if the questions are ordered by id."},{"metadata":{"trusted":true,"_uuid":"40bfe90a1ff321c78bd7cdbecdd045f17a0358ad"},"cell_type":"code","source":"df_train_qid = df_train.copy()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2faf07182183c19ecdd3ef0e9c9118c6a867d208"},"cell_type":"code","source":"df_train_qid['qid_base_ten'] = df_train_qid['qid'].apply(lambda x : int(x, 16))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5a7d2e450ef558f18bf28ddb144bfcc937cc1ef9"},"cell_type":"code","source":"df_train_qid.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"684657ac1d6ba0b77e0ebd1c487424646db4a009"},"cell_type":"code","source":"min_qid = df_train_qid['qid_base_ten'].min()\nmax_qid = df_train_qid['qid_base_ten'].max()\ndf_train_qid['qid_base_ten_normalized'] = df_train_qid['qid_base_ten'].apply(lambda x : (x - min_qid)/min_qid)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7da268d4145b1eb889783456cfc2bbe543e21c55"},"cell_type":"code","source":"plt.figure(figsize=(18, 8));\nplt.scatter(x=df_train_qid['qid_base_ten_normalized'][:100], y=df_train_qid.index[:100]);\nplt.xlabel('qid_base_ten_normalized');\nplt.ylabel('Question index in df_train_qid');","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2b05794f9f01195ae858fd94bac427835334ef8b"},"cell_type":"markdown","source":"As I suspected, questions are indeed sorted by ascending question id in our train dataset. Let's see if it is the same in the test one."},{"metadata":{"trusted":true,"_uuid":"d00809011248c4242799f4dddcab516aafb55471"},"cell_type":"code","source":"df_test_qid = df_test.copy()\n\ndf_test_qid['qid_base_ten'] = df_test_qid['qid'].apply(lambda x : int(x, 16))\n\ndf_test_qid['qid_base_ten_normalized'] = df_test_qid['qid_base_ten'].apply(lambda x : (x - min_qid)/min_qid)\n\nplt.figure(figsize=(18, 8));\nplt.scatter(x=df_test_qid['qid_base_ten_normalized'][:100], y=df_test_qid.index[:100]);\nplt.xlabel('qid_base_ten_normalized');\nplt.ylabel('Question index in df_test_qid');","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6322575dbe36fc90bb130d078f13ecd6a54d47f2"},"cell_type":"markdown","source":"Here again, questions are sorted by ascending question id ! Now I wonder if I can know how Quora has made its train an test datasets. Is it with a random (and stratified?) train.test split, or a simple split based on the id?\n\nTo get the answer, I have merged the train and test dataframes, with the 'qid_base_ten_normalized' column, sorted by ascending 'qid_base_ten_normalized' and reset the index."},{"metadata":{"trusted":true,"_uuid":"4b4e6291867ebe7eb16fc5fa7490d3fed33aee4b"},"cell_type":"code","source":"df_train_qid.drop('target', axis=1, inplace=True)\ndf_train_qid['test_or_train'] = 'train'\ndf_test_qid['test_or_train'] = 'test'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"f07f9cd139da5b599976abab2376c257a6b55357"},"cell_type":"code","source":"df_qid = pd.concat([df_train_qid, df_test_qid]).sort_values('qid_base_ten_normalized').reset_index()\ndf_qid.drop('index', axis=1, inplace=True)\ndf_qid.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"scrolled":false,"_uuid":"ad8f1e52af557dd6ea8f09dc6e397c9001b2a029"},"cell_type":"code","source":"df_qid_train = df_qid[df_qid['test_or_train']=='train']\ndf_qid_test = df_qid[df_qid['test_or_train']=='test']\n\nplt.figure(figsize=(18, 8));\nplt.scatter(x=df_qid_train['qid_base_ten_normalized'], y=df_qid_train.index, label='Train');\nplt.scatter(x=df_qid_test['qid_base_ten_normalized'], y=df_qid_test.index, label='Test',s=5);\nplt.xlabel('qid_base_ten_normalized');\nplt.ylabel('Question index');\nplt.title('qid_base_ten_normalized for train and test datasets')\nplt.legend();","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3c0c6a597e9ed7680a38ee76ae9ae43f42cd3e69"},"cell_type":"markdown","source":"So the question ids range of the test dataset is sthe same as the question ids range for the train one. The test and train datasets come as expected from a random train/test split on a single dataset.\nThe figure below confirms the 'random' choice of the elemnts for the test dataset."},{"metadata":{"trusted":true,"_uuid":"fc743ea3640d16eec975832a61e00a730a3d00a5"},"cell_type":"code","source":"plt.figure(figsize=(18, 8));\nplt.scatter(x=df_qid_train['qid_base_ten_normalized'][:1500], y=df_qid_train.index[:1500], label='Train');\nplt.scatter(x=df_qid_test['qid_base_ten_normalized'][:50], y=df_qid_test.index[:50], label='Test',s=150, marker='d');\nplt.xlabel('qid_base_ten_normalized');\nplt.ylabel('Question index');\nplt.title('qid_base_ten_normalized for the first 1500 train points and 50 test points')\nplt.legend();","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ed139a69ec6e1803287c0a424a574c23cbab624c"},"cell_type":"markdown","source":"Here, we still can't figure out if the train/test plit has been done in a stratified way!\n\nThat's all for this bonus part on questions id, it is not that useful, but it was working on it was fun!\n\n![FolksURL](https://media.giphy.com/media/upg0i1m4DLe5q/giphy.gif \"Folks\")"},{"metadata":{"_uuid":"27e8e345d25f40e0cc119fb09ee89bf5475c14d2"},"cell_type":"markdown","source":"   # Machine Learning"},{"metadata":{"_uuid":"9097715e52645a5f57b503d3079d5ecd0701bd59"},"cell_type":"markdown","source":"Now is the most difficult part for me. As I said, i've only been learning python and machine learning for 6 weeks. So if you have already read until this point, thank you, and do not hesitate to give me advices on how I could improve my kernel! "},{"metadata":{"_uuid":"6a27dba133ca54b699fa71dea616246453cc351e"},"cell_type":"markdown","source":"## Splitting into train and test with `sklearn.model_selection.train_test_split`"},{"metadata":{"trusted":true,"_uuid":"4b7d5c95e114edcc98d72a322fc9bc158e4d3ec9"},"cell_type":"code","source":"from sklearn.model_selection import train_test_split","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f650f9eefbf04823d35ad8814c1b286e38d13d44"},"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(X['question_text'], y, test_size=.2, random_state=42, stratify=y)\nX_train.shape, y_train.shape, X_test.shape, y_test.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9478e9b9529d2e0565a49442fe7df3013220eaf9"},"cell_type":"markdown","source":"## Importing utils from sklearn"},{"metadata":{"trusted":true,"_uuid":"c54a4be8d3a06209ab8998ba960409412d40c116"},"cell_type":"code","source":"from sklearn.pipeline import Pipeline\nfrom sklearn.model_selection import GridSearchCV\n\nfrom sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer\nfrom sklearn.decomposition import TruncatedSVD\n\nfrom sklearn.metrics import confusion_matrix\nfrom sklearn.metrics import classification_report\nfrom sklearn.metrics import f1_score","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f19e522495d8e9f0b4d674f7f7911a0901dac202"},"cell_type":"markdown","source":"## Baseline model"},{"metadata":{"_uuid":"5bee06c787faccdb5c6af684392ac1e825c9fdde"},"cell_type":"markdown","source":"Let's create a first, simple model, that will be my baseline model. I have to chosen to use a TfidFVectorizer and a LogisticRegression on my raw data (ie: no preprocessing)"},{"metadata":{"trusted":true,"_uuid":"f2e41110477c939bd8ea96b9e075f060e8d8b80c"},"cell_type":"code","source":"from sklearn.feature_extraction.text import TfidfVectorizer\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.pipeline import Pipeline","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f481ab44611e4765cf86299491c2adf1787e36d6"},"cell_type":"code","source":"tfidf = TfidfVectorizer(stop_words='english', ngram_range=(1, 3))\nlr = LogisticRegression()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"10915130643ccd45776d56aaa76cd4a99e20c0da"},"cell_type":"code","source":"pipe_baseline = Pipeline([('tfidf', tfidf), ('lr', lr)])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e32b82d2d36fc6d4cb1122b11bd38e0f0827246f"},"cell_type":"code","source":"pipe_baseline.fit(X_train, y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"af8dba6abaf291c4a3f6d5d576abfa4bc4cd2cff"},"cell_type":"code","source":"y_pred = pipe_baseline.predict(X_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f6bcc7b553bc7ca776bb7cc79b641ad74cb26bf1"},"cell_type":"code","source":"cm = confusion_matrix(y_test, y_pred)\n\nax = plt.gca()\nsns.heatmap(cm, cmap='Blues', cbar=False, annot=True, xticklabels=y_test.unique(), yticklabels=y_test.unique(), ax=ax);\nax.set_xlabel('y_pred');\nax.set_ylabel('y_true');\nax.set_title('Confusion Matrix');","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"89cf10a84e8dd8b0abaebd8c87cdf34d1cb4ea61"},"cell_type":"code","source":"cr = classification_report(y_test, y_pred)\nprint(cr)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1e6d46762bb8e3270f74cbcd74f83da9a976ee3c"},"cell_type":"markdown","source":"Ok that's a beginning! Lets work with predict_proba to find the best threshold that optimizes our f1_score for insincere questions (target = 1)"},{"metadata":{"trusted":true,"_uuid":"32e20de86cb10897f2dea9166cecf138f2c738e2"},"cell_type":"code","source":"y_prob = pipe_baseline.predict_proba(X_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0a9254b444042f16829c11985bf92223600d46fc"},"cell_type":"code","source":"best_threshold = 0\nf1=0\nfor i in np.arange(.1, .51, 0.01):\n    y_pred = [1 if proba>i else 0 for proba in y_prob[:, 1]]\n    f1score = f1_score(y_pred, y_test)\n    if f1score>f1:\n        best_threshold = i\n        f1=f1score\n        \ny_pred = [1 if proba>best_threshold else 0 for proba in y_prob[:, 1]]\nf1 = f1_score(y_pred, y_test)\nprint('The best threshold is {}, with an f1_score of {}'.format(best_threshold, f1))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"76e28231b8e6c9c6c52ac500fe1d6289ca95faa1"},"cell_type":"markdown","source":"With a simple `LogisticRegression` and a `TfidfVectorizer` we already have an f1 score of 0.59!\n\nNext step is to try to improve this model, tuning the LogisticRegression C parameter for instance.\n\n![StartURL](https://media.giphy.com/media/3oAt1TznOzEcx3MssU/giphy.gif \"GettingStarted\")"},{"metadata":{"trusted":true,"_uuid":"ddf5fb4a610472ff5901a0f5c0cb764f7e401eb2"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}