{"cells":[{"metadata":{"_uuid":"01f361ddc47e0b386595316fe3d7f4dabbd260db"},"cell_type":"markdown","source":" # Should you Clean your Data ? \n\n### A question that has appeared is whether or not one should apply preprocessing to the text. In fact, most popular kernels appear not to have done a lot, and still reached good scores (0.68-ish)\n\n** In this kernel, I will use a simple Deep Learning model and compare its performance on pre-processed texts with usual methods and on raw texts. **\n\n#### The model is the following one :\n* GloVe Embedding\n* Bidirectional GRU\n* Attention \n* Dense \n\nI did not bother tunning it because it is good enough to highlight the point I'm trying to make\n\n\n#### Feel free to give any feedback, it is always appreciated. (plz upvote !)"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport keras\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport codecs\nimport unidecode\nimport re\nimport spacy\nfrom nltk.corpus import stopwords\nfrom time import time\n\nfrom keras import backend as K\nfrom keras.engine.topology import Layer\nfrom keras import initializers, regularizers, constraints, optimizers, layers\nfrom keras.models import Model\nfrom keras.layers import Dense, Embedding, Dropout, Bidirectional, CuDNNGRU, GlobalMaxPool1D, Input\n\n\nimport sys\nimport warnings\n\nif not sys.warnoptions:\n    warnings.simplefilter(\"ignore\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"63530632439404a85540565eb31c6390bddb33e9"},"cell_type":"markdown","source":"## Loading data"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"df = pd.read_csv(\"../input/train.csv\")\nprint(\"Number of texts: \", df.shape[0])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"96012bbe09075239c432cf48a204ca2a35f85c16"},"cell_type":"code","source":"for i in range(5):\n    print(df['question_text'][df.index[i]])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"258ba5f4c4cb16bd6597d4d2ea81feaa21c60a85"},"cell_type":"markdown","source":"** So far, texts look quite complicated, I'm going to apply some usual text processing techniques to simplify them **\n* Dealing with contractions ('t, 've and other stuff)\n* Removing numbers and special characters\n* Lowering letters\n* Removing Stopwords\n* Lemmatization (keeping only the simple form of the word)"},{"metadata":{"_uuid":"9032f827658e0209171a738add697ce7776c3e7a"},"cell_type":"markdown","source":"## Treating Texts"},{"metadata":{"trusted":true,"_uuid":"275e02486adb44bc55c2714d0c3f3707b8f99e6d"},"cell_type":"code","source":"from nltk import WordNetLemmatizer\n\nwnl = WordNetLemmatizer()\n\ncontraction_mapping = {\"ain't\": \"is not\", \"aren't\": \"are not\",\"can't\": \"cannot\", \"can't've\": \"cannot have\", \"'cause\": \"because\", \"could've\": \"could have\", \"couldn't\": \"could not\", \"couldn't've\": \"could not have\",\"didn't\": \"did not\",  \"doesn't\": \"does not\", \"don't\": \"do not\", \"hadn't\": \"had not\", \"hadn't've\": \"had not have\", \"hasn't\": \"has not\", \"haven't\": \"have not\",  \"he'd\": \"he would\", \"he'd've\": \"he would have\", \"he'll\": \"he will\", \"he'll've\": \"he will have\", \"he's\": \"he is\", \"how'd\": \"how did\", \"how'd'y\": \"how do you\", \"how'll\": \"how will\", \"how's\": \"how is\",  \"I'd\": \"I would\", \"I'd've\": \"I would have\", \"I'll\": \"I will\", \"I'll've\": \"I will have\",\"I'm\": \"I am\", \"I've\": \"I have\", \"i'd\": \"i would\", \"i'd've\": \"i would have\", \"i'll\": \"i will\", \"i'll've\": \"i will have\",\"i'm\": \"i am\", \"i've\": \"i have\", \"isn't\": \"is not\", \"it'd\": \"it would\", \"it'd've\": \"it would have\", \"it'll\": \"it will\", \"it'll've\": \"it will have\",\"it's\": \"it is\", \"let's\": \"let us\", \"ma'am\": \"madam\", \"mayn't\": \"may not\", \"might've\": \"might have\",\"mightn't\": \"might not\",\"mightn't've\": \"might not have\", \"must've\": \"must have\", \"mustn't\": \"must not\", \"mustn't've\": \"must not have\", \"needn't\": \"need not\", \"needn't've\": \"need not have\",\"o'clock\": \"of the clock\", \"oughtn't\": \"ought not\", \"oughtn't've\": \"ought not have\", \"shan't\": \"shall not\",\"sha'n't\": \"shall not\", \"shan't've\": \"shall not have\", \"she'd\": \"she would\", \"she'd've\": \"she would have\", \"she'll\": \"she will\", \"she'll've\": \"she will have\", \"she's\": \"she is\", \"should've\": \"should have\", \"shouldn't\": \"should not\", \"shouldn't've\": \"should not have\", \"so've\": \"so have\",\"so's\": \"so as\", \"this's\": \"this is\",\"that'd\": \"that would\", \"that'd've\": \"that would have\",\"that's\": \"that is\", \"there'd\": \"there would\", \"there'd've\": \"there would have\",\"there's\": \"there is\", \"here's\": \"here is\",\"they'd\": \"they would\", \"they'd've\": \"they would have\", \"they'll\": \"they will\", \"they'll've\": \"they will have\", \"they're\": \"they are\", \"they've\": \"they have\", \"to've\": \"to have\", \"wasn't\": \"was not\", \"we'd\": \"we would\", \"we'd've\": \"we would have\", \"we'll\": \"we will\", \"we'll've\": \"we will have\", \"we're\": \"we are\", \"we've\": \"we have\", \"weren't\": \"were not\", \"what'll\": \"what will\", \"what'll've\": \"what will have\", \"what're\": \"what are\", \"what's\": \"what is\", \"what've\": \"what have\", \"when's\": \"when is\", \"when've\": \"when have\", \"where'd\": \"where did\", \"where's\": \"where is\", \"where've\": \"where have\", \"who'll\": \"who will\", \"who'll've\": \"who will have\", \"who's\": \"who is\", \"who've\": \"who have\", \"why's\": \"why is\", \"why've\": \"why have\", \"will've\": \"will have\", \"won't\": \"will not\", \"won't've\": \"will not have\", \"would've\": \"would have\", \"wouldn't\": \"would not\", \"wouldn't've\": \"would not have\", \"y'all\": \"you all\", \"y'all'd\": \"you all would\",\"y'all'd've\": \"you all would have\",\"y'all're\": \"you all are\",\"y'all've\": \"you all have\",\"you'd\": \"you would\", \"you'd've\": \"you would have\", \"you'll\": \"you will\", \"you'll've\": \"you will have\", \"you're\": \"you are\", \"you've\": \"you have\" } \n\nstop_words = set(stopwords.words('english'))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c69deec8fc27d3e72acadbf2ded5e1b1bbe87602"},"cell_type":"code","source":"def treat_text(text):\n    # Decoding\n    try:\n        decoded = unidecode.unidecode(codecs.decode(text, 'unicode_escape'))\n    except:\n        decoded = unidecode.unidecode(text)\n        \n    # Handling Apostrophes\n    apostrophe_handled = re.sub(\"’\", \"'\", decoded)\n    text = ' '.join([contraction_mapping[t] if t in contraction_mapping else t for t in apostrophe_handled.split(\" \")])\n    \n    # Keeping letters + lowerring\n    text = re.findall(r\"[a-zA-Z]+\", text.lower())\n    \n    # Removing stopwords\n    text = [word for word in text if (word not in stop_words and len(word)>2)]\n    \n    # Lemming\n    text = [wnl.lemmatize(word) for word in text]\n    \n    # Removing repetitions\n    text = re.sub(r'(.)\\1+', r'\\1\\1', ' '.join(text))\n    \n    return text","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1b9bd8a9edfc242b388943da14f34160a0648aca","scrolled":true},"cell_type":"code","source":"t0 = time()\nprint('Cleaning data ... ')\ndf['treated_text'] = df['question_text'].transform(treat_text)\nprint(f\"Data cleaned in {round(time() - t0, 1)} seconds\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"77dc37c01e2db1269ad0694496c06ea56c32f91a"},"cell_type":"code","source":"for i in range(5):\n    print(df['treated_text'][df.index[i]])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7e3439e3097565d22ce2625389f20e33f860abd4"},"cell_type":"markdown","source":"Harder to understand, but much simpler !"},{"metadata":{"_uuid":"734199049d49ae5229951fa2dbd3c4aca3ad8355"},"cell_type":"markdown","source":"## Text lengths"},{"metadata":{"trusted":true,"_uuid":"e08b5cc3ac58950dfe8f0611c5e79d833606a241"},"cell_type":"code","source":"df['length'] = df['question_text'].transform(lambda x: len(x.split(' ')) // 5 * 5)\ndf['treated_length'] = df['treated_text'].transform(lambda x: len(x.split(' ')) // 5 * 5)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"27dac805983bd0efeb537b37ea62761f5c61716b"},"cell_type":"code","source":"plt.figure(figsize=(12,8))\nsns.countplot(df['length'])\nplt.title('Length repartiton (rounded down to the 5)')\nplt.yscale('log')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b5069a478cd65d4fffc49974a5494697326ac3f6"},"cell_type":"code","source":"plt.figure(figsize=(12,8))\nsns.countplot(df['treated_length'])\nplt.title('Length repartiton (rounded down to the 5)')\nplt.yscale('log')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8ad9eb5aba7fb8af734c7210fa5a020163537a3a"},"cell_type":"markdown","source":"Treating text lowers the length of texts, and therefore allows us to make a model with less parameters and a shorter training time."},{"metadata":{"trusted":true,"_uuid":"52fd946fd9993987d0f78a4bb91ab12696305864"},"cell_type":"code","source":"max_len = 70\nmax_len_treated = 40","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"87fc62fa686f54ac3531eddfadedb639d15c0ffd"},"cell_type":"markdown","source":"## Tokenizer"},{"metadata":{"trusted":true,"_uuid":"363b9ece53b8e1800d03b188a50ddf84d0cb12a5"},"cell_type":"code","source":"def make_tokenizer(texts, len_voc):\n    from keras.preprocessing.text import Tokenizer\n    t = Tokenizer(num_words=len_voc)\n    t.fit_on_texts(texts)\n    return t","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"68847d69312b393b5b81da8c3409b4ea5841dfdb"},"cell_type":"code","source":"len_voc = 50000","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"017eb151a1d1e927a491879bf068fabd0a7edc4d"},"cell_type":"code","source":"tokenizer = make_tokenizer(df['question_text'], len_voc)\ntokenizer_treated = make_tokenizer(df['treated_text'], len_voc)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fbc447350c0ff1e2a8c736a4a1e9214ee0b53663"},"cell_type":"markdown","source":"## Train/Test split"},{"metadata":{"trusted":true,"_uuid":"372ba361304e05398001f3f8a3a1fa9d19a250d4"},"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\ndf_train, df_test = train_test_split(df, test_size=0.1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"46e8882712c42236e854b6017933ca423381b712"},"cell_type":"markdown","source":"## Data for the Network"},{"metadata":{"trusted":true,"_uuid":"ff437d91e275fb426389691441725f1f6b52f93a"},"cell_type":"code","source":"X_train = tokenizer.texts_to_sequences(df_train['question_text'])\nX_test = tokenizer.texts_to_sequences(df_test['question_text'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e1a800650e7c2f9a58b5aeabd79b1bf3a54e3145"},"cell_type":"code","source":"X_train_treated = tokenizer_treated.texts_to_sequences(df_train['treated_text'])\nX_test_treated = tokenizer_treated.texts_to_sequences(df_test['treated_text'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0e8159172e82cf408207e6f5ffa40f7543318c35"},"cell_type":"code","source":"from keras.preprocessing.sequence import pad_sequences\n\nX_train = pad_sequences(X_train, maxlen=max_len)\nX_test = pad_sequences(X_test, maxlen=max_len)\n\nX_train_treated = pad_sequences(X_train_treated, maxlen=max_len_treated)\nX_test_treated = pad_sequences(X_test_treated, maxlen=max_len_treated)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ba5a1b8109dee2c9fbc628d5da4a7c3447d42fb8"},"cell_type":"code","source":"y_train = df_train['target'].values\ny_test = df_test['target'].values","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"57533e8b1621dd25eff1196e9a183e2ef94a13bd"},"cell_type":"markdown","source":"## Loading pre-trained word vectors"},{"metadata":{"trusted":true,"_uuid":"0928686188e81132531669c1f32ab2bc4e50703c"},"cell_type":"code","source":"def get_coefs(word,*arr): \n    return word, np.asarray(arr, dtype='float32')\n\ndef load_embedding(file):\n    if file == '../input/embeddings/wiki-news-300d-1M/wiki-news-300d-1M.vec':\n        embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(file) if len(o)>100)\n    else:\n        embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(file, encoding='latin'))\n    return embeddings_index","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1c8e464ea78e11328dc7ec0d65d5d2a5fd0276a5"},"cell_type":"code","source":"def make_embedding_matrix(embedding, tokenizer, len_voc):\n    all_embs = np.stack(embedding.values())\n    emb_mean,emb_std = all_embs.mean(), all_embs.std()\n    embed_size = all_embs.shape[1]\n    word_index = tokenizer.word_index\n    embedding_matrix = np.random.normal(emb_mean, emb_std, (len_voc, embed_size))\n    \n    for word, i in word_index.items():\n        if i >= len_voc:\n            continue\n        embedding_vector = embedding.get(word)\n        if embedding_vector is not None: \n            embedding_matrix[i] = embedding_vector\n    \n    return embedding_matrix","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6f8027d0909333f521cbe2caa9ecd3719789161f"},"cell_type":"code","source":"glove = load_embedding('../input/embeddings/glove.840B.300d/glove.840B.300d.txt')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"997e2d47f50f8b358c31427c1d446de363a5f021"},"cell_type":"code","source":"embed_mat = make_embedding_matrix(glove, tokenizer, len_voc)\nembed_mat_treated = make_embedding_matrix(glove, tokenizer_treated, len_voc)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"81969da3e097faf75122672f3a0cce7b386b3213"},"cell_type":"markdown","source":" ## Attention Layer\n> Code from Khoi Ngyuen, check here : https://www.kaggle.com/suicaokhoailang/lstm-attention-baseline-0-652-lb"},{"metadata":{"trusted":true,"_uuid":"9b08dc6250893f59713abdb2359d470cd031797f"},"cell_type":"code","source":"class Attention(Layer):\n    def __init__(self, step_dim, W_regularizer=None, b_regularizer=None, W_constraint=None, b_constraint=None, bias=True, **kwargs):\n        self.supports_masking = True\n        self.init = initializers.get('glorot_uniform')\n        self.W_regularizer = regularizers.get(W_regularizer)\n        self.b_regularizer = regularizers.get(b_regularizer)\n        self.W_constraint = constraints.get(W_constraint)\n        self.b_constraint = constraints.get(b_constraint)\n        self.bias = bias\n        self.step_dim = step_dim\n        self.features_dim = 0\n        super(Attention, self).__init__(**kwargs)\n        \n    def build(self, input_shape):\n        assert len(input_shape) == 3\n        self.W = self.add_weight((input_shape[-1],), initializer=self.init, name='{}_W'.format(self.name), regularizer=self.W_regularizer, constraint=self.W_constraint)\n        self.features_dim = input_shape[-1]\n        if self.bias:\n            self.b = self.add_weight((input_shape[1],), initializer='zero', name='{}_b'.format(self.name), regularizer=self.b_regularizer, constraint=self.b_constraint)\n        else:\n            self.b = None\n        self.built = True\n\n    def compute_mask(self, input, input_mask=None):\n        return None\n\n    def call(self, x, mask=None):\n        features_dim = self.features_dim\n        step_dim = self.step_dim\n        eij = K.reshape(K.dot(K.reshape(x, (-1, features_dim)), K.reshape(self.W, (features_dim, 1))), (-1, step_dim))\n        if self.bias: eij += self.b\n        eij = K.tanh(eij)\n        a = K.exp(eij)\n        if mask is not None: a *= K.cast(mask, K.floatx())\n        a /= K.cast(K.sum(a, axis=1, keepdims=True) + K.epsilon(), K.floatx())\n        a = K.expand_dims(a)\n        weighted_input = x * a\n        return K.sum(weighted_input, axis=1)\n\n    def compute_output_shape(self, input_shape):\n        return input_shape[0],  self.features_dim","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2a5f324273d8e4726a6f0f9206170845d5ead890"},"cell_type":"markdown","source":"## Making model"},{"metadata":{"trusted":true,"_uuid":"4584596a85f8380a0c25eea1e6364f672b114e25"},"cell_type":"code","source":"def make_model(embedding_matrix, max_len, len_voc=50000, embed_size=300):\n    inp = Input(shape=(max_len,))\n    x = Embedding(len_voc, embed_size, weights=[embedding_matrix], trainable=False)(inp)\n    x = Bidirectional(CuDNNGRU(64, return_sequences=True))(x)\n    x = Attention(max_len)(x)\n    x = Dense(1, activation=\"sigmoid\")(x)\n    model = Model(inputs=inp, outputs=x)\n    model.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])\n    return model","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"54d59f934f48046cca805f2d7684dba540e7eaa1"},"cell_type":"code","source":"model = make_model(embed_mat, max_len)\nmodel_treated = make_model(embed_mat_treated, max_len_treated)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"032df52e96dda468647fc10a138c3005127a01b6"},"cell_type":"code","source":"model.summary()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"80580a6745b44b8b63ab725672c8ffafbab57e75"},"cell_type":"markdown","source":"### Fitting"},{"metadata":{"trusted":true,"_uuid":"8e1e47353ba3171ba1d798ee7c7c3e36577729c8"},"cell_type":"code","source":"model.fit(X_train, y_train, batch_size=1024, epochs=3, validation_data=[X_test, y_test])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"028cf6ef1b94a8ec4c5120d42d1fa6ae861205c8"},"cell_type":"code","source":"model_treated.fit(X_train_treated, y_train, batch_size=1024, epochs=3, validation_data=[X_test_treated, y_test])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2aebfe6e5889091b2d97f446373f5a1e2d4f015e"},"cell_type":"markdown","source":"### Predictions"},{"metadata":{"trusted":true,"_uuid":"35bd2078e02117d7fca467cc7068be4e3ae41822"},"cell_type":"code","source":"pred_train = model.predict([X_train], batch_size=256, verbose=1)\npred_test = model.predict([X_test], batch_size=256, verbose=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"01a9cb883cffa99c9fdacfc0c5bd226b9a5d71d2"},"cell_type":"code","source":"pred_train_treated = model_treated.predict([X_train_treated], batch_size=256, verbose=1)\npred_test_treated = model_treated.predict([X_test_treated], batch_size=256, verbose=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"603d53b9394c43f2afefc6f46e0abee52d8755fd"},"cell_type":"markdown","source":"### Tweaking threshold"},{"metadata":{"trusted":true,"_uuid":"217d8ac6f655b35c7c0c82bffcc76ba50fee51dd"},"cell_type":"code","source":"def tweak_threshold(pred, truth):\n    from sklearn.metrics import f1_score\n    scores = []\n    for thresh in np.arange(0.1, 0.501, 0.01):\n        thresh = np.round(thresh, 2)\n        score = f1_score(truth, (pred>thresh).astype(int))\n        scores.append(score)\n    return round(np.max(scores), 4)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"77bdaceb60a5840c1561d8b9560ad7b29316f9b3"},"cell_type":"code","source":"print(f\"Scored {tweak_threshold(pred_train, y_train)} without text treatment on train data\")\nprint(f\"Scored {tweak_threshold(pred_test, y_test)} without text treatment on test data\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3347d752a8b671019b448d092234bd786fa06bfb"},"cell_type":"code","source":"print(f\"Scored {tweak_threshold(pred_train_treated, y_train)} with text treatment on train data\")\nprint(f\"Scored {tweak_threshold(pred_test_treated, y_test)} with text treatment on test data\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"50698a4b9b2cf21d6370d8e88d90d265fb3af1fc"},"cell_type":"markdown","source":"** The model without treatment appears to significantly outperform the other one **\n\n  I do believe that it is because GloVe Word Vectors are able to capture the information better when the words are not processed.\n  \n  In fact, word vectors are able to deal with number, most of special characters, words starting with an upper case letter (etc ...)\n  \n  Moreover, Keras' Tokenizer does the most basic preprocessing steps (lowercasing / punctuation removal)\n   \n## Conclusion : Text cleaning is a waste of time\n### Well,  at least, the steps I chose are ...\nOne should focus on optimizing the model instead, and perhaps doing some data augmentation.\n "}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}