{"cells":[{"metadata":{"_uuid":"6f5ea96dfcc682c232d6cbb6cebd8b48ecabf8fe"},"cell_type":"markdown","source":"### Preface\n\nHello . This is basically cutting and pasting from the amazing kernels of this competition. Please notify me if I don't attribute something correctly.\n\n* https://www.kaggle.com/gmhost/gru-capsule\n* How to: Preprocessing when using embeddings\nhttps://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings\n* Improve your Score with some Text Preprocessing https://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing\n* Simple attention layer taken from https://github.com/mttk/rnn-classifier/blob/master/model.py\n* https://www.kaggle.com/ziliwang/baseline-pytorch-bilstm\n* https://www.kaggle.com/hengzheng/pytorch-starter\n\n**UPDATE**: I seems that the shuffling the data doesn't add the features in the correct order. To address this issue I added a custom dataset class that can return indexes so that they can be accessed while training and properly put each feature with the corresponding sample. The training time though is increased, so you might need to make the model lighter in order to submit results."},{"metadata":{"_uuid":"710ed17d0c57bd287be0ee3b2782a53a54510561"},"cell_type":"markdown","source":"## IMPORTS "},{"metadata":{"_uuid":"abb7e3c30b8a412a50c6b451c49939e3cf4bc11b","scrolled":true,"trusted":true},"cell_type":"code","source":"import time\nimport random\nimport pandas as pd\nimport numpy as np\nimport gc\nimport re\nimport torch\nfrom torchtext import data\nimport spacy\nfrom tqdm import tqdm_notebook, tnrange\nfrom tqdm.auto import tqdm\n\ntqdm.pandas(desc='Progress')\nfrom collections import Counter\nfrom textblob import TextBlob\nfrom nltk import word_tokenize\n\nimport lightgbm as lgb\nimport xgboost as xgb\n\nimport torch.nn as nn\nimport torch.optim as optim\nimport torch.nn.functional as F\nfrom torch.utils.data import Dataset, DataLoader\nfrom torch.nn.utils.rnn import pack_padded_sequence, pad_packed_sequence\nfrom torch.autograd import Variable\nfrom torchtext.data import Example\nfrom sklearn.metrics import f1_score\nimport torchtext\nimport os \n\nfrom keras.preprocessing.text import Tokenizer\nfrom keras.preprocessing.sequence import pad_sequences\n\n# cross validation and metrics\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.metrics import f1_score\nfrom sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer\nfrom sklearn.model_selection import train_test_split\n# from torch.optim.optimizer import Optimizer\nfrom unidecode import unidecode","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9a4ff5590a6f152dc1bec5aeca79aef10218f7de"},"cell_type":"markdown","source":"### Basic Parameters"},{"metadata":{"_uuid":"deee49df5ca1c4413f71677939e26aa1ff784e44","scrolled":true,"trusted":true},"cell_type":"code","source":"embed_size = 400 # how big is each word vector\nmax_features = 120000 # how many unique words to use (i.e num rows in embedding vector)\nmaxlen = 70 # max number of words in a question to use\nbatch_size = 512 # how many samples to process at once\nn_epochs = 5 # how many times to iterate over all samples\nn_splits = 10 # Number of K-fold Splits\n\nSEED = 229","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"654cbe3c8a1f2a618a2441afe00df3b4a89e0a58"},"cell_type":"markdown","source":"### Ensure determinism in the results\n\nA common headache in this competition is the lack of determinism in the results due to cudnn. The following Kernel has a solution in Pytorch.\n\nSee https://www.kaggle.com/hengzheng/pytorch-starter. "},{"metadata":{"_uuid":"58bbf87335799247586aaed16531f4d28d10ed4a","scrolled":true,"trusted":true},"cell_type":"code","source":"def seed_everything(seed=SEED):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True\nseed_everything()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c890692644acce2dc4f6e2f929d6d294faca4ad2"},"cell_type":"markdown","source":"### Code for Loading Embeddings\n\nFunctions taken from the kernel:https://www.kaggle.com/gmhost/gru-capsule\n"},{"metadata":{"_uuid":"7026ee1d913f54dd4b560f654efdb9f833581cd3","scrolled":true,"trusted":true},"cell_type":"code","source":"## FUNCTIONS TAKEN FROM https://www.kaggle.com/gmhost/gru-capsule\n'''\ndef load_embeddings(word_index, EMBEDDING_FILE):\n    def get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\n    embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE))[:400]\n    \n    all_embs = np.stack(embeddings_index.values())\n    emb_mean,emb_std = all_embs.mean(), all_embs.std()\n    print('File {}: all_embs.mean = {}, all_embs.std = {}'.format(EMBEDDING_FILE, emb_mean, emb_std))\n    embed_size = all_embs.shape[1]\n\n    # word_index = tokenizer.word_index\n    nb_words = min(max_features, len(word_index))\n    # Why random embedding for OOV? what if use mean?\n    embedding_matrix = np.random.normal(emb_mean, emb_std, (nb_words, embed_size))\n    #embedding_matrix = np.random.normal(emb_mean, 0, (nb_words, embed_size)) # std 0\n    for word, i in word_index.items():\n        if i >= max_features: continue\n        embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: embedding_matrix[i] = embedding_vector\n            \n    return embedding_matrix \n\n\nload_glove = lambda word_index: load_embeddings(word_index, EMBEDDING_FILE = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt')\nload_fasttext = lambda word_index: load_embeddings(word_index, EMBEDDING_FILE = '../input/embeddings/wiki-news-300d-1M/wiki-news-300d-1M.vec')\nload_para = lambda word_index: load_embeddings(word_index, EMBEDDING_FILE = '../input/embeddings/paragram_300_sl999/paragram_300_sl999.txt')\n'''","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e4ec7ef83a30425fded108e6b1b99323652e4380"},"cell_type":"code","source":"def load_glove(word_index):\n    EMBEDDING_FILE = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt'\n    def get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\n    embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE) if o.split(\" \")[0] in word_index)\n\n    all_embs = np.stack(embeddings_index.values())\n    emb_mean,emb_std = all_embs.mean(), all_embs.std()\n    embed_size = all_embs.shape[1]\n\n    embedding_matrix = np.random.normal(emb_mean, emb_std, (max_features, embed_size))\n    for word, i in word_index.items():\n        if i >= max_features: continue\n        embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: embedding_matrix[i] = embedding_vector\n            \n    return embedding_matrix \n    \ndef load_fasttext(word_index):    \n    EMBEDDING_FILE = '../input/embeddings/wiki-news-300d-1M/wiki-news-300d-1M.vec'\n    def get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\n    embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE) if len(o)>100 and o.split(\" \")[0] in word_index )\n\n    all_embs = np.stack(embeddings_index.values())\n    emb_mean,emb_std = all_embs.mean(), all_embs.std()\n    embed_size = all_embs.shape[1]\n\n    embedding_matrix = np.random.normal(emb_mean, emb_std, (max_features, embed_size))\n    for word, i in word_index.items():\n        if i >= max_features: continue\n        embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: embedding_matrix[i] = embedding_vector\n\n    return embedding_matrix\n\ndef load_para(word_index):\n    EMBEDDING_FILE = '../input/embeddings/paragram_300_sl999/paragram_300_sl999.txt'\n    def get_coefs(word,*arr): return word, np.asarray(arr, dtype='float32')\n    embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(EMBEDDING_FILE, encoding=\"utf8\", errors='ignore') if len(o)>100 and o.split(\" \")[0] in word_index)\n\n    all_embs = np.stack(embeddings_index.values())\n    emb_mean,emb_std = all_embs.mean(), all_embs.std()\n    embed_size = all_embs.shape[1]\n    \n    embedding_matrix = np.random.normal(emb_mean, emb_std, (max_features, embed_size))\n    for word, i in word_index.items():\n        if i >= max_features: continue\n        embedding_vector = embeddings_index.get(word)\n        if embedding_vector is not None: embedding_matrix[i] = embedding_vector\n    \n    return embedding_matrix","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ea10c8e218a1280faa9802bcb7f1117c89ec96f9"},"cell_type":"markdown","source":"## LOAD PROCESSED TRAINING DATA FROM DISK"},{"metadata":{"_uuid":"173753f0178464d2ba26baf22899884d76d1c83d","scrolled":true,"trusted":true},"cell_type":"code","source":"df_train = pd.read_csv(\"../input/train.csv\")\ndf_test = pd.read_csv(\"../input/test.csv\")\ndf = pd.concat([df_train ,df_test],sort=True)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cd5693800c98f697743f3f28ca81435461718f8f"},"cell_type":"markdown","source":"TF-IDF features: "},{"metadata":{"trusted":true,"_uuid":"363575b1f57afad25fa39f3677caa7b0dce9239e"},"cell_type":"code","source":"q = df[\"question_text\"].fillna(\"na_\").values\n\nchar_vector = TfidfVectorizer(\n    ngram_range=(2,4),\n    max_features=20000,\n    stop_words='english',\n    analyzer='char_wb',\n    token_pattern=r'\\w{1,}',\n    strip_accents='unicode',\n    sublinear_tf=True, \n    max_df=0.98,\n    min_df=2\n)\nchar_vector.fit(q[:85000])\n\nchar_vector = char_vector.transform(q).tocsr()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3a5d03ade4c2cc41f6f6abbad8f7dd45360b31ff"},"cell_type":"code","source":"word_vector = TfidfVectorizer(\n    ngram_range=(1,1), \n    max_features=9000,\n    sublinear_tf=True, \n    strip_accents='unicode', \n    analyzer='word', \n    token_pattern=\"\\w{1,}\", \n    stop_words=\"english\",\n    max_df=0.95,\n    min_df=2\n)\nword_vector.fit(q)\n\nword_vector = word_vector.transform(q).tocsr()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ce767562ab6c74de588af35d99e7061b61f45880"},"cell_type":"code","source":"del q\ngc.collect()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0f75559b6fa28c27ecbf145121309378afc884ee","scrolled":true,"trusted":true},"cell_type":"code","source":"def build_vocab(texts):\n    sentences = texts.apply(lambda x: x.split()).values\n    vocab = {}\n    for sentence in sentences:\n        for word in sentence:\n            try:\n                vocab[word] += 1\n            except KeyError:\n                vocab[word] = 1\n    return vocab\nvocab = build_vocab(df['question_text'])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"559bbbc5d30b33443396a8b817de74aa64930431"},"cell_type":"code","source":"list(vocab.keys())[:10]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5cb425ffbf1f79c1edc4cad3da15a1c1aa53edca","scrolled":true,"trusted":true},"cell_type":"code","source":"sin = len(df_train[df_train[\"target\"]==0])\ninsin = len(df_train[df_train[\"target\"]==1])\npersin = (sin/(sin+insin))*100\nperinsin = (insin/(sin+insin))*100            \nprint(\"# Sincere questions: {:,}({:.2f}%) and # Insincere questions: {:,}({:.2f}%)\".format(sin,persin,insin,perinsin))\n# print(\"Sinsere:{}% Insincere: {}%\".format(round(persin,2),round(perinsin,2)))\nprint(\"# Test samples: {:,}({:.3f}% of train samples)\".format(len(df_test),len(df_test)/len(df_train)))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"07e9890ec0b490cef57565f7dff953aa56ebd3dc"},"cell_type":"markdown","source":"## Normalization\n\nBorrowed from:\n* How to: Preprocessing when using embeddings\nhttps://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings\n* Improve your Score with some Text Preprocessing https://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing"},{"metadata":{"trusted":true,"_uuid":"762e4b376da83113038772902bb1b79bb0622573"},"cell_type":"code","source":"def get_wired(d):\n    wired = []\n    for v in d:\n        m = re.match('[\\w{}]+', v)\n        if not m or m[0] != v:\n            wired += [v]\n    return wired\n\nwired = get_wired(vocab)\nlen(wired)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"672a32a5e949bdfa44417a66f2030df758fdef74"},"cell_type":"code","source":"contraction_mapping = {\"ain't\": \"is not\", \"aren't\": \"are not\",\"can't\": \"cannot\", \"'cause\": \"because\", \"could've\": \"could have\", \"couldn't\": \"could not\", \"didn't\": \"did not\",  \"doesn't\": \"does not\", \"don't\": \"do not\", \"hadn't\": \"had not\", \"hasn't\": \"has not\", \"haven't\": \"have not\", \"he'd\": \"he would\",\"he'll\": \"he will\", \"he's\": \"he is\", \"how'd\": \"how did\", \"how'd'y\": \"how do you\", \"how'll\": \"how will\", \"how's\": \"how is\",  \"I'd\": \"I would\", \"I'd've\": \"I would have\", \"I'll\": \"I will\", \"I'll've\": \"I will have\",\"I'm\": \"I am\", \"I've\": \"I have\", \"i'd\": \"i would\", \"i'd've\": \"i would have\", \"i'll\": \"i will\",  \"i'll've\": \"i will have\",\"i'm\": \"i am\", \"i've\": \"i have\", \"isn't\": \"is not\", \"it'd\": \"it would\", \"it'd've\": \"it would have\", \"it'll\": \"it will\", \"it'll've\": \"it will have\",\"it's\": \"it is\", \"let's\": \"let us\", \"ma'am\": \"madam\", \"mayn't\": \"may not\", \"might've\": \"might have\",\"mightn't\": \"might not\",\"mightn't've\": \"might not have\", \"must've\": \"must have\", \"mustn't\": \"must not\", \"mustn't've\": \"must not have\", \"needn't\": \"need not\", \"needn't've\": \"need not have\",\"o'clock\": \"of the clock\", \"oughtn't\": \"ought not\", \"oughtn't've\": \"ought not have\", \"shan't\": \"shall not\", \"sha'n't\": \"shall not\", \"shan't've\": \"shall not have\", \"she'd\": \"she would\", \"she'd've\": \"she would have\", \"she'll\": \"she will\", \"she'll've\": \"she will have\", \"she's\": \"she is\", \"should've\": \"should have\", \"shouldn't\": \"should not\", \"shouldn't've\": \"should not have\", \"so've\": \"so have\",\"so's\": \"so as\", \"this's\": \"this is\",\"that'd\": \"that would\", \"that'd've\": \"that would have\", \"that's\": \"that is\", \"there'd\": \"there would\", \"there'd've\": \"there would have\", \"there's\": \"there is\", \"here's\": \"here is\",\"they'd\": \"they would\", \"they'd've\": \"they would have\", \"they'll\": \"they will\", \"they'll've\": \"they will have\", \"they're\": \"they are\", \"they've\": \"they have\", \"to've\": \"to have\", \"wasn't\": \"was not\", \"we'd\": \"we would\", \"we'd've\": \"we would have\", \"we'll\": \"we will\", \"we'll've\": \"we will have\", \"we're\": \"we are\", \"we've\": \"we have\", \"weren't\": \"were not\", \"what'll\": \"what will\", \"what'll've\": \"what will have\", \"what're\": \"what are\",  \"what's\": \"what is\", \"what've\": \"what have\", \"when's\": \"when is\", \"when've\": \"when have\", \"where'd\": \"where did\", \"where's\": \"where is\", \"where've\": \"where have\", \"who'll\": \"who will\", \"who'll've\": \"who will have\", \"who's\": \"who is\", \"who've\": \"who have\", \"why's\": \"why is\", \"why've\": \"why have\", \"will've\": \"will have\", \"won't\": \"will not\", \"won't've\": \"will not have\", \"would've\": \"would have\", \"wouldn't\": \"would not\", \"wouldn't've\": \"would not have\", \"y'all\": \"you all\", \"y'all'd\": \"you all would\",\"y'all'd've\": \"you all would have\",\"y'all're\": \"you all are\",\"y'all've\": \"you all have\",\"you'd\": \"you would\", \"you'd've\": \"you would have\", \"you'll\": \"you will\", \"you'll've\": \"you will have\", \"you're\": \"you are\", \"you've\": \"you have\" }\npuncts = [',', '.', '\"', ':', ')', '(', '-', '!', '?', '|', ';', \"'\", '$', '&', '/', '[', ']', '>', '%', '=', '#', '*', '+', '\\\\', '•',  '~', '@', '£', '·', '_', '{', '}', '©', '^', '®', '`',  '<', '→', '°', '€', '™', '›',  '♥', '←', '×', '§', '″', '′', 'Â', '█', '½', 'à', '…', '“', '★', '”', '–', '●', 'â', '►', '−', '¢', '²', '¬', '░', '¶', '↑', '±', '¿', '▾', '═', '¦', '║', '―', '¥', '▓', '—', '‹', '─', '▒', '：', '¼', '⊕', '▼', '▪', '†', '■', '’', '▀', '¨', '▄', '♫', '☆', 'é', '¯', '♦', '¤', '▲', 'è', '¸', '¾', 'Ã', '⋅', '‘', '∞', '∙', '）', '↓', '、', '│', '（', '»', '，', '♪', '╩', '╚', '³', '・', '╦', '╣', '╔', '╗', '▬', '❤', 'ï', 'Ø', '¹', '≤', '‡', '√', ]\nmispell_dict = {'colour': 'color', 'centre': 'center', 'favourite': 'favorite', 'travelling': 'traveling', 'counselling': 'counseling', 'theatre': 'theater', 'cancelled': 'canceled', 'labour': 'labor', 'organisation': 'organization', 'wwii': 'world war 2', 'citicise': 'criticize', 'youtu ': 'youtube ', 'Qoura': 'Quora', 'sallary': 'salary', 'Whta': 'What', 'narcisist': 'narcissist', 'howdo': 'how do', 'whatare': 'what are', 'howcan': 'how can', 'howmuch': 'how much', 'howmany': 'how many', 'whydo': 'why do', 'doI': 'do I', 'theBest': 'the best', 'howdoes': 'how does', 'mastrubation': 'masturbation', 'mastrubate': 'masturbate', \"mastrubating\": 'masturbating', 'pennis': 'penis', 'Etherium': 'Ethereum', 'narcissit': 'narcissist', 'bigdata': 'big data', '2k17': '2017', '2k18': '2018', 'qouta': 'quota', 'exboyfriend': 'ex boyfriend', 'airhostess': 'air hostess', \"whst\": 'what', 'watsapp': 'whatsapp', 'demonitisation': 'demonetization', 'demonitization': 'demonetization', 'demonetisation': 'demonetization'}","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"28ebb28ba78972bb8d4fee9b53437045542d20fb","scrolled":true,"trusted":true},"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler\n\ndef known_contractions(embed):\n    known = []\n    for contract in contraction_mapping:\n        if contract in embed:\n            known.append(contract)\n    return known\n\n\ndef clean_contractions(text, mapping):\n    specials = [\"’\", \"‘\", \"´\", \"`\"]\n    for s in specials:\n        text = text.replace(s, \"'\")\n    text = ' '.join([mapping[t] if t in mapping else t for t in text.split(\" \")])\n    return text\n\n\ndef correct_spelling(x, dic):\n    for word in dic.keys():\n        x = x.replace(word, dic[word])\n    return x\n\n\ndef unknown_punct(embed, punct):\n    unknown = ''\n    for p in punct:\n        if p not in embed:\n            unknown += p\n            unknown += ' '\n    return unknown\n\n\ndef clean_numbers(x):\n    x = re.sub('[0-9]{5,}', '#####', x)\n    x = re.sub('[0-9]{4}', '####', x)\n    x = re.sub('[0-9]{3}', '###', x)\n    x = re.sub('[0-9]{2}', '##', x)\n    return x\n\n\ndef clean_special_chars(text, punct, mapping):\n    for p in mapping:\n        text = text.replace(p, mapping[p])\n    \n    for p in punct:\n        text = text.replace(p, f' {p} ')\n    \n    specials = {'\\u200b': ' ', '…': ' ... ', '\\ufeff': '', 'करना': '', 'है': ''}  # Other special characters that I have to deal with in last\n    for s in specials:\n        text = text.replace(s, specials[s])\n    \n    return text\n\n\ndef add_lower(embedding, vocab):\n    count = 0\n    for word in vocab:\n        if word in embedding and word.lower() not in embedding:  \n            embedding[word.lower()] = embedding[word]\n            count += 1\n    print(f\"Added {count} words to embedding\")\n\n\ndef clean_text(x):\n    x = str(x)\n    for punct in puncts:\n        x = x.replace(punct, f' {punct} ')\n    return x\n\n\ndef clean_numbers(x):\n    x = re.sub('[0-9]{5,}', '#####', x)\n    x = re.sub('[0-9]{4}', '####', x)\n    x = re.sub('[0-9]{3}', '###', x)\n    x = re.sub('[0-9]{2}', '##', x)\n    return x\n\n\ndef _get_mispell(mispell_dict):\n    mispell_re = re.compile('(%s)' % '|'.join(mispell_dict.keys()))\n    return mispell_dict, mispell_re\n\n\nmispellings, mispellings_re = _get_mispell(mispell_dict)\ndef replace_typical_misspell(text):\n    def replace(match):\n        return mispellings[match.group(0)]\n    return mispellings_re.sub(replace, text)\n\n\n# Extra feature part taken from https://github.com/wongchunghang/toxic-comment-challenge-lstm/blob/master/toxic_comment_9872_model.ipynb\ndef add_features(df):\n    df['question_text'] = df['question_text'].progress_apply(lambda x:str(x))\n    df['total_length'] = df['question_text'].progress_apply(len)\n    df['capitals'] = df['question_text'].progress_apply(lambda comment: sum(1 for c in comment if c.isupper()))\n    df['caps_vs_length'] = df.progress_apply(lambda row: float(row['capitals'])/float(row['total_length']), axis=1)\n    df['num_words'] = df.question_text.str.count('\\S+')\n    df['num_unique_words'] = df['question_text'].progress_apply(lambda comment: len(set(w for w in comment.split())))\n    df['words_vs_unique'] = df['num_unique_words'] / df['num_words']\n    df['char_vector'] = char_vector\n    df['word_vector'] = word_vector\n    \n    df['question_text'] = df['question_text'].progress_apply(lambda x: x.lower())\n    df['question_text'] = df['question_text'].progress_apply(clean_text)\n    df['question_text'] = df['question_text'].progress_apply(clean_numbers)\n    df['question_text'] = df['question_text'].progress_apply(replace_typical_misspell)\n    return df","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"63cb21525251b060aeb309e7be4b48772f8720f5","scrolled":true,"trusted":true},"cell_type":"code","source":" ## fill up the missing values\nX = df[\"question_text\"].fillna(\"_##_\").values\n\n###################### Add Features ###############################\n#  https://github.com/wongchunghang/toxic-comment-challenge-lstm/blob/master/toxic_comment_9872_model.ipynb\ndf = add_features(df)\n\nfeatures = df[['caps_vs_length', 'words_vs_unique']].fillna(0)\n\nss = StandardScaler()\nss.fit(features)\nfeatures = ss.transform(features)\n###########################################################################\n\n## Tokenize the sentences\ntokenizer = Tokenizer(num_words=max_features)\ntokenizer.fit_on_texts(list(X))\nX = tokenizer.texts_to_sequences(X)\n\n## Pad the sentences \nX = pad_sequences(X, maxlen=maxlen)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ec318b64da43e88946717d5115dc3cebfb082b13"},"cell_type":"code","source":"X_train, X_test = X[:len(df_train)], X[len(df_train):]\nfeatures_train, features_test = features[:len(df_train)], features[len(df_train):]\ndf_train, df_test = df.iloc[:len(df_train)], df.iloc[len(df_train):]\n\n## Get the target values\ny_train = df_train['target'].values\n\n#     # Splitting to training and a final test set    \n#     train_X, x_test_f, train_y, y_test_f = train_test_split(list(zip(train_X,features)), train_y, test_size=0.2, random_state=SEED)    \n#     train_X, features = zip(*train_X)\n#     x_test_f, features_t = zip(*x_test_f)    \n\n#shuffling the data\nnp.random.seed(SEED)\nsample = np.random.permutation(len(X_train))\n\nX_train = X_train[sample]\ny_train = y_train[sample]\nfeatures_train = features_train[sample]\n\n#     return train_X, test_X, train_y, features, test_features, tokenizer.word_index\n#     return train_X, test_X, train_y, x_test_f,y_test_f,features, test_features, features_t, tokenizer.word_index\n#     return train_X, test_X, train_y, tokenizer.word_index","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"5264b1a511613e0cb9cffc11e93e906f932be191"},"cell_type":"markdown","source":"### SAVE DATASET TO DISK"},{"metadata":{"_uuid":"7dc2704001695c0d691fec74cc111308da3fc340","scrolled":true,"trusted":true},"cell_type":"code","source":"np.save(\"x_train\", X_train)\nnp.save(\"x_test\", X_test)\nnp.save(\"y_train\", y_train)\n\nnp.save(\"features\", features_train)\nnp.save(\"test_features\",features_test)\nnp.save(\"word_index.npy\",tokenizer.word_index)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ada19e5103b3c5fc5feead877dfd897c1445b969"},"cell_type":"code","source":"del X_train\ndel X_test\ndel y_train\ndel features_train\ndel features_test\ndel tokenizer.word_index\ngc.collect()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ef622350c3ac00bcbea516ccf7dffbeca7b4cc39"},"cell_type":"markdown","source":"### LOAD DATASET FROM DISK"},{"metadata":{"_uuid":"6e1f33f0d744e86cfbcc21b8e696bb568ba1c454","scrolled":true,"trusted":true},"cell_type":"code","source":"X_train = np.load(\"x_train.npy\")\nX_test = np.load(\"x_test.npy\")\ny_train = np.load(\"y_train.npy\")\nfeatures_train = np.load(\"features.npy\")\nfeatures_test = np.load(\"test_features.npy\")\nword_index = np.load(\"word_index.npy\").item()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6fe70227e9d6faf70741abe3326d4c3bb544c87d"},"cell_type":"code","source":"input_train, input_valid, y_train, y_val = train_test_split(X_train, y_train, train_size=.9, random_state=SEED)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6db7b1d3f22a6e35699d9f0dbb51343e249d9ee7","scrolled":true,"trusted":true},"cell_type":"code","source":"features.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e5c51a8329d569d13b9f0369ebb98ca8e2e55440"},"cell_type":"markdown","source":"### Load Embeddings\n\nTwo embedding matrices have been used. Glove, and paragram. The mean of the two is used as the final embedding matrix"},{"metadata":{"_uuid":"6a5f4502324d369ff6faa3692accee4f8a233005","scrolled":true,"trusted":true},"cell_type":"code","source":"# missing entries in the embedding are set using np.random.normal so we have to seed here too\nseed_everything()\n\nglove_embeddings = load_glove(word_index)\nparagram_embeddings = load_para(word_index)\nfasttext_embeddings = load_fasttext(word_index)\n\nembedding_matrix = np.mean([glove_embeddings, paragram_embeddings, fasttext_embeddings], axis=0)\n\n# vocab = build_vocab(df['question_text'])\n# add_lower(embedding_matrix, vocab)\ndel glove_embeddings, paragram_embeddings, fasttext_embeddings\ngc.collect()\n\nnp.shape(embedding_matrix)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d3b707858357eec18403f97368a40c43987b40e8"},"cell_type":"markdown","source":"### Use Stratified K Fold to improve results"},{"metadata":{"_uuid":"fa55c64890964220ded1762236e091e59e502e53","scrolled":true,"trusted":true},"cell_type":"code","source":"# splits = list(StratifiedKFold(n_splits=n_splits, shuffle=True, random_state=SEED).split(X_train, y_train))\n# splits[:3]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"da3299db24754d71851d7f634ee3b96532323e88"},"cell_type":"markdown","source":"## Model and training\n\nHere we use `LGBT` instead of `BiLSTM` just because it is faster. \n\nDocuments: \n\n- [Document](https://gormanalysis.com/gradient-boosting-explained/)\n- [Kernel](https://www.kaggle.com/sjb1988/lgb-python-basic-features-only)\n- [Another kernel (copied here)](https://www.kaggle.com/jaguar00/xgboost-baseline)"},{"metadata":{"trusted":true,"_uuid":"70e58226e0522aec0901c4add13d93cbe18bf296"},"cell_type":"code","source":"'''\ndef train_model(x, y, lgb_params, \n                number_of_folds=10, \n                evaluation_metric='auc', \n                save_feature_importances=False, \n                early_stopping_rounds=50, \n                num_round = 50,\n                identifier_columns=['MachineIdentifier'],\n                single_fold=False):\n    cross_validator = StratifiedKFold(n_splits=number_of_folds,\n                                  random_state=random_state,\n                                  shuffle=shuffle)\n    \n    validation_scores = []\n    classifier_models = []\n    feature_importance_df = pd.DataFrame()\n    for fold_index, (train_index, validation_index) in enumerate(cross_validator.split(x, y)):\n        x_train, x_validation = x.iloc[train_index], x.iloc[validation_index]\n        y_train, y_validation = y.iloc[train_index], y.iloc[validation_index]\n    \n        x_train.drop(identifier_columns, axis=1, inplace=True)\n        validation_identifier_data = x_validation[identifier_columns]\n        x_validation.drop(identifier_columns, axis=1, inplace=True)\n        x_train_columns = x_train.columns\n        trn_data = lgb.Dataset(x_train,\n                       label=y_train,\n                       # categorical_feature=categorical_columns\n                       )\n        del x_train\n        del y_train\n        val_data = lgb.Dataset(x_validation,\n                               label=y_validation,\n                               # categorical_feature=categorical_columns\n                               )\n        classifier_model = lgb.train(lgb_params,\n                                     trn_data,\n                                      num_round,\n                                     valid_sets=[trn_data, val_data],\n                                     verbose_eval=100,\n                                     early_stopping_rounds=early_stopping_rounds\n                                     )\n\n        classifier_models.append(classifier_model)\n        \n        predictions = classifier_model.predict(x_validation, num_iteration=classifier_model.best_iteration)\n        false_positive_rate, recall, thresholds = metrics.roc_curve(y_validation, predictions)\n        score = metrics.auc(false_positive_rate, recall)\n        validation_scores.append(score)\n        \n        fold_importance_df = pd.DataFrame()\n        fold_importance_df[\"feature\"] = x_train_columns\n        fold_importance_df[\"importance\"] = classifier_model.feature_importance(importance_type='gain')\n        fold_importance_df[\"fold\"] = fold_index + 1\n        feature_importance_df = pd.concat([feature_importance_df, fold_importance_df], axis=0)\n\n        if single_fold:\n            break\n    if save_feature_importances:\n        cols = (feature_importance_df[[\"feature\", \"importance\"]]\n                .groupby(\"feature\")\n                .mean()\n                .sort_values(by=\"importance\", ascending=False)[:1000].index)\n\n        best_features = feature_importance_df.loc[feature_importance_df.feature.isin(cols)]\n\n        plt.figure(figsize=(14, 25))\n        sns.barplot(x=\"importance\",\n                    y=\"feature\",\n                    data=best_features.sort_values(by=\"importance\",\n                                                   ascending=False))\n        plt.title('LightGBM Features (avg over folds)')\n        plt.tight_layout()\n        plt.savefig('lgbm_importances.png')\n\n        # mean_gain = feature_importances[['gain', 'feature']].groupby('feature').mean()\n        # feature_importances['mean_gain'] = feature_importances['feature'].map(mean_gain['gain'])\n        #\n        # temp = feature_importances.sort_values('mean_gain', ascending=False)\n        best_features.sort_values(by=\"importance\", ascending=False) \\\n            .groupby(\"feature\") \\\n            .mean() \\\n            .sort_values(by=\"importance\", ascending=False) \\\n            .to_csv('feature_importances_new.csv', index=True)\n\n    score = sum(validation_scores) / len(validation_scores)\n    return classifier_models, score\n'''","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bd3d7c19553cba2b8da051877ea81aa7cda5d418"},"cell_type":"code","source":"'''reference: some settings inspired by Toxic competition kernels'''\ndef build_xgb(train_X, train_y, valid_X, valid_y=None, subsample=0.75):\n\n    xgtrain = xgb.DMatrix(train_X, label=train_y)\n    if valid_y is not None:\n        xgvalid = xgb.DMatrix(valid_X, label=valid_y)\n    else:\n        xgvalid = None\n    \n    model_params = {}\n    # binary 0 or 1\n    model_params['objective'] = 'binary:logistic'\n    # eta is the learning_rate, [default=0.3]\n    model_params['eta'] = 0.3\n    # depth of the tree, deeper more complex.\n    model_params['max_depth'] = 3\n    # 0 [default] print running messages, 1 means silent mode\n    model_params['silent'] = 0\n    model_params['eval_metric'] = 'auc'\n    # will give up further partitioning [default=1]\n    model_params['min_child_weight'] = 1\n    # subsample ratio for the training instance\n    model_params['subsample'] = subsample\n    # subsample ratio of columns when constructing each tree\n    model_params['colsample_bytree'] = subsample\n    # random seed\n    model_params['seed'] = SEED\n    # imbalance data ratio\n    #model_params['scale_pos_weight'] = \n    \n    # convert params to list\n    model_params = list(model_params.items())\n    \n    return xgtrain, xgvalid, model_params","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2730c31c58827635ca5033b8a73c5b96bd53a989"},"cell_type":"code","source":"def train_xgboost(xgtrain, xgvalid, model_params, num_rounds=500, patience=20):\n    \n    if xgvalid is not None:\n        # watchlist what information should be printed. specify validation monitoring\n        watchlist = [ (xgtrain, 'train'), (xgvalid, 'test') ]\n        #early_stopping_rounds = stop if performance does not improve for k rounds\n        model = xgb.train(model_params, xgtrain, num_rounds, watchlist, early_stopping_rounds=patience)\n    else:\n        model = xgb.train(model_params, xgtrain, num_rounds)\n    \n    return model","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3a718d3ef342b1b547cbc9919777dcb66f161b61"},"cell_type":"code","source":"# # params from https://www.kaggle.com/fabiendaniel/detecting-malwares-with-lgbm\n# params = {'num_leaves': 128,\n#          'min_data_in_leaf': 42,\n#          'objective': 'binary',\n#          'max_depth': -1,\n#          'learning_rate': 0.05,\n#          \"boosting\": \"gbdt\",\n#          \"feature_fraction\": 0.8,\n#          \"bagging_freq\": 5,\n#          \"bagging_fraction\": 0.8,\n#          \"bagging_seed\": 11,\n#          \"lambda_l1\": 0.15,\n#          \"lambda_l2\": 0.15,\n#          \"random_state\": 42,          \n#          \"verbosity\": -1}\n\nbase_params = {   \n        'boosting_type': 'gbdt',\n        'objective': 'binary',\n        'metric': 'auc',\n        'nthread': 4,\n        'learning_rate': 0.05,\n        'max_depth': 3,\n        'num_leaves': 40,\n        'sub_feature': 0.9,\n        'sub_row':0.9,\n        'bagging_freq': 1,\n        'lambda_l1': 0.1,\n        'lambda_l2': 0.1,\n        'random_state': SEED\n        }","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2266c29ff854c8ad5cb0917d42f4f4162520513c"},"cell_type":"code","source":"%%time\n# models, validation_score = train_model(train.drop('HasDetections', axis=1),\n#                                       train_y, base_params,\n#                                       num_round=5000,\n#                                       single_fold=stop_after_one_fold,\n#                                       save_feature_importances=True)\nprint('train the model')\nxgtrain, xgvalid, model_params = build_xgb(input_train, y_train ,input_valid, y_val)\nmodel = train_xgboost(xgtrain, xgvalid, model_params)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1d2187d4bbf48350eaf365b6dd8b027d5be69e0c"},"cell_type":"markdown","source":"### Find final Thresshold\n\nBorrowed from: https://www.kaggle.com/ziliwang/baseline-pytorch-bilstm"},{"metadata":{"_uuid":"dc4a4c681294ba06526fed4e871cfe8639cf25e4","scrolled":true,"trusted":true},"cell_type":"code","source":"def bestThresshold(y_train,train_preds):\n    tmp = [0,0,0] # idx, cur, max\n    delta = 0\n    for tmp[0] in tqdm(np.arange(0.1, 0.501, 0.01)):\n        tmp[1] = f1_score(y_train, np.array(train_preds)>tmp[0])\n        if tmp[1] > tmp[2]:\n            delta = tmp[0]\n            tmp[2] = tmp[1]\n    print('best threshold is {:.4f} with F1 score: {:.4f}'.format(delta, tmp[2]))\n    return delta\n\ntrain_preds = np.zeros((input_train.shape[0], 1))\ntrain_preds[:,0] = model.predict(xgb.DMatrix(input_train), ntree_limit=model.best_ntree_limit)\ndelta = bestThresshold(y_train,train_preds)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"626232f581b847c0a7e524d10e758f1558a3b50c"},"cell_type":"code","source":"print('predict results')\ntest_preds = np.zeros(( X_test.shape[0], 1 ))\ntest_preds[:,0] = model.predict(xgb.DMatrix(X_test), ntree_limit=model.best_ntree_limit)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d17dd7b0a92ec98134bf8996fd210edaadf7bed6","scrolled":true,"trusted":true},"cell_type":"code","source":"submission = df_test[['qid']].copy()\nsubmission['prediction'] = (test_preds > delta).astype(int)\nsubmission.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a1a07a65e47b5dcefaac040f3f633dc077e8f61e","scrolled":true,"trusted":true},"cell_type":"code","source":"!head submission.csv","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}