{"cells":[{"metadata":{"_uuid":"530c1abff3f072cfcc5b30598b9b38dc56355e81"},"cell_type":"markdown","source":"# Preface    \n\n**Text preprocessing and embeddings** - I'm using some of the methods that are not presented in different notebooks at the point of time.   \nFor example: lower only first letter in sentence, TweetTokenizer, normalise unicode data, few-steps embedding concatenation.\n\n**Model** - as the base I use the Hung The Nguyen PyTorch model - [Pytorch starter](https://www.kaggle.com/hung96ad/pytorch-starter). I adapt it to the concept described by Alexander Burmistrov  [Toxic Comment Classification Challenge 3rd place](https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52644) .\n"},{"metadata":{"_uuid":"a0ee4b882e716e79f45718505b95365e6c96b1a3"},"cell_type":"markdown","source":"### Quick summary:\n**1. Text preprocessing**:\n  \n  * replace words/characters based on dictionary ex. don't -> do not, Brexit -> leave EU (*@Dieter, @Theo Viel*)\n  * change first letter in sentence to lower (size of first letter do matter in pretrained embedding, it's better to lower only the first letter in sentence instead of all of the words in text),   \n  * remove apostrophe and 's ending form word. ex. AI's -> AI (there isn't \"AI's\" vector in pretrained glove but there is \"AI\" vector),\n  * replace digits with mask, ex. 'Marek gets 3333' -> 'Marek gets ####'\"\n  * use TweetTokenizer() from NLTK for splitting words - I think it's the best tokenizer available for unformal text,   \n  * normalise unicode data to remove umlauts, accents etc.\n  * padding=\"pre\", truncating=\"post\"\n    \n**2. Text additional features (apply MinMaxScaler)**:\n  * unique words rate,\n  * rate of all-caps words,\n  * sentence length rate (number of word / max sequence length parameter),     \n  \n** 3. Embeddings:**      \n*ref Alexander Burmistrov*\n  * concatenated fasttext and glove twitter embeddings (I've prepare similar function for average embegings too),\n  * Glove vector is used by itself if there is no Fasttext vector but not the other way around. \n  * If there is no vector for word I'm looking for, transform word to lowercase and look again, \n  * Words without word vectors are replaced with a word vector for a word \"something\".\n  * Added additional value that was set to 1 if a word was written in all capital letters and 0 otherwise,      \n\n** 4. Architecture  -> PyTorch**    \n*ref Alexander Burmistrov, Hung The Nguyen :*\n\n1) Concatenated fasttext and glove twitter embeddings.    \n2) SpatialDropout1D(0.1)   \n3) First Layer of RNN: Bidirectional LSTM with a kernel size 64 -> LSTM    \n4) Attention of LSTM layer -> LSTM_Atten    \n5) Second layer of RNN: Bidirectional GRU with a kernel size 128 -> GRU    \n6) Attention of GRU layer -> GRU_Atten    \n7) A concatenation of the: [LSTM, LSTM_Atten, GRU, GRU_Atten, Text additional features]    \n8)  Dense layers:    Dense(192, relu, Dropout 0,1) -> Dense(64, relu, Dropout 0,1) -> Dense(1, sigmoid).      \n\n**Loss:** Binary Cross Entropy    \n**Optimizer:** Adam,"},{"metadata":{"_uuid":"50fa61e0bf88fa41b06515afe35b3e00c4ec14d6"},"cell_type":"markdown","source":"### References:    \n* [@Dieter](http://https://www.kaggle.com/christofhenkel) -  [How to: Preprocessing when using embeddings](https://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings)\n* [@Theo Viel](https://www.kaggle.com/theoviel) -  [Improve your score with some text preprocessing](https://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing)\n* [@Alexander Burmistrov](https://www.kaggle.com/mrboor) -  [Toxic Comment Classification Challenge 3rd place](https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52644)    \n* [@Hung The Nguyen](https://www.kaggle.com/hung96ad) - [Pytorch starter](https://www.kaggle.com/hung96ad/pytorch-starter)"},{"metadata":{"_uuid":"0825f726aebba039c2b47be6e42f63824cf0275b"},"cell_type":"markdown","source":"# Notebook"},{"metadata":{"trusted":true,"_uuid":"1ddffb5d75f2b9ba8380304cf907262595da2e92"},"cell_type":"code","source":"import warnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6f292b02cf7d70ec87168e4bafc928c373868b9c"},"cell_type":"markdown","source":"### Load data"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\ndef load_data():\n    #load and shuffle training set \n    train_df = pd.read_csv(\"../input/train.csv\")\n    test_df = pd.read_csv(\"../input/test.csv\")\n    \n    print(\"Train shape : \",train_df.shape)\n    print(\"Test shape : \",test_df.shape)\n    list_sentences_train = list(train_df[\"question_text\"].fillna(\"NAN_WORD\").values)\n    list_sentences_test = list(test_df[\"question_text\"].fillna(\"NAN_WORD\").values)\n    \n    return list_sentences_train, list_sentences_test","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"665e24d66093d9532876fbdada0aae077f6b5cb9","scrolled":true},"cell_type":"code","source":"list_sentences_train, list_sentences_test = load_data()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"98d6fe3f4e92ec528b9e0f1202a2745bd8c34672"},"cell_type":"markdown","source":"### What is the percentage of insincere questions in training set?"},{"metadata":{"trusted":true,"_uuid":"2e2c923b15da2dd904c8d3822b475f2656876fd5"},"cell_type":"code","source":"from collections import Counter\ntmp = Counter([i for i in pd.read_csv(\"../input/train.csv\")['target'].tolist()])\nprint(\"The percentage of insincere questions in train set is {:.2%} .\".format(tmp[1]/(tmp[1]+tmp[0])))\n\ntrain_1_prc = tmp[1]/(tmp[1]+tmp[0])\ndel(tmp)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"49590fa5bfb24f72ef34b86fe0b73cf28c4ba75e"},"cell_type":"markdown","source":"### Choose MAX_SEQUENCE_LENGTH"},{"metadata":{"trusted":true,"scrolled":false,"_uuid":"fa49c1bab2c04cc8de2851fba8abdbbac8464558"},"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n%matplotlib inline  \n\nsns.set(rc={'figure.figsize':(12,6)})\nsns.distplot([len(i.split()) for i in list_sentences_train+list_sentences_test])\nplt.title('Distribution of the length of the question_text - all examples')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f418d856d24d4439a78fda51c8003ff745d66ba8"},"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nimport numpy as np\n%matplotlib inline  \n\nsns.set(rc={'figure.figsize':(12,6)})\nsns.distplot(np.array([len(i.split()) for i in list_sentences_train])[np.array([i==1 for i in pd.read_csv(\"../input/train.csv\")['target'].tolist()])])\nplt.title('Distribution of the length of the question_text - only insincere questions')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"950694e29e7d7a419c957064713848d89babb6f8"},"cell_type":"markdown","source":"I think that max length equal 70 words will be good enough."},{"metadata":{"trusted":true,"_uuid":"3bc76b6b9a2f8f94d2138b939cd24a3aa47e6de5"},"cell_type":"code","source":"MAX_SEQUENCE_LENGTH = 70","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"dcdde049de71d1a22c88f31d08c94ee697843270"},"cell_type":"markdown","source":"### Define functions "},{"metadata":{"trusted":true,"_uuid":"d492abded4c02e45ea23e404693c28cf90b4068d"},"cell_type":"code","source":"def change_word(text: str,dict_:dict) -> str:\n    \"\"\"\n    Function that replace words base on dictionary.\n    \n    Parameters\n    ----------\n    text : str\n        Input text\n    dict_: dict\n        Dictionary with pairs {pattern: repl,..} ex. {\"'\":\"'\", \"‘\":\"'}\n        \n    Returns\n    -------\n    text : processed text \n    \n    Examples\n    --------\n    >>> change_word(text=\"sample ? text\", dict_={\"?\":\"question\", \"‘\":\"'\"})\n    'sample question text'\n    \"\"\"\n    for s in dict_.items():\n        text = text.replace(s[0],s[1])\n    return text","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3c599a662314b66a87025f8df2d4d5cf4af6fb15"},"cell_type":"code","source":"import re\ndef lower_first_in_sentence(text: str) -> str:\n    \"\"\"\n    Function that change first letter in sentence to lower. \n    \n    Parameters\n    ----------\n    text : str\n        Input text\n        \n    Returns\n    -------\n    text : processed text \n    \n    Examples\n    --------\n    >>> lower_first_in_sentence(\"Matt is smart - claims John. Yes, I think so. Wait.. he is not.\")\n    'matt is smart - claims John. yes, I think so. wait.. he is not.'\n    \"\"\"\n    spl = re.compile('(\\?+ *|\\!+ *|\\.+ *)')\n    def lower_firts(txt): return txt if (len(txt)<2 or txt[0] ==\"I\") else (txt[0].lower() + txt[1:])\n    return ''.join([lower_firts(i) for i in re.split(spl, text)])","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9aac37b7366028f25f02f76f9f8df9c72cf6d4e3","scrolled":true},"cell_type":"code","source":"import re\ndef strip(word:str ) -> str: \n    \"\"\"\n    Function that removes 's and solo apostrophe ' from end of the word. ex. AI's -> AI \n    \n    Parameters\n    ----------\n    text : str\n        Input text\n        \n    Returns\n    -------\n    text : processed text \n    \n    Examples\n    --------\n    >>> strip(\"exactly AI's VC's Donald'\")\n    'exactly AI VC Donald '\n    \"\"\"\n\n    return re.sub(\"('$ |'$|'s |'s)\",' ',word) if len(word)>2 else word","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c8a1baaa531178c1ad0db4989d99f831b7db0fb3"},"cell_type":"code","source":"def clean_numbers(x):\n    \"\"\"\n    Function that replaces digits\n    \"\"\"\n    x = re.sub('[0-9]{5,}', '#####', x)\n    x = re.sub('[0-9]{4}', '####', x)\n    x = re.sub('[0-9]{3}', '###', x)\n    x = re.sub('[0-9]{2}', '##', x)\n    return x","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b001af4cd2191c63b8a82b4c0170ad686665ab4a"},"cell_type":"code","source":"def load_embed(file):\n    def get_coefs(word,*arr): \n        return word, np.asarray(arr, dtype='float32')\n    \n    if file == '../input/embeddings/wiki-news-300d-1M/wiki-news-300d-1M.vec':\n        embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(file) if len(o)>100)\n    else:\n        embeddings_index = dict(get_coefs(*o.split(\" \")) for o in open(file, encoding='latin'))\n        \n    return embeddings_index","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3f86ef243e3292db7e52d13d64e1e10a6beb560a"},"cell_type":"code","source":"def check_coverage(vocab, embeddings_index):\n    known_words = {}\n    unknown_words = {}\n    nb_known_words = 0\n    nb_unknown_words = 0\n    for word in vocab.keys():\n        try:\n            known_words[word] = embeddings_index[word]\n            nb_known_words += vocab[word]\n        except:\n            unknown_words[word] = vocab[word]\n            nb_unknown_words += vocab[word]\n            pass\n\n    print('Found embeddings for {:.2%} of vocab'.format(len(known_words) / len(vocab)))\n    print('Found embeddings for  {:.2%} of all text'.format(nb_known_words / (nb_known_words + nb_unknown_words)))\n    unknown_words = sorted(unknown_words.items(), key=operator.itemgetter(1))[::-1]\n\n    return unknown_words","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1f1586c3301977c5b7c85a14cfb5741ca5ae6f81"},"cell_type":"code","source":"import numpy as np\nfrom tqdm import tqdm\n\ndef concat_embed(first_embed, second_embed, word_index):\n    \"\"\"\n    Function that concat two embeddings and apply rules described here: \n      * concatenate first embedding and second embedding,\n      * first vector is used by itself if there is no second vector but not the other way around. \n      * If there is no first vector for word you looking for, transform word to lowercase and look again, \n      * Words without word vectors are replaced with a word vector for a word \"something\".\n      * Add additional value that was set to 1 if a word was written in all capital letters and 0 otherwise,      \n    \n    Parameters\n    ----------\n    first_embed : dict\n        First embedding\n\n    second_embed : dict\n        Second embedding\n        \n    word_index : dict\n        Dictionary that contains words you are looking for in a form {'ID':word} ex. {1:'the',2:'ok'...} \n        word_index = {t: i+1 for i,t in enumerate(vocab)}\n        \n    Returns\n    -------\n    wv_matrix : array of vectors (shape: nb_words x WV_DIM)\n    WV_DIM : embedding size\n    nb_words : number of words in embedding\n    \n    Examples:\n    tbc\n    --------\n\n    \"\"\"\n\n    WV_DIM=first_embed['I'].shape[0]+second_embed['I'].shape[0]+1\n    nb_words = len(word_index)+1\n\n    wv_matrix = np.zeros(shape=(nb_words, WV_DIM))\n    for word, i in tqdm(word_index.items()):\n        cap_flag = word.isupper() and word!='I'\n        cap = np.where(cap_flag,np.array([1]),np.array([0]))\n        if word in first_embed:\n            if word in second_embed:\n                #print('1')\n                wv_matrix[i] = np.hstack((second_embed[word],first_embed[word],cap))\n            else:\n                wv_matrix[i] = np.hstack((second_embed['something'],first_embed[word],cap))\n        else:\n            if word.lower() in first_embed:\n                if word.lower() in second_embed:\n                    wv_matrix[i] = np.hstack((second_embed[word.lower()],first_embed[word.lower()],cap))\n                else:\n                    wv_matrix[i] = np.hstack((second_embed['something'],first_embed[word.lower()],cap))\n            else:\n                wv_matrix[i] = np.hstack((second_embed['something'],first_embed['something'],cap))\n            \n    return wv_matrix, WV_DIM, nb_words","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"73ad3c65b845cfdae395a259a1e8ae1423178410"},"cell_type":"code","source":"import numpy as np\nfrom tqdm import tqdm\n\ndef avg_embed(first_embed, second_embed, word_index):\n    \"\"\"\n    Function that average two embeddings and apply rules described here: \n      * concatenate first embedding and second embedding,\n      * first vector is used by itself if there is no second vector but not the other way around. \n      * If there is no first vector for word you looking for, transform word to lowercase and look again, \n      * Words without word vectors are replaced with a word vector for a word \"something\".\n      * Add additional value that was set to 1 if a word was written in all capital letters and 0 otherwise,      \n    \n    Parameters\n    ----------\n    first_embed : dict\n        First embedding\n\n    second_embed : dict\n        Second embedding\n        \n    word_index : dict\n        Dictionary that contains words you are looking for in a form {'ID':word} ex. {1:'the',2:'ok'...} \n        word_index = {t: i+1 for i,t in enumerate(vocab)}\n        \n    Returns\n    -------\n    wv_matrix : array of vectors (shape: nb_words x WV_DIM)\n    WV_DIM : embedding size\n    nb_words : number of words in embedding\n    \n    Examples:\n    tbc\n    --------\n\n\n    \"\"\"\n    WV_DIM=np.mean([second_embed['I'],first_embed['I']], axis = 0).shape[0]+1\n    nb_words = len(word_index)+1\n\n    wv_matrix = np.zeros(shape=(nb_words, WV_DIM))\n    for word, i in tqdm(word_index.items()):\n        cap_flag = word.isupper() and word!='I'\n        cap = np.where(cap_flag,np.array([1]),np.array([0]))\n        if word in first_embed:\n            if word in second_embed:\n                #print('1')\n                wv_matrix[i] = np.hstack((np.mean([second_embed[word],first_embed[word]], axis = 0),cap))\n            else:\n                wv_matrix[i] = np.hstack((np.mean([second_embed['something'],first_embed[word]], axis = 0),cap))\n        else:\n            if word.lower() in first_embed:\n                if word.lower() in second_embed:\n                    wv_matrix[i] = np.hstack((np.mean([second_embed[word.lower()],first_embed[word.lower()]], axis = 0),cap))\n                else:\n                    wv_matrix[i] = np.hstack((np.mean([second_embed['something'],first_embed[word.lower()]], axis = 0),cap))\n            else:\n                wv_matrix[i] = np.hstack((np.mean([second_embed['something'],first_embed['something']], axis = 0),cap))        \n    return wv_matrix, WV_DIM, nb_words","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"38d6342698c9fbb5c8a10a3db5032849870951fb"},"cell_type":"code","source":"from multiprocessing import Pool, cpu_count\nprint(\"Number of available cpu cores: {}\".format(cpu_count()))\n\ndef process_in_parallel(function, list_):\n    with Pool(cpu_count()) as p:\n        tmp = p.map(function, list_)\n    return tmp","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"76bdf63e27e70ec6a510598080f58c758adeaf2a"},"cell_type":"markdown","source":"### Define pseudo functions"},{"metadata":{"trusted":true,"_uuid":"f0df5338b336c7f304246d20726482df71762a9f"},"cell_type":"code","source":"import re\nfrom tqdm import tqdm\nfrom collections import Counter, OrderedDict\nimport operator\nimport unicodedata as ud\n\nfrom nltk.tokenize import TweetTokenizer\n\ndef process_questions(list_sentences, dict_, exclude):\n    \"\"\"\n    Function that applies text preprocessing\n    \"\"\"\n    sleep(0.5)\n    print(\"Lower first in sentence\")\n    list_sentences = process_in_parallel(lower_first_in_sentence, list_sentences)\n    #list_sentences = [lower_first_in_sentence(s) for s in tqdm(list_sentences)]\n    \n    print(\"Change words - using dictionary\")\n    list_sentences = [change_word(s,dict_) for s in list_sentences]\n    \n    print(\"Remove 's and solo apostrophe ' from end of the word. ex. AI's -> AI \")\n    list_sentences = process_in_parallel(strip, list_sentences)\n    #list_sentences = [strip(s) for s in tqdm(list_sentences)]\n    \n    print(\"Replace digits with mask ex. 'Marek gets 3333' -> 'Marek gets ####'\")\n    list_sentences = process_in_parallel(clean_numbers,list_sentences)\n    #list_sentences = [clean_numbers(s) for s in tqdm(list_sentences)]\n    \n    print(\"Normalise unicode data to remove umlauts, accents etc.\")\n    #https://gist.github.com/j4mie/557354\n    list_sentences = [ud.normalize('NFKD', i).encode('ASCII', 'ignore') for i in list_sentences]\n    \n    print(\"Use TweetTokenizer() from NLTK for splitting words.\")\n    tokenizer = TweetTokenizer()\n    list_sentences = process_in_parallel(tokenizer.tokenize,list_sentences)\n    #list_sentences = [tokenizer.tokenize(i) for i in tqdm(list_sentences)]\n    \n    print(\"Create vocab\")\n    dictionary = Counter([item for sublist in list_sentences for item in sublist])\n    dictionary = OrderedDict(sorted(dictionary.items(), key=operator.itemgetter(1), reverse = True))\n\n    #exclude = [',','.','\"','(',')','[',']',\"'\",'’',\"\\\\\",'{','}','…','..']                      \n    print(\"Exclude some punctuations: {}\".format(\" \".join(exclude)))\n    list_sentences = [list(filter(lambda x: x not in exclude,i)) for i in list_sentences]\n\n    return list_sentences, dictionary","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"cb5aa9a2f866d4af44b9e5ebae20fecd67197944"},"cell_type":"code","source":"from time import sleep\nfrom tqdm import tqdm\nfrom sklearn.preprocessing import MinMaxScaler\n\ndef text_features(list_sentences, MAX_SEQUENCE_LENGTH):\n    \"\"\"\n    Function creates matrix with additional text features.\n    Column 1: Unique words rate\n    Column 2: Rate of all-caps words\n    Column 3: Sentence length rate (number of word / max sequence length parameter)\n    \"\"\"\n    print(\"Column 1: Unique words rate\\nColumn 2: Rate of all-caps words\\nColumn 3: Sentence length rate (number of word / max sequence length parameter)\")\n    sleep(0.2)\n    #\"Unique words rate\" \n    def uwr(seq): return len(set(seq))/len(seq) if len(seq)>0 else 0\n    #\"Rate of all-caps words\"\n    def acw(seq): return len(list(filter(lambda x: x.isupper() and x!='I', seq)))/len(seq) if len(seq)>0 else 0\n    #\"sentence length rate\" \n    def slr(seq): return len(seq)/MAX_SEQUENCE_LENGTH if len(seq)>0 else 0\n    \n    uwr_f = [uwr(i) for i in tqdm(list_sentences)]\n    acw_f = [acw(i) for i in tqdm(list_sentences)]\n    slr_f = [slr(i) for i in tqdm(list_sentences)]\n    scaler = MinMaxScaler()\n    return scaler.fit_transform(np.array([uwr_f,acw_f,slr_f]).transpose())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"af1d7fc148ae6d2a10bee987af29d4aa52036851"},"cell_type":"markdown","source":"## Let's get work started    \n    \n### Prepare sequences and additional fetures\n"},{"metadata":{"_uuid":"e7f5d44d96745bdbd46db062b2f97928556378e3"},"cell_type":"markdown","source":"Define dictionaries that maps words and characters to change"},{"metadata":{"trusted":true,"_uuid":"40156facb8e4bc0d8958dbee296e3ac2b36add1d"},"cell_type":"code","source":"signs = {\"'\":\"'\", \"‘\":\"'\",\"´\": \"'\", \"°\": \"\",\"`\": \"'\", '“': '\"', '”': '\"', '“': '\"',\n         \"₹\": \"e\",\"€\": \"e\", \"™\": \"tm\", \"√\": \" sqrt \", \"×\": \"x\", \"²\": \"2\",\"—\": \"-\", \n         \"–\": \"-\", \"’\": \"'\", \"_\": \"-\", \"£\": \"e\",'∞': 'infinity', 'θ': 'theta', '÷': '/', \n         'α': 'alpha', '•': '.', 'à': 'a', '−': '-','β': 'beta', '∅': '', '³': '3', 'π': 'pi'}\n\n\nmisspelled = { 'colour': 'color', 'centre': 'center', 'favourite': 'favorite', 'travelling': 'traveling', \n                 'counselling': 'counseling', 'theatre': 'theater', 'cancelled': 'canceled', 'labour': 'labor',\n                 'organisation': 'organization', 'wwii': 'world war 2', 'citicise': 'criticize', \n                 'youtu ': 'youtube ', 'Quorans':'Quora','Qoura': 'Quora', 'sallary': 'salary', 'Whta': 'What', \n                 'narcisist': 'narcissist', 'howdo': 'how do', 'whatare': 'what are', 'howcan': 'how can', \n                 'howmuch': 'how much', 'howmany': 'how many', 'whydo': 'why do', 'doI': 'do I', \n                 'theBest': 'the best', 'howdoes': 'how does', 'mastrubation': 'masturbation', \n                 'mastrubate': 'masturbate', \"mastrubating\": 'masturbating', 'pennis': 'penis', \n                 'Etherium': 'Ethereum', 'narcissit': 'narcissist', 'bigdata': 'big data', '2k17': '2017', \n                 '2k18': '2018', 'qouta': 'quota', 'exboyfriend': 'ex-boyfriend', 'airhostess': 'air hostess', \n                 \"whst\": 'what', 'watsapp': 'whatsapp', 'demonitisation': 'demonetization', \n                 'demonitization': 'demonetization', 'demonetisation': 'demonetization',\n                  \"’\":\"'\",\"‘\":\"'\",\"´\":\"'\",\"`\":\"'\",'9/11':'terrorism',\"Quoran\":\"Koran\",'1/2':'half',\n                  'cryptocurrencies':'cryptocurrency',\"Brexit\":'leave EU', \"Blockchain\":\"blockchain\",'..':''}\n\ncontractions = {\"ain't\": \"is not\", \"aren't\": \"are not\",\"can't\": \"cannot\", \"'cause\": \"because\", \n                       \"could've\": \"could have\", \"couldn't\": \"could not\", \"didn't\": \"did not\",  \n                       \"doesn't\": \"does not\", \"don't\": \"do not\", \"hadn't\": \"had not\", \"hasn't\": \"has not\", \n                       \"haven't\": \"have not\", \"he'd\": \"he would\",\"he'll\": \"he will\", \"he's\": \"he is\", \n                       \"how'd\": \"how did\", \"how'd'y\": \"how do you\", \"how'll\": \"how will\", \"how's\": \"how is\",  \n                       \"I'd\": \"I would\", \"I'd've\": \"I would have\", \"I'll\": \"I will\", \"I'll've\": \"I will have\",\n                       \"I'm\": \"I am\", \"I've\": \"I have\", \"i'd\": \"i would\", \"i'd've\": \"i would have\", \n                       \"i'll\": \"i will\",  \"i'll've\": \"i will have\",\"i'm\": \"i am\", \"i've\": \"i have\", \n                       \"isn't\": \"is not\", \"it'd\": \"it would\", \"it'd've\": \"it would have\", \"it'll\": \"it will\", \n                       \"it'll've\": \"it will have\",\"it's\": \"it is\", \"let's\": \"let us\", \"ma'am\": \"madam\", \n                       \"mayn't\": \"may not\", \"might've\": \"might have\",\"mightn't\": \"might not\",\n                       \"mightn't've\": \"might not have\", \"must've\": \"must have\", \"mustn't\": \"must not\", \n                       \"mustn't've\": \"must not have\", \"needn't\": \"need not\", \"needn't've\": \"need not have\",\n                       \"o'clock\": \"of the clock\", \"oughtn't\": \"ought not\", \"oughtn't've\": \"ought not have\", \n                       \"shan't\": \"shall not\", \"sha'n't\": \"shall not\", \"shan't've\": \"shall not have\", \n                       \"she'd\": \"she would\", \"she'd've\": \"she would have\", \"she'll\": \"she will\", \n                       \"she'll've\": \"she will have\", \"she's\": \"she is\", \"should've\": \"should have\", \n                       \"shouldn't\": \"should not\", \"shouldn't've\": \"should not have\", \"so've\": \"so have\",\n                       \"so's\": \"so as\", \"this's\": \"this is\",\"that'd\": \"that would\", \n                       \"that'd've\": \"that would have\", \"that's\": \"that is\", \"there'd\": \"there would\", \n                       \"there'd've\": \"there would have\", \"there's\": \"there is\", \"here's\": \"here is\",\n                       \"they'd\": \"they would\", \"they'd've\": \"they would have\", \"they'll\": \"they will\", \n                       \"they'll've\": \"they will have\", \"they're\": \"they are\", \"they've\": \"they have\", \n                       \"to've\": \"to have\", \"wasn't\": \"was not\", \"we'd\": \"we would\", \"we'd've\": \"we would have\", \n                       \"we'll\": \"we will\", \"we'll've\": \"we will have\", \"we're\": \"we are\", \"we've\": \"we have\", \n                       \"weren't\": \"were not\", \"what'll\": \"what will\", \"what'll've\": \"what will have\", \n                       \"what're\": \"what are\",  \"what's\": \"what is\", \"what've\": \"what have\", \"when's\": \"when is\",\n                       \"when've\": \"when have\", \"where'd\": \"where did\", \"where's\": \"where is\", \"where've\": \"where have\", \n                       \"who'll\": \"who will\", \"who'll've\": \"who will have\", \"who's\": \"who is\", \"who've\": \"who have\", \n                       \"why's\": \"why is\", \"why've\": \"why have\", \"will've\": \"will have\", \"won't\": \"will not\", \n                       \"won't've\": \"will not have\", \"would've\": \"would have\", \"wouldn't\": \"would not\", \n                       \"wouldn't've\": \"would not have\", \"y'all\": \"you all\", \"y'all'd\": \"you all would\",\n                       \"y'all'd've\": \"you all would have\",\"y'all're\": \"you all are\",\"y'all've\": \"you all have\",\n                       \"you'd\": \"you would\", \"you'd've\": \"you would have\", \"you'll\": \"you will\", \n                       \"you'll've\": \"you will have\", \"you're\": \"you are\", \"you've\": \"you have\" ,\n                       \"Isn't\":\"is not\", \"\\u200b\":\"\", \"It's\": \"it is\",\"I'm\": \"I am\",\"don't\":\"do not\"}\n\ndict_= {}\ndict_.update(signs)\ndict_.update(misspelled)\ndict_.update(contractions)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4f6a4dbc6b448be794bef4c70f543ae55154cc39"},"cell_type":"code","source":"exclude_list = [',', '.', '\"', ':', ')', '(', '-', '!', '?', '|', ';', \"'\", '$', '&', '/', '[', ']', '>', '%', '=', '#', '*', '+', '\\\\', '•',  '~', '@', '£', \n '·', '_', '{', '}', '©', '^', '®', '`',  '<', '→', '°', '€', '™', '›',  '♥', '←', '×', '§', '″', '′', 'Â', '█', '½', 'à', '…', \n '“', '★', '”', '–', '●', 'â', '►', '−', '¢', '²', '¬', '░', '¶', '↑', '±', '¿', '▾', '═', '¦', '║', '―', '¥', '▓', '—', '‹', '─', \n '▒', '：', '¼', '⊕', '▼', '▪', '†', '■', '’', '▀', '¨', '▄', '♫', '☆', 'é', '¯', '♦', '¤', '▲', 'è', '¸', '¾', 'Ã', '⋅', '‘', '∞', \n '∙', '）', '↓', '、', '│', '（', '»', '，', '♪', '╩', '╚', '³', '・', '╦', '╣', '╔', '╗', '▬', '❤', 'ï', 'Ø', '¹', '≤', '‡', '√', ]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0f7d27142f4846fd550fbe46ae2f3d1f1842969e"},"cell_type":"code","source":"#Apply text preprocessing\nsentences, vocab = process_questions(list_sentences_train+list_sentences_test, dict_, exclude_list)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"scrolled":false,"_uuid":"00c2cb94fb0af5d4e92ea81e7ddf200df93a5b20"},"cell_type":"code","source":"#Prepare matrix with additional features\nadditional_features = text_features(sentences, MAX_SEQUENCE_LENGTH)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a54659144afa83f9019662ad33f5fc11a5506110"},"cell_type":"code","source":"#Prepare embedding\n\nimport gc\nfrom time import sleep\nimport pandas as pd\nfrom keras.preprocessing.sequence import pad_sequences\n\ndef prepare_sequences():\n\n    word_index = {t: i+1 for i,t in enumerate(vocab)}\n    \n    glove = '../input/embeddings/glove.840B.300d/glove.840B.300d.txt'\n    wiki_news = '../input/embeddings/wiki-news-300d-1M/wiki-news-300d-1M.vec'\n    \n    print(\"Extracting GloVe embedding\")\n    embed_glove = load_embed(glove)\n\n    print(\"Extracting FastText embedding\")\n    embed_fasttext = load_embed(wiki_news)\n\n    print(\"Glove coverage: \")\n    oov_glove = check_coverage(vocab, embed_glove)\n\n    print(\"FastText coverage: \")\n    oov_fasttext = check_coverage(vocab, embed_fasttext)\n\n    print(\"Create embedding for word vocabulary\")\n    sleep(0.2)\n    wv_matrix, WV_DIM, nb_words = concat_embed(embed_glove, embed_fasttext, word_index)\n    \n    #del(vocab)\n    del(globals()['vocab'])\n    gc.collect()\n    del(embed_glove,embed_fasttext)\n    gc.collect()\n    sleep(5)\n\n    print(\"Create sequences\")\n    sleep(0.2)\n    sequences = [[word_index.get(t, 0) for t in sentence]\n                 for sentence in tqdm(sentences[:len(list_sentences_train)])]\n    test_sequences = [[word_index.get(t, 0)  for t in sentence] \n                      for sentence in tqdm(sentences[len(list_sentences_train):])]\n\n    print(\"Assign additional features to train / test list\")\n    sleep(0.2)    \n    additional_features_train = additional_features[:len(list_sentences_train),:]\n    additional_features_test = additional_features[len(list_sentences_train):,:]\n\n    del(globals()['list_sentences_train'])\n    del(globals()['list_sentences_test'])\n    \n    print(\"Pad sequences\")\n    sleep(0.2)\n    train_data = pad_sequences(sequences, maxlen=MAX_SEQUENCE_LENGTH, padding=\"pre\", truncating=\"post\")\n    #list_classes = [\"toxic\", \"severe_toxic\", \"obscene\", \"threat\", \"insult\", \"identity_hate\"]\n    train_df = pd.read_csv(\"../input/train.csv\")\n    train_y = train_df['target'].values\n    print('Shape of data tensor:', train_data.shape)\n    print('Shape of label tensor:', train_y.shape)\n\n    test_data = pad_sequences(test_sequences, maxlen=MAX_SEQUENCE_LENGTH, padding=\"pre\",truncating=\"post\")\n    print('Shape of test_data tensor:', test_data.shape)\n\n    gc.collect()\n    sleep(5)\n    return word_index, wv_matrix, WV_DIM, nb_words, train_data, train_y, test_data, additional_features_train, additional_features_test","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9306bee27c3e932dddccbd39ca6474b74bdbc6ab"},"cell_type":"code","source":"word_index, wv_matrix, WV_DIM, nb_words, train_data, train_y, test_data, train_data_a, test_data_a = prepare_sequences()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6ddf0f3fade4a64138a0162b20f92dcc4114ceac"},"cell_type":"code","source":"import gc\nfrom time import sleep\ngc.collect()\nsleep(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1fb023d90c393d0e513fbd0d70003689ab9219e6"},"cell_type":"markdown","source":"### Build model"},{"metadata":{"_uuid":"1c64be12be1ae2f3ab999e53673e4fca18710a01"},"cell_type":"markdown","source":"https://www.kaggle.com/hung96ad/pytorch-starter"},{"metadata":{"trusted":true,"_uuid":"f51f11f97c4b1f94944387ed57260814cb66d10b","scrolled":true},"cell_type":"code","source":"embed_size = WV_DIM # how big is each word vector\n#max_features = 95000 # how many unique words to use (i.e num rows in embedding vector)\nnb_words = nb_words #number of unique words\nmaxlen = MAX_SEQUENCE_LENGTH # max number of words in a question to use\n\nbatch_size = 2048\ntrain_epochs = 6\n\nSEED = 666\n\nprint(\"embed_size : {embed_size},\\nnb_words : {nb_words},\\nmaxlen : {maxlen},\\nbatch_size : {batch_size}, \\\n      \\ntrain_epochs : {train_epochs},\\nSEED : {SEED}\".format(\n    **{'embed_size':embed_size, 'nb_words':nb_words,'maxlen':maxlen,'batch_size':batch_size,'train_epochs':train_epochs,'SEED':SEED}))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ac652a3b22d2bf9bdace5d10e445732569b36e94"},"cell_type":"code","source":"import torch\nimport torch.nn as nn\nimport torch.utils.data","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"806308d51dc51863cfdce4761b9d44b0162b2023"},"cell_type":"code","source":"import random\nimport os\nimport torch\n\ndef seed_torch(seed=666):\n    random.seed(seed)\n    os.environ['PYTHONHASHSEED'] = str(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    torch.cuda.manual_seed(seed)\n    torch.backends.cudnn.deterministic = True","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0c5039da7534fe00f412c808aca33536315af080"},"cell_type":"code","source":"class Attention(nn.Module):\n    def __init__(self, feature_dim, step_dim, bias=True, **kwargs):\n        super(Attention, self).__init__(**kwargs)\n        \n        self.supports_masking = True\n\n        self.bias = bias\n        self.feature_dim = feature_dim\n        self.step_dim = step_dim\n        self.features_dim = 0\n        \n        weight = torch.zeros(feature_dim, 1)\n        nn.init.xavier_uniform_(weight)\n        self.weight = nn.Parameter(weight)\n        \n        if bias:\n            self.b = nn.Parameter(torch.zeros(step_dim))\n        \n    def forward(self, x, mask=None):\n        feature_dim = self.feature_dim\n        step_dim = self.step_dim\n\n        eij = torch.mm(\n            x.contiguous().view(-1, feature_dim), \n            self.weight\n        ).view(-1, step_dim)\n        \n        if self.bias:\n            eij = eij + self.b\n            \n        eij = torch.tanh(eij)\n        a = torch.exp(eij)\n        \n        if mask is not None:\n            a = a * mask\n\n        a = a / torch.sum(a, 1, keepdim=True) + 1e-10\n\n        weighted_input = x * torch.unsqueeze(a, -1)\n        return torch.sum(weighted_input, 1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5698ef1d84f4f6c7c25b8471d1e72abda18a69da"},"cell_type":"code","source":"from torch.nn import * \n\nclass NeuralNet(nn.Module):\n    def __init__(self):\n        super(NeuralNet, self).__init__()\n        \n        hidden_size = 64\n        \n        self.embedding = nn.Embedding(nb_words, embed_size)\n        self.embedding.weight = nn.Parameter(torch.tensor(wv_matrix, dtype=torch.float32))\n        self.embedding.weight.requires_grad = False\n        \n        self.embedding_dropout = nn.Dropout2d(0.1)\n        self.lstm = nn.LSTM(embed_size, hidden_size, bidirectional=True, batch_first=True)\n        self.gru = nn.GRU(hidden_size*2, hidden_size, bidirectional=True, batch_first=True)\n        \n        self.lstm_attention = Attention(hidden_size*2, maxlen)\n        self.gru_attention = Attention(hidden_size*2, maxlen)\n    \n        self.AvgPool1d = nn.AdaptiveAvgPool1d(1)\n        self.MaxPool1d = nn.AdaptiveMaxPool1d(1)\n\n        self.linear = nn.Linear(399, 192)\n        self.relu = nn.ReLU()\n        self.dropout = nn.Dropout(0.1)\n        self.linear2 = nn.Linear(192, 64)\n        self.out = nn.Linear(64, 1)\n        self.out_act = nn.Sigmoid()\n    \n    def forward(self, x, x_a):\n        h_embedding = self.embedding(x)\n        h_embedding = torch.squeeze(self.embedding_dropout(torch.unsqueeze(h_embedding, 0)))\n        \n        h_lstm, _ = self.lstm(h_embedding)\n        h_gru, _ = self.gru(h_lstm)\n        \n        h_lstm_atten = self.lstm_attention(h_lstm)\n        h_gru_atten = self.gru_attention(h_gru)\n        \n        avg_pool = torch.squeeze(self.AvgPool1d(h_gru))\n        max_pool = torch.squeeze(self.MaxPool1d(h_gru))\n        #avg_pool = torch.mean(h_gru, 1)\n        #max_pool, _ = torch.max(h_gru, 1)\n        #print(h_gru_atten.shape)\n        #print(avg_pool.shape)\n        #print(max_pool.shape)\n        #print(x_a.shape)\n\n        conc = torch.cat((h_gru_atten,h_lstm_atten, avg_pool, max_pool, x_a), 1)\n        #print(conc.shape)\n        conc = self.relu(self.linear(conc))\n        conc = self.dropout(conc)\n        conc = self.relu(self.linear2(conc))\n        conc = self.dropout(conc)       \n        out = self.out(conc)\n        y = self.out_act(out)\n        return y","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d4098c6d0e5a66b0d89efc8e0876896d862b39c8"},"cell_type":"code","source":"def sigmoid(x):\n    return 1 / (1 + np.exp(-x))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1d6b2c728872b920b5bb23aa052a22adada601e8"},"cell_type":"code","source":"from sklearn.metrics import f1_score, roc_auc_score\n\ndef threshold_search(y_true, y_proba):\n    best_threshold = 0\n    best_score = 0\n    for threshold in tqdm([i/100 for i in range(10,90)]):\n        score = f1_score(y_true=y_true, y_pred=y_proba > threshold)\n        if score > best_score:\n            best_threshold = threshold\n            best_score = score\n    search_result = {'threshold': best_threshold, 'f1': best_score}\n    return search_result","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0297f648b3fcca218637f78b33f403dbc2959661"},"cell_type":"markdown","source":"### Train model and predict"},{"metadata":{"trusted":true,"_uuid":"c66e412f4ed0bfe4c4b08bd5e8ab46adc09e00cf","scrolled":true},"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV, StratifiedKFold\nsplits = list(StratifiedKFold(n_splits=5, shuffle=True, random_state=SEED).split(train_data, train_y))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0b89bfdd66429c5b828e74f6a6c4e88c0e74ab74","scrolled":true},"cell_type":"code","source":"import warnings\nwarnings.filterwarnings('always')\n\nimport time\n\ntrain_preds = np.zeros((len(train_data)))\ntest_preds = np.zeros((len(test_data)))\n\nseed_torch(SEED)\n\nx_test_cuda = torch.tensor(test_data, dtype=torch.long).cuda()\nx_test_a_cuda = torch.tensor(test_data_a, dtype=torch.float32).cuda()\n\ntest_dataset = torch.utils.data.TensorDataset(x_test_cuda,x_test_a_cuda)\ntest_loader = torch.utils.data.DataLoader(test_dataset, batch_size=batch_size, shuffle=False)\n\nloss_kFold = []\n\nfor i, (train_idx, valid_idx) in enumerate(splits):\n    x_train_fold = torch.tensor(train_data[train_idx], dtype=torch.long).cuda()\n    x_train_a_fold = torch.tensor(train_data_a[train_idx], dtype=torch.float32).cuda()\n    y_train_fold = torch.tensor(train_y[train_idx, np.newaxis], dtype=torch.float32).cuda()\n    x_val_fold = torch.tensor(train_data[valid_idx], dtype=torch.long).cuda()\n    x_val_a_fold = torch.tensor(train_data_a[valid_idx], dtype=torch.float32).cuda()\n    y_val_fold = torch.tensor(train_y[valid_idx, np.newaxis], dtype=torch.float32).cuda()\n    \n    model = NeuralNet()\n    model.cuda()\n    \n    loss_fn = torch.nn.BCELoss()\n    optimizer = torch.optim.Adam(model.parameters(), lr=0.001, betas=(0.9, 0.999))\n    \n    train = torch.utils.data.TensorDataset(x_train_fold,x_train_a_fold, y_train_fold)\n    valid = torch.utils.data.TensorDataset(x_val_fold, x_val_a_fold,y_val_fold)\n    \n    train_loader = torch.utils.data.DataLoader(train, batch_size=batch_size, shuffle=True)\n    valid_loader = torch.utils.data.DataLoader(valid, batch_size=batch_size, shuffle=False)\n    \n    print(f'Fold {i + 1}')\n    loss_epoch = []\n    for epoch in range(train_epochs):\n        start_time = time.time()\n        \n        model.train()\n        avg_loss = 0.\n        losses = []\n        for x_batch,x_a_batch, y_batch in tqdm(train_loader, disable=True):\n            \n            optimizer.zero_grad()\n            # (1) Forward\n            y_pred = model(x_batch,x_a_batch)\n            # (2) Compute diff\n            loss = loss_fn(y_pred, y_batch)\n            # (3) Compute gradients\n            loss.backward()\n            # (4) update weights\n            optimizer.step()\n            avg_loss += loss.item() / len(train_loader)\n            losses.append(loss.data.cpu().numpy())\n        \n        loss_epoch.append(losses)\n        \n        model.eval()\n        valid_preds_fold = np.zeros((x_val_fold.size(0)))\n        test_preds_fold = np.zeros(len(test_data))\n        avg_val_loss = 0.\n        for i, (x_batch, x_2_batch, y_batch) in enumerate(valid_loader):\n            y_pred = model(x_batch,x_2_batch).detach()\n            avg_val_loss += loss_fn(y_pred, y_batch).item() / len(valid_loader)\n            valid_preds_fold[i * batch_size:(i+1) * batch_size] = y_pred.cpu().numpy()[:, 0]\n        \n        elapsed_time = time.time() - start_time \n        print('Epoch {}/{} \\t loss={:.4f} \\t val_loss={:.4f} \\t time={:.2f}s'.format(\n            epoch + 1, train_epochs, avg_loss, avg_val_loss, elapsed_time))\n        \n    for i, (x_batch,x_a_batch) in enumerate(test_loader):\n        y_pred = model(x_batch,x_a_batch).detach()\n        test_preds_fold[i * batch_size:(i+1) * batch_size] = y_pred.cpu().numpy()[:, 0]\n\n    train_preds[valid_idx] = valid_preds_fold\n    print(threshold_search(train_y[valid_idx], train_preds[valid_idx]))\n    \n    test_preds += test_preds_fold / len(splits)  \n    \n    loss_kFold.append(loss_epoch)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5eef469062e215e5e0a17191abd5023eb246c0f4"},"cell_type":"code","source":"search_result = threshold_search(train_y, train_preds)\nsearch_result","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"0768fb1fa65b7490fbc643b202a2084bea56e433"},"cell_type":"code","source":"test_df = pd.read_csv(\"../input/test.csv\")\nsub = pd.DataFrame({\"qid\": test_df[\"qid\"].values})\nsub['prediction'] = (test_preds > search_result['threshold']).astype(int)\nsub.to_csv(\"submission.csv\", index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a67c1ba44319ab1d31d4289d2951af451c4e05b2"},"cell_type":"code","source":"import warnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n%matplotlib inline  \n\nsns.set(rc={'figure.figsize':(12,6)})\nsns.distplot(test_preds)\nplt.title('Distribution of the test predictions probability')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c06314fd5ab87f3a7c41221a05256cd6a733b774"},"cell_type":"code","source":"from collections import Counter\nprint(\"The percentage of insincere questions in train set is {:.2%} .\".format(train_1_prc))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"45f75c34b10f679a4fd1e403bf4d343069d512f0","scrolled":true},"cell_type":"code","source":"test_pred_1_prc = Counter(sub[\"prediction\"])[1]/len(sub[\"prediction\"])\nprint(\"The percentage of predicted insincere questions in test set is {:.2%} .\".format(test_pred_1_prc))","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}