{"cells":[{"metadata":{"_uuid":"f0924bc3eb5bd2b99ce3d7e666d24db348435ff5"},"cell_type":"markdown","source":"# Quora Insincere Questions - Starter without NLP Frameworks"},{"metadata":{"_uuid":"7531604b711e95e6c97503f558319db6eb7b2d8c"},"cell_type":"markdown","source":"This Python Notebook is an humble effort, without using any time-consuming NLP framework, to tackle the Kaggle challenge on Quora Insincere Questions.\n\nUsing the given training dataset, the Notebook prepares a list of one words, two consecutive words, and two alternate words along with their corresponding insincerity indices.\n\nThen, using this list as a reference, the Notebook computes insincerity index for each question in the given test dataset."},{"metadata":{"_uuid":"f7beee06ea0bf70505534f547cc3abc8137a262e"},"cell_type":"markdown","source":"Import necessary packages."},{"metadata":{"trusted":false,"_uuid":"a349b623f046fe4052c95a3a1cd406d9b5e8a89a"},"cell_type":"code","source":"import pandas as pd\nfrom itertools import chain\nfrom collections import Counter","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"23e8b7ae774fa5a3a2560fe7c0907e551b2ac8c2"},"cell_type":"markdown","source":"Make dataframes from csv files of training dataset and test dataaset."},{"metadata":{"trusted":false,"_uuid":"1dd0ac99671ea0af2a11773557cc3b886679dafb"},"cell_type":"code","source":"train_df = pd.read_csv('../input/train.csv')\ntest_df = pd.read_csv('../input/test.csv')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"93ad2c7f56db0691250c47e129ef642ff7826441"},"cell_type":"markdown","source":"Define a method to print something in between its title  and separator."},{"metadata":{"trusted":false,"_uuid":"c1fe3fb0ca2b5252ec6e467f8493ce9b92cb4718"},"cell_type":"code","source":"def print_block(sometitle, someblock):\n    '''\n    Print something in between its title  and separator.\n    '''\n    print(sometitle)\n    print(\"\\n\")\n    print(someblock)\n    print(\"\\n\" + \"=\"*80 + \"\\n\")","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3cfc8f1f5fcf3a524e182b35701d9cc864debc38"},"cell_type":"markdown","source":"Define a method to trim a text, and split it into one-words and two-words."},{"metadata":{"trusted":true,"_uuid":"e6c353e99f6390d77372bfa0e3b9d6ff3a42d1b5"},"cell_type":"code","source":"def text_to_words(orig_text):\n    '''\n    Trim a text, and split it into one-words and two-words.\n    '''\n    word_vect = orig_text.lower()\n    word_vect = word_vect.replace(\"  \",\" \")\n    remove_char = ['.', '?', '!', ',', ';', '(', ')', '\"', \"\"\"''\"\"\"]\n    for i in range(9):\n        word_vect = word_vect.replace(remove_char[i], \"\")\n        \n    # Make a list of one words.\n    word_vect = word_vect.split()\n    word_count = len(word_vect)\n    \n    # Add two consecutive words.\n    if word_count > 2:\n        for i in range(word_count - 2):\n            word_vect.append(word_vect[i] + \" \" + word_vect[i+1])\n    \n    # Add two alternate words.\n    if word_count > 3:\n        for i in range(word_count - 3):\n            word_vect.append(word_vect[i] + \" \" + word_vect[i+2])\n    \n    return word_vect","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4298b220bba6261466de954507aedc9efe72dc69"},"cell_type":"markdown","source":"Define a method to prepare a dataframe of words and label-indices, based on a given dataframe of texts and lables, "},{"metadata":{"trusted":false,"_uuid":"11057db882a697bda85dd6a98513a2d91ef84109"},"cell_type":"code","source":"def df_insincerity(df):\n    '''\n    Prepare a dataframe of words and label-indices,\n    based on a given dataframe of texts and lables.\n    '''\n    any_count = len(df)\n    df_any = df['question_text'].apply(lambda x: text_to_words(x))\n    print_block(\"Words in first five questions\", df_any[:5])\n    df_any = Counter(chain.from_iterable(df_any[i] for i in range(any_count)))\n    print_block(\"5 most common words\", Counter.most_common(df_any)[:5])\n    \n    insincere_count = len(df[df['target']==1])    \n    df_insincere = df[df['target']==1]['question_text'].apply(\n                   lambda x: text_to_words(x)).reset_index(drop=True)\n    print_block(\"Words in first five insincere questions\", df_insincere[:5])    \n    df_insincere = Counter(chain.from_iterable(df_insincere[i] for\n                   i in range(insincere_count)))\n    print_block(\"5 most common words from insincere questions\",\n                Counter.most_common(df_insincere)[:5])\n    \n    sincere_count = len(df[df['target']==0])    \n    df_sincere = df[df['target']==0]['question_text'].apply(\n                 lambda x: text_to_words(x)).reset_index(drop=True)\n    print_block(\"Words in first five sincere questions\", df_sincere[:5])    \n    df_sincere = Counter(chain.from_iterable(df_sincere[i] for\n                 i in range(sincere_count)))\n    print_block(\"5 most common words from sincere questions\",\n                Counter.most_common(df_sincere)[:5])\n    \n    df_insincerity_a = {k: round(df_insincere[k]/insincere_count -\n                                 df_sincere[k]/sincere_count,\n                                 2) for k in df_any.keys()}    \n    df_insincerity_a = Counter({k: float(round(v, 2)) for k, v in\n                       df_insincerity_a.items() if abs(v) > 0.0001})\n    df_insincerity_a = pd.DataFrame(Counter.most_common(df_insincerity_a),\n                       columns=[\"word\",\"insincerity\"])\n    print_block(\"5 most insincere words\", df_insincerity_a[:5])\n    print_block(\"5 least insincere words\", df_insincerity_a[-5:])\n    \n    return df_insincerity_a","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f84caa73a2c9a704515c9dd5ffd513a68ef17b66"},"cell_type":"markdown","source":"Prepare a dataframe of words and label-indices, based on the given training dataset."},{"metadata":{"scrolled":false,"trusted":false,"_uuid":"48736a4af595054ac324056290b11eceb516748c"},"cell_type":"code","source":"df_insincerity = df_insincerity(train_df)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3e5b27048341eec24a79085121d66b6d52277677"},"cell_type":"markdown","source":"Define a way to compute insincerity index of a question, based on the prepared list of words with insincerity indices. "},{"metadata":{"trusted":false,"_uuid":"d090cae18d941723943a3427f882e3842d171b72"},"cell_type":"code","source":"def question_insincerity(quest):\n    '''\n    Compute insincerity index of a question,\n    based on the prepared list of words with insincerity indices.\n    '''\n    df_questword = pd.DataFrame(text_to_words(quest), columns=[\"word\"])\n    df_questword = pd.merge(df_questword, df_insincerity,\n                   on='word', how='inner')\n    \n    if sum(df_questword['insincerity']) > 0:\n        return 1\n    else:\n        return 0","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"784bcb0fca635966b8efd27eeb42aa4a2f0e1d59"},"cell_type":"markdown","source":"Compute insincerity indices for the test questions, based on the prepared list of words with insincerity indices."},{"metadata":{"scrolled":false,"trusted":false,"_uuid":"c5c9869d90f01f48090903e71e98678cfe11590d"},"cell_type":"code","source":"test_df['prediction'] = test_df['question_text'].apply(\n                        lambda x: question_insincerity(x))\nprint(test_df['prediction'][:50])","execution_count":null,"outputs":[]},{"metadata":{"trusted":false,"_uuid":"3bb838cc7034e4087dd996450fbd28e5cb16d337"},"cell_type":"markdown","source":"Make csv file of the predicted results as needed for the submission."},{"metadata":{"trusted":false,"_uuid":"3ef29be89a944d2ff3a273b6c2e9f4abfeb602a1"},"cell_type":"code","source":"test_df[['qid','prediction']].to_csv('submission.csv', index = False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"68ca03d3a3b2b977128fc86b53bcaed91e0e66a6"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}