{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "29122681-33cb-35f3-ef27-e1bfe4efd4a2"
      },
      "source": [
        "This notebook checks the training set for full duplicates that might be wrongly classified.\n",
        "This can easily be further expanded.\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "937a071e-259f-72f4-7b73-dd8d0332f5d5"
      },
      "outputs": [],
      "source": [
        "import numpy as np # linear algebra\n",
        "import pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n",
        "import re\n",
        "import string\n",
        "from copy import deepcopy\n",
        "\n",
        "pd.set_option('expand_frame_repr', False)\n",
        "pd.set_option('display.max_colwidth', -1)\n",
        "\n",
        "train_df = pd.read_csv('../input/train.csv')\n",
        "test_df = pd.read_csv('../input/test.csv')\n",
        "# samplesub_df = pd.read_csv(\"../input/sample_submission.csv')\n",
        "\n",
        "print('Train shape', train_df.shape)\n",
        "print('Test shape', test_df.shape)\n"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "be67e773-adb7-ac66-1cad-4755f3257581"
      },
      "source": [
        "Now let's look for rows where both question 1 and question 2 are identical.\n",
        "Additionally we look for NaN values in these fields.\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "7a4f81c1-f2ba-67eb-de41-8c474a0e6c3c"
      },
      "outputs": [],
      "source": [
        "train_full_duplies = train_df[train_df['question1'] == train_df['question2']]\n",
        "print('Fully equal training questions - no processing: ', train_full_duplies.shape[0])\n",
        "\n",
        "train_drop = train_df.dropna(how=\"any\")\n",
        "# train_df[~train_df['id'].isin(train_drop['id'])]\n",
        "train_nan = train_df[train_df['id'].isin(train_drop['id']) == False]\n",
        "print('Training questions with NaN - no processing: ', train_nan.shape[0])\n",
        "train_no_nan = train_drop\n",
        "train_no_nan_raw = deepcopy(train_drop) # we need this later\n",
        "train_drop = None\n",
        "\n",
        "for index, row in train_nan.iterrows():\n",
        "    print('Train_NaN: ', row['id'], row['qid1'], row['qid2'], ' Q1: ', row['question1'], ' Q2: ', row['question2'],\n",
        "          ' is duplicate: ', row['is_duplicate'])\n"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "ff2e3e78-7669-c44d-915b-b8c7f9f02f13"
      },
      "source": [
        "This shows that there are no directly comparable question duplicates in the train data. Let's check again with a bit processing.\n",
        "\n",
        "We make all questions lower case and remove the punctuation chars.\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "f40c5684-85ff-e82c-0f4c-47b232a7ad11"
      },
      "outputs": [],
      "source": [
        "def remove_punct(val):\n",
        "    # remove all punctuation chars\n",
        "    regex = re.compile('[%s]' % re.escape(string.punctuation))\n",
        "    sentence = regex.sub('', val).lower()\n",
        "    \n",
        "    return sentence\n",
        "\n",
        "def clean_dataframe(data):\n",
        "    # first remove punctuation than make lowercase\n",
        "    for col in ['question1', 'question2']:\n",
        "        data[col] = data[col].apply(remove_punct)\n",
        "\n",
        "    return data\n",
        "\n",
        "train_data_clean = clean_dataframe(train_no_nan)\n",
        "# print(train_data_clean.head(5))\n",
        "\n",
        "train_full_duplies_punct = train_data_clean[train_data_clean['question1'] == train_data_clean['question2']]\n",
        "print('Fully equal training questions punctuation removed: ', train_full_duplies_punct.shape[0])\n"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "036a3aa5-1927-bf40-aa06-8b02b06b16b5"
      },
      "source": [
        "Maybe we have removed a bit too much - let's get a closer look. We display both versions of the questions - with and without punctuation.\n",
        "\n",
        "We display all those question pairs that are marked as is_duplicate = 0 from the ones found above and compare them with the raw version (with punctuation).\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "f275732b-d517-b9ad-9142-d590975fca64"
      },
      "outputs": [],
      "source": [
        "i_wrong_class = 0\n",
        "for index, row in train_full_duplies_punct.iterrows():\n",
        "    if row['is_duplicate'] == 0:\n",
        "        i_wrong_class += 1\n",
        "        print('Train-duplicates potentially wrongly classified: ', row['id'], row['qid1'], row['qid2'], ' [Q1]: ', row['question1'], ' [Q2]: ',\n",
        "              row['question2'], ' [is duplicate]: ', row['is_duplicate'], '\\n')\n",
        "        raw_row = train_no_nan_raw.loc[(train_no_nan_raw['qid1'] == row['qid1']) & (train_no_nan_raw['is_duplicate'] == 0)]\n",
        "        print('Orig: [raw Q1]:', format(raw_row['question1']), ' [raw Q2]: ', format(raw_row['question2']), '\\n' )\n",
        "\n",
        "print ('Total of potentially wrongly classified in training set: ', format(i_wrong_class))"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "_cell_guid": "8102fd82-b0bb-9f91-ea5a-8087df59f2e0"
      },
      "source": [
        "We can see that some of these 69 pairs returned are indeed wrongly marked as not duplicate. \n",
        "\n",
        "My favorite: \"...*I am 17 now, how can I \"earn\" my first house or Lamborghini within 5 years? (One of my hobbies is Animation if that matters.)*...\"\n",
        "\n",
        "I left the select from the raw rows in a way that this also shows when questions exists multiple times in the data set.\n",
        "\n",
        "Next step: check for fully duplicate question pairs shared between the training and test set."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "_cell_guid": "96123b69-25d5-8bb9-c5f9-3c834b1b2427"
      },
      "outputs": [],
      "source": [
        ""
      ]
    }
  ],
  "metadata": {
    "_change_revision": 0,
    "_is_fork": false,
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "codemirror_mode": {
        "name": "ipython",
        "version": 3
      },
      "file_extension": ".py",
      "mimetype": "text/x-python",
      "name": "python",
      "nbconvert_exporter": "python",
      "pygments_lexer": "ipython3",
      "version": "3.6.0"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}