{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"**Hi !**\n\nThis kernel is an improved extract from my first kernel, focusing on questions id and how Quora probably built the train and test datasets. I don't know what to do of this work, if you have any idea feel free to share it!\n"},{"metadata":{"_uuid":"d307dca2d3ffb2bf5edcaa3edbbd0dd3bbfe3536"},"cell_type":"markdown","source":"I was wondering wether the test dataset was an extract from the train one, or completely different. To answer this existential question, I've been working on the 'qid' column.\n\nThe question id is an hexadecimal number. The first step here is to extract this value, and see if the questions are ordered by id."},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\n\nimport seaborn as sns\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nprint(os.listdir(\"../input\"))\n\nimport warnings\nwarnings.filterwarnings('ignore')\n\n# Any results you write to the current directory are saved as output.\n\nsns.set()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"94b0be36a83d7c0303982d19a500e3c096b713f8","_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"filepath_train = os.path.join('..', 'input', 'train.csv')\nfilepath_test = os.path.join('..', 'input', 'test.csv')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ffeba9ca8bcfbb0102de09041db4ca9550ce09ca"},"cell_type":"code","source":"df_train = pd.read_csv(filepath_train)\ndf_test = pd.read_csv(filepath_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f5c6cd8a09574ab0d79df07317e9ea0e9168fd8b"},"cell_type":"code","source":"df_train.shape, df_test.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"942ded7d207a94c78da5fe1ad214adfec2cd0fbf","scrolled":true},"cell_type":"code","source":"df_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"40bfe90a1ff321c78bd7cdbecdd045f17a0358ad"},"cell_type":"code","source":"df_train_qid = df_train.copy()\n\ndf_train_qid['qid_base_ten'] = df_train_qid['qid'].apply(lambda x : int(x, 16))\ndf_train_qid.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"684657ac1d6ba0b77e0ebd1c487424646db4a009"},"cell_type":"code","source":"min_qid = df_train_qid['qid_base_ten'].min()\nmax_qid = df_train_qid['qid_base_ten'].max()\ndf_train_qid['qid_base_ten_normalized'] = df_train_qid['qid_base_ten'].apply(lambda x : (x - min_qid)/min_qid)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7da268d4145b1eb889783456cfc2bbe543e21c55"},"cell_type":"code","source":"plt.figure(figsize=(18, 8));\nplt.scatter(x=df_train_qid['qid_base_ten_normalized'][:100], y=df_train_qid.index[:100]);\nplt.xlabel('qid_base_ten_normalized');\nplt.ylabel('Question index in df_train_qid');","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2b05794f9f01195ae858fd94bac427835334ef8b"},"cell_type":"markdown","source":"As I suspected, questions are indeed sorted by ascending question id in our train dataset. Let's see if it is the same in the test one."},{"metadata":{"trusted":true,"_uuid":"d00809011248c4242799f4dddcab516aafb55471"},"cell_type":"code","source":"df_test_qid = df_test.copy()\n\ndf_test_qid['qid_base_ten'] = df_test_qid['qid'].apply(lambda x : int(x, 16))\n\ndf_test_qid['qid_base_ten_normalized'] = df_test_qid['qid_base_ten'].apply(lambda x : (x - min_qid)/min_qid)\n\nplt.figure(figsize=(18, 8));\nplt.scatter(x=df_test_qid['qid_base_ten_normalized'][:100], y=df_test_qid.index[:100]);\nplt.xlabel('qid_base_ten_normalized');\nplt.ylabel('Question index in df_test_qid');","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6322575dbe36fc90bb130d078f13ecd6a54d47f2"},"cell_type":"markdown","source":"Here again, questions are sorted by ascending question id ! Now I wonder if I can know how Quora has made its train an test datasets. Is it with a random (and stratified?) train.test split, or a simple split based on the id?\n\nTo get the answer, I have merged the train and test dataframes, with the 'qid_base_ten_normalized' column, sorted by ascending 'qid_base_ten_normalized' and reset the index."},{"metadata":{"trusted":true,"_uuid":"4b4e6291867ebe7eb16fc5fa7490d3fed33aee4b"},"cell_type":"code","source":"df_train_qid.drop('target', axis=1, inplace=True)\ndf_train_qid['test_or_train'] = 'train'\ndf_test_qid['test_or_train'] = 'test'","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"scrolled":true,"_uuid":"f07f9cd139da5b599976abab2376c257a6b55357"},"cell_type":"code","source":"df_qid = pd.concat([df_train_qid, df_test_qid]).sort_values('qid_base_ten_normalized').reset_index()\ndf_qid.drop('index', axis=1, inplace=True)\ndf_qid.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"scrolled":false,"_uuid":"ad8f1e52af557dd6ea8f09dc6e397c9001b2a029"},"cell_type":"code","source":"df_qid_train = df_qid[df_qid['test_or_train']=='train']\ndf_qid_test = df_qid[df_qid['test_or_train']=='test']\n\nplt.figure(figsize=(18, 8));\nplt.scatter(x=df_qid_train['qid_base_ten_normalized'], y=df_qid_train.index, label='Train', s=300);\nplt.scatter(x=df_qid_test['qid_base_ten_normalized'], y=df_qid_test.index, label='Test',s=2);\nplt.xlabel('qid_base_ten_normalized');\nplt.ylabel('Question index');\nplt.title('qid_base_ten_normalized for train and test datasets')\nplt.legend();","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3c0c6a597e9ed7680a38ee76ae9ae43f42cd3e69"},"cell_type":"markdown","source":"So the question ids range of the test dataset is the same as the question ids range for the train dataset. The test and train datasets come as expected from a random train/test split on a single dataset.\nThe figure below confirms the 'random' choice of the elements for the test dataset."},{"metadata":{"trusted":true,"_uuid":"fc743ea3640d16eec975832a61e00a730a3d00a5"},"cell_type":"code","source":"plt.figure(figsize=(18, 8));\nplt.scatter(x=df_qid_train['qid_base_ten_normalized'][:1500], y=df_qid_train.index[:1500], label='Train');\nplt.scatter(x=df_qid_test['qid_base_ten_normalized'][:50], y=df_qid_test.index[:50], label='Test',s=150, marker='d');\nplt.xlabel('qid_base_ten_normalized');\nplt.ylabel('Question index');\nplt.title('qid_base_ten_normalized for the first 1500 train points and 50 test points')\nplt.legend();","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ed139a69ec6e1803287c0a424a574c23cbab624c"},"cell_type":"markdown","source":"Here, we still can't figure out if the train/test plit was done using the 'stratify' parameter of `train_test_split()`!\n"},{"metadata":{"_uuid":"1703ea2605f4e700243b390a255ef8fbec3f1c95"},"cell_type":"markdown","source":"# Working on 'distance' between questions"},{"metadata":{"trusted":true,"_uuid":"b2b1dfbdbfbe2e59d341c6722eb906831a8b2642"},"cell_type":"code","source":"df_train_qid.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4fc478b8ed61bbcd11b37aa0eb9292d86ba88aa4"},"cell_type":"code","source":"df_train_0 = df_train_qid[df_train['target']==0][:5000].drop('test_or_train', axis=1)\ndf_train_1 = df_train_qid[df_train['target']==1][:5000].drop('test_or_train', axis=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a4a2b5e364af2f916e3ae0bdf3b49c4a57e7188d"},"cell_type":"code","source":"df_train_0['qid_distance'] = 0\n\nfor i in range(1, df_train_0.shape[0]):\n    df_train_0.iloc[i, 4] = df_train_0.iloc[i, 3] - df_train_0.iloc[i-1, 3]\ndf_train_0.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"76ef8ca81d34f24a891cf8feb436b79b1db167fc"},"cell_type":"code","source":"df_train_1['qid_distance'] = 0\n\nfor i in range(1, df_train_1.shape[0]):\n    df_train_1.iloc[i, 4] = df_train_1.iloc[i, 3] - df_train_1.iloc[i-1, 3]\ndf_train_1.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fb3036fc5d390ce93fae7c3f1fc639ada3d487b9"},"cell_type":"code","source":"print('The mean \\'qid_distance\\' for sincere questions is {} with a standard deviation of {}'.format(round(df_train_0['qid_distance'].mean(), 3), round(df_train_0['qid_distance'].std(), 3)))\nprint('The mean \\'qid_distance\\' for insincere questions is {} with a standard deviation of {}'.format(round(df_train_1['qid_distance'].mean(), 3), round(df_train_1['qid_distance'].std(), 3)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"91aa152f14bab5f727bdcca9948be8d5fb106d0e"},"cell_type":"code","source":"fig, axes = plt.subplots(1, 2, figsize=(18, 8));\nsns.distplot(df_train_1['qid_distance'], ax=axes[0], color='red', label='Insincere questions', kde=False);\naxes[0].legend();\nsns.distplot(df_train_0['qid_distance'], ax=axes[1], label='Sincere questions', kde=False);\naxes[1].legend();","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"de419fadef15182c5530608667bebe833ce65612"},"cell_type":"code","source":"plt.figure(figsize=(18, 8));\nsns.boxplot(df_train_1['qid_distance'], color='red');\nplt.title('Insincere questions');\nplt.figure(figsize=(18, 8));\nsns.boxplot(df_train_0['qid_distance']);\nplt.title('Sincere questions');","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"740d59bddc8fd6dd3a723b11dce93a8c8f173169"},"cell_type":"code","source":"df_train_qid.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1c270ad36ceae4907afea37dad9516b5697654bd"},"cell_type":"code","source":"df_train_distance = df_train_qid[:100000].drop(['qid_base_ten', 'test_or_train'], axis=1).copy()\ndf_train_distance['qid_distance'] = 0\nfor i in range(1, df_train_distance['qid'].shape[0]):\n    df_train_distance.iloc[i, 3] = df_train_distance.iloc[i, 2] - df_train_distance.iloc[i-1, 2]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f3544073fdbe9f08c0067fed799af70a24ca099b"},"cell_type":"code","source":"df_train_distance.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"43cd1fd7505c2449670f2d4de381f55b447ed094"},"cell_type":"code","source":"df_train_distance_0 = df_train_distance[df_train['target']==0]\ndf_train_distance_1 = df_train_distance[df_train['target']==1]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b40c74853b2c81af1f69bf69e5d1a70feea8a853"},"cell_type":"code","source":"ax=plt.gca()\nsns.distplot(df_train_distance_0['qid_distance'], ax=ax, color='red', label='Insincere questions');\nsns.distplot(df_train_distance_1['qid_distance'], ax=ax, label='Sincere questions');\nax.legend();","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"044c2fa2940f526e1eccd682b601baa4d40256ec"},"cell_type":"code","source":"sns.boxplot(x=df_train_distance['qid_distance'], y=df_train['target'], orient='h');","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e75c123bcbb965d4287f882af83fd1a29d3b7560"},"cell_type":"markdown","source":"That's all for this bonus part on questions id, it is not that useful, but it was working on it was fun!\n\n![FolksURL](https://media.giphy.com/media/upg0i1m4DLe5q/giphy.gif \"Folks\")"},{"metadata":{"trusted":true,"_uuid":"5b2335933431bc4b0f7e822075db60808e790325"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}