{"cells":[{"metadata":{"_uuid":"f95d4b52716ed50f31abbea2105306aab1008d12"},"cell_type":"markdown","source":"A lot of insightful info has been found by kagglers. This kernel won't repeat work that has been done but focuses on one single point which may worth making a feature down the track.\n\nWhile doing sanity check, I was a bit suprised price, as a crucial part of advertisement, is missing many values. Will that reduce the probability that the product would be sold?"},{"metadata":{"collapsed":true,"trusted":true,"_uuid":"d4a07cfe96b9fe96ed455fe6c5befd3ca69d0e6f"},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n%matplotlib inline\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt","execution_count":8,"outputs":[]},{"metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","trusted":true},"cell_type":"code","source":"# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.","execution_count":2,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true,"collapsed":true},"cell_type":"code","source":"train = pd.read_csv('../input/train.csv')","execution_count":4,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5255ebccd5c1b4a488ed23b86ac04886f383e06c"},"cell_type":"code","source":"train.shape","execution_count":5,"outputs":[]},{"metadata":{"_uuid":"0679bb658defb7abc64bb83c526a32903ff87b6b"},"cell_type":"markdown","source":"There are many price data missing."},{"metadata":{"trusted":true,"_uuid":"9ce3441f19564d61a4d4f458d23823ea640ee620"},"cell_type":"code","source":"train.isnull().sum()","execution_count":6,"outputs":[]},{"metadata":{"_uuid":"cfed10b6da9b15a06f299588606302ebe155b02f"},"cell_type":"markdown","source":"Missing rate:"},{"metadata":{"trusted":true,"_uuid":"3e70c74e0d3d69cf451076ab8fedf2bf0265042b"},"cell_type":"code","source":"train.isnull().sum() / train.shape[0] * 100","execution_count":26,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"df89f8759ebd04d7733d197c2e9ef781138cbd06"},"cell_type":"code","source":"train.describe()","execution_count":10,"outputs":[]},{"metadata":{"_uuid":"d798159c7bbc211ab7145611df47b24b9b6e5c87"},"cell_type":"markdown","source":"Before heading to the illustration of products missing price, let's first have a quick look at the replation between price and deal probability.  Having the max price will suppress most points to the left of the chart and make it hard to read. We plot only 75% of the data."},{"metadata":{"trusted":true,"_uuid":"83b8cec02b42db0981d4437bd5513c962b18ced0"},"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15, 8))\ntrain.loc[train.price < 7e3, ['price', 'deal_probability']].plot(kind='scatter', x='price', y='deal_probability', ax=ax, alpha=0.1, color='r')\nax.grid()\nplt.show()","execution_count":15,"outputs":[]},{"metadata":{"_uuid":"31722047ee0ba041f62b9adafff63aeaa0ecc826"},"cell_type":"markdown","source":"The plot above actually makes sense that most prices are populated around times of hundreds as appearing as the stripes in the chart.\n\nThen quite suprisingly again, price-missing products actually have relatively higher chance to be sold."},{"metadata":{"trusted":true,"_uuid":"95e1ca8fd53abec2acd04ad6b0569b9fa1d2d330"},"cell_type":"code","source":"fig, ax = plt.subplots(figsize=(15, 8), nrows=2)\ntrain.loc[train.price.isnull(), 'deal_probability'].plot(kind='hist', bins=20, ax=ax[0], color='r', grid=True)\nax[0].set_title('No price')\ntrain.loc[train.price < 7e3, 'deal_probability'].plot(kind='hist', bins=20, ax=ax[1], color='b', grid=True)\nax[1].set_title('Price < 7,000')\nax[1].set_xlabel('Deal probability')\nplt.show()","execution_count":24,"outputs":[]},{"metadata":{"collapsed":true,"trusted":true,"_uuid":"c48744bbecd8f9f2a8059a70ef227a693044901e"},"cell_type":"markdown","source":"Wondering what percentage of price data are mising in test set."},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"814dc3fdff5d6e30fdcf3bc0dd10ce56bb624fd3"},"cell_type":"code","source":"test = pd.read_csv('../input/test.csv')","execution_count":27,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"61f33792b14952e40a1e7528f43da549395a05d7"},"cell_type":"code","source":"test.isnull().sum() / test.shape[0] * 100","execution_count":28,"outputs":[]},{"metadata":{"_uuid":"f7b4b97f90476486c82111528b018aefba28129f"},"cell_type":"markdown","source":"The missing rate 6% is slightly higer than 5.6% for the train set."},{"metadata":{"_uuid":"db4846a96b7afd5036b1e95a59dc502a068eccc5"},"cell_type":"markdown","source":"Anyway, this is a data quality issue that should be addressed. Maybe it means price negotiable, which probably explains the higher deal probability."},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"d6f9095216bec4e4f8c0d69f9cbad3bd6775d0e2"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.5","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}