{"cells":[{"metadata":{"_uuid":"524d13941856ffa94c0aa41f689f15f272fcce11"},"cell_type":"markdown","source":"This notebook was motivated by this great [notebook](https://www.kaggle.com/yuliagm/be-careful-about-ips-as-a-signal/notebook) from @yulia .  She discuss there a potential difference between train and test data.    Here is a simple way to see that difference."},{"metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.\n\nimport matplotlib.pyplot as plt\n%matplotlib inline","execution_count":12,"outputs":[]},{"metadata":{"_uuid":"0f49603c3f650e6074b7efcf38e8e46d55f753ef"},"cell_type":"markdown","source":"Let's load some data."},{"metadata":{"collapsed":true,"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"dtypes = {\n        'ip'            : 'uint32',\n        'app'           : 'uint16',\n        'device'        : 'uint16',\n        'os'            : 'uint16',\n        'channel'       : 'uint16',\n        'is_attributed' : 'uint8',\n        'click_id'      : 'uint32'\n        }\n\ntrain = pd.read_csv('../input/train.csv', dtype=dtypes, usecols=['ip', 'is_attributed'])\ntest = pd.read_csv('../input/test.csv', dtype=dtypes, usecols=['ip'])\n","execution_count":2,"outputs":[]},{"metadata":{"_uuid":"750a00d14f144d32207e0b78920d968135932ee6"},"cell_type":"markdown","source":"Let's now look at the downlad rate per ip in the train data."},{"metadata":{"trusted":true,"_uuid":"59ba92655b589e89a5c679bf2abc869e7a35b158"},"cell_type":"code","source":"df = train.groupby('ip').is_attributed.mean().to_frame().reset_index()\n\ndf.head()","execution_count":3,"outputs":[]},{"metadata":{"_uuid":"be5002e7c53cdad19285094ae7d4edb48cc6c33e"},"cell_type":"markdown","source":"One way to display numeric data and spot trend is to use a moving average.  Let's try it here."},{"metadata":{"trusted":true,"_uuid":"93b02886c43603b8e88cabeb02cf85b9f6c9a873"},"cell_type":"code","source":"df['roll'] = df.is_attributed.rolling(window=1000).mean()\nplt.plot(df.ip, df.roll)","execution_count":6,"outputs":[]},{"metadata":{"_uuid":"7989fb3c56790dbd971681fad6f82e6b6cc639f5"},"cell_type":"markdown","source":"There is a clear cut split around 130,000, and a lighter split around 220,000.  Below the first split ip have about 0.02 app download rate, then the rate climbs above 0.25.  This is a 10x increase, worth eploring further.\n\nA little trial and error leads to an identification of where the split is."},{"metadata":{"trusted":true,"_uuid":"15beabf4bc4605ee9291ebd7cb5fc1859375ad41"},"cell_type":"code","source":"df1 = df[(df.ip >= 120000) & (df.ip <= 130000)]\nplt.plot(df1.ip, df1.roll)","execution_count":14,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e4e8e8d4303a8fa862a4318c215440e5f2cb33ae"},"cell_type":"code","source":"df1 = df[(df.ip >= 126000) & (df.ip <= 126700)]\nplt.plot(df1.ip, df1.roll)","execution_count":15,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e090952cb41f9fe61879e981af8757d755ae19ae"},"cell_type":"code","source":"df1 = df[(df.ip >= 126400) & (df.ip <= 126500)]\nplt.plot(df1.ip, df1.roll)","execution_count":16,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5e3efeaa3fa354a9c6444951efcc5ff615f6e43e"},"cell_type":"code","source":"df1 = df[(df.ip >= 126415) & (df.ip <= 126425)]\nplt.plot(df1.ip, df1.roll)","execution_count":17,"outputs":[]},{"metadata":{"_uuid":"be6a03a756f572ee02a37e41770ba5576cd53833"},"cell_type":"markdown","source":"The rate starts to climb at 126420.\n\nLet's now look at where the ips present in test are compared to that split:"},{"metadata":{"trusted":true,"_uuid":"467b33dbc31d2e903e517a171c10a63c531498b5"},"cell_type":"code","source":"test.ip.max()","execution_count":9,"outputs":[]},{"metadata":{"_uuid":"c8460cfc47605b5700dcf59ee2056e7eb3b24a2e"},"cell_type":"markdown","source":"All the test ip are below the split!\n\nI thought this was significant, which is why I share it."},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"9b3c276a2676bd6edf2760de1849746b12240799"},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"6919e7b15d0b79e87f8464f4089446254755e123"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.4","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}