{"cells":[{"metadata":{"_uuid":"6fb86d10f5924f1944295a232fb1e7b346c5d046"},"cell_type":"markdown","source":"Have you taken closer look at target variable raw values? If not, then this small notebook most probably will give you some new insights ;)"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\ndf = pd.read_csv('../input/train.csv')\nprint('Train data:', df.shape)","execution_count":1,"outputs":[]},{"metadata":{"_uuid":"3cb60016ec3f1d436c57700dd208e2321738505e"},"cell_type":"markdown","source":"Let's take all entries with parent category 'Услуги' (services) and see, what do we have..."},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"df = df[df['parent_category_name'] == 'Услуги']\nprint('Train data for \"services\" parent catogory:', df.shape)\nprint('Distinct values for target variable in this category:', df['deal_probability'].value_counts().shape[0])\ndf['deal_probability'].value_counts()\n","execution_count":10,"outputs":[]},{"metadata":{"_uuid":"e46044ea596f46c368ad93206cfee30a3cea174b"},"cell_type":"markdown","source":"What we see:\n* There's only 23 distinct values out of 64385 in total - which is quite few.\n* Those values are actually all possible A/B, where A and B are integers <=8 - they are: 0, 1/8, 1/7, 1/6, 1/5, 2/8, ... , 6/7, 7/8, 1\n\nLet's dig deeper and group them by `param_2`..."},{"metadata":{"trusted":true,"_uuid":"f61efe1c864f2ac78f710d0c527006a494ee3af0"},"cell_type":"code","source":"dfg = df.groupby(['param_2','deal_probability'])['item_id'].count().reset_index()\ndfg.columns = ['param_2','deal_probability','count']\ndfg","execution_count":8,"outputs":[]},{"metadata":{"_uuid":"b41274dabd2f7ace8db5b9ee484276da57a7f939"},"cell_type":"markdown","source":"It's enough to take a quick look of previous table to see the pattern.\n\nE.g. `Cоздание и продвижение сайтов` has deal probabilities 0/6, 1/6, ... 5/6, 6/6. `Автосервис` has deal probabilities 0/4, 1/4, ... 4/4. And so on.\n\nSo my guess about target variable meaning here is that some model was used to predict deal probability (or it was obtained in some other way), then results were grouped by `param_2` (this parent category doesn't have `param_3` so this is lowest possible grouping here), then each group split into N bins according to their deal probabilities and finally indexes of bins were normalized by dividing by N - and these normalized bin indexes are given to us as target variable.\n\nHomework questions for you:\n* How are number of bins N chosen for each `param_2` group?\n* What were algorithms for obtaining target variable in other `parent_category_name` groups?\n* How can this insight been used to get better score? ;)"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.5","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}