{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","collapsed":true,"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":false},"cell_type":"markdown","source":"The description of the data files from the data page:\n\n* train.csv - Train data.\n* test.csv - Test data. Same schema as the train data, minus deal_probability.\n* train_active.csv - Supplemental data from ads that were displayed during the same period as train.csv. Same schema as the train data, minus deal_probability.\n* test_active.csv - Supplemental data from ads that were displayed during the same period as test.csv. Same schema as the train data, minus deal_probability.\n* periods_train.csv - Supplemental data showing the dates when the ads from train_active.csv were activated and when they where displayed.\n* periods_test.csv - Supplemental data showing the dates when the ads from test_active.csv were activated and when they where displayed. Same schema as periods_train.csv, except that the item ids map to an ad in test_active.csv.\n* train_jpg.zip - Images from the ads in train.csv.\n* test_jpg.zip - Images from the ads in test.csv.\n* sample_submission.csv - A sample submission in the correct format.\n\nLet us start with the train file."},{"metadata":{"trusted":true,"_uuid":"3f394ec7ffa6e0ca8e5f1758a37a2bff1eda9be7"},"cell_type":"code","source":"train_df = pd.read_csv(\"../input/train.csv\", parse_dates=[\"activation_date\"])\ntest_df = pd.read_csv(\"../input/test.csv\", parse_dates=[\"activation_date\"])\ntrain_df.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9ec2812799956c4a0d3f9162ca87e99d0956d434"},"cell_type":"markdown","source":"The train dataset description is as follows:\n\n* item_id - Ad id.\n* user_id - User id.\n* region - Ad region.\n* city - Ad city.\n* parent_category_name - Top level ad category as classified by Avito's ad model.\n* category_name - Fine grain ad category as classified by Avito's ad model.\n* param_1 - Optional parameter from Avito's ad model.\n* param_2 - Optional parameter from Avito's ad model.\n* param_3 - Optional parameter from Avito's ad model.\n* title - Ad title.\n* description - Ad description.\n* price - Ad price.\n* item_seq_number - Ad sequential number for user.\n* activation_date- Date ad was placed.\n* user_type - User type.\n* image - Id code of image. Ties to a jpg file in train_jpg. Not every ad has an image.\n* image_top_1 - Avito's classification code for the image.\n* deal_probability - The target variable. This is the likelihood that an ad actually sold something. It's not possible to verify every transaction with certainty, so this column's value can be any float from zero to one.\n\nSo deal probability is our target variable and  is a float value between 0 and 1 as per the data page. Let us have a look at it. "},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"9f2786faae88b675c91c6be5193fe8dad08bff72"},"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3cd8ddfe14602880f6722dc6bb385290bc2d7124"},"cell_type":"code","source":"sns.distplot(train_df.deal_probability.values)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"03bdfff50160971b00b28d3a51233b912fbb6500"},"cell_type":"code","source":"plt.figure(figsize=(8,6))\nplt.scatter( range(train_df.shape[0]),np.sort(train_df['deal_probability'].values))\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b8d8da3339b19690e17fe51c606c341837c429d1"},"cell_type":"code","source":"plt.figure(figsize=(12,8))\nsns.barplot(y=train_df.region, x=\"deal_probability\", data=train_df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6c856690453004d6c89b49d833e3d1692b990a43"},"cell_type":"code","source":"train_df.city.value_counts()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ab35ef856db5f0f93e770a83950e01d2ba056e67"},"cell_type":"code","source":"plt.figure(figsize=(12,8))\nsns.barplot(x=train_df.parent_category_name, y=\"deal_probability\", data=train_df)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"63b45fa531df779dfbd9ff5d1a2b6038fed660d5"},"cell_type":"code","source":"plt.figure(figsize=(12,8))\nsns.boxplot(x=\"parent_category_name\", y=\"deal_probability\", data=train_df)\nplt.ylabel('Deal probability', fontsize=12)\nplt.xlabel('Parent Category', fontsize=12)\nplt.title(\"Deal probability by parent category\", fontsize=14)\nplt.xticks(rotation='vertical')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b30dce0bb978b7b6ba8482a1830f20c40d71e954"},"cell_type":"code","source":"plt.figure(figsize=(15,5))\nsns.countplot(x=\"category_name\",data=train_df)\nplt.xticks(rotation='vertical')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"20a7e084ab14c80f02d4d117c57afa67d4338255"},"cell_type":"code","source":"plt.figure(figsize=(15,5))\nsns.countplot(x=\"user_type\",data=train_df)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"da7fe5b2e76d77bc7497bcb7cc3830fedf265b4f"},"cell_type":"markdown","source":"more in pipeline, if you like it please upvote for me.\n\nThank you :)"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}