{"cells":[{"metadata":{"_cell_guid":"7950e986-49fe-41bd-bd55-59b0730626e3","_uuid":"6c051e18bbcc88808e776859e31c6850c4fae940"},"cell_type":"markdown","source":"![](http://www.ewdn.com/wp-content/uploads/sites/6/2017/02/logo-avito.png)"},{"metadata":{"_cell_guid":"36fd7c6c-03b1-41b1-a0e2-66af0d17f000","_uuid":"a16bbfa5009c63231decbf538730958f4cf99c6c"},"cell_type":"markdown","source":"# More To Come. Stay Tuned. !!\nIf there are any suggestions/changes you would like to see in the Kernel please let me know :). Appreciate every ounce of help!\n\n**This notebook will always be a work in progress**. Please leave any comments about further improvements to the notebook! Any feedback or constructive criticism is greatly appreciated!. **If you like it or it helps you , you can upvote and/or leave a comment :).**\n\nI am using [Yandex Translate](https://translate.yandex.com/?lang=ru-en&text=Челябинск) for converting russian language to english language.\n"},{"metadata":{"_uuid":"88a729e47fe641234f4f1925b1b7c3bd226159b8"},"cell_type":"markdown","source":"- <a href='#intro'>1. Introduction</a>  \n- <a href='#rtd'>2. Retrieving the Data</a>\n     - <a href='#ll'>2.1 Load libraries</a>\n     - <a href='#rrtd'>2.2 Read the Data</a>\n- <a href='#god'>3. Glimpse of Data</a>\n     - <a href='#oot'>3.1 Overview of tables</a>\n     - <a href='#sootd'>3.2 Statistical overview of the Data</a>\n- <a href='#dp'>4. Data preparation</a>\n     - <a href='#cfmd'> 4.1 Check for missing data</a>\n- <a href='#de'>5. Data Exploration</a>\n    - <a href='#hadodp'>5.1 Histogram and distribution of deal probability</a>\n    - <a href='#hadoap'>5.2 Histogram and distribution of Ad price</a>\n    - <a href='#dodar'>5.3 Distribution of differnet Ad regions</a>\n    - <a href='#t'>5.4 Top 5</a> \n        - <a href='#ttat'>5.4.1 Top 5 Ad titles</a>\n        - <a href='#ttac'>5.4.2 Top 5 Ad city</a>\n        - <a href='#ttara'>5.4.3 Top 10 Ad regions</a>\n        - <a href='#tttlcalm'>5.4.4 Top 10 Fine grain ad category as classified by Avito's ad mode</a>\n        - <a href='#ttlcmsd'>5.4.5 Top 10 Top level ad category as classified by Avito's ad model</a>\n    - <a href='#pvdpp'>5.5 Price V.S. Deal probability</a>\n    - <a href='#5-6'>5.6 Deal probability V.S.  Price for regions</a>\n    - <a href='#dout'>5.7 Distribution of user type</a>\n    - <a href='#mdodadr'>5.8 Monthly distribution of Ad prices in different regions </a>\n    - <a href='#dorpdp'>5.9 Distribution of regions, per deal probability</a>\n    - <a href='#tkfad'>5.10 Top Keywords from Ad description</a>\n    - <a href='#toapaa'>5.10 Time Series Analysis</a>\n    - <a href='#5-11'>5.11 Ad sequential number for user V.S. deal probability</a>\n    - <a href='#5-12'>5.12 Deal probability V.S. Ad sequential number for user for regions</a>\n    - <a href='#5-13'>5.13 Number of words in description column</a>\n    - <a href='#vdcgv'>5.14 Venn Diagrams(Common Features values in training and test data)</a>\n    - <a href='#toapaa'>5.15 Time series Analysis</a>\n        - <a href='#toadp'>5.15.1 Trend of Ad price</a>\n        - <a href='#paetsd'>5.15.2 Price average every two days</a>\n        - <a href='#paeed'>5.15.3 Price average every day</a>\n        - <a href='#dapevtr'>5.15.4 deal probability average every two days</a>\n        - <a href='#dapevtttr'>5.15.5 Deal probability average every  days</a>\n        - <a href='#tnodawdjshhs'>5.15.6 Total number of days a Ad was dispalyed when it was posted on particular day</a>\n        - <a href='#fodsbdir'>5.15.7  frequency and pattern of ad activation date in train and test data</a>\n- <a href='#5-15'>6. Feature Engineering</a>    \n    - <a href='#5-15-1'>6.1 Features created from activation_date </a>\n    - <a href='#5-15-2'>6.2 Featrues created from Ad description</a>\n    - <a href='#5-15-3'>6.3 Converting Categorical featues to numericals features</a>\n    - <a href='#5-15-4'>6.4  Removing Unwanted features</a>\n- <a href='#5-16'>7 Pearson Correlation of features</a>\n- <a href='#5-17'>8 Feature importance via Gradient Boosting model</a>\n- <a href='#5-18'>9 Decision Tree Visualisation</a>\n- <a href='#bsc'>10. Brief summary and conclusion </a>"},{"metadata":{"_cell_guid":"41829d79-700f-4db4-863a-ea098b2d4c89","_uuid":"e194ee6814bf372fe5adc3ca089cf6b8590bdbb6"},"cell_type":"markdown","source":"# <a id='intro'>1. Introduction</a>  "},{"metadata":{"_cell_guid":"fbd66860-d6bf-41aa-9ebf-5c77c82d787d","_uuid":"9a513f5ea6c14fbe516d3b06f366c79fd1751f01"},"cell_type":"markdown","source":"Avito, Russia’s largest classified advertisements website, is deeply familiar with this problem. Sellers on their platform sometimes feel frustrated with both too little demand (indicating something is wrong with the product or the product listing) or too much demand (indicating a hot item with a good description was underpriced).\n\nIn their fourth Kaggle competition, Avito is challenging you to predict demand for an online advertisement based on its full description (title, description, images, etc.), its context (geographically where it was posted, similar ads already posted) and historical demand for similar ads in similar contexts. With this information, Avito can inform sellers on how to best optimize their listing and provide some indication of how much interest they should realistically expect to receive."},{"metadata":{"_cell_guid":"ea86dbe2-8db9-4486-82df-6bb56a4fbb83","_uuid":"746605b4858f9c058ea711397e5322648d82facb"},"cell_type":"markdown","source":"# <a id='rtd'>2. Retrieving the Data</a>\n## <a id='ll'>2.1 Load libraries</a>"},{"metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"import pandas as pd # package for high-performance, easy-to-use data structures and data analysis\nimport numpy as np # fundamental package for scientific computing with Python\nimport matplotlib\nimport matplotlib.pyplot as plt # for plotting\nimport seaborn as sns # for making plots with seaborn\ncolor = sns.color_palette()\nimport plotly.offline as py\npy.init_notebook_mode(connected=True)\nfrom plotly.offline import init_notebook_mode, iplot\ninit_notebook_mode(connected=True)\nimport plotly.graph_objs as go\nimport plotly.offline as offline\noffline.init_notebook_mode()\nimport plotly.tools as tls\nimport squarify\nfrom mpl_toolkits.basemap import Basemap\nfrom numpy import array\nfrom matplotlib import cm\n\nfrom sklearn import preprocessing\n# Supress unnecessary warnings so that presentation looks clean\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\n# Print all rows and columns\npd.set_option('display.max_columns', None)\npd.set_option('display.max_rows', None)\n\nfrom nltk.corpus import stopwords\nfrom textblob import TextBlob\nimport datetime as dt\nimport warnings\nimport string\nimport time\n# stop_words = []\nstop_words = list(set(stopwords.words('russian')))\nwarnings.filterwarnings('ignore')\npunctuation = string.punctuation\n\n# Plotting Decision tree\nfrom sklearn import tree\nfrom IPython.display import Image as PImage\nfrom subprocess import check_call\nfrom PIL import Image, ImageDraw, ImageFont\nimport re\n\n# Venn diagram\nfrom matplotlib_venn import venn2","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"a5bae02a-0957-45b3-9070-a5ca666d6abd","_uuid":"a65c908f4932c6b82bd55f5c6f0d950f505907f3"},"cell_type":"markdown","source":"## <a id='rrtd'>2.2 Read the Data</a>"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"print(\"Reading Data......\")\nperiods_test = pd.read_csv('../input/periods_test.csv', parse_dates=[\"activation_date\", \"date_from\", \"date_to\"])\nperiods_train = pd.read_csv('../input/../input/periods_train.csv', parse_dates=[\"activation_date\", \"date_from\", \"date_to\"])\ntest = pd.read_csv('../input/test.csv')\ntrain = pd.read_csv('../input/train.csv')\nprint(\"Reading Done....\")\n# train_active = pd.read_csv('../input/train_active.csv')\n# test_active = pd.read_csv('../input/test_active.csv')","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"2ad51a30-db88-4ea5-9430-5077212e81ca","_uuid":"24e8932ef3663d3f080e8cdabdb37885c9d23849","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"print(\"size of train data\", train.shape)\nprint(\"size of test data\", test.shape)\nprint(\"size of periods_train data\", periods_train.shape)\nprint(\"size of periods_test data\", periods_test.shape)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"94666ad1-d315-498f-bc52-150a427cfc6e","_uuid":"2726a6037513175e52687ad902fb256d30988d90"},"cell_type":"markdown","source":"# <a id='god'>3. Glimpse of Data</a>\n## <a id='oot'>3.1 Overview of tables</a>"},{"metadata":{"_cell_guid":"bbc19b3a-f020-44c4-a6c9-ff07a67420b9","_uuid":"fc2e96f7e72659b1fdddb7331284862e20c844f3"},"cell_type":"markdown","source":"**train data**"},{"metadata":{"scrolled":false,"_cell_guid":"29a5ad42-d921-45f4-a832-97a904882430","_uuid":"09bdd50325c8973ce2c88bf0320b965244ce9b87","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"train.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"0cf6d4b2-a1a3-4bfb-a56d-f65477e104dc","_uuid":"92322e701f45a17ead64327643e6f7c9bd8ba15f"},"cell_type":"markdown","source":"**test data**"},{"metadata":{"scrolled":false,"_cell_guid":"5129f680-1e36-4fa9-9c2a-038e5dfd9db0","_uuid":"3f2a48b586fe698253f75a921aa7dba8410345a0","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"test.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"cad87ca4-84a1-4466-9599-fced20f17dcb","_uuid":"7b9072f522438d986f7a2fa0cb0fee722497b492"},"cell_type":"markdown","source":"**periods train data**"},{"metadata":{"_cell_guid":"78de07a3-7e3c-4673-8148-9aa4d138f7c4","_uuid":"073f7f0141cd30be91c292439142a16201cbd0aa","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"periods_train.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"22da34ee-7b5f-4529-ba45-031acbfe8c99","_uuid":"7b032ea9c52ec61fce42049f98caa19739b723f8"},"cell_type":"markdown","source":"**periods test data**"},{"metadata":{"_cell_guid":"5f812f08-30cb-4d5d-8964-9ff194235f40","_uuid":"ff3756f4736b1cfc0f78c7413a52c672886a0c6b","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"periods_test.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"cfa24183-51bb-48ee-a416-dddacf613a8e","_uuid":"df89d3f3268e70e8553887563dd816ab5aeb8eae"},"cell_type":"markdown","source":"## <a id='sootd'>3.2 Statistical overview of the Data</a>"},{"metadata":{"_cell_guid":"74cedad3-bd9b-49b0-a8a0-9bd489eaedec","_uuid":"5a51b23064ee654ff0aafbf2328bc9f6b766cfda"},"cell_type":"markdown","source":"**Training Data some little info**"},{"metadata":{"_cell_guid":"536993aa-6af2-4105-808c-d3a1c58dcf18","_uuid":"f2bc9628f3772ed972976a4ae01ad3c5df29a637","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"train.info()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"450b201a-c95b-47ca-a8b3-03f29456452c","_uuid":"6c6e7c355de4a1809afb5ee56992bbb9f6f2751c"},"cell_type":"markdown","source":"**Little description of training data for numerical features**"},{"metadata":{"_cell_guid":"ed949735-9b31-4ce8-9560-f7c425f200a1","_uuid":"42973ca59fa1ab57ea951f85291056c4858a5909","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"train.describe()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"b389699e-4708-4f58-ac46-cb28096c938c","_uuid":"5e7247c2106760f37a8d6165deaa568b9fcb55ab"},"cell_type":"markdown","source":"**Little description of training data for categorical features**"},{"metadata":{"_cell_guid":"fa0e4f1e-8967-4b39-be6f-d8d5628175b6","_uuid":"c7b72f76685801710492155a95bdfaba79870b10","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"train.describe(include=[\"O\"])","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"849a8f47-e3cc-46c2-8ae4-b09f16320ddb","_uuid":"cd8030c8758270083fc80b493ed72e4aaed9a62f"},"cell_type":"markdown","source":"# <a id='dp'>4. Data preparation</a>\n ## <a id='cfmd'> 4.1 Check for missing data</a>"},{"metadata":{"_cell_guid":"2c026c36-08e0-4ba1-bbe3-03c022894fcb","_uuid":"78fa29a97a0a70839435942c5b3d74b13b4fc206"},"cell_type":"markdown","source":"**checking missing data in training data **"},{"metadata":{"_cell_guid":"58913142-5158-4961-af1f-fa8ed4f8ea7e","_uuid":"95cc0c66b501d5b2f741fd697008363dd8a1899f","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"# checking missing data in training data \ntotal = train.isnull().sum().sort_values(ascending = False)\npercent = (train.isnull().sum()/train.isnull().count()*100).sort_values(ascending = False)\nmissing_train_data  = pd.concat([total, percent], axis=1, keys=['Total', 'Percent'])\nmissing_train_data.head(10)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"f373ba0c-0789-437c-8537-3a198b7999a9","_uuid":"b13b7b4ddf4ff404e0fed3bb26670e42b4e4864a"},"cell_type":"markdown","source":"**checking missing data in periods training data **"},{"metadata":{"_cell_guid":"e8c47508-9a7e-480a-9bb9-26354ca50978","_uuid":"21169500b4ffa5e215fe5e971493af840d733c73","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"# checking missing data in periods data \ntotal = periods_train.isnull().sum().sort_values(ascending = False)\npercent = (periods_train.isnull().sum()/periods_train.isnull().count()*100).sort_values(ascending = False)\nmissing_train_data  = pd.concat([total, percent], axis=1, keys=['Total', 'Percent'])\nmissing_train_data.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"fe29a079-12b2-4546-9a21-07e75c771165","_uuid":"fde9902eaedb5b96b622a9a23edf0a9b321967d4"},"cell_type":"markdown","source":"# <a id='de'>5. Data Exploration</a>"},{"metadata":{"_cell_guid":"7548bbf0-8d2e-4965-9f86-f62a4d56a9ea","_uuid":"ffea4f3a719a6d77730086ebfbbb457db738343a"},"cell_type":"markdown","source":"## <a id='hadodp'>5.1 Histogram and distribution of deal probability</a>"},{"metadata":{"_cell_guid":"f0bb689a-9162-49ad-9c29-dc7bb37dfb6c","_uuid":"6f105fcba729925b2cbd1673ca8867d7b3afd9fb","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"plt.figure(figsize = (12, 8))\n\nsns.distplot(train['deal_probability'])\nplt.xlabel('likelihood that an ad actually sold something', fontsize=12)\nplt.title(\"Histogram of likelihood that an ad actually sold something\")\nplt.show() \nplt.figure(figsize = (12, 8))\nplt.scatter(range(train.shape[0]), np.sort(train.deal_probability.values))\nplt.xlabel('likelihood that an ad actually sold something', fontsize=12)\nplt.title(\"Distribution of likelihood that an ad actually sold something\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"a7a44f0a-32ea-4c56-8018-778d81481d09","_uuid":"70f1fd2a013e145d3720b89bbeff95751a6bafbe"},"cell_type":"markdown","source":"## <a id='hadoap'>5.2 Histogram and distribution of Ad price</a>"},{"metadata":{"_cell_guid":"3ad0bf3a-da6b-43d2-8029-046f4ea9804a","_uuid":"86548eea50a24a638915e5e9c80bfa430ff07a61","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"plt.figure(figsize = (12, 8))\n\nsns.distplot(train['price'].dropna())\nplt.xlabel('Ad price', fontsize=12)\nplt.title(\"Histogram of Ad price\")\nplt.show() \nplt.figure(figsize = (12, 8))\nplt.scatter(range(train.shape[0]), np.sort(train.price.values))\nplt.xlabel('Ad price', fontsize=12)\nplt.title(\"Distribution of Ad price\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"0bc08ec7-fa4f-441e-b757-74f4c23a1a8a","_uuid":"016254a728db26d08c07831fd3b2e010a3b59ef8","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"train['deal_class'] = train['deal_probability'].apply(lambda x: \">=0.5\" if x >=0.5 else \"<0.5\")\ntemp = train['deal_class'].value_counts()\nlabels = temp.index\nsizes = (temp / temp.sum())*100\ntrace = go.Pie(labels=labels, values=sizes, hoverinfo='label+percent')\nlayout = go.Layout(title='Distribution of deal class')\ndata = [trace]\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig)\n\ndel train['deal_class']","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"85ee71af-1f3d-41b5-89c4-60ca820e049e","_uuid":"156329242266ea4b593aa2d4cae19f615ad4862b"},"cell_type":"markdown","source":"* Approx 88 % training data having less than 0.5 deal probabilty. Remaining 12 % having probability more than or equal to 0.5."},{"metadata":{"_cell_guid":"dc93a468-e986-4cae-8baf-befab939b791","_uuid":"299932d940a398e265f65a10e12d66dff9c14de2","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"from io import StringIO\n\nconversion = StringIO(\"\"\"\nregion,region_english\nСвердловская область, Sverdlovsk oblast\nСамарская область, Samara oblast\nРостовская область, Rostov oblast\nТатарстан, Tatarstan\nВолгоградская область, Volgograd oblast\nНижегородская область, Nizhny Novgorod oblast\nПермский край, Perm Krai\nОренбургская область, Orenburg oblast\nХанты-Мансийский АО, Khanty-Mansi Autonomous Okrug\nТюменская область, Tyumen oblast\nБашкортостан, Bashkortostan\nКраснодарский край, Krasnodar Krai\nНовосибирская область, Novosibirsk oblast\nОмская область, Omsk oblast\nБелгородская область, Belgorod oblast\nЧелябинская область, Chelyabinsk oblast\nВоронежская область, Voronezh oblast\nКемеровская область, Kemerovo oblast\nСаратовская область, Saratov oblast\nВладимирская область, Vladimir oblast\nКалининградская область, Kaliningrad oblast\nКрасноярский край, Krasnoyarsk Krai\nЯрославская область, Yaroslavl oblast\nУдмуртия, Udmurtia\nАлтайский край, Altai Krai\nИркутская область, Irkutsk oblast\nСтавропольский край, Stavropol Krai\nТульская область, Tula oblast\n\"\"\")\n\nconversion = pd.read_csv(conversion)\ntrain = pd.merge(train, conversion, how=\"left\", on=\"region\")\n","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"37830567-3f43-4d9a-a5f3-379d835cbec3","_uuid":"af094915b01be5ce64434dcd5ad2319ef0d44afc"},"cell_type":"markdown","source":"## <a id='dodar'>5.3 Distribution of differnet Ad regions</a>"},{"metadata":{"_cell_guid":"457a0bc5-b5dc-4123-9bcc-eb8621dcba2a","_uuid":"129267d1cb7b6bd0865075001d50846648544ea6","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"temp = train['region_english'].value_counts()\nlabels = temp.index\nsizes = (temp / temp.sum())*100\ntrace = go.Pie(labels=labels, values=sizes, hoverinfo='label+percent')\nlayout = go.Layout(title='Distribution of differnet Ad regions')\ndata = [trace]\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"6668b1b2-aba3-4a7d-af7f-f0b73fa19baa","_uuid":"409177353ab12e9fff9b15091d90c681e0759518"},"cell_type":"markdown","source":"## <a id='t'>5.4 Top 5</a> "},{"metadata":{"_cell_guid":"060b72d1-0fb0-498c-a498-044f8961dee6","_uuid":"e7b98b7a3c82190e303dc88106534d2c470833d2"},"cell_type":"markdown","source":"## <a id='ttat'>5.4.1 Top 5 Ad titles</a>"},{"metadata":{"_cell_guid":"fb66c1e4-67bd-48c2-924c-9eff4f277ef7","_uuid":"9610f700f67907e39775945c26120deecd3b8c17","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"temp = train[\"title\"].value_counts().head(20)\nprint(\"Top 5 Ad titles :\\n\", temp.head(5))\nprint(\"Total Ad titles : \",len(train[\"title\"]))\ntrace = go.Bar(\n    x = temp.index,\n    y = temp.values,\n)\ndata = [trace]\nlayout = go.Layout(\n    title = \"Top Ad titles\",\n    xaxis=dict(\n        title='Ad title',\n        tickfont=dict(\n            size=14,\n            color='rgb(107, 107, 107)'\n        )\n    ),\n    yaxis=dict(\n        title='Count of Ad titles',\n        titlefont=dict(\n            size=16,\n            color='rgb(107, 107, 107)'\n        ),\n        tickfont=dict(\n            size=14,\n            color='rgb(107, 107, 107)'\n        )\n)\n)\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"a2b2f978-0b73-4ee1-9671-c36da4ade277","_uuid":"d0f58577679b938c2736253ef70d3be46c539047"},"cell_type":"markdown","source":"* ** Top 5 Ad titles are :**\n  *  Платье(Dress)\n  *  Туфли (Shoes) \n  *  Куртка(Jacket) \n  * Пальто (Coat) \n  * Джинсы(Jeans) \n   "},{"metadata":{"_cell_guid":"d8330b2d-5282-42c1-9dab-7f7f990e001f","_uuid":"8ac446112b748883291b259193fd1401cdf1a02b"},"cell_type":"markdown","source":"## <a id='ttac'>5.4.2 Top 5 Ad city</a>"},{"metadata":{"scrolled":true,"_cell_guid":"813b92c0-f82f-4323-be7d-4b8ea85bdd10","_uuid":"b3b8af0f9ba4075ccd7a1a2d07e1f52b113ce4b2","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"temp = train[\"city\"].value_counts().head(20)\nprint('Top 5 Ad cities :\\n', temp.head(5))\nprint(\"Total Ad cities : \",len(train[\"title\"]))\ntrace = go.Bar(\n    x = temp.index,\n    y = temp.values,\n)\ndata = [trace]\nlayout = go.Layout(\n    title = \"Top Ad city\",\n    xaxis=dict(\n        title='Ad title name',\n        tickfont=dict(\n            size=14,\n            color='rgb(107, 107, 107)'\n        )\n    ),\n    yaxis=dict(\n        title='Count of Ad cities',\n        titlefont=dict(\n            size=16,\n            color='rgb(107, 107, 107)'\n        ),\n        tickfont=dict(\n            size=14,\n            color='rgb(107, 107, 107)'\n        )\n)\n)\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"66293720-430d-4cd9-b85f-98be1b6a6f94","_uuid":"bc67c45efa0c14391706238c5966a17c4cad3ea2"},"cell_type":"markdown","source":"* **Top 5 Ad cities :**\n  * Краснодар (Krasnodar)         \n  * Екатеринбург (Yekaterinburg)\n  * Новосибирск (Novosibirsk) \n  * Ростов-на-Дону  (Rostov-on-don) \n  * Нижний Новгород  (Nizhny Novgorod) "},{"metadata":{"_cell_guid":"ba2841c3-3b2f-4ee5-b1ab-95fe898d3463","_uuid":"dd355c87503a4f19e1b60cda110ba392f8e87d7d"},"cell_type":"markdown","source":"## <a id='ttara'>5.4.3 Top 5 Ad regions</a>"},{"metadata":{"_cell_guid":"a82e07e5-0bd7-4abb-a3d7-43d3c353a128","_uuid":"6a7e94b3c1579f117352b0d27df679b3033e3a49","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"temp = train[\"region_english\"].value_counts().head(20)\nprint('Top 5 Ad regions :\\n',temp.head(5))\nprint(\"Total Ad regions : \",len(train[\"title\"]))\ntrace = go.Bar(\n    x = temp.index,\n    y = temp.values,\n)\ndata = [trace]\nlayout = go.Layout(\n    title = \"Top Ad regions\",\n    xaxis=dict(\n        title='Ad region name',\n        tickfont=dict(\n            size=14,\n            color='rgb(107, 107, 107)'\n        )\n    ),\n    yaxis=dict(\n        title='Count of Ad regions',\n        titlefont=dict(\n            size=16,\n            color='rgb(107, 107, 107)'\n        ),\n        tickfont=dict(\n            size=14,\n            color='rgb(107, 107, 107)'\n        )\n)\n)\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"0e06c10a-fe97-4a1a-a1b7-d058a0c3d359","_uuid":"604a836ae410c7e0e56870069679943590ce584f"},"cell_type":"markdown","source":"* ** Top 5 Ad regions :**\n  * Krasnodar Krai\n  * Sverdlovsk oblast  \n  * Rostov oblast \n  * Tatarstan \n  * Chelyabinsk oblast "},{"metadata":{"_cell_guid":"0443c72d-1f40-4401-97ff-759e5eccab73","_uuid":"11193dc0123809bfbaa773e5e3e5568a574edf67"},"cell_type":"markdown","source":"## <a id='tttlcalm'>5.4.4 Top 5 Fine grain ad category as classified by Avito's ad mode</a>"},{"metadata":{"_cell_guid":"102706d5-91b8-4e6d-a23f-bf82ba371bf8","_uuid":"d2859da42831f39bfde71518db4f99ccf4391216","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"#https://www.kaggle.com/sudalairajkumar/simple-exploration-notebook-avito\nconversion = StringIO(\"\"\"\ncategory_name,category_name_english\n\"Одежда, обувь, аксессуары\",\"Clothing, shoes, accessories\"\nДетская одежда и обувь,Children's clothing and shoes\nТовары для детей и игрушки,Children's products and toys\nКвартиры,Apartments\nТелефоны,Phones\nМебель и интерьер,Furniture and interior\nПредложение услуг,Offer services\nАвтомобили,Cars\nРемонт и строительство,Repair and construction\nБытовая техника,Appliances\nТовары для компьютера,Products for computer\n\"Дома, дачи, коттеджи\",\"Houses, villas, cottages\"\nКрасота и здоровье,Health and beauty\nАудио и видео,Audio and video\nСпорт и отдых,Sports and recreation\nКоллекционирование,Collecting\nОборудование для бизнеса,Equipment for business\nЗемельные участки,Land\nЧасы и украшения,Watches and jewelry\nКниги и журналы,Books and magazines\nСобаки,Dogs\n\"Игры, приставки и программы\",\"Games, consoles and software\"\nДругие животные,Other animals\nВелосипеды,Bikes\nНоутбуки,Laptops\nКошки,Cats\nГрузовики и спецтехника,Trucks and buses\nПосуда и товары для кухни,Tableware and goods for kitchen\nРастения,Plants\nПланшеты и электронные книги,Tablets and e-books\nТовары для животных,Pet products\nКомнаты,Room\nФототехника,Photo\nКоммерческая недвижимость,Commercial property\nГаражи и машиноместа,Garages and Parking spaces\nМузыкальные инструменты,Musical instruments\nОргтехника и расходники,Office equipment and consumables\nПтицы,Birds\nПродукты питания,Food\nМотоциклы и мототехника,Motorcycles and bikes\nНастольные компьютеры,Desktop computers\nАквариум,Aquarium\nОхота и рыбалка,Hunting and fishing\nБилеты и путешествия,Tickets and travel\nВодный транспорт,Water transport\nГотовый бизнес,Ready business\nНедвижимость за рубежом,Property abroad\n\"\"\")\n\nconversion = pd.read_csv(conversion)\ntrain = pd.merge(train, conversion, on=\"category_name\", how=\"left\")","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"1e036033-f807-4e84-8d21-7de178f8d3a2","_uuid":"e9f908099c76e46367287a28c0777fbeebec0cbe","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"temp = train[\"category_name_english\"].value_counts().head(20)\nprint(\"Top 5 Fine grain ad category as classified by Avito's ad mode : \\n\", temp.head(5))\nprint(\"Total Fine grain ad category as classified by Avito's ad mode : \",len(train[\"title\"]))\ntrace = go.Bar(\n    x = temp.index,\n    y = temp.values,\n)\ndata = [trace]\nlayout = go.Layout(\n    title = \"Top Fine grain ad category as classified by Avito's ad mode\",\n    xaxis=dict(\n        title='Fine grain ad category as classified by Avitos ad mode',\n        tickfont=dict(\n            size=14,\n            color='rgb(107, 107, 107)'\n        )\n    ),\n    yaxis=dict(\n        title='Count of Fine grain ad category',\n        titlefont=dict(\n            size=16,\n            color='rgb(107, 107, 107)'\n        ),\n        tickfont=dict(\n            size=14,\n            color='rgb(107, 107, 107)'\n        )\n)\n)\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"7fa98c52-7a0e-479b-9830-e2af2934b084","_uuid":"aae9d5862e69cbcc928a9650e32a5f38d6856cae"},"cell_type":"markdown","source":"* **Top 5 Fine grain ad category as classified by Avito's ad mode :**\n * Clothing, shoes and accessories \n * Children clothing and shoes \n * Childrens product and toys  \n * Apartments \n * Phones "},{"metadata":{"_cell_guid":"87cf6c2b-2667-4eae-b4fd-e23276c33a61","_uuid":"62538b034df1e364d1a33caf35bb39f01e493a56","_kg_hide-input":false},"cell_type":"markdown","source":"## <a id='ttlcmsd'>5.4.5 Top 5 Top level ad category as classified by Avito's ad model</a>"},{"metadata":{"_cell_guid":"54cb6760-7cba-494f-a890-3fe5d6847c30","_uuid":"6f0b727d23efb98cf796aba3e41d2d08a7cfc284","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"conversion = StringIO(\"\"\"\nparent_category_name,parent_category_name_english\nЛичные вещи,Personal belongings\nДля дома и дачи,For the home and garden\nБытовая электроника,Consumer electronics\nНедвижимость,Real estate\nХобби и отдых,Hobbies & leisure\nТранспорт,Transport\nУслуги,Services\nЖивотные,Animals\nДля бизнеса,For business\n\"\"\")\n\nconversion = pd.read_csv(conversion)\ntrain = pd.merge(train, conversion, on=\"parent_category_name\", how=\"left\")","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"81845a2f-73cb-4e32-974b-4ee05de701bb","_uuid":"2ba429cfc3d8dc752086f684be1782d96e41365c","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"temp = train[\"parent_category_name_english\"].value_counts()\nprint(\"Total Top level ad category as classified by Avito's ad model : \",len(train[\"title\"]))\ntrace = go.Bar(\n    x = temp.index,\n    y = (temp / temp.sum())*100,\n)\ndata = [trace]\nlayout = go.Layout(\n    title = \"Top level ad category as classified by Avito's ad model\",\n    xaxis=dict(\n        title='Top level ad category as classified by Avitos ad model',\n        tickfont=dict(\n            size=14,\n            color='rgb(107, 107, 107)'\n        )\n    ),\n    yaxis=dict(\n        title='Count of Top level ad category in %',\n        titlefont=dict(\n            size=16,\n            color='rgb(107, 107, 107)'\n        ),\n        tickfont=dict(\n            size=14,\n            color='rgb(107, 107, 107)'\n        )\n)\n)\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"4f8078f4-bac7-4cc4-8188-02650bfe03b0","_uuid":"d2832d2b4db6889f85828c921ddd65a023e2e068"},"cell_type":"markdown","source":"*** Top 5 Top level ad category as classified by Avito's ad model :**\n* Personal belongings - 46 %\n* For the home and garden - 12 %\n* Consumer electronics - 12 %\n* Real estate - 10 %\n* Hobbies & leisure - 6 %"},{"metadata":{"_cell_guid":"f13e58f3-092f-412f-9a46-a9e74ea55076","_uuid":"e5a1cff5af4e51a3f5a46e12341aeb7d7653095e"},"cell_type":"markdown","source":"## <a id='pvdpp'>5.5 Price V.S. Deal probability</a>"},{"metadata":{"_cell_guid":"266fc091-530b-4dd3-855f-89e96675ee8a","_uuid":"a36550f3f177f94c58dd85b2cf961fe6802a0a55","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"plt.figure(figsize=(15,6))\nplt.scatter(np.log(train.price), train.deal_probability)\nplt.xlabel('Ad price')\nplt.ylabel('deal probability')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"42f81024-f38c-4444-b536-eb321a973821","_uuid":"a46739e71d5606d3aa3b1e56da11591c5dd8b2ab"},"cell_type":"markdown","source":"## <a id='5-6'>5.6 Deal probability V.S.  Price for regions</a>"},{"metadata":{"_cell_guid":"b7ebc438-0626-4dcb-887b-9564420a9f3f","_uuid":"760080bccb4163fd25bcfd162259171518b4e4c2","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"populated_states = train[:100]\n\ndata = [go.Scatter(\n    y = populated_states['deal_probability'],\n    x = populated_states['price'],\n    mode='markers+text',\n    marker=dict(\n        size= np.log(populated_states.price) - 2,\n        color=populated_states['deal_probability'],\n        colorscale='Portland',\n        showscale=True\n    ),\n    text=populated_states['region_english'],\n    textposition=[\"top center\"]\n)]\nlayout = go.Layout(\n    title='Deal probability V.S.  Price for regions',\n    xaxis= dict(title='Ad price'),\n    yaxis=dict(title='deal probability')\n)\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"50b6f2a8-aff7-43e2-8a97-75c8162b369b","_uuid":"33a61da9c224a814d77f7dc5bca0159a6a254453"},"cell_type":"markdown","source":"## <a id='dout'>5.7 Distribution of user type</a>"},{"metadata":{"_cell_guid":"3d0e2567-faa1-4093-8212-0f56124a6036","_uuid":"c5cd9b9295870b3f1a658fab27af807e06027e91","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"temp = train['user_type'].value_counts()\nlabels = temp.index\nsizes = (temp / temp.sum())*100\ntrace = go.Pie(labels=labels, values=sizes, hoverinfo='label+percent')\nlayout = go.Layout(title='Distribution of user type')\ndata = [trace]\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"d210cb76-47f0-4421-83a1-8a3a4b543e1f","_uuid":"d0d09412cb75f94d2f1b310dc08398dd7323fdb2"},"cell_type":"markdown","source":"* **Distribution of user types :**\n  * Private users constitutes 71.6 % data\n  * Comapny users constitutes 23.1 % data\n  * Shop users constitutes 5.35 % data"},{"metadata":{"_cell_guid":"eb0241a4-e61a-4e60-9bbc-b707a5cb4790","_uuid":"468157b0a8a9a4ea852114ac6520b8b121cb6f8e"},"cell_type":"markdown","source":"## <a id='mdodadr'>5.8 Monthly distribution of Ad prices in different regions </a>"},{"metadata":{"_cell_guid":"b3d1532c-d2b9-482c-819c-7eb0aaae4887","_uuid":"4a1cafe52d57a5a1d8ce2cc676e741b41dda6a7a","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"train['activation_date'] = pd.to_datetime(train['activation_date'])\ntrain['month'] = train.activation_date.dt.month\npr = train.groupby(['region_english', 'month'])['price'].mean().unstack()\n#pr = pr.sort_values([12], ascending=False)\nf, ax = plt.subplots(figsize=(15, 20)) \npr = pr.fillna(0)\ntemp = sns.heatmap(pr, cmap='Reds')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"ca8fbaa7-5f90-4988-93cd-71b16ae4f838","_uuid":"36ca70133a81577d6ce9eff6cf8e5b62b25cfe41"},"cell_type":"markdown","source":"* Highest Ad prices is in **Irkutsk oblast** region followed by **Krasnodar Krai** region."},{"metadata":{"_cell_guid":"03a263da-62a0-4b29-889b-46da600c1d57","_uuid":"ac06536487c186aeb05a24a570ec4b258ab1dbb0"},"cell_type":"markdown","source":"## <a id='dorpdp'>5.9 Distribution of regions, per deal probability</a>"},{"metadata":{"_cell_guid":"27e51a9f-d14d-4d7f-93c9-ebd9bbc037de","_uuid":"bc3e0aa0d45b4192464fc03ecf658ca8861ec535","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"plt.figure(figsize=(17,8))\nboxplot = sns.boxplot(x=\"region_english\", y=\"deal_probability\", data=train)\nboxplot.set(xlabel='', ylabel='')\nplt.title('Distribution of regions, per deal probability', fontsize=17)\nplt.xticks(rotation=80, fontsize=17)\nplt.yticks(fontsize=17)\nplt.xlabel('region name')\nplt.ylabel('deal probability')\nplt.show()\n\n_, ax = plt.subplots(figsize=(17, 8))\nsns.violinplot(ax=ax, x=\"region_english\", y=\"deal_probability\", data=train)\nplt.title('Distribution of regions, per deal probability', fontsize=17)\nplt.xticks(rotation=80, fontsize=17)\nplt.yticks(fontsize=17)\nplt.xlabel('region name')\nplt.ylabel('deal probability')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"5ad2030d-15b9-467f-99f0-1ba7781cfdfb","_uuid":"d65b7964de77ec62cb5867892f1e7bc2940f1591"},"cell_type":"markdown","source":"## <a id='tkfad'>5.10 Top Keywords from Ad description</a>"},{"metadata":{"_cell_guid":"1a08ac80-a2ce-46fd-8838-c7af20738571","_uuid":"61ad8cbf3ec10f0891c39e50e9d3a16cf62e46f4","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"from wordcloud import WordCloud, STOPWORDS\nnames = test[\"description\"][~pd.isnull(test[\"description\"])]\n#print(names)\nwordcloud = WordCloud(max_font_size=50, width=600, height=300).generate(' '.join(names))\nplt.figure(figsize=(15,8))\nplt.imshow(wordcloud)\nplt.title(\"Wordcloud for Ad description\", fontsize=35)\nplt.axis(\"off\")\nplt.show() ","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"d70538b7-c3f9-47c6-b305-192a2d32b39c","_uuid":"a9b0467c8176c80afb0e408281d929a596ce7147"},"cell_type":"markdown","source":"## <a id='5-11'>5.11 Ad sequential number for user V.S. deal probability</a>"},{"metadata":{"_cell_guid":"028f9991-3773-4f0f-a486-0a3debc4efb3","_uuid":"b1ee571bbcc15fd2e5aac9b71ac510d5f21c3eeb","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"plt.figure(figsize=(15,6))\nplt.scatter(train.item_seq_number, train.deal_probability)\nplt.xlabel('Ad sequential number for user')\nplt.ylabel('deal probability')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"2714ab06-d373-40f1-94ea-8c880663b720","_uuid":"73529954e11d43ccf4b245639a20e444638bd8a7"},"cell_type":"markdown","source":"## <a id='5-12'>5.12 Deal probability V.S. Ad sequential number for user for regions</a>"},{"metadata":{"_kg_hide-input":true,"trusted":true,"collapsed":true,"_uuid":"f91b41dded568d400ce869c4b5ae265825e278f1"},"cell_type":"code","source":"populated_regions = train[:50]\n\ndata = [go.Scatter(\n    y = populated_regions['deal_probability'],\n    x = populated_regions['item_seq_number'],\n    mode='markers+text',\n    marker=dict(\n        size= np.log(populated_regions.price) - 2,\n        color=populated_regions['deal_probability'],\n        colorscale='Portland',\n        showscale=True\n    ),\n    text=populated_regions['region_english'],\n    textposition=[\"top center\"]\n)]\nlayout = go.Layout(\n    title='Deal probability V.S. Ad sequential number for user for regions',\n    xaxis= dict(title='Ad sequential number for user'),\n    yaxis=dict(title='deal probability')\n)\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"e14b88b9-a8bc-4da3-8155-cc4c72ef1223","_uuid":"fc9f8eb6f204cae89cd454c3390bf9e086b54a67"},"cell_type":"markdown","source":"## <a id='5-13'>5.13 Number of words in description column</a>"},{"metadata":{"_cell_guid":"55582e44-3861-4ee9-86ad-afff24dc5617","_uuid":"d538c4d920831774539f3d93bb4dabd75db04405","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"train[\"description\"].fillna(\"NA\", inplace=True)\ntrain[\"desc_numOfWords\"] = train[\"description\"].apply(lambda x: len(x.split()))\ntemp = train[\"desc_numOfWords\"].value_counts().head(80)\ntrace = go.Bar(\n    x = temp.index,\n    y = temp.values,\n)\ndata = [trace]\nlayout = go.Layout(\n    title = \"Number of words in description column\",\n)\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig)\ndel train[\"desc_numOfWords\"]","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"e3ceebd8-d4c7-47c3-bdf9-55b101af72f1","_uuid":"206d37f8f83b0d2028e421a18a4543637accc05c"},"cell_type":"markdown","source":"## <a id='vdcgv'>5.14 Venn Diagram(Common Features values in training and test data)</a>"},{"metadata":{"_cell_guid":"0218228a-2272-4c4b-a3ba-2a933c4ad92c","_uuid":"e1a9934d4f2ccfc175f3acbbccd08799924475e6"},"cell_type":"markdown","source":"* **A Venn diagram uses overlapping circles or other shapes to illustrate the logical relationships between two or more sets of items. Often, they serve to graphically organize things, highlighting how the items are similar and different.**"},{"metadata":{"_kg_hide-input":true,"trusted":true,"collapsed":true,"_uuid":"57381a12c7adfe7c70fddd1e5639d0bc25bd0120"},"cell_type":"code","source":"plt.figure(figsize=(23,13))\n\nplt.subplot(321)\nvenn2([set(train.item_id.unique()), set(test.item_id.unique())], set_labels = ('Train set', 'Test set') )\nplt.title(\"Common items in training and test data\", fontsize=15)\n#plt.show()\n\n#plt.figure(figsize=(15,8))\nplt.subplot(322)\nvenn2([set(train.user_id.unique()), set(test.user_id.unique())], set_labels = ('Train set', 'Test set') )\nplt.title(\"Common users in training and test data\", fontsize=15)\n#plt.show()\n\n#plt.figure(figsize=(15,8))\nplt.subplot(323)\nvenn2([set(train.title.unique()), set(test.title.unique())], set_labels = ('Train set', 'Test set') )\nplt.title(\"Common Ad titles in training and test data\", fontsize=15)\n#plt.show()\n\n#plt.figure(figsize=(15,8))\nplt.subplot(324)\nvenn2([set(train.item_seq_number.unique()), set(test.item_seq_number.unique())], set_labels = ('Train set', 'Test set') )\nplt.title(\"Common Ad sequential number of user in training and test data\", fontsize=15)\n#plt.show()\n\n#plt.figure(figsize=(15,8))\nplt.subplot(325)\nvenn2([set(train.activation_date.unique()), set(test.activation_date.unique())], set_labels = ('Train set', 'Test set') )\nplt.title(\"Common Ad Activation dates in training and test data\", fontsize=15)\n#plt.show()\n\n#plt.figure(figsize=(15,8))\nplt.subplot(326)\nvenn2([set(train.image.unique()), set(test.image.unique())], set_labels = ('Train set', 'Test set') )\nplt.title(\"Common images in training and test data\", fontsize=15)\n\nplt.subplots_adjust(wspace = 0.5, hspace = 0.5,\n                    top = 0.9)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"0fa37430-cf26-4291-b37a-118b70170b84","_uuid":"a89cdf7816f9fd753330b1cfa422bde227e26a3e"},"cell_type":"markdown","source":"* **Common items in training and test data :** 0\n* **Common users in training and test data :** Approx. 68k\n* **CommonAd titles in training and test data :** Approx. 64K\n* **CommonAd sequential number of users in training and test data :** Approx. 10k\n* **CommonAd activation dates in training and test data :** 0\n* **Common images in training and test data :** 1"},{"metadata":{"_cell_guid":"a1d66eb2-7f94-4b9d-a6f4-abd21c0730de","_uuid":"23d4e3a9bbf16bcb467c529d17b9b525ccace41b"},"cell_type":"markdown","source":" ## <a id='toapaa'>5.15 Time series Analysis</a>"},{"metadata":{"_cell_guid":"53b7f664-b25a-46db-9cb0-27d154bf9fd5","_uuid":"87492339692d508548e927fa1a0761cfbc3bf755"},"cell_type":"markdown","source":"## <a id='toadp'>5.15.1 Trend of Ad price</a>"},{"metadata":{"_cell_guid":"2c8af19d-7dd7-4555-84df-595d2b5c2ade","_uuid":"65032f308c348be3b239939a9093764de4aef58c","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"train.posted_time = pd.to_datetime(train['activation_date'])\ntrain.index = pd.to_datetime(train['activation_date'])\nplt.figure(figsize = (12, 8))\nax = train['price'].resample('w').sum().plot()\n#ax = kiva_loans_data['funded_amount'].resample('w').sum().plot()\nax.set_ylabel('Price')\nax.set_xlabel('day-month')\nax.set_xlim((pd.to_datetime(train['activation_date'].min()), \n             pd.to_datetime(train['activation_date'].max())))\nax.legend([\"Ad Price\"])\nplt.title('Trend of Ad price')","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"31dfdaf8-d7b0-4f14-9757-f2fe9a4d07cc","_uuid":"c4f6cc32c156f705c8a45596949eb4ef54b32d0f","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"df = pd.read_csv('../input/train.csv').set_index('activation_date')\ndf.index = pd.to_datetime(df.index)\n#df.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"25b07952-66e4-43e2-a800-acbfea439ede","_uuid":"d9aa964815cd7da5f16a363a3aa9957acad1bfaa","collapsed":true,"trusted":false},"cell_type":"code","source":"df.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"c331bb41-4a12-46e7-ab76-93b03591e683","_uuid":"41d3e46b3d16f0e705b10f5f53e919be82182fce"},"cell_type":"markdown","source":"## <a id='paetsd'>5.15.2 Price average every two days</a>"},{"metadata":{"_cell_guid":"02d55bfe-5e29-44d6-8587-e209e355c7ee","_uuid":"a3cb4d256c95a89d77f19895f29dabbe83836e5a","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"df[\"price\"].resample(\"2D\").apply([np.mean]).plot()\nplt.title(\"Price average every two days\")\nplt.ylabel(\"Price\")","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"ca7ce1e5-a96e-451a-859f-e27d8477665c","_uuid":"05cc8c1e35b964bd40364ee92920ecb8c092fbcd"},"cell_type":"markdown","source":"## <a id='paeed'>5.15.3 Price average every day</a>"},{"metadata":{"_cell_guid":"2c4b9db7-adff-41d0-a43c-8857efecd504","_uuid":"2ac5f983a9a4af7c8942244f45a3b145e72c82bc","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"df[\"price\"].resample(\"D\").apply([np.mean]).plot()\nplt.title(\"Price average every day\")\nplt.ylabel(\"Price\")","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"b7088fde-9605-455d-9a02-f5d0b6f1cd2f","_uuid":"6fa18f467a4da03413e66c6c2a02eb5979bccf38"},"cell_type":"markdown","source":"## <a id='dapevtr'>5.15.4 Deal probability average every two days</a>"},{"metadata":{"_cell_guid":"3fe9921c-bff9-43ca-ba16-9c0084b356d3","_uuid":"80112edc6961f64532fba8eae8f9a826aa471313","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"df[\"deal_probability\"].resample(\"2D\").apply([np.mean]).plot()\nplt.title(\"deal probability average every two days\")\nplt.ylabel(\"deal probability\")","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"d0f8067a-dca5-4d02-b7ed-6589443d10fa","_uuid":"5b0fec5c571705f9dd1013e1e23b880388779244"},"cell_type":"markdown","source":"## <a id='dapevtttr'>5.15.5 Deal probability average every  days</a>"},{"metadata":{"_cell_guid":"1ff8bd18-5512-43cf-95b6-ed8af1c971ab","_uuid":"71538f757b609c057328336312d566307dde8dd7","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"df[\"deal_probability\"].resample(\"D\").apply([np.mean]).plot()\nplt.title(\"deal probability average every day\")\nplt.ylabel(\"deal probability\")","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"54123ce3-0605-48d5-9fd2-e15158b4438b","_uuid":"639536fd3f43a988a1a83281917928ff6cceb8e9"},"cell_type":"markdown","source":"## <a id='tnodawdjshhs'>5.15.6 Total number of days a Ad was dispalyed when it was posted on particular day</a>"},{"metadata":{"_cell_guid":"95295268-8f4d-48f5-9418-d38723786c9c","_uuid":"c185b02c704244bcb998a8543db26efb76ce4a28","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"periods_train['total_days'] = periods_train['date_to'] - periods_train['date_from']\nperiods_test['total_days'] = periods_test['date_to'] - periods_test['date_from']","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"99f422f5-7268-4ca4-9094-ced8710a1e2f","_uuid":"fe250148f29f16a49f3c89f05bc4b1dc8b66cd1d","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"periods_train.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"6629a639-abc1-4b95-828a-c1b96733077c","_uuid":"a5bbc4e54f2edae41487bcf37e4706625d84a1c8","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"periods_train['total_days_value'] = periods_train['total_days'].dt.days\n#periods_train['total_days'], _ = zip(*periods_train['total_days'].map(lambda x: x.split(' ')))","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"c26101e4-28a0-4a6d-a6fc-187fd8e0de42","_uuid":"c13e196664dd444f8406bfd31e92fc81c6fbf64f","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"periods_train.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"752db9e3-887d-456c-b4a2-ec7cbef673d2","_uuid":"f96f6ece999c4af6f4e3870c53bea735da000cb6","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"#df = periods_train.set_index('activation_date')\nperiods_train.index = pd.to_datetime(periods_train['activation_date'])\nplt.figure(figsize = (12, 8))\nax = periods_train['total_days_value'].resample('w').sum().plot()\n#ax = kiva_loans_data['funded_amount'].resample('w').sum().plot()\nax.set_ylabel('Total days')\nax.set_xlabel('day-month')\nax.set_xlim((pd.to_datetime(periods_train['activation_date'].min()), \n             pd.to_datetime(periods_train['activation_date'].max())))\nax.legend([\"Total number of days\"])\nplt.title('Total number of days a Ad was dispalyed when it was posted on particular day')","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"082a8c55-5fa1-4dff-9802-e9014fb4fe91","_uuid":"994cf36248927062971a9c475ece1b2ce11f4926"},"cell_type":"markdown","source":"## <a id='fodsbdir'>5.15.7  frequency and pattern of ad activation date in train and test data</a>\n"},{"metadata":{"_cell_guid":"318630c1-83f1-4b98-a67e-38a5161e51ce","_uuid":"f93d2442dd5eba6da20dd4213b8eafa7bd94df00","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"temp = train[\"activation_date\"].value_counts()\ntemp1 = test[\"activation_date\"].value_counts()\ntrace0 = go.Bar(\n    x = temp.index,\n    y = (temp / temp.sum())*100,\n    name = 'Ad activation dates in training data'\n)\ntrace1 = go.Bar(\n    x = temp1.index,\n    y = (temp1 / temp1.sum())*100,\n    name = 'Ad activation dates in test data'\n)\ndata = [trace0, trace1]\nlayout = go.Layout(\n    title = \"frequency and pattern of ad activation date in train and test data\",\n    xaxis=dict(\n        title='Ad Activation Date',\n        tickfont=dict(\n            size=14,\n            color='rgb(107, 107, 107)'\n        )\n    ),\n    yaxis=dict(\n        title='frequency in %',\n        titlefont=dict(\n            size=16,\n            color='rgb(107, 107, 107)'\n        ),\n        tickfont=dict(\n            size=14,\n            color='rgb(107, 107, 107)'\n        )\n)\n)\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"d9cb4719-f218-422a-bfd4-efd3cea76298","_uuid":"e8f305741cb0c92da41e216096c972bc4beb9b6f"},"cell_type":"markdown","source":"* Most Ad activation date range is from 15 march 2017 to 28 March 2017 in training data and in test data the range is 12 April 2017 to 19 April 2017"},{"metadata":{"_cell_guid":"6cb68066-bcf2-450c-a9fe-0c6bcfb3fbd2","_uuid":"f22ff88cd0df2397a69a28a307be6a7f41ae4008"},"cell_type":"markdown","source":"# <a id='5-15'>6. Feature Engineering</a>"},{"metadata":{"_cell_guid":"6a609935-5365-41f4-b363-6faf51e38da9","_uuid":"50c13fba09708b30442e54fb22d3a1a740e54708"},"cell_type":"markdown","source":"## <a id='5-15-1'>6.1 Features created from activation_date </a>"},{"metadata":{"_cell_guid":"9c763d77-5d0a-469c-9e81-7fb28654d213","_uuid":"8df00c08272a5094dfdaecbe7081f69a68e8a3f2"},"cell_type":"markdown","source":"**Total 4 features created from activation_date :**\n\n - month\n - weekday\n - month_day\n - year_day"},{"metadata":{"_cell_guid":"5f7ef7d4-62a7-4bc1-b1dd-7b322f0144fe","_uuid":"cc7f5ee41086eb68f3dc7c631c90356c9d68b1c6","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"train['activation_date'] = pd.to_datetime(train['activation_date'])\ntest['activation_date'] = pd.to_datetime(test['activation_date'])\n\n# ***** Train data *******\ntrain[\"month\"] = train[\"activation_date\"].dt.month\ntrain['weekday'] = train['activation_date'].dt.weekday\ntrain[\"month_Day\"] = train['activation_date'].dt.day\ntrain[\"year_Day\"] = train['activation_date'].dt.dayofyear\n\n# ***** Test data *******\ntest[\"month\"] = test[\"activation_date\"].dt.month\ntest['weekday'] = test['activation_date'].dt.weekday\ntest[\"month_Day\"] = test['activation_date'].dt.day\ntest[\"year_Day\"] = test['activation_date'].dt.dayofyear \n","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"04cd4099-2c5b-4bff-bc6b-0a73468c285d","_uuid":"6c51edef15354e618a360db7bd671590696e55d5"},"cell_type":"markdown","source":"## <a id='5-15-2'>6.2 Featrues created from Ad description</a>"},{"metadata":{"_cell_guid":"aae0e8f9-a487-48bd-80f5-5ef9c1184b1b","_uuid":"36dc1f2ec9319cd743c7e24dd970fca08801a264"},"cell_type":"markdown","source":"Here we created total  features :\n* **char_count**\n* **word_count**\n* **word_density**\n* **punctuation_count**\n* **title_word_count**\n* **upper_case_word_count**\n* **stopword_count**"},{"metadata":{"_cell_guid":"ccd42558-60a4-4541-9d4c-282f0b8addc9","_uuid":"38e655864b2fd826ab33b474bca8b225b1b6c4c7","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"start_time=time.time()\n\n# *******train data *********\n#train['d_length'] = train['description'].apply(lambda x: len(str(x))) \ntrain['char_count'] = train['description'].apply(len)\ntrain['word_count'] = train['description'].apply(lambda x: len(x.split()))\ntrain['word_density'] = train['char_count'] / (train['word_count']+1)\ntrain['punctuation_count'] = train['description'].apply(lambda x: len(\"\".join(_ for _ in x if _ in punctuation))) \ntrain['title_word_count'] = train['description'].apply(lambda x: len([wrd for wrd in x.split() if wrd.istitle()]))\ntrain['upper_case_word_count'] = train['description'].apply(lambda x: len([wrd for wrd in x.split() if wrd.isupper()]))\ntrain['stopword_count'] = train['description'].apply(lambda x: len([wrd for wrd in x.split() if wrd.lower() in stop_words]))\n\n# *******test data *********\n#test['d_length'] = test['description'].apply(lambda x: len(str(x))) \ntest['char_count'] = test['description'].apply(len)\ntest['word_count'] = test['description'].apply(lambda x: len(x.split()))\ntest['word_density'] = test['char_count'] / (test['word_count']+1)\ntest['punctuation_count'] = test['description'].apply(lambda x: len(\"\".join(_ for _ in x if _ in punctuation))) \ntest['title_word_count'] = test['description'].apply(lambda x: len([wrd for wrd in x.split() if wrd.istitle()]))\ntest['upper_case_word_count'] = test['description'].apply(lambda x: len([wrd for wrd in x.split() if wrd.isupper()]))\ntest['stopword_count'] = test['description'].apply(lambda x: len([wrd for wrd in x.split() if wrd.lower() in stop_words]))\n\nend_time=time.time()\nprint(\"total time in the cuurent cell \",end_time-start_time,\"s\")","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"4e41a25a-74c5-4922-86fa-a8ab8d2959fd","_uuid":"a406cd6e8f00ab33e0bfb5d3ab19d3578c8abb55","collapsed":true},"cell_type":"markdown","source":"## <a id='5-15-3'>6.3 Converting Categorical featues to numericals features</a>"},{"metadata":{"_cell_guid":"58ce2b9a-d772-4206-9e43-426eef941590","_uuid":"ddb1371ef9ab4974d8ccaf4b2a02585c6135bcac"},"cell_type":"markdown","source":"**Categorical variables are :**\n   - region\n   - city\n   - parent_category_name\n   - category_name\n   - user_type\n   - param_1\n   - param_2\n   - param_3"},{"metadata":{"_cell_guid":"44368063-65aa-49e4-b580-4eb7b2c5ec67","_uuid":"524053db93600ad2116655e4a6f8bb3dba7d2a52","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"# Label encode the categorical variables\ncat_vars = [\"region\", \"city\", \"parent_category_name\", \"category_name\", \"user_type\", \"param_1\", \"param_2\", \"param_3\"]\nfor col in cat_vars:\n    lb = preprocessing.LabelEncoder()\n    lb.fit(list(train[col].values.astype('str')) + list(test[col].values.astype('str')))\n    train[col] = lb.transform(list(train[col].values.astype('str')))\n    test[col] = lb.transform(list(test[col].values.astype('str')))","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"b923c4b6-7b02-4408-8a2b-5245ceed4931","_uuid":"1b0e261d790082df710b2cf7d0b4df3265f310dc"},"cell_type":"markdown","source":"## <a id='5-15-4'>6.4  Removing Unwanted features</a>"},{"metadata":{"_cell_guid":"618ee72e-0b34-4ab0-8ba5-006c98e8f682","_uuid":"dbb0ab06ada8ab2e789ecca36f50b4a194e14d0e"},"cell_type":"markdown","source":"- item_id\n- user_id\n- title\n- description\n- activation_date\n- image\n- region_en\n- parent_category_name_en\n- category_name_en\n- deal_probability"},{"metadata":{"_cell_guid":"07828b3c-1552-46b6-98a9-c34706fa4407","_uuid":"38ccd0d92c38a65f1fcb2a490b3ba8299cb0d74e","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"# Target and ID variables \nyt = train[\"deal_probability\"].values\ntest_id = test[\"item_id\"].values\n\ncols_to_drop = [\"item_id\", \"user_id\", \"title\", \"description\", \"activation_date\", \"image\"]\ntrain_df = train.drop(cols_to_drop + [\"region_english\", \"parent_category_name_english\", \"category_name_english\", \"deal_probability\"], axis=1)\n","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"d6d51c62-884a-4e05-95d9-a4b6df18721c","_uuid":"95157de3ba342e136825c8f3e2537c3dfa24ce26"},"cell_type":"markdown","source":"# <a id='5-16'>7. Pearson Correlation of features</a>"},{"metadata":{"_cell_guid":"7ce21065-9866-470e-9737-eacb1ee0d4e2","_uuid":"2ed0853a30935ea53bc088bc7ce55e7c76721fc6","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"data = [\n    go.Heatmap(\n        z= train_df.corr().values,\n        x=train_df.columns.values,\n        y=train_df.columns.values,\n        colorscale='Viridis',\n        reversescale = False,\n        text = True ,\n        opacity = 1.0 )\n]\n\nlayout = go.Layout(\n    title='Pearson Correlation of features',\n    xaxis = dict(ticks='', nticks=36),\n    yaxis = dict(ticks='' ),\n    width = 900, height = 700)\n\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig, filename='labelled-heatmap')","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"23644382-0f02-4bed-87bf-34a1cc29f467","_uuid":"144d4e0c035fe44e6895474048691bdb31cf5d48"},"cell_type":"markdown","source":"# <a id='5-17'>8. Feature importance via Gradient Boosting model</a>"},{"metadata":{"_cell_guid":"ef8a6457-395d-4d25-8355-5fe64068f082","_uuid":"aecea828674f07ac111cbf3eaf952ec94ab33360","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"start_time=time.time()\n\ntrain_df.fillna(-1, inplace = True)\nfrom sklearn.ensemble import GradientBoostingRegressor\ngb = GradientBoostingRegressor()\ngb.fit(train_df, yt)\nfeatures = train_df.columns.values\n\nend_time=time.time()\nprint(\"total time in the cuurent cell \",end_time-start_time,\"s\")","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"f2a7d218-e95f-4e59-8d3b-dc2abd717f6d","_uuid":"a85d383f81bb0dc1506ee6f2eedb6f80e92e6c69","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"# Scatter plot \ntrace = go.Scatter(\n    y = gb.feature_importances_,\n    x = features,\n    mode='markers',\n    marker=dict(\n        sizemode = 'diameter',\n        sizeref = 1,\n        size = 13,\n        color = gb.feature_importances_,\n        colorscale='Portland',\n        showscale=True\n    ),\n    text = features\n)\ndata = [trace]\n\nlayout= go.Layout(\n    autosize= True,\n    title= 'Gradient Boosting Machine Feature Importance',\n    hovermode= 'closest',\n     xaxis= dict(\n         ticklen= 5,\n         showgrid=False,\n        zeroline=False,\n        showline=False\n     ),\n    yaxis=dict(\n        title= 'Feature Importance',\n        showgrid=False,\n        zeroline=False,\n        ticklen= 5,\n        gridwidth= 2\n    ),\n    showlegend= False\n)\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"da1626fa-8f39-47cc-b5d4-511454706a84","_uuid":"40caa26c34fa1a0a3c4fc964699fe6d77a237ac0"},"cell_type":"markdown","source":"# <a id='5-18'>9. Decision Tree Visualisation</a>"},{"metadata":{"_cell_guid":"8d915a26-efd3-47d3-954f-335a21ed97b0","_uuid":"42d0ac1f6a7cc29bb34fd625dffef50289cb460f","collapsed":true,"_kg_hide-input":true,"trusted":false},"cell_type":"code","source":"decision_tree = tree.DecisionTreeRegressor(max_depth = 3)\ndecision_tree.fit(train_df, yt)\n\n# Export our trained model as a .dot file\nwith open(\"tree1.dot\", 'w') as f:\n     f = tree.export_graphviz(decision_tree,\n                              out_file=f,\n                              max_depth = 4,\n                              impurity = False,\n                              feature_names = train_df.columns.values,\n                              class_names = ['No', 'Yes'],\n                              rounded = True,\n                              filled= True )\n        \n#Convert .dot to .png to allow display in web notebook\ncheck_call(['dot','-Tpng','tree1.dot','-o','tree1.png'])\n\n# Annotating chart with PIL\nimg = Image.open(\"tree1.png\")\ndraw = ImageDraw.Draw(img)\nimg.save('sample-out.png')\nPImage(\"sample-out.png\",)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"f8ffbcec-e749-4ccf-a707-992f5c22bd3c","_uuid":"e2a45e824f44f539dbd4da6342c130553eb4ba4a"},"cell_type":"markdown","source":"# <a id='bsc'>10. Brief summary and conclusion :</a>"},{"metadata":{"_uuid":"24ff87f96e8950498a4e2ea731fa5761b7fea91c"},"cell_type":"markdown","source":"\n\n\n* Approx 88 % training data having less than 0.5 deal probabilty. Remaining 12 % having probability more than or equal to 0.5.\n* ** Top 5 Ad titles are :**\n  *  Платье(Dress)\n  *  Туфли (Shoes) \n  *  Куртка(Jacket) \n  * Пальто (Coat) \n  * Джинсы(Jeans) \n* **Top 5 Ad cities :**\n  * Краснодар (Krasnodar)         \n  * Екатеринбург (Yekaterinburg)\n  * Новосибирск (Novosibirsk) \n  * Ростов-на-Дону  (Rostov-on-don) \n  * Нижний Новгород  (Nizhny Novgorod) \n* ** Top 5 Ad regions :**\n  * Krasnodar Krai\n  * Sverdlovsk oblast  \n  * Rostov oblast \n  * Tatarstan \n  * Chelyabinsk oblast \n* **Top 5 Fine grain ad category as classified by Avito's ad mode :**\n  * Clothing, shoes and accessories \n  * Children clothing and shoes \n  * Childrens product and toys  \n  * Apartments \n  * Phones \n* ** Top 5 Top level ad category as classified by Avito's ad model :**\n   * Personal belongings - 46 %\n   * For the home and garden - 12 %\n   * Consumer electronics - 12 %\n   * Real estate - 10 %\n   * Hobbies & leisure - 6 %\n* **Distribution of user types :**\n  * Private users constitutes 71.6 % data\n  * Comapny users constitutes 23.1 % data\n  * Shop users constitutes 5.35 % data\n* Common Features values in training and test data : \n   * **Common items in training and test data :** 0\n   * **Common users in training and test data :** Approx. 68k\n   * **CommonAd titles in training and test data :** Approx. 64K\n   * **CommonAd sequential number of users in training and test data :** Approx. 10k\n   * **CommonAd activation dates in training and test data :** 0\n   * **Common images in training and test data :** 1\n* Highest Ad prices is in **Irkutsk oblast** region followed by **Krasnodar Krai** region.  \n* Most Ad activation date range is from 15 march 2017 to 28 March 2017 in training data and in test data the range is 12 April 2017 to 19 April 2017"},{"metadata":{"_cell_guid":"e5a208fb-fa52-429a-8009-bd4d58369df1","_uuid":"a909ea107188ee9950c04a12fe758aad531ab8e4"},"cell_type":"markdown","source":"## This is only a brief summary if want more details please go through my Notebook."},{"metadata":{"_cell_guid":"8d12503a-1b07-4cf9-a10a-dfe4af87ebce","_uuid":"5b97ec94ac9088ccf43665c9216ff67826161937"},"cell_type":"markdown","source":"# More to come. Stayed Tuned !!"}],"metadata":{"language_info":{"name":"python","version":"3.6.5","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"}},"nbformat":4,"nbformat_minor":1}