{"cells":[{"metadata":{"_cell_guid":"e9400bef-70e4-4451-8c33-c3b78595d81c","_uuid":"8ef47a8800abd5411dac954cad4a57f2734e92a8"},"cell_type":"markdown","source":"# Welcome to my new Kernel. \n\nLink to my all kernels  <a href=\"https://www.kaggle.com/kabure/kernels\">HERE </a>\n\n## I will try a deep understanding of Avito's Dataset.\n\n<i> * English is not my first language, so sorry for any error </i>\n\n## I will try to answer some questions. \n**Some of them is like: **<br>\nWe have null values in any column?  <br>\nAre the price normal distributed? <br>\nAll ads came from the same category?  <br>\nAll city's are equal?  <br>\nWe have a equal distribuition of params? <br>\nWhich region are most frequent? <br>\nThe most frequent regions, have the same price and deal probability? <br>\nThe price and deal probability are correlated?  <br>\nAre all features important to we predict the deal probability? <br>\n\n\n"},{"metadata":{"_cell_guid":"42dd48b1-8baa-49a3-b4fa-3c930c7e8ae6","_uuid":"f728ba7ab39770c9b1487e09a113643bf500430c","collapsed":true},"cell_type":"markdown","source":"<i>English is not my first language, so sorry for any error. </i>"},{"metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","collapsed":true,"trusted":true},"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport seaborn as sns","execution_count":1,"outputs":[]},{"metadata":{"_cell_guid":"e4c615bc-6e38-4ffb-8671-c0607ee49e5a","_uuid":"550b62e9dfab89b170725e31ce4a60a444ea8a6f"},"cell_type":"markdown","source":"<h2>Importing datasets</h2>"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"df_train = pd.read_csv(\"../input/train.csv\", parse_dates=['activation_date'])\n\nprint(\"Shape train: \", df_train.shape)","execution_count":2,"outputs":[]},{"metadata":{"_cell_guid":"63119678-79eb-46b5-b009-6240b434c581","_uuid":"79c5950be42c0ad449719827d5d29bb8735fa52a","trusted":true},"cell_type":"code","source":"df_train.info()","execution_count":3,"outputs":[]},{"metadata":{"_cell_guid":"9b88156a-35a0-451d-81cb-f7ae965c364d","_uuid":"7014e167b63dbce28a6dd096b224d3adebcb1b6d"},"cell_type":"markdown","source":"<h2>Looking percentual of null to each column</h2>"},{"metadata":{"_cell_guid":"b3bdf1b1-d157-447c-a666-bc4aeca232e2","_uuid":"6d153454ea9bc68205583d6c5552ef047239bb0e","_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"is_null = df_train.isnull().sum() / len(df_train) * 100\nprint(\"NaN values in train Dataset\")\nprint(is_null[is_null > 0].sort_values(ascending=False))","execution_count":4,"outputs":[]},{"metadata":{"_cell_guid":"8c5c0bea-82e8-4398-8166-ae70f5bee9b5","_uuid":"04a1443bbd2730c7a9f831ab868d37416ead98d1"},"cell_type":"markdown","source":"<h2>Visualing the distribuition of the unique values by each feature</h2>"},{"metadata":{"_cell_guid":"f8fa47a3-0849-4fa6-8079-72f635e0eafa","_uuid":"e6e4b47b817e77cd96ad93a2edd0e99e41bca528","_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16, 5))\n\ncols = df_train.columns\n\nuniques = [len(df_train[col].unique()) for col in cols]\nsns.set(font_scale=1.2)\nax = sns.barplot(cols, uniques, palette='hls', log=True)\nax.set(xlabel='Feature', ylabel='log(unique count)', title='Number of unique per feature')\nfor p, uniq in zip(ax.patches, uniques):\n    height = p.get_height()\n    ax.text(p.get_x()+p.get_width()/2.,\n            height + 10,\n            uniq,\n            ha=\"center\") \nax.set_xticklabels(ax.get_xticklabels(),rotation=45)\nplt.show()","execution_count":5,"outputs":[]},{"metadata":{"_cell_guid":"5db6caca-1ded-4ebc-a581-f2cec7aa2364","_uuid":"6df981bc574cd9fa465eade0f1ae85458163129b"},"cell_type":"markdown","source":"<h2>Converting some words from russian to English</h2>"},{"metadata":{"_cell_guid":"081cd9c3-b2b5-4c5b-8e17-0d85d00faea6","_uuid":"929f6d23e40f1d30afa3405e1888c1cc2fe6d033","_kg_hide-input":true,"collapsed":true,"trusted":true},"cell_type":"code","source":"# I have copied some dictionary's from another fellows Kernels.\nparent_category_name_map = {\"Личные вещи\" : \"Personal belongings\",\n                            \"Для дома и дачи\" : \"For the home and garden\",\n                            \"Бытовая электроника\" : \"Consumer electronics\",\n                            \"Недвижимость\" : \"Real estate\",\n                            \"Хобби и отдых\" : \"Hobbies & leisure\",\n                            \"Транспорт\" : \"Transport\",\n                            \"Услуги\" : \"Services\",\n                            \"Животные\" : \"Animals\",\n                            \"Для бизнеса\" : \"For business\"}\n\nregion_map = {\"Свердловская область\" : \"Sverdlovsk oblast\",\n            \"Самарская область\" : \"Samara oblast\",\n            \"Ростовская область\" : \"Rostov oblast\",\n            \"Татарстан\" : \"Tatarstan\",\n            \"Волгоградская область\" : \"Volgograd oblast\",\n            \"Нижегородская область\" : \"Nizhny Novgorod oblast\",\n            \"Пермский край\" : \"Perm Krai\",\n            \"Оренбургская область\" : \"Orenburg oblast\",\n            \"Ханты-Мансийский АО\" : \"Khanty-Mansi Autonomous Okrug\",\n            \"Тюменская область\" : \"Tyumen oblast\",\n            \"Башкортостан\" : \"Bashkortostan\",\n            \"Краснодарский край\" : \"Krasnodar Krai\",\n            \"Новосибирская область\" : \"Novosibirsk oblast\",\n            \"Омская область\" : \"Omsk oblast\",\n            \"Белгородская область\" : \"Belgorod oblast\",\n            \"Челябинская область\" : \"Chelyabinsk oblast\",\n            \"Воронежская область\" : \"Voronezh oblast\",\n            \"Кемеровская область\" : \"Kemerovo oblast\",\n            \"Саратовская область\" : \"Saratov oblast\",\n            \"Владимирская область\" : \"Vladimir oblast\",\n            \"Калининградская область\" : \"Kaliningrad oblast\",\n            \"Красноярский край\" : \"Krasnoyarsk Krai\",\n            \"Ярославская область\" : \"Yaroslavl oblast\",\n            \"Удмуртия\" : \"Udmurtia\",\n            \"Алтайский край\" : \"Altai Krai\",\n            \"Иркутская область\" : \"Irkutsk oblast\",\n            \"Ставропольский край\" : \"Stavropol Krai\",\n            \"Тульская область\" : \"Tula oblast\"}\n\n\ncategory_map = {\"Одежда, обувь, аксессуары\":\"Clothing, shoes, accessories\",\n\"Детская одежда и обувь\":\"Children's clothing and shoes\",\n\"Товары для детей и игрушки\":\"Children's products and toys\",\n\"Квартиры\":\"Apartments\",\n\"Телефоны\":\"Phones\",\n\"Мебель и интерьер\":\"Furniture and interior\",\n\"Предложение услуг\":\"Offer services\",\n\"Автомобили\":\"Cars\",\n\"Ремонт и строительство\":\"Repair and construction\",\n\"Бытовая техника\":\"Appliances\",\n\"Товары для компьютера\":\"Products for computer\",\n\"Дома, дачи, коттеджи\":\"Houses, villas, cottages\",\n\"Красота и здоровье\":\"Health and beauty\",\n\"Аудио и видео\":\"Audio and video\",\n\"Спорт и отдых\":\"Sports and recreation\",\n\"Коллекционирование\":\"Collecting\",\n\"Оборудование для бизнеса\":\"Equipment for business\",\n\"Земельные участки\":\"Land\",\n\"Часы и украшения\":\"Watches and jewelry\",\n\"Книги и журналы\":\"Books and magazines\",\n\"Собаки\":\"Dogs\",\n\"Игры, приставки и программы\":\"Games, consoles and software\",\n\"Другие животные\":\"Other animals\",\n\"Велосипеды\":\"Bikes\",\n\"Ноутбуки\":\"Laptops\",\n\"Кошки\":\"Cats\",\n\"Грузовики и спецтехника\":\"Trucks and buses\",\n\"Посуда и товары для кухни\":\"Tableware and goods for kitchen\",\n\"Растения\":\"Plants\",\n\"Планшеты и электронные книги\":\"Tablets and e-books\",\n\"Товары для животных\":\"Pet products\",\n\"Комнаты\":\"Room\",\n\"Фототехника\":\"Photo\",\n\"Коммерческая недвижимость\":\"Commercial property\",\n\"Гаражи и машиноместа\":\"Garages and Parking spaces\",\n\"Музыкальные инструменты\":\"Musical instruments\",\n\"Оргтехника и расходники\":\"Office equipment and consumables\",\n\"Птицы\":\"Birds\",\n\"Продукты питания\":\"Food\",\n\"Мотоциклы и мототехника\":\"Motorcycles and bikes\",\n\"Настольные компьютеры\":\"Desktop computers\",\n\"Аквариум\":\"Aquarium\",\n\"Охота и рыбалка\":\"Hunting and fishing\",\n\"Билеты и путешествия\":\"Tickets and travel\",\n\"Водный транспорт\":\"Water transport\",\n\"Готовый бизнес\":\"Ready business\",\n\"Недвижимость за рубежом\":\"Property abroad\"}\n\nparams_top35_map = {'Женская одежда':\"Women's clothing\",\n                    'Для девочек':'For girls',\n                    'Для мальчиков':'For boys',\n                    'Продам':'Selling',\n                    'С пробегом':'With mileage',\n                    'Аксессуары':'Accessories',\n                    'Мужская одежда':\"Men's Clothing\",\n                    'Другое':'Other','Игрушки':'Toys',\n                    'Детские коляски':'Baby carriages', \n                    'Сдам':'Rent',\n                    'Ремонт, строительство':'Repair, construction',\n                    'Стройматериалы':'Building materials',\n                    'iPhone':'iPhone',\n                    'Кровати, диваны и кресла':'Beds, sofas and armchairs',\n                    'Инструменты':'Instruments',\n                    'Для кухни':'For kitchen',\n                    'Комплектующие':'Accessories',\n                    'Детская мебель':\"Children's furniture\",\n                    'Шкафы и комоды':'Cabinets and chests of drawers',\n                    'Приборы и аксессуары':'Devices and accessories',\n                    'Для дома':'For home',\n                    'Транспорт, перевозки':'Transport, transportation',\n                    'Товары для кормления':'Feeding products',\n                    'Samsung':'Samsung',\n                    'Сниму':'Hire',\n                    'Книги':'Books',\n                    'Телевизоры и проекторы':'Televisions and projectors',\n                    'Велосипеды и самокаты':'Bicycles and scooters',\n                    'Предметы интерьера, искусство':'Interior items, art',\n                    'Другая':'Other','Косметика':'Cosmetics',\n                    'Постельные принадлежности':'Bed dress',\n                    'С/х животные' :'Farm animals','Столы и стулья':'Tables and chairs'}","execution_count":6,"outputs":[]},{"metadata":{"_cell_guid":"e10a532a-14be-4a01-bbf3-72cf5991300d","_uuid":"fa5f614c6de9af82a8458bf3e90ab99489039ce7","collapsed":true,"trusted":true},"cell_type":"code","source":"df_train['region_en'] = df_train['region'].apply(lambda x : region_map[x])\ndf_train['parent_category_name_en'] = df_train['parent_category_name'].apply(lambda x : parent_category_name_map[x])\ndf_train['category_name_en'] = df_train['category_name'].apply(lambda x : category_map[x])\n\ndel df_train['region']\ndel df_train['parent_category_name']\ndel df_train['category_name']","execution_count":7,"outputs":[]},{"metadata":{"_cell_guid":"94df53c0-91ab-4734-aef8-1689ba5f9c82","_uuid":"346f702816fb3407f07c846eab8bd78b2fbbc163"},"cell_type":"markdown","source":"<h2>Let's look how the data appears</h2>"},{"metadata":{"_cell_guid":"4ca3e604-f3f7-4ed6-b751-f3b62fe75364","_uuid":"c5ec8208f607780e89ca73b5567bb6e0dc5cad23","trusted":true},"cell_type":"code","source":"df_train.head()","execution_count":8,"outputs":[]},{"metadata":{"_cell_guid":"d89e9a48-8d61-44b7-9b60-3d8c54489acb","_uuid":"a9715980e7cd211e178900e08f3e946f7983f6d4"},"cell_type":"markdown","source":""},{"metadata":{"_cell_guid":"43b8d7f4-93a7-40a4-ab37-ad0e2db5b7ac","_uuid":"7b5657190becc9d2145b69d11fc8b37b54e0ebfd"},"cell_type":"markdown","source":"Let's start exploring the distribuition of price and deal probability that will be one of the most important elements to guide our exploration"},{"metadata":{"_cell_guid":"071c0511-ba19-4534-b266-f403ba8f7aef","_uuid":"f02412b79b1cdad8fb5b49855ec9fef23006c334"},"cell_type":"markdown","source":"<h2>I will start taking a look at  the Deal Probability that is the target feature </h2>"},{"metadata":{"_cell_guid":"c06abce7-cb5d-4cac-9b88-28586895329b","_uuid":"db9c3077273eb59686c07eac9de77bd6b374af32","_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(14,5))\n\nplt.subplot(1,2,1)\nax = sns.distplot(df_train[\"deal_probability\"].values, bins=100, kde=False)\nax.set_xlabel('Deal Probility', fontsize=15)\nax.set_ylabel('Deal Probility', fontsize=15)\nax.set_title(\"Deal Probability Histogram\", fontsize=20)\n\nplt.subplot(1,2,2)\nplt.scatter(range(df_train.shape[0]), np.sort(df_train['deal_probability'].values))\nplt.xlabel('Index', fontsize=15)\nplt.ylabel('Deal Probability', fontsize=15)\nplt.title(\"Deal Probability Distribution\", fontsize=20)\nplt.xticks(rotation=45)\nplt.show()\n","execution_count":9,"outputs":[]},{"metadata":{"_cell_guid":"813bfc61-b347-42ec-994c-80209faffbdf","_uuid":"b38a8fac9b0c34e5a2160be1e7f1cedfe566c286"},"cell_type":"markdown","source":"To a better understanding of the  dataset and a further exploration, I will create a new feature that will be the deal probability categorical. \n"},{"metadata":{"_cell_guid":"3ea3db6f-7040-48d4-a5d4-ad6b2af885e0","_uuid":"229f9ec4a69be25440cfca586a01ec0b01d2935f"},"cell_type":"markdown","source":"<h2>Seting the deal probability categorical</h2>"},{"metadata":{"_cell_guid":"0b18e27a-9ae6-40dd-919c-f72193d0a580","_uuid":"b96c347f06984370b55a4e170fc223b7b98e2b5d","_kg_hide-input":true,"collapsed":true,"trusted":true},"cell_type":"code","source":"interval = (-0.99, .02, .05, .1, .15, .2, .35, .50, .70,.85,2)\ncats = ['0 -.02%', '.02%-.05%', '.05-.10', '.10-.15', '.15-.20', '.20-.35', '.35-.50', '.50-.70', '.70-.85','.85+']\n\ndf_train[\"deal_prob_cat\"] = pd.cut(df_train.deal_probability, interval, labels=cats)","execution_count":10,"outputs":[]},{"metadata":{"_cell_guid":"992008df-11c1-424f-a369-a1c309f9efa8","_uuid":"66fc9bdf9190fd1f0e8227824a5515e727299f48"},"cell_type":"markdown","source":"<h2>Now, let's do a count of this categorical feature to see the distribuition of each value </h2>"},{"metadata":{"_cell_guid":"c79083d6-c0b8-44b4-918e-e1f01bde5bdc","_uuid":"92c1f34519fab6f4fc51f7220e4560dd76dd9dc3","_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"prob_cat_percent = df_train[\"deal_prob_cat\"].value_counts() / len(df_train['deal_probability'])* 100\n\nplt.figure(figsize=(12,5))\ng = sns.barplot(prob_cat_percent.index, prob_cat_percent.values)\ng.set_xlabel('The Deal Probability Categorical Dist',fontsize=16)\ng.set_ylabel('% of frequency',fontsize=16)\ng.set_title('Deal Probability Categorical % Frequency',fontsize= 20)\nplt.show()","execution_count":11,"outputs":[]},{"metadata":{"_cell_guid":"0baec035-637f-407f-a36a-c9b78a61391f","_uuid":"f612d174bd7966a6bdbe559b978a4decbc02abaa"},"cell_type":"markdown","source":"We can see that that more than 65% of our target have from zero to .02%of Deal Probability"},{"metadata":{"_cell_guid":"eaa9a098-3fb2-4170-a811-7ff87c963bee","_uuid":"b71ba20fbaa959ae638651ddb213773b6b9faaa4"},"cell_type":"markdown","source":"<h2>Deal Probability x Price Log Distribuition </h2>"},{"metadata":{"_cell_guid":"13556be4-c234-49c5-9f6e-b5f7485f71a3","_uuid":"936663ce5a3f3ee3b98e76b90634dbba96aef75f"},"cell_type":"markdown","source":"Now, let's use the categorical probability to verify if the price have the same behavior to each value in our categories'"},{"metadata":{"_cell_guid":"b70ca3ce-2b5b-46e6-9aa4-d097dc5c9e08","_uuid":"b0743a8b89b70ad839a0fb883b718787f25516f3","_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"df_train['price_log'] = np.log(df_train['price'] + 1)\n\nplt.figure(figsize=(12,5))\n\ng = sns.boxplot(x='deal_prob_cat', y='price_log', data=df_train)\ng.set_xlabel('The Deal Probability Categorical Dist',fontsize=16)\ng.set_ylabel('Price Log Dist',fontsize=16)\ng.set_title('Looking the Price Log of each deal_prob_cat',fontsize= 20)\n\nplt.show()","execution_count":12,"outputs":[]},{"metadata":{"_cell_guid":"86d7685b-3be7-4596-894e-e1b18f8e29ba","_uuid":"de051d2476842e0f2061fd0e71759b043b6505e2"},"cell_type":"markdown","source":"We can see an interesting behavior of price on the lowest deal probabilities category that is different of another two lowest values. Also, we can clearly see that low probabilities have a lowest prices and is the most frequent deal probability interval in the dataset.  Later we will explore this further."},{"metadata":{"_cell_guid":"377a7b0a-c488-4533-8d10-4ea9e86b59da","_uuid":"c0c4910ab5e51993015d98b11e384550335db8d0","_kg_hide-input":true},"cell_type":"markdown","source":"<h2>Let's take a first look at Price Feature</h2>"},{"metadata":{"_cell_guid":"62e547a0-7761-449f-b920-9827dbdac210","_uuid":"f6cf829d1398576f1da338782718eadae77dd0b5","_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(12,5))\n\nplt.subplot(1,2,1)\ng = sns.distplot(np.log(df_train['price'].dropna() + 1))\ng.set_xlabel('Price Log', fontsize=15)\ng.set_ylabel('Probility', fontsize=15)\ng.set_title(\"Price Histogram\", fontsize=20)\n\nplt.subplot(1,2,2)\nplt.scatter(range(df_train.shape[0]), np.sort(np.log(df_train['price']+1).values))\nplt.xlabel('Index ', fontsize=15)\nplt.ylabel('Price Log', fontsize=15)\nplt.title(\"Price Log Distribution\", fontsize=20)\nplt.xticks(rotation=45)\nplt.show()\n\nplt.show()","execution_count":13,"outputs":[]},{"metadata":{"_cell_guid":"8fe56fdf-1a0e-42ef-8de1-8750c1abc440","_uuid":"597cd3e66984ba7cc4f0b190b0942546354909b1"},"cell_type":"markdown","source":"We can see that a great part of our data is under 10 Price log "},{"metadata":{"_cell_guid":"e8b393db-bfbf-44f1-8c74-e407beb4aef3","_uuid":"19175f5207a91a78bde4e5b8d6bc2b375bf5d962"},"cell_type":"markdown","source":"<h2>Taking a look at user type feature</h2>"},{"metadata":{"_cell_guid":"99ee2c2a-b1c4-43fd-8e76-a9d1ef5a29a7","_uuid":"0ad8fe7c9cdadd2bab3086c1bce0498e20681661","_kg_hide-input":true,"scrolled":false,"trusted":true},"cell_type":"code","source":"print(\"User Type % Proportion\")\nprint(round(df_train['user_type'].value_counts() / len(df_train) * 100, 2))\n\nplt.figure(figsize=(16,5))\n\nplt.subplot(1,3,1)\ng = sns.countplot(x='user_type', data=df_train, )\ng.set_xlabel('User Type',fontsize=16)\ng.set_ylabel('Count',fontsize=16)\ng.set_title('User Type Count',fontsize= 20)\n\nplt.subplot(1,3,2)\ng1 = sns.boxplot(x='user_type', y='deal_probability', data=df_train)\ng1.set_xlabel('User Type',fontsize=16)\ng1.set_ylabel('Deal probability',fontsize=16)\ng1.set_title('User Type x Prob Dist',fontsize= 20)\n\nplt.subplot(1,3,3)\ng1 = sns.boxplot(x='user_type', y='price_log', data=df_train)\ng1.set_xlabel('User Type',fontsize=16)\ng1.set_ylabel('Price Log',fontsize=16)\ng1.set_title('User Type x Price Log',fontsize= 20)\n\nplt.show()","execution_count":14,"outputs":[]},{"metadata":{"_cell_guid":"e7d99215-2927-44db-99bb-1a191c6177c6","_uuid":"e1a0d4956c8ba9622f9a008fab17ee5d2bee14d9"},"cell_type":"markdown","source":"Highest frequent user type  in Ad is Privat. And it also have a high Deal Probability and lowest price values. "},{"metadata":{"_cell_guid":"97d065c6-2f25-4c3f-9ab4-850d484bff9f","_uuid":"f61ded712d1d0bc45aff9770a7264199da84ce85"},"cell_type":"markdown","source":"**Let's start our powerful Heat table, using deal prob cat and user_type**"},{"metadata":{"_cell_guid":"f0029601-037c-45d9-a1b3-81622bf4be5f","_uuid":"7f5b9c8ed80a251bb3b76ddcc3f6471920d0744d","_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"cols = ['user_type','deal_prob_cat']\ncolmap = sns.light_palette(\"green\", as_cmap=True)\npd.crosstab(df_train[cols[0]], df_train[cols[1]]).style.background_gradient(cmap = colmap)","execution_count":15,"outputs":[]},{"metadata":{"_cell_guid":"926e75e1-cca8-4765-82d2-2d256f39e59f","_uuid":"c05cd7692af25888c33f0498f1d51825209c7844"},"cell_type":"markdown","source":"Intetresting value distribuition...  We have a highest "},{"metadata":{"_cell_guid":"88ff70c3-ab9c-426a-b47b-6cdca3ebac2e","_uuid":"0bd7acc5884f477e911a6f8c01e27f446c4995dd"},"cell_type":"markdown","source":"<h2>Parent Category Name Feature: </h2>\n- Count\n- Crossed with deal prob\n- Crossed with price"},{"metadata":{"_cell_guid":"9c35a70b-18cc-479f-94c5-f3c705a51949","_uuid":"2d2d72f45321daec4118a982b6b0000e12d24792","_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(12,14))\n\nplt.subplot(3,1,1)\ng = sns.countplot(x='parent_category_name_en', data=df_train)\ng.set_xlabel('User Type',fontsize=16)\ng.set_ylabel('Count',fontsize=16)\ng.set_title('Category of Ad',fontsize= 20)\ng.set_xticklabels(g.get_xticklabels(),rotation=45)\n\nplt.subplot(3,1,2)\ng1 = sns.boxplot(x='parent_category_name_en',y='deal_probability', data=df_train)\ng1.set_xlabel(\"Category's Name\",fontsize=16)\ng1.set_ylabel('Deal Probability',fontsize=16)\ng1.set_title('Category of Ad',fontsize= 20)\ng1.set_xticklabels(g.get_xticklabels(),rotation=45)\n\nplt.subplot(3,1,3)\ng2 = sns.boxplot(x='parent_category_name_en', y='price_log', data=df_train)\ng2.set_xlabel(\"Category's Name\",fontsize=16)\ng2.set_ylabel('Price Log',fontsize=16)\ng2.set_title('Category of Ad',fontsize= 20)\ng2.set_xticklabels(g1.get_xticklabels(),rotation=45)\n\nplt.subplots_adjust(hspace = 0.7,top = 0.9)\n\nplt.show()","execution_count":16,"outputs":[]},{"metadata":{"_cell_guid":"95d8a95b-f49f-45d5-a1d8-50a9362dd4ca","_uuid":"f093d22593f0785799e829fd3dd883c5f9698be8"},"cell_type":"markdown","source":"Let's use our new features to understand better the distribuition of each category'"},{"metadata":{"_cell_guid":"80c92186-2f90-4bb8-b030-1a79b7f2917b","_uuid":"04d29b3589f17c5ef5266fefca1ac4ee9778e147","_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"cols = ['parent_category_name_en','deal_prob_cat']\ncm = sns.light_palette(\"green\", as_cmap=True)\npd.crosstab(df_train[cols[0]], df_train[cols[1]]).style.background_gradient(cmap = cm)","execution_count":17,"outputs":[]},{"metadata":{"_cell_guid":"00b5a898-1d14-455f-9285-455a3c9c3221","_uuid":"08627e493a055d692fe53d0de3ad7cfa5ad00c9e"},"cell_type":"markdown","source":"Very interesting and meaningful crosstab."},{"metadata":{"_cell_guid":"82ac092d-423b-4f37-b6f7-cfdbf78320c9","_uuid":"a439ece1e1270b5f0ea4326b3d6292a436180299"},"cell_type":"markdown","source":"<h2>Region Feature</h2>\n- Count\n- Crossed with deal prob\n- Crossed with price"},{"metadata":{"_cell_guid":"6d4cfb15-59b1-4aa2-955c-5f4dbb9fe884","_uuid":"a36c34461571166984bd14ece1952730b35878fd","_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16,20))\nplt.subplot(3,1,1)\ng = sns.countplot(x='region_en', data=df_train)\ng.set_xlabel('Ad Regions',fontsize=16)\ng.set_ylabel('Count',fontsize=16)\ng.set_title('Ad Regions Count',fontsize= 20)\ng.set_xticklabels(g.get_xticklabels(),rotation=90)\n\nplt.subplot(3,1,2)\ng1 = sns.boxplot(x='region_en', y='deal_probability',data=df_train, orient='')\ng1.set_xlabel('Ad Regions',fontsize=16)\ng1.set_ylabel('Deal Probability',fontsize=16)\ng1.set_title('Ad Regions Deal Prob Distribuition',fontsize= 20)\ng1.set_xticklabels(g1.get_xticklabels(),rotation=90)\n\nplt.subplot(3,1,3)\ng2 = sns.boxplot(x='region_en', y='price_log',data=df_train, orient='')\ng2.set_xlabel('Ad Regions',fontsize=16)\ng2.set_ylabel('Price Log Distribuition',fontsize=16)\ng2.set_title('Ad Regions Price Distribuition',fontsize= 20)\ng2.set_xticklabels(g2.get_xticklabels(),rotation=90)\n\nplt.subplots_adjust(hspace = 0.7,top = 0.9)\n\nplt.show()","execution_count":18,"outputs":[]},{"metadata":{"_cell_guid":"58eb9880-ae32-4cc3-905c-62a7cf3cd8fb","_uuid":"861163985906eb8d30a50818f670f40725dae44a"},"cell_type":"markdown","source":"We can see a city with the clear highest frequency, but almost all cities with the same price statistics"},{"metadata":{"_cell_guid":"d02fee69-6d96-411a-9f86-f853a831255c","_uuid":"bbac81a8bdfed9a9af6d328f6881000a31384986"},"cell_type":"markdown","source":"<h3>Let's take a look at our crosstab with region and deal prob categorys'</h3>"},{"metadata":{"_cell_guid":"5c6aed2c-524b-423d-b615-dbe8058e92fa","_uuid":"46888eae68a8a546e614c7899f89d22126a2647a","trusted":true},"cell_type":"code","source":"cols = ['region_en','deal_prob_cat']\ncm = sns.light_palette(\"green\", as_cmap=True)\npd.crosstab(df_train[cols[0]], df_train[cols[1]]).style.background_gradient(cmap = cm)","execution_count":19,"outputs":[]},{"metadata":{"_cell_guid":"960d465b-0568-4447-894f-f5909cf3e0a1","_uuid":"8bec9db6a2e98f93bddb0e8a7a27369904061ea0"},"cell_type":"markdown","source":"## I will do a test showing the count of citys with high and lowest deal probability to we verify if they are the same"},{"metadata":{"_cell_guid":"d4cf87cf-7c6f-49fd-a776-4957143ddfc6","_uuid":"5c737c3bdf3fe320b3c86a754deb10e3d1bb7793","_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"lower_probs = ['0 -.02%', '.02%-.05%', '.05-.10']\nhigher_probs = ['.70-.85', '.85+']\n\nplt.figure(figsize=(15,12))\n\nplt.subplot(2,1,1)\ng = sns.countplot(x='region_en', data=df_train[df_train.deal_prob_cat.isin(lower_probs)])\ng.set_title(\"Region with lower deal probs\", fontsize=20)\ng.set_xticklabels(g.get_xticklabels(),rotation=90)\ng.set_xlabel(\"Regions\", fontsize=15)\ng.set_ylabel(\"Count\", fontsize=15)\n\n\nplt.subplot(2,1,2)\ng1 = sns.countplot(x='region_en', data=df_train[df_train.deal_prob_cat.isin(higher_probs)])\ng1.set_title(\"Regions with higher deal probability\", fontsize=20)\ng1.set_xticklabels(g1.get_xticklabels(),rotation=90)\ng1.set_xlabel(\"Regions\",fontsize=15)\ng1.set_ylabel(\"Count\", fontsize=15)\n\nplt.subplots_adjust(hspace = 0.8,top = 0.9)\n\nplt.show()","execution_count":20,"outputs":[]},{"metadata":{"_cell_guid":"206e64ca-ad2f-4aa7-9449-9448409f20d4","_uuid":"165d8875f7976963068e5b617d57d66d8958cb51"},"cell_type":"markdown","source":"Very interesting graphic.  We can see that the regions are differents and also have different distribuitions"},{"metadata":{"_cell_guid":"9702059d-304a-4e39-9477-9e38bf2b8dfa","_uuid":"b2853291f601565a0d0f11478899b3dde0c9cb23"},"cell_type":"markdown","source":"  ## Taking Advantage, let's take quick look at city feature\n"},{"metadata":{"_cell_guid":"4662a3ba-92a1-4a10-80c4-8f1c8401f1c1","_uuid":"f5ca2cdf4738974f6bb472dcb94c60d6983c6554","_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"city_count = df_train['city'].value_counts()[:35].index.values\n\nplt.figure(figsize=(15,6))\n\ng = sns.countplot(x='city', data=df_train[df_train.city.isin(city_count)])\ng.set_xlabel('Ad Citys',fontsize=16)\ng.set_ylabel('Count',fontsize=16)\ng.set_title(\"Ad City's Count\",fontsize= 20)\ng.set_xticklabels(g.get_xticklabels(),rotation=90)\n\nplt.show()","execution_count":21,"outputs":[]},{"metadata":{"_cell_guid":"45e2dbe1-a856-4434-802f-06fe9609cf33","_uuid":"5d1131470832ba055366d4ca2db6dd2293e8a539","_kg_hide-input":true,"collapsed":true,"trusted":true},"cell_type":"code","source":"cities_top_35 = ['Краснодар', 'Екатеринбург', 'Новосибирск', 'Ростов-на-Дону',\n       'Нижний Новгород', 'Челябинск', 'Пермь', 'Казань', 'Самара', 'Омск',\n       'Уфа', 'Красноярск', 'Воронеж', 'Волгоград', 'Саратов', 'Тюмень',\n       'Калининград', 'Барнаул', 'Ярославль', 'Иркутск', 'Оренбург', 'Сочи',\n       'Ижевск', 'Тольятти', 'Кемерово', 'Белгород', 'Тула', 'Ставрополь',\n       'Набережные Челны', 'Новокузнецк', 'Владимир', 'Сургут', 'Магнитогорск',\n       'Нижний Тагил', 'Новороссийск']","execution_count":22,"outputs":[]},{"metadata":{"_cell_guid":"6e6c01d3-cfab-4738-b726-0af768b8ddb7","_uuid":"e96f06c7e751b35370238424571736621e290859"},"cell_type":"markdown","source":"## Let's take a look at the distribuition of Deal Probability and Price of the top 35 Cities"},{"metadata":{"_cell_guid":"12f370ca-b0c9-4374-b61a-f3456cb55ce6","_uuid":"1d941dc6e90f73e4e78fdcf9f9c1a9159d8abf31","_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16,10))\n\nplt.subplot(2,1,1)\ng = sns.boxplot(x='city', y='deal_probability', data=df_train[df_train.city.isin(cities_top_35)])\ng.set_xlabel(\"\", fontsize=15)\ng.set_ylabel(\"Deal Probability\", fontsize=15)\ng.set_title(\"Deal Probability of TOP 35 City's\", fontsize=20)\ng.set_xticklabels(g.get_xticklabels(),rotation=90)\n\nplt.subplot(2,1,2)\ng1 = sns.boxplot(x='city', y='price_log', data=df_train[df_train.city.isin(cities_top_35)])\ng1.set_xlabel(\"TOP 35 City's\", fontsize=15)\ng1.set_ylabel(\"Price Log\", fontsize=15)\ng1.set_title(\"Price Log of TOP 35 City's\", fontsize=20)\ng1.set_xticklabels(g1.get_xticklabels(),rotation=90)\n\nplt.subplots_adjust(hspace = 0.7,top = 0.9)\n\nplt.show()","execution_count":23,"outputs":[]},{"metadata":{"_cell_guid":"e3a54640-6ab6-42be-a950-7b7477a3c23b","_uuid":"a6f18ee0ed789dd93540b0ccc4793d5c11fb469e"},"cell_type":"markdown","source":"We can see that just on city have a different pattern at pricces, but in deal probability we can consider a normal distribuition\n"},{"metadata":{"_cell_guid":"6a353b70-b64c-4a6c-b6ef-879cb3d33b1e","_uuid":"770a74606ef3a870704b8c26a645c8d77623971f"},"cell_type":"markdown","source":"<h2> Category name distribuitions </h2>"},{"metadata":{"_cell_guid":"994b88ea-0bb2-4370-874e-83fcda6499a9","_uuid":"80645aa9a4dcb1d8ec7039a91b562b873a13ee8e","_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16,12))\n\nplt.subplot(2,1,1)\ng = sns.countplot(x='category_name_en', data=df_train)\ng.set_xticklabels(g.get_xticklabels(),rotation=90)\ng.set_xlabel('Category Names', fontsize=15)\ng.set_ylabel('Count', fontsize=15)\ng.set_title('Category Name Count', fontsize=20)\n\nplt.subplot(2,1,2)\ng1 = sns.boxplot(x='category_name_en', y='price_log', data=df_train)\ng1.set_xticklabels(g1.get_xticklabels(),rotation=90)\ng1.set_xlabel('Category Names', fontsize=15)\ng1.set_ylabel('Price log Dist', fontsize=15)\ng1.set_title('Category Name Count', fontsize=20)\n\nplt.subplots_adjust(hspace = 0.9,top = 0.9)\n\nplt.show()","execution_count":24,"outputs":[]},{"metadata":{"_cell_guid":"dd08199a-8341-4077-83e5-f3b8e124016a","_uuid":"1eb30874f108cbde266c2a4ff63f7e56e1c81da7"},"cell_type":"markdown","source":"We can see that land, apartments, houses, cars and trucks have a highest mean price. I will verify the deal prob using the categorical feature"},{"metadata":{"_cell_guid":"24f41923-4c77-4566-8768-1ea092e294ed","_uuid":"88bd388966560aaf5ee526b4f85177da5d7ef1ab"},"cell_type":"markdown","source":"<h3>Lets take a look at the heat table of categorical deal prob and category names</h3>\n"},{"metadata":{"_cell_guid":"7851a1ba-05e3-4e13-a0f1-513041e10e62","_uuid":"dd6eed9fd902a8c9e84eff6a46f3137ff27a9e94","trusted":true},"cell_type":"code","source":"cols = ['category_name_en','deal_prob_cat']\ncm = sns.light_palette(\"green\", as_cmap=True)\npd.crosstab(df_train[cols[0]], df_train[cols[1]]).style.background_gradient(cmap = cm)","execution_count":25,"outputs":[]},{"metadata":{"_cell_guid":"2b74f400-8ea9-4ef9-92c8-27d6be07e9be","_uuid":"e99147c43c579f9d63fb665d439df0c871f12fdb"},"cell_type":"markdown","source":"very Interesting and meaningful heat table. I will create a subset of the principal values"},{"metadata":{"_cell_guid":"f1744836-d88d-4d5e-ad23-13346919766a","_uuid":"68dc8b5018444b1b0126fbf7e8d65928e35df7a3","_kg_hide-input":true,"collapsed":true},"cell_type":"markdown","source":"## Now we will know take a look at param_1 feature"},{"metadata":{"_cell_guid":"11c74cd1-4990-4c14-acda-b5cb4e2c26f1","_uuid":"b80143ae83052ef7baa9e2f8677c74e30e05c02a","_kg_hide-input":true,"collapsed":true,"trusted":true},"cell_type":"code","source":"params = df_train.param_1.value_counts().head(35)\n\nparams.index = [\"Women's Clothing\", 'For Girls', 'For Boys', 'Selling',\n                'With mileage', 'Accessories', \"Men's clothing\", 'Other', 'Toys',\n                'Baby carriages', 'Rent', 'Repair, construction', 'Building materials',\n                'iPhone', 'Beds, sofas and armchairs', 'Tools', 'For the kitchen',\n                'Accessories', \"Children's Furniture\", 'Cabinets and Chests', \n                'Devices and accessories', 'For the house', 'Transport, transportation',\n                'Nursing Items', 'Samsung', 'Hire', 'Books',\n                'TVs and projectors', 'Bicycles and scooters',\n                'Interior items, art', 'Other', 'Cosmetics',\n                'Bedding', 'Farm animals', 'Tables and chairs']","execution_count":26,"outputs":[]},{"metadata":{"_cell_guid":"579696fa-0a0e-4f73-b0cd-649c3072c2e3","_uuid":"6d507e2fbec633a8f5e4ce0b86979bd0e71444fe","_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16,5))\n\ng = sns.barplot(x=params.index, y=params.values)\ng.set_xlabel(\"Translated params\", fontsize=15)\ng.set_ylabel(\"Count\", fontsize=15)\ng.set_title(\"Most Frequent params in Ads\", fontsize=20)\ng.set_xticklabels(g.get_xticklabels(),rotation=90)\n\nplt.show()","execution_count":27,"outputs":[]},{"metadata":{"_cell_guid":"1e579ec1-7b37-4137-9ad2-664ac79961cc","_uuid":"4adf79cfcabeb7c6646d8e351489dee8e64ae754"},"cell_type":"markdown","source":"Wow, it's very insightful. \nIt's very clear to see that womens, \"for girls\",  \"for boys\" and \"selling\" have the highest % of params.<br>\nLet's take a look how representative is the top 5 values"},{"metadata":{"_cell_guid":"fc6a5c78-b230-46af-b1d1-aef0517e27fe","_uuid":"5aebf5776f649db29d22f24f44027e2eff98ae71","_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print(\"The top five Ad params in %\")\nprint(round((params / len(df_train) * 100).head(n=5),2))","execution_count":28,"outputs":[]},{"metadata":{"_cell_guid":"9bbe4784-ad9c-4cf4-9aeb-8e4a5e10c27c","_uuid":"8663a19b4bd97c4ead52a10935a6bbc84c1f3a8e"},"cell_type":"markdown","source":"This top 5 values represents 44.62% of total all showed values. Later I will explore this further."},{"metadata":{"_cell_guid":"04ad6e0c-9a70-4de8-aa96-8daeb577db06","_uuid":"741722f96422e904e2898335b2c49868f9714707"},"cell_type":"markdown","source":"### I will try understand the top 35 most frequent values in param_1"},{"metadata":{"_cell_guid":"22a64587-c1c1-400e-b3c4-f5754471d5e8","_uuid":"7d8b8b8dd043c00528f2c3c264c07e9926edd2b6"},"cell_type":"markdown","source":"- This 35 values represents 74.37% of total data frequency"},{"metadata":{"_cell_guid":"1f65ec21-c97e-42a4-bf4b-5ee3002178d0","_uuid":"548bdc2833b59b28ac6f8272849966b9000634c2","trusted":true},"cell_type":"code","source":"# I used the google translate to translate this all words in params\nrussian_param_names = [\"Женская одежда\",\"Для девочек\",\"Для мальчиков\",\"Продам\", \"С пробегом\",\"Аксессуары\",\n\"Мужская одежда\",\"Другое\",\"Игрушки\",\"Детские коляски\",\"Сдам\",\"Ремонт, строительство\",\"Стройматериалы\",\n\"iPhone\",\"Кровати, диваны и кресла\",\"Инструменты\",\"Для кухни\",\"Комплектующие\",\"Детская мебель\",\"Шкафы и комоды\",\n\"Приборы и аксессуары\",\"Для дома\",\"Транспорт, перевозки\",\"Товары для кормления\",\"Samsung\",\"Сниму\",\n\"Книги\",\"Телевизоры и проекторы\",\"Велосипеды и самокаты\",\"Предметы интерьера, искусство\",\"Другая\",\n\"Косметика\",\"Постельные принадлежности\",\"С/х животные\",\"Столы и стулья\"]\n\nsubset_param = df_train[df_train.param_1.isin(russian_param_names)]\n\nsubset_param['param_en'] = subset_param['param_1'].apply(lambda x : params_top35_map[x])","execution_count":29,"outputs":[]},{"metadata":{"_cell_guid":"3ddfddf7-cb98-4ba3-b8e0-afbb9da7cf0c","_uuid":"a9730457914faed270b90bc8e7ebacb514730983"},"cell_type":"markdown","source":"## Visualing the top 35 param_1 values by Prices and deal probability"},{"metadata":{"_cell_guid":"2ac6aa49-ff53-4a9a-9000-9c2629ca8dab","_uuid":"8000ad9de5d5361c1dc6fa072639abc4a382fdd7","_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"plt.figure(figsize=(16,12))\nplt.subplot(2,1,1)\ng = sns.boxplot(x='param_en', y='price_log', data=subset_param)\ng.set_xlabel(\"\", fontsize=15)\ng.set_ylabel(\"Price Dist(log)\", fontsize=15)\ng.set_title(\"Price of TOP 35 params_1\", fontsize=20)\ng.set_xticklabels(g.get_xticklabels(),rotation=90)\n\nplt.subplot(2,1,2)\ng1 = sns.boxplot(x='param_en', y='deal_probability', data=subset_param)\ng1.set_xlabel(\"TOP 35 Params\", fontsize=15)\ng1.set_ylabel(\"Deal Probability\", fontsize=15)\ng1.set_title(\"Deal Probability of TOP 35 params_1\", fontsize=20)\ng1.set_xticklabels(g.get_xticklabels(),rotation=90)\n\nplt.subplots_adjust(hspace = 0.9,top = 0.9)\n\nplt.show()","execution_count":30,"outputs":[]},{"metadata":{"_cell_guid":"5dd30e09-6dcd-4df3-b1f4-8df269863abe","_uuid":"2340af3ca67b493a8ca6e6117d091a5ce6b8cd42"},"cell_type":"markdown","source":"Very interesting values in dataset, we can verify that some params have diffferences in price and deal probability  feature. It's a very meaningful graphic.\n\n# Let's take a look at a Crosstab function that help to understand our target"},{"metadata":{"_cell_guid":"8c8b91e9-f20d-4819-9b54-fce3e4b80e80","_uuid":"a2633d93a55070e2c852dcd3e1837b95c679a4e7","trusted":true},"cell_type":"code","source":"cols = ['param_en','deal_prob_cat']\ncm = sns.light_palette(\"green\", as_cmap=True)\npd.crosstab(subset_param[cols[0]], subset_param[cols[1]]).style.background_gradient(cmap = cm)","execution_count":31,"outputs":[]},{"metadata":{"_cell_guid":"d23a9ab1-a6b3-4109-8820-4549c5d9ba3f","_uuid":"e5efffce194c21bf410bc2f4d61b65fa8f78d94e"},"cell_type":"markdown","source":"## Let's take a look at the Activation Date' - How we have a little number of dates, I will extract just the day."},{"metadata":{"_cell_guid":"c0058cd0-d9c3-417d-8f8c-546b7e746589","_uuid":"6f2c998ed6dbc31f8c13604c506f8470070d61ff","_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"df_train['day'] = df_train['activation_date'].dt.day\n\ntime_count = df_train['day'].value_counts()\n\nplt.figure(figsize=(16,5))\n\ng = sns.barplot(time_count.index, time_count.values)\ng.set_xlabel(\"Date Distribuition of Dataset\", fontsize=15)\ng.set_ylabel(\"Count\", fontsize=15)\ng.set_title(\"Ads Date of Avitos dataset\", fontsize=20)\n\nplt.show()","execution_count":32,"outputs":[]},{"metadata":{"_cell_guid":"84466d41-76e0-4973-b690-aeb98046c8ee","_uuid":"2d5ad416bb4b2e340e94ea2cf5709bc8a435b469","collapsed":true},"cell_type":"markdown","source":"Let's try see the probs by day"},{"metadata":{"_cell_guid":"98c27fe5-c3d0-4cfc-ac46-01eafa836af3","_uuid":"26c9d4b7a4dd04c46c1f9715fc7c144aa9e770f5","trusted":true},"cell_type":"code","source":"cols = ['day','deal_prob_cat']\ncm = sns.light_palette(\"green\", as_cmap=True)\npd.crosstab(df_train[cols[0]], df_train[cols[1]]).style.background_gradient(cmap = cm)","execution_count":33,"outputs":[]},{"metadata":{"_cell_guid":"5a617e1a-37ea-4c4c-b2da-59f80e032138","_uuid":"4fa4b4dff9773e446a330781311df42ce8c64f96"},"cell_type":"markdown","source":"Might we would not consider this feature to this job, but after we can do a measure of his importance"},{"metadata":{"_cell_guid":"350475e3-d242-475f-9170-f15de856befd","_uuid":"84a2a965695ae5a396a7b7f168de80f33920c475"},"cell_type":"markdown","source":"## Let's do some new feature engineering with the title and description feature\n- importing some necessary librarys"},{"metadata":{"scrolled":false,"_cell_guid":"723ea8ee-b489-4f7d-b815-1cb9549862e1","_uuid":"af67a4ab89dfd506dfc142475161d312cce7f4e1","_kg_hide-input":true,"collapsed":true,"trusted":true},"cell_type":"code","source":"#nlp\nimport string\nimport re    #for any necessary regex\nimport nltk\nimport spacy\nfrom nltk import pos_tag\nfrom nltk.stem.wordnet import WordNetLemmatizer \nfrom nltk.tokenize import word_tokenize\n# Tweet tokenizer does not split at apostophes which is what we want\nfrom nltk.tokenize import TweetTokenizer   \nfrom nltk.corpus import stopwords\n","execution_count":34,"outputs":[]},{"metadata":{"_cell_guid":"7fa97c86-7421-488d-972e-f8745e61f06f","_uuid":"16ff17babc11e76413ff6988e97f86d23152d025"},"cell_type":"markdown","source":"## Now, let's create our new variables using the title and description\n"},{"metadata":{"_cell_guid":"5ea739af-5980-48cf-a56b-617b366e114b","_uuid":"cd67f50b015dd80473ab4d5fb5836028ef5c9497","_kg_hide-input":true,"collapsed":true,"trusted":true},"cell_type":"code","source":"#Word count in each comment:\ndf_train['count_word'] = df_train[\"title\"].apply(lambda x: len(str(x).split()))\ndf_train['count_word_desc'] = df_train[\"description\"].apply(lambda x: len(str(x).split()))\n\n#Unique word count\ndf_train['count_unique_word'] = df_train[\"title\"].apply(lambda x: len(set(str(x).split())))\ndf_train['count_unique_word_desc']= df_train[\"description\"].apply(lambda x: len(set(str(x).split())))\n\n#Letter count\ndf_train['count_letters'] = df_train[\"title\"].apply(lambda x: len(str(x)))\ndf_train['count_letters_desc']= df_train[\"description\"].apply(lambda x: len(str(x)))\n\n#punctuation count\ndf_train[\"count_punctuations\"] = df_train[\"title\"].apply(lambda x: len([c for c in str(x) if c in string.punctuation]))\ndf_train[\"count_punctuations_desc\"] = df_train[\"description\"].apply(lambda x: len([c for c in str(x) if c in string.punctuation]))\n\n#upper case words count\ndf_train[\"count_words_upper\"] = df_train[\"title\"].apply(lambda x: len([w for w in str(x).split() if w.isupper()]))\ndf_train[\"count_words_upper_desc\"] = df_train[\"description\"].apply(lambda x: len([w for w in str(x).split() if w.isupper()]))\n\n#title case words count\ndf_train[\"count_words_title\"] = df_train[\"title\"].apply(lambda x: len([w for w in str(x).split() if w.istitle()]))\ndf_train[\"count_words_title_desc\"] = df_train[\"description\"].apply(lambda x: len([w for w in str(x).split() if w.istitle()]))\n\n#Average length of the words\ndf_train[\"mean_word_len\"] = df_train[\"title\"].apply(lambda x: np.mean([len(w) for w in str(x).split()]))\ndf_train[\"mean_word_len_desc\"] = df_train[\"description\"].apply(lambda x: np.mean([len(w) for w in str(x).split()]))","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"cbf3de8e-c5c9-45e0-b731-223bdbb4ec81","_uuid":"045010a783079bc817ef37f2da802fd9c72533ff"},"cell_type":"markdown","source":"### Plot the distribuition of our new features"},{"metadata":{"_cell_guid":"fc468fac-1e16-47b0-afeb-f377c1c39d69","_uuid":"845ecf1b128592631ff20025b42fd0f5035ced7b","_kg_hide-input":true,"trusted":true,"collapsed":true},"cell_type":"code","source":"plt.figure(figsize = (12,18))\n\nplt.subplot(421)\ng1 = sns.distplot(np.log(df_train['count_word']), \n                  hist=False, label='Title')\ng1 = sns.distplot(np.log(df_train['count_word_desc']), \n                  hist=False, label='Description')\ng1.set_title(\"COUNT WORDS DISTRIBUITION\", fontsize=16)\n\nplt.subplot(422)\ng2 = sns.distplot(np.log(df_train['count_unique_word']),\n                  hist=False, label='Title')\ng2 = sns.distplot(np.log(df_train['count_unique_word_desc']), \n                  hist=False, label='Description')\ng2.set_title(\"COUNT UNIQUE DISTRIBUITION\", fontsize=16)\n\nplt.subplot(423)\ng3 = sns.distplot(np.log(df_train['count_letters']), \n                  hist=False, label='Title')\ng3 = sns.distplot(np.log(df_train['count_letters_desc']), \n                  hist=False, label='Description')\ng3.set_title(\"COUNT LETTERS DISTRIBUITION\", fontsize=16)\n\nplt.subplot(424)\ng4 = sns.distplot(np.log(df_train[\"count_punctuations\"]), \n                  hist=False, label='Title')\ng4 = sns.distplot(np.log(df_train[\"count_punctuations_desc\"]), \n                  hist=False, label='Description')\ng4.set_xlim([-2,50])\ng4.set_title('COUNT PONCTUATIONS DISTRIBUITION', fontsize=16)\n\nplt.subplot(425)\ng5 = sns.distplot(np.log(df_train[\"count_words_upper\"] + 1) , \n                  hist=False, label='Title')\ng5 = sns.distplot(np.log(df_train[\"count_words_upper_desc\"] + 1) , \n                  hist=False, label='Description')\ng5.set_title('COUNT WORDS UPPER DISTRIBUITION', fontsize=16)\n\nplt.subplot(426)\ng6 = sns.distplot(np.log(df_train[\"count_words_title\"] + 1), \n                  hist=False, label='Title')\ng6 = sns.distplot(np.log(df_train[\"count_words_title_desc\"]  + 1), \n                  hist=False, label='Tags')\ng6.set_title('WORDS DISTRIBUITION', fontsize=16)\n\nplt.subplot(427)\ng7 = sns.distplot(np.log(df_train[\"mean_word_len\"]  + 1), \n                  hist=False, label='Title')\ng7 = sns.distplot(np.log(df_train[\"mean_word_len_desc\"] + 1), \n                  hist=False, label='Description')\ng7.set_xlim([-2,100])\ng7.set_title('MEAN WORD LEN DISTRIBUITION', fontsize=16)\n\nplt.subplots_adjust(wspace = 0.2, hspace = 0.4,top = 0.9)\nplt.legend()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"32663c72-484b-450e-a28a-627129db0cee","_uuid":"e2e68c6a7aa86b99882b8d75a60a28672b601d68"},"cell_type":"markdown","source":"## Lets explore the distribuition of title word count by each deal probability categorical"},{"metadata":{"_cell_guid":"65d1ee19-1979-45f0-86d3-5e0ffa81d9a8","_uuid":"c79329fceb571512debb4e475c50333d4070ffde","_kg_hide-input":true,"trusted":true,"collapsed":true},"cell_type":"code","source":"df_train['count_word_log'] = np.log(df_train['count_word'])\n\n(sns\n  .FacetGrid(df_train, \n             hue='deal_prob_cat', \n             size=5, aspect=2)\n  .map(sns.kdeplot, 'count_word_log', shade=True)\n .add_legend()\n)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"23a59474-ffde-4e99-955f-90a25a89b0d3","_uuid":"1626ffda81430793dc5c168994b19385bab0390a"},"cell_type":"markdown","source":"## Now the description word count by each deal probability categorical"},{"metadata":{"_cell_guid":"13cc6c27-5e60-463a-9853-1b54ca582079","_uuid":"afe068af7ed33cd8bd1ff911c3c86f316a25662d","_kg_hide-input":true,"trusted":true,"collapsed":true},"cell_type":"code","source":"df_train['count_word_desc_log'] = np.log(df_train['count_word_desc'])\n\n(sns\n  .FacetGrid(df_train, \n             hue='deal_prob_cat', \n             size=5, aspect=2)\n  .map(sns.kdeplot, 'count_word_desc_log', shade=True)\n .add_legend()\n)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"22383338-1910-4a40-bf27-3e40b126389f","_uuid":"98ebfa99ef30adc9d60b9fb8a4dc1116aa966dc9"},"cell_type":"markdown","source":"## Also, let's look the unique words count to title and description.'"},{"metadata":{"_cell_guid":"a52a28e1-1be4-46cd-9904-51a1166514ec","_uuid":"1a475a3e51b702b8965deea3bab7b4f3ad6f3e4c","_kg_hide-input":true,"trusted":true,"collapsed":true},"cell_type":"code","source":"df_train['count_unique_word_log'] = np.log(df_train['count_unique_word'])\n\n(sns\n  .FacetGrid(df_train, \n             hue='deal_prob_cat', \n             size=5, aspect=2)\n  .map(sns.kdeplot, 'count_unique_word_log', shade=True)\n .add_legend()\n)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"9dade370-f15a-4322-8c2a-51560b1656ef","_uuid":"c21c3a6c90dc40a0d05a03d60ab5689e7a9d558c"},"cell_type":"markdown","source":"- Count unique word of description"},{"metadata":{"_cell_guid":"7ac4a657-5be6-41bc-8c03-c335829bb901","_uuid":"3efbab352bbf1b9343ad622565bf5c639773c1ad","_kg_hide-input":true,"trusted":true,"collapsed":true},"cell_type":"code","source":"df_train['count_unique_word_desc_log'] = np.log(df_train['count_unique_word_desc'])\n\n(sns\n  .FacetGrid(df_train, \n             hue='deal_prob_cat', \n             size=5, aspect=2)\n  .map(sns.kdeplot, 'count_unique_word_desc_log', shade=True)\n .add_legend()\n)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"1d7fbd26-178b-4815-b18b-f98962854a94","_uuid":"c45d050ce1ab9408efa68718d728dc8fdcff1f88"},"cell_type":"markdown","source":"- We can suppose that the ads with lowest number of unique values have a lower change to be sold. might, it will be an excellent feature to predict the deal probability. I will verify this later."},{"metadata":{"_cell_guid":"260a2eb4-fe13-4d85-bce3-cbbb85230f0b","_uuid":"e6dc3b4e2884101300e111251ec9c742920e093d"},"cell_type":"markdown","source":"## Let's take a look at Wordcloud's with russian words :O"},{"metadata":{"_cell_guid":"aeb149fa-c910-4681-882f-12a163722a3a","_uuid":"8aca6407ab3c0b8da9248a29b7f3fa40220f2d77","collapsed":true,"trusted":true},"cell_type":"code","source":"stopWords = set(stopwords.words('russian'))","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"ba18229e-a811-4f21-9dd9-3ce73fa87a2e","_uuid":"b0a9a1164be8b9a5df911020e6a51c7baa8c37b6","_kg_hide-input":true,"trusted":true,"collapsed":true},"cell_type":"code","source":"from wordcloud import WordCloud, STOPWORDS\n\nwordcloud = WordCloud(\n                          background_color='white',\n                          stopwords=stopWords,\n                          max_words=1500,\n                          max_font_size=200, \n                          width=1000, height=800,\n                          random_state=42,\n                         ).generate(\" \".join(df_train['title'].astype(str)))\n\nprint(wordcloud)\nfig = plt.figure(figsize = (12,14))\nplt.imshow(wordcloud)\nplt.title(\"WORD CLOUD - TITLE\",fontsize=25)\nplt.axis('off')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"e2f2269b-3539-4aed-bdfd-da4c81844992","_uuid":"a3e4c6c98d0e50d8e47eec431778cee778aa73ca"},"cell_type":"markdown","source":"I really don't understand, but seems meaningful LOL"},{"metadata":{"_cell_guid":"81b8a6bb-ddd5-4304-898e-758af3a30471","_uuid":"3c8d9d74b2c03c808bc64c2920c971bad64e8770"},"cell_type":"markdown","source":"## Now let's seee the word cloud of description"},{"metadata":{"_cell_guid":"66149e98-a4e5-4b0e-8584-e5a278b4545b","_uuid":"e338a670ff9cc076616fbca5aaa87b5566fee745","_kg_hide-input":true,"trusted":true,"collapsed":true},"cell_type":"code","source":"wordcloud = WordCloud(\n                          background_color='white',\n                          stopwords=stopWords,\n                          max_words=1000,\n                          max_font_size=200, \n                          width=1000, height=800,\n                          random_state=42,\n                         ).generate(\" \".join(df_train['description'].astype(str)))\n\nprint(wordcloud)\nfig = plt.figure(figsize = (12,14))\nplt.imshow(wordcloud)\nplt.title(\"WORD CLOUD - DESCRIPTION\",fontsize=25)\nplt.axis('off')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"9bd7c572-0959-4d75-839c-33f1368e237c","_uuid":"aef33870e173a6dc2d4aa0a180970d356c8cfac9"},"cell_type":"markdown","source":"To a russian, I think that will be very interesting "},{"metadata":{"_cell_guid":"14e572c0-0f2d-49b1-bc48-c53e23fef98f","_uuid":"5081fa343975506fc44984a055560f16b7b73165"},"cell_type":"markdown","source":"## Also, ploting the param_1 WordCloud"},{"metadata":{"_cell_guid":"1cae5941-9d9e-4404-9e0c-c1785d98727a","_uuid":"d8ab06ea8739f6bd29a9a53352f93e684e399cec","_kg_hide-input":true,"trusted":true,"collapsed":true},"cell_type":"code","source":"wordcloud = WordCloud(\n                          background_color='white',\n                          stopwords=stopWords,\n                          max_words=1500,\n                          max_font_size=250, \n                          width=1000, height=800,\n                          random_state=42,\n                         ).generate(\" \".join(df_train['param_1'].astype(str)))\n\nprint(wordcloud)\nfig = plt.figure(figsize = (12,14))\nplt.imshow(wordcloud)\nplt.title(\"WORD CLOUD - PARAM_1\",fontsize=25)\nplt.axis('off')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"4cf480ab-3d1c-45c3-9187-bf68a86b805a","_uuid":"4905e49b5068f6dbe3567e68d05a8519b0b43e95"},"cell_type":"markdown","source":"- Cool word clouds =D"},{"metadata":{"_cell_guid":"48f7fb13-454c-4359-bbf7-2747fde8df5e","_uuid":"af15842bc3c8036a13a3ccf409f982fa0cd4af37"},"cell_type":"markdown","source":"## Let's explore further the description and title features, because might it can be interesting to our purpose'"},{"metadata":{"_cell_guid":"839c7ec9-0d23-46a6-8dc4-69ead66ce898","_uuid":"ead252fb3b5e5d7b0f2687713895976817b1d9f3","_kg_hide-input":true,"collapsed":true,"trusted":true},"cell_type":"code","source":"title_freq = df_train.title.value_counts()[:35]\n\ntitle_freq.index = [\"Dress\",\"Shoes\",\"Jacket\",\"Coat\",\"Jeans\",\"Overalls\",\"Sneakers\",\"Costume\",\n                    \"Boots\",\"Sandals\",\"Skirt\",\"Blouse\",\"Windbreaker\",\"I'll rent one apartment\",\n                    \"Boots\",\"Wedding Dress\",\"Shirt\",\"A bag\",\"Stroller\",\"Blouse\",\"Sandals\",\"Sofa\",\n                    \"Pants\",\"Cloak\",\"Ankle Booties\",\"A bike\",\"The plot is 10 hundred. (IZhS)\",\"Sneakers\",\n                    \"Jacket demi-season\",\"Hire a house\",\"Selling dress\",\"A jacket\",\n                    \"I'll rent a 2-room apartment\",\"T-shirt\",\"Footwear\"]","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"725c2ae6-025b-48f9-891d-b15ce4726622","_uuid":"1df4f345ea7cf919abef8b58487f8879a1a1584a","_kg_hide-input":true,"trusted":true,"collapsed":true},"cell_type":"code","source":"plt.figure(figsize=(16,6))\n\ng = sns.barplot(title_freq.index, title_freq.values)\ng.set_xlabel(\"TOP 35 Titles\", fontsize=15)\ng.set_ylabel(\"Count\", fontsize=15)\ng.set_title(\"TOP 35 titles frequency\", fontsize=20)\ng.set_xticklabels(g.get_xticklabels(),rotation=90)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"dbbca9c4-fa35-48ab-aee2-d80790d33ba5","_uuid":"2ad40504257b33bfb892323bdc71b8d51cf4c936"},"cell_type":"markdown","source":"- The top 5 values are: <br>\n1 - Dress <br>\n2 - Shoes  <br>\n3 - Jacket <br>\n4 - Coat  <br>\n5 - Jeans <br>\n\nIt can clearly show to us that clothing have a high influence in deal probability of Avito's Store"},{"metadata":{"_cell_guid":"d418eb61-cc76-499a-b585-ff4bcebba09a","_uuid":"5f15f028400f69ca93b87cbb8be2f2c45658c073","_kg_hide-input":true},"cell_type":"markdown","source":"## Looking the title by deal probability and Price"},{"metadata":{"_cell_guid":"3be7ed81-c9e8-42f6-8d88-04c89630de60","_uuid":"49460597b8687bffa8874dda8f43946ed20f9274","_kg_hide-input":true,"trusted":true,"collapsed":true},"cell_type":"code","source":"title_freq = df_train.title.value_counts()[:35]\n\nplt.figure(figsize=(16,12))\n\nplt.subplot(2,1,1)\ng = sns.boxplot(x='title', y='price_log', \n                data=df_train[df_train.title.isin(title_freq.index.values)])\ng.set_xlabel(\"\", fontsize=15)\ng.set_ylabel(\"Price Log\", fontsize=15)\ng.set_title(\"TOP 35 titles by Price_log\", fontsize=20)\ng.set_xticklabels(g.get_xticklabels(),rotation=90)\n\nplt.subplot(2,1,2)\ng1 = sns.boxplot(x='title', y='deal_probability', \n                data=df_train[df_train.title.isin(title_freq.index.values)])\ng1.set_xlabel(\"TOP 35 Titles\", fontsize=15)\ng1.set_ylabel(\"Deal Probability\", fontsize=15)\ng1.set_title(\"TOP 35 titles by Deal Probability\", fontsize=20)\ng1.set_xticklabels(g1.get_xticklabels(),rotation=90)\n\nplt.subplots_adjust(hspace = 0.7,top = 0.9)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79531767-bc7f-4b1c-9d4c-321a0a78fa8c","_uuid":"f0ff3bf3cb80870c9e1442219ea023544e77e876"},"cell_type":"markdown","source":"Very interesting values. We can  verify that we have a clear different prices in same categorys. Also the deal probability. "},{"metadata":{"_cell_guid":"da6322e6-b7ee-4bf6-803f-36b3a473b740","_uuid":"3c8dc6b676c0b82a835513915e55b0db20ff35ae","trusted":true,"collapsed":true},"cell_type":"code","source":"cols = ['title','deal_prob_cat']\ncm = sns.light_palette(\"green\", as_cmap=True)\npd.crosstab(df_train[df_train.title.isin(title_freq.index.values)][cols[0]], \n            df_train[df_train.title.isin(title_freq.index.values)][cols[1]]).style.background_gradient(cmap = cm)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"d7149330-fb42-40d6-a10c-4fab9920e828","_uuid":"601e36a630d3dab007ee228f38a095b582fdb495"},"cell_type":"markdown","source":"Very cool heat table. I will try convert this names to a better understant."},{"metadata":{"_cell_guid":"b5433b3f-1be6-4c16-b56c-8c840c907c7c","_uuid":"b5523c2f5aa39a77e60ac13a90195f7b37f5d7ab"},"cell_type":"markdown","source":"<h2>Ploting a squarify of Parent Category name and deal probability mean by each category</h2>"},{"metadata":{"_cell_guid":"0ab26a93-bcc2-4821-92ca-62eade731e31","_uuid":"3ca036b39b82af5d2e5260d26ee41e8eb545cbd2","_kg_hide-input":true,"trusted":true,"collapsed":true},"cell_type":"code","source":"import squarify \n\nplt.figure(figsize = (16,5)) \n\nplt.subplot(1,2,1)\ngrouped_prob_cat = np.log1p(df_train.groupby(['parent_category_name_en']).mean())\ngrouped_prob_cat['cat'] = grouped_prob_cat.index\ncurrent_palette = sns.color_palette()\nsquarify.plot(sizes = grouped_prob_cat['deal_probability'], \\\n              label = grouped_prob_cat.index, alpha = 0.8,color = current_palette)\nplt.rc('font', size = 12)\nplt.axis('off')\nplt.title(\"SQUARIFY OF CATEGORY AND DEAL PROB MEAN \")\n\nplt.subplot(1,2,2)\ngrouped_prob_cat = np.log1p(df_train.groupby(['parent_category_name_en']).mean())\ngrouped_prob_cat['cat'] = grouped_prob_cat.index\ncurrent_palette = sns.color_palette()\nsquarify.plot(sizes = grouped_prob_cat['price'], \\\n              label = grouped_prob_cat.index, alpha = 0.8,color = current_palette)\nplt.rc('font', size = 12)\nplt.axis('off')\nplt.title(\"SQUARIFY OF CATEGORY AND PRICE MEAN \")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"8a6985a0-601d-452f-aa58-27e8d45a7d0c","_uuid":"b6b4daad57c619ee6f33c8e22a778854b6688b37"},"cell_type":"markdown","source":"##    Ploting the user_type mean deal probability and mean price by each "},{"metadata":{"_cell_guid":"8e0e68aa-8f88-4394-a7be-aff867185196","_uuid":"0562292cd56443526fb4d3d42bb3ee3e84387962","_kg_hide-input":true,"trusted":true,"collapsed":true},"cell_type":"code","source":"plt.figure(figsize = (16,5)) \n\nplt.subplot(1,2,1)\ngrouped_prob_cat = np.log1p(df_train.groupby(['user_type']).mean())\ngrouped_prob_cat['cat'] = grouped_prob_cat.index\ncurrent_palette = sns.color_palette()\nsquarify.plot(sizes = grouped_prob_cat['deal_probability'].values, \n              label = grouped_prob_cat.index, alpha = 0.8,color = current_palette)\nplt.rc('font', size = 12)\nplt.axis('off')\nplt.title(\"SQUARIFY OF USER TYPE filled by Deal probability\")\n\nplt.subplot(1,2,2)\ngrouped_prob_cat = np.log1p(df_train.groupby(['user_type']).sum())\ngrouped_prob_cat['cat'] = grouped_prob_cat.index\ncurrent_palette = sns.color_palette()\nsquarify.plot(sizes = grouped_prob_cat['price'].values, \n              label = grouped_prob_cat.index, alpha = 0.8,color = current_palette)\nplt.rc('font', size = 12)\nplt.axis('off')\nplt.title(\"SQUARIFY weighted by Mean Price\")\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"4f15a754-6b8c-4942-af56-1ea52d5e4335","_uuid":"6e88937bc6938074633b81dc59c0184a65cb50d0"},"cell_type":"markdown","source":"I am having some error when I run my kernels, so I am doing some tests"},{"metadata":{"_cell_guid":"f18cc5f8-7399-42ae-95d1-440e28e21772","_uuid":"f62e86812b1b8b4cd7e4b2358d2d70eff1830009"},"cell_type":"markdown","source":""},{"metadata":{"_cell_guid":"5881a2ca-44aa-41ad-a07d-b5a3d6a1d4e8","_uuid":"09eb7250018519be274edb6775504884be854258","collapsed":true},"cell_type":"markdown","source":"cols_agg = ['region', 'city', 'parent_category_name', 'category_name',\n            'image_top_1', 'user_type','item_seq_number'];\n\nfor col in tqdm(cols_agg):\n    gp = df_train.groupby([col])['deal_probability']\n    mean = gp.mean()\n    std  = gp.std()\n    df_train[col + '_deal_probab_avg'] = df_train[col].map(mean)\n    df_train[col + '_deal_probab_std'] = df_train[col].map(std)\n\nfor col in tqdm(cols_agg):\n    gp = df_train.groupby([col])['price']\n    mean = gp.mean()\n    df_train[col + '_price_avg'] = df_train[col].map(mean)"},{"metadata":{"_cell_guid":"9027b638-3c53-4383-888f-170134d7366f","_uuid":"10486e121066a7677f592dcf76162f5e47c1d004","collapsed":true},"cell_type":"markdown","source":"\nI will continue doing this analysis! If you like this kernel, votes up my to keep me motivated =) "},{"metadata":{"_cell_guid":"53d71df8-3509-4201-b089-47963736c8b4","_uuid":"e621ffea029f7390bfc5e6ea4f14daa5f37631a4","collapsed":true,"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"36e13a5a-0809-4844-a58f-7eead297b087","_uuid":"2e2e153a243685d0c6ed5d6b3f098cfb3bcac704"},"cell_type":"markdown","source":"I am using some of \"new techniqes\" to preprocessing that I saw in some Kaggle Kernels"},{"metadata":{"_cell_guid":"eaa7ea6e-48f4-41f3-9fec-43107589e5ed","_uuid":"cf23b0b46293057f6e35e8be8f3a6232b5b5fbc7","collapsed":true,"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.5","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}