{"cells":[{"metadata":{"_uuid":"277a0f5492ebe8429e8831ad14fe7f3013378e30"},"cell_type":"markdown","source":"# What Makes a Good Deal?\nWhat makes an item offered for sale a good deal is a combination of **price** and **item type** (laptop, car, dress, table, ...).  Demand on some items may vary seasonally and geographically too, affecting the price accordingly.\n\nIn our training and test data we have the item's `price` and its geographical location (`city`, `region`).  We can identify its type either using the `category_name`feature, one of the `param_x` features, or using image recognition on the ad's `image`.\n\nLet's see how these categorical features interact with `price` affecting `deal_probability` and pick one of them to cross with `price` as our new Good Deal feature.\n\nFirst, we need to clean the `price` variable and remove outliers and missing values."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"collapsed":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"\"\"\"Copy Keras pre-trained model files to work directory from:\nhttps://www.kaggle.com/gaborfodor/keras-pretrained-models\n\nCode from: https://www.kaggle.com/classtag/extract-avito-image-features-via-keras-vgg16/notebook\n\"\"\"\nimport os\n\ncache_dir = os.path.expanduser(os.path.join('~', '.keras'))\nif not os.path.exists(cache_dir):\n    os.makedirs(cache_dir)\n\n# Create symbolic links for trained models.\n# Thanks to Lem Lordje Ko for the idea\n# https://www.kaggle.com/lemonkoala/pretrained-keras-models-symlinked-not-copied\nmodels_symlink = os.path.join(cache_dir, 'models')\nif not os.path.exists(models_symlink):\n    os.symlink('/kaggle/input/keras-pretrained-models/', models_symlink)\n\nimages_dir = os.path.expanduser(os.path.join('~', 'avito_images'))\nif not os.path.exists(images_dir):\n    os.makedirs(images_dir)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","collapsed":true,"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true,"_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"\"\"\"Extract images from Avito's advertisement image zip archive.\n\nCode adapted from: https://www.kaggle.com/classtag/extract-avito-image-features-via-keras-vgg16/notebook\n\"\"\"\nimport zipfile\n\nNUM_IMAGES_TO_EXTRACT = 100\n\nwith zipfile.ZipFile('../input/avito-demand-prediction/train_jpg.zip', 'r') as train_zip:\n    files_in_zip = sorted(train_zip.namelist())\n    for idx, file in enumerate(files_in_zip[:NUM_IMAGES_TO_EXTRACT]):\n        if file.endswith('.jpg'):\n            train_zip.extract(file, path=file.split('/')[3])\n\n!mv *.jpg/data/competition_files/train_jpg/* ~/avito_images\n!rm -rf *.jpg","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true,"_kg_hide-input":true,"_kg_hide-output":true,"_uuid":"22dcadb289b9bf0acfbe465e579a493105cfde7f"},"cell_type":"code","source":"def get_avito_image(image_id):\n    \"\"\"Read image file without extracting it from Avito's training images archive file and return it.\n    \n    Args:\n        image_id (str): Id of the image to read.\n        \n    Returns:\n       file: Image file.  Can be read with plt.imread() \n    \"\"\"\n    with zipfile.ZipFile('../input/avito-demand-prediction/train_jpg.zip', 'r') as train_zip:\n        archive_member = os.path.join('data/competition_files/train_jpg', image_id + '.jpg')\n        member = train_zip.open(archive_member)\n    return member\n\ndef is_image_extracted(image_id):\n    \"\"\"Check if extracted copy of the image exists.\n    \n    Args:\n        image_id (str): Image id\n        \n    Returns:\n        bool: True if extracted copy of the image exists.\n    \"\"\"\n    dir_path = os.path.expanduser('~/avito_images')\n    return os.path.exists(os.path.join(dir_path, 'data/competition_files/train_jpg' , image_id + '.jpg'))\n\n\ndef extract_avito_images(image_ids):\n    \"\"\"Extract images from Avito training images and return extracted images path.\n    \n    If extracted copies of all images exist, returns their path instead.\n    Args:\n        image_ids (list): Image ids to extract.\n    \n    Returns:\n        list: Extracted images path.\n    \"\"\"\n    ex = []\n    if all([is_image_extracted(x) for x in image_ids]):\n        dir_path = os.path.expanduser('~/avito_images')\n        for image_id in image_ids:\n            ex.append(os.path.join(dir_path, 'data/competition_files/train_jpg' , image_id + '.jpg'))\n        return ex\n    \n    with zipfile.ZipFile('../input/avito-demand-prediction/train_jpg.zip', 'r') as train_zip:\n        for image_id in image_ids:\n            archive_member = os.path.join('data/competition_files/train_jpg', image_id + '.jpg')\n            dest = os.path.expanduser('~/avito_images')\n            if os.path.exists(os.path.join(dest, archive_member)):\n                e = os.path.join(dest, archive_member)\n            else:        \n                e = train_zip.extract(archive_member, dest)\n            ex.append(e)\n    return ex\n\n\ndef show_avito_images(image_ids):\n    \"\"\"Plot images from Avito's image data set.\n    \n    Args:\n        image_id (list): Ids of the images to plot.\n        \n    Returns:\n        matplotlib axis\n    \"\"\"\n    ncols = 3\n    nrows = len(image_ids) // ncols + 1 - (len(image_ids) % ncols == 0)\n    images_path = extract_avito_images(image_ids)\n    fig, axes = plt.subplots(nrows, ncols, figsize=(20, 20))\n    for i, (image_id, image_path) in enumerate(zip(image_ids, images_path)):\n        image = Image.open(image_path)\n        ax = axes[i//ncols, i%ncols]\n        ax.imshow(image)\n        ax.set_title(image_id)\n        ax.axis('off')\n    plt.show()\n    return ax","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true,"collapsed":true,"_uuid":"6d4f6ed5dcf3169780dc00dd8ea9d2352e347618","_kg_hide-output":true},"cell_type":"code","source":"# Dictionary of category names (Russian to English)\nparent_cat_rus_eng = {\n    'Личные вещи': 'Personal items',\n    'Для дома и дачи': 'For home and cottages',\n    'Бытовая электроника': 'Consumer electronics',\n    'Недвижимость': 'Real estate',\n    'Хобби и отдых': 'Hobbies and recreation',\n    'Транспорт': 'Transportation',\n    'Услуги': 'Services',\n    'Животные':'Animals',\n    'Для бизнеса': 'For business'\n}\n\n# Dictionary of parent category names (Russian to English)\ncat_rus_en = {\n    'Автомобили': 'Cars',\n     'Аквариум': 'Aquarium',\n     'Аудио и видео': 'Audio and video',\n     'Билеты и путешествия': 'Tickets and travel',\n     'Бытовая техника': 'Household appliances',\n     'Велосипеды': 'Bicycles',\n     'Водный транспорт': 'Water transport',\n     'Гаражи и машиноместа': 'Garages and Parking spaces',\n     'Готовый бизнес': 'Ready business',\n     'Грузовики и спецтехника': 'Trucks and machinery',\n     'Детская одежда и обувь': \"Children's clothing and footwear\",\n     'Дома, дачи, коттеджи': 'Houses, cottages, cottages',\n     'Другие животные': 'Other animals',\n     'Земельные участки': 'Land plots',\n     'Игры, приставки и программы': 'Games, consoles and programs',\n     'Квартиры': 'Apartments',\n     'Книги и журналы': 'Books and magazines',\n     'Коллекционирование': 'Collecting',\n     'Коммерческая недвижимость': 'Commercial real estate',\n     'Комнаты': 'Rooms',\n     'Кошки': 'Cats',\n     'Красота и здоровье': 'Health and beauty',\n     'Мебель и интерьер': 'Furniture and interior',\n     'Мотоциклы и мототехника': 'Motorcycles and motor vehicles',\n     'Музыкальные инструменты': 'Musical instruments',\n     'Настольные компьютеры': 'Desktop computers',\n     'Недвижимость за рубежом': 'Real estate abroad',\n     'Ноутбуки': 'Laptops',\n     'Оборудование для бизнеса': 'Business equipment',\n     'Одежда, обувь, аксессуары': 'Clothes, shoes, accessories',\n     'Оргтехника и расходники': 'Office equipment and consumables',\n     'Охота и рыбалка': 'Hunting and fishing',\n     'Планшеты и электронные книги': 'Tablets and e-books',\n     'Посуда и товары для кухни': 'Kitchen utensils and goods',\n     'Предложение услуг': 'Offer of services',\n     'Продукты питания': 'Food',\n     'Птицы': 'Birds',\n     'Растения': 'Plants',\n     'Ремонт и строительство': 'Repair and construction',\n     'Собаки': 'Dogs',\n     'Спорт и отдых': 'Sport and recreation',\n     'Телефоны': 'Phones',\n     'Товары для детей и игрушки': 'Goods for children and toys',\n     'Товары для животных': 'Animal products',\n     'Товары для компьютера': 'Goods for computer',\n     'Фототехника': 'Photo equipment',\n     'Часы и украшения': 'Watches and jewelry'\n}","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"8d19fd9415a6949999f1a39e7b1ceeccf44c663c","_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport os\nimport matplotlib.pyplot as plt\nimport matplotlib\nimport seaborn as sns\nsns.set()\nplt.rc('font', size=16)\n\n%matplotlib inline","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"775f60c2f372cc11c5252774346ca10e8e3b86cf","_kg_hide-input":true,"_kg_hide-output":true,"collapsed":true},"cell_type":"code","source":"usecols = ['item_id', 'parent_category_name', 'category_name', 'price', 'image']\ntrain_df = pd.read_csv('../input/avito-demand-prediction/train.csv', usecols=usecols)\ntest_df = pd.read_csv('../input/avito-demand-prediction/test.csv', usecols=usecols)\ndf = pd.concat([train_df, test_df], axis=0)\n\ndf.category_name = df.category_name.map(lambda x: cat_rus_en[x])\ndf.parent_category_name = df.parent_category_name.map(lambda x: parent_cat_rus_eng[x])\ndf.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8a387790199559acbb9a050154fbefb721fb5d2f"},"cell_type":"markdown","source":"# Remove Extreme Prices\nWe have about 6% of the ads where the price is missing or is less than 10  &#8381; (Russian Rubles) (~ 0.16 USD).  These ads will be ignored from our analysis since they probably don't represent a reasonable asking price. \n\nHere is the distribution of the top 10 category names these cheap ads fall under, and their percentage make up."},{"metadata":{"_kg_hide-input":true,"trusted":true,"scrolled":false,"_uuid":"46bd6ebaee19ad43f34d2a61c6f093bdac3451b4","collapsed":true},"cell_type":"code","source":"cheap = df[df.price < 10].category_name.value_counts().map(lambda x: '{:.2f}%'.format(x/df.shape[0] * 100))\nprint('Categories of items < 10 \\u20BD (top 10)')\ncheap.head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3db2fdd526218b423d4ebd3075f3ea17a21a4e49"},"cell_type":"markdown","source":"Also, about 6% of the ads have price tage above 1 mllion &#8381; (~ 16,000 USD).  These are mosly real estate and cars with high price variance.  We will ignore these ads as well.\n\nHere is the top 10 category distribution for these items and their percentage among all listed ads."},{"metadata":{"_kg_hide-input":true,"trusted":true,"_uuid":"5e7435ea129a82647bb9a9af6fd6094a5a11ceba","collapsed":true},"cell_type":"code","source":"expensive = df[df.price > 1000000].category_name.value_counts().map(lambda x: '{:.2f}%'.format(x/df.shape[0] * 100))\nprint('Categories of items > 1M \\u20BD (top 10)')\nexpensive.head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8063a9c823691f23b3b3b8a2b5f5bdfc93d2cfa9"},"cell_type":"markdown","source":"Overall, we drop 12% of the ads for this exercise.  This may seem a lot of wasted data, but I believe it's necessary to simplify this analysis and this feature creation."},{"metadata":{"trusted":true,"_uuid":"f0267383b4ea4d758ea7ea1aa22956fdb8ed0833","_kg_hide-input":true,"scrolled":true,"collapsed":true},"cell_type":"code","source":"na_count = df.price.isna().sum()\nzero_count = df[df.price < 10].price.count()\nmill_count = df[df.price > 1000000].price.count()\ntotal_count = df.shape[0]\nprint('Ads with missing price:\\t{:,}\\t({:.2f}%)'.format(na_count, na_count/total_count*100))\nprint('Ads with price < 10 \\u20BD:\\t{:,}\\t({:.2f}%)'.format(zero_count, zero_count/total_count*100))\nprint('Ads with price > 1M \\u20BD:\\t{:,}\\t({:.2f}%)'.format(mill_count, mill_count/total_count*100))\ntotal_dropped = na_count + zero_count + mill_count\nprint('---------------------------------------')\nprint('Total dropped:\\t\\t{:,}\\t({:.2f}%)'.format(total_dropped, total_dropped/total_count*100))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c420e4aa47a354073400a6f52cd2f57fd55c1be9"},"cell_type":"markdown","source":"# Price Binning\nA good item price binning should be based on the price distribution.  A histogram for the log value of the price for the remaining ads is shown below.  The bin boundaries are shown as vertical lines."},{"metadata":{"_kg_hide-input":true,"trusted":true,"scrolled":true,"_uuid":"fc16d5c1b385903d6a18435e9d5a9de5a01c933e","collapsed":true},"cell_type":"code","source":"dff = df[(df.price > 10) & (df.price < 1000000)] # Drop missing price and < 10 Ruble\nprice_log = np.log(dff.price)\nout, bins = pd.cut(price_log, bins=100, retbins=True, labels=False)\n\nplt.figure(figsize=(16, 5))\nax = sns.distplot(price_log, axlabel='Log(price)')\nplt.title('Histogram of the Logarithmic value of the price.  Prices under 10 \\u20BD and above 1M \\u20BD are dropped.\\nVertical lines show bin boundaries.')\nplt.vlines(bins, color='g', ymin=0, ymax=0.1, alpha=0.3)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9c6f78ad60c4795078b8cef285219703ebd44ff8"},"cell_type":"markdown","source":"# Price Variance Within category_name\nHigh price variance within an item category (`category_name`) may indicate the existance of implicit sub-categories causing the wide price range.  This behavior is undesirable for the Good Deal feature we are building here.  It is hard to have a good price estimate for an item if we only know the broad category it falls in.  For example, a `consumer_electronics` category will have a wide price range.  It's hard to estimate how good this deal is without knowing its subcategory (MacBookPro, speaker, TV, ...).\n\nFor a qualitative measure, the two violin plots below show how the log value of the price ( $Log(price)$ ) varies within each `category_name`.  The variation is large for about half of the categories.  Keep in mind that we are using a log scale for the price."},{"metadata":{"trusted":true,"_uuid":"4d33cea495ef360c7908d65eeb3df89fc35d9ec2","_kg_hide-input":true,"scrolled":false,"collapsed":true},"cell_type":"code","source":"cats = dff.category_name.unique()\ndfa = dff[dff.category_name.isin(cats[0: 23])][['category_name', 'price']]\ncategory_name = dfa.category_name\nprice_log = np.log(dfa.price)\n\nplt.figure(figsize=(20, 8))\nax = sns.violinplot(x=category_name, y=price_log, scale='width', palette='Set3')\nplt.ylim([1, 15])\nplt.xticks(rotation=40, fontsize=14)\nplt.ylabel('Log(price) \\u20BD')\nplt.title('Price variance within categories (1 of 2)')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"3834f0363a47f661363e607e36f3a1648a388f10","_kg_hide-input":true,"scrolled":false,"collapsed":true},"cell_type":"code","source":"dfb = dff[dff.category_name.isin(cats[23: 47])][['category_name', 'price']]\ncategory_name = dfb.category_name\nprice_log = np.log(dfb.price)\nplt.figure(figsize=(20, 8))\n\nax = sns.violinplot(x=category_name, y=price_log, scale='width', palette='Set3')\nplt.ylim([1, 15])\nplt.xticks(rotation=40, fontsize=14)\nplt.title('Price variance within categories (2 of 2)')\nplt.ylabel('Log(price) \\u20BD')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"802aed8ae46895d5cb7fee239f7ecc4a6dd806a1"},"cell_type":"markdown","source":"And then let's see how the price ( $Log(price)$ ) varies within the `parent_category_name` in the violin plot below.  Again, the variation looks relatively high for almost all the parent categories."},{"metadata":{"trusted":true,"_uuid":"88e084069fe28660fcb3f776c073d0ebfdc15881","_kg_hide-input":true,"collapsed":true},"cell_type":"code","source":"parent_category = dff.parent_category_name\ncategory = dff.category_name\nprice_log = np.log(dff.price)\n\nplt.figure(figsize=(20, 8))\nax = sns.violinplot(x=parent_category, y=price_log, scale='width', palette='Set3')\nplt.ylim([1, 15])\nplt.xticks(rotation=40, fontsize=14)\nplt.title('Price variance within parent categories')\nplt.ylabel('Log(price) \\u20BD')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"42cc0e6d4b68ffd58cb394d8943f9575f9798f9e"},"cell_type":"markdown","source":"For an objective measure, let's calculate the [coefficient of variation (CV)][cv] for our 47 categories.  A CV value greater than 1 suggests high relative variability in our variable.\n\nAs you can see blow, almost all categories have CV value greater than 1.\n\n[cv]: http://www.statisticshowto.com/probability-and-statistics/how-to-find-a-coefficient-of-variation/"},{"metadata":{"trusted":true,"_uuid":"7b8cf8af28c411036f1eae390fc1a9eb67dfe15f","_kg_hide-input":true,"scrolled":false,"collapsed":true},"cell_type":"code","source":"print('Coefficient of variation (CV) for prices in different categories (category_name).')\ndffd = dff.groupby('category_name')['price'].apply(lambda x: np.std(x)/np.mean(x)).sort_values(ascending=False)\ndffd","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7998c51f500b54bcf644d42a05e03c2e79822112"},"cell_type":"markdown","source":"----------------"},{"metadata":{"_uuid":"6669d7e28b867dbc8e308c20a45e7c9969e81562"},"cell_type":"markdown","source":"# Use Ad Image to Identify Item Category\nIn [an earlier kernel][1] I compared three popular image recognition models on a small sample of images and found that [InceptionV3][inceptionv3] performed well.  This also aligns with [the models' accuracy as reported by Keras][2].  Therefore, we will be using InceptionV3 to recognize the item in the ad's image.\n\nDue to computational resources limitations here on Kaggle, I ran this step locally to recognize the items in about 1.85 million images.  You can check out the recognized image items in this kernel's data set (`train_image_labels.csv`, `test_image_labels.csv`).\n\nLet's see how well InceptionV3 did on this image recognition job.  Here are some images with their recognized labels.\n\n[1]: https://www.kaggle.com/wesamelshamy/high-correlation-feature-image-classification-conf\n[2]: https://keras.io/applications/#documentation-for-individual-models\n[inceptionv3]: https://keras.io/applications/#inceptionv3"},{"metadata":{"trusted":true,"collapsed":true,"_uuid":"2a538cc998b1d72511512b09a1f8b0ae7e44ca22","_kg_hide-input":true,"_kg_hide-output":true},"cell_type":"code","source":"train_image_labels = pd.read_csv('../input/avito-images-recognized/train_image_labels.csv', index_col='image_id')\ntest_image_labels = pd.read_csv('../input/avito-images-recognized/test_image_labels.csv', index_col='image_id')\nall_image_labels = pd.concat([train_image_labels, test_image_labels], axis=0)","execution_count":null,"outputs":[]},{"metadata":{"_kg_hide-input":true,"trusted":true,"scrolled":false,"_uuid":"d6298d0264bbecac0ed4c23c1821d7f99ec935d9","collapsed":true},"cell_type":"code","source":"from PIL import Image\n\ndir_iter = os.scandir(os.path.expanduser('~/avito_images'))\nfig, axes = plt.subplots(5, 5, figsize=(20, 20))\nfor i in range(25):\n    e = next(dir_iter)\n    img = Image.open(e.path)\n    img = img.resize((360, 360))\n    ax = axes[i//5, i%5]\n    ax.imshow(img)\n    im_id = e.name.split('.')[0]\n    im_label = train_image_labels.loc[im_id]['image_label']\n    ax.set_title(im_label, fontsize=24)\n    ax.axis('off')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"397065c17e186dfc0bb6b5209ab7671535d1b353"},"cell_type":"markdown","source":"Overall, there were 996 identified items in the training and test datasets combined.  This is a good improvement over the 47 `category_name` classes we had earlier.\n\nLet's see how the price varies within each one of these identified items."},{"metadata":{"_uuid":"5d42995b8d8f1f7b3e862336d2cac104c57efe13"},"cell_type":"markdown","source":"# Price Variance Within Identified Items\nHere are the recognized categories with the top 10 CV values.  We have 996 such categories."},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"trusted":true,"_uuid":"d7aa0cdf2aa92f17715ee55ec1598740ebd0fb88","scrolled":true,"collapsed":true},"cell_type":"code","source":"print('Coefficient of variation (CV) for prices in different recognized image categories.')\ndfl = dff.merge(all_image_labels, left_on='image', right_index=True, how='left')\ndfd = dfl.groupby('image_label')['price'].apply(lambda x: np.std(x)/np.mean(x)).sort_values(ascending=False)\ndfd.head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0eccd9305143a9e036fc08aea78b11824d4fb86c"},"cell_type":"markdown","source":"We now compare price variance in the recognized image categories and price variance in `category_name` categories.  The histogram below shows the distribution of the Coefficient of Variation (CV) for both cases.\n\nA good distribution in our case should be right skewed (very thin and long right tail) indicating low price variance within categories."},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true,"collapsed":true,"_uuid":"d788a0529a870b110a0de8ffd449193ab32dd1ea"},"cell_type":"code","source":"dfd2 = dfl.groupby('image_label')['price'].apply(lambda x: np.std(np.log(x))/np.mean(np.log(x)))\ndffd2 = dfl.groupby('category_name')['price'].apply(lambda x: np.std(np.log(x))/np.mean(np.log(x)))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f438864a6346c59bb87ee95f3064d986e18a01db","scrolled":false,"_kg_hide-input":true,"_kg_hide-output":false,"collapsed":true},"cell_type":"code","source":"plt.figure(figsize=(16, 5))\nsns.set()\nax = sns.distplot(dfd2.values, color='darkred', hist_kws={'alpha': 0.3})\nax = sns.distplot(dffd2.values, color='darkgreen', hist_kws={'alpha': 0.3}, ax=ax)\nax.set_title('CV of image_lable and category_name')\nax.set_xlabel('CV of $Log(price)$')\nax.legend(['image label', 'category_name'])\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4d358db83bdf6bf9dc6ec7fe5f3fa18c84e05a39"},"cell_type":"markdown","source":"# Price Variance in param_2\nFinally, subcategory features `param_1`, `param_2` and `param_3` are other candidates for use with `price` to predict `deal_probability`.\n\nOne thing to note is that `param_2` and `param_3` are missing from about half the ads.  Also, `param_3` takes on values like {*small, 40-42 (XS), 39,..*} which could overlap with other categories.\n\nAgain, here is a histogram comparing CV distributions for all three cases (`image_label`, `category_name`, `param_2`).  The more right skewed the distribution is (longer and thinner right tail) the better.  That indicates lower CV values for more categories."},{"metadata":{"trusted":true,"_kg_hide-input":true,"_kg_hide-output":false,"_uuid":"205578f79112a3c5bafdad933a98446ea633c216","collapsed":true},"cell_type":"code","source":"train_all = pd.read_csv('../input/avito-demand-prediction/train.csv', usecols=['param_1', 'param_2', 'param_3', 'price', 'deal_probability'])\nparam2v = train_all[(train_all.price > 0) & (train_all.price < 1000000)].groupby('param_2')['price'].apply(lambda x: np.std(np.log(x))/np.mean(np.log(x)))\n\nplt.figure(figsize=(16, 5))\nsns.set()\nax = sns.distplot(dfd2.values, color='darkred', hist_kws={'alpha': 0.3})\nax = sns.distplot(dffd2.values, color='darkgreen', hist_kws={'alpha': 0.3}, ax=ax)\nax = sns.distplot(param2v.values, color='navy', hist_kws={'alpha': 0.3}, ax=ax)\n\nax.set_title('CV of image_label, category_name, and param_2')\nax.set_xlabel('CV of $Log(price)$')\nax.legend(['image label', 'category_name', 'param_2'])\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b23e416d7a705178e7143ceaf6a9e7e5b266a32a"},"cell_type":"markdown","source":"`param_2` has the lowest CV values among the three, by a big margin."},{"metadata":{"_uuid":"d34ba4a3aaa88f1d57decda60c4edd4ab8bb0c23"},"cell_type":"markdown","source":"# Effect of price on deal_probability\nNow that we know that `param_2` has the best potential for identifying a good deal along with the price, let's see how their interaction affects `deal_probability`.\n\nIn the figure below we can clearly see the deal probability value drops as the price within the same `param_2` category goes up, making the deal less attractive."},{"metadata":{"_kg_hide-input":true,"_kg_hide-output":false,"trusted":true,"_uuid":"4ee951df60a6eaf69e5da7f435866e050f1e910e","collapsed":true},"cell_type":"code","source":"shoes = train_all[train_all.param_2 == 'Обувь'][['price', 'deal_probability']]\nouter = train_all[train_all.param_2 == 'Верхняя одежда'][['price', 'deal_probability']]\ndress = train_all[train_all.param_2 == 'Платья и юбки'][['price', 'deal_probability']]\nknit = train_all[train_all.param_2 == 'Трикотаж'][['price', 'deal_probability']]\n\nplt.figure(figsize=(16, 5))\nfontsize=18\n\nax1 = plt.subplot(121)\nplt.scatter(shoes.price, shoes.deal_probability, s=3)\nx = plt.setp(ax1.get_yticklabels(), fontsize=fontsize)\nplt.xlim([0, 6000])\nplt.xlabel('Price \\u20BD', fontsize=fontsize)\nplt.xticks(fontsize=fontsize)\nplt.yticks(fontsize=fontsize)\nplt.ylabel('Deal probability', fontsize=fontsize)\nax = plt.title('Shoes', fontsize=fontsize)\n\nax2 = plt.subplot(122)\nplt.scatter(outer.price, outer.deal_probability, s=3)\nplt.setp(ax2.get_yticklabels(), visible=False)\nplt.xlim([0, 7000])\nplt.xlabel('Price \\u20BD', fontsize=fontsize)\nplt.xticks(fontsize=fontsize)\nax = plt.title('Outerwear', fontsize=fontsize)\n\n# ax3 = plt.subplot(133)\n# plt.scatter(dress.price, dress.deal_probability, s=3)\n# x = plt.setp(ax3.get_yticklabels(), visible=False)\n# plt.xlim([0, 6000])\n# plt.xlabel('Price \\u20BD')\n# ax = plt.title('Dresses')\n\n# ax4 = plt.subplot(144, sharex=ax1, sharey=ax1)\n# plt.scatter(knit.price, knit.deal_probability)\n# x = plt.setp(ax4.get_yticklabels(), visible=False)\n# plt.xlim([0, 2000])\n# plt.xlabel('Price \\u20BD')\n# ax = plt.title('Knitwear')\n\nplt.tight_layout()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fb115dfa5abef5ce3b825e4ee8c3af8e18d73966"},"cell_type":"markdown","source":"# What Does This All Mean?\nHaving `param_2` and `price` as separate features in decision tree models help the model as we can see in [kxx's xbg model][kxx], [Lathwal's model][lath], and [Himanshu's LGM model][him] among many others.  They usually rank in the top 10 predictive features.\n\n**Having a feature that is a corss of `param_2` and bucketized `price` should help more in predicting `deal_probability`.**\n\nThat was my analysis.  I don't think sharing a complete wokring model with code and everything helps us learn.  Now you can implement this feature yourself and share with us what you find.\n\nCheers\n\n[kxx]: https://www.kaggle.com/kailex/xgb-text2vec-tfidf-0-2237\n[lath]: https://www.kaggle.com/codename007/avito-eda-fe-time-series-dt-visualization\n[him]: https://www.kaggle.com/him4318/avito-lightgbm-with-ridge-feature-v-2-0"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.5","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}