{"cells":[{"cell_type":"markdown","metadata":{},"source":"This script initially forked from [Predict hotel type with pandas](https://www.kaggle.com/dvasyukova/expedia-hotel-recommendations/predict-hotel-type-with-pandas/run/217488)\nis implementing some ideas from [R version of most popular local hotels](https://www.kaggle.com/signochastic/expedia-hotel-recommendations/apr-23/run/216500) by [Reda](https://www.kaggle.com/signochastic)\n\nThe main idea is to fill `test.csv` with most popular hotels:\n\n1. grouped by `orig_destination_distance`\n2. grouped by `srch_destination_id`\n3. from all the dataset.\n\nIn order to get the most popular hotels, I use a function `sum_and_count` from Reda's [script](https://www.kaggle.com/signochastic/expedia-hotel-recommendations/apr-23/run/216500) by [Reda](https://www.kaggle.com/signochastic) that give to each hotel cluster a relevance based on the number of bookings in each group.\n\n    def sum_and_count(x):\n        return np.sum(x) * 0.8456 + len(x) * (1 - 0.8456)\n        \n**Scores 0.45824** however I've limited the train rows used in order to get the script run in less than 10 minutes. You need to run it over the full train dataset removing the nrows option in `read_csv` to get this score."},{"cell_type":"code","execution_count":null,"metadata":{"collapsed":false},"outputs":[],"source":"import pandas as pd\nimport numpy as np"},{"cell_type":"code","execution_count":null,"metadata":{"collapsed":false},"outputs":[],"source":"expedia_train = pd.read_csv('../input/train.csv',\n                            usecols=[\"is_booking\",\n                                     \"orig_destination_distance\",\n                                     \"hotel_cluster\",\n                                     \"srch_destination_id\"],\n                            dtype={'is_booking':np.int8,\n                                   \"orig_destination_distance\":np.float64,\n                                   \"hotel_cluster\":np.int8,\n                                   \"srch_destination_id\":np.int32},\n                            nrows=1000000\n                           )\nexpedia_test = pd.read_csv('../input/test.csv',\n                           usecols=[\"orig_destination_distance\",\n                                    \"srch_destination_id\"],\n                            dtype={\"orig_destination_distance\":np.float64,\n                                   \"srch_destination_id\":np.int32})"},{"cell_type":"markdown","metadata":{},"source":"### Scoring hotel clusters\nNow we give to each hotel cluster a relevance based on booking events."},{"cell_type":"code","execution_count":null,"metadata":{"collapsed":false},"outputs":[],"source":"def sum_and_count(x):\n    return np.sum(x) * 0.8456 + len(x) * (1 - 0.8456)\n\ndest_id_hotel_cluster_count = expedia_train.groupby(['orig_destination_distance', 'hotel_cluster']).is_booking.apply(sum_and_count).reset_index()\ndest_id_hotel_cluster_count1 = expedia_train.groupby(['srch_destination_id', 'hotel_cluster']).is_booking.apply(sum_and_count).reset_index()\ndest_id_hotel_cluster_count.head()"},{"cell_type":"markdown","metadata":{},"source":"### Ordering hotel clusters by score for each group\nNow we define a function to get most popular hotels based on the relevance score calculated before."},{"cell_type":"code","execution_count":null,"metadata":{"collapsed":false},"outputs":[],"source":"def most_popular(group):\n    a = group.values.astype(np.int8)\n    # order hotel_clusters by score then reverse it and take the 5 first\n    clusters = a[:, 0][a[:,1].argsort()[::-1]][:5]\n    return np.array_str(clusters)[1:-1]# remove square brackets\n\ndest_top_five = dest_id_hotel_cluster_count.groupby('orig_destination_distance')['hotel_cluster', 'is_booking'].apply(most_popular).reset_index()\ndest_top_five1 = dest_id_hotel_cluster_count1.groupby('srch_destination_id')['hotel_cluster', 'is_booking'].apply(most_popular).reset_index()\ndest_top_five.head()"},{"cell_type":"markdown","metadata":{},"source":"### Merging into the test data\n\nI'm filling `test.csv` in specific order:\n\n1. grouped by `orig_destination_distance`\n2. grouped by `srch_destination_id`\n3. from all the dataset."},{"cell_type":"code","execution_count":null,"metadata":{"collapsed":false},"outputs":[],"source":"dd = expedia_test.merge(dest_top_five, how='left', on='orig_destination_distance')\ndd1 = expedia_test.merge(dest_top_five1, how='left', on='srch_destination_id')\n\n# fill null values of dd with dd1 ones\ndd[0][dd[0].isnull()] = dd1[0][dd[0].isnull()]\n\n# fill remaining null values of dd with most popular hotels from the full dataset\nmost_pop_all = expedia_train.groupby('hotel_cluster')['is_booking'].sum().nlargest(5).index\nmost_pop_all = np.array_str(most_pop_all)[1:-1]\ndd[0].fillna(most_pop_all, inplace=True)\n"},{"cell_type":"code","execution_count":null,"metadata":{"collapsed":false},"outputs":[],"source":"### Write the `csv` submission file\ndd[0].to_frame(\"hotel_cluster\").to_csv('pandas_version_of_most_popular_hotels.csv', header=True, index_label='id')"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"}},"nbformat":4,"nbformat_minor":0}