{"cells":[{"metadata":{"_cell_guid":"09fe787a-a457-4d9a-8ae5-8d76678d9384","_uuid":"3d1460fd9482cf2456ba9e3e6208ba887f0c1fff","trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Draw inline\n%matplotlib inline\n\n# Set figure aesthetics\nsns.set_style(\"white\", {'ytick.major.size': 10.0})\nsns.set_context(\"poster\", font_scale=1.1)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"8c27e048-61f9-4c60-bdeb-b03b912b145f","_uuid":"aa43d6db6d8e5baf87d7d4f115789c69acb97c74"},"cell_type":"markdown","source":"I wanted to take a look at the user data we have for this competition so I made this little notebook to share my findings and discuss about those. At the moment I've started with the basic user data, I'll take a look at sessions and the other *csv* files later on this month.\n\nPlease, feel free to comment with anything you think it can be improved or fixed. I am not a professional in this field and there will be mistakes or things that can be *improved*. This is the flow I took and there are some plots not really interesting but I thought on keeping it in case someone see something interesting.\n\nLet's see the data!"},{"metadata":{"_cell_guid":"1efc14a0-4406-4b3b-a10a-f160ca269701","_uuid":"5ef165c3569280ea2ccfd5c8e055c433e4df9ec9"},"cell_type":"markdown","source":"## Data Exploration"},{"metadata":{"_cell_guid":"bc1b841a-c425-455b-a085-a613efa61c2f","_uuid":"d1e0c1ab34068043cd7e64666d96f08bdaea055d"},"cell_type":"markdown","source":"Generally, when I start with a Data Science project I'm looking to answer the following questions:\n\n- Is there any mistakes in the data?\n- Does the data have peculiar behavior?\n- Do I need to fix or remove any of the data to be more realistic?"},{"metadata":{"_cell_guid":"b7cfa2af-274f-462b-8647-29620d42821d","_uuid":"342064d5c667b53648ac8432e48db30462ceacdd","trusted":true},"cell_type":"code","source":"# Load the data into DataFrames\ntrain_users = pd.read_csv('⁩../input/train_users_2.csv')\ntest_users = pd.read_csv('../input/test_users.csv')","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"d70fb51c-102a-42a9-8aaf-51f04583fa65","_uuid":"aacbf2cb6106198c5d1ca0e5bd749dfa8bfbd5f4","trusted":false},"cell_type":"code","source":"print(\"We have\", train_users.shape[0], \"users in the training set and\", \n      test_users.shape[0], \"in the test set.\")\nprint(\"In total we have\", train_users.shape[0] + test_users.shape[0], \"users.\")","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"160f8fe6-f655-4635-807e-8021a8cab7fc","_uuid":"38da91560a565c3eefb67fc9002bae9ab30e3c39"},"cell_type":"markdown","source":"Let's get those together so we can work with all the data."},{"metadata":{"_cell_guid":"136fe63f-99a7-4a5f-b0c0-9857af015b83","_uuid":"c07b61a89b771b0363beb95a66bdd42990ab1fa8","trusted":false},"cell_type":"code","source":"# Merge train and test users\nusers = pd.concat((train_users, test_users), axis=0, ignore_index=True)\n\n# Remove ID's since now we are not interested in making predictions\nusers.drop('id',axis=1, inplace=True)\n\nusers.head()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"229ea3e7-626e-40c1-9b04-96585338efa0","_uuid":"4f5543d3f271c78ff6fbfcea37c44257ccf9a964"},"cell_type":"markdown","source":"The data seems to be in an ussable format so the next important thing is to take a look at the missing data."},{"metadata":{"_cell_guid":"2a3f2858-3853-413f-b820-7b60a642229e","_uuid":"19c2bf7a8d0a16c96f8e0ed3c01078b931fdc7f4"},"cell_type":"markdown","source":"### Missing Data"},{"metadata":{"_cell_guid":"9eb0206b-d462-4423-bb84-3ed6dca3e9d5","_uuid":"84dc7fa0ef5eee9daab0959ff67169f98e47a5fa"},"cell_type":"markdown","source":"Usually the missing data comes in the way of *NaN*, but if we take a look at the DataFrame printed above we can see at the `gender` column some values being `-unknown-`. We will need to transform those values into *NaN* first:"},{"metadata":{"_cell_guid":"0b510172-4df4-42d9-96dd-c2ba70e00595","_uuid":"0104bcc438501e4a1d48576c5df733139bba586f","trusted":false},"cell_type":"code","source":"users.gender.replace('-unknown-', np.nan, inplace=True)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"6ea3c069-62c7-4105-bc14-06e8f3963bbb","_uuid":"d22a47d41ba29bfb0eb26b8407809fead88c4f68"},"cell_type":"markdown","source":"Now let's see how much data we are missing. For this purpose let's compute the NaN percentage of each feature."},{"metadata":{"_cell_guid":"97394712-8970-4eb4-b58c-d14f6cc1c6fd","_uuid":"94133efc4659d2955bb5160a13dd9812c070fecc","trusted":false},"cell_type":"code","source":"users_nan = (users.isnull().sum() / users.shape[0]) * 100\nusers_nan[users_nan > 0].drop('country_destination')","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"3b12334c-2c21-49b3-b3c9-e19bff1d88f3","_uuid":"fa13731412721166a8580497e67ececedb2719a5"},"cell_type":"markdown","source":"We have quite a lot of *NaN* in the `age` and `gender` wich will yield in lesser performance of the classifiers we will build. The feature `date_first_booking` has a 58% of NaN values because this feature is not present at the tests users, and therefore, we won't need it at the *modeling* part."},{"metadata":{"_cell_guid":"e8fa9304-098f-456e-a922-e3d590a112a9","_uuid":"dce772fe9876f0f71b25d877519d31d3d1d14bac","trusted":false},"cell_type":"code","source":"print(\"Just for the sake of curiosity; we have\", \n      int((train_users.date_first_booking.isnull().sum() / train_users.shape[0]) * 100), \n      \"% of missing values at date_first_booking in the training data\")","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"a8a01dd9-c5ab-4f45-8500-4f07d8d8b7c1","_uuid":"66d2d88642700276c9789209bc0201c93099c80f"},"cell_type":"markdown","source":"The other feature with a high rate of *NaN* was `age`. Let's see:"},{"metadata":{"_cell_guid":"ba62ad44-5302-4e7c-8374-d90f02237659","_uuid":"bd9c071277400cbb67f5cbc764526996727688c1","trusted":false},"cell_type":"code","source":"users.age.describe()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"1dcd44ac-61c6-4a1d-a895-8940b1288ef1","_uuid":"86bdb5a79fc8fc8ae56ce7bf55bdec7db99c4d7b"},"cell_type":"markdown","source":"There is some inconsistency in the age of some users as we can see above. It could be because the `age` inpout field was not sanitized or there was some mistakes handlig the data."},{"metadata":{"_cell_guid":"b1527f23-34fb-46e5-9349-27c558572662","_uuid":"539ca16fbb11cbf47ec8301b45567c09a77ad3ec","trusted":false},"cell_type":"code","source":"print(sum(users.age > 122))\nprint(sum(users.age < 18))","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"4022b543-ad8b-447b-bfe8-76c66f8949ff","_uuid":"d207206bf46ddafcd5c1614dbc36f09e081fa8bd"},"cell_type":"markdown","source":"So far, do we have 801 users with [the longest confirmed human lifespan record](https://en.wikipedia.org/wiki/Jeanne_Calment) and 176 little *gangsters* breaking the [Aribnb Eligibility Terms](https://www.airbnb.com/terms)?"},{"metadata":{"_cell_guid":"303a381d-6380-4e3c-8e6f-fd58c48fa322","_uuid":"61dec3883b624e4e6a7a93308713fa819ee95108","trusted":false},"cell_type":"code","source":"users[users.age > 122]['age'].describe()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"0f1042a9-a35c-4187-9f1a-f368686e47df","_uuid":"a449197e9c83adb2e5864878d79a0d54e0423101"},"cell_type":"markdown","source":"It's seems that the weird values are caused by the appearance of 2014. I didn't figured why, but I supose that might be related with a wrong input being added with the new users."},{"metadata":{"_cell_guid":"6901f9b3-e1bf-479a-9933-3293d5efc794","_uuid":"ceee3ffbd5178ba96d71bd3a150957e6b5bae0a5","trusted":false},"cell_type":"code","source":"users[users.age < 18]['age'].describe()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"bee598b5-863b-4b7c-a542-319fb24c7c75","_uuid":"385da813d1cc7dd35420ccb5da0d11e4197213dd"},"cell_type":"markdown","source":"The young users seems to be under an acceptable range being the 50% of those users above 16 years old. \nWe will need to hande the outliers. The simple thing that came to my mind it's to set an acceptance range and put those out of it to NaN."},{"metadata":{"_cell_guid":"a216acbc-3c31-40da-ba54-2dcc57cfc624","_uuid":"d14234386c6f5bd507283d00dd93206807979ed1","trusted":false},"cell_type":"code","source":"users.loc[users.age > 95, 'age'] = np.nan\nusers.loc[users.age < 13, 'age'] = np.nan","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"51dbd932-79e1-4788-9264-562673ad535d","_uuid":"e2b5a7cc550a706ca7e4318f9f9bb616b387f88f"},"cell_type":"markdown","source":"### Data Types"},{"metadata":{"_cell_guid":"be9fbea6-09ea-43f4-bda4-db5e6ab41133","_uuid":"2876afa6d582c0617170ce06dee067506af3ab47"},"cell_type":"markdown","source":"Let's treat each feature as what they are. This means we need to transform into categorical those features that we treas as categories and the same with the dates:"},{"metadata":{"_cell_guid":"dadd5f3f-b389-46a6-8dc2-56e3c31efe84","_uuid":"1afa0bbce6c52c170ceb1ee54dc3c8a63d878130","trusted":false},"cell_type":"code","source":"categorical_features = [\n    'affiliate_channel',\n    'affiliate_provider',\n    'country_destination',\n    'first_affiliate_tracked',\n    'first_browser',\n    'first_device_type',\n    'gender',\n    'language',\n    'signup_app',\n    'signup_method'\n]\n\nfor categorical_feature in categorical_features:\n    users[categorical_feature] = users[categorical_feature].astype('category')","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"8061aca3-be0f-4828-9e0e-131408ed4b21","_uuid":"3edf488ce0cd56fc8bb2cb0232480576a22c7f80","trusted":false},"cell_type":"code","source":"users['date_account_created'] = pd.to_datetime(users['date_account_created'])\nusers['date_first_booking'] = pd.to_datetime(users['date_first_booking'])\nusers['date_first_active'] = pd.to_datetime((users.timestamp_first_active // 1000000), format='%Y%m%d')","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"783094ac-bd01-4242-a9ff-432ce355c006","_uuid":"ca3f2e5d5e3b7864b8e3254f1d63421a1b31e19b"},"cell_type":"markdown","source":"### Visualizing the Data"},{"metadata":{"_cell_guid":"1da0385d-0683-4e80-bd89-f363310c2a54","_uuid":"153b2d3bb5a987b1c3d83389fd9386eb5dedbd2b"},"cell_type":"markdown","source":"Usually, looking at tables, percentiles, means, and other several measures at this state is rarely useful unless you know very well your data.\n\nFor me, it's usually better to visualize the data in some way. Visualization makes me see the outliers and errors immediately!"},{"metadata":{"_cell_guid":"c2aa319c-4b2a-420a-983e-59dc5f873cd4","_uuid":"52f1b22ec4d616686f4187616d8213b44db7f84e"},"cell_type":"markdown","source":"#### Gender"},{"metadata":{"_cell_guid":"f5daecf1-d75e-490c-a25e-67b33fb65871","_uuid":"08eefc3ca539f93d9778f0ce074d3772d8f3b038","trusted":false},"cell_type":"code","source":"users.gender.value_counts(dropna=False).plot(kind='bar', color='#FD5C64', rot=0)\nplt.xlabel('Gender')\nsns.despine()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"feeeb90a-7e1d-40f8-b847-b48105e141a2","_uuid":"0951a3eb85d12c0870a5aceaf2d8a544a6ff58a5"},"cell_type":"markdown","source":"As we've seen before at this plot we can see the ammount of missing data in perspective. Also, notice that there is a slight difference between user gender.\n\nNext thing it might be interesting to see if there is any gender preferences when travelling:"},{"metadata":{"_cell_guid":"e09d8fe0-f74f-4711-8125-bbdf9b951ed8","_uuid":"be3294eb5580837247048641f1ddce8f33f8b176","trusted":false},"cell_type":"code","source":"women = sum(users['gender'] == 'FEMALE')\nmen = sum(users['gender'] == 'MALE')\n\nfemale_destinations = users.loc[users['gender'] == 'FEMALE', 'country_destination'].value_counts() / women * 100\nmale_destinations = users.loc[users['gender'] == 'MALE', 'country_destination'].value_counts() / men * 100\n\n# Bar width\nwidth = 0.4\n\nmale_destinations.plot(kind='bar', width=width, color='#4DD3C9', position=0, label='Male', rot=0)\nfemale_destinations.plot(kind='bar', width=width, color='#FFA35D', position=1, label='Female', rot=0)\n\nplt.legend()\nplt.xlabel('Destination Country')\nplt.ylabel('Percentage')\n\nsns.despine()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"53296296-fd46-4f5a-b84a-ab7612e13919","_uuid":"c405940922d181ef6ee02ad9fcb12b1fd482b6de"},"cell_type":"markdown","source":"There are no big differences between the 2 main genders, so this plot it's not really ussefull except to know the relative destination frecuency of the countries. Let's see it clear here:"},{"metadata":{"_cell_guid":"4195f02a-4731-4a08-b180-aebb7fc537bc","_uuid":"bed781ca26c067e8beada47e1d0cc073c33809a3","trusted":false},"cell_type":"code","source":"destination_percentage = users.country_destination.value_counts() / users.shape[0] * 100\ndestination_percentage.plot(kind='bar',color='#FD5C64', rot=0)\n# Using seaborn can also be plotted\n# sns.countplot(x=\"country_destination\", data=users, order=list(users.country_destination.value_counts().keys()))\nplt.xlabel('Destination Country')\nplt.ylabel('Percentage')\nsns.despine()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"637e1e7d-a4b6-4352-8bc3-173d18202704","_uuid":"3deb1eaddf1e2a9f5a59b853c354424e9a6aa729"},"cell_type":"markdown","source":"The first thing we can see that if there is a reservation, it's likely to be inside the US. But there is a 45% of people that never did a reservation."},{"metadata":{"_cell_guid":"b9c6297b-469e-4ea3-bbe2-5ac820e73af9","_uuid":"0e60f98cf249dc2d7ca5242994026818528e70a2"},"cell_type":"markdown","source":"#### Age"},{"metadata":{"_cell_guid":"4868eb4f-3589-4aa3-8340-86e351399634","_uuid":"d5b127f34a78fdb32f74b16f3af7fdd4975ea2d7"},"cell_type":"markdown","source":"Now that I know there is no difference between male and female reservations at first sight I'll dig into the age."},{"metadata":{"_cell_guid":"995c078d-d050-423c-9ce6-13b8bb391c9f","_uuid":"303485b23ad0a7ce6df526bf8723853b23743d70","trusted":false},"cell_type":"code","source":"sns.distplot(users.age.dropna(), color='#FD5C64')\nplt.xlabel('Age')\nsns.despine()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"3da9e81e-7df2-4a25-b420-23664220f71a","_uuid":"e8d447bc8a7866b58e5b1374e0dd3f4a7e5b31c3"},"cell_type":"markdown","source":"As expected, the common age to travel is between 25 and 40. Let's see if, for example, older people travel in a different way. Let's pick an arbitrary age to split into two groups. Maybe 45?"},{"metadata":{"_cell_guid":"47d57b08-e22e-4316-8b3c-f238a5823b65","_uuid":"18a9703d789df0fea2ac724145a98e110a3cd297","trusted":false},"cell_type":"code","source":"age = 45\n\nyounger = sum(users.loc[users['age'] < age, 'country_destination'].value_counts())\nolder = sum(users.loc[users['age'] > age, 'country_destination'].value_counts())\n\nyounger_destinations = users.loc[users['age'] < age, 'country_destination'].value_counts() / younger * 100\nolder_destinations = users.loc[users['age'] > age, 'country_destination'].value_counts() / older * 100\n\nyounger_destinations.plot(kind='bar', width=width, color='#63EA55', position=0, label='Youngers', rot=0)\nolder_destinations.plot(kind='bar', width=width, color='#4DD3C9', position=1, label='Olders', rot=0)\n\nplt.legend()\nplt.xlabel('Destination Country')\nplt.ylabel('Percentage')\n\nsns.despine()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"7abbc588-59dc-44c0-9cb3-f9adb08df05d","_uuid":"48f7017faa6532baaefe4d1b9b0d139d62f3ec16"},"cell_type":"markdown","source":"We can see that the young people tends to stay in the US, and the older people choose to travel outside the country. Of vourse, there are no big differences between them and we must remember that we do not have the 42% of the ages. \n\nThe first thing I thought when reading the problem was the importance of the native lenguage when choosing the destination country. So let's see how manny users use english as main language:"},{"metadata":{"_cell_guid":"6532c13c-a8c7-4e87-9e04-1b3ffdb9f149","_uuid":"07b173e4f5395104960537c82b261fe8b4a0a303","trusted":false},"cell_type":"code","source":"print((sum(users.language == 'en') / users.shape[0])*100)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"219b0e1a-8fe6-41f5-88fd-0d88b6d9415c","_uuid":"09c2526d2e578a74077578a87ae0eaf68280f945"},"cell_type":"markdown","source":"With the 96% of users using English as their language, it is understandable that a lot of people stay in the US. Someone maybe thinking, if the language is important, why not travel to GB? We need to remember that there is also a lot of factor we are not acounting so making assumpions or predictions like that might be dangerous."},{"metadata":{"_cell_guid":"f888b123-0c9c-4781-a68b-686067d41315","_uuid":"ab8e590ed01f30cc757e8c274c75d8914341b124"},"cell_type":"markdown","source":"#### Dates"},{"metadata":{"_cell_guid":"49baa414-d26d-437b-acf9-9f86e0d323bc","_uuid":"463bfe68f37fb3db4fcd1d3c627d361658b574ab"},"cell_type":"markdown","source":"To see the dates of our users and the timespan of them, let's plot the number of accounts created by time:"},{"metadata":{"_cell_guid":"05d9ed73-82c7-4e57-a789-b60b2ba613e9","_uuid":"52e89a8a211c8c0a16f20da604e7a23ebff699ac","trusted":false},"cell_type":"code","source":"sns.set_style(\"whitegrid\", {'axes.edgecolor': '0'})\nsns.set_context(\"poster\", font_scale=1.1)\nusers.date_account_created.value_counts().plot(kind='line', linewidth=1.2, color='#FD5C64')","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"94c8f5d2-e1c9-49a1-a0da-1cec876200d3","_uuid":"4f98079ded2634c45def003e0a758fe01341eae6"},"cell_type":"markdown","source":"It's appreciable how fast Airbnb has grown over the last 3 years. Does this correlate with the date when the user was active for the first time? It should be very similar, so doing this is a way to check the data!"},{"metadata":{"_cell_guid":"739183b4-9afa-4b88-8ad9-bc0c756ff498","_uuid":"2d0408da67a044f77396b6a42204650b756d653b","trusted":false},"cell_type":"code","source":"users.date_first_active.value_counts().plot(kind='line', linewidth=1.2, color='#FD5C64')","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"b298473c-20e6-426f-bd9a-650e081981ec","_uuid":"128e0dd4a7e7a59587458aa03c967fa778755a74"},"cell_type":"markdown","source":"We can se that's almost the same as `date_account_created`, and also, notice the small peaks. We can, either smooth the graph or dig into those peaks. Let's dig in:"},{"metadata":{"_cell_guid":"218b0fa8-0e1c-4ef1-80cb-47f53f689989","_uuid":"d24a03a8525f2b1cc19585d21ecf511439cfd15b","trusted":false},"cell_type":"code","source":"users_2013 = users[users['date_first_active'] > pd.to_datetime(20130101, format='%Y%m%d')]\nusers_2013 = users_2013[users_2013['date_first_active'] < pd.to_datetime(20140101, format='%Y%m%d')]\nusers_2013.date_first_active.value_counts().plot(kind='line', linewidth=2, color='#FD5C64')\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"ae17d3fd-3d19-4f4a-9db9-b513bbcf30a8","_uuid":"75a0d09ea6d41d6f83f12d4c66d1088a121bce27"},"cell_type":"markdown","source":"At first sight we can see a small pattern, there are some peaks at the same distance. Looking more closely:"},{"metadata":{"_cell_guid":"fd53f9b6-72ca-4f64-9efb-0540e6477a9d","_uuid":"f83968385e5c8ec958beed377521db3b29cc1c91","trusted":false},"cell_type":"code","source":"weekdays = []\nfor date in users.date_account_created:\n    weekdays.append(date.weekday())\nweekdays = pd.Series(weekdays)","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"c4ef4f1e-2592-452a-a93f-0f9d507f8701","_uuid":"af39a3562478d16c922ed1206a7f2711a6bffcbf","trusted":false},"cell_type":"code","source":"sns.barplot(x = weekdays.value_counts().index, y=weekdays.value_counts().values, order=range(0,7))\nplt.xlabel('Week Day')\nsns.despine()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"21727864-1c2d-47e5-a394-45f434c9d978","_uuid":"7c0f49e412772c5595a83becf6e034caf395dc02"},"cell_type":"markdown","source":"The local minimums where the Sundays(where the people use less *the Internet*), and it's usually to hit a maximum at Tuesdays!\n\nThe last date related plot I want to see is the next:"},{"metadata":{"_cell_guid":"507cb13b-6e49-44ad-8b2d-d177293691c3","_uuid":"7ec0d9db1ed6a139fcc974e1f522c6371911c4c2","trusted":false},"cell_type":"code","source":"date = pd.to_datetime(20140101, format='%Y%m%d')\n\nbefore = sum(users.loc[users['date_first_active'] < date, 'country_destination'].value_counts())\nafter = sum(users.loc[users['date_first_active'] > date, 'country_destination'].value_counts())\nbefore_destinations = users.loc[users['date_first_active'] < date, \n                                'country_destination'].value_counts() / before * 100\nafter_destinations = users.loc[users['date_first_active'] > date, \n                               'country_destination'].value_counts() / after * 100\nbefore_destinations.plot(kind='bar', width=width, color='#63EA55', position=0, label='Before 2014', rot=0)\nafter_destinations.plot(kind='bar', width=width, color='#4DD3C9', position=1, label='After 2014', rot=0)\n\nplt.legend()\nplt.xlabel('Destination Country')\nplt.ylabel('Percentage')\n\nsns.despine()\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"e4d733f7-93f1-4561-9591-80c0050077a3","_uuid":"40aeb7d8faf1f03040cec3a844fc7450dc4d5854"},"cell_type":"markdown","source":"It's a clean comparision of usual destinations then and now, where we can see how the new users, register more and book less, and when they book they stay at the US."},{"metadata":{"_cell_guid":"aae61139-bbc0-43b3-82e5-91305005db62","_uuid":"a02ae3b98e53ae4902d62ad3a8fa2d435550eec8"},"cell_type":"markdown","source":"I'll make more plots about the devices and singup methods/flow later this week. I hope you all have enjoyed this little analysis that despine not being very rellevant to make the predictions, it is to understand the problem and the user behaviour. \n\nAgain, criticism is welcomed!\n\n                                                                                David Gasquez"}],"metadata":{"_change_revision":0,"_is_fork":false,"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}