{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# H&M Recommendation - New Features for Customer ID\n\nThank you for your checking this notebook.\n\nThis is my notebook for \"H&M Personalized Fashion Recommendations\" competition [(Link)](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/overview) to check and add new features to customer ID based on purchasing history. The idea is coming from my this [notebook](https://www.kaggle.com/code/hechtjp/h-m-eda-rule-base-by-customer-age/notebook) which showed the improvement by grouping of customer's age. I would like to check further potential of grouping of customers based on other features.\n\nIf you think this notebook is interesting, please leave your comment or question and I appreciate your upvote as well. :) \n\n<a id='top'></a>\n## Contents\n1. [Import Library & Set Config](#config)\n2. [Load Data](#load)\n3. [Check and add new features to customer ID](#add)\n4. [EDA of recent popular articles in each customer's features](#eda)\n5. [Conclution](#conclution)\n6. [Reference](#ref)","metadata":{}},{"cell_type":"markdown","source":"<a id='config'></a>\n\n---\n## 1. Import Library & Set Config\n---\n\n[Back to Contents](#top)","metadata":{}},{"cell_type":"code","source":"# === General ===\nimport sys, warnings, time, os, copy, gc, re, random, pickle, cudf\nwarnings.filterwarnings('ignore')\nfrom IPython.display import display\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\n# pd.set_option('display.max_rows', 50)\n# pd.set_option('display.max_columns', None)\n# pd.set_option(\"display.max_colwidth\", 10000)\nimport seaborn as sns\nsns.set()\nfrom pandas.io.json import json_normalize\nfrom pprint import pprint\nfrom pathlib import Path\nfrom tqdm import tqdm\ntqdm.pandas()\nfrom collections import Counter\nfrom datetime import datetime, timedelta","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DEBUG = False\nPATH_INPUT = r'../input/h-and-m-personalized-fashion-recommendations/'","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='load'></a>\n\n---\n## 2. Load Data\n---\n\n[Back to Contents](#top)","metadata":{}},{"cell_type":"code","source":"def display_df(df, head=3):\n    print(f'The shape of df is {df.shape}.\\n')\n    display(df.head(head))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dfArticles = cudf.read_csv(PATH_INPUT + 'articles.csv', \n                           usecols=['article_id', \"index_name\", \"perceived_colour_master_name\"],\n                           dtype={'article_id': 'int32', 'index_name': 'string', 'perceived_colour_master_name': 'string'}\n                           )\ndisplay_df(dfArticles, head=3)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dfCustomers = cudf.read_csv(PATH_INPUT + 'customers.csv', \n                            usecols=['customer_id', 'age'],\n                            dtype={'age': 'int16', 'customer_id': 'string'})\n\nlistBin = [-1, 19, 29, 39, 49, 59, 69, 119]\ndfCustomers['age_bins'] = cudf.cut(dfCustomers['age'], listBin)\ndisplay_df(dfCustomers, head=3)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dfTransactions = cudf.read_csv(PATH_INPUT + 'transactions_train.csv',\n                               dtype={'article_id': 'int32', 't_dat': 'string',\n                                      'customer_id': 'string', 'price': 'float32',\n                                      'sales_channel_id': 'string'})\n\ndfTransactions['t_dat'] = cudf.to_datetime(dfTransactions['t_dat'])\ndfTransactions.set_index('t_dat', inplace=True)\ndisplay_df(dfTransactions, head=3)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if DEBUG:\n    dfTransactions = dfTransactions.loc['2020-09-15' : '2020-09-21']\n    print(f'****** Under debugging *****\\n')\n    display_df(dfTransactions, head=3)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='add'></a>\n\n---\n## 3. Check and add new features to customer ID\n\n- Based on purchasing history of each customer ID, create new features.\n- Scope is sales_channel_id, price, index_name & perceived_colour_master_name.\n---\n\n[Back to Contents](#top)","metadata":{}},{"cell_type":"code","source":"dfCustomers = dfCustomers.to_pandas()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Add total number of purchasing.\ndfTemp = dfTransactions.groupby(['customer_id']).count().reset_index()\ndfTemp = dfTemp.to_pandas()\ndfCustomers = dfCustomers.merge(dfTemp[['customer_id', 'article_id']], on='customer_id', how='left')\ndfCustomers = dfCustomers.rename(columns={'article_id': 'count_all'})\ndfCustomers['count_all'] = dfCustomers['count_all'].fillna(0)\ndisplay_df(dfCustomers, head=3)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Add how many % using sales channel 1.\ndfTemp = dfTransactions.groupby(['customer_id', 'sales_channel_id']).count().reset_index()\ndfTemp = dfTemp.to_pandas()\ndfTemp = dfTemp[dfTemp['sales_channel_id'] == '1']\ndfCustomers = dfCustomers.merge(dfTemp[['customer_id', 'article_id']], on='customer_id', how='left')\ndfCustomers = dfCustomers.rename(columns={'article_id': 'count_sales_ch1'})\ndfCustomers['count_sales_ch1'] = dfCustomers['count_sales_ch1'].fillna(0)\n\ndfCustomers['share_sales_ch1'] = dfCustomers['count_sales_ch1'] / dfCustomers['count_all']\ndfCustomers['share_sales_ch1'] = dfCustomers['share_sales_ch1'].fillna(0)\ndisplay_df(dfCustomers, head=3)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Add total purchasing values & avg. purchasing price.\n\ndfTemp = dfTransactions.groupby(['customer_id']).sum().reset_index()\ndfTemp = dfTemp.to_pandas()\ndfCustomers = dfCustomers.merge(dfTemp[['customer_id', 'price']], on='customer_id', how='left')\ndfCustomers = dfCustomers.rename(columns={'price': 'sum_price'})\ndfCustomers['sum_price'] = dfCustomers['sum_price'].fillna(0)\n\ndfTemp = dfTransactions.groupby(['customer_id']).mean().reset_index()\ndfTemp = dfTemp.to_pandas()\ndfCustomers = dfCustomers.merge(dfTemp[['customer_id', 'price']], on='customer_id', how='left')\ndfCustomers = dfCustomers.rename(columns={'price': 'mean_price'})\ndfCustomers['mean_price'] = dfCustomers['mean_price'].fillna(0)\ndisplay_df(dfCustomers, head=3)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dfTransactions = dfTransactions.reset_index().merge(dfArticles, on='article_id', how='left')\ndfTransactions","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Add how many % of black articles customer purchased.\n\ndfTemp = dfTransactions.groupby(['customer_id', 'perceived_colour_master_name']).count().reset_index()\ndfTemp = dfTemp.to_pandas()\ndfTemp = dfTemp[dfTemp['perceived_colour_master_name'] == 'Black']\ndfCustomers = dfCustomers.merge(dfTemp[['customer_id', 'article_id']], on='customer_id', how='left')\ndfCustomers = dfCustomers.rename(columns={'article_id': 'count_Black'})\ndfCustomers['count_Black'] = dfCustomers['count_Black'].fillna(0)\n\ndfCustomers['share_Black'] = dfCustomers['count_Black'] / dfCustomers['count_all']\ndfCustomers['share_Black'] = dfCustomers['share_Black'].fillna(0)\ndisplay_df(dfCustomers, head=3)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Add how many % of white articles customer purchased.\n\ndfTemp = dfTransactions.groupby(['customer_id', 'perceived_colour_master_name']).count().reset_index()\ndfTemp = dfTemp.to_pandas()\ndfTemp = dfTemp[dfTemp['perceived_colour_master_name'] == 'White']\ndfCustomers = dfCustomers.merge(dfTemp[['customer_id', 'article_id']], on='customer_id', how='left')\ndfCustomers = dfCustomers.rename(columns={'article_id': 'count_White'})\ndfCustomers['count_White'] = dfCustomers['count_White'].fillna(0)\n\ndfCustomers['share_White'] = dfCustomers['count_White'] / dfCustomers['count_all']\ndfCustomers['share_White'] = dfCustomers['share_White'].fillna(0)\ndisplay_df(dfCustomers, head=3)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Add how many % of Menswear customer purchased.\n\ndfTemp = dfTransactions.groupby(['customer_id', 'index_name']).count().reset_index()\ndfTemp = dfTemp.to_pandas()\ndfTemp = dfTemp[dfTemp['index_name'] == 'Menswear']\ndfCustomers = dfCustomers.merge(dfTemp[['customer_id', 'article_id']], on='customer_id', how='left')\ndfCustomers = dfCustomers.rename(columns={'article_id': 'count_Menswear'})\ndfCustomers['count_Menswear'] = dfCustomers['count_Menswear'].fillna(0)\n\ndfCustomers['share_Menswear'] = dfCustomers['count_Menswear'] / dfCustomers['count_all']\ndfCustomers['share_Menswear'] = dfCustomers['share_Menswear'].fillna(0)\n\ndisplay_df(dfCustomers, head=3)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Add how many % of Divided customer purchased.\n\ndfTemp = dfTransactions.groupby(['customer_id', 'index_name']).count().reset_index()\ndfTemp = dfTemp.to_pandas()\ndfTemp = dfTemp[dfTemp['index_name'] == 'Divided']\ndfCustomers = dfCustomers.merge(dfTemp[['customer_id', 'article_id']], on='customer_id', how='left')\ndfCustomers = dfCustomers.rename(columns={'article_id': 'count_Divided'})\ndfCustomers['count_Divided'] = dfCustomers['count_Divided'].fillna(0)\n\ndfCustomers['share_Divided'] = dfCustomers['count_Divided'] / dfCustomers['count_all']\ndfCustomers['share_Divided'] = dfCustomers['share_Divided'].fillna(0)\n\ndisplay_df(dfCustomers, head=3)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dfCustomers.to_csv(f'customers_addFeatures.csv', index=False)\nprint(f'Saved customers_addFeatures.csv.')\ndfCustomers.describe()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='eda'></a>\n\n---\n## 4. EDA of recent popular articles in each customer's features\n\n- Check the latest popular articles in each groups based on customer's features btw. 2020-09-01 and 2020-09-21.\n- Compare that whether is there any difference btw. ages.\n\n---\n\n[Back to Contents](#top)","metadata":{}},{"cell_type":"code","source":"# Filtered dfTransactions by target date and merge features from dfCustomers.\n\ndfRecent = dfTransactions.set_index('t_dat').loc['2020-09-01' : '2020-09-21']\ndfRecent = dfRecent.to_pandas()\ndfRecent = dfRecent.merge(dfCustomers[['customer_id', 'age_bins', 'share_sales_ch1', 'sum_price', 'mean_price', 'share_Black', 'share_White', 'share_Menswear', 'share_Divided']], on='customer_id', how='inner')\ndisplay_df(dfRecent, head=3)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create dictionaly of top 100 articles in each features of customers by age bins.\n\nlistUniBins = dfRecent['age_bins'].unique().tolist()\nlistScopes = ['share_sales_ch1', 'sum_price', 'mean_price', 'share_Black', 'share_White', 'share_Menswear', 'share_Divided']\n\ndictAge = {}\nfor uniBin in listUniBins:\n    if str(uniBin) == 'nan':\n        dfTemp = dfRecent[dfRecent['age_bins'].isnull()]\n    else:\n        dfTemp = dfRecent[dfRecent['age_bins'] == uniBin]\n    \n    dictScope = {}\n    for scope in listScopes:\n        dfTemp[scope + '_bins'] = pd.cut(dfTemp[scope], 5)\n        listScopeBins = dfTemp[scope + '_bins'].unique().tolist()\n        dfTemp2 = dfTemp.groupby([scope + '_bins', 'article_id']).count().reset_index().rename(columns={'customer_id': 'counts'})\n        dfTemp2 = dfTemp2.sort_values(by='counts', ascending=False)\n        dict100 = {}\n        for x in listScopeBins:\n            dfTemp3 = dfTemp2[dfTemp2[scope + '_bins'] == x]\n            dict100[x] = dfTemp3.head(100)['article_id'].values.tolist()\n        dictScope[scope] = dict100\n            \n    dictAge[uniBin] = dictScope","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualize how many articles are same btw. each bins of scope features in each age bins.\n\nfor uniBin in listUniBins:\n    for scope in listScopes:\n        dictBins = dictAge[uniBin][scope]\n        df100 = pd.DataFrame([dictBins]).T.rename(columns={0:'top100'})\n        df100 = df100.sort_index()\n        \n        for index in df100.index:\n            df100[index] = [len(set(df100.at[index, 'top100']) & set(df100.at[x, 'top100']))/100 for x in df100.index]\n            \n        df100 = df100.drop(columns='top100')\n        plt.figure(figsize=(10, 6))\n        plt.title(f'age: {uniBin}, scope: {scope}')\n        sns.heatmap(df100, annot=True, cbar=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id='conclution'></a>\n\n---\n\n## 5. Conclution\n\nThank you for your reading through this Notebook!\n\nIf you think this notebook is interesting for you, please do click upvote :)\n\n---\n\n[Back to Contents](#top)","metadata":{}},{"cell_type":"markdown","source":"<a id='ref'></a>\n\n---\n## 6. Reference\n\n---\n\n[Back to Contents](#top)","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}