{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## H&M Personalized Fashion Recommendations\nIn this competition, H&M Group invited us to develop product recommendations based on data from previous transactions, as well as from customer and product meta data. The available meta data spans from simple data, such as garment type and customer age, to text data from product descriptions, to image data from garment images.\n\n## My approach\nHere with this given data I am going to approach EDA concept. Before proceeding let me tell you... what is EDA?\n\nExploratory Data Analysis: this is unavoidable and one of the major step to fine-tune the given data set(s) in a different form of analysis to understand the insights of the key characteristics of various entities of the data set like column(s), row(s) by applying Pandas, NumPy, Statistical Methods, and Data visualization packages. ","metadata":{}},{"cell_type":"code","source":"#Importing all the required liabraries\nimport pandas as pd\nimport numpy as np\nimport os\n\nimport sys, warnings, time, os, copy, gc, re, random, pickle\nwarnings.filterwarnings('ignore')\nfrom IPython.display import display\n\n\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport seaborn as sns\nimport matplotlib.image as mpimg\nsns.set()\nfrom pandas.io.json import json_normalize\nfrom pprint import pprint\nfrom pathlib import Path\nfrom tqdm import tqdm\ntqdm.pandas()\nfrom collections import Counter\nfrom datetime import datetime, timedelta","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:30:06.199769Z","iopub.execute_input":"2022-03-23T03:30:06.200191Z","iopub.status.idle":"2022-03-23T03:30:06.218356Z","shell.execute_reply.started":"2022-03-23T03:30:06.200154Z","shell.execute_reply":"2022-03-23T03:30:06.217409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Importing all the required dataset and analyze the dataset","metadata":{}},{"cell_type":"markdown","source":"**Articles Dataset**","metadata":{}},{"cell_type":"code","source":"#Importing articles dataset\narticles = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/articles.csv\")\narticles.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:30:06.305322Z","iopub.execute_input":"2022-03-23T03:30:06.305896Z","iopub.status.idle":"2022-03-23T03:30:07.433236Z","shell.execute_reply.started":"2022-03-23T03:30:06.305855Z","shell.execute_reply":"2022-03-23T03:30:07.431968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Analyzing the columns and the data types\nprint(articles.info(), articles.shape)","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:30:07.435232Z","iopub.execute_input":"2022-03-23T03:30:07.435493Z","iopub.status.idle":"2022-03-23T03:30:07.519147Z","shell.execute_reply.started":"2022-03-23T03:30:07.435462Z","shell.execute_reply":"2022-03-23T03:30:07.518310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Analyze the grament sections with the product group name and index group name\nplt.subplots(figsize=(15,10))\nax=sns.histplot(data=articles, y='product_group_name',hue='index_group_name', multiple=\"stack\")\nax.set_xlabel('counts')\nax.set_ylabel('group name')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:30:07.520922Z","iopub.execute_input":"2022-03-23T03:30:07.521544Z","iopub.status.idle":"2022-03-23T03:30:08.326382Z","shell.execute_reply.started":"2022-03-23T03:30:07.521497Z","shell.execute_reply":"2022-03-23T03:30:08.325284Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From above we can see that the purchase of 'Garment upper body', 'Garment lower body' and 'Garment full body' is heigher.","metadata":{}},{"cell_type":"code","source":"#Analyze the grament sections with the product group\nplt.subplots(figsize=(15,10))\nax=sns.histplot(data=articles, y='index_name')\nax.set_xlabel('counts')\nax.set_ylabel('index name')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:30:08.329103Z","iopub.execute_input":"2022-03-23T03:30:08.329724Z","iopub.status.idle":"2022-03-23T03:30:08.704866Z","shell.execute_reply.started":"2022-03-23T03:30:08.329659Z","shell.execute_reply":"2022-03-23T03:30:08.703647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From above we can see that 'Ladieswear' selling is leading in this 'Index Name' section. But also thhere are some sub-groups for index group. Lets analyze that also.","metadata":{}},{"cell_type":"code","source":"articles.groupby(['index_group_name', 'index_name']).count()['article_id']","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:30:08.706236Z","iopub.execute_input":"2022-03-23T03:30:08.706619Z","iopub.status.idle":"2022-03-23T03:30:08.811675Z","shell.execute_reply.started":"2022-03-23T03:30:08.706533Z","shell.execute_reply":"2022-03-23T03:30:08.810636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Similarly we can see the subgroups in prduct group and product index also.","metadata":{}},{"cell_type":"code","source":"articles.groupby(['product_group_name','product_type_name']).count()['article_id']","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:30:08.813357Z","iopub.execute_input":"2022-03-23T03:30:08.813743Z","iopub.status.idle":"2022-03-23T03:30:08.922301Z","shell.execute_reply.started":"2022-03-23T03:30:08.813694Z","shell.execute_reply":"2022-03-23T03:30:08.921270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And look at the product group-product structure. Accessories are really various, the most numerious: bags, earrings and hats. However, trousers prevail.","metadata":{}},{"cell_type":"code","source":"articles.describe()","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:30:08.923788Z","iopub.execute_input":"2022-03-23T03:30:08.924115Z","iopub.status.idle":"2022-03-23T03:30:08.997592Z","shell.execute_reply.started":"2022-03-23T03:30:08.924070Z","shell.execute_reply":"2022-03-23T03:30:08.996664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Transactions Dataset**","metadata":{}},{"cell_type":"code","source":"#Importing transactions dataset\ntrans=pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\")\ntrans.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:30:08.999390Z","iopub.execute_input":"2022-03-23T03:30:08.999723Z","iopub.status.idle":"2022-03-23T03:31:00.328367Z","shell.execute_reply.started":"2022-03-23T03:30:08.999686Z","shell.execute_reply":"2022-03-23T03:31:00.327382Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Analyzing the columns and the data types\nprint(trans.info(), trans.shape)","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:31:00.330036Z","iopub.execute_input":"2022-03-23T03:31:00.330422Z","iopub.status.idle":"2022-03-23T03:31:00.343694Z","shell.execute_reply.started":"2022-03-23T03:31:00.330383Z","shell.execute_reply":"2022-03-23T03:31:00.342450Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Lets analyze the transactions completed by the customer\nTrans_Per_Customer = trans.groupby('customer_id').count()\nTrans_Per_Customer.sort_values(by='price',ascending=False)['price'][:20]\n","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:31:00.348291Z","iopub.execute_input":"2022-03-23T03:31:00.348988Z","iopub.status.idle":"2022-03-23T03:31:16.450880Z","shell.execute_reply.started":"2022-03-23T03:31:00.348931Z","shell.execute_reply":"2022-03-23T03:31:16.450048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From above we can get the priority customers for the H&M. But from above we may not get the top list product catagories for our priority customers. Lets merge Articles and Transaction datasets to get a better idea ","metadata":{}},{"cell_type":"code","source":"trans.shape","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:31:16.452069Z","iopub.execute_input":"2022-03-23T03:31:16.452449Z","iopub.status.idle":"2022-03-23T03:31:16.458925Z","shell.execute_reply.started":"2022-03-23T03:31:16.452418Z","shell.execute_reply":"2022-03-23T03:31:16.457312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"art_sub=articles[['article_id','prod_name','product_type_name','product_group_name','index_name']]\ntrans_art=trans[['t_dat','customer_id','article_id','price']]\ntrans_art=trans_art.merge(art_sub,on='article_id', how='left')\ntrans_art.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:31:16.460505Z","iopub.execute_input":"2022-03-23T03:31:16.460770Z","iopub.status.idle":"2022-03-23T03:31:29.438727Z","shell.execute_reply.started":"2022-03-23T03:31:16.460740Z","shell.execute_reply":"2022-03-23T03:31:29.437443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"trans_art_cust=trans_art.groupby('customer_id').count()","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:31:29.440120Z","iopub.execute_input":"2022-03-23T03:31:29.440518Z","iopub.status.idle":"2022-03-23T03:31:58.259470Z","shell.execute_reply.started":"2022-03-23T03:31:29.440476Z","shell.execute_reply":"2022-03-23T03:31:58.258330Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles_index = trans_art[['index_name', 'price']].groupby('index_name').mean()\nsns.set_style(\"darkgrid\")\nf, ax = plt.subplots(figsize=(10,5))\nax = sns.barplot(x=articles_index.price, y=articles_index.index, color='orange', alpha=0.8)\nax.set_xlabel('Price by index')\nax.set_ylabel('Index')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:31:58.261548Z","iopub.execute_input":"2022-03-23T03:31:58.261943Z","iopub.status.idle":"2022-03-23T03:32:01.892677Z","shell.execute_reply.started":"2022-03-23T03:31:58.261896Z","shell.execute_reply":"2022-03-23T03:32:01.891468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The index with the highest mean price is Ladieswear. With the lowest - children.","metadata":{}},{"cell_type":"code","source":"articles_index = trans_art[['product_group_name', 'price']].groupby('product_group_name').mean()\nsns.set_style(\"darkgrid\")\nf, ax = plt.subplots(figsize=(10,5))\nax = sns.barplot(x=articles_index.price, y=articles_index.index, color='orange', alpha=0.8)\nax.set_xlabel('Price by product group')\nax.set_ylabel('Product group')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:32:01.894244Z","iopub.execute_input":"2022-03-23T03:32:01.894485Z","iopub.status.idle":"2022-03-23T03:32:05.707696Z","shell.execute_reply.started":"2022-03-23T03:32:01.894456Z","shell.execute_reply":"2022-03-23T03:32:05.706790Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"**Customers Dataset**","metadata":{}},{"cell_type":"code","source":"#importing the customer dataset\ncustomers = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/customers.csv\")\nprint('shape : ',customers.shape)\ncustomers.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:32:05.708931Z","iopub.execute_input":"2022-03-23T03:32:05.709259Z","iopub.status.idle":"2022-03-23T03:32:10.903712Z","shell.execute_reply.started":"2022-03-23T03:32:05.709213Z","shell.execute_reply":"2022-03-23T03:32:10.902628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#We can check if we haave any duplicate data for any customer\ncustomers['customer_id'].shape[0]-customers['customer_id'].nunique()","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:32:10.905351Z","iopub.execute_input":"2022-03-23T03:32:10.905639Z","iopub.status.idle":"2022-03-23T03:32:11.559623Z","shell.execute_reply.started":"2022-03-23T03:32:10.905594Z","shell.execute_reply":"2022-03-23T03:32:11.558366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#We may have many customers for a single postal code. lets analyze that\npostal_cust=customers.groupby('postal_code').count().sort_values('customer_id',ascending=False)\npostal_cust.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:32:11.561394Z","iopub.execute_input":"2022-03-23T03:32:11.561770Z","iopub.status.idle":"2022-03-23T03:32:13.517690Z","shell.execute_reply.started":"2022-03-23T03:32:11.561721Z","shell.execute_reply":"2022-03-23T03:32:13.516612Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#With this customers table we also can get a clear idea of customer age group \nplt.subplots(figsize=(25,25))\nca=sns.histplot(data=customers,x='age')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:32:13.519119Z","iopub.execute_input":"2022-03-23T03:32:13.519363Z","iopub.status.idle":"2022-03-23T03:32:15.399610Z","shell.execute_reply.started":"2022-03-23T03:32:13.519334Z","shell.execute_reply":"2022-03-23T03:32:15.398391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From above graph we can clearly understand the most of our customers are from 18-28 and also 45-55.","metadata":{}},{"cell_type":"code","source":"#We have one columns where we can get the number of sutomers with club member status. lets analyse the data\nplt.subplots(figsize=(10,10))\nca=sns.histplot(data=customers,x='club_member_status')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:32:15.401013Z","iopub.execute_input":"2022-03-23T03:32:15.401691Z","iopub.status.idle":"2022-03-23T03:32:16.503944Z","shell.execute_reply.started":"2022-03-23T03:32:15.401645Z","shell.execute_reply":"2022-03-23T03:32:16.502650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From above we can clearly undersatnd that most of our customer has active cumber status","metadata":{}},{"cell_type":"code","source":"#Do our customer like the notification that we send?\nplt.subplots(figsize=(10,10))\nca=sns.histplot(data=customers,x='fashion_news_frequency')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:32:16.505680Z","iopub.execute_input":"2022-03-23T03:32:16.506041Z","iopub.status.idle":"2022-03-23T03:32:17.627727Z","shell.execute_reply.started":"2022-03-23T03:32:16.505997Z","shell.execute_reply":"2022-03-23T03:32:17.626562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"H&M need to check with the fashin notifiction team as most of our customer dont like the notificatios","metadata":{}},{"cell_type":"markdown","source":"**Images with description and price**","metadata":{}},{"cell_type":"code","source":"#Lets check our higher range clothes\nmax_price_ids = trans[trans.t_dat==trans.t_dat.max()].sort_values('price', ascending=False).iloc[:5][['article_id', 'price']]\n\nf, ax = plt.subplots(1, 5, figsize=(20,10))\ni = 0\nfor _, data in max_price_ids.iterrows():\n    desc = articles[articles['article_id'] == data['article_id']]['detail_desc'].iloc[0]\n    desc_list = desc.split(' ')\n    for j, elem in enumerate(desc_list):\n        if j > 0 and j % 5 == 0:\n            desc_list[j] = desc_list[j] + '\\n'\n    desc = ' '.join(desc_list)\n    img = mpimg.imread(f'../input/h-and-m-personalized-fashion-recommendations/images/0{str(data.article_id)[:2]}/0{int(data.article_id)}.jpg')\n    ax[i].imshow(img)\n    ax[i].set_title(f'price: {data.price:.2f}')\n    ax[i].set_xticks([], [])\n    ax[i].set_yticks([], [])\n    ax[i].grid(False)\n    ax[i].set_xlabel(desc, fontsize=10)\n    i += 1\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:32:17.629299Z","iopub.execute_input":"2022-03-23T03:32:17.629523Z","iopub.status.idle":"2022-03-23T03:32:23.140365Z","shell.execute_reply.started":"2022-03-23T03:32:17.629497Z","shell.execute_reply":"2022-03-23T03:32:23.139357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#ets check our lower range clothes\nmin_price_ids = trans[trans.t_dat==trans.t_dat.min()].sort_values('price', ascending=True).iloc[:5][['article_id', 'price']]\n\nf, ax = plt.subplots(1, 5, figsize=(20,10))\ni = 0\nfor _, data in min_price_ids.iterrows():\n    desc = articles[articles['article_id'] == data['article_id']]['detail_desc'].iloc[0]\n    desc_list = desc.split(' ')\n    for j, elem in enumerate(desc_list):\n        if j > 0 and j % 4 == 0:\n            desc_list[j] = desc_list[j] + '\\n'\n    desc = ' '.join(desc_list)\n    img = mpimg.imread(f'../input/h-and-m-personalized-fashion-recommendations/images/0{str(data.article_id)[:2]}/0{int(data.article_id)}.jpg')\n    ax[i].imshow(img)\n    ax[i].set_title(f'price: {data.price:.4f}')\n    ax[i].set_xlabel(desc, fontsize=10)\n    ax[i].set_xticks([], [])\n    ax[i].set_yticks([], [])\n    ax[i].grid(False)\n    i += 1\nplt.axis('off')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:32:23.142110Z","iopub.execute_input":"2022-03-23T03:32:23.142515Z","iopub.status.idle":"2022-03-23T03:32:28.847128Z","shell.execute_reply.started":"2022-03-23T03:32:23.142461Z","shell.execute_reply":"2022-03-23T03:32:28.846106Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Predictions**\n","metadata":{}},{"cell_type":"code","source":"#trans['t_dat'] = pd.to_datetime(trans['t_dat'])\n#trans.set_index('t_dat', inplace=True)\n#trans=pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:32:28.848527Z","iopub.execute_input":"2022-03-23T03:32:28.848846Z","iopub.status.idle":"2022-03-23T03:32:28.852973Z","shell.execute_reply.started":"2022-03-23T03:32:28.848808Z","shell.execute_reply":"2022-03-23T03:32:28.852013Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"listBin = [-1, 19, 29, 39, 49, 59, 69, 119]\ncustomers['age_bins'] = pd.cut(customers['age'], listBin)\nN = 12\nlistUniBins = customers['age_bins'].unique().tolist()\nfor uniBin in listUniBins:\n    df  = trans[['t_dat', 'customer_id', 'article_id']]\n    df['customer_id'].astype('string')\n    if str(uniBin) == 'nan':\n        customersTemp = customers[customers['age_bins'].isnull()]\n    else:\n        customersTemp = customers[customers['age_bins'] == uniBin]\n    \n    customersTemp = customersTemp.drop(['age_bins'], axis=1)\n    #customersTemp = pd.from_pandas(customersTemp)\n    \n    df = df.merge(customersTemp[['customer_id', 'age']], on='customer_id', how='inner')\n    print(f'The shape of scope transaction for {uniBin} is {df.shape}. \\n')\n    hex_to_int = lambda x: int(x, 16)\n    #df[['A', 'B', 'C']] = df[['A', 'B', 'C']].applymap(hex_to_int)\n    #df ['customer_id'] = df ['customer_id'].str[-16:].astype('int64')\n    df ['customer_id'] = df ['customer_id'].apply(lambda x: int(x, base=16))\n    df['t_dat'] = pd.to_datetime(df['t_dat'])\n   \n    last_ts = df['t_dat'].max()\n\n    tmp = df[['t_dat']]\n    tmp['dow'] = tmp['t_dat'].dt.dayofweek\n    tmp['ldbw'] = tmp['t_dat'] - pd.TimedeltaIndex(tmp['dow'] - 1, unit='D')\n    tmp.loc[tmp['dow'] >=2 , 'ldbw'] = tmp.loc[tmp['dow'] >=2 , 'ldbw'] + pd.TimedeltaIndex(np.ones(len(tmp.loc[tmp['dow'] >=2])) * 7, unit='D')\n\n    df['ldbw'] = tmp['ldbw'].values\n    \n    weekly_sales = df.drop('customer_id', axis=1).groupby(['ldbw', 'article_id']).count().reset_index()\n    weekly_sales = weekly_sales.rename(columns={'t_dat': 'count'})\n    \n    df = df.merge(weekly_sales, on=['ldbw', 'article_id'], how = 'left')\n    \n    weekly_sales = weekly_sales.reset_index().set_index('article_id')\n\n    df = df.merge(\n        weekly_sales.loc[weekly_sales['ldbw']==last_ts, ['count']],\n        on='article_id', suffixes=(\"\", \"_targ\"))\n\n    df['count_targ'].fillna(0, inplace=True)\n    del weekly_sales\n    \n    df['quotient'] = df['count_targ'] / df['count']\n    \n    target_sales = df.drop('customer_id', axis=1).groupby('article_id')['quotient'].sum()\n    general_pred = target_sales.nlargest(N).index.tolist()\n    general_pred = ['0' + str(article_id) for article_id in general_pred]\n    general_pred_str =  ' '.join(general_pred)\n    del target_sales\n    \n    purchase_dict = {}\n\n    tmp = df\n    tmp['x'] = ((last_ts - tmp['t_dat']) / np.timedelta64(1, 'D')).astype(int)\n    tmp['dummy_1'] = 1 \n    tmp['x'] = tmp[[\"x\", \"dummy_1\"]].max(axis=1)\n\n    a, b, c, d = 2.5e4, 1.5e5, 2e-1, 1e3\n    tmp['y'] = a / np.sqrt(tmp['x']) + b * np.exp(-c*tmp['x']) - d\n\n    tmp['dummy_0'] = 0 \n    tmp['y'] = tmp[[\"y\", \"dummy_0\"]].max(axis=1)\n    tmp['value'] = tmp['quotient'] * tmp['y'] \n\n    tmp = tmp.groupby(['customer_id', 'article_id']).agg({'value': 'sum'})\n    tmp = tmp.reset_index()\n\n    tmp = tmp.loc[tmp['value'] > 0]\n    tmp['rank'] = tmp.groupby(\"customer_id\")[\"value\"].rank(\"dense\", ascending=False)\n    tmp = tmp.loc[tmp['rank'] <= 12]\n\n    purchase_df = tmp.sort_values(['customer_id', 'value'], ascending = False).reset_index(drop = True)\n    purchase_df['prediction'] = '0' + purchase_df['article_id'].astype(str) + ' '\n    purchase_df = purchase_df.groupby('customer_id').agg({'prediction': sum}).reset_index()\n    purchase_df['prediction'] = purchase_df['prediction'].str.strip()\n    purchase_df = pd.DataFrame(purchase_df)\n    \n    sub  = pd.read_csv('../input/h-and-m-personalized-fashion-recommendations/sample_submission.csv',\n                            usecols= ['customer_id'], \n                            dtype={'customer_id': 'string'})\n    \n    numCustomers = sub.shape[0]\n    \n    sub = sub.merge(customersTemp[['customer_id', 'age']], on='customer_id', how='inner')\n\n    #sub['customer_id2'] = sub['customer_id'].str[-16:].str.hex_to_int().astype('int64')\n    sub['customer_id2'] = sub['customer_id']\n    sub = sub.merge(purchase_df, left_on = 'customer_id2', right_on = 'customer_id', how = 'left',\n                   suffixes = ('', '_ignored'))\n\n    #sub = sub.to_pandas()\n    sub['prediction'] = sub['prediction'].fillna(general_pred_str)\n    sub['prediction'] = sub['prediction'] + ' ' +  general_pred_str\n    sub['prediction'] = sub['prediction'].str.strip()\n    sub['prediction'] = sub['prediction'].str[:131]\n    sub = sub[['customer_id', 'prediction']]\n    sub.to_csv(f'submission_' + str(uniBin) + '.csv',index=False)\n    print(f'Saved prediction for {uniBin}. The shape is {sub.shape}. \\n')\n    print('-'*50)\nprint('Finished.\\n')\nprint('='*50)","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:32:28.855014Z","iopub.execute_input":"2022-03-23T03:32:28.855538Z","iopub.status.idle":"2022-03-23T03:36:32.794392Z","shell.execute_reply.started":"2022-03-23T03:32:28.855490Z","shell.execute_reply":"2022-03-23T03:36:32.793419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i, uniBin in enumerate(listUniBins):\n    dfTemp  = pd.read_csv(f'submission_' + str(uniBin) + '.csv')\n    if i == 0:\n        dfSub = dfTemp\n    else:\n        dfSub = pd.concat([dfSub, dfTemp], axis=0)\n\nassert dfSub.shape[0] == numCustomers, f'The number of dfSub rows is not correct. {dfSub.shape[0]} vs {numCustomers}.'\n\ndfSub.to_csv(f'submission.csv', index=False)\nprint(f'Saved submission.csv.')","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:36:32.796103Z","iopub.execute_input":"2022-03-23T03:36:32.796810Z","iopub.status.idle":"2022-03-23T03:36:42.734797Z","shell.execute_reply.started":"2022-03-23T03:36:32.796762Z","shell.execute_reply":"2022-03-23T03:36:42.733683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dfCheck = pd.read_csv('./submission.csv')\ndfCheck.head(5)","metadata":{"execution":{"iopub.status.busy":"2022-03-23T03:36:42.736299Z","iopub.execute_input":"2022-03-23T03:36:42.736661Z","iopub.status.idle":"2022-03-23T03:36:45.301464Z","shell.execute_reply.started":"2022-03-23T03:36:42.736612Z","shell.execute_reply":"2022-03-23T03:36:45.300350Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Thank you for watching my analysis.**","metadata":{}}]}