{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 0.Overview\n**Edit ver 2**:\n\nFor a simple tutorial, this notebook use test data with `full_sort_topk` function. Due to its limit, we are missing last item. So i create new notebooks with custom code, in order to use all items for submitted recommendation:\n* Detail about `full_sort_topk` and custom function: https://www.kaggle.com/astrung/recbole-using-all-items-for-prediction\n* New custom function with interaction data: https://www.kaggle.com/astrung/sequential-model-fixed-missing-last-item\n* New custom function with interaction data + item features: https://www.kaggle.com/code/astrung/lstm-model-with-item-infor-fix-missing-last-item/notebook\n\nPlease check new notebook for better score and upvote if it help\n\n**Edit ver 1**:\n\nI have created a new version for adding user features in sequential model. It looks like by using item features, sequential get more accurate( i got higher score). If you want to improve your sequential model, please check my new notebook: https://www.kaggle.com/astrung/recbole-lstm-sequential-with-item-features\n\n\nThis notebook demonstrate how to use LSTM for recomendation system.\nI am using Recbole as an open source, as it has so many built-in models for recommendation(CNN, GRU-LSTM, Context-aware, Graph). In this notebook, we tried to use GRU/LSTM model for testing effect of sequential model for recommendation.\n\nDue to memory limit and faster testing purpose, we will just use data in 2020.\n\nIf you want to use with all of interactions in all time, i have created a new atomic dataset here for you: https://www.kaggle.com/astrung/hm-atomic-interation\n\nWe also have other limit: we only train model and predict with users who buy more than 40 items and items which is bought by more than 40 people.\n\nWe will follow below steps for creating model:\n\n1. In order to use Recbole, we create atomic file from interaction data\n2. Because we only use Recbole model for predicting with users who buy more than 40 items, other users will need to fill by default recomendation items. We create most viewed items in last month as defautl recomendation\n3. We create dataset and train model in recbole.\n4. We create prediction result by trained model\n5. We combine recomendation result from most viewed items in last month and Recbole predicted model.\n\nI will explain more detail in following cells.\n\n","metadata":{}},{"cell_type":"code","source":"!pip install recbole","metadata":{"execution":{"iopub.status.busy":"2022-03-06T10:47:28.476509Z","iopub.execute_input":"2022-03-06T10:47:28.477117Z","iopub.status.idle":"2022-03-06T10:47:49.605304Z","shell.execute_reply.started":"2022-03-06T10:47:28.476947Z","shell.execute_reply":"2022-03-06T10:47:49.604197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1. Create atomic file","metadata":{}},{"cell_type":"code","source":"import pandas as pd\ndf = pd.read_csv(r\"/kaggle/input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\", \n                 dtype={'article_id': 'str'})\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-06T10:47:49.607884Z","iopub.execute_input":"2022-03-06T10:47:49.608133Z","iopub.status.idle":"2022-03-06T10:48:54.890926Z","shell.execute_reply.started":"2022-03-06T10:47:49.608102Z","shell.execute_reply":"2022-03-06T10:48:54.890053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['t_dat'] = pd.to_datetime(df['t_dat'], format=\"%Y-%m-%d\")\ndf","metadata":{"execution":{"iopub.status.busy":"2022-03-06T10:48:54.892532Z","iopub.execute_input":"2022-03-06T10:48:54.893011Z","iopub.status.idle":"2022-03-06T10:49:01.110911Z","shell.execute_reply.started":"2022-03-06T10:48:54.892914Z","shell.execute_reply":"2022-03-06T10:49:01.109783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\ndf['timestamp'] = df.t_dat.values.astype(np.int64) // 10 ** 9\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-06T10:49:01.113364Z","iopub.execute_input":"2022-03-06T10:49:01.113598Z","iopub.status.idle":"2022-03-06T10:49:01.706201Z","shell.execute_reply.started":"2022-03-06T10:49:01.113555Z","shell.execute_reply":"2022-03-06T10:49:01.704788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**We fill with data in only 2020(timestapm > > 1585620000) and create inter file**\nFor anyone need instruction about inter file, please check below links:\n* https://recbole.io/docs/user_guide/data_intro.html\n* https://recbole.io/docs/user_guide/data/atomic_files.html","metadata":{}},{"cell_type":"code","source":"temp = df[df['timestamp'] > 1585620000][['customer_id', 'article_id', 'timestamp']].rename(\n    columns={'customer_id': 'user_id:token', 'article_id': 'item_id:token', 'timestamp': 'timestamp:float'})\ntemp","metadata":{"execution":{"iopub.status.busy":"2022-03-06T10:49:01.708149Z","iopub.execute_input":"2022-03-06T10:49:01.708799Z","iopub.status.idle":"2022-03-06T10:49:03.476808Z","shell.execute_reply.started":"2022-03-06T10:49:01.708755Z","shell.execute_reply":"2022-03-06T10:49:03.475861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We save atomic file in dataset format for using with recbole","metadata":{}},{"cell_type":"code","source":"!mkdir /kaggle/working/recbox_data\ntemp.to_csv('/kaggle/working/recbox_data/recbox_data.inter', index=False, sep='\\t')","metadata":{"execution":{"iopub.status.busy":"2022-03-06T10:49:03.478539Z","iopub.execute_input":"2022-03-06T10:49:03.478951Z","iopub.status.idle":"2022-03-06T10:49:40.370065Z","shell.execute_reply.started":"2022-03-06T10:49:03.478891Z","shell.execute_reply":"2022-03-06T10:49:40.368832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. We create defautl recomendation for user who can not be predicted by sequential model.\nI use this approach in notebook: https://www.kaggle.com/hervind/h-m-faster-trending-products-weekly You can check it for more detail information. I will juse copy only code here","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd","metadata":{"execution":{"iopub.status.busy":"2022-03-06T10:49:40.371937Z","iopub.execute_input":"2022-03-06T10:49:40.372504Z","iopub.status.idle":"2022-03-06T10:49:44.138941Z","shell.execute_reply.started":"2022-03-06T10:49:40.372444Z","shell.execute_reply":"2022-03-06T10:49:44.137905Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub0 = pd.read_csv('../input/hm-pre-recommendation/submissio_byfone_chris.csv').sort_values('customer_id').reset_index(drop=True)\nsub1 = pd.read_csv('../input/hm-pre-recommendation/submission_trending.csv').sort_values('customer_id').reset_index(drop=True)\nsub2 = pd.read_csv('../input/hm-pre-recommendation/submission_exponential_decay.csv').sort_values('customer_id').reset_index(drop=True)\n\nsub0.shape, sub1.shape, sub2.shape","metadata":{"execution":{"iopub.status.busy":"2022-03-06T10:49:44.140808Z","iopub.execute_input":"2022-03-06T10:49:44.141119Z","iopub.status.idle":"2022-03-06T10:49:49.009148Z","shell.execute_reply.started":"2022-03-06T10:49:44.141079Z","shell.execute_reply":"2022-03-06T10:49:49.008139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub0.columns = ['customer_id', 'prediction0']\nsub0['prediction1'] = sub1['prediction']\nsub0['prediction2'] = sub2['prediction']\ndel sub1, sub2\nsub0.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-06T10:49:49.011039Z","iopub.execute_input":"2022-03-06T10:49:49.011372Z","iopub.status.idle":"2022-03-06T10:50:01.291616Z","shell.execute_reply.started":"2022-03-06T10:49:49.011321Z","shell.execute_reply":"2022-03-06T10:50:01.290486Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def cust_blend(dt, W = [1,1,1]):\n    #Global ensemble weights\n    #W = [1.15,0.95,0.85]\n    \n    #Create a list of all model predictions\n    REC = []\n    REC.append(dt['prediction0'].split())\n    REC.append(dt['prediction1'].split())\n    REC.append(dt['prediction2'].split())\n    \n    #Create a dictionary of items recommended. \n    #Assign a weight according the order of appearance and multiply by global weights\n    res = {}\n    for M in range(len(REC)):\n        for n, v in enumerate(REC[M]):\n            if v in res:\n                res[v] += (W[M]/(n+1))\n            else:\n                res[v] = (W[M]/(n+1))\n    \n    # Sort dictionary by item weights\n    res = list(dict(sorted(res.items(), key=lambda item: -item[1])).keys())\n    \n    # Return the top 12 itens only\n    return ' '.join(res[:12])\n\nsub0['prediction'] = sub0.apply(cust_blend, W = [1.05,1.00,0.95], axis=1)\nsub0.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-06T10:50:01.296021Z","iopub.execute_input":"2022-03-06T10:50:01.296515Z","iopub.status.idle":"2022-03-06T10:50:01.509497Z","shell.execute_reply.started":"2022-03-06T10:50:01.296475Z","shell.execute_reply":"2022-03-06T10:50:01.508564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del sub0['prediction0']\ndel sub0['prediction1']\ndel sub0['prediction2']\nsub0.to_csv(f'submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-03-06T10:50:01.510918Z","iopub.execute_input":"2022-03-06T10:50:01.511968Z","iopub.status.idle":"2022-03-06T10:50:01.619184Z","shell.execute_reply.started":"2022-03-06T10:50:01.511916Z","shell.execute_reply":"2022-03-06T10:50:01.617691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del sub0\ndel df","metadata":{"execution":{"iopub.status.busy":"2022-03-06T10:50:01.620303Z","iopub.execute_input":"2022-03-06T10:50:01.620686Z","iopub.status.idle":"2022-03-06T10:50:01.68948Z","shell.execute_reply.started":"2022-03-06T10:50:01.620652Z","shell.execute_reply":"2022-03-06T10:50:01.688653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. Create dataset and train model with Recbole\n\nFor anyone need instruction document, please check this link: https://recbole.io/docs/user_guide/usage/use_modules.html","metadata":{}},{"cell_type":"code","source":"import logging\nfrom logging import getLogger\nfrom recbole.config import Config\nfrom recbole.data import create_dataset, data_preparation\nfrom recbole.model.sequential_recommender import GRU4Rec\nfrom recbole.trainer import Trainer\nfrom recbole.utils import init_seed, init_logger","metadata":{"execution":{"iopub.status.busy":"2022-03-06T10:50:55.811569Z","iopub.execute_input":"2022-03-06T10:50:55.812326Z","iopub.status.idle":"2022-03-06T10:50:57.73877Z","shell.execute_reply.started":"2022-03-06T10:50:55.81226Z","shell.execute_reply":"2022-03-06T10:50:57.737589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"parameter_dict = {\n    'data_path': '/kaggle/working',\n    'USER_ID_FIELD': 'user_id',\n    'ITEM_ID_FIELD': 'item_id',\n    'TIME_FIELD': 'timestamp',\n    'user_inter_num_interval': \"[30,inf)\",\n    'item_inter_num_interval': \"[40,inf)\",\n    'load_col': {'inter': ['user_id', 'item_id', 'timestamp']},\n    'neg_sampling': None,\n    'epochs': 50,\n    'eval_args': {\n        'split': {'RS': [9, 0, 1]},\n        'group_by': 'user',\n        'order': 'TO',\n        'mode': 'full'}\n}\n\nconfig = Config(model='GRU4Rec', dataset='recbox_data', config_dict=parameter_dict)\n\n# init random seed\ninit_seed(config['seed'], config['reproducibility'])\n\n# logger initialization\ninit_logger(config)\nlogger = getLogger()\n# Create handlers\nc_handler = logging.StreamHandler()\nc_handler.setLevel(logging.INFO)\nlogger.addHandler(c_handler)\n\n# write config info into log\nlogger.info(config)","metadata":{"execution":{"iopub.status.busy":"2022-03-06T10:50:57.740925Z","iopub.execute_input":"2022-03-06T10:50:57.741338Z","iopub.status.idle":"2022-03-06T10:50:58.420824Z","shell.execute_reply.started":"2022-03-06T10:50:57.741292Z","shell.execute_reply":"2022-03-06T10:50:58.419583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset = create_dataset(config)\nlogger.info(dataset)","metadata":{"execution":{"iopub.status.busy":"2022-03-06T10:50:58.422104Z","iopub.execute_input":"2022-03-06T10:50:58.422449Z","iopub.status.idle":"2022-03-06T10:52:27.676728Z","shell.execute_reply.started":"2022-03-06T10:50:58.422406Z","shell.execute_reply":"2022-03-06T10:52:27.674917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# dataset splitting\ntrain_data, valid_data, test_data = data_preparation(config, dataset)","metadata":{"execution":{"iopub.status.busy":"2022-03-06T10:52:27.67935Z","iopub.execute_input":"2022-03-06T10:52:27.680207Z","iopub.status.idle":"2022-03-06T10:52:54.494908Z","shell.execute_reply.started":"2022-03-06T10:52:27.68006Z","shell.execute_reply":"2022-03-06T10:52:54.494078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# model loading and initialization\nmodel = GRU4Rec(config, train_data.dataset).to(config['device'])\nlogger.info(model)\n\n# trainer loading and initialization\ntrainer = Trainer(config, model)\n\n# model training\nbest_valid_score, best_valid_result = trainer.fit(train_data)","metadata":{"scrolled":true,"execution":{"iopub.status.busy":"2022-03-06T10:52:54.496644Z","iopub.execute_input":"2022-03-06T10:52:54.49695Z","iopub.status.idle":"2022-03-06T11:06:44.84024Z","shell.execute_reply.started":"2022-03-06T10:52:54.496908Z","shell.execute_reply":"2022-03-06T11:06:44.839349Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Create recommendation result from trained model\n\nI note document here for any one want to customize it: https://recbole.io/docs/user_guide/usage/case_study.html","metadata":{}},{"cell_type":"code","source":"from recbole.utils.case_study import full_sort_topk\nexternal_user_ids = dataset.id2token(\n    dataset.uid_field, list(range(dataset.user_num)))[1:]#fist element in array is 'PAD'(default of Recbole) ->remove it ","metadata":{"execution":{"iopub.status.busy":"2022-03-06T11:06:44.841429Z","iopub.execute_input":"2022-03-06T11:06:44.841697Z","iopub.status.idle":"2022-03-06T11:06:44.869434Z","shell.execute_reply.started":"2022-03-06T11:06:44.841659Z","shell.execute_reply":"2022-03-06T11:06:44.868586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"topk_items = []\nfor internal_user_id in list(range(dataset.user_num))[1:]:\n    _, topk_iid_list = full_sort_topk([internal_user_id], model, test_data, k=12, device=config['device'])\n    last_topk_iid_list = topk_iid_list[-1]\n    external_item_list = dataset.id2token(dataset.iid_field, last_topk_iid_list.cpu()).tolist()\n    topk_items.append(external_item_list)\nprint(len(topk_items))","metadata":{"execution":{"iopub.status.busy":"2022-03-06T11:06:44.873904Z","iopub.execute_input":"2022-03-06T11:06:44.876153Z","iopub.status.idle":"2022-03-06T11:07:42.830516Z","shell.execute_reply.started":"2022-03-06T11:06:44.876112Z","shell.execute_reply":"2022-03-06T11:07:42.829504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"external_item_str = [' '.join(x) for x in topk_items]\nresult = pd.DataFrame(external_user_ids, columns=['customer_id'])\nresult['prediction'] = external_item_str\nresult.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-06T11:07:42.832266Z","iopub.execute_input":"2022-03-06T11:07:42.832961Z","iopub.status.idle":"2022-03-06T11:07:42.868943Z","shell.execute_reply.started":"2022-03-06T11:07:42.83289Z","shell.execute_reply":"2022-03-06T11:07:42.867703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5. Combine result from most bought items and GRU model","metadata":{}},{"cell_type":"code","source":"submit_df = pd.read_csv('submission.csv')\nsubmit_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-03-06T11:07:42.870754Z","iopub.execute_input":"2022-03-06T11:07:42.871293Z","iopub.status.idle":"2022-03-06T11:07:46.475298Z","shell.execute_reply.started":"2022-03-06T11:07:42.871229Z","shell.execute_reply":"2022-03-06T11:07:46.474039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submit_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-06T11:07:46.479086Z","iopub.execute_input":"2022-03-06T11:07:46.479391Z","iopub.status.idle":"2022-03-06T11:07:46.496133Z","shell.execute_reply.started":"2022-03-06T11:07:46.479362Z","shell.execute_reply":"2022-03-06T11:07:46.494507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submit_df = pd.merge(submit_df, result, on='customer_id', how='outer')\nsubmit_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-06T11:08:00.919686Z","iopub.execute_input":"2022-03-06T11:08:00.92007Z","iopub.status.idle":"2022-03-06T11:08:01.963787Z","shell.execute_reply.started":"2022-03-06T11:08:00.920041Z","shell.execute_reply":"2022-03-06T11:08:01.962796Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submit_df = submit_df.fillna(-1)\nsubmit_df['prediction'] = submit_df.apply(\n    lambda x: x['prediction_y'] if x['prediction_y'] != -1 else x['prediction_x'], axis=1)\nsubmit_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-06T11:08:17.718536Z","iopub.execute_input":"2022-03-06T11:08:17.718859Z","iopub.status.idle":"2022-03-06T11:08:47.154309Z","shell.execute_reply.started":"2022-03-06T11:08:17.718828Z","shell.execute_reply":"2022-03-06T11:08:47.153205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submit_df[submit_df['prediction_y'] != -1]","metadata":{"execution":{"iopub.status.busy":"2022-03-06T11:09:21.078845Z","iopub.execute_input":"2022-03-06T11:09:21.079814Z","iopub.status.idle":"2022-03-06T11:09:21.518756Z","shell.execute_reply.started":"2022-03-06T11:09:21.079758Z","shell.execute_reply":"2022-03-06T11:09:21.517751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submit_df = submit_df.drop(columns=['prediction_y', 'prediction_x'])\nsubmit_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-06T11:09:33.668649Z","iopub.execute_input":"2022-03-06T11:09:33.669287Z","iopub.status.idle":"2022-03-06T11:09:33.805517Z","shell.execute_reply.started":"2022-03-06T11:09:33.66924Z","shell.execute_reply":"2022-03-06T11:09:33.804412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submit_df.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-03-06T11:09:38.442105Z","iopub.execute_input":"2022-03-06T11:09:38.442393Z","iopub.status.idle":"2022-03-06T11:09:49.458488Z","shell.execute_reply.started":"2022-03-06T11:09:38.442362Z","shell.execute_reply":"2022-03-06T11:09:49.457249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}