{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**Edit**: \nI have create new notebooks for applying our customize function for using all items as input for recommendation:\n* Using only interactions: https://www.kaggle.com/astrung/sequential-model-fixed-missing-last-item\n* Using interactions with item features: https://www.kaggle.com/code/astrung/lstm-model-with-item-infor-fix-missing-last-item\n\n- - -\n\nIn my previous [notebook](https://www.kaggle.com/code/astrung/recbole-lstm-sequential-for-recomendation-tutorial) about sequential model with Recbole, someone asked me about the mechanism of test data when using `full_sort_topk` for prediction submitted recommendation in this [comment](https://www.kaggle.com/code/astrung/recbole-lstm-sequential-for-recomendation-tutorial/comments#1723707) and this [comment](https://www.kaggle.com/code/astrung/recbole-lstm-sequential-for-recomendation-tutorial/comments#1723707), and they have a doubt about whether we are using all items for getting final recommendation. Most of people has 2 questions about using `full_sort_topk` with test data:\n1. Do items in test data are used as input features for getting recommendation ?\n2. If test data is necessary for getting recommendation in Recbole API, how can we get recommendation without splitting into train/test data?\n\nIn this notebook i will answer all questions:\n1. Yes. In sequential models, items in test data is used as input features, but not last items. As a example, if user X have 3 items in test data(A, B, C) and 5 items in train data(a,b,c,d,e), test data will generate 3 sample rows for evaluating performance on user X:\n* Row 1: Input features: `a,b,c,d,e,0,0`. Output features: `A`. `0` is a pad item\n* Row 2: Input features: `a,b,c,d,e,A,0`. Output features: `B`.\n* Row 3: Input features: `a,b,c,d,e,A,B`. Output features: `C`.\n\nIn my previous notebook, i use last row result as recommendation, **so we still using nearly all of items as input for recommendation, except last item(item C)**. Our recommendation in previous notebooks may be not perfect, but it is simple as a tutorial for anyone want to start.\n\n**Note: This mechanism is only for sequential model in recbole. For other types of model, it isn't correct - it won't use items in test data for getting recommendation. If you have requests for explaining for other model, please upvote and comment. I will explain it in other notebook**\n\nIn first session of this notebook, i will dig into test data to prove this conclusion.\n\n2. Yes, we can get recommendation by using all of items as input features, without splitting train/test. In order to do this, you need to modify recbole code:\n* Fist, you copy last row in dataset(input features have all items, except last one), then add last item into input features.\n* Then you predict directly from model api, without using [full_sort_score or full_sort_topk](https://recbole.io/docs/user_guide/usage/case_study.html)\n\nIn second session of this notebook, i will show you how to do that.\n\nOk, let start","metadata":{}},{"cell_type":"markdown","source":"# I. How test items are used in test data.\n\nFor each item in test data, it will be generated as a sample row. As a example, if user X have 3 items in test data(A, B, C) and 5 items in train data(a,b,c,d,e), test data will generate 3 sample rows for evaluating performance on user X:\n* Row 1: Input features: `a,b,c,d,e,0,0`. Output label: `A`. `0` is a pad item\n* Row 2: Input features: `a,b,c,d,e,A,0`. Output label: `B`.\n* Row 3: Input features: `a,b,c,d,e,A,B`. Output label: `C`.\n\nFor proving it, we will create a dataset, then extract input features and label in test data.\n\n### 1. Let create test data in recbole","metadata":{}},{"cell_type":"code","source":"!pip install recbole","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-03-19T07:15:02.070886Z","iopub.execute_input":"2022-03-19T07:15:02.071208Z","iopub.status.idle":"2022-03-19T07:15:17.427504Z","shell.execute_reply.started":"2022-03-19T07:15:02.071127Z","shell.execute_reply":"2022-03-19T07:15:17.426536Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\ndf = pd.read_csv(r\"/kaggle/input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\", \n                 dtype={'article_id': 'str'})\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-19T07:15:17.429692Z","iopub.execute_input":"2022-03-19T07:15:17.429943Z","iopub.status.idle":"2022-03-19T07:16:16.933214Z","shell.execute_reply.started":"2022-03-19T07:15:17.429911Z","shell.execute_reply":"2022-03-19T07:16:16.932588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\ndf['t_dat'] = pd.to_datetime(df['t_dat'], format=\"%Y-%m-%d\")\ndf['timestamp'] = df.t_dat.values.astype(np.int64) // 10 ** 9\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-19T07:16:16.934353Z","iopub.execute_input":"2022-03-19T07:16:16.934632Z","iopub.status.idle":"2022-03-19T07:16:23.239899Z","shell.execute_reply.started":"2022-03-19T07:16:16.934597Z","shell.execute_reply":"2022-03-19T07:16:23.239231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = df[df['timestamp'] > 1585620000][['customer_id', 'article_id', 'timestamp']].rename(\n    columns={'customer_id': 'user_id:token', 'article_id': 'item_id:token', 'timestamp': 'timestamp:float'})\ntemp","metadata":{"execution":{"iopub.status.busy":"2022-03-19T07:16:23.242057Z","iopub.execute_input":"2022-03-19T07:16:23.24231Z","iopub.status.idle":"2022-03-19T07:16:24.81193Z","shell.execute_reply.started":"2022-03-19T07:16:23.242276Z","shell.execute_reply":"2022-03-19T07:16:24.811301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Create data file in recbole format","metadata":{}},{"cell_type":"code","source":"!mkdir /kaggle/working/recbox_data\ntemp.to_csv('/kaggle/working/recbox_data/recbox_data.inter', index=False, sep='\\t')","metadata":{"execution":{"iopub.status.busy":"2022-03-19T07:16:24.815554Z","iopub.execute_input":"2022-03-19T07:16:24.818561Z","iopub.status.idle":"2022-03-19T07:17:00.440341Z","shell.execute_reply.started":"2022-03-19T07:16:24.818521Z","shell.execute_reply":"2022-03-19T07:17:00.439473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import gc\ndel temp\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-03-19T07:17:00.442024Z","iopub.execute_input":"2022-03-19T07:17:00.44228Z","iopub.status.idle":"2022-03-19T07:17:00.551153Z","shell.execute_reply.started":"2022-03-19T07:17:00.442243Z","shell.execute_reply":"2022-03-19T07:17:00.550452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import logging\nfrom logging import getLogger\nfrom recbole.config import Config\nfrom recbole.data import create_dataset, data_preparation\nfrom recbole.model.sequential_recommender import GRU4Rec\nfrom recbole.trainer import Trainer\nfrom recbole.utils import init_seed, init_logger","metadata":{"execution":{"iopub.status.busy":"2022-03-19T07:17:00.552624Z","iopub.execute_input":"2022-03-19T07:17:00.553038Z","iopub.status.idle":"2022-03-19T07:17:02.98159Z","shell.execute_reply.started":"2022-03-19T07:17:00.552996Z","shell.execute_reply":"2022-03-19T07:17:02.98082Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"parameter_dict = {\n    'data_path': '/kaggle/working',\n    'USER_ID_FIELD': 'user_id',\n    'ITEM_ID_FIELD': 'item_id',\n    'TIME_FIELD': 'timestamp',\n    'user_inter_num_interval': \"[40,inf)\",\n    'item_inter_num_interval': \"[40,inf)\",\n    'load_col': {'inter': ['user_id', 'item_id', 'timestamp']},\n    'neg_sampling': None,\n    'epochs': 2,\n    'eval_args': {\n        'split': {'RS': [9, 0, 1]},\n        'group_by': 'user',\n        'order': 'TO',\n        'mode': 'full'}\n}\nconfig = Config(model='GRU4Rec', dataset='recbox_data', config_dict=parameter_dict)\n\n# init random seed\ninit_seed(config['seed'], config['reproducibility'])\n\n# logger initialization\ninit_logger(config)\nlogger = getLogger()\n# Create handlers\nc_handler = logging.StreamHandler()\nc_handler.setLevel(logging.INFO)\nlogger.addHandler(c_handler)\n\n# write config info into log\n# logger.info(config)","metadata":{"execution":{"iopub.status.busy":"2022-03-19T07:17:02.983051Z","iopub.execute_input":"2022-03-19T07:17:02.983313Z","iopub.status.idle":"2022-03-19T07:17:03.13864Z","shell.execute_reply.started":"2022-03-19T07:17:02.983278Z","shell.execute_reply":"2022-03-19T07:17:03.137884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let start spliting train data and test data in recbole","metadata":{}},{"cell_type":"code","source":"dataset = create_dataset(config)\nlogger.info(dataset)","metadata":{"execution":{"iopub.status.busy":"2022-03-19T07:17:03.140106Z","iopub.execute_input":"2022-03-19T07:17:03.140339Z","iopub.status.idle":"2022-03-19T07:18:19.70185Z","shell.execute_reply.started":"2022-03-19T07:17:03.140307Z","shell.execute_reply":"2022-03-19T07:18:19.701228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# dataset splitting\ntrain_data, valid_data, test_data = data_preparation(config, dataset)","metadata":{"execution":{"iopub.status.busy":"2022-03-19T07:18:19.704432Z","iopub.execute_input":"2022-03-19T07:18:19.704679Z","iopub.status.idle":"2022-03-19T07:18:39.021192Z","shell.execute_reply.started":"2022-03-19T07:18:19.704645Z","shell.execute_reply":"2022-03-19T07:18:39.020572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2. Let extract sample rows from test data.\n\nWe will check items of user `0109ad0b5a76924a1b58be677409bb601cc8bead9a87b8ce5b08a4a1f5bc71ef`. \n\nWe except that last items of this user will be used as label in test data\n","metadata":{}},{"cell_type":"code","source":"last_item_ids = df[df.customer_id == '0109ad0b5a76924a1b58be677409bb601cc8bead9a87b8ce5b08a4a1f5bc71ef'\n                  ].tail(10).article_id.values\ndf[df.customer_id == '0109ad0b5a76924a1b58be677409bb601cc8bead9a87b8ce5b08a4a1f5bc71ef'].tail(10)","metadata":{"execution":{"iopub.status.busy":"2022-03-19T07:18:39.030998Z","iopub.execute_input":"2022-03-19T07:18:39.031566Z","iopub.status.idle":"2022-03-19T07:18:48.211015Z","shell.execute_reply.started":"2022-03-19T07:18:39.031527Z","shell.execute_reply":"2022-03-19T07:18:48.21026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"last_item_ids","metadata":{"execution":{"iopub.status.busy":"2022-03-19T08:01:57.51256Z","iopub.execute_input":"2022-03-19T08:01:57.513135Z","iopub.status.idle":"2022-03-19T08:01:57.522122Z","shell.execute_reply.started":"2022-03-19T08:01:57.513098Z","shell.execute_reply":"2022-03-19T08:01:57.521355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Recbole use an internal ids for identify user_id and item_id, so let convert this user_id and his items into internal ids.\n* customer_id: `0109ad0b5a76924a1b58be677409bb601cc8bead9a87b8ce5b08a4a1f5bc71ef` -> internal user id: 2\n* last bought item_id: [..., '0698286004', '0861478002', '0901955001'] -> internal item id: [..., 3237, 4377, 4559]","metadata":{}},{"cell_type":"code","source":"test_data.dataset.token2id(test_data.dataset.uid_field, \n                           '0109ad0b5a76924a1b58be677409bb601cc8bead9a87b8ce5b08a4a1f5bc71ef')","metadata":{"execution":{"iopub.status.busy":"2022-03-19T07:56:12.424527Z","iopub.execute_input":"2022-03-19T07:56:12.424787Z","iopub.status.idle":"2022-03-19T07:56:12.430958Z","shell.execute_reply.started":"2022-03-19T07:56:12.42476Z","shell.execute_reply":"2022-03-19T07:56:12.430115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(dataset.token2id(dataset.iid_field, last_item_ids))","metadata":{"execution":{"iopub.status.busy":"2022-03-19T08:02:33.042454Z","iopub.execute_input":"2022-03-19T08:02:33.042747Z","iopub.status.idle":"2022-03-19T08:02:33.052111Z","shell.execute_reply.started":"2022-03-19T08:02:33.042716Z","shell.execute_reply":"2022-03-19T08:02:33.051172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Now let extract input features and labels in our test data.\nMy extracted code is copy from [this source](https://recbole.io/docs/_modules/recbole/utils/case_study.html#full_sort_scores)**","metadata":{}},{"cell_type":"code","source":"input_features = test_data.dataset[np.isin(test_data.dataset[test_data.dataset.uid_field].numpy(), [2])]\ninput_features","metadata":{"execution":{"iopub.status.busy":"2022-03-19T07:18:48.212158Z","iopub.execute_input":"2022-03-19T07:18:48.212848Z","iopub.status.idle":"2022-03-19T07:18:48.225174Z","shell.execute_reply.started":"2022-03-19T07:18:48.212806Z","shell.execute_reply":"2022-03-19T07:18:48.224388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **item_id in above interaction is used as label item**\n* **item_id_list in above interaction is used as feature items**\n\nLet check it","metadata":{}},{"cell_type":"code","source":"print(\"test label: \" + str(input_features['item_id']))\nprint(\"last 10 items from origin dataset: \" + str(dataset.token2id(dataset.iid_field, last_item_ids)))","metadata":{"execution":{"iopub.status.busy":"2022-03-19T08:09:07.458382Z","iopub.execute_input":"2022-03-19T08:09:07.459079Z","iopub.status.idle":"2022-03-19T08:09:07.470377Z","shell.execute_reply.started":"2022-03-19T08:09:07.459035Z","shell.execute_reply":"2022-03-19T08:09:07.469481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we expected, in last 10 items, 5 last items are used as label item. So for evaluating this user, we will have 5 sample rows in test data: \n* Input feature: ? -> Output: 6745\n* Input feature: ? -> Output: 3237\n* Input feature: ? -> Output: 3237\n* Input feature: ? -> Output: 4377\n* Input feature: ? -> Output: 4559\n\nNow, let check input features in **item_id_list**","metadata":{}},{"cell_type":"code","source":"input_features['item_id_list']","metadata":{"execution":{"iopub.status.busy":"2022-03-19T07:18:48.242849Z","iopub.execute_input":"2022-03-19T07:18:48.243541Z","iopub.status.idle":"2022-03-19T07:18:48.252864Z","shell.execute_reply.started":"2022-03-19T07:18:48.243504Z","shell.execute_reply":"2022-03-19T07:18:48.25217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see:\n* For 1st row, it uses all items in training as input features.\n* For 2nd row, it uses all items in training + first label as input features\n* For 3rd row, it uses all items in training + first label + second label as input features\n* ...\n* For last row, it uses all items except last item as input features.\n\nIn my previous notebooks([here](https://www.kaggle.com/code/astrung/lstm-sequential-modelwith-item-features-tutorial) and [here](https://www.kaggle.com/code/astrung/recbole-lstm-sequential-for-recomendation-tutorial/notebook)), **I use last row result for recommendation, so we are missing information from last item. **\n\nSo now let fix it- find a new way for using all items","metadata":{}},{"cell_type":"markdown","source":"# 2. Custom code for using all items in recommendation","metadata":{}},{"cell_type":"markdown","source":"We have seen that last row is missing only last item, so fixxing ideal is simple now:\n* copy last row, add last item into it as a new interation(a row in test dataset)\n* make prediction with new interation\n\nSo now let train a dummy model for testing it\n\n### 1. Make dummy model","metadata":{}},{"cell_type":"code","source":"# model loading and initialization\nmodel = GRU4Rec(config, train_data.dataset).to(config['device'])\nlogger.info(model)\n\n# trainer loading and initialization\ntrainer = Trainer(config, model)\n\n# model training\nbest_valid_score, best_valid_result = trainer.fit(train_data)","metadata":{"execution":{"iopub.status.busy":"2022-03-19T07:18:48.254291Z","iopub.execute_input":"2022-03-19T07:18:48.254774Z","iopub.status.idle":"2022-03-19T07:19:18.345435Z","shell.execute_reply.started":"2022-03-19T07:18:48.254738Z","shell.execute_reply":"2022-03-19T07:19:18.344692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.eval()","metadata":{"execution":{"iopub.status.busy":"2022-03-19T07:19:18.346914Z","iopub.execute_input":"2022-03-19T07:19:18.347146Z","iopub.status.idle":"2022-03-19T07:19:18.353692Z","shell.execute_reply.started":"2022-03-19T07:19:18.347113Z","shell.execute_reply":"2022-03-19T07:19:18.352872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"input_features['item_id_list'].shape","metadata":{"execution":{"iopub.status.busy":"2022-03-19T07:19:18.354828Z","iopub.execute_input":"2022-03-19T07:19:18.355241Z","iopub.status.idle":"2022-03-19T07:19:18.365557Z","shell.execute_reply.started":"2022-03-19T07:19:18.355205Z","shell.execute_reply":"2022-03-19T07:19:18.364825Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our sequence items is always have fix length(50). So if we have more than 50 items, we need to drop earlier items, and if there are less than 50 items, we need to add a padding(0) into input item features. As example:\n* If last row input = [3, 4, 7,..., 20, 0, 0 ,0] (47 items < 50 item, so we have padding), after adding id=30, we will have input = [3, 4, 7,..., 20, 30, 0 ,0] Now our sequence lenght = 48 items.\n* If last row input = [3, 4, 7,..., 20, 9, 10 ,12] (50 items), after adding id=30, we will have input = [4, 7,..., 9, 10, 12, 30] (drop first item and add last item).Now our sequence lenght still = 50 items.\n\nNow let implement it.\n\nFirst let extract last row from all interation when internal_user_id = 2 ","metadata":{}},{"cell_type":"code","source":"index = np.isin(dataset[dataset.uid_field].numpy(), [2])\ninput_interaction = dataset[index]\ninput_interaction","metadata":{"execution":{"iopub.status.busy":"2022-03-19T08:34:57.762933Z","iopub.execute_input":"2022-03-19T08:34:57.763906Z","iopub.status.idle":"2022-03-19T08:34:57.788119Z","shell.execute_reply.started":"2022-03-19T08:34:57.763854Z","shell.execute_reply":"2022-03-19T08:34:57.787359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let add last item into sequences, and make new interaction.\nWe also need to edit sequence lenght (without padding)","metadata":{}},{"cell_type":"code","source":"import torch\nfrom recbole.data.interaction import Interaction\n\ndef add_last_item(old_interaction, last_item_id, max_len=50):\n    new_seq_items = old_interaction['item_id_list'][-1]\n    if old_interaction['item_length'][-1].item() < max_len:\n        new_seq_items[input_interaction['item_length'][-1].item()] = last_item_id\n    else:\n        new_seq_items = torch.roll(new_seq_items, -1)\n        new_seq_items[-1] = last_item_id\n    return new_seq_items.view(1, len(new_seq_items))\n\ntest = {\n            'item_id_list': add_last_item(input_interaction, input_interaction['item_id'][-1].item(), model.max_seq_length),\n            'item_length': torch.tensor(\n                [input_interaction['item_length'][-1].item() + 1\n                 if input_interaction['item_length'][-1].item() < model.max_seq_length else model.max_seq_length])\n        }\nnew_inter = Interaction(test)\nnew_inter","metadata":{"execution":{"iopub.status.busy":"2022-03-19T08:39:53.561619Z","iopub.execute_input":"2022-03-19T08:39:53.562149Z","iopub.status.idle":"2022-03-19T08:39:53.573441Z","shell.execute_reply.started":"2022-03-19T08:39:53.562108Z","shell.execute_reply":"2022-03-19T08:39:53.572725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Interaction for GRU4Rec model need to have only `item_id_list` and `item_lenght`. You can drop other key.\nIf you want more information, you can check [GRU4Rec code](https://recbole.io/docs/_modules/recbole/model/sequential_recommender/gru4rec.html#GRU4Rec)\n\nThen we can apply the remaining prediction code from [full_sort_scores](https://recbole.io/docs/_modules/recbole/utils/case_study.html#full_sort_scores)\n","metadata":{}},{"cell_type":"code","source":"new_inter = new_inter.to(config['device'])\nnew_scores = model.full_sort_predict(new_inter)\nnew_scores = new_scores.view(-1, test_data.dataset.item_num)\nnew_scores[:, 0] = -np.inf  # set scores of [pad] to -inf","metadata":{"execution":{"iopub.status.busy":"2022-03-19T08:46:00.688153Z","iopub.execute_input":"2022-03-19T08:46:00.68842Z","iopub.status.idle":"2022-03-19T08:46:00.713138Z","shell.execute_reply.started":"2022-03-19T08:46:00.688378Z","shell.execute_reply":"2022-03-19T08:46:00.712452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So now by combining all fragments,we have a new function for predicting with all item in dataset. You can use this custom code for all sequential model","metadata":{}},{"cell_type":"code","source":"import torch\nfrom recbole.data.interaction import Interaction\n\ndef add_last_item(old_interaction, last_item_id, max_len=50):\n    new_seq_items = old_interaction['item_id_list'][-1]\n    if old_interaction['item_length'][-1].item() < max_len:\n        new_seq_items[old_interaction['item_length'][-1].item()] = last_item_id\n    else:\n        new_seq_items = torch.roll(new_seq_items, -1)\n        new_seq_items[-1] = last_item_id\n    return new_seq_items.view(1, len(new_seq_items))\n\ndef predict_for_all_item(external_user_id, dataset, model):\n    model.eval()\n    with torch.no_grad():\n        uid_series = dataset.token2id(dataset.uid_field, [external_user_id])\n        index = np.isin(dataset[dataset.uid_field].numpy(), uid_series)\n        input_interaction = dataset[index]\n        test = {\n            'item_id_list': add_last_item(input_interaction, \n                                          input_interaction['item_id'][-1].item(), model.max_seq_length),\n            'item_length': torch.tensor(\n                [input_interaction['item_length'][-1].item() + 1\n                 if input_interaction['item_length'][-1].item() < model.max_seq_length else model.max_seq_length])\n        }\n        new_inter = Interaction(test)\n        new_inter = new_inter.to(config['device'])\n        new_scores = model.full_sort_predict(new_inter)\n        new_scores = new_scores.view(-1, test_data.dataset.item_num)\n        new_scores[:, 0] = -np.inf  # set scores of [pad] to -inf\n    return torch.topk(new_scores, 10)","metadata":{"execution":{"iopub.status.busy":"2022-03-19T08:50:43.785866Z","iopub.execute_input":"2022-03-19T08:50:43.786142Z","iopub.status.idle":"2022-03-19T08:50:43.795781Z","shell.execute_reply.started":"2022-03-19T08:50:43.786113Z","shell.execute_reply":"2022-03-19T08:50:43.794849Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predict_for_all_item('0109ad0b5a76924a1b58be677409bb601cc8bead9a87b8ce5b08a4a1f5bc71ef', \n                     dataset, model) # we feed directly origin dataset, not train data or test data","metadata":{"execution":{"iopub.status.busy":"2022-03-19T08:49:01.592098Z","iopub.execute_input":"2022-03-19T08:49:01.592846Z","iopub.status.idle":"2022-03-19T08:49:01.648005Z","shell.execute_reply.started":"2022-03-19T08:49:01.592811Z","shell.execute_reply":"2022-03-19T08:49:01.647317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Congratulation !!!.Now you can use all data as train set, don't need for a test set, but still can predict directly from dataset without testset.Now let apply it into our previous notebook.\n\nI have create new notebooks for applying our customize function for using all items as input for recommendation:\n* Using only interactions: https://www.kaggle.com/astrung/sequential-model-fixed-missing-last-item\n* Using interactions with item features: https://www.kaggle.com/code/astrung/lstm-model-with-item-infor-fix-missing-last-item\n\nPlease check and upvote it","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}