{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# I am trying to understand the methods of the public notebooks and note in this notebook.\n\nI am a beginner for the Recommended system.\n\n**I hope it is helpful for u.**\n\n**Continuous update!!!**","metadata":{}},{"cell_type":"markdown","source":"**First all. I use the sample of the all data base on  https://www.kaggle.com/paweljankiewicz/hm-create-dataset-samples.**\n\n**It sample 0.1%, 1%, 5% of the all data. It save a lot of time when u are learning.**\n\nYou can use it directly. https://www.kaggle.com/datasets/hjimbean/hm-samples","metadata":{}},{"cell_type":"markdown","source":"# 1. 【LB: 0.0220】 base on https://www.kaggle.com/code/hengzheng/time-is-our-best-friend-v2/notebook","metadata":{}},{"cell_type":"markdown","source":"## 1.1 data processing","metadata":{}},{"cell_type":"markdown","source":"It only use the article_id in transactions_train.csv. It dont use the detail information like the images of goods, the customer age or artcles detail something...","metadata":{}},{"cell_type":"markdown","source":"### This method only uses data from the last three weeks.\n\ntransactions_1w is the data from 2020-09-15 to 2020-09-22\n\ntransactions_2w is the data from 2020-09-07 to 2020-09-14\n\ntransactions_3w is the data from 2020-08-31 to 2020-09-06","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\nfrom pathlib import Path\n\n# data_path = Path('/kaggle/input/h-and-m-personalized-fashion-recommendations/')\ndata_path = Path('../input/hm-samples/')\n\ntransactions = pd.read_csv(\n    data_path / 'transactions_train_sample1.csv',\n    # set dtype or pandas will drop the leading '0' and convert to int\n    dtype={'article_id': str} \n)\n\nsubmission = pd.read_csv('../input/h-and-m-personalized-fashion-recommendations/sample_submission.csv')\n\ntransactions['t_dat'] = pd.to_datetime(transactions['t_dat'])\n\ntransactions_3w = transactions[transactions['t_dat'] >= pd.to_datetime('2020-08-31')].copy()\ntransactions_2w = transactions[transactions['t_dat'] >= pd.to_datetime('2020-09-07')].copy()\ntransactions_1w = transactions[transactions['t_dat'] >= pd.to_datetime('2020-09-15')].copy()","metadata":{"execution":{"iopub.status.busy":"2022-04-09T08:26:50.492961Z","iopub.execute_input":"2022-04-09T08:26:50.493899Z","iopub.status.idle":"2022-04-09T08:26:57.509616Z","shell.execute_reply.started":"2022-04-09T08:26:50.493788Z","shell.execute_reply":"2022-04-09T08:26:57.508331Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### the shape of data is same with transactions_train.csv","metadata":{}},{"cell_type":"code","source":"transactions.head()","metadata":{"execution":{"iopub.status.busy":"2022-04-09T08:33:22.189133Z","iopub.execute_input":"2022-04-09T08:33:22.189433Z","iopub.status.idle":"2022-04-09T08:33:22.202448Z","shell.execute_reply.started":"2022-04-09T08:33:22.189401Z","shell.execute_reply":"2022-04-09T08:33:22.201391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"transactions_1w.head()","metadata":{"execution":{"iopub.status.busy":"2022-04-09T08:33:26.129735Z","iopub.execute_input":"2022-04-09T08:33:26.130298Z","iopub.status.idle":"2022-04-09T08:33:26.144352Z","shell.execute_reply.started":"2022-04-09T08:33:26.130256Z","shell.execute_reply":"2022-04-09T08:33:26.143256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### turn this data into a dict，like：{customer_id: {article_id: count}}\n\ncount: this week the same customer_id buy how many times of the same article_id","metadata":{}},{"cell_type":"code","source":"purchase_dict_3w = {}\n\nfor i,x in enumerate(zip(transactions_3w['customer_id'], transactions_3w['article_id'])):\n    cust_id, art_id = x\n    if cust_id not in purchase_dict_3w:\n        purchase_dict_3w[cust_id] = {}\n    \n    if art_id not in purchase_dict_3w[cust_id]:\n        purchase_dict_3w[cust_id][art_id] = 0\n    \n    purchase_dict_3w[cust_id][art_id] += 1\n    \nprint(len(purchase_dict_3w))\n\ndummy_list_3w = list((transactions_3w['article_id'].value_counts()).index)[:12]\n\npurchase_dict_2w = {}\n\nfor i,x in enumerate(zip(transactions_2w['customer_id'], transactions_2w['article_id'])):\n    cust_id, art_id = x\n    if cust_id not in purchase_dict_2w:\n        purchase_dict_2w[cust_id] = {}\n    \n    if art_id not in purchase_dict_2w[cust_id]:\n        purchase_dict_2w[cust_id][art_id] = 0\n    \n    purchase_dict_2w[cust_id][art_id] += 1\n    \nprint(len(purchase_dict_2w))\n\ndummy_list_2w = list((transactions_2w['article_id'].value_counts()).index)[:12]\n\n\npurchase_dict_1w = {}\n\nfor i,x in enumerate(zip(transactions_1w['customer_id'], transactions_1w['article_id'])):\n    cust_id, art_id = x\n    if cust_id not in purchase_dict_1w:\n        purchase_dict_1w[cust_id] = {}\n    \n    if art_id not in purchase_dict_1w[cust_id]:\n        purchase_dict_1w[cust_id][art_id] = 0\n    \n    purchase_dict_1w[cust_id][art_id] += 1\n    \nprint(len(purchase_dict_1w))\n\ndummy_list_1w = list((transactions_1w['article_id'].value_counts()).index)[:12]","metadata":{"execution":{"iopub.status.busy":"2022-04-09T08:37:00.510051Z","iopub.execute_input":"2022-04-09T08:37:00.510361Z","iopub.status.idle":"2022-04-09T08:37:00.537832Z","shell.execute_reply.started":"2022-04-09T08:37:00.510328Z","shell.execute_reply":"2022-04-09T08:37:00.536846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**dummy_list_3w, purchase_dict_2w, purchase_dict_1w** is Top 12 Most Popular Items of the week.","metadata":{}},{"cell_type":"code","source":"purchase_dict_3w['045fc8403924222479a2263f84258b9bf7b02cb6ddfe2d495bd8a43583165981']","metadata":{"execution":{"iopub.status.busy":"2022-04-09T08:37:00.644356Z","iopub.execute_input":"2022-04-09T08:37:00.644644Z","iopub.status.idle":"2022-04-09T08:37:00.650865Z","shell.execute_reply.started":"2022-04-09T08:37:00.644614Z","shell.execute_reply":"2022-04-09T08:37:00.650146Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 1.2 predict","metadata":{}},{"cell_type":"markdown","source":"If you purchased something in the last week, it will be sorted and recommended according to the number of items you purchased. If you purchased less than 12 items, it will be recommended and supplemented according to the most popular items in the last week, and the data of the following weeks will not be seen. \n\nIf you didn't buy anything in the last week, look at the last two weeks, then look at the last three weeks. If you haven't bought anything at all, it will recommend you the most popular products in the last week.","metadata":{}},{"cell_type":"code","source":"not_so_fancy_but_fast_benchmark = submission[['customer_id']]\nprediction_list = []\n\ndummy_list = list((transactions_1w['article_id'].value_counts()).index)[:12]\ndummy_pred = ' '.join(dummy_list)\n\nfor i, cust_id in enumerate(submission['customer_id'].values.reshape((-1,))):\n    if cust_id in purchase_dict_1w:\n        l = sorted((purchase_dict_1w[cust_id]).items(), key=lambda x: x[1], reverse=True)\n        l = [y[0] for y in l]\n        if len(l)>12:\n            s = ' '.join(l[:12])\n        else:\n            s = ' '.join(l+dummy_list_1w[:(12-len(l))])\n    elif cust_id in purchase_dict_2w:\n        l = sorted((purchase_dict_2w[cust_id]).items(), key=lambda x: x[1], reverse=True)\n        l = [y[0] for y in l]\n        if len(l)>12:\n            s = ' '.join(l[:12])\n        else:\n            s = ' '.join(l+dummy_list_2w[:(12-len(l))])\n    elif cust_id in purchase_dict_3w:\n        l = sorted((purchase_dict_3w[cust_id]).items(), key=lambda x: x[1], reverse=True)\n        l = [y[0] for y in l]\n        if len(l)>12:\n            s = ' '.join(l[:12])\n        else:\n            s = ' '.join(l+dummy_list_3w[:(12-len(l))])\n    else:\n        s = dummy_pred\n    prediction_list.append(s)\n\nnot_so_fancy_but_fast_benchmark['prediction'] = prediction_list\nprint(not_so_fancy_but_fast_benchmark.shape)\nnot_so_fancy_but_fast_benchmark.head()","metadata":{},"execution_count":null,"outputs":[]}]}