{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Radek posted about this [here](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/309220), and linked to a GitHub repo with the code.\n\nI just transferred that code here to Kaggle notebooks, that's all.","metadata":{}},{"cell_type":"markdown","source":"## Introduction\nThis notebook is an article with code analysis of \"[Radek's LGBMRanker starter-pack](https://www.kaggle.com/code/marcogorelli/radek-s-lgbmranker-starter-pack)\" This notebook is a code analysis of \"[Radek's LGBMRanker starter-pack]().\n\nIt provides an overview of each process in Japanese. (Please refer to the comment-outs.)\n\nAs a new user of LGBMRanker, the preprocessing of this data set was very difficult for me.\nTherefore, in this notebook, I focused on the preprocessing of the data before rank learning in particular.\n\nFor more information on how to use LGBMRanker and how to put together a submission file after making predictions, you may want to refer to the following notebook.\n\nhttps://www.kaggle.com/code/kimurayut/gbm-ranking\n\nAlso, the following discussion should help you understand the data to be put into the study.\n\nhttps://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288#1728274","metadata":{}},{"cell_type":"markdown","source":"## はじめに\nこちらのノートブックは「[Radek's LGBMRanker starter-pack](https://www.kaggle.com/code/marcogorelli/radek-s-lgbmranker-starter-pack)」のコード分析をした記事です。\n\n日本語で各処理の概要を解説しています。（コメントアウトを参考ください。）\n\nLGBMRankerを初めて使う私にとって今回のデータセットの前処理は非常に難しいものでした。\nこのノートブックでは特にランク学習を行う前のデータの前処理にフォーカスをして分析しました。\n\nLGBMRankerの使い方や、予測を行った後どのように提出ファイルとしてまとめるかについては以下のノートブックを参考にすると良いかもしれません。\n\nhttps://www.kaggle.com/code/kimurayut/gbm-ranking\n\nまた、学習に投入するデータについては以下のディスカッションの内容を参考にすると理解できるはずです。\n\nhttps://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/discussion/307288#1728274\n","metadata":{}},{"cell_type":"code","source":"import numpy as np\n\ndef apk(actual, predicted, k=10):\n    \"\"\"\n    Computes the average precision at k.\n\n    This function computes the average prescision at k between two lists of\n    items.\n\n    Parameters\n    ----------\n    actual : list\n             A list of elements that are to be predicted (order doesn't matter)\n    predicted : list\n                A list of predicted elements (order does matter)\n    k : int, optional\n        The maximum number of predicted elements\n\n    Returns\n    -------\n    score : double\n            The average precision at k over the input lists\n\n    \"\"\"\n    if len(predicted)>k:\n        predicted = predicted[:k]\n\n    score = 0.0\n    num_hits = 0.0\n\n    for i,p in enumerate(predicted):\n        if p in actual and p not in predicted[:i]:\n            num_hits += 1.0\n            score += num_hits / (i+1.0)\n\n    if not actual:\n        return 0.0\n\n    return score / min(len(actual), k)\n\ndef mapk(actual, predicted, k=10):\n    \"\"\"\n    Computes the mean average precision at k.\n\n    This function computes the mean average prescision at k between two lists\n    of lists of items.\n\n    Parameters\n    ----------\n    actual : list\n             A list of lists of elements that are to be predicted \n             (order doesn't matter in the lists)\n    predicted : list\n                A list of lists of predicted elements\n                (order matters in the lists)\n    k : int, optional\n        The maximum number of predicted elements\n\n    Returns\n    -------\n    score : double\n            The mean average precision at k over the input lists\n\n    \"\"\"\n    return np.mean([apk(a,p,k) for a,p in zip(actual, predicted)])","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:37:39.980332Z","iopub.execute_input":"2022-04-20T22:37:39.980686Z","iopub.status.idle":"2022-04-20T22:37:40.011026Z","shell.execute_reply.started":"2022-04-20T22:37:39.980600Z","shell.execute_reply":"2022-04-20T22:37:40.010324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.base import BaseEstimator, TransformerMixin\nimport numpy as np\n\n# https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations/discussion/308635\ndef customer_hex_id_to_int(series):\n    return series.str[-16:].apply(hex_id_to_int)\n\ndef hex_id_to_int(str):\n    return int(str[-16:], 16)\n\ndef article_id_str_to_int(series):\n    return series.astype('int32')\n\ndef article_id_int_to_str(series):\n    return '0' + series.astype('str')\n\nclass Categorize(BaseEstimator, TransformerMixin):\n    def __init__(self, min_examples=0):\n        self.min_examples = min_examples\n        self.categories = []\n        \n    def fit(self, X):\n        for i in range(X.shape[1]):\n            vc = X.iloc[:, i].value_counts()\n            self.categories.append(vc[vc > self.min_examples].index.tolist())\n        return self\n\n    def transform(self, X):\n        data = {X.columns[i]: pd.Categorical(X.iloc[:, i], categories=self.categories[i]).codes for i in range(X.shape[1])}\n        return pd.DataFrame(data=data)\n\n\ndef calculate_apk(list_of_preds, list_of_gts):\n    # for fast validation this can be changed to operate on dicts of {'cust_id_int': [art_id_int, ...]}\n    # using 'data/val_week_purchases_by_cust.pkl'\n    apks = []\n    for preds, gt in zip(list_of_preds, list_of_gts):\n        apks.append(apk(gt, preds, k=12))\n    return np.mean(apks)\n\ndef eval_sub(sub_csv, skip_cust_with_no_purchases=True):\n    sub=pd.read_csv(sub_csv)\n    validation_set=pd.read_parquet('data/validation_ground_truth.parquet')\n\n    apks = []\n\n    no_purchases_pattern = []\n    for pred, gt in zip(sub.prediction.str.split(), validation_set.prediction.str.split()):\n        if skip_cust_with_no_purchases and (gt == no_purchases_pattern): continue\n        apks.append(apk(gt, pred, k=12))\n    return np.mean(apks)","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:37:40.013477Z","iopub.execute_input":"2022-04-20T22:37:40.013803Z","iopub.status.idle":"2022-04-20T22:37:41.023710Z","shell.execute_reply.started":"2022-04-20T22:37:40.013757Z","shell.execute_reply":"2022-04-20T22:37:41.022799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:37:41.025056Z","iopub.execute_input":"2022-04-20T22:37:41.025363Z","iopub.status.idle":"2022-04-20T22:37:41.030699Z","shell.execute_reply.started":"2022-04-20T22:37:41.025312Z","shell.execute_reply":"2022-04-20T22:37:41.029685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\ntransactions = pd.read_parquet('../input/warmup/transactions_train.parquet')\ncustomers = pd.read_parquet('../input/warmup/customers.parquet')\narticles = pd.read_parquet('../input/warmup/articles.parquet')\n\n# sample = 0.05\n# transactions = pd.read_parquet(f'data/transactions_train_sample_{sample}.parquet')\n# customers = pd.read_parquet(f'data/customers_sample_{sample}.parquet')\n# articles = pd.read_parquet(f'data/articles_train_sample_{sample}.parquet')","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:37:41.033504Z","iopub.execute_input":"2022-04-20T22:37:41.033832Z","iopub.status.idle":"2022-04-20T22:37:48.335612Z","shell.execute_reply.started":"2022-04-20T22:37:41.033780Z","shell.execute_reply":"2022-04-20T22:37:48.334315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"transactions.week.max() - 10","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:37:48.336954Z","iopub.execute_input":"2022-04-20T22:37:48.337495Z","iopub.status.idle":"2022-04-20T22:37:48.377150Z","shell.execute_reply.started":"2022-04-20T22:37:48.337445Z","shell.execute_reply":"2022-04-20T22:37:48.376257Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 後にテストデータを用意するために使用\n# Used to prepare test data after\ntest_week = transactions.week.max() + 1 #105\n\n# week94よりも大きいweekのトランザクションを保存\n# Save transactions for a week greater than week94\ntransactions = transactions[transactions.week > transactions.week.max() - 10] ","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:37:48.378420Z","iopub.execute_input":"2022-04-20T22:37:48.379269Z","iopub.status.idle":"2022-04-20T22:37:48.636391Z","shell.execute_reply.started":"2022-04-20T22:37:48.379224Z","shell.execute_reply":"2022-04-20T22:37:48.635461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Generating candidates","metadata":{}},{"cell_type":"markdown","source":"### Last purchase candidates","metadata":{}},{"cell_type":"code","source":"%%time\n# 各カスタマーが購入したweekを抽出\n# Extract weeks purchased by each customer\nc2weeks = transactions.groupby('customer_id')['week'].unique()","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:37:48.637933Z","iopub.execute_input":"2022-04-20T22:37:48.638440Z","iopub.status.idle":"2022-04-20T22:38:13.155091Z","shell.execute_reply.started":"2022-04-20T22:37:48.638404Z","shell.execute_reply":"2022-04-20T22:38:13.153838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 各weekの始まりと終わりの日付を確認\n# Check the beginning and end dates of each WEEK\ntransactions.groupby('week')['t_dat'].agg(['min', 'max'])","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:13.156660Z","iopub.execute_input":"2022-04-20T22:38:13.156933Z","iopub.status.idle":"2022-04-20T22:38:13.244200Z","shell.execute_reply.started":"2022-04-20T22:38:13.156902Z","shell.execute_reply":"2022-04-20T22:38:13.243234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"c2weeks","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:13.245308Z","iopub.execute_input":"2022-04-20T22:38:13.245538Z","iopub.status.idle":"2022-04-20T22:38:13.255844Z","shell.execute_reply.started":"2022-04-20T22:38:13.245511Z","shell.execute_reply":"2022-04-20T22:38:13.254764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n# 顧客が購入したweekの１周ずらした値を辞書型で持つ\n# Have the value of one round shift of the WEEK purchased by the customer in dictionary type.\nc2weeks2shifted_weeks = {}\n\nfor c_id, weeks in c2weeks.items():\n    # c_id...顧客ID\n    # weeks...顧客が購入したweekの配列\n    c2weeks2shifted_weeks[c_id] = {}\n    for i in range(weeks.shape[0]-1):\n        # 顧客が購入したweekの１周ずらした値を設定\n        # Set the value of one round shift of the WEEK purchased by the customer.\n        c2weeks2shifted_weeks[c_id][weeks[i]] = weeks[i+1]\n    c2weeks2shifted_weeks[c_id][weeks[-1]] = test_week #最後のweekには必ずtest_week(今回は105)が設定される The last week is always set to test_week (105 in this case)","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:13.259091Z","iopub.execute_input":"2022-04-20T22:38:13.259718Z","iopub.status.idle":"2022-04-20T22:38:14.581509Z","shell.execute_reply.started":"2022-04-20T22:38:13.259684Z","shell.execute_reply":"2022-04-20T22:38:14.580713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"c2weeks2shifted_weeks[28847241659200][95]","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:14.582622Z","iopub.execute_input":"2022-04-20T22:38:14.583018Z","iopub.status.idle":"2022-04-20T22:38:14.588616Z","shell.execute_reply.started":"2022-04-20T22:38:14.582987Z","shell.execute_reply":"2022-04-20T22:38:14.587603Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"candidates_last_purchase = transactions.copy()","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:14.590229Z","iopub.execute_input":"2022-04-20T22:38:14.590452Z","iopub.status.idle":"2022-04-20T22:38:14.644162Z","shell.execute_reply.started":"2022-04-20T22:38:14.590419Z","shell.execute_reply":"2022-04-20T22:38:14.643124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nweeks = []\nfor i, (c_id, week) in enumerate(zip(transactions['customer_id'], transactions['week'])):\n    # Set the week one round off from the week of purchase (but the last week is always set to 105)\n    # 購入した週から１周ずらした週を設定（ただし最後の週は必ず105が設定される）\n    weeks.append(c2weeks2shifted_weeks[c_id][week]) \n    \n#列情報を全て１周ずらした形で上書き\n# Overwrite all column information with one round shift. \ncandidates_last_purchase.week=weeks ","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:14.645673Z","iopub.execute_input":"2022-04-20T22:38:14.646057Z","iopub.status.idle":"2022-04-20T22:38:27.082588Z","shell.execute_reply.started":"2022-04-20T22:38:14.645999Z","shell.execute_reply":"2022-04-20T22:38:27.081726Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 顧客IDが272412481300040の人の情報\n# Information about the person whose customer ID is 272412481300040\ncandidates_last_purchase[candidates_last_purchase['customer_id']==272412481300040]","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:27.083738Z","iopub.execute_input":"2022-04-20T22:38:27.084678Z","iopub.status.idle":"2022-04-20T22:38:27.104974Z","shell.execute_reply.started":"2022-04-20T22:38:27.084636Z","shell.execute_reply":"2022-04-20T22:38:27.104105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 顧客IDが272412481300040の人の情報\n# Information about the person whose customer ID is 272412481300040\ntransactions[transactions['customer_id']==272412481300040]","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:27.106188Z","iopub.execute_input":"2022-04-20T22:38:27.106490Z","iopub.status.idle":"2022-04-20T22:38:27.125223Z","shell.execute_reply.started":"2022-04-20T22:38:27.106462Z","shell.execute_reply":"2022-04-20T22:38:27.124412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"★ 全ての顧客が最後に買った週はweekが105に設定されているためそれを、最終購入日を抽出するための条件として後々使用する？","metadata":{}},{"cell_type":"markdown","source":"### Bestsellers candidates","metadata":{}},{"cell_type":"code","source":"# week、article_id毎に売り上げの平均を算出\n# Average sales per week, per article_id\nmean_price = transactions \\\n    .groupby(['week', 'article_id'])['price'].mean()","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:27.126359Z","iopub.execute_input":"2022-04-20T22:38:27.126668Z","iopub.status.idle":"2022-04-20T22:38:27.374114Z","shell.execute_reply.started":"2022-04-20T22:38:27.126637Z","shell.execute_reply":"2022-04-20T22:38:27.373232Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mean_price","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:27.436649Z","iopub.execute_input":"2022-04-20T22:38:27.436952Z","iopub.status.idle":"2022-04-20T22:38:27.446840Z","shell.execute_reply.started":"2022-04-20T22:38:27.436923Z","shell.execute_reply":"2022-04-20T22:38:27.445988Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"transactions.groupby('week')['article_id'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:27.448416Z","iopub.execute_input":"2022-04-20T22:38:27.449238Z","iopub.status.idle":"2022-04-20T22:38:28.426796Z","shell.execute_reply.started":"2022-04-20T22:38:27.449191Z","shell.execute_reply":"2022-04-20T22:38:28.425850Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"transactions \\\n    .groupby('week')['article_id'].value_counts() \\\n    .groupby('week').rank(method='dense', ascending=False) \\\n    .groupby('week').head(12).rename('bestseller_rank').astype('int8')","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:28.428429Z","iopub.execute_input":"2022-04-20T22:38:28.429437Z","iopub.status.idle":"2022-04-20T22:38:29.437518Z","shell.execute_reply.started":"2022-04-20T22:38:28.429393Z","shell.execute_reply":"2022-04-20T22:38:29.436091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 週毎の商品購入回数をもとにランク付けを行い、上位12の商品のみ抽出\n# Ranked based on number of product purchases per week, only top 12 products selected\nsales = transactions \\\n    .groupby('week')['article_id'].value_counts() \\\n    .groupby('week').rank(method='dense', ascending=False) \\\n    .groupby('week').head(12).rename('bestseller_rank').astype('int8') ","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:29.439113Z","iopub.execute_input":"2022-04-20T22:38:29.439447Z","iopub.status.idle":"2022-04-20T22:38:30.372215Z","shell.execute_reply.started":"2022-04-20T22:38:29.439405Z","shell.execute_reply":"2022-04-20T22:38:30.371315Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sales","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:30.374109Z","iopub.execute_input":"2022-04-20T22:38:30.374808Z","iopub.status.idle":"2022-04-20T22:38:30.385263Z","shell.execute_reply.started":"2022-04-20T22:38:30.374757Z","shell.execute_reply":"2022-04-20T22:38:30.384218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# week95の上位12の商品\n# top 12 products of week95\nsales.loc[95]","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:30.386694Z","iopub.execute_input":"2022-04-20T22:38:30.387122Z","iopub.status.idle":"2022-04-20T22:38:30.405475Z","shell.execute_reply.started":"2022-04-20T22:38:30.387086Z","shell.execute_reply":"2022-04-20T22:38:30.404685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#売り上げ上位12の商品テーブルと平均販売価格のテーブルを結合\n#Combine the top 12 products sold table with the average selling price table\nbestsellers_previous_week = pd.merge(sales, mean_price, on=['week', 'article_id']).reset_index()","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:30.407205Z","iopub.execute_input":"2022-04-20T22:38:30.407443Z","iopub.status.idle":"2022-04-20T22:38:30.453141Z","shell.execute_reply.started":"2022-04-20T22:38:30.407417Z","shell.execute_reply":"2022-04-20T22:38:30.452119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bestsellers_previous_week","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:30.454824Z","iopub.execute_input":"2022-04-20T22:38:30.456836Z","iopub.status.idle":"2022-04-20T22:38:30.470496Z","shell.execute_reply.started":"2022-04-20T22:38:30.456798Z","shell.execute_reply":"2022-04-20T22:38:30.469519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# １周分足し上げる\n# Add up one round\nbestsellers_previous_week.week += 1","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:30.471685Z","iopub.execute_input":"2022-04-20T22:38:30.472405Z","iopub.status.idle":"2022-04-20T22:38:30.479221Z","shell.execute_reply.started":"2022-04-20T22:38:30.472361Z","shell.execute_reply":"2022-04-20T22:38:30.478381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# week96のデータ確認\n# Check data for week96\nbestsellers_previous_week.pipe(lambda df: df[df['week']==96])","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:30.480478Z","iopub.execute_input":"2022-04-20T22:38:30.480996Z","iopub.status.idle":"2022-04-20T22:38:30.502728Z","shell.execute_reply.started":"2022-04-20T22:38:30.480886Z","shell.execute_reply":"2022-04-20T22:38:30.501597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 週ごとに各顧客の一番最初のトランザクションを格納\n# Store first transaction for each customer per week\nunique_transactions = transactions \\\n    .groupby(['week', 'customer_id']) \\\n    .head(1) \\\n    .drop(columns=['article_id', 'price']) \\\n    .copy()","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:30.504306Z","iopub.execute_input":"2022-04-20T22:38:30.504609Z","iopub.status.idle":"2022-04-20T22:38:31.118547Z","shell.execute_reply.started":"2022-04-20T22:38:30.504578Z","shell.execute_reply":"2022-04-20T22:38:31.117578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"unique_transactions","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:31.124962Z","iopub.execute_input":"2022-04-20T22:38:31.125317Z","iopub.status.idle":"2022-04-20T22:38:31.140772Z","shell.execute_reply.started":"2022-04-20T22:38:31.125280Z","shell.execute_reply":"2022-04-20T22:38:31.139828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 重複している行を削除（週ごとに同じ顧客は存在しなくなる）\n# Remove duplicate rows (same customer no longer exists from week to week)\ntransactions.drop_duplicates(['week', 'customer_id'])","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:31.142067Z","iopub.execute_input":"2022-04-20T22:38:31.142363Z","iopub.status.idle":"2022-04-20T22:38:31.428463Z","shell.execute_reply.started":"2022-04-20T22:38:31.142327Z","shell.execute_reply":"2022-04-20T22:38:31.427495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\ncandidates_bestsellers = pd.merge(\n    #週ごとに各顧客の一番最初のトランザクション\n    #First transaction for each customer per week\n    unique_transactions, \n    #売り上げ上位12の商品テーブルと平均販売価格のテーブルを結合したテーブル\n    #Table combining the top 12 products sold table and the average selling price table\n    bestsellers_previous_week, \n    on='week',\n)","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:31.429736Z","iopub.execute_input":"2022-04-20T22:38:31.429967Z","iopub.status.idle":"2022-04-20T22:38:32.121601Z","shell.execute_reply.started":"2022-04-20T22:38:31.429940Z","shell.execute_reply":"2022-04-20T22:38:32.120708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"candidates_bestsellers","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:32.122850Z","iopub.execute_input":"2022-04-20T22:38:32.123084Z","iopub.status.idle":"2022-04-20T22:38:32.141516Z","shell.execute_reply.started":"2022-04-20T22:38:32.123057Z","shell.execute_reply":"2022-04-20T22:38:32.140531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 全てのweekで重複している顧客IDを削除する\n# Remove duplicate customer IDs in all WEEKS\ntest_set_transactions = unique_transactions.drop_duplicates('customer_id').reset_index(drop=True)\n\n#全てweekを105にする\n#All to 105 for the week.\ntest_set_transactions.week = test_week \n","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:32.143181Z","iopub.execute_input":"2022-04-20T22:38:32.143520Z","iopub.status.idle":"2022-04-20T22:38:32.246076Z","shell.execute_reply.started":"2022-04-20T22:38:32.143476Z","shell.execute_reply":"2022-04-20T22:38:32.245130Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_set_transactions","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:32.248146Z","iopub.execute_input":"2022-04-20T22:38:32.248475Z","iopub.status.idle":"2022-04-20T22:38:32.263154Z","shell.execute_reply.started":"2022-04-20T22:38:32.248443Z","shell.execute_reply":"2022-04-20T22:38:32.262140Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"candidates_bestsellers_test_week = pd.merge(\n    # 全てのトランザクションにtest_week(105)を設定したテーブル\n    # Table with test_week(105) for all transactions\n    test_set_transactions, \n    # 売り上げ上位12の商品テーブルと平均販売価格のテーブルを結合したテーブル\n    # Table combining the top 12 products sold table and the average selling price table\n    bestsellers_previous_week, \n    on='week'\n)","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:32.264392Z","iopub.execute_input":"2022-04-20T22:38:32.264836Z","iopub.status.idle":"2022-04-20T22:38:32.726685Z","shell.execute_reply.started":"2022-04-20T22:38:32.264803Z","shell.execute_reply":"2022-04-20T22:38:32.725576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"candidates_bestsellers_test_week","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:32.728002Z","iopub.execute_input":"2022-04-20T22:38:32.728230Z","iopub.status.idle":"2022-04-20T22:38:32.745328Z","shell.execute_reply.started":"2022-04-20T22:38:32.728204Z","shell.execute_reply":"2022-04-20T22:38:32.744437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 週ごとに各顧客の一番最初のトランザクションが保持されているテーブルと\n# 全てのweekで重複している顧客IDを削除し、全レコードのweekを105にしたテーブルを縦に結合\n# with a table that holds the very first transaction for each customer by week, and\n# Vertically join the table with all records in week 105, removing duplicate customer IDs in all weeks.\ncandidates_bestsellers = pd.concat([candidates_bestsellers, candidates_bestsellers_test_week])\ncandidates_bestsellers.drop(columns='bestseller_rank', inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:32.746770Z","iopub.execute_input":"2022-04-20T22:38:32.747151Z","iopub.status.idle":"2022-04-20T22:38:33.783419Z","shell.execute_reply.started":"2022-04-20T22:38:32.747108Z","shell.execute_reply":"2022-04-20T22:38:33.782416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"candidates_bestsellers","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:33.785097Z","iopub.execute_input":"2022-04-20T22:38:33.785412Z","iopub.status.idle":"2022-04-20T22:38:33.804750Z","shell.execute_reply.started":"2022-04-20T22:38:33.785371Z","shell.execute_reply":"2022-04-20T22:38:33.803695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Combining transactions and candidates / negative examples","metadata":{}},{"cell_type":"code","source":"# 全てのトランザクションに対し購入フラグ(1)を設定\n# Set purchase flag (1) for all transactions\ntransactions['purchased'] = 1","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:33.806366Z","iopub.execute_input":"2022-04-20T22:38:33.806581Z","iopub.status.idle":"2022-04-20T22:38:33.817829Z","shell.execute_reply.started":"2022-04-20T22:38:33.806554Z","shell.execute_reply":"2022-04-20T22:38:33.816776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# transactions..全てのトランザクション\n# candidates_last_purchase..transactionsのweek列を一周ずらしたテーブル\n# candidates_bestsellers...週ごとに各顧客の一番最初のトランザクションが保持されているテーブルと\n# 全てのweekで重複している顧客IDを削除し、全レコードのweekを105にしたテーブルを縦に結合したテーブル\n# All transactions and\n# a table holding the first transaction for each customer by week, and\n# Vertically join the table with all records in week 105, removing duplicate customer IDs in all weeks.\ndata = pd.concat([transactions, candidates_last_purchase, candidates_bestsellers])\n\n# 全てのトランザクションに対し欠損値対応、このときcandidates_last_purchase、candidates_bestsellersから参照したデータには、\n# 購入フラグ(0)が設定される\n# Missing values for all transactions, data referenced from candidates_last_purchase and candidates_bestsellers will be set to Purchase flag (0) is set\ndata.purchased.fillna(0, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:33.819144Z","iopub.execute_input":"2022-04-20T22:38:33.819573Z","iopub.status.idle":"2022-04-20T22:38:34.414469Z","shell.execute_reply.started":"2022-04-20T22:38:33.819542Z","shell.execute_reply":"2022-04-20T22:38:34.413471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:34.415774Z","iopub.execute_input":"2022-04-20T22:38:34.416352Z","iopub.status.idle":"2022-04-20T22:38:34.435094Z","shell.execute_reply.started":"2022-04-20T22:38:34.416303Z","shell.execute_reply":"2022-04-20T22:38:34.434187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 顧客IDと商品ID、weekで重複しているレコードを削除\n# Delete duplicate records with customer ID, product ID, and week\ndata.drop_duplicates(['customer_id', 'article_id', 'week'], inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:34.436242Z","iopub.execute_input":"2022-04-20T22:38:34.436837Z","iopub.status.idle":"2022-04-20T22:38:41.406765Z","shell.execute_reply.started":"2022-04-20T22:38:34.436800Z","shell.execute_reply":"2022-04-20T22:38:41.405695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.purchased.mean()","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:41.408206Z","iopub.execute_input":"2022-04-20T22:38:41.408468Z","iopub.status.idle":"2022-04-20T22:38:41.467287Z","shell.execute_reply.started":"2022-04-20T22:38:41.408438Z","shell.execute_reply":"2022-04-20T22:38:41.466295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Add bestseller information","metadata":{}},{"cell_type":"code","source":"# ランク情報を付与\n# Assign rank information\ndata = pd.merge(\n    data,\n    bestsellers_previous_week[['week', 'article_id', 'bestseller_rank']],\n    on=['week', 'article_id'],\n    how='left'\n)","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:41.468922Z","iopub.execute_input":"2022-04-20T22:38:41.469266Z","iopub.status.idle":"2022-04-20T22:38:44.936211Z","shell.execute_reply.started":"2022-04-20T22:38:41.469214Z","shell.execute_reply":"2022-04-20T22:38:44.935362Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:44.937547Z","iopub.execute_input":"2022-04-20T22:38:44.937772Z","iopub.status.idle":"2022-04-20T22:38:44.961451Z","shell.execute_reply.started":"2022-04-20T22:38:44.937746Z","shell.execute_reply":"2022-04-20T22:38:44.960123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 一番古いweekは対象外にする\n# Exclude the oldest WEEK from the list.\ndata = data[data.week != data.week.min()]","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:44.963163Z","iopub.execute_input":"2022-04-20T22:38:44.963610Z","iopub.status.idle":"2022-04-20T22:38:46.578077Z","shell.execute_reply.started":"2022-04-20T22:38:44.963565Z","shell.execute_reply":"2022-04-20T22:38:46.577202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:46.581584Z","iopub.execute_input":"2022-04-20T22:38:46.581861Z","iopub.status.idle":"2022-04-20T22:38:46.603794Z","shell.execute_reply.started":"2022-04-20T22:38:46.581829Z","shell.execute_reply":"2022-04-20T22:38:46.602839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ランク情報がないものは999で補完\n# 999 completes those without rank information\ndata.bestseller_rank.fillna(999, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:46.605362Z","iopub.execute_input":"2022-04-20T22:38:46.605927Z","iopub.status.idle":"2022-04-20T22:38:46.651176Z","shell.execute_reply.started":"2022-04-20T22:38:46.605887Z","shell.execute_reply":"2022-04-20T22:38:46.650533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 商品と顧客の特徴情報を追加\n# Add product and customer feature information\ndata = pd.merge(data, articles, on='article_id', how='left')\ndata = pd.merge(data, customers, on='customer_id', how='left')","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:38:46.652316Z","iopub.execute_input":"2022-04-20T22:38:46.652516Z","iopub.status.idle":"2022-04-20T22:39:06.578203Z","shell.execute_reply.started":"2022-04-20T22:38:46.652491Z","shell.execute_reply":"2022-04-20T22:39:06.577238Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:39:06.579489Z","iopub.execute_input":"2022-04-20T22:39:06.579712Z","iopub.status.idle":"2022-04-20T22:39:10.283492Z","shell.execute_reply.started":"2022-04-20T22:39:06.579686Z","shell.execute_reply":"2022-04-20T22:39:10.282620Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# weekと顧客IDで並び替え\n# Sort by week and customer ID\ndata.sort_values(['week', 'customer_id'], inplace=True)\ndata.reset_index(drop=True, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:39:10.284688Z","iopub.execute_input":"2022-04-20T22:39:10.284899Z","iopub.status.idle":"2022-04-20T22:39:16.113487Z","shell.execute_reply.started":"2022-04-20T22:39:10.284873Z","shell.execute_reply":"2022-04-20T22:39:16.112350Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# week105以外のデータをトレーニングデータにする\n# Use data other than week105 as training data\ntrain = data[data.week != test_week]\n\n# week105のデータをテストデータにする\n# week105 data to test data.\ntest = data[data.week==test_week].drop_duplicates(['customer_id', 'article_id', 'sales_channel_id']).copy()","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:39:16.115226Z","iopub.execute_input":"2022-04-20T22:39:16.115467Z","iopub.status.idle":"2022-04-20T22:39:22.553669Z","shell.execute_reply.started":"2022-04-20T22:39:16.115438Z","shell.execute_reply":"2022-04-20T22:39:22.552784Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.groupby(['week', 'customer_id'])['article_id'].count()","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:39:22.555613Z","iopub.execute_input":"2022-04-20T22:39:22.555947Z","iopub.status.idle":"2022-04-20T22:39:23.475788Z","shell.execute_reply.started":"2022-04-20T22:39:22.555904Z","shell.execute_reply":"2022-04-20T22:39:23.474841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# LGBMRankerのgroupに登録するためのクエリ(week毎に顧客が何回買い物をしたかわかるクエリ)\n# Query to register in LGBMRanker's group (query to see how many times a customer has shopped each week)\ntrain_baskets = train.groupby(['week', 'customer_id'])['article_id'].count().values","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:39:23.477088Z","iopub.execute_input":"2022-04-20T22:39:23.477318Z","iopub.status.idle":"2022-04-20T22:39:24.411854Z","shell.execute_reply.started":"2022-04-20T22:39:23.477290Z","shell.execute_reply":"2022-04-20T22:39:24.410774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"columns_to_use = ['article_id', 'product_type_no', 'graphical_appearance_no', 'colour_group_code', 'perceived_colour_value_id',\n'perceived_colour_master_id', 'department_no', 'index_code',\n'index_group_no', 'section_no', 'garment_group_no', 'FN', 'Active',\n'club_member_status', 'fashion_news_frequency', 'age', 'postal_code', 'bestseller_rank']","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:39:24.413292Z","iopub.execute_input":"2022-04-20T22:39:24.413631Z","iopub.status.idle":"2022-04-20T22:39:24.418626Z","shell.execute_reply.started":"2022-04-20T22:39:24.413586Z","shell.execute_reply":"2022-04-20T22:39:24.417428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\ntrain_X = train[columns_to_use]\ntrain_y = train['purchased']\n\ntest_X = test[columns_to_use]","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:39:24.419991Z","iopub.execute_input":"2022-04-20T22:39:24.420351Z","iopub.status.idle":"2022-04-20T22:39:24.968545Z","shell.execute_reply.started":"2022-04-20T22:39:24.420313Z","shell.execute_reply":"2022-04-20T22:39:24.967764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Model training","metadata":{}},{"cell_type":"code","source":"from lightgbm.sklearn import LGBMRanker","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:39:24.969950Z","iopub.execute_input":"2022-04-20T22:39:24.970296Z","iopub.status.idle":"2022-04-20T22:39:26.056795Z","shell.execute_reply.started":"2022-04-20T22:39:24.970253Z","shell.execute_reply":"2022-04-20T22:39:26.055824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ranker = LGBMRanker(\n    objective=\"lambdarank\",\n    metric=\"ndcg\",\n    boosting_type=\"dart\",\n    n_estimators=1,\n    importance_type='gain',\n    verbose=10\n)","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:39:26.058483Z","iopub.execute_input":"2022-04-20T22:39:26.058733Z","iopub.status.idle":"2022-04-20T22:39:26.063892Z","shell.execute_reply.started":"2022-04-20T22:39:26.058704Z","shell.execute_reply":"2022-04-20T22:39:26.062640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nranker = ranker.fit(\n    train_X,\n    train_y,\n    group=train_baskets,\n)","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:39:26.065851Z","iopub.execute_input":"2022-04-20T22:39:26.066235Z","iopub.status.idle":"2022-04-20T22:39:37.052153Z","shell.execute_reply.started":"2022-04-20T22:39:26.066152Z","shell.execute_reply":"2022-04-20T22:39:37.050867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in ranker.feature_importances_.argsort()[::-1]:\n    print(columns_to_use[i], ranker.feature_importances_[i]/ranker.feature_importances_.sum())","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:39:37.053397Z","iopub.execute_input":"2022-04-20T22:39:37.053625Z","iopub.status.idle":"2022-04-20T22:39:37.066653Z","shell.execute_reply.started":"2022-04-20T22:39:37.053595Z","shell.execute_reply":"2022-04-20T22:39:37.066010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Calculate predictions","metadata":{}},{"cell_type":"code","source":"%time\n\ntest['preds'] = ranker.predict(test_X)\n\nc_id2predicted_article_ids = test \\\n    .sort_values(['customer_id', 'preds'], ascending=False) \\\n    .groupby('customer_id')['article_id'].apply(list).to_dict()\n\nbestsellers_last_week = \\\n    bestsellers_previous_week[bestsellers_previous_week.week == bestsellers_previous_week.week.max()]['article_id'].tolist()","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:39:37.067655Z","iopub.execute_input":"2022-04-20T22:39:37.067954Z","iopub.status.idle":"2022-04-20T22:39:53.030585Z","shell.execute_reply.started":"2022-04-20T22:39:37.067910Z","shell.execute_reply":"2022-04-20T22:39:53.029497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create submission","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv('/kaggle/input/h-and-m-personalized-fashion-recommendations/sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:39:53.031961Z","iopub.execute_input":"2022-04-20T22:39:53.032294Z","iopub.status.idle":"2022-04-20T22:39:58.396470Z","shell.execute_reply.started":"2022-04-20T22:39:53.032254Z","shell.execute_reply":"2022-04-20T22:39:58.395525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\npreds = []\nfor c_id in customer_hex_id_to_int(sub.customer_id):\n    pred = c_id2predicted_article_ids.get(c_id, [])\n    pred = pred + bestsellers_last_week\n    preds.append(pred[:12])","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:39:58.397820Z","iopub.execute_input":"2022-04-20T22:39:58.398078Z","iopub.status.idle":"2022-04-20T22:40:05.785755Z","shell.execute_reply.started":"2022-04-20T22:39:58.398047Z","shell.execute_reply":"2022-04-20T22:40:05.784668Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preds = [' '.join(['0' + str(p) for p in ps]) for ps in preds]\nsub.prediction = preds","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:40:05.787802Z","iopub.execute_input":"2022-04-20T22:40:05.788351Z","iopub.status.idle":"2022-04-20T22:40:12.173740Z","shell.execute_reply.started":"2022-04-20T22:40:05.788296Z","shell.execute_reply":"2022-04-20T22:40:12.172809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_name = 'basic_model_submission'\nsub.to_csv(f'{sub_name}.csv.gz', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-04-20T22:40:12.174919Z","iopub.execute_input":"2022-04-20T22:40:12.175196Z","iopub.status.idle":"2022-04-20T22:40:39.989584Z","shell.execute_reply.started":"2022-04-20T22:40:12.175165Z","shell.execute_reply":"2022-04-20T22:40:39.988387Z"},"trusted":true},"execution_count":null,"outputs":[]}]}