{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"We know that popularity-based recommendation is strong. But, which period should we use for counting popularity? As we can't know what is popular and what is not in the test period, we have to make a best guess based on the sales in the training period. If the period is too short like the last one day, it could be noisy. But if it's too long like the last one year, it could vague the hot trend.\n\nIn this notebook, I investigate which is the optimal period for counting popularity. Again, as we don't have the popularity in the test period, I instead use validation period to confirm its optimality.\n\n# Summary\n It is the best to use only the last day to count popularity.","metadata":{}},{"cell_type":"markdown","source":"# Setups","metadata":{}},{"cell_type":"code","source":"from datetime import datetime, date, timedelta\nfrom typing import List, Tuple, Dict, Union\n\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom scipy import stats\nfrom tqdm import tqdm\n\nsns.set()","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-04-12T22:43:46.179965Z","iopub.execute_input":"2022-04-12T22:43:46.180375Z","iopub.status.idle":"2022-04-12T22:43:47.488212Z","shell.execute_reply.started":"2022-04-12T22:43:46.180280Z","shell.execute_reply":"2022-04-12T22:43:47.487417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Configs","metadata":{}},{"cell_type":"code","source":"LAST_DAY = date(2020,9,22)\nFIRST_DAY = date(2018,9,20)\nLAST_DAY_TRAIN = (LAST_DAY - timedelta(1*7))\ndf = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/customers.csv\")\nCUSTOMERID2INDEX = dict(zip(df[\"customer_id\"], df.index))\nINDEX2CUSTOMERID = dict(zip(df.index, df[\"customer_id\"]))\ndel df","metadata":{"execution":{"iopub.status.busy":"2022-04-12T22:43:47.489803Z","iopub.execute_input":"2022-04-12T22:43:47.490066Z","iopub.status.idle":"2022-04-12T22:43:54.443294Z","shell.execute_reply.started":"2022-04-12T22:43:47.490037Z","shell.execute_reply":"2022-04-12T22:43:54.442458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loading Data","metadata":{"execution":{"iopub.status.busy":"2022-04-12T01:06:17.982186Z","iopub.execute_input":"2022-04-12T01:06:17.983099Z","iopub.status.idle":"2022-04-12T01:06:25.712392Z","shell.execute_reply.started":"2022-04-12T01:06:17.983044Z","shell.execute_reply":"2022-04-12T01:06:25.711441Z"}}},{"cell_type":"code","source":"def load_efficient_df() -> pd.DataFrame:\n    df = pd.read_csv(\n        \"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\",\n        dtype={\"t_dat\": \"object\", \"customer_id\": \"object\", \"article_id\": \"object\", \"price\": float, \"sales_channel_id\": int},\n    )\n    df['t_dat'] = pd.to_datetime(df['t_dat'])\n    # For reducing memory usage\n    df[\"customer_id\"] = df[\"customer_id\"].map(CUSTOMERID2INDEX).astype('int32')\n    df['article_id'] = df['article_id'].astype('int32')\n    return df","metadata":{"execution":{"iopub.status.busy":"2022-04-12T22:43:54.444765Z","iopub.execute_input":"2022-04-12T22:43:54.445107Z","iopub.status.idle":"2022-04-12T22:43:54.452728Z","shell.execute_reply.started":"2022-04-12T22:43:54.445056Z","shell.execute_reply":"2022-04-12T22:43:54.451661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = load_efficient_df()\nN_ARTICLES = df[\"article_id\"].nunique()\ndisplay(df)","metadata":{"execution":{"iopub.status.busy":"2022-04-12T22:43:54.455563Z","iopub.execute_input":"2022-04-12T22:43:54.455890Z","iopub.status.idle":"2022-04-12T22:45:30.978543Z","shell.execute_reply.started":"2022-04-12T22:43:54.455846Z","shell.execute_reply":"2022-04-12T22:45:30.977463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Counting Popularity with Different Periods","metadata":{}},{"cell_type":"markdown","source":"Here is the \"target\" popularity, which is counted in the last week of the training data. This is what we want to predict/approximate in this experiment. ","metadata":{}},{"cell_type":"code","source":"ranks = df.query(f\"t_dat > '{LAST_DAY_TRAIN}'\").groupby(\"article_id\")[\"customer_id\"].agg(pd.Series.nunique) \\\n                .rank(ascending=False)\nranks.name = \"target\"\ndisplay(ranks)","metadata":{"execution":{"iopub.status.busy":"2022-04-12T22:45:30.980401Z","iopub.execute_input":"2022-04-12T22:45:30.980751Z","iopub.status.idle":"2022-04-12T22:45:32.686231Z","shell.execute_reply.started":"2022-04-12T22:45:30.980704Z","shell.execute_reply":"2022-04-12T22:45:32.685328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We are interested in the popular articles. Here is how they fluctuate in ranks by changing the length of the period from 1 day to 4 weeks.","metadata":{}},{"cell_type":"code","source":"deltas = list(range(1, 4*7+1))\nfor delta in tqdm(deltas):\n    past = df.query(f\"t_dat > '{LAST_DAY_TRAIN - timedelta(delta)}' and t_dat <= '{LAST_DAY_TRAIN}'\") \\\n        .groupby(\"article_id\")[\"customer_id\"].agg(pd.Series.nunique).rank(ascending=False)\n    past.name = f\"{delta}D\"\n    ranks = pd.merge(ranks, past, on=\"article_id\", how=\"left\").sort_values(\"target\")[:12]\ndisplay(ranks)","metadata":{"execution":{"iopub.status.busy":"2022-04-12T22:45:32.687671Z","iopub.execute_input":"2022-04-12T22:45:32.688028Z","iopub.status.idle":"2022-04-12T22:46:22.898820Z","shell.execute_reply.started":"2022-04-12T22:45:32.687961Z","shell.execute_reply":"2022-04-12T22:46:22.897961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(12):\n    plt.plot(range(28, -1, -1), ranks.iloc[i].values[::-1].clip(0,24))\nplt.gca().invert_yaxis()\nplt.xlabel(\"Length of counting period [days]\")\nplt.ylabel(\"Rank\")\nplt.title(\"Change in rankings of top 12 articles\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-04-12T22:46:22.900037Z","iopub.execute_input":"2022-04-12T22:46:22.900246Z","iopub.status.idle":"2022-04-12T22:46:23.217178Z","shell.execute_reply.started":"2022-04-12T22:46:22.900220Z","shell.execute_reply":"2022-04-12T22:46:23.216090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To measure how different each ranking is to the target ranking, I use the following two metrics:\n- Spearman's rank correlation coefficient (equivalent to Pearson's $r$ of ranks)\n- Rank-weighted reciprocal rank (I extended the idea of mean reciprocal rank)\n\nNOTE: It is also possible to validate by MAP@12, the true evaluation metric, but I omit it here for brevity. Calculating MAP@12 involves arbitrary choice of which customers to give popularity-based recommendations and which to other algorithms.","metadata":{}},{"cell_type":"code","source":"def spearmanr(y_true, y_pred):\n    \"\"\"Spearman's rank correlation coefficient\"\"\"\n    score, _ = stats.pearsonr(y_true, y_pred)\n    return score\n\ndef rwrr(y_true, y_pred):\n    \"\"\"Rank-weighted reciprocal rank\"\"\"\n    score = 0\n    for i, rank in enumerate(y_pred):\n        score += (1 / rank) / (i+1)\n    return score","metadata":{"execution":{"iopub.status.busy":"2022-04-12T22:46:23.218986Z","iopub.execute_input":"2022-04-12T22:46:23.219328Z","iopub.status.idle":"2022-04-12T22:46:23.226267Z","shell.execute_reply.started":"2022-04-12T22:46:23.219281Z","shell.execute_reply":"2022-04-12T22:46:23.225256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y1s, y2s = [], []\nfor delta in deltas:\n    y1s.append(spearmanr(ranks[\"target\"], ranks[f\"{delta}D\"]))\n    y2s.append(rwrr(ranks[\"target\"], ranks[f\"{delta}D\"]))\n","metadata":{"execution":{"iopub.status.busy":"2022-04-12T22:46:23.228127Z","iopub.execute_input":"2022-04-12T22:46:23.228640Z","iopub.status.idle":"2022-04-12T22:46:23.252007Z","shell.execute_reply.started":"2022-04-12T22:46:23.228599Z","shell.execute_reply":"2022-04-12T22:46:23.251276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(deltas, y1s)\nplt.xlabel(\"Length of counting period [days]\")\nplt.ylabel(\"Rank correlation (Spearman)\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-04-12T22:46:23.254276Z","iopub.execute_input":"2022-04-12T22:46:23.255162Z","iopub.status.idle":"2022-04-12T22:46:23.486890Z","shell.execute_reply.started":"2022-04-12T22:46:23.255123Z","shell.execute_reply":"2022-04-12T22:46:23.485909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(deltas, y2s)\nplt.xlabel(\"Length of counting period [days]\")\nplt.ylabel(\"Rank-weighted reciprocal rank\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-04-12T22:46:23.489667Z","iopub.execute_input":"2022-04-12T22:46:23.490008Z","iopub.status.idle":"2022-04-12T22:46:23.731059Z","shell.execute_reply.started":"2022-04-12T22:46:23.489910Z","shell.execute_reply":"2022-04-12T22:46:23.730024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Both metrics tell us it is the best to **use only the last day** to count popularity.","metadata":{}}]}