{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Validation Notebook for My Candidate ReRank Model\nIn this notebook, we attempted to develop a two-stage model which includes the candidate generation model (Covisitation Matrix) and Ranking Model. \nThis practice is widely used in big tech company since the candidate generation \n\nIt should be noted that the candidate generation model should target for high recall while the ranking model should target for to rank the most relevent item first.\n","metadata":{"papermill":{"duration":0.005761,"end_time":"2022-11-10T16:03:24.627071","exception":false,"start_time":"2022-11-10T16:03:24.62131","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# Introduction of Notebook\n\n\n## Step 1: Model Training\n### Step 1.1 - Loading Training Data\nThe training data in this notebook is extracted by this logic: <br>\n`train_df = train_df[train_df['session']%10 == 1]` <br>\n\nThe label of the training data is stored in the test_labels.parquet, which **contain the label for both training and testing data (For quick experimet only).**\n\n\n### Step 1.2 - Feature Engineering on Training Data\nAll the feature is pre-calculated and saved in parquet files. All parquet file is saved in this kaggle dataset `/kaggle/input/otto-validation` <br>\nHere we only do a simple joining between the training data & pre-calculated features\n\n\n### Step 1.3 - Model Training on Training data\nWe train the LGBM ranking using the training data.\n\n<br>\n\n## Step 2: Model Inference\n### Step 2.1 - Loading testing data\nThe training data in this notebook is extracted by this logic: <br>\n`test_df = test_df[test_df['session']%10 == 0]`\n\n\n### Step 2.2 - Generate Candidates\nWe use the logic from the Chris Deotte's Candidate ReRank Model to generate 40 potential candidates. We have pre-calculated all the co-visitation dictionary and saved in the kaggle dataset \n`/kaggle/input/otto-covisitation-matrix-parquet-files` <br>\nIn this stage, we will generate 40 candidates per session and pass into the ranker for the final ranking. <br>\n\nPlease refer to Chris notebook for detailed logic explaination\nhttps://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575 <br>\n\n### Step 2.3 - Ranker Model\n1. The recommmended aid from the candidate generation phase will merge with the test aid for ranking.\n2. Feature preprocessing step is exactly the same with the training pipeline\n3. The test dataframe will be passed for prediction and the scores will be merged with the dataframe\n4. Sort each session in the dataframe by score from high to low\n5. Using `groupby('session').last(20) to extract top 20 result\n\n\n### Step 2.4 - Export to CSV\n<br>\n\n## Step 3: Model Evaluation\nSame logic with the Chris notebook: https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565","metadata":{}},{"cell_type":"markdown","source":"# Credits\n1. The validation part of the notebook comes from Chris Deotte Notebook: https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565 <br>\n2. The covisition matrix calculation comes from Chris Deotte Notebook: https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575 <br>\n3. The code of LGBM ranker is inspired by RADEK OSMULSKI in his Notebook: https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker <br>\n\n# Important Note!\nThis notebook is still in development phase so it is using the local validation dataset instead of the full training and testing dataset","metadata":{"execution":{"iopub.status.busy":"2023-01-07T23:01:16.145377Z","iopub.execute_input":"2023-01-07T23:01:16.145999Z","iopub.status.idle":"2023-01-07T23:01:16.158308Z","shell.execute_reply.started":"2023-01-07T23:01:16.145950Z","shell.execute_reply":"2023-01-07T23:01:16.156429Z"}}},{"cell_type":"code","source":"! pip install polars","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:48:30.369387Z","iopub.execute_input":"2023-01-08T23:48:30.369952Z","iopub.status.idle":"2023-01-08T23:48:43.413600Z","shell.execute_reply.started":"2023-01-08T23:48:30.369829Z","shell.execute_reply":"2023-01-08T23:48:43.412304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"VER = 6\n\nimport pandas as pd, numpy as np\nfrom tqdm.notebook import tqdm\nimport os, sys, pickle, glob, gc\nfrom collections import Counter\nimport itertools\nimport polars as pl\nfrom gensim.models import Word2Vec\n\n# Balance of type weighting タイプごとの重み付けバランス\n# 0:clicks 1:carts 2:orders\ntype_weight = {0:0.5,\n               1:9,\n               2:0.5}\ntype_weight_multipliers = type_weight\n\n# Use top X for clicks, carts and orders Top何位までを使うか\nclicks_th = 15 # クリック数\ncarts_th  = 20 # カート数\norders_th = 20 # 購入数\n\nVER = 7\n\ntype_labels = {'clicks':0, 'carts':1, 'orders':2}","metadata":{"papermill":{"duration":2.845087,"end_time":"2022-11-10T16:03:27.484916","exception":false,"start_time":"2022-11-10T16:03:24.639829","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-01-08T23:48:43.416024Z","iopub.execute_input":"2023-01-08T23:48:43.416429Z","iopub.status.idle":"2023-01-08T23:48:43.679260Z","shell.execute_reply.started":"2023-01-08T23:48:43.416390Z","shell.execute_reply":"2023-01-08T23:48:43.678025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 1: Model Training","metadata":{}},{"cell_type":"markdown","source":"## Step 1.1 - Generate Training Data","metadata":{}},{"cell_type":"markdown","source":"There are two types of aid we will reocommned to the users: \n1. Data from the training data set (Actual aid in the dataset [i.e. the actual behavior of the user. User may cart the item that they clicked or they may click again for same item])\n2. Data from candiate generation (Recommended aid from the candidate generation logic [i.e. The item that is related to the actual user behavior])\n\nTherefore, our ranker should be able to rank these two items at the same time.\n\nWe want to simulate the this data for inference time with recommended candidates. Therefore, we need to generate some recommended item  in the training size to make sure training set has same distribution as inference time. <br>\n","metadata":{}},{"cell_type":"markdown","source":"### Generating Training data from the test set (Actual behavior)","metadata":{}},{"cell_type":"code","source":"# Getting the actual aid from dataset\n\ndef load_train_data_sampled():    \n    dfs = []\n    for e, chunk_file in enumerate(glob.glob('../input/otto-validation/test_parquet/*')):\n        chunk = pd.read_parquet(chunk_file)\n        chunk.ts = (chunk.ts/1000).astype('int32')\n        chunk['type'] = chunk['type'].map(type_labels).astype('int8')\n        dfs.append(chunk)\n        \n    train_df = pd.concat(dfs).reset_index(drop=True) #.astype({\"ts\": \"datetime64[ms]\"})\n\n    # Using different sample as the candidate generation for training\n    train_df = train_df[train_df['session']%10 == 1]\n    return train_df\n\ntrain_df = load_train_data_sampled()\nprint('Sampled Training data has shape',train_df.shape)\n\n\n# This indicate that this aid is actual behaviors\ntrain_df['real_action'] = 1\n\n# CG stands for candidate generation. Since the aid here is actual user behavior, they should have no ranking\ntrain_df['CG_ranking'] = 0\n\n# Only focus on click action first\ntrain_df_click = train_df[train_df['type'] == 0]\n\n# Calculate the last three aid for the embedding calculation\ntrain_df_click['aid_last'] = train_df_click.groupby(['session']).aid.shift(1).bfill()\ntrain_df_click['aid_second_last'] = train_df_click.groupby(['session']).aid.shift(2).bfill()\ntrain_df_click['aid_third_last'] = train_df_click.groupby(['session']).aid.shift(3).bfill()","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:48:43.680713Z","iopub.execute_input":"2023-01-08T23:48:43.682321Z","iopub.status.idle":"2023-01-08T23:48:46.972085Z","shell.execute_reply.started":"2023-01-08T23:48:43.682282Z","shell.execute_reply":"2023-01-08T23:48:46.970737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df_click = pl.from_pandas(train_df_click)","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:49:15.845842Z","iopub.execute_input":"2023-01-08T23:49:15.846328Z","iopub.status.idle":"2023-01-08T23:49:15.879193Z","shell.execute_reply.started":"2023-01-08T23:49:15.846288Z","shell.execute_reply":"2023-01-08T23:49:15.878032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df_click = train_df_click.with_columns([\n    pl.col('aid_last').cast(pl.Int32),\n    pl.col('aid_second_last').cast(pl.Int32),\n    pl.col('aid_third_last').cast(pl.Int32),\n])","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:49:15.880898Z","iopub.execute_input":"2023-01-08T23:49:15.881261Z","iopub.status.idle":"2023-01-08T23:49:15.899377Z","shell.execute_reply.started":"2023-01-08T23:49:15.881229Z","shell.execute_reply":"2023-01-08T23:49:15.898402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Generating Training data from candidate generation","metadata":{}},{"cell_type":"markdown","source":"The below block of code is exactly the same code we used for candidate generation based on the candidate ReRank Model. <br>\nYou may refer here for the detailed logic: https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575","metadata":{}},{"cell_type":"code","source":"%%time\n# Generating recommended aid from the actual aid based on candidate generation\ntop_clicks = train_df.loc[train_df['type']== 0,'aid'].value_counts().index.values[:20] \n\n\n# Improved speed for 2X using polars. \ndef pqt_to_dict(path):\n    return pl.read_parquet(path).groupby('aid_x').agg(pl.col('aid_y').list()).to_pandas().set_index('aid_x').aid_y.apply(list).to_dict()\n\nDISK_PIECES = 4\n\n# LOAD THREE CO-VISITATION MATRICES\ntop_20_clicks = pqt_to_dict(f'/kaggle/input/otto-covisitation-matrix-parquet-files/top_20_clicks_v{VER}_0.pqt')\n\nfor k in range(1,DISK_PIECES): \n    top_20_clicks.update(pd.read_parquet(f'/kaggle/input/otto-covisitation-matrix-parquet-files/top_20_clicks_v{VER}_{k}.pqt') ) \n\ndef suggest_clicks(df):\n    # USER HISTORY AIDS AND TYPES\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    # RERANK CANDIDATES USING WEIGHTS\n    if len(unique_aids)>=20:\n        weights=np.logspace(0.1,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        # RERANK BASED ON REPEAT ITEMS AND TYPE OF ITEMS\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    # USE \"CLICKS\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(40) if aid2 not in unique_aids]    \n    result = unique_aids + top_aids2#[:20 - len(unique_aids)]\n    # USE TOP20 TEST CLICKS\n    return result + list(top_clicks)#[:20-len(result)]\n\n\npred_df_clicks = train_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(\n    lambda x: suggest_clicks(x)\n)\n\ntrain_df_click_recommended = pl.from_pandas(pd.DataFrame(pred_df_clicks, columns = ['aid']).reset_index()).explode(\"aid\")\ntrain_df_click_recommended = train_df_click_recommended.with_columns([\n    pl.lit(0).alias('ts').cast(pl.Int32),    \n    pl.lit(0).alias('type').cast(pl.Int8),\n    pl.lit(0).alias('real_action').cast(pl.Int64),\n    pl.col('session').cast(pl.Int32),\n    pl.col('aid').cast(pl.Int32),\n])\nn_col_after_join = train_df_click_recommended.groupby('session').agg([\n    pl.col('aid').cumcount().alias('CG_ranking')]).select(\n    pl.col('CG_ranking').explode().cast(pl.Int64))\ntrain_df_click_recommended = pl.concat([train_df_click_recommended, n_col_after_join], how=\"horizontal\")","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:49:15.903385Z","iopub.execute_input":"2023-01-08T23:49:15.904536Z","iopub.status.idle":"2023-01-08T23:49:59.234840Z","shell.execute_reply.started":"2023-01-08T23:49:15.904486Z","shell.execute_reply":"2023-01-08T23:49:59.233520Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Step 1.2 - Feature Engineering","metadata":{}},{"cell_type":"markdown","source":"### Calculating the sparse feature","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:38:44.903033Z","iopub.execute_input":"2023-01-08T23:38:44.903470Z","iopub.status.idle":"2023-01-08T23:38:44.908465Z","shell.execute_reply.started":"2023-01-08T23:38:44.903437Z","shell.execute_reply":"2023-01-08T23:38:44.907193Z"}}},{"cell_type":"code","source":"model = Word2Vec.load(\"/kaggle/input/ottoprecalculatedfeatureparquet/word2vec.model\")\nembedding_weight = np.load('/kaggle/input/ottoprecalculatedfeatureparquet/word2vec.model.wv.vectors.npy')\nembedding_weight_neg = np.load('/kaggle/input/ottoprecalculatedfeatureparquet/word2vec.model.syn1neg.npy')\n\nembedding_weigh_dict_df = pl.from_pandas(pd.DataFrame(embedding_weight, columns = ['Embedding_' + str(x) for x in range(32)]).reset_index().rename(columns = {'index':'aid'})).with_columns(pl.col('aid').cast(pl.Int32))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n\n\n# Calculating the embedding of last three actions aid\ntrain_df_click = train_df_click.join(embedding_weigh_dict_df, on = 'aid', how = 'left').join(embedding_weigh_dict_df, right_on = 'aid', left_on = 'aid_last', how = 'left', suffix = 'last_1').join(embedding_weigh_dict_df, right_on = 'aid', left_on = 'aid_second_last', how = 'left', suffix = 'last_2').join(embedding_weigh_dict_df, right_on = 'aid', left_on = 'aid_third_last', suffix = 'last_3').select(pl.exclude(['aid_last', 'aid_second_last', 'aid_third_last']))","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:49:59.236317Z","iopub.execute_input":"2023-01-08T23:49:59.236764Z","iopub.status.idle":"2023-01-08T23:50:00.605923Z","shell.execute_reply.started":"2023-01-08T23:49:59.236726Z","shell.execute_reply":"2023-01-08T23:50:00.604558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculating the embedding of last three actions aid for the recommended action. We all use the last three actual aid embedding for the last three aid\ntrain_df_click_last_action_embedding = train_df_click.sort(['session', 'ts']).groupby(['session']).last().select(pl.exclude(['aid', 'ts','type','real_action', 'CG_ranking', 'aid_last', 'aid_second_last', 'aid_third_last'] + ['Embedding_' + str(i) for i in range(32)]))\n\ntrain_df_click_recommended = train_df_click_recommended.join(embedding_weigh_dict_df, on = 'aid', how = 'left').join(train_df_click_last_action_embedding, on = 'session', how = 'left')","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:50:00.608349Z","iopub.execute_input":"2023-01-08T23:50:00.608896Z","iopub.status.idle":"2023-01-08T23:50:03.109992Z","shell.execute_reply.started":"2023-01-08T23:50:00.608848Z","shell.execute_reply":"2023-01-08T23:50:03.108864Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Calculating the dense feature","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:38:57.098170Z","iopub.execute_input":"2023-01-08T23:38:57.098604Z","iopub.status.idle":"2023-01-08T23:38:57.103904Z","shell.execute_reply.started":"2023-01-08T23:38:57.098567Z","shell.execute_reply":"2023-01-08T23:38:57.102479Z"}}},{"cell_type":"code","source":"# Creating feature by joining with pre-computed feature\naid_global_counter_all_types = pl.read_parquet('/kaggle/input/ottoprecalculatedfeatureparquet/aid_global_counter_all_types.pqt')\n\naid_global_user_counter_all_types = pl.read_parquet('/kaggle/input/ottoprecalculatedfeatureparquet/aid_global_user_counter_all_types.pqt')\n\naid_global_user_counter_all_types_time_weighted = pl.read_parquet('/kaggle/input/ottoprecalculatedfeatureparquet/aid_global_user_counter_all_types_time_weighted (1).pqt')","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:50:03.111058Z","iopub.execute_input":"2023-01-08T23:50:03.111432Z","iopub.status.idle":"2023-01-08T23:50:03.508221Z","shell.execute_reply.started":"2023-01-08T23:50:03.111400Z","shell.execute_reply":"2023-01-08T23:50:03.507099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df_click = train_df_click.join(aid_global_counter_all_types, on='aid', suffix ='_global_counter').join(aid_global_user_counter_all_types, on='aid', suffix ='_user_counter').join(aid_global_user_counter_all_types_time_weighted, on='aid', suffix ='_timed_global_counter')\ntrain_df_click_recommended = train_df_click_recommended.join(aid_global_counter_all_types, on='aid', suffix ='_global_counter').join(aid_global_user_counter_all_types, on='aid', suffix ='_user_counter').join(aid_global_user_counter_all_types_time_weighted, on='aid', suffix ='_timed_global_counter')","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:50:03.509319Z","iopub.execute_input":"2023-01-08T23:50:03.509671Z","iopub.status.idle":"2023-01-08T23:50:12.362388Z","shell.execute_reply.started":"2023-01-08T23:50:03.509638Z","shell.execute_reply":"2023-01-08T23:50:12.361267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Merging two training dataset together","metadata":{}},{"cell_type":"code","source":"# Merging two type of data as training data\ntrain_df_click_all = pl.concat([train_df_click_recommended, train_df_click], how = 'vertical')","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:50:12.363625Z","iopub.execute_input":"2023-01-08T23:50:12.363995Z","iopub.status.idle":"2023-01-08T23:50:15.592657Z","shell.execute_reply.started":"2023-01-08T23:50:12.363948Z","shell.execute_reply":"2023-01-08T23:50:15.591528Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 1.3 - Model Training","metadata":{}},{"cell_type":"code","source":"# Merging the Ground Truth label with training dataset\n# Using negative downsampling of 50%\n\ntrain_labels = pd.read_parquet('../input/otto-validation/test_labels.parquet')\ntrain_labels['type'] = train_labels['type'].map(type_labels).astype('int8')\ntrain_labels = pl.from_pandas(train_labels)\ntrain_labels.head()\ntrain_labels = train_labels.explode('ground_truth').with_columns([pl.col('ground_truth').alias('aid'), pl.lit(1).alias('label')]).with_columns([\n    pl.col('ground_truth').cast(pl.Int32),\n    pl.col('session').cast(pl.Int32),\n    pl.col('aid').cast(pl.Int32),\n])\n\ntrain_df_click_all_sampled =  train_df_click_all.sample(n= int(len(train_df_click_all)*0.5))\n\ntrain_df_click_all_sampled = train_df_click_all_sampled.join(train_labels, how='left', on=['session', 'type', 'aid']).with_column(pl.col('label').fill_null(0))","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:50:15.593840Z","iopub.execute_input":"2023-01-08T23:50:15.595089Z","iopub.status.idle":"2023-01-08T23:50:20.764525Z","shell.execute_reply.started":"2023-01-08T23:50:15.595043Z","shell.execute_reply":"2023-01-08T23:50:20.763124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n# Training the model. Seems LGBMRanker train pretty fast. Should be able to add more features\nfrom lightgbm.sklearn import LGBMRanker\n\nranker = LGBMRanker(\n    objective=\"lambdarank\",\n    metric=\"ndcg\",\n    boosting_type=\"dart\",\n    n_estimators=20,\n    importance_type='gain',\n)\n\nfeature_cols = [\n 'ts',\n 'real_action',\n 'CG_ranking',\n 'orders',\n 'clicks',\n 'carts',\n 'carts_user_counter',\n 'clicks_user_counter',\n 'orders_user_counter',\n 'carts_timed_global_counter',\n 'orders_timed_global_counter',\n 'clicks_timed_global_counter',]\n\ntarget = 'label'\n\ndef get_session_lenghts(df):\n    return df.groupby('session').agg([\n        pl.col('session').count().alias('session_length')\n    ])['session_length'].to_numpy()\n\nsession_lengths_train = get_session_lenghts(train_df_click_all_sampled)\n\nranker = ranker.fit(\n    train_df_click_all_sampled[feature_cols].to_pandas(),\n    train_df_click_all_sampled[target].to_pandas(),\n    group=session_lengths_train,\n)","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:50:20.766298Z","iopub.execute_input":"2023-01-08T23:50:20.766650Z","iopub.status.idle":"2023-01-08T23:50:29.771814Z","shell.execute_reply.started":"2023-01-08T23:50:20.766620Z","shell.execute_reply":"2023-01-08T23:50:29.770640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n# Step 2: Model Inference","metadata":{}},{"cell_type":"markdown","source":"## Step 2.1: Loading testing data and pre-calculated co-visitation Matrix\nFor quicker experiment, we will only use 1/10 of the validation test_parquet \n\nWe use another set of session to mimic the test set pattern. \n\nWe extract a different set of sesion using the function  test_df[test_df['session']%10 == 0]","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[]}},{"cell_type":"code","source":"def load_test():    \n    dfs = []\n    for e, chunk_file in enumerate(glob.glob('../input/otto-validation/test_parquet/*')):\n        chunk = pd.read_parquet(chunk_file)\n        chunk.ts = (chunk.ts/1000).astype('int32')\n        chunk['type'] = chunk['type'].map(type_labels).astype('int8')\n        dfs.append(chunk)\n    return pd.concat(dfs).reset_index(drop=True) #.astype({\"ts\": \"datetime64[ms]\"})\n\ntest_df = load_test()\ntest_df = test_df[test_df['session']%10 == 0]\nprint('Test data has shape',test_df.shape)\ntest_df.head()","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[],"execution":{"iopub.status.busy":"2023-01-08T23:50:29.776140Z","iopub.execute_input":"2023-01-08T23:50:29.779036Z","iopub.status.idle":"2023-01-08T23:50:31.390232Z","shell.execute_reply.started":"2023-01-08T23:50:29.778959Z","shell.execute_reply":"2023-01-08T23:50:31.389250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have pre-calcaulted the covisitatioin matrix and saved them as the parquet file. Here we only load the result matrix from kaggle dataset to save time<br>\nTo understand how to generate the co-vistation matrix, you can see here:\nhttps://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575","metadata":{}},{"cell_type":"code","source":"%%time\n# Improved speed for 2X using polars. \ndef pqt_to_dict(path):\n    return pl.read_parquet(path).groupby('aid_x').agg(pl.col('aid_y').list()).to_pandas().set_index('aid_x').aid_y.apply(list).to_dict()\n\nDISK_PIECES = 4\n\n# LOAD THREE CO-VISITATION MATRICES\ntop_20_clicks = pqt_to_dict(f'/kaggle/input/otto-covisitation-matrix-parquet-files/top_20_clicks_v{VER}_0.pqt')\n\nfor k in range(1,DISK_PIECES): \n    top_20_clicks.update(pd.read_parquet(f'/kaggle/input/otto-covisitation-matrix-parquet-files/top_20_clicks_v{VER}_{k}.pqt') ) \n\n\ntop_20_buys = pqt_to_dict(f'/kaggle/input/otto-covisitation-matrix-parquet-files/top_15_carts_orders_v{VER}_0.pqt') \n\nfor k in range(1,DISK_PIECES): \n    top_20_buys.update( pqt_to_dict( f'/kaggle/input/otto-covisitation-matrix-parquet-files/top_15_carts_orders_v{VER}_{k}.pqt') )\n\ntop_20_buy2buy = pqt_to_dict(f'/kaggle/input/otto-covisitation-matrix-parquet-files/top_15_buy2buy_v{VER}_0.pqt') \n\nprint('Here are size of our 3 co-visitation matrices:')\nprint( len( top_20_clicks ), len( top_20_buy2buy ), len( top_20_buys ) )","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[],"execution":{"iopub.status.busy":"2023-01-08T23:50:31.392116Z","iopub.execute_input":"2023-01-08T23:50:31.392894Z","iopub.status.idle":"2023-01-08T23:50:58.070696Z","shell.execute_reply.started":"2023-01-08T23:50:31.392846Z","shell.execute_reply":"2023-01-08T23:50:58.069321Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 2.2 Candidate Generation using ReRank Model\n\nSame logic with the candidate ReRank model. There is a changes made for the clicks suggestion (See the comment below)","metadata":{}},{"cell_type":"code","source":"top_clicks = test_df.loc[test_df['type']== 0,'aid'].value_counts().index.values[:20] \ntop_carts = test_df.loc[test_df['type']== 1,'aid'].value_counts().index.values[:20]\ntop_orders = test_df.loc[test_df['type']== 2,'aid'].value_counts().index.values[:20]","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:50:58.072618Z","iopub.execute_input":"2023-01-08T23:50:58.073043Z","iopub.status.idle":"2023-01-08T23:50:58.147957Z","shell.execute_reply.started":"2023-01-08T23:50:58.073005Z","shell.execute_reply":"2023-01-08T23:50:58.146598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def suggest_clicks(df):\n    # USER HISTORY AIDS AND TYPES\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    # RERANK CANDIDATES USING WEIGHTS\n    if len(unique_aids)>=20:\n        weights=np.logspace(0.1,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        # RERANK BASED ON REPEAT ITEMS AND TYPE OF ITEMS\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    # USE \"CLICKS\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(20) if aid2 not in unique_aids]    \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    # USE TOP20 TEST CLICKS\n    return result + list(top_clicks)[:20-len(result)]","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[],"execution":{"iopub.status.busy":"2023-01-08T23:50:58.150083Z","iopub.execute_input":"2023-01-08T23:50:58.150513Z","iopub.status.idle":"2023-01-08T23:50:58.191504Z","shell.execute_reply.started":"2023-01-08T23:50:58.150474Z","shell.execute_reply":"2023-01-08T23:50:58.190394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I have re-write the function and made the changes below\n\n1. Changing recommended aid number from 20 to 40\nThe reason is that if only 20 is recommended, the ranker actually will not boost perofrmance since the LeaderBoard Recall score is calcualted **regardless** of the order. <br>\nTherefore, here we recommnd more candidate to increase the **Recall metrics (i.e. Total coverage of ground truth aid that is recommeneded in the candidate recommendation)**\n\n2. aid in test set will not be recommened\nThe aid in test set will be handled seperately since they have ts information and should not have CG_ranking (Candidate generation ranking) feature info.","metadata":{}},{"cell_type":"code","source":"top_clicks = test_df.loc[test_df['type']== 0,'aid'].value_counts().index.values[:40] \ndef suggest_clicks_40_candidates(df):\n    # USER HISTORY AIDS AND TYPES\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    # RERANK CANDIDATES USING WEIGHTS\n    if len(unique_aids)>=40:\n        weights=np.logspace(0.1,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        # RERANK BASED ON REPEAT ITEMS AND TYPE OF ITEMS\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        sorted_aids = [k for k,v in aids_temp.most_common(40)]\n        return sorted_aids\n    # USE \"CLICKS\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(40) if aid2 not in unique_aids]    \n    result = top_aids2[:40]\n    # USE TOP20 TEST CLICKS\n    return result + list(top_clicks)[:40-len(result)]","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:50:58.193419Z","iopub.execute_input":"2023-01-08T23:50:58.194228Z","iopub.status.idle":"2023-01-08T23:50:58.264527Z","shell.execute_reply.started":"2023-01-08T23:50:58.194188Z","shell.execute_reply":"2023-01-08T23:50:58.263199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def suggest_carts(df):\n    # User history aids and types\n    aids = df.aid.tolist()\n    types = df.type.tolist()\n    \n    # UNIQUE AIDS AND UNIQUE BUYS\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    df = df.loc[(df['type'] == 0)|(df['type'] == 1)]\n    unique_buys = list(dict.fromkeys(df.aid.tolist()[::-1]))\n    \n    # Rerank candidates using weights\n    if len(unique_aids) >= 20:\n        weights=np.logspace(0.5,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        \n        # Rerank based on repeat items and types of items\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        \n        # Rerank candidates using\"top_20_carts\" co-visitation matrix\n        aids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_buys if aid in top_20_buys]))\n        for aid in aids2: aids_temp[aid] += 0.1\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    \n    # Use \"cart order\" and \"clicks\" co-visitation matrices\n    aids1 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    aids2 = list(itertools.chain(*[top_20_buys[aid]*2 for aid in unique_aids if aid in top_20_buys]))\n    \n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids1+aids2).most_common(20) if aid2 not in unique_aids] \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    \n    # USE TOP20 TEST ORDERS\n\n    return result + list(top_carts)[:20-len(result)]","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:50:58.266301Z","iopub.execute_input":"2023-01-08T23:50:58.266662Z","iopub.status.idle":"2023-01-08T23:50:58.314124Z","shell.execute_reply.started":"2023-01-08T23:50:58.266629Z","shell.execute_reply":"2023-01-08T23:50:58.313019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def suggest_buys(df):\n    \n    # USER HISTORY AIDS AND TYPES\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    # UNIQUE AIDS AND UNIQUE BUYS\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    df = df.loc[(df['type']==1)|(df['type']==2)]\n    unique_buys = list(dict.fromkeys( df.aid.tolist()[::-1] ))\n    # RERANK CANDIDATES USING WEIGHTS\n    if len(unique_aids)>=20:\n        weights=np.logspace(0.5,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        # RERANK BASED ON REPEAT ITEMS AND TYPE OF ITEMS\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        # RERANK CANDIDATES USING \"BUY2BUY\" CO-VISITATION MATRIX\n        aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n        for aid in aids3: aids_temp[aid] += 0.1\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    # USE \"CART ORDER\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_aids if aid in top_20_buys]))\n    # USE \"BUY2BUY\" CO-VISITATION MATRIX\n    aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2+aids3).most_common(20) if aid2 not in unique_aids] \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    # USE TOP20 TEST ORDERS\n    return result + list(top_orders)[:20-len(result)]","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:50:58.315703Z","iopub.execute_input":"2023-01-08T23:50:58.316574Z","iopub.status.idle":"2023-01-08T23:50:58.331005Z","shell.execute_reply.started":"2023-01-08T23:50:58.316536Z","shell.execute_reply":"2023-01-08T23:50:58.329554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Generate candidates for each action using co-visitation matrix","metadata":{}},{"cell_type":"code","source":"from pandarallel import pandarallel","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:50:58.332882Z","iopub.execute_input":"2023-01-08T23:50:58.333300Z","iopub.status.idle":"2023-01-08T23:50:58.398214Z","shell.execute_reply.started":"2023-01-08T23:50:58.333265Z","shell.execute_reply":"2023-01-08T23:50:58.396851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Using pandarallel to accelerate\nfrom pandarallel import pandarallel\npandarallel.initialize(nb_workers = 4,progress_bar=True)","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:50:58.402504Z","iopub.execute_input":"2023-01-08T23:50:58.403260Z","iopub.status.idle":"2023-01-08T23:50:58.412802Z","shell.execute_reply.started":"2023-01-08T23:50:58.403207Z","shell.execute_reply":"2023-01-08T23:50:58.411457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Improved speed for 2X using pandarallel\n# pred_df_clicks = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).parallel_apply(\n#     lambda x: suggest_clicks(x)\n# )\n\n# Improved speed for 2X using pandarallel\npred_df_clicks = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).parallel_apply(\n    lambda x: suggest_clicks_40_candidates(x)\n)\n\npred_df_buys = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).parallel_apply(\n    lambda x: suggest_buys(x)\n)\n\npred_df_carts = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).parallel_apply(\n    lambda x: suggest_carts(x)\n)","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:50:58.414652Z","iopub.execute_input":"2023-01-08T23:50:58.415085Z","iopub.status.idle":"2023-01-08T23:56:11.040920Z","shell.execute_reply.started":"2023-01-08T23:50:58.415047Z","shell.execute_reply":"2023-01-08T23:56:11.039369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 2.3 Candidate Ranking using LGBM Ranker","metadata":{}},{"cell_type":"markdown","source":"To run the ranker, we need to combine two sources of data\n1. aid provided by the test set\n2. aid recommended in the candidate generation stage\n\nThese two types of aid should be handled seperately since some of their feature is different (e.g. aid from candidate generation does not haev ts information)","metadata":{}},{"cell_type":"markdown","source":"### Handle aid from test set","metadata":{}},{"cell_type":"code","source":"# Extracting the testing data action for clicks only\ntest_df_click = test_df[test_df['type'] == 0]\ntest_df_click['real_action'] = 1\ntest_df_click['CG_ranking'] = 0\n\n# Calculate the last three aid for the embedding calculation\ntest_df_click['aid_last'] = test_df_click.groupby(['session']).aid.shift(1).bfill()\ntest_df_click['aid_second_last'] = test_df_click.groupby(['session']).aid.shift(2).bfill()\ntest_df_click['aid_third_last'] = test_df_click.groupby(['session']).aid.shift(3).bfill()\n\ntest_df_click = pl.from_pandas(test_df_click)\n\ntest_df_click = test_df_click.with_columns([\n    pl.col('aid_last').cast(pl.Int32),\n    pl.col('aid_second_last').cast(pl.Int32),\n    pl.col('aid_third_last').cast(pl.Int32),\n])","metadata":{"execution":{"iopub.status.busy":"2023-01-09T00:01:51.386753Z","iopub.execute_input":"2023-01-09T00:01:51.387932Z","iopub.status.idle":"2023-01-09T00:01:51.626475Z","shell.execute_reply.started":"2023-01-09T00:01:51.387875Z","shell.execute_reply":"2023-01-09T00:01:51.625000Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Handle aid from candidate recommenedation","metadata":{}},{"cell_type":"code","source":"# Extracting the candiddate generated from the testing data\nclicks_candidate_df = pl.from_pandas(pd.DataFrame(pred_df_clicks, columns = ['aid']).reset_index())\nclicks_candidate_df = clicks_candidate_df.explode('aid')\nclicks_candidate_df = clicks_candidate_df.with_columns([\n    pl.lit(0).alias('ts').cast(pl.Int32),    \n    pl.lit(0).alias('type').cast(pl.Int8),\n    pl.lit(0).alias('real_action').cast(pl.Int64),\n    pl.col('session').cast(pl.Int32),\n    pl.col('aid').cast(pl.Int32),\n])\nn_col_after_join = clicks_candidate_df.groupby('session').agg([\n    pl.col('aid').cumcount().alias('CG_ranking')]).select(\n    pl.col('CG_ranking').explode().cast(pl.Int64))\n\ntest_df_click_recommended = pl.concat([clicks_candidate_df, n_col_after_join], how=\"horizontal\")\n","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:56:11.090546Z","iopub.execute_input":"2023-01-08T23:56:11.090893Z","iopub.status.idle":"2023-01-08T23:56:13.487606Z","shell.execute_reply.started":"2023-01-08T23:56:11.090861Z","shell.execute_reply":"2023-01-08T23:56:13.486203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Feature Calculation for inference data","metadata":{}},{"cell_type":"code","source":"%%time\n\n# Calculating the embedding of last three actions aid\ntest_df_click = test_df_click.join(embedding_weigh_dict_df, on = 'aid', how = 'left').join(embedding_weigh_dict_df, right_on = 'aid', left_on = 'aid_last', how = 'left', suffix = 'last_1').join(embedding_weigh_dict_df, right_on = 'aid', left_on = 'aid_second_last', how = 'left', suffix = 'last_2').join(embedding_weigh_dict_df, right_on = 'aid', left_on = 'aid_third_last', suffix = 'last_3').select(pl.exclude(['aid_last', 'aid_second_last', 'aid_third_last']))","metadata":{"execution":{"iopub.status.busy":"2023-01-09T00:01:55.951598Z","iopub.execute_input":"2023-01-09T00:01:55.952089Z","iopub.status.idle":"2023-01-09T00:01:57.925122Z","shell.execute_reply.started":"2023-01-09T00:01:55.952050Z","shell.execute_reply":"2023-01-09T00:01:57.923887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculating the embedding of last three actions aid for the recommended action. We all use the last three actual aid embedding for the last three aid\ntrain_df_click_last_action_embedding = test_df_click.sort(['session', 'ts']).groupby(['session']).last().select(pl.exclude(['aid', 'ts','type','real_action', 'CG_ranking', 'aid_last', 'aid_second_last', 'aid_third_last'] + ['Embedding_' + str(i) for i in range(32)]))\n\ntest_df_click_recommended = test_df_click_recommended.join(embedding_weigh_dict_df, on = 'aid', how = 'left').join(train_df_click_last_action_embedding, on = 'session', how = 'left')","metadata":{"execution":{"iopub.status.busy":"2023-01-09T00:02:01.318340Z","iopub.execute_input":"2023-01-09T00:02:01.319298Z","iopub.status.idle":"2023-01-09T00:02:06.329095Z","shell.execute_reply.started":"2023-01-09T00:02:01.319238Z","shell.execute_reply":"2023-01-09T00:02:06.328057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\n# Calculating the feature on the inference data\ntest_df_click = test_df_click.join(aid_global_counter_all_types, on='aid', suffix ='_global_counter').join(aid_global_user_counter_all_types, on='aid', suffix ='_user_counter').join(aid_global_user_counter_all_types_time_weighted, on='aid', suffix ='_timed_global_counter')\ntest_df_click_recommended = test_df_click_recommended.join(aid_global_counter_all_types, on='aid', suffix ='_global_counter').join(aid_global_user_counter_all_types, on='aid', suffix ='_user_counter').join(aid_global_user_counter_all_types_time_weighted, on='aid', suffix ='_timed_global_counter')\n\n# Combining the actual test data aid and recommended aid\ntest_df_click_all = pl.concat([test_df_click, test_df_click_recommended], how = 'vertical')","metadata":{"execution":{"iopub.status.busy":"2023-01-09T00:02:06.578069Z","iopub.execute_input":"2023-01-09T00:02:06.578524Z","iopub.status.idle":"2023-01-09T00:02:28.857497Z","shell.execute_reply.started":"2023-01-09T00:02:06.578488Z","shell.execute_reply":"2023-01-09T00:02:28.856427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Model inference\nscores = ranker.predict(test_df_click_all[feature_cols].to_pandas())\n\n# Appending the model score to the original dataframe\ntest_df_click_all = test_df_click_all.with_columns(pl.Series(name='score', values=scores))\n\n# Getting the top 20 candidates from the prediction\nclicks_pred_df = test_df_click_all.sort(['session', 'score'], reverse=True).groupby('session').agg([\n    pl.col('aid').limit(20).list().alias('labels')\n])\n\n# Converting to pandas format and making it align with result format\nclicks_pred_df = clicks_pred_df.with_columns(\npl.col('session') + '_clicks'\n).to_pandas()","metadata":{"execution":{"iopub.status.busy":"2023-01-09T00:02:28.859583Z","iopub.execute_input":"2023-01-09T00:02:28.860276Z","iopub.status.idle":"2023-01-09T00:02:36.436603Z","shell.execute_reply.started":"2023-01-09T00:02:28.860236Z","shell.execute_reply":"2023-01-09T00:02:36.434935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 2.4 Exporting to csv","metadata":{}},{"cell_type":"code","source":"# clicks_pred_df = pd.DataFrame(pred_df_clicks.add_suffix(\"_clicks\"), columns=[\"labels\"]).reset_index()\norders_pred_df = pd.DataFrame(pred_df_buys.add_suffix(\"_orders\"), columns=[\"labels\"]).reset_index()\ncarts_pred_df = pd.DataFrame(pred_df_carts.add_suffix(\"_carts\"), columns=[\"labels\"]).reset_index()","metadata":{"execution":{"iopub.status.busy":"2023-01-09T00:03:08.028483Z","iopub.execute_input":"2023-01-09T00:03:08.029228Z","iopub.status.idle":"2023-01-09T00:03:08.441973Z","shell.execute_reply.started":"2023-01-09T00:03:08.029189Z","shell.execute_reply":"2023-01-09T00:03:08.440684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_df = pd.concat([clicks_pred_df, orders_pred_df, carts_pred_df])\npred_df.columns = [\"session_type\", \"labels\"]\npred_df[\"labels\"] = pred_df.labels.apply(lambda x: \" \".join(map(str,x)))\npred_df.to_csv(\"validation_preds.csv\", index=False)\npred_df.head()","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[],"execution":{"iopub.status.busy":"2023-01-09T00:03:12.341828Z","iopub.execute_input":"2023-01-09T00:03:12.343176Z","iopub.status.idle":"2023-01-09T00:03:21.398720Z","shell.execute_reply.started":"2023-01-09T00:03:12.343126Z","shell.execute_reply":"2023-01-09T00:03:21.397474Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 3: Model Evaluation","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[]}},{"cell_type":"markdown","source":"This code is from Chris. It aims to calculate the local Recall score of the ranking. See the notebook link here: https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565","metadata":{}},{"cell_type":"code","source":"# # FREE MEMORY\n# del pred_df_clicks, pred_df_buys, clicks_pred_df, orders_pred_df, carts_pred_df\n# del top_20_clicks, top_20_buy2buy, top_20_buys, top_clicks, top_orders, test_df\n# _ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2023-01-08T23:56:14.125404Z","iopub.status.idle":"2023-01-08T23:56:14.126579Z","shell.execute_reply.started":"2023-01-08T23:56:14.126140Z","shell.execute_reply":"2023-01-08T23:56:14.126163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# COMPUTE METRIC\nscore = 0\nweights = {'clicks': 0.10, 'carts': 0.30, 'orders': 0.60}\nfor t in ['clicks','carts','orders']:\n    sub = pred_df.loc[pred_df.session_type.str.contains(t)].copy()\n    sub['session'] = sub.session_type.apply(lambda x: int(x.split('_')[0]))\n    sub.labels = sub.labels.apply(lambda x: [int(i) for i in x.split(' ')])\n    test_labels = pd.read_parquet('../input/otto-validation/test_labels.parquet')\n    test_labels = test_labels.loc[test_labels['type']==t]\n    test_labels = test_labels.merge(sub, how='left', on=['session'])\n    test_labels = test_labels.dropna()\n    test_labels['hits'] = test_labels.apply(lambda df: len(set(df.ground_truth).intersection(set(df.labels))), axis=1)\n    test_labels['gt_count'] = test_labels.ground_truth.str.len().clip(0,20)\n    recall = test_labels['hits'].sum() / test_labels['gt_count'].sum()\n    score += weights[t]*recall\n    print(f'{t} recall =',recall)\n    \nprint('=============')\nprint('Overall Recall =',score)\nprint('=============')","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[],"execution":{"iopub.status.busy":"2023-01-09T00:03:27.565402Z","iopub.execute_input":"2023-01-09T00:03:27.565828Z","iopub.status.idle":"2023-01-09T00:03:45.327141Z","shell.execute_reply.started":"2023-01-09T00:03:27.565786Z","shell.execute_reply":"2023-01-09T00:03:45.326082Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Previous performance of the ReRank Model\n\nclicks recall = 0.3896982547336502 <br>\ncarts recall = 0.4105610333097339<br>\norders recall = 0.6519093116387601 <br>\n=============<br>\nOverall Recall = 0.5532837224495413 <br>\n=============<br>\n\n\n**The performance drop for the click behavior, we may need to optimize the model or put more weight to the aid of actual behavior.**","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}