{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Hyper Parameter Tuning","metadata":{}},{"cell_type":"markdown","source":"## The ideas and research in this work are drawn from the works of\n* https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575 \n* https://www.kaggle.com/code/utm529fg/otto-tuning-candidate-rerank-model-lb-0-577","metadata":{}},{"cell_type":"code","source":"# Balance of type weighting タイプごとの重み付けバランス\n# 0:clicks 1:carts 2:orders\ntype_weight = {0:0.5,\n               1:9,\n               2:0.5}\ntype_weight_multipliers = type_weight\n\n# Use top X for clicks, carts and orders Top何位までを使うか\nclicks_th = 15 # クリック数\ncarts_th  = 20 # カート数\norders_th = 20 # 購入数\n\nVER = 5","metadata":{"execution":{"iopub.status.busy":"2023-01-06T10:17:59.293047Z","iopub.execute_input":"2023-01-06T10:17:59.294124Z","iopub.status.idle":"2023-01-06T10:17:59.299030Z","shell.execute_reply.started":"2023-01-06T10:17:59.294083Z","shell.execute_reply":"2023-01-06T10:17:59.298128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The following is an appropriation of CHRIS DEOTTE's notebook. Thank you!\n### https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\n---","metadata":{}},{"cell_type":"markdown","source":"# Candidate ReRank Model using Handcrafted Rules\nIn this notebook, we present a \"candidate rerank\" model using handcrafted rules. We can improve this model by engineering features, merging them unto items and users, and training a reranker model (such as XGB) to choose our final 20. Furthermore to tune and improve this notebook, we should build a local CV scheme to experiment new logic and/or models.\n\n UPDATE: I published a notebook to compute validation score [here][10] using Radek's scheme described [here][11].\n\nNote in this competition, a \"session\" actually means a unique \"user\". So our task is to predict what each of the `1,671,803` test \"users\" (i.e. \"sessions\") will do in the future. For each test \"user\" (i.e. \"session\") we must predict what they will `click`, `cart`, and `order` during the remainder of the week long test period.\n\n### Step 1 - Generate Candidates\nFor each test user, we generate possible choices, i.e. candidates. In this notebook, we generate candidates from 5 sources:\n* User history of clicks, carts, orders\n* Most popular 20 clicks, carts, orders during test week\n* Co-visitation matrix of click/cart/order to cart/order with type weighting\n* Co-visitation matrix of cart/order to cart/order called buy2buy\n* Co-visitation matrix of click/cart/order to clicks with time weighting\n\n### Step 2 - ReRank and Choose 20\nGiven the list of candidates, we must select 20 to be our predictions. In this notebook, we do this with a set of handcrafted rules. We can improve our predictions by training an XGBoost model to select for us. Our handcrafted rules give priority to:\n* Most recent previously visited items\n* Items previously visited multiple times\n* Items previously in cart or order\n* Co-visitation matrix of cart/order to cart/order\n* Current popular items\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/c_r_model.png)\n  \n# Credits\nWe thank many Kagglers who have shared ideas. We use co-visitation matrix idea from Vladimir [here][1]. We use groupby sort logic from Sinan in comment section [here][4]. We use duplicate prediction removal logic from Radek [here][5]. We use multiple visit logic from Pietro [here][2]. We use type weighting logic from Ingvaras [here][3]. We use leaky test data from my previous notebook [here][4]. And some ideas may have originated from Tawara [here][6] and KJ [here][7]. We use Colum2131's parquets [here][8]. Above image is from Ravi's discussion about candidate rerank models [here][9]\n\n[1]: https://www.kaggle.com/code/vslaykovsky/co-visitation-matrix\n[2]: https://www.kaggle.com/code/pietromaldini1/multiple-clicks-vs-latest-items\n[3]: https://www.kaggle.com/code/ingvarasgalinskas/item-type-vs-multiple-clicks-vs-latest-items\n[4]: https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost\n[5]: https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\n[6]: https://www.kaggle.com/code/ttahara/otto-mors-aid-frequency-baseline\n[7]: https://www.kaggle.com/code/whitelily/co-occurrence-baseline\n[8]: https://www.kaggle.com/datasets/columbia2131/otto-chunk-data-inparquet-format\n[9]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\n[10]: https://www.kaggle.com/cdeotte/compute-validation-score-cv-564\n[11]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991","metadata":{}},{"cell_type":"markdown","source":"# Step 1 - Candidate Generation with RAPIDS\nFor candidate generation, we build three co-visitation matrices. One computes the popularity of cart/order given a user's previous click/cart/order. We apply type weighting to this matrix. One computes the popularity of cart/order given a user's previous cart/order. We call this \"buy2buy\" matrix. One computes the popularity of clicks given a user previously click/cart/order.  We apply time weighting to this matrix. We will use RAPIDS cuDF GPU to compute these matrices quickly!","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"!pip install pandarallel\n!pip install pyarrow\n!pip install fastparquet\n!pip install ipywidgets","metadata":{"execution":{"iopub.status.busy":"2023-01-06T10:17:59.300597Z","iopub.execute_input":"2023-01-06T10:17:59.300917Z","iopub.status.idle":"2023-01-06T10:18:13.078824Z","shell.execute_reply.started":"2023-01-06T10:17:59.300889Z","shell.execute_reply":"2023-01-06T10:18:13.077692Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd, numpy as np\nfrom tqdm.notebook import tqdm\nimport os, sys, pickle, glob, gc\nfrom collections import Counter\nfrom pandarallel import pandarallel\nimport itertools\npandarallel.initialize(progress_bar=False)\npandarallel.initialize(use_memory_fs=False)","metadata":{"execution":{"iopub.status.busy":"2023-01-06T10:58:55.864144Z","iopub.execute_input":"2023-01-06T10:58:55.865147Z","iopub.status.idle":"2023-01-06T10:58:55.871659Z","shell.execute_reply.started":"2023-01-06T10:58:55.865108Z","shell.execute_reply":"2023-01-06T10:58:55.870915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 2 - ReRank (choose 20) using handcrafted rules\nFor description of the handcrafted rules, read this notebook's intro.","metadata":{"jupyter":{"source_hidden":true}}},{"cell_type":"code","source":"type_labels={'clicks':0,'carts':1,'orders':2}\ndef load_test():    \n    dfs = []\n    for e, chunk_file in enumerate(glob.glob('../input/otto-chunk-data-inparquet-format/test_parquet/*')):\n        chunk = pd.read_parquet(chunk_file)\n        chunk.ts = (chunk.ts/1000).astype('int32')\n        chunk['type'] = chunk['type'].map(type_labels).astype('int8')\n        dfs.append(chunk)\n    return pd.concat(dfs).reset_index(drop=True) #.astype({\"ts\": \"datetime64[ms]\"})\n\ntest_df = load_test()\nprint('Test data has shape',test_df.shape)\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-06T10:18:13.603641Z","iopub.execute_input":"2023-01-06T10:18:13.604081Z","iopub.status.idle":"2023-01-06T10:18:14.760273Z","shell.execute_reply.started":"2023-01-06T10:18:13.604051Z","shell.execute_reply":"2023-01-06T10:18:14.759444Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nDISK_PIECES=4\npath='/kaggle/input/rankmode-and-covisitations-dataset/'\ndef pqt_to_dict(df):\n    return df.groupby('aid_x').aid_y.apply(list).to_dict()\n\n# LOAD THREE CO-VISITATION MATRICES\ntop_20_clicks = pqt_to_dict( pd.read_parquet(path+f'top_20_clicks_v{VER}_0.pqt') )\n\nfor k in range(1,DISK_PIECES): \n    top_20_clicks.update( pqt_to_dict( pd.read_parquet(path+f'top_20_clicks_v{VER}_{k}.pqt') ) )\n\n\ntop_20_buys = pqt_to_dict( pd.read_parquet(path+f'top_15_carts_orders_v{VER}_0.pqt') )\n\nfor k in range(1,DISK_PIECES): \n    top_20_buys.update( pqt_to_dict( pd.read_parquet(path+f'top_15_carts_orders_v{VER}_{k}.pqt') ) )\n\ntop_20_buy2buy = pqt_to_dict( pd.read_parquet(path+f'top_15_buy2buy_v{VER}_0.pqt') )\n\n# TOP CLICKS AND ORDERS IN TEST\n#top_clicks = test_df.loc[test_df['type']=='clicks','aid'].value_counts().index.values[:20]\n#top_orders = test_df.loc[test_df['type']=='orders','aid'].value_counts().index.values[:20]\n\nprint('Here are size of our 3 co-visitation matrices:')\nprint( len( top_20_clicks ), len( top_20_buy2buy ), len( top_20_buys ) )","metadata":{"execution":{"iopub.status.busy":"2023-01-06T10:18:14.762680Z","iopub.execute_input":"2023-01-06T10:18:14.763377Z","iopub.status.idle":"2023-01-06T10:21:22.295502Z","shell.execute_reply.started":"2023-01-06T10:18:14.763345Z","shell.execute_reply":"2023-01-06T10:21:22.294051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"top_clicks = test_df.loc[test_df['type']== 0,'aid'].value_counts().index.values[:20] \ntop_carts = test_df.loc[test_df['type']== 1,'aid'].value_counts().index.values[:20]\ntop_orders = test_df.loc[test_df['type']== 2,'aid'].value_counts().index.values[:20]","metadata":{"execution":{"iopub.status.busy":"2023-01-06T10:21:22.297037Z","iopub.execute_input":"2023-01-06T10:21:22.297455Z","iopub.status.idle":"2023-01-06T10:21:22.798682Z","shell.execute_reply.started":"2023-01-06T10:21:22.297414Z","shell.execute_reply":"2023-01-06T10:21:22.797629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def suggest_clicks(df):\n    # USER HISTORY AIDS AND TYPES\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    # RERANK CANDIDATES USING WEIGHTS\n    if len(unique_aids)>=20:\n        weights=np.logspace(0.1,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        # RERANK BASED ON REPEAT ITEMS AND TYPE OF ITEMS\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    # USE \"CLICKS\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(20) if aid2 not in unique_aids]    \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    # USE TOP20 TEST CLICKS\n    return result + list(top_clicks)[:20-len(result)]","metadata":{"execution":{"iopub.status.busy":"2023-01-06T11:04:11.047527Z","iopub.execute_input":"2023-01-06T11:04:11.047964Z","iopub.status.idle":"2023-01-06T11:04:11.059542Z","shell.execute_reply.started":"2023-01-06T11:04:11.047928Z","shell.execute_reply":"2023-01-06T11:04:11.058541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def suggest_carts(df):\n    # User history aids and types\n    aids = df.aid.tolist()\n    types = df.type.tolist()\n    \n    # UNIQUE AIDS AND UNIQUE BUYS\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    df = df.loc[(df['type'] == 0)|(df['type'] == 1)]\n    unique_buys = list(dict.fromkeys(df.aid.tolist()[::-1]))\n    \n    # Rerank candidates using weights\n    if len(unique_aids) >= 20:\n        weights=np.logspace(0.5,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        \n        # Rerank based on repeat items and types of items\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        \n        # Rerank candidates using\"top_20_carts\" co-visitation matrix\n        aids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_buys if aid in top_20_buys]))\n        for aid in aids2: aids_temp[aid] += 0.1\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    \n    # Use \"cart order\" and \"clicks\" co-visitation matrices\n    aids1 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    aids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_aids if aid in top_20_buys]))\n    \n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids1+aids2).most_common(20) if aid2 not in unique_aids] \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    \n    # USE TOP20 TEST ORDERS\n    return result + list(top_carts)[:20-len(result)]","metadata":{"execution":{"iopub.status.busy":"2023-01-06T11:04:11.300527Z","iopub.execute_input":"2023-01-06T11:04:11.301247Z","iopub.status.idle":"2023-01-06T11:04:11.314041Z","shell.execute_reply.started":"2023-01-06T11:04:11.301206Z","shell.execute_reply":"2023-01-06T11:04:11.312753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def suggest_buys(df):\n    # USER HISTORY AIDS AND TYPES\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    # UNIQUE AIDS AND UNIQUE BUYS\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    df = df.loc[(df['type']==1)|(df['type']==2)]\n    unique_buys = list(dict.fromkeys( df.aid.tolist()[::-1] ))\n    # RERANK CANDIDATES USING WEIGHTS\n    if len(unique_aids)>=20:\n        weights=np.logspace(0.5,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        # RERANK BASED ON REPEAT ITEMS AND TYPE OF ITEMS\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        # RERANK CANDIDATES USING \"BUY2BUY\" CO-VISITATION MATRIX\n        aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n        for aid in aids3: aids_temp[aid] += 0.1\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    # USE \"CART ORDER\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_aids if aid in top_20_buys]))\n    # USE \"BUY2BUY\" CO-VISITATION MATRIX\n    aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2+aids3).most_common(20) if aid2 not in unique_aids] \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    # USE TOP20 TEST ORDERS\n    return result + list(top_orders)[:20-len(result)]","metadata":{"execution":{"iopub.status.busy":"2023-01-06T11:04:11.538920Z","iopub.execute_input":"2023-01-06T11:04:11.539955Z","iopub.status.idle":"2023-01-06T11:04:11.552486Z","shell.execute_reply.started":"2023-01-06T11:04:11.539896Z","shell.execute_reply":"2023-01-06T11:04:11.551481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create Submission CSV\nInferring test data with Pandas groupby is slow. We need to accelerate the following code.","metadata":{}},{"cell_type":"code","source":"%%time\n\npred_df_clicks = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).parallel_apply(\n    lambda x: suggest_clicks(x)\n)\n\npred_df_carts = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).parallel_apply(\n    lambda x: suggest_carts(x)\n)\n\npred_df_buys = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).parallel_apply(\n    lambda x: suggest_carts(x)\n)","metadata":{"execution":{"iopub.status.busy":"2023-01-06T11:04:12.163734Z","iopub.execute_input":"2023-01-06T11:04:12.164949Z","iopub.status.idle":"2023-01-06T11:16:13.272299Z","shell.execute_reply.started":"2023-01-06T11:04:12.164907Z","shell.execute_reply":"2023-01-06T11:16:13.271178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clicks_pred_df = pd.DataFrame(pred_df_clicks.add_suffix(\"_clicks\"), columns=[\"labels\"]).reset_index()\norders_pred_df = pd.DataFrame(pred_df_buys.add_suffix(\"_orders\"), columns=[\"labels\"]).reset_index()\ncarts_pred_df = pd.DataFrame(pred_df_carts.add_suffix(\"_carts\"), columns=[\"labels\"]).reset_index()","metadata":{"execution":{"iopub.status.busy":"2023-01-06T11:17:56.148275Z","iopub.execute_input":"2023-01-06T11:17:56.149257Z","iopub.status.idle":"2023-01-06T11:18:01.085757Z","shell.execute_reply.started":"2023-01-06T11:17:56.149217Z","shell.execute_reply":"2023-01-06T11:18:01.084695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_df = pd.concat([clicks_pred_df, orders_pred_df, carts_pred_df])\npred_df.columns = [\"session_type\", \"labels\"]\npred_df[\"labels\"] = pred_df.labels.apply(lambda x: \" \".join(map(str,x)))\npred_df.to_csv(\"submission.csv\", index=False)\npred_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-06T11:18:01.087579Z","iopub.execute_input":"2023-01-06T11:18:01.087942Z","iopub.status.idle":"2023-01-06T11:18:53.004844Z","shell.execute_reply.started":"2023-01-06T11:18:01.087911Z","shell.execute_reply":"2023-01-06T11:18:53.003884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}