{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":38760,"databundleVersionId":4493939,"sourceType":"competition"},{"sourceId":4436180,"sourceType":"datasetVersion","datasetId":2597726}],"dockerImageVersionId":30302,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style=\"font-family:verdana;\"> <center>📝OTTO: co-visitation matrix functionalization📝</center></h1>","metadata":{}},{"cell_type":"markdown","source":"# References\n**Great notebooks! Thank you🙇**\n- https://www.kaggle.com/code/tetsuro731/duplicate-fix-otto-tuning-pipeline2-lb-0-577\n- https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\n- https://www.kaggle.com/code/adaubas/otto-fast-handcrafted-model-recall-20\n- https://www.kaggle.com/code/vslaykovsky/co-visitation-matrix","metadata":{}},{"cell_type":"markdown","source":"# In this notebook\n- Easy-to-use co-visitation matrix fuctionalization\n- Enable to create co-visitation matrix by common function `co_visitation_matrix`\n- Augments as follows\n    - `attention_period`: period threshold for co-visitation matrix\n    - `drop_th_sess_num`: cut session with small samples\n    - `attention_type`: matrix axis ex. carts vs order\n    - `weights_func`: variable weight function\n    - `save_top_k`: save top k samples of matrix","metadata":{}},{"cell_type":"markdown","source":"# 📌Hyper Parameters","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:14px; font-family:verdana;\">\n    📌For hyper parameters, read following notebook: <a href=\"https://www.kaggle.com/code/tetsuro731/duplicate-fix-otto-tuning-pipeline2-lb-0-577#Step-1---Candidate-Generation-with-RAPIDS\" target=\"_blank\" style=\"color:#ff7f7f;\">[duplicate fix]OTTO: Tuning-pipeline2 [LB 0.577] by tetsuro731</a>\n</div>","metadata":{}},{"cell_type":"code","source":"# Balance of type weighting\n# 0:clicks 1:carts 2:orders\ntype_weight = {0:0.5,\n               1:9,\n               2:0.5}\n\n# Use top X for clicks, carts and orders\nclicks_th = 15\ncarts_th  = 20\norders_th = 20","metadata":{"execution":{"iopub.status.busy":"2026-01-05T15:57:49.307939Z","iopub.execute_input":"2026-01-05T15:57:49.308320Z","iopub.status.idle":"2026-01-05T15:57:49.336607Z","shell.execute_reply.started":"2026-01-05T15:57:49.308230Z","shell.execute_reply":"2026-01-05T15:57:49.335745Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 📌Candidate ReRank Model using Handcrafted Rules\n<div class=\"alert alert-block alert-info\" style=\"font-size:14px; font-family:verdana;\">\n    📌For following commentary, read following notebook: <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\" style=\"color:#ff7f7f;\">Candidate ReRank Model - [LB 0.575] by cdeotte</a>\n</div>\nIn this notebook, we present a \"candidate rerank\" model using handcrafted rules. We can improve this model by engineering features, merging them unto items and users, and training a reranker model (such as XGB) to choose our final 20. Furthermore to tune and improve this notebook, we should build a local CV scheme to experiment new logic and/or models.\n\nUPDATE: I published a notebook to compute validation score [here][10] using Radek's scheme described [here][11].\n\nNote in this competition, a \"session\" actually means a unique \"user\". So our task is to predict what each of the `1,671,803` test \"users\" (i.e. \"sessions\") will do in the future. For each test \"user\" (i.e. \"session\") we must predict what they will `click`, `cart`, and `order` during the remainder of the week long test period.\n\n### Step 1 - Generate Candidates\nFor each test user, we generate possible choices, i.e. candidates. In this notebook, we generate candidates from 5 sources:\n* User history of clicks, carts, orders\n* Most popular 20 clicks, carts, orders during test week\n* Co-visitation matrix of click/cart/order to cart/order with type weighting\n* Co-visitation matrix of cart/order to cart/order called buy2buy\n* Co-visitation matrix of click/cart/order to clicks with time weighting\n\n### Step 2 - ReRank and Choose 20\nGiven the list of candidates, we must select 20 to be our predictions. In this notebook, we do this with a set of handcrafted rules. We can improve our predictions by training an XGBoost model to select for us. Our handcrafted rules give priority to:\n* Most recent previously visited items\n* Items previously visited multiple times\n* Items previously in cart or order\n* Co-visitation matrix of cart/order to cart/order\n* Current popular items\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/c_r_model.png)\n  \n# Credits\nWe thank many Kagglers who have shared ideas. We use co-visitation matrix idea from Vladimir [here][1]. We use groupby sort logic from Sinan in comment section [here][4]. We use duplicate prediction removal logic from Radek [here][5]. We use multiple visit logic from Pietro [here][2]. We use type weighting logic from Ingvaras [here][3]. We use leaky test data from my previous notebook [here][4]. And some ideas may have originated from Tawara [here][6] and KJ [here][7]. We use Colum2131's parquets [here][8]. Above image is from Ravi's discussion about candidate rerank models [here][9]\n\n[1]: https://www.kaggle.com/code/vslaykovsky/co-visitation-matrix\n[2]: https://www.kaggle.com/code/pietromaldini1/multiple-clicks-vs-latest-items\n[3]: https://www.kaggle.com/code/ingvarasgalinskas/item-type-vs-multiple-clicks-vs-latest-items\n[4]: https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost\n[5]: https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\n[6]: https://www.kaggle.com/code/ttahara/otto-mors-aid-frequency-baseline\n[7]: https://www.kaggle.com/code/whitelily/co-occurrence-baseline\n[8]: https://www.kaggle.com/datasets/columbia2131/otto-chunk-data-inparquet-format\n[9]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\n[10]: https://www.kaggle.com/cdeotte/compute-validation-score-cv-564\n[11]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991","metadata":{}},{"cell_type":"markdown","source":"# 📌Step 1 - Candidate Generation with RAPIDS\nFor candidate generation, we build three co-visitation matrices. One computes the popularity of cart/order given a user's previous click/cart/order. We apply type weighting to this matrix. One computes the popularity of cart/order given a user's previous cart/order. We call this \"buy2buy\" matrix. One computes the popularity of clicks given a user previously click/cart/order.  We apply time weighting to this matrix. We will use RAPIDS cuDF GPU to compute these matrices quickly!","metadata":{}},{"cell_type":"markdown","source":"# Import libraries","metadata":{}},{"cell_type":"code","source":"import pandas as pd, numpy as np\nfrom tqdm.notebook import tqdm\nimport os, sys, pickle, glob, gc\nfrom collections import Counter\nimport cudf, itertools\nprint('We will use RAPIDS version',cudf.__version__)","metadata":{"execution":{"iopub.status.busy":"2026-01-05T15:57:49.338923Z","iopub.execute_input":"2026-01-05T15:57:49.339243Z","iopub.status.idle":"2026-01-05T15:57:52.036088Z","shell.execute_reply.started":"2026-01-05T15:57:49.339209Z","shell.execute_reply":"2026-01-05T15:57:52.035189Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📌Compute Three Co-visitation Matrices with RAPIDS\nWe will compute 3 co-visitation matrices using RAPIDS cuDF on GPU. This is 30x faster than using Pandas CPU like other public notebooks! For maximum speed, set the variable `DISK_PIECES` to the smallest number possible based on the GPU you are using without incurring memory errors. If you run this code offline with 32GB GPU ram, then you can use `DISK_PIECES = 1` and compute each co-visitation matrix in almost 1 minute! Kaggle's GPU only has 16GB ram, so we use `DISK_PIECES = 4` and it takes an amazing 3 minutes each! Below are some of the tricks to speed up computation\n* Use RAPIDS cuDF GPU instead of Pandas CPU\n* Read disk once and save in CPU RAM for later GPU multiple use\n* Process largest amount of data possible on GPU at one time\n* Merge data in two stages. Multiple small to single medium. Multiple medium to single large.\n* Write result as parquet instead of dictionary","metadata":{}},{"cell_type":"code","source":"%%time\n# CACHE FUNCTIONS\ndef read_file(f):\n    return cudf.DataFrame( data_cache[f] )\n\ndef read_file_to_cache(f):\n    df = pd.read_parquet(f)\n    df.ts = (df.ts/1000).astype('int32')\n    df['type'] = df['type'].map(type_labels).astype('int8')\n    return df\n\n# CACHE THE DATA ON CPU BEFORE PROCESSING ON GPU\ndata_cache = {}\ntype_labels = {'clicks':0, 'carts':1, 'orders':2}\nfiles = glob.glob('../input/otto-chunk-data-inparquet-format/*_parquet/*')\nfor f in files: data_cache[f] = read_file_to_cache(f)\n\n# CHUNK PARAMETERS\nREAD_CT = 5\nOUTER_CHUNK_NUM = 6\nCHUNK_NUM = int( np.ceil( len(files) / OUTER_CHUNK_NUM ))\nprint(f'We will process {len(files)} files, in groups of {READ_CT} and chunks of {CHUNK_NUM}.')","metadata":{"execution":{"iopub.status.busy":"2026-01-05T15:57:52.037144Z","iopub.execute_input":"2026-01-05T15:57:52.037423Z","iopub.status.idle":"2026-01-05T15:58:35.667573Z","shell.execute_reply.started":"2026-01-05T15:57:52.037398Z","shell.execute_reply":"2026-01-05T15:58:35.666669Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 📔Matrix weight function","metadata":{}},{"cell_type":"code","source":"def weights_func_carts_order(df):\n    df['wgt'] = df.type_y.map(type_weight)\n    return df\n\ndef weights_func_orders(df):\n    df['wgt'] = 1\n    return df\n\ndef weights_func_clicks(df):\n    df['wgt'] = 1 + 3*(df.ts_x - 1659304800)/(1662328791-1659304800)\n    # 1659304800 : minimum timestamp\n    # 1662328791 : maximum timestamp\n    return df","metadata":{"execution":{"iopub.status.busy":"2026-01-05T15:58:35.668936Z","iopub.execute_input":"2026-01-05T15:58:35.669209Z","iopub.status.idle":"2026-01-05T15:58:35.674757Z","shell.execute_reply.started":"2026-01-05T15:58:35.669184Z","shell.execute_reply":"2026-01-05T15:58:35.673761Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def drop_dupli_cols_for_clicks(df):\n    df = df[['session', 'aid_x', 'aid_y','ts_x']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n    return df\n\ndef drop_dupli_cols_for_carts_orders(df):\n    df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y','type_y'])\n    return df","metadata":{"execution":{"iopub.status.busy":"2026-01-05T15:58:35.675676Z","iopub.execute_input":"2026-01-05T15:58:35.675946Z","iopub.status.idle":"2026-01-05T15:58:35.685991Z","shell.execute_reply.started":"2026-01-05T15:58:35.675923Z","shell.execute_reply":"2026-01-05T15:58:35.685230Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 📔Co-visitation matrix function","metadata":{}},{"cell_type":"code","source":"def co_visitation_matrix(\n    files: list,\n    read_part_id: int,\n    read_size: float,\n    weights_func,\n    drop_dupli_func,\n    outer_chunk_num=6,\n    attention_period=24*60*60,\n    drop_th_sess_num=30,\n    attention_type=None,\n    save_top_k=20\n):\n    # MERGE IS FASTEST PROCESSING CHUNKS WITHIN CHUNKS\n    # => OUTER CHUNKS\n    for outer_chk_id in range(outer_chunk_num):\n        start_chk_id = outer_chk_id * CHUNK_NUM\n        end_chk_id = min( (outer_chk_id+1)*CHUNK_NUM, len(files) )\n        \n        for inner_chk_id in range(start_chk_id, end_chk_id, READ_CT):\n            df = [read_file(files[inner_chk_id])]\n            # read inner chunk files\n            for i in range(1, READ_CT): \n                if (inner_chk_id+i) < end_chk_id:\n                    df.append( read_file(files[(inner_chk_id+i)]) )\n\n            df = cudf.concat(df,ignore_index=True,axis=0)\n            # ** SPECIFY BY ARG **\n            if attention_type is not None:\n                df = df.loc[df['type'].isin(attention_type)]\n            df = df.sort_values(['session','ts'],ascending=[True,False])\n\n            # USE TAIL OF SESSION\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            # ** SPECIFY BY ARG **\n            df = df.loc[df.n < drop_th_sess_num].drop('n',axis=1)\n\n            # CREATE PAIRS\n            df = df.merge(df,on='session')\n            # ** SPECIFY BY ARG **\n            df = df.loc[ ((df.ts_x - df.ts_y).abs()< attention_period) & (df.aid_x != df.aid_y) ]\n\n            # MEMORY MANAGEMENT COMPUTE IN PARTS\n            # ** SPECIFY BY ARG **\n            df = df.loc[(df.aid_x >= read_part_id*read_size)&(df.aid_x < (read_part_id+1)*read_size)]\n\n            # ASSIGN WEIGHTS\n            # ** SPECIFY BY ARG **\n            df = drop_dupli_func(df)\n            df = weights_func(df)\n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n\n            # COMBINE INNER CHUNKS\n            if inner_chk_id == start_chk_id:\n                df_tmp_inner_chk = df\n            else:\n                df_tmp_inner_chk = df_tmp_inner_chk.add(df, fill_value=0)\n            print(inner_chk_id,', ',end='')\n\n        print()\n\n        # COMBINE OUTER CHUNKS\n        if start_chk_id == 0:\n            co_vis_matrix = df_tmp_inner_chk\n        else:\n            co_vis_matrix = co_vis_matrix.add(df_tmp_inner_chk, fill_value=0)\n\n        del df_tmp_inner_chk, df\n        gc.collect()\n        \n    # CONVERT MATRIX TO DICTIONARY\n    co_vis_matrix = co_vis_matrix.reset_index()\n    co_vis_matrix = co_vis_matrix.sort_values(['aid_x','wgt'],ascending=[True,False])\n    \n    # SAVE TOP K\n    co_vis_matrix = co_vis_matrix.reset_index(drop=True)\n    co_vis_matrix['n'] = co_vis_matrix.groupby('aid_x').aid_y.cumcount()\n    co_vis_matrix = co_vis_matrix.loc[co_vis_matrix.n<save_top_k].drop('n',axis=1)\n\n    return co_vis_matrix","metadata":{"execution":{"iopub.status.busy":"2026-01-05T15:58:35.687201Z","iopub.execute_input":"2026-01-05T15:58:35.687449Z","iopub.status.idle":"2026-01-05T15:58:35.701402Z","shell.execute_reply.started":"2026-01-05T15:58:35.687427Z","shell.execute_reply":"2026-01-05T15:58:35.700757Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 1) \"Carts Orders\" Co-visitation Matrix - Type Weighted","metadata":{}},{"cell_type":"code","source":"%%time\n# USE SMALLEST DISK_PIECES POSSIBLE WITHOUT MEMORY ERROR\nDISK_PIECES = 4\nSIZE = 1.86e6/DISK_PIECES\n\n# COMPUTE IN PARTS FOR MEMORY MANGEMENT\nfor PART in range(DISK_PIECES):\n    print()\n    print('### DISK PART',PART+1)\n    \n    co_vis_matrix = co_visitation_matrix(\n        files,\n        PART, # read settings\n        SIZE, # read settings\n        weights_func_carts_order,\n        drop_dupli_cols_for_carts_orders,\n        outer_chunk_num=6,\n        attention_period=24*60*60,\n        drop_th_sess_num=30,\n        attention_type=None,\n        save_top_k=carts_th\n    )\n\n    # SAVE PART TO DISK (convert to pandas first uses less memory)\n    co_vis_matrix.to_pandas().to_parquet(f'top_{carts_th}_carts_orders_{PART}.pqt')","metadata":{"execution":{"iopub.status.busy":"2026-01-05T15:58:35.703752Z","iopub.execute_input":"2026-01-05T15:58:35.703978Z","iopub.status.idle":"2026-01-05T16:01:51.523295Z","shell.execute_reply.started":"2026-01-05T15:58:35.703958Z","shell.execute_reply":"2026-01-05T16:01:51.522323Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2) \"Buy2Buy\" Co-visitation Matrix","metadata":{}},{"cell_type":"code","source":"%%time\n# USE SMALLEST DISK_PIECES POSSIBLE WITHOUT MEMORY ERROR\nDISK_PIECES = 1\nSIZE = 1.86e6/DISK_PIECES\n\n# COMPUTE IN PARTS FOR MEMORY MANGEMENT\nfor PART in range(DISK_PIECES):\n    print()\n    print('### DISK PART',PART+1)\n    \n    co_vis_matrix = co_visitation_matrix(\n        files,\n        PART, # read settings\n        SIZE, # read settings\n        weights_func_orders,\n        drop_dupli_cols_for_carts_orders,\n        outer_chunk_num=6,\n        attention_period=7*24*60*60,\n        drop_th_sess_num=30,\n        attention_type=[1, 2],# ONLY WANT CARTS AND ORDER\n        save_top_k=orders_th\n    )\n\n    # SAVE PART TO DISK (convert to pandas first uses less memory)\n    co_vis_matrix.to_pandas().to_parquet(f'top_{orders_th}_buy2buy_{PART}.pqt')\n    # SAVE PART TO DISK (convert to pandas first uses less memory)\n    co_vis_matrix.to_pandas().to_parquet(f'top_{orders_th}_buy2buy_{PART}.pqt')","metadata":{"execution":{"iopub.status.busy":"2026-01-05T16:01:51.524802Z","iopub.execute_input":"2026-01-05T16:01:51.525174Z","iopub.status.idle":"2026-01-05T16:02:20.693620Z","shell.execute_reply.started":"2026-01-05T16:01:51.525138Z","shell.execute_reply":"2026-01-05T16:02:20.692621Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3) \"Clicks\" Co-visitation Matrix - Time Weighted","metadata":{}},{"cell_type":"code","source":"%%time\n# USE SMALLEST DISK_PIECES POSSIBLE WITHOUT MEMORY ERROR\nDISK_PIECES = 4\nSIZE = 1.86e6/DISK_PIECES\n\n# COMPUTE IN PARTS FOR MEMORY MANGEMENT\nfor PART in range(DISK_PIECES):\n    print()\n    print('### DISK PART',PART+1)\n\n    co_vis_matrix = co_visitation_matrix(\n        files,\n        PART, # read settings\n        SIZE, # read settings\n        weights_func_clicks,\n        drop_dupli_cols_for_clicks,\n        outer_chunk_num=6,\n        attention_period=24*60*60,\n        drop_th_sess_num=30,\n        attention_type=None,\n        save_top_k=clicks_th\n    )\n\n    # SAVE PART TO DISK (convert to pandas first uses less memory)\n    co_vis_matrix.to_pandas().to_parquet(f'top_{clicks_th}_clicks_{PART}.pqt')","metadata":{"execution":{"iopub.status.busy":"2026-01-05T16:02:20.694986Z","iopub.execute_input":"2026-01-05T16:02:20.695611Z","iopub.status.idle":"2026-01-05T16:05:29.883207Z","shell.execute_reply.started":"2026-01-05T16:02:20.695569Z","shell.execute_reply":"2026-01-05T16:05:29.882275Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# FREE MEMORY\ndel data_cache, co_vis_matrix\n_ = gc.collect()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-01-05T16:05:29.884471Z","iopub.execute_input":"2026-01-05T16:05:29.884858Z","iopub.status.idle":"2026-01-05T16:05:30.041202Z","shell.execute_reply.started":"2026-01-05T16:05:29.884827Z","shell.execute_reply":"2026-01-05T16:05:30.040136Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 📌ReRank (choose 20) using handcrafted rules\n<div class=\"alert alert-block alert-info\" style=\"font-size:14px; font-family:verdana;\">\n    📌For description of the handcrafted rules, read following notebook: <a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\" style=\"color:#ff7f7f;\">Candidate ReRank Model - [LB 0.575] by cdeotte</a>\n</div>","metadata":{}},{"cell_type":"code","source":"def load_test():    \n    dfs = []\n    for e, chunk_file in enumerate(glob.glob('../input/otto-chunk-data-inparquet-format/test_parquet/*')):\n        chunk = pd.read_parquet(chunk_file)\n        chunk.ts = (chunk.ts/1000).astype('int32')\n        chunk['type'] = chunk['type'].map(type_labels).astype('int8')\n        dfs.append(chunk)\n    return pd.concat(dfs).reset_index(drop=True) #.astype({\"ts\": \"datetime64[ms]\"})\n\ntest_df = load_test()\nprint('Test data has shape',test_df.shape)\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2026-01-05T16:05:30.042452Z","iopub.execute_input":"2026-01-05T16:05:30.042745Z","iopub.status.idle":"2026-01-05T16:05:31.154690Z","shell.execute_reply.started":"2026-01-05T16:05:30.042693Z","shell.execute_reply":"2026-01-05T16:05:31.153424Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time\ndef pqt_to_dict(df):\n    return df.groupby('aid_x').aid_y.apply(list).to_dict()\n\n# LOAD THREE CO-VISITATION MATRICES\ntop_k_clicks = pqt_to_dict( pd.read_parquet(f'top_{clicks_th}_clicks_0.pqt') )\n\nfor k in range(1,DISK_PIECES): \n    top_k_clicks.update( pqt_to_dict( pd.read_parquet(f'top_{clicks_th}_clicks_{k}.pqt') ) )\n\n\ntop_k_buys = pqt_to_dict( pd.read_parquet(f'top_{carts_th}_carts_orders_0.pqt') )\n\nfor k in range(1,DISK_PIECES): \n    top_k_buys.update( pqt_to_dict( pd.read_parquet(f'top_{carts_th}_carts_orders_{k}.pqt') ) )\n\ntop_k_buy2buy = pqt_to_dict( pd.read_parquet(f'top_{orders_th}_buy2buy_0.pqt') )\n\n# TOP CLICKS AND ORDERS IN TEST\n#top_clicks = test_df.loc[test_df['type']=='clicks','aid'].value_counts().index.values[:20]\n#top_orders = test_df.loc[test_df['type']=='orders','aid'].value_counts().index.values[:20]\n\nprint('Here are size of our 3 co-visitation matrices:')\nprint( len( top_k_clicks ), len( top_k_buy2buy ), len( top_k_buys ) )","metadata":{"execution":{"iopub.status.busy":"2026-01-05T16:05:31.155984Z","iopub.execute_input":"2026-01-05T16:05:31.156429Z","iopub.status.idle":"2026-01-05T16:06:50.412157Z","shell.execute_reply.started":"2026-01-05T16:05:31.156393Z","shell.execute_reply":"2026-01-05T16:06:50.411235Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"top_clicks = test_df.loc[test_df['type']== 0,'aid'].value_counts().index.values[:20] \ntop_carts = test_df.loc[test_df['type']== 1,'aid'].value_counts().index.values[:20]\ntop_orders = test_df.loc[test_df['type']== 2,'aid'].value_counts().index.values[:20]","metadata":{"execution":{"iopub.status.busy":"2026-01-05T16:06:50.413245Z","iopub.execute_input":"2026-01-05T16:06:50.413499Z","iopub.status.idle":"2026-01-05T16:06:50.776934Z","shell.execute_reply.started":"2026-01-05T16:06:50.413463Z","shell.execute_reply":"2026-01-05T16:06:50.775877Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"small_top_clicks = top_clicks[:20]\ndef suggest_clicks(df):\n    # USER HISTORY AIDS AND TYPES\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    # RERANK CANDIDATES USING WEIGHTS\n    if len(unique_aids)>=20:\n        weights=np.logspace(0.1,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        # RERANK BASED ON REPEAT ITEMS AND TYPE OF ITEMS\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight[t]\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    # USE \"CLICKS\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_k_clicks[aid] for aid in unique_aids if aid in top_k_clicks]))\n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(20) if aid2 not in unique_aids]    \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    # USE TOP20 TEST CLICKS\n    #return result + list(top_clicks)[:20-len(result)]\n    # FIXED BY tetsuro731, remove duplicate\n    set_result = set(result)\n    return result + [i for i in small_top_clicks if i not in set_result][:20 - len(result)]","metadata":{"execution":{"iopub.status.busy":"2026-01-05T16:06:50.778173Z","iopub.execute_input":"2026-01-05T16:06:50.778446Z","iopub.status.idle":"2026-01-05T16:06:50.820615Z","shell.execute_reply.started":"2026-01-05T16:06:50.778421Z","shell.execute_reply":"2026-01-05T16:06:50.819759Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"small_top_carts = top_carts[:20]\ndef suggest_carts(df):\n    # User history aids and types\n    aids = df.aid.tolist()\n    types = df.type.tolist()\n    \n    # UNIQUE AIDS AND UNIQUE BUYS\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    df = df.loc[(df['type'] == 0)|(df['type'] == 1)]\n    unique_buys = list(dict.fromkeys(df.aid.tolist()[::-1]))\n    \n    # Rerank candidates using weights\n    if len(unique_aids) >= 20:\n        weights=np.logspace(0.5,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        \n        # Rerank based on repeat items and types of items\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight[t]\n        \n        # Rerank candidates using\"top_20_carts\" co-visitation matrix\n        aids2 = list(itertools.chain(*[top_k_buys[aid] for aid in unique_buys if aid in top_k_buys]))\n        for aid in aids2: aids_temp[aid] += 0.1\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    \n    # Use \"cart order\" and \"clicks\" co-visitation matrices\n    aids1 = list(itertools.chain(*[top_k_clicks[aid] for aid in unique_aids if aid in top_k_clicks]))\n    aids2 = list(itertools.chain(*[top_k_buys[aid] for aid in unique_aids if aid in top_k_buys]))\n    \n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids1+aids2).most_common(20) if aid2 not in unique_aids] \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    \n    # USE TOP20 TEST ORDERS\n    #return result + list(top_carts)[:20-len(result)]\n    # FIXED BY tetsuro731, remove duplicate\n    set_result = set(result)\n    return result + [i for i in small_top_carts if i not in set_result][:20 - len(result)]","metadata":{"execution":{"iopub.status.busy":"2026-01-05T16:06:50.821784Z","iopub.execute_input":"2026-01-05T16:06:50.822053Z","iopub.status.idle":"2026-01-05T16:06:50.838900Z","shell.execute_reply.started":"2026-01-05T16:06:50.822015Z","shell.execute_reply":"2026-01-05T16:06:50.837997Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"small_top_orders = top_orders[:20]\ndef suggest_buys(df):\n    # USER HISTORY AIDS AND TYPES\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    # UNIQUE AIDS AND UNIQUE BUYS\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    df = df.loc[(df['type']==1)|(df['type']==2)]\n    unique_buys = list(dict.fromkeys( df.aid.tolist()[::-1] ))\n    # RERANK CANDIDATES USING WEIGHTS\n    if len(unique_aids)>=20:\n        weights=np.logspace(0.5,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        # RERANK BASED ON REPEAT ITEMS AND TYPE OF ITEMS\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight[t]\n        # RERANK CANDIDATES USING \"BUY2BUY\" CO-VISITATION MATRIX\n        aids3 = list(itertools.chain(*[top_k_buy2buy[aid] for aid in unique_buys if aid in top_k_buy2buy]))\n        for aid in aids3: aids_temp[aid] += 0.1\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    # USE \"CART ORDER\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_k_buys[aid] for aid in unique_aids if aid in top_k_buys]))\n    # USE \"BUY2BUY\" CO-VISITATION MATRIX\n    aids3 = list(itertools.chain(*[top_k_buy2buy[aid] for aid in unique_buys if aid in top_k_buy2buy]))\n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2+aids3).most_common(20) if aid2 not in unique_aids] \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    # USE TOP20 TEST ORDERS\n    #return result + list(top_orders)[:20-len(result)]\n    # FIXED BY tetsuro731, remove duplicate\n    set_result = set(result)\n    return result + [i for i in small_top_orders if i not in set_result][:20 - len(result)]","metadata":{"execution":{"iopub.status.busy":"2026-01-05T16:06:50.839960Z","iopub.execute_input":"2026-01-05T16:06:50.840191Z","iopub.status.idle":"2026-01-05T16:06:50.853822Z","shell.execute_reply.started":"2026-01-05T16:06:50.840169Z","shell.execute_reply":"2026-01-05T16:06:50.853076Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Create Submission CSV\nInferring test data with Pandas groupby is slow. We need to accelerate the following code.","metadata":{}},{"cell_type":"code","source":"%%time\n\npred_df_clicks = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(\n    lambda x: suggest_clicks(x)\n)\ndel small_top_clicks\ngc.collect()\n\npred_df_carts = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(\n    lambda x: suggest_carts(x)\n)\ndel small_top_carts\ngc.collect()\n\npred_df_buys = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(\n    lambda x: suggest_buys(x)\n)\ndel small_top_orders\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2026-01-05T16:06:50.855004Z","iopub.execute_input":"2026-01-05T16:06:50.855732Z","iopub.status.idle":"2026-01-05T16:50:20.311075Z","shell.execute_reply.started":"2026-01-05T16:06:50.855679Z","shell.execute_reply":"2026-01-05T16:50:20.310137Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"clicks_pred_df = pd.DataFrame(pred_df_clicks.add_suffix(\"_clicks\"), columns=[\"labels\"]).reset_index()\norders_pred_df = pd.DataFrame(pred_df_buys.add_suffix(\"_orders\"), columns=[\"labels\"]).reset_index()\ncarts_pred_df = pd.DataFrame(pred_df_carts.add_suffix(\"_carts\"), columns=[\"labels\"]).reset_index()","metadata":{"execution":{"iopub.status.busy":"2026-01-05T16:50:20.312207Z","iopub.execute_input":"2026-01-05T16:50:20.312477Z","iopub.status.idle":"2026-01-05T16:50:23.232261Z","shell.execute_reply.started":"2026-01-05T16:50:20.312452Z","shell.execute_reply":"2026-01-05T16:50:23.231524Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"pred_df = pd.concat([clicks_pred_df, orders_pred_df, carts_pred_df])\npred_df.columns = [\"session_type\", \"labels\"]\npred_df[\"labels\"] = pred_df.labels.apply(lambda x: \" \".join(map(str,x)))\npred_df.to_csv(\"submission.csv\", index=False)\npred_df.head()","metadata":{"execution":{"iopub.status.busy":"2026-01-05T16:50:23.233184Z","iopub.execute_input":"2026-01-05T16:50:23.233427Z","iopub.status.idle":"2026-01-05T16:51:06.193014Z","shell.execute_reply.started":"2026-01-05T16:50:23.233404Z","shell.execute_reply":"2026-01-05T16:51:06.192120Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}