{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Multiple clicks vs latest items\n**❗❗❗This notebook is a fork of Chris Deotte's [notebook](https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost).❗❗❗**\n\nIn this notebook I added a small change in the prediction logic when we have more than 20 unique aids.\n\nInstead of plain prediction of latest items in the session I gave a weight to each item seen in the session that was depending on the position inside the session, in particular I used np.logspace(...,base=2)-1 to generate the weights.\n\nAfter assigning the weights I create a dictionary that assigned to each item the sum of the weights of its occurences inside the session.\n\nThe last step was suggesting the items ordering them in descending order given their weight.\n\nThe intuition behind this approach is quite easy, items visited more times are likely to be more interesting to the user, thus they should have a higher weight (the hard part is computing this weight). \n\nI have many ideas to try:\n1. Using directly counts (likely not too good)\n2. Filtering session aids older than 24hours (similar to how the co visitation matrix is used)\n3. Using a weighting technique that considers the time distance between the items seen in the session\n4. Analyzing loops(user clicks an item, and after some clicks goes back to that same item with no cart nor order) inside sessions to see if some can be dropped (coming back to the original item may be an indicator that the items inside the loop were less interesting than the item that started the loop).\n\nIf you find this notebook interesting remember to upvote first \nVladimir's [noteboob](https://www.kaggle.com/code/vslaykovsky/co-visitation-matrix) and Chris's [notebook](https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost).\n\nTo find more easily the modified parts I left comments like this one in the code\n\n**###CHANGED PART-----------------------------------------------------------------------------**\n\n# Original Chris notebook Introduction\nThis notebook is a fork of Vladimir's notebook [here][1]. Please upvote his original notebook.\n\nIn Kaggle's OTTO – Multi-Objective Recommender System Competition, we are given both the train data (from present) and the test data (from future). The test data contains partial sequences of sessions. Therefore we can use these partial sequences to train our models in addition to using the train data. In this notebook we begin with Vladimir's Co-visitation Matrix notebook [here][1] and add test data to the training data. We make one change to his code below. Where he loads train parquets with `'../input/otto-chunk-data-inparquet-format/train_parquet/*'` , we change this to include test with `'../input/otto-chunk-data-inparquet-format/*_parquet/*'`.\n\nThis is an experiment to see if using test data during training will boost LB score. The original notebook achieved LB 0.539. Let's see what this notebook achieves. UPDATE: Experiment was successful and boosted LB `+0.003`, so it works!\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/leak.png)\n\nNote this method of using test data cannot be used in real life because some of the data we are training with occurs in the future of some of the inference data. The inference test data is from a one week period. Therefore when we infer events occurring the first day of the week, we have the advantage of using events occurring in the last 6 days of the week which is impossible in real life. Discussion [here][3]\n\n# Notes\nVersion 1 confirms that using leaky test data during training boosts our LB score. We will now try some additional experiments to boost LB with co-visitation matrix. The best way to experiment is to set up a local CV scheme and try experiments on local CV. But just for fun, I will try a few experiments directly on LB. The details of the experiments are explained below. Let's see how they perform on LB...\n\n* **Version 0 LB 0.539** Original notebook using only train data (and no test data).\n* **Version 1 LB 0.541** Train with both test and train data. Boost LB from 0.539 to 0.541. So using test data helps!\n* **Version 2 LB 0.541** Train with test data as weight 2 and train data weight 1 via modifying `all_pairs[aid1][aid2] +=w`. LB slightly better.\n* **Version 3 LB 0.546** Train data plus test data spans 5 weeks. Use weight 4 for most recent data and weight 1 for least recent with formula: `w = 1 + 3*(ts - 1659304800025)/(1662328791563-1659304800025)`. We also remove duplicate recommendations from prediction using `set` on `tail(20)`. Achieved LB 0.546!\n* **Version 4 LB 0.556** Same as version 3 except we use top40. And we remove duplicate recommendations more efficiently using `list(dict.fromkeys(AIDS[::-1]))` trick from Radek1 notebook [here][2]. Achieved LB 0.556!\n* **Version 5** Has BUG, please ignore\n* **Version 6 LB 0.555** Same as version 4 except use equal weight for train and test data. I'm curious if unequal weight is helping or not. I think most of the boost between version 4 and version 1 is because of removing duplicates during inference. (Using top40 instead of top20 only boosts `+0.001`). Let's see how equal weights does. UPDATE: using weights introduced in version 3 boost `+0.001`.\n* **Version 7 LB 0.557** In this experiment, we make Co-visitation Matrices for each type. When making Co-visitation matrix for type \"orders\" we boost weight of pair to w=5 when second item is an \"order\". For \"carts\", we boost to w=5. For \"clicks\" we do not boost. UPDATE: it boost LB `+0.001`.\n* **Version 8 LB 0.550** Since version 7 seemed to work, in this experiment, for \"orders\" top40 we use w=100 for \"order\" and w=10 for \"cart\". Then for \"carts\" top40 we use w=10 for \"order\" and w=100 for \"cart\". For \"clicks\" top40 we use w=1 for all. We also add our time weight from version 3. UPDATE: it hurt LB `-0.006`.\n* **Version 9 LB 562** We will use a suggestion from Sinan Calisir in the comments. When inferring test data we will sort by `session` and `ts` but not `type` before groupby apply. Let's try this with our best version 7 notebook. Woohoo! It achieved LB 0.562! Thanks Sinan!\n* **Version 10 LB ???** We will try one more time to adjust the weights of \"orders\" and \"carts\". For the \"order\" top40 we will use w=10. And for the \"cart\" top40 we will use w=4. For \"clicks\" w=1. We will also use time weight from version 3. Let's see how this does...\n\nNote if we had a local CV, then we could set up a grid search and find the optimal weights to use with our co-visitation matrix (with regard to importance of time, clicks, orders, carts and whatever else we can think of). And we could quickly find what other modifications can boost our model's performance.\n\n[1]: https://www.kaggle.com/code/vslaykovsky/co-visitation-matrix\n[2]: https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\n[3]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/363939","metadata":{}},{"cell_type":"markdown","source":"# OTTO: Co-visitation Matrix\n\nThere exist products that are frequently viewed and bought together. Here we leverage this idea by computing a co-visitation matrix of products. It's done in the following way:\n\n1. First we look at all pairs of events within the same session that are close to each other in time (< 1 day). We compute co-visitation matrix $M_{aid1,aid2}$ by counting global number of event pairs for each pair across all sessions.\n2. For each $aid1$ we find top 20 most frequent aid2:  `aid2=argsort(M[aid])[-20:]`\n3. We produce test results by concatenating `tail(20)` of test session events (see https://www.kaggle.com/code/simamumu/old-test-data-last-20-aid-get-lb0-947) with the most likely recommendations from co-visitation matrix. These recommendations are generated from session AIDs and `aid2` from the step 2\n\n\n**Please, smash that thumbs up button and subscribe if you like this notebook!**","metadata":{}},{"cell_type":"markdown","source":"## Utils, imports","metadata":{}},{"cell_type":"code","source":"### import numpy as np\nfrom collections import defaultdict\nimport pandas as pd\nfrom tqdm.notebook import tqdm\nimport glob\nimport numpy as np\nimport multiprocessing\nimport os\nimport pickle\n\nimport glob\nfrom collections import Counter\n\nDEBUG=False   \nSAMPLING = 1  # Reduce it to improve performance","metadata":{"execution":{"iopub.status.busy":"2022-11-06T11:39:12.672114Z","iopub.execute_input":"2022-11-06T11:39:12.673585Z","iopub.status.idle":"2022-11-06T11:39:12.811301Z","shell.execute_reply.started":"2022-11-06T11:39:12.673458Z","shell.execute_reply":"2022-11-06T11:39:12.810035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"TOP_20_CACHE = '../input/otto-pickles/top_40_aids_v4.pkl'\n","metadata":{"execution":{"iopub.status.busy":"2022-11-06T11:39:12.813745Z","iopub.execute_input":"2022-11-06T11:39:12.814175Z","iopub.status.idle":"2022-11-06T11:39:12.819468Z","shell.execute_reply.started":"2022-11-06T11:39:12.81413Z","shell.execute_reply":"2022-11-06T11:39:12.818388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Generate AID pairs","metadata":{}},{"cell_type":"code","source":"import sys\ndef gen_pairs(df):\n    df = df.query('session % @SAMPLING == 0').groupby('session', as_index=False, sort=False).apply(lambda g: g.tail(30)).reset_index(drop=True)\n    df = pd.merge(df, df, on='session')\n    pairs = df.query('abs(ts_x - ts_y) < 24 * 60 * 60 * 1000 and aid_x != aid_y')[['session', 'aid_x', 'aid_y', 'ts_x', 'type_y']]\\\n        .drop_duplicates(['session', 'aid_x', 'aid_y'])\n    return pairs[['aid_x', 'aid_y', 'ts_x', 'type_y']].values\n    \n\ndef gen_aid_pairs():\n    all_pairs = defaultdict(lambda: Counter())\n    with tqdm(glob.glob('../input/otto-chunk-data-inparquet-format/*_parquet/*'), desc='Chunks') as prog:\n        with multiprocessing.Pool(4) as p:\n            for idx, chunk_file in enumerate(prog):\n                chunk = pd.read_parquet(chunk_file)#.drop(columns=['type'])\n                pair_chunks = p.map(gen_pairs, np.array_split(chunk.head(100000000 if not DEBUG else 10000), 120))            \n                for pairs in pair_chunks:\n                    for aid1, aid2, ts, typ in pairs:\n                        w = 1 + 3*(ts - 1659304800025)/(1662328791563-1659304800025)\n                        # HERE WE CAN BOOST WEIGHT, i.e. IF TYP==\"ORDERS\": W *= 10.0\n                        # THEN SAVE THIS MATRIX AS THE \"ORDERS\" MATRIX\n                        # WE CAN MAKE 3 DIFFERENT CO-VISITATION MATRICES\n                        all_pairs[aid1][aid2] +=w \n                prog.set_description(f'Mem: {sys.getsizeof(object) // (2 ** 20)}MB')\n\n                if DEBUG and idx >= 2:\n                    break\n                del chunk, pair_chunks\n                gc.collect()\n    return all_pairs\n        \nif os.path.exists(TOP_20_CACHE):\n    print('Reading top20 AIDs from cache')\n    top_20 = pickle.load(open(TOP_20_CACHE, 'rb'))\nelse:\n    all_pairs = gen_aid_pairs()\n    df_top_20 = []\n    for aid, cnt in tqdm(all_pairs.items()):\n        df_top_20.append({'aid1': aid, 'aid2': [aid2 for aid2, freq in cnt.most_common(20)]})\n\n    df_top_20 = pd.DataFrame(df_top_20).set_index('aid1')\n    top_20 = df_top_20.aid2.to_dict()\n    import pickle\n    with open('top_20_aids.pkl', 'wb') as f:\n        pickle.dump(top_20, f)\n        \nlen(top_20)","metadata":{"execution":{"iopub.status.busy":"2022-11-06T11:39:12.821027Z","iopub.execute_input":"2022-11-06T11:39:12.822553Z","iopub.status.idle":"2022-11-06T11:40:12.517774Z","shell.execute_reply.started":"2022-11-06T11:39:12.822503Z","shell.execute_reply":"2022-11-06T11:40:12.516175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i, (k, v) in enumerate(top_20.items()):\n    print(k, v)\n    if i > 10:\n        break","metadata":{"execution":{"iopub.status.busy":"2022-11-06T11:40:12.522699Z","iopub.execute_input":"2022-11-06T11:40:12.523294Z","iopub.status.idle":"2022-11-06T11:40:12.536428Z","shell.execute_reply.started":"2022-11-06T11:40:12.523257Z","shell.execute_reply":"2022-11-06T11:40:12.534194Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Test set inference","metadata":{}},{"cell_type":"code","source":"def load_test():    \n    dfs = []\n    for e, chunk_file in enumerate(tqdm(glob.glob('../input/otto-chunk-data-inparquet-format/test_parquet/*'))):\n        chunk = pd.read_parquet(chunk_file)\n        dfs.append(chunk)\n\n    return pd.concat(dfs).reset_index(drop=True).astype({\"ts\": \"datetime64[ms]\"})","metadata":{"execution":{"iopub.status.busy":"2022-11-06T11:40:12.539732Z","iopub.execute_input":"2022-11-06T11:40:12.540571Z","iopub.status.idle":"2022-11-06T11:40:12.550763Z","shell.execute_reply.started":"2022-11-06T11:40:12.540515Z","shell.execute_reply":"2022-11-06T11:40:12.549005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df = load_test()","metadata":{"execution":{"iopub.status.busy":"2022-11-06T11:40:12.552496Z","iopub.execute_input":"2022-11-06T11:40:12.55301Z","iopub.status.idle":"2022-11-06T11:40:15.641281Z","shell.execute_reply.started":"2022-11-06T11:40:12.552959Z","shell.execute_reply":"2022-11-06T11:40:15.639999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"##################\n# BELOW IS CODE ADDED BY CHRIS\n\ntop_20_orders = pickle.load(open('../input/otto-pickles-2/top_40_orders_v6.pkl', 'rb'))\ntop_20_carts = pickle.load(open('../input/otto-pickles-2/top_40_carts_v7.pkl', 'rb'))\n","metadata":{"execution":{"iopub.status.busy":"2022-11-06T11:40:15.643076Z","iopub.execute_input":"2022-11-06T11:40:15.644017Z","iopub.status.idle":"2022-11-06T11:40:52.897355Z","shell.execute_reply.started":"2022-11-06T11:40:15.643977Z","shell.execute_reply":"2022-11-06T11:40:52.895853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import itertools\nfrom collections import defaultdict\n\ntype_weight_multipliers_click = {'clicks': 1, 'carts': 4, 'orders': 2}\n\ntype_weight_multipliers_cart = {'clicks': 2, 'carts': 3, 'orders': 1}\n\ntype_weight_multipliers_order = {'clicks': 1, 'carts': 4, 'orders': 2}\n\n\ndef suggest_aids(df):\n    # REMOVE DUPLICATE AIDS AND REVERSE ORDER OF LIST\n    \n    ###CHANGED PART-----------------------------------------------------------------------------\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    if len(unique_aids)<20:\n        aids=unique_aids\n    else:\n        # THIS IS A BASIC IDEA -> you can think about exponential weight decay, or even using df.ts and time-dependent weight decay\n        weights=np.linspace(0.3, 1, len(aids), endpoint=True) \n        aids_temp=defaultdict(lambda :0)\n        #to each aids give the sum of the weights of its occurences within the session\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid]+= w * type_weight_multipliers_click[t]\n        #order in descending order the aids depending on its weight\n        sorted_aids=[k for k, v in sorted(aids_temp.items(), key=lambda item: -item[1])]\n        \n        #just a sanity check to see if they get correctly sorted\n        #if len(unique_aids)<22: \n        #    print(aids_temp)\n        #    print(sorted_aids)\n        #    break\n        \n        #get first 20 aids with highest weight (visited more recently or visited frequently enough to overcome not being recent)\n        return sorted_aids[:20]\n    ###END CHANGED PART-----------------------------------------------------------------------------\n    \n    # Append it with AIDs from the co-visitation matrix. \n    aids2 = list(itertools.chain(*[top_20[aid] for aid in aids if aid in top_20]))\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(20) if aid2 not in aids]        \n    return list(aids) + top_aids2[:20 - len(aids)]\n\n\ndef suggest_orders(df):\n    # REMOVE DUPLICATE AIDS AND REVERSE ORDER OF LIST\n    ###CHANGED PART-----------------------------------------------------------------------------\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    if len(unique_aids)<20:\n        aids=unique_aids\n    else:\n        weights=np.linspace(0.3, 1, len(aids), endpoint=True) \n        aids_temp=defaultdict(lambda :0)\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid]+= w * type_weight_multipliers_order[t]\n\n        sorted_aids=[k for k, v in sorted(aids_temp.items(), key=lambda item: -item[1])]\n        return sorted_aids[:20]\n    ###END CHANGED PART-----------------------------------------------------------------------------\n    \n    # Append it with AIDs from the co-visitation matrix. \n    aids2 = list(itertools.chain(*[top_20_orders[aid] for aid in aids if aid in top_20_orders]))\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(20) if aid2 not in aids]        \n    return list(aids) + top_aids2[:20 - len(aids)]\n\ndef suggest_carts(df):\n    # REMOVE DUPLICATE AIDS AND REVERSE ORDER OF LIST\n    ###CHANGED PART-----------------------------------------------------------------------------\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    if len(unique_aids)<20:\n        aids=unique_aids\n    else:\n        weights=np.linspace(0.3, 1, len(aids), endpoint=True) \n        aids_temp=defaultdict(lambda :0)\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid]+= w * type_weight_multipliers_cart[t]\n        sorted_aids=[k for k, v in sorted(aids_temp.items(), key=lambda item: -item[1])]\n        return sorted_aids[:20]\n    ###END CHANGED PART-----------------------------------------------------------------------------\n    \n    # Append it with AIDs from the co-visitation matrix. \n    aids2 = list(itertools.chain(*[top_20_carts[aid] for aid in aids if aid in top_20_carts]))\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(20) if aid2 not in aids]        \n    return list(aids) + top_aids2[:20 - len(aids)]","metadata":{"execution":{"iopub.status.busy":"2022-11-06T11:43:19.443995Z","iopub.execute_input":"2022-11-06T11:43:19.444469Z","iopub.status.idle":"2022-11-06T11:43:19.472035Z","shell.execute_reply.started":"2022-11-06T11:43:19.444432Z","shell.execute_reply":"2022-11-06T11:43:19.470081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tqdm.pandas() #enable progress_apply in pandas","metadata":{"execution":{"iopub.status.busy":"2022-11-06T11:43:57.617969Z","iopub.execute_input":"2022-11-06T11:43:57.618564Z","iopub.status.idle":"2022-11-06T11:43:57.625526Z","shell.execute_reply.started":"2022-11-06T11:43:57.618494Z","shell.execute_reply":"2022-11-06T11:43:57.624021Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_df = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).progress_apply(\n    lambda x: suggest_aids(x)\n)\n\n##################\n# BELOW IS CODE ADDED BY CHRIS\n\npred_df_orders = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).progress_apply(\n    lambda x: suggest_orders(x)\n)\n\npred_df_carts = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).progress_apply(\n    lambda x: suggest_carts(x)\n)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-11-06T11:43:58.474114Z","iopub.execute_input":"2022-11-06T11:43:58.475134Z","iopub.status.idle":"2022-11-06T12:05:17.722381Z","shell.execute_reply.started":"2022-11-06T11:43:58.475077Z","shell.execute_reply":"2022-11-06T12:05:17.721038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clicks_pred_df = pd.DataFrame(pred_df.add_suffix(\"_clicks\"), columns=[\"labels\"]).reset_index()\norders_pred_df = pd.DataFrame(pred_df_orders.add_suffix(\"_orders\"), columns=[\"labels\"]).reset_index()\ncarts_pred_df = pd.DataFrame(pred_df_carts.add_suffix(\"_carts\"), columns=[\"labels\"]).reset_index()","metadata":{"execution":{"iopub.status.busy":"2022-11-06T12:08:29.272974Z","iopub.execute_input":"2022-11-06T12:08:29.273505Z","iopub.status.idle":"2022-11-06T12:08:32.728889Z","shell.execute_reply.started":"2022-11-06T12:08:29.273465Z","shell.execute_reply":"2022-11-06T12:08:32.727542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_df","metadata":{"execution":{"iopub.status.busy":"2022-11-06T12:08:32.731223Z","iopub.execute_input":"2022-11-06T12:08:32.73163Z","iopub.status.idle":"2022-11-06T12:08:32.745612Z","shell.execute_reply.started":"2022-11-06T12:08:32.731596Z","shell.execute_reply":"2022-11-06T12:08:32.7444Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_df = pd.concat(\n    [clicks_pred_df, orders_pred_df, carts_pred_df]\n)\npred_df.columns = [\"session_type\", \"labels\"]\npred_df[\"labels\"] = pred_df.labels.apply(lambda x: \" \".join(map(str,x)))\npred_df.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-11-06T12:08:32.747288Z","iopub.execute_input":"2022-11-06T12:08:32.747784Z","iopub.status.idle":"2022-11-06T12:09:28.781479Z","shell.execute_reply.started":"2022-11-06T12:08:32.747737Z","shell.execute_reply":"2022-11-06T12:09:28.780504Z"},"trusted":true},"execution_count":null,"outputs":[]}]}