{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This notebook is the most interesting of my three public notebooks for OTTO competition.\n\nHere I'll run  \n* faster `suggest_clicks` and `suggest_buys` than [here](https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565).\n* faster validation recall function than [here](https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565).\n\nYou will see that those 3 functions run in 1 minute each.\n\nHow ?  \nBy using parallel CPU and lists instead of pandas dataframes.\n\nThose lists were created one time [here](https://www.kaggle.com/code/adaubas/otto-prepare-valid-test) and are available  [here](https://www.kaggle.com/datasets/adaubas/otto-valid-test-list).\n\nCo-visitation matrices were created [here](https://www.kaggle.com/code/adaubas/otto-co-visisitation-matrices) and are available [here](https://www.kaggle.com/datasets/adaubas/otto-co-visitation-matrices).\nThere were created with GPU like in Chris Deotte's notebook.\n\n\nWhy ?\n* I need to do validations as fast as possible\n* I do not have a 20's CPU computer ;-) and I'm playing with only Kaggle's tools.\n\nThank's to [Chris Deotte](https://www.kaggle.com/cdeotte) for this [notebook](https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565)\n\n\n**Version 2** : first public version<br>\n**Version 4** look at cell n°15 & 17 : update to compute faster RECALL@20 (only 30 seconds for clicks + carts + orders). Instead of doing `intersection` like in `otto_metric`'s function, we count zeroes in 20 * 46 columns (there is a zero if there's no diff between candidates column n°i and ground truth column n°j for a session), with :\n* 20 : max number of candidates for each session\n* 46 : max orders in ground truth   *(with carts it is 105 instead of 46)*<br>\n\n**Version 6** : cells 3 & 25 corrections. Thank's to [YuZhang](https://www.kaggle.com/yuzhang0422). Cells n°15 & 17 are splitted from 15 to 21 to see time execution.\n\nI tried we `numba` first, but `isin` function is not available.\n\n\n## Credits\nWe thank many Kagglers who have shared ideas. We use co-visitation matrix idea from Vladimir [here][1]. We use groupby sort logic from Sinan in comment section [here][4]. We use duplicate prediction removal logic from Radek [here][5]. We use multiple visit logic from Pietro [here][2]. We use type weighting logic from Ingvaras [here][3]. We use leaky test data from my previous notebook [here][4]. And some ideas may have originated from Tawara [here][6] and KJ [here][7]. We use Colum2131's parquets [here][8]. Above image is from Ravi's discussion about candidate rerank models [here][9]\n\n[1]: https://www.kaggle.com/code/vslaykovsky/co-visitation-matrix\n[2]: https://www.kaggle.com/code/pietromaldini1/multiple-clicks-vs-latest-items\n[3]: https://www.kaggle.com/code/ingvarasgalinskas/item-type-vs-multiple-clicks-vs-latest-items\n[4]: https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost\n[5]: https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\n[6]: https://www.kaggle.com/code/ttahara/otto-mors-aid-frequency-baseline\n[7]: https://www.kaggle.com/code/whitelily/co-occurrence-baseline\n[8]: https://www.kaggle.com/datasets/columbia2131/otto-chunk-data-inparquet-format\n[9]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721","metadata":{}},{"cell_type":"code","source":"VER = 1\nimport pandas as pd, numpy as np\nimport pickle, glob, gc\n\nfrom collections import Counter\nimport itertools\n\n# multiprocessing \nimport psutil\nN_CORES = psutil.cpu_count()     # Available CPU cores\nprint(f\"N Cores : {N_CORES}\")\nfrom multiprocessing import Pool","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:46:03.153637Z","iopub.execute_input":"2022-12-12T18:46:03.154101Z","iopub.status.idle":"2022-12-12T18:46:03.162750Z","shell.execute_reply.started":"2022-12-12T18:46:03.154062Z","shell.execute_reply":"2022-12-12T18:46:03.161375Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Validation","metadata":{}},{"cell_type":"code","source":"type_labels = {'clicks':0, 'carts':1, 'orders':2}\n\ndef load_test(files):    \n    dfs = []\n    for e, chunk_file in enumerate(glob.glob(files)):\n        chunk = pd.read_parquet(chunk_file)\n        chunk.ts = (chunk.ts/1000).astype('int32')\n        chunk['type'] = chunk['type'].map(type_labels).astype('int8')\n        dfs.append(chunk)\n    return pd.concat(dfs).reset_index(drop=True) #.astype({\"ts\": \"datetime64[ms]\"})\n\nvalid = load_test('../input/otto-validation/test_parquet/*')\nprint('Valid data has shape',valid.shape)\nvalid.head()","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-12-12T18:46:03.165029Z","iopub.execute_input":"2022-12-12T18:46:03.165473Z","iopub.status.idle":"2022-12-12T18:46:07.118796Z","shell.execute_reply.started":"2022-12-12T18:46:03.165437Z","shell.execute_reply":"2022-12-12T18:46:07.117303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\nDISK_PIECES = 4\n# LOAD THREE CO-VISITATION MATRICES\ndef pqt_to_dict(df):\n    return df.groupby('aid_x').aid_y.apply(list).to_dict()\n\ntop_20_clicks = pqt_to_dict( pd.read_parquet(f'../input/otto-co-visitation-matrices/top_20_valid_clicks_v{VER}_0.pqt') )\nfor k in range(1, DISK_PIECES): \n    top_20_clicks.update( pqt_to_dict( pd.read_parquet(f'../input/otto-co-visitation-matrices/top_20_valid_clicks_v{VER}_{k}.pqt') ) )\n\n\ntop_20_buys = pqt_to_dict( pd.read_parquet(f'../input/otto-co-visitation-matrices/top_15_valid_carts_orders_v{VER}_0.pqt') )\nfor k in range(1, DISK_PIECES): \n    top_20_buys.update( pqt_to_dict( pd.read_parquet(f'../input/otto-co-visitation-matrices/top_15_valid_carts_orders_v{VER}_{k}.pqt') ) )\n    \ntop_20_buy2buy = pqt_to_dict( pd.read_parquet(f'../input/otto-co-visitation-matrices/top_15_valid_buy2buy_v{VER}_0.pqt') )\n\n# TOP CLICKS AND ORDERS IN TEST\ntop_clicks = valid.loc[valid['type']==0, 'aid'].value_counts().index.values[:20]\ntop_orders = valid.loc[valid['type']==2, 'aid'].value_counts().index.values[:20]\n\nprint('Here are size of our 3 co-visitation matrices:')\nprint( len( top_20_clicks ), len( top_20_buy2buy ), len( top_20_buys ) )","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:46:07.120462Z","iopub.execute_input":"2022-12-12T18:46:07.120891Z","iopub.status.idle":"2022-12-12T18:48:19.654041Z","shell.execute_reply.started":"2022-12-12T18:46:07.120855Z","shell.execute_reply":"2022-12-12T18:48:19.652313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def df_parallelize_run(func, t_split):\n    \n    num_cores = np.min([N_CORES, len(t_split)])\n    pool = Pool(num_cores)\n    df = pool.map(func, t_split)\n    pool.close()\n    pool.join()\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:48:19.656923Z","iopub.execute_input":"2022-12-12T18:48:19.657391Z","iopub.status.idle":"2022-12-12T18:48:19.666067Z","shell.execute_reply.started":"2022-12-12T18:48:19.657351Z","shell.execute_reply":"2022-12-12T18:48:19.664223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nPIECES = 5\nvalid_bysession_list = []\nfor PART in range(PIECES):\n    with open(f'../input/otto-valid-test-list/valid_group_tolist_{PART}_{VER}.pkl', 'rb') as f:\n        valid_bysession_list.extend(pickle.load(f))\nprint(len(valid_bysession_list))","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:48:19.668433Z","iopub.execute_input":"2022-12-12T18:48:19.669012Z","iopub.status.idle":"2022-12-12T18:48:34.532805Z","shell.execute_reply.started":"2022-12-12T18:48:19.668971Z","shell.execute_reply":"2022-12-12T18:48:34.531239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#type_weight_multipliers = {'clicks': 1, 'carts': 6, 'orders': 3}\ntype_weight_multipliers = {0: 1, 1: 6, 2: 3}\n\ndef suggest_clicks(df):\n    \n    session = df[0]\n    aids = df[1]\n    types = df[2]\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    # RERANK CANDIDATES USING WEIGHTS\n    if len(unique_aids)>=20:\n        weights=np.logspace(0.1,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        # RERANK BASED ON REPEAT ITEMS AND TYPE OF ITEMS\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return session, sorted_aids\n    \n    # USE \"CLICKS\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(20) if aid2 not in unique_aids]    \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    \n    # USE TOP20 TEST CLICKS\n    return session, result + list(top_clicks)[:20-len(result)]","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:48:34.534751Z","iopub.execute_input":"2022-12-12T18:48:34.536533Z","iopub.status.idle":"2022-12-12T18:48:34.554057Z","shell.execute_reply.started":"2022-12-12T18:48:34.536461Z","shell.execute_reply":"2022-12-12T18:48:34.552008Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n# Predict on all sessions in parallel\ntemp = df_parallelize_run(suggest_clicks, valid_bysession_list)\nval_clicks = pd.Series([f[1]  for f in temp], index=[f[0] for f in temp])","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:48:34.556190Z","iopub.execute_input":"2022-12-12T18:48:34.557041Z","iopub.status.idle":"2022-12-12T18:49:46.795720Z","shell.execute_reply.started":"2022-12-12T18:48:34.556982Z","shell.execute_reply":"2022-12-12T18:49:46.793949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def suggest_buys(df):\n    # USE USER HISTORY AIDS AND TYPES\n    session = df[0]\n    aids = df[1]\n    types = df[2]\n\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    unique_buys = list(dict.fromkeys( [f for i, f in enumerate(aids) if types[i] in [1, 2]][::-1] ))\n\n    # RERANK CANDIDATES USING WEIGHTS\n    if len(unique_aids)>=20:\n        \n        weights=np.logspace(0.5,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        # RERANK BASED ON REPEAT ITEMS AND TYPE OF ITEMS\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        # RERANK CANDIDATES USING \"BUY2BUY\" CO-VISITATION MATRIX\n        aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n        for aid in aids3: aids_temp[aid] += 0.1\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return session, sorted_aids\n            \n    # USE \"CART ORDER\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_aids if aid in top_20_buys]))\n    # USE \"BUY2BUY\" CO-VISITATION MATRIX\n    aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2 + aids3).most_common(20) if aid2 not in unique_aids] \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    # USE TOP20 TEST ORDERS\n    return session, result + list(top_orders)[:20-len(result)]","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:49:46.797892Z","iopub.execute_input":"2022-12-12T18:49:46.799058Z","iopub.status.idle":"2022-12-12T18:49:46.817322Z","shell.execute_reply.started":"2022-12-12T18:49:46.798993Z","shell.execute_reply":"2022-12-12T18:49:46.815706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n# Predict on all sessions in parallel\ntemp = df_parallelize_run(suggest_buys, valid_bysession_list)\nval_buys = pd.Series([f[1]  for f in temp], index=[f[0] for f in temp])","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:49:46.823222Z","iopub.execute_input":"2022-12-12T18:49:46.824002Z","iopub.status.idle":"2022-12-12T18:51:04.624147Z","shell.execute_reply.started":"2022-12-12T18:49:46.823958Z","shell.execute_reply":"2022-12-12T18:51:04.622588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nvalid_labels = pd.read_parquet('../input/otto-validation/test_labels.parquet')","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:51:04.625780Z","iopub.execute_input":"2022-12-12T18:51:04.626224Z","iopub.status.idle":"2022-12-12T18:51:06.459892Z","shell.execute_reply.started":"2022-12-12T18:51:04.626181Z","shell.execute_reply":"2022-12-12T18:51:06.458800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"benchmark = {\"clicks\":0.5255597442145808, \"carts\":0.4093328152483512, \"orders\":0.6487936598117477, \"all\":.5646320148830121}\nweights = {'clicks': 0.10, 'carts': 0.30, 'orders': 0.60}\n\nvalid_labels = pd.read_parquet('../input/otto-validation/test_labels.parquet')\n\n\ndef hits(b):\n    # b[0] : session id\n    # b[1] : ground truth\n    # b[2] : aids prediction \n    return b[0], len(set(b[1]).intersection(set(b[2]))), np.clip(len(b[1]), 0, 20)\n\ndef otto_metric_piece(values, typ, verbose=True):\n    \n    c1 = pd.DataFrame(values, columns=[\"labels\"]).reset_index().rename({\"index\":\"session\"}, axis=1)\n    a = valid_labels.loc[valid_labels['type']==typ].merge(c1, how='left', on=['session'])\n\n    b=[[a0, a1, a2] for a0, a1, a2 in zip(a[\"session\"], a[\"ground_truth\"], a[\"labels\"])]\n    c = df_parallelize_run(hits, b)\n    c = np.array(c)\n    \n    recall = c[:,1].sum() / c[:,2].sum()\n    \n    print('{} recall = {:.5f} (vs {:.5f} in benchmark)'.format(typ ,recall, benchmark[typ]))\n    \n    return recall\n\ndef otto_metric(clicks, carts, orders, verbose = True):\n    \n    score = 0\n    score += weights[\"clicks\"] * otto_metric_piece(clicks, \"clicks\", verbose = verbose)\n    score += weights[\"carts\"] * otto_metric_piece(carts, \"carts\", verbose = verbose)\n    score += weights[\"orders\"] * otto_metric_piece(orders, \"orders\", verbose = verbose)\n    \n    if verbose:\n        print('=============')\n        print('Overall Recall = {:.5f} (vs {:.5f} in benchmark)'.format(score, benchmark[\"all\"]))\n        print('=============')\n    \n    return score","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:51:06.461242Z","iopub.execute_input":"2022-12-12T18:51:06.462653Z","iopub.status.idle":"2022-12-12T18:51:08.015702Z","shell.execute_reply.started":"2022-12-12T18:51:06.462605Z","shell.execute_reply":"2022-12-12T18:51:08.014377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n_ = otto_metric_piece(val_buys, \"orders\")","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:51:08.018453Z","iopub.execute_input":"2022-12-12T18:51:08.019038Z","iopub.status.idle":"2022-12-12T18:51:16.756863Z","shell.execute_reply.started":"2022-12-12T18:51:08.018978Z","shell.execute_reply":"2022-12-12T18:51:16.755052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n_ = otto_metric_piece(val_buys, \"carts\")","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:51:16.760483Z","iopub.execute_input":"2022-12-12T18:51:16.761189Z","iopub.status.idle":"2022-12-12T18:51:31.849484Z","shell.execute_reply.started":"2022-12-12T18:51:16.761092Z","shell.execute_reply":"2022-12-12T18:51:31.847213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n_ = otto_metric(val_clicks, val_buys, val_buys)","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:51:31.852615Z","iopub.execute_input":"2022-12-12T18:51:31.853175Z","iopub.status.idle":"2022-12-12T18:52:42.734873Z","shell.execute_reply.started":"2022-12-12T18:51:31.853088Z","shell.execute_reply":"2022-12-12T18:52:42.732891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Faster but result is not exactly the same\nInstead of doing `intesect` on each row, we will count zeroes in `columns differences` to compute RECALL@20.\n\nI'm not able to explain why result is different yet. I don't know if `intersect` method is more or less correct than this `columns difference` method. If you have an idea, please tell me.","metadata":{}},{"cell_type":"code","source":"%%time\n\n# Two columns difference\n# in each column, aids\n# if aids are the same, columns difference gives 0\n# and then, count zeroes\ndef comp(a, c):\n    n = a - c\n    return n[n==0].shape[0]\n\nfor typ in [\"carts\", \"orders\"]:\n\n    # Session number of all sessions candidates\n    session_list = np.array([ f[0] for f in temp ], dtype = np.int32).flatten()\n    \n    # candidates for each session, filled with some -1 to complete lines to have a rectangular matrix\n    preds = np.array([ f[1] + [-1]*(20-len(f[1] )) for f in temp ], dtype = np.int32)\n    # keep only candidates for sessions which are in ground truth sessions\n    preds = preds[np.isin(session_list, valid_labels.loc[(valid_labels['type'] == typ), [\"session\"]].values.flatten())] \n    \n    # Ground Truth\n    gtv = valid_labels.loc[valid_labels['type'] == typ].ground_truth.values\n\n    # How many actions on each session ground truth ?\n    gtv_lens = np.array([len(e) for e in list(gtv)])\n    gtv_lens.max()\n    \n    # Main loop\n    cpt = 0\n    # for each column in ground truth\n    for i in range(gtv_lens.max()):\n        y_val = np.array([e[i] for e in list(gtv[gtv_lens > i])])\n        # for each column in candidates\n        for j in range(20):\n            # number of matching candidates with ground_truth\n            cpt += comp(preds[gtv_lens > i, j], y_val)\n    \n    print(\"{} recall {:.5f} (versus benchmark {:.5f})\".format(typ, cpt / gtv_lens.sum(), benchmark[typ]))","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:52:42.737686Z","iopub.execute_input":"2022-12-12T18:52:42.739743Z","iopub.status.idle":"2022-12-12T18:53:13.777877Z","shell.execute_reply.started":"2022-12-12T18:52:42.739550Z","shell.execute_reply":"2022-12-12T18:53:13.776526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"typ = \"carts\"\n# Session number of all sessions candidates\nsession_list = np.array([ f[0] for f in temp ], dtype = np.int32).flatten()\n    \n# candidates for each session, filled with some -1 to complete lines to have a rectangular matrix\npreds = np.array([ f[1] + [-1]*(20-len(f[1] )) for f in temp ], dtype = np.int32)","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:53:13.779983Z","iopub.execute_input":"2022-12-12T18:53:13.781185Z","iopub.status.idle":"2022-12-12T18:53:31.420491Z","shell.execute_reply.started":"2022-12-12T18:53:13.781100Z","shell.execute_reply":"2022-12-12T18:53:31.419119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# keep only candidates for sessions which are in ground truth sessions\npreds = preds[np.isin(session_list, valid_labels.loc[(valid_labels['type'] == typ), [\"session\"]].values.flatten())] \n    \n# Ground Truth\ngtv = valid_labels.loc[valid_labels['type'] == typ].ground_truth.values\n\n# How many actions on each session ground truth ?\ngtv_lens = np.array([len(e) for e in list(gtv)])\ngtv_lens.max()\n    \n# Main loop\ncpt = 0\n# for each column in ground truth\nfor i in range(gtv_lens.max()):\n    y_val = np.array([e[i] for e in list(gtv[gtv_lens > i])])\n    # for each column in candidates\n    for j in range(20):\n        # number of matching candidates with ground_truth\n        cpt += comp(preds[gtv_lens > i, j], y_val)\n    \nprint(\"{} recall {:.5f} (versus benchmark {:.5f})\".format(typ, cpt / gtv_lens.sum(), benchmark[typ]))","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:53:31.422267Z","iopub.execute_input":"2022-12-12T18:53:31.422688Z","iopub.status.idle":"2022-12-12T18:53:35.107963Z","shell.execute_reply.started":"2022-12-12T18:53:31.422650Z","shell.execute_reply":"2022-12-12T18:53:35.106618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"typ = \"orders\"\n# Session number of all sessions candidates\nsession_list = np.array([ f[0] for f in temp ], dtype = np.int32).flatten()\n    \n# candidates for each session, filled with some -1 to complete lines to have a rectangular matrix\npreds = np.array([ f[1] + [-1]*(20-len(f[1] )) for f in temp ], dtype = np.int32)","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:53:35.110080Z","iopub.execute_input":"2022-12-12T18:53:35.110763Z","iopub.status.idle":"2022-12-12T18:53:42.469003Z","shell.execute_reply.started":"2022-12-12T18:53:35.110718Z","shell.execute_reply":"2022-12-12T18:53:42.467613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# keep only candidates for sessions which are in ground truth sessions\npreds = preds[np.isin(session_list, valid_labels.loc[(valid_labels['type'] == typ), [\"session\"]].values.flatten())] \n    \n# Ground Truth\ngtv = valid_labels.loc[valid_labels['type'] == typ].ground_truth.values\n\n# How many actions on each session ground truth ?\ngtv_lens = np.array([len(e) for e in list(gtv)])\ngtv_lens.max()\n    \n# Main loop\ncpt = 0\n# for each column in ground truth\nfor i in range(gtv_lens.max()):\n    y_val = np.array([e[i] for e in list(gtv[gtv_lens > i])])\n    # for each column in candidates\n    for j in range(20):\n        # number of matching candidates with ground_truth\n        cpt += comp(preds[gtv_lens > i, j], y_val)\n    \nprint(\"{} recall {:.5f} (versus benchmark {:.5f})\".format(typ, cpt / gtv_lens.sum(), benchmark[typ]))","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:53:42.470920Z","iopub.execute_input":"2022-12-12T18:53:42.472436Z","iopub.status.idle":"2022-12-12T18:53:44.123830Z","shell.execute_reply.started":"2022-12-12T18:53:42.472353Z","shell.execute_reply":"2022-12-12T18:53:44.122405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = df_parallelize_run(suggest_clicks, valid_bysession_list)\nval_clicks = pd.Series([f[1]  for f in temp], index=[f[0] for f in temp])","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:53:44.125329Z","iopub.execute_input":"2022-12-12T18:53:44.125717Z","iopub.status.idle":"2022-12-12T18:55:21.674105Z","shell.execute_reply.started":"2022-12-12T18:53:44.125683Z","shell.execute_reply":"2022-12-12T18:55:21.672651Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfor typ in [\"clicks\"]:\n\n    # Session number of all sessions candidates\n    session_list = np.array([ f[0] for f in temp ], dtype = np.int32).flatten()\n    \n    # candidates for each session, filled with some -1 to complete lines to have a rectangular matrix\n    preds = np.array([ f[1] + [-1]*(20-len(f[1] )) for f in temp ], dtype = np.int32)\n    # keep only candidates for sessions which are in ground truth sessions\n    preds = preds[np.isin(session_list, valid_labels.loc[(valid_labels['type'] == typ), [\"session\"]].values.flatten())] \n    \n    # Ground Truth\n    gtv = valid_labels.loc[valid_labels['type'] == typ].ground_truth.values\n\n    # How many actions on each session ground truth ?\n    gtv_lens = np.array([len(e) for e in list(gtv)])\n    gtv_lens.max()\n    \n    # Main loop\n    cpt = 0\n    # for each column in ground truth\n    for i in range(gtv_lens.max()):\n        y_val = np.array([e[i] for e in list(gtv[gtv_lens > i])])\n        # for each column in candidates\n        for j in range(20):\n            # number of matching candidates with ground_truth\n            cpt += comp(preds[gtv_lens > i, j], y_val)\n    \n    print(\"{} recall {:.5f} (versus benchmark {:.5f})\".format(typ, cpt / gtv_lens.sum(), benchmark[typ]))","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:55:21.676397Z","iopub.execute_input":"2022-12-12T18:55:21.677546Z","iopub.status.idle":"2022-12-12T18:55:31.726053Z","shell.execute_reply.started":"2022-12-12T18:55:21.677486Z","shell.execute_reply":"2022-12-12T18:55:31.724382Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del temp\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:55:31.728674Z","iopub.execute_input":"2022-12-12T18:55:31.729374Z","iopub.status.idle":"2022-12-12T18:55:40.789495Z","shell.execute_reply.started":"2022-12-12T18:55:31.729305Z","shell.execute_reply":"2022-12-12T18:55:40.788053Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# FREE MEMORY\ndel valid_bysession_list, val_clicks, val_buys\ndel top_20_clicks, top_20_buy2buy, top_20_buys, top_clicks, top_orders, valid\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-12-12T18:55:40.791368Z","iopub.execute_input":"2022-12-12T18:55:40.791910Z","iopub.status.idle":"2022-12-12T18:55:48.800339Z","shell.execute_reply.started":"2022-12-12T18:55:40.791868Z","shell.execute_reply":"2022-12-12T18:55:48.798836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Test\nHere a submission file is created, the same than [here](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575) but faster.","metadata":{}},{"cell_type":"code","source":"test = load_test('../input/otto-chunk-data-inparquet-format/test_parquet/*')\nprint('Test data has shape',test.shape)\ntest.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-03T15:58:43.578830Z","iopub.execute_input":"2022-12-03T15:58:43.579216Z","iopub.status.idle":"2022-12-03T15:58:46.240465Z","shell.execute_reply.started":"2022-12-03T15:58:43.579185Z","shell.execute_reply":"2022-12-03T15:58:46.239198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\ntop_20_clicks = pqt_to_dict( pd.read_parquet(f'../input/otto-co-visitation-matrices/top_20_test_clicks_v{VER}_0.pqt') )\nfor k in range(1, DISK_PIECES): \n    top_20_clicks.update( pqt_to_dict( pd.read_parquet(f'../input/otto-co-visitation-matrices/top_20_test_clicks_v{VER}_{k}.pqt') ) )\n\n\ntop_20_buys = pqt_to_dict( pd.read_parquet(f'../input/otto-co-visitation-matrices/top_15_test_carts_orders_v{VER}_0.pqt') )\nfor k in range(1, DISK_PIECES): \n    top_20_buys.update( pqt_to_dict( pd.read_parquet(f'../input/otto-co-visitation-matrices/top_15_test_carts_orders_v{VER}_{k}.pqt') ) )\n    \ntop_20_buy2buy = pqt_to_dict( pd.read_parquet(f'../input/otto-co-visitation-matrices/top_15_test_buy2buy_v{VER}_0.pqt') )\n\n# TOP CLICKS AND ORDERS IN TEST\ntop_clicks = test.loc[test['type']==0, 'aid'].value_counts().index.values[:20]\ntop_orders = test.loc[test['type']==2, 'aid'].value_counts().index.values[:20]\n\nprint('Here are size of our 3 co-visitation matrices:')\nprint( len( top_20_clicks ), len( top_20_buy2buy ), len( top_20_buys ) )","metadata":{"execution":{"iopub.status.busy":"2022-12-03T15:59:05.343698Z","iopub.execute_input":"2022-12-03T15:59:05.344920Z","iopub.status.idle":"2022-12-03T16:01:06.228759Z","shell.execute_reply.started":"2022-12-03T15:59:05.344881Z","shell.execute_reply":"2022-12-03T16:01:06.227346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nPIECES = 5\ntest_bysession_list = []\nfor PART in range(PIECES):\n    with open(f'../input/otto-valid-test-list/test_group_tolist_{PART}_{VER}.pkl', 'rb') as f:\n        test_bysession_list.extend(pickle.load(f))\nprint(len(test_bysession_list))","metadata":{"execution":{"iopub.status.busy":"2022-12-03T16:01:06.230957Z","iopub.execute_input":"2022-12-03T16:01:06.231326Z","iopub.status.idle":"2022-12-03T16:01:18.909638Z","shell.execute_reply.started":"2022-12-03T16:01:06.231268Z","shell.execute_reply":"2022-12-03T16:01:18.908126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n# Predict on all sessions in parallel\ntemp = df_parallelize_run(suggest_clicks, test_bysession_list)\nclicks_pred_df = pd.Series([f[1] for f in temp], index=[f[0] for f in temp])\nclicks_pred_df = clicks_pred_df.add_suffix(\"_clicks\")\nclicks_pred_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-03T16:12:25.041062Z","iopub.execute_input":"2022-12-03T16:12:25.041536Z","iopub.status.idle":"2022-12-03T16:13:23.048697Z","shell.execute_reply.started":"2022-12-03T16:12:25.041495Z","shell.execute_reply":"2022-12-03T16:13:23.047374Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\n# Predict on all sessions in parallel\ntemp = df_parallelize_run(suggest_buys, test_bysession_list)\nbuys_pred_df = pd.Series([f[1] for f in temp], index=[f[0] for f in temp])\norders_pred_df = buys_pred_df.add_suffix(\"_orders\")\ncarts_pred_df = buys_pred_df.add_suffix(\"_carts\")","metadata":{"execution":{"iopub.status.busy":"2022-12-03T16:13:23.051249Z","iopub.execute_input":"2022-12-03T16:13:23.051683Z","iopub.status.idle":"2022-12-03T16:14:22.887680Z","shell.execute_reply.started":"2022-12-03T16:13:23.051630Z","shell.execute_reply":"2022-12-03T16:14:22.886263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_df = pd.concat([clicks_pred_df, orders_pred_df, carts_pred_df]).reset_index()\npred_df.columns = [\"session_type\", \"labels\"]\npred_df[\"labels\"] = pred_df.labels.apply(lambda x: \" \".join(map(str,x)))\npred_df.to_csv(\"submission.csv\", index=False)\npred_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-03T16:14:22.889591Z","iopub.execute_input":"2022-12-03T16:14:22.890313Z","iopub.status.idle":"2022-12-03T16:15:04.893240Z","shell.execute_reply.started":"2022-12-03T16:14:22.890259Z","shell.execute_reply":"2022-12-03T16:15:04.892095Z"},"trusted":true},"execution_count":null,"outputs":[]}]}