{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":38760,"databundleVersionId":4493939,"sourceType":"competition"},{"sourceId":4933424,"sourceType":"datasetVersion","datasetId":2860873}],"dockerImageVersionId":30382,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Rules Only Model Achieves LB 0.590\nIn this notebook, we show how a \"rules only\" model can achieve **LB 0.590** submission for Kaggle's Otto competition. This simple notebook loads 20 covisit matrices which were created with RAPIDS cuDF and achieves Top 50 final private LB! If we add a GBT reranker, then we can boost the LB score to single model **LB 0.601** !! \n \nThis notebook is similar to my original public notebook [here][1] with 17 additional covisit matrices added. And new logic to incorporate the new covisit matrices. There is a discussion about this \"rules only\" notebook [here][2]\n\n[1]: https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\n[2]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/383013","metadata":{"papermill":{"duration":0.005829,"end_time":"2022-11-03T16:49:27.399833","exception":false,"start_time":"2022-11-03T16:49:27.394004","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np, gc, glob\nimport multiprocessing, os, pickle\nfrom collections import Counter\n\nITEM_CT = 20","metadata":{"papermill":{"duration":0.126841,"end_time":"2022-11-03T16:49:27.538248","exception":false,"start_time":"2022-11-03T16:49:27.411407","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-01-12T12:18:06.361207Z","iopub.execute_input":"2024-01-12T12:18:06.362523Z","iopub.status.idle":"2024-01-12T12:18:06.401794Z","shell.execute_reply.started":"2024-01-12T12:18:06.362385Z","shell.execute_reply":"2024-01-12T12:18:06.400652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Click candidates\n\nBelow are a description of the 20 covisit matrices used in this notebook. All covisit matrices were computed quickly using RAPIDS cuDF (in another notebook and uploaded to Kaggle dataset). Example code to compute covist with cuDF is shown [here][1]:  \n* **top_20** - this covisit matrix is in my original notebook\n* **top_20b** - all covisit pair counts are consecutive items. See code below.  \n    `df['k'] = np.arange(len(df))`  \n    `df = df.merge(df, on=['session'])`  \n    `df = df.loc[ (df.k_y - df.k_x).abs()==1 ]`  \n* **top_20_test2** - Use most recent 2 weeks data with time decay.\n\n[1]: https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575","metadata":{}},{"cell_type":"code","source":"%%time\nPATH = '/kaggle/input/otto-covisit-matrix/'\ntop_20 = pickle.load(open(PATH+'top_40_aids_v104.pkl', 'rb')) \n#top_20b = pickle.load(open(PATH+'top_40_aids_v23.pkl', 'rb')) \n#top_20_test2 = pickle.load(open(PATH+'top_40_aids_v34.pkl', 'rb'))","metadata":{"execution":{"iopub.status.busy":"2023-11-02T08:33:14.570557Z","iopub.execute_input":"2023-11-02T08:33:14.571192Z","iopub.status.idle":"2023-11-02T08:34:20.967610Z","shell.execute_reply.started":"2023-11-02T08:33:14.571153Z","shell.execute_reply":"2023-11-02T08:34:20.966126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list(top_20.keys())[0]","metadata":{"execution":{"iopub.status.busy":"2023-11-02T08:37:04.701445Z","iopub.execute_input":"2023-11-02T08:37:04.701924Z","iopub.status.idle":"2023-11-02T08:37:04.836749Z","shell.execute_reply.started":"2023-11-02T08:37:04.701887Z","shell.execute_reply":"2023-11-02T08:37:04.835276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"top_20[404474]","metadata":{"execution":{"iopub.status.busy":"2023-11-02T08:37:16.248950Z","iopub.execute_input":"2023-11-02T08:37:16.249379Z","iopub.status.idle":"2023-11-02T08:37:16.259056Z","shell.execute_reply.started":"2023-11-02T08:37:16.249345Z","shell.execute_reply":"2023-11-02T08:37:16.257525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Most Popular Items from test.csv","metadata":{"papermill":{"duration":0.006685,"end_time":"2022-11-03T16:50:19.618761","exception":false,"start_time":"2022-11-03T16:50:19.612076","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# COMPUTED FROM TEST DATA\ntop_clicks = [1460571,  485256,  108125,  986164, 1551213,  754412,  554660,\n        832192,  579690,   33343, 1006198,  688602,   29735,  329725,\n        184976, 1019736,  496180,  861401,  944778,  659399, 1043508,\n       1022566,  811371, 1604220,  836852,  471073,  819288, 1264313,\n        508883, 1751274,  620545,  959208,  717965,  332654, 1731920,\n        544144,  147526, 1116095, 1294924,  102345, 1645990, 1497089,\n        558573,   95488, 1196256,  199409, 1110150, 1146575, 1236775,\n        137514, 1030009,  435253, 1800674,  881286, 1609228, 1286213,\n        337471,  670066,  831165, 1685214, 1673641,  909449, 1260564,\n       1099100,  995962,  612920, 1647563, 1462420, 1741695, 1281615,\n       1603001, 1722991,  442293,  206735, 1219503,  166037,  799923,\n       1469891,  557072, 1156699,  111891, 1624436, 1782099, 1639229,\n        530377, 1197632, 1140985,  152547,  247240, 1449873, 1825743,\n        901817, 1420240, 1733943,  542343,  680375,  406358,  147278,\n       1627951,  836707]","metadata":{"execution":{"iopub.status.busy":"2023-10-17T11:36:07.305940Z","iopub.execute_input":"2023-10-17T11:36:07.306188Z","iopub.status.idle":"2023-10-17T11:36:07.313236Z","shell.execute_reply.started":"2023-10-17T11:36:07.306165Z","shell.execute_reply":"2023-10-17T11:36:07.312206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Rules to Suggest Clicks","metadata":{}},{"cell_type":"code","source":"import itertools\n\ntype_weight_multipliers = {0: 1, 1: 6, 2: 3}\n\ndef suggest_clicks_old(df):\n    # USER HISTORY AIDS AND TYPES\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    # RERANK CANDIDATES USING WEIGHTS\n    if len(unique_aids)>=20:\n        weights=np.logspace(0.1,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        # RERANK BASED ON REPEAT ITEMS AND TYPE OF ITEMS\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    # USE \"CLICKS\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(20) if aid2 not in unique_aids]    \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    # USE TOP20 TEST CLICKS\n    return result + list(top_clicks)[:20-len(result)]\n\n\ndef suggest_aids(df):\n    session = df[0]\n    aids = df[1]\n    types = df[2]\n    real_session_last_item = df[6]\n        \n    \n    #### ALL UNIQUE ITEMS IN USER HISTORY\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    \n    #### LAST ITEM FROM EACH REAL SESSION IN USER HISTORY\n    unique_aids3 = list(dict.fromkeys( [f for i, f in enumerate(aids) if real_session_last_item[i] == 1][::-1] ))\n    \n    if len(unique_aids)>=15:\n        weights=np.logspace(0.1,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n         \n        ##### START NEW PART >= 15\n        aids3 = list(itertools.chain(*[candidates[\"top_20b\"][aid][:15] for aid in unique_aids3 if aid in candidates[\"top_20b\"]]))\n        for i,aid in enumerate(aids3):\n            aids_temp[aid] += 0.3\n        ##### END NEW PART >= 15    \n        \n        result = [k for k,v in aids_temp.most_common(ITEM_CT)]\n        return session, (result + top_clicks[:ITEM_CT-len(result)])[:ITEM_CT]\n    \n    aids_temp = Counter() \n    \n    \n    ##### START NEW PART < 15\n    for i, a in enumerate(unique_aids):\n        w0 = np.max([1 - (0.35 * i), 0.001]) #Weight aid order starting from the last one. \n        if a in top_20:\n            for j, aj in enumerate(candidates[\"top_20\"][a]):\n                w1 = np.max([1 - (0.005 * j), 0.01]) #Weight the candidate aid from the dict\n                aids_temp[aj] += (w0*w1)\n            \n    aids3 = list(itertools.chain(*[candidates[\"top_20b\"][aid][:20] for aid in unique_aids[:2] if aid in candidates[\"top_20b\"]]))\n    for i,aid in enumerate(aids3):\n        aids_temp[aid] += 1\n        if i%20==0: aids_temp[aid] += 1\n    ##### END NEW PART < 15a\n        \n    top_aids2 = [k for k,v in aids_temp.most_common(ITEM_CT) if k not in unique_aids]\n    result = unique_aids + top_aids2[:ITEM_CT - len(unique_aids)]\n    return session, (result + top_clicks[:ITEM_CT-len(result)])[:ITEM_CT]","metadata":{"papermill":{"duration":0.019192,"end_time":"2022-11-03T16:50:23.026501","exception":false,"start_time":"2022-11-03T16:50:23.007309","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-10-17T11:36:07.315128Z","iopub.execute_input":"2023-10-17T11:36:07.315365Z","iopub.status.idle":"2023-10-17T11:36:07.330837Z","shell.execute_reply.started":"2023-10-17T11:36:07.315340Z","shell.execute_reply":"2023-10-17T11:36:07.329866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Use Parallel Processing\nThe following fast code to run my `suggest clicks` and `suggest carts/orders` quickly is from Aldparis notebook [here][1]\n\n[1]: https://www.kaggle.com/code/adaubas/otto-fast-handcrafted-model-recall-20","metadata":{}},{"cell_type":"code","source":"import psutil\nN_CORES = psutil.cpu_count()     \nprint(f\"N Cores : {N_CORES}\")\nfrom multiprocessing import Pool","metadata":{"execution":{"iopub.status.busy":"2023-10-17T11:36:07.331860Z","iopub.execute_input":"2023-10-17T11:36:07.332122Z","iopub.status.idle":"2023-10-17T11:36:07.346084Z","shell.execute_reply.started":"2023-10-17T11:36:07.332098Z","shell.execute_reply":"2023-10-17T11:36:07.345349Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"N_CORES = N_CORES//2\ndef df_parallelize_run(func, t_split):\n    \n    num_cores = np.min([N_CORES, len(t_split)])\n    pool = Pool(num_cores)\n    df = pool.map(func, t_split)\n    pool.close()\n    pool.join()\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2023-10-17T11:36:07.346928Z","iopub.execute_input":"2023-10-17T11:36:07.347185Z","iopub.status.idle":"2023-10-17T11:36:07.357098Z","shell.execute_reply.started":"2023-10-17T11:36:07.347163Z","shell.execute_reply":"2023-10-17T11:36:07.356333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Kaggle's Test Data\nWe converted Kaggle's test data into lists for each session. Code by Aldparis showing how to do this is [here][1]\n\n[1]: https://www.kaggle.com/code/adaubas/otto-prepare-valid-test","metadata":{}},{"cell_type":"code","source":"%%time\nPIECES = 10\ntest_bysession_list = []\nfor PART in range(PIECES):\n    with open(f'{PATH}test_group_tolist_{PART}_1.pkl', 'rb') as f:\n        test_bysession_list.extend(pickle.load(f))\nprint(len(test_bysession_list))","metadata":{"execution":{"iopub.status.busy":"2023-10-17T11:36:07.357956Z","iopub.execute_input":"2023-10-17T11:36:07.358209Z","iopub.status.idle":"2023-10-17T11:37:12.320610Z","shell.execute_reply.started":"2023-10-17T11:36:07.358186Z","shell.execute_reply":"2023-10-17T11:37:12.319564Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Suggest Clicks","metadata":{}},{"cell_type":"code","source":"def create_prediction_with_matrix(test_list, candidate_matrix):\n    global candidates\n    candidates = {\"top_20\": top_20, \"top_20b\": top_20b, \"top_20_test2\": top_20_test2}\n    candidates = {key: ({} if key not in [\"top_20\", candidate_matrix] else value) for key, value in candidates.items()}\n    \n    print(f\"using candidates: {['top_20', candidate_matrix]}\")\n    \n    print(\"suggesting aids\")\n    temp = df_parallelize_run(suggest_aids, test_bysession_list)\n    \n    print(\"creating click_df\")\n    click_df = pd.Series([f[1]  for f in temp], index=[f[0] for f in temp])\n\n    print(\"creating pred_df\")\n    pred_df = pd.DataFrame(click_df.add_suffix(\"_clicks\"), columns=[\"labels\"]).reset_index()\n    \n    pred_df.columns = [\"session_type\", \"labels\"]\n    pred_df[\"labels\"] = pred_df.labels.apply(lambda x: \" \".join(map(str,x)))\n    \n    print(\"creating submission\")\n    pred_df.to_csv(f\"top20+{candidate_matrix}.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2023-10-17T11:52:53.546359Z","iopub.execute_input":"2023-10-17T11:52:53.546732Z","iopub.status.idle":"2023-10-17T11:52:53.554969Z","shell.execute_reply.started":"2023-10-17T11:52:53.546698Z","shell.execute_reply":"2023-10-17T11:52:53.554310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for candidate in candidates:\n    create_prediction_with_matrix(test_bysession_list, candidate)","metadata":{"execution":{"iopub.status.busy":"2023-10-17T11:53:01.853144Z","iopub.execute_input":"2023-10-17T11:53:01.853528Z","iopub.status.idle":"2023-10-17T11:58:39.045903Z","shell.execute_reply.started":"2023-10-17T11:53:01.853494Z","shell.execute_reply":"2023-10-17T11:58:39.044649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}