{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Rules Only Model Achieves LB 0.590\nIn this notebook, we show how a \"rules only\" model can achieve **LB 0.590** submission for Kaggle's Otto competition. This simple notebook loads 20 covisit matrices which were created with RAPIDS cuDF and achieves Top 50 final private LB! If we add a GBT reranker, then we can boost the LB score to single model **LB 0.601** !! \n \nThis notebook is similar to my original public notebook [here][1] with 17 additional covisit matrices added. And new logic to incorporate the new covisit matrices. There is a discussion about this \"rules only\" notebook [here][2]\n\n[1]: https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\n[2]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/383013","metadata":{"papermill":{"duration":0.005829,"end_time":"2022-11-03T16:49:27.399833","exception":false,"start_time":"2022-11-03T16:49:27.394004","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# Uses 20 Covisit Matrices!\n\nBelow are a description of the 20 covisit matrices used in this notebook. All covisit matrices were computed quickly using RAPIDS cuDF (in another notebook and uploaded to Kaggle dataset). Example code to compute covist with cuDF is shown [here][1]:  \n* **top_20** - this covisit matrix is in my original notebook\n* **top_20b** - all covisit pair counts are consecutive items. See code below.  \n    `df['k'] = np.arange(len(df))`  \n    `df = df.merge(df, on=['session'])`  \n    `df = df.loc[ (df.k_y - df.k_x).abs()==1 ]`  \n* **top_20c** - all covisit pair counts are `(df.k_y - df.k_x).abs()<=2`\n* **top_20d** - all covisit pairs are carts/orders and forward at most 3 consecutive  \n    `df = df.loc[df['type'].isin(['carts','orders'])]`  \n    `df = df.merge(df, on=['session'])`  \n    `df = df.loc[ (df.k_y - df.k_x > 0) & (df.k_y - df.k_x <= 3) ]`  \n* **top_20e** - all covisit pairs are `(df.k_y - df.k_x).abs()<=3` and have time decay with  \n    `df['wgt'] = (1/2)**( (df.ts_x - df.ts_y).abs() /60/60)`  \n* **top_20f** - same as above but `(df.k_y - df.k_x).abs()<=6`  \n* **top_20_orders** - this covisit matrix is in my original notebook\n* **top_20_buy2buy** - this covisit matrix is in my original notebook\n* **top_20_buy2buy2** - use most recent 3 weeks data and only carts/orders. Apply time decay shown above.\n* **top_20_test** - use most recent 3 weeks data. Only forward in time pairs. Use clicks/carts/orders to carts/orders. Add time decay  \n    `df = df.loc[df.ts >= LAST_3_WEEKS ]`  \n    `df2 = df.loc[df['type'].isin(['carts','orders'])]`  \n    `df = df.merge(df2, on=['session'])`  \n    `df = df.loc[ df.ts_y - df.ts_x > 0 ]`  \n    `df['wgt'] = (1/2)**( (df.ts_x - df.ts_y).abs() /60/60)`  \n* **top_20_test2** - Use most recent 2 weeks data with time decay.\n* **top_20_buy** - Limit to forward 2 hours. Use clicks/carts/orders to carts/orders. Apply time decay.\n* **top_20_new** - Use cold start users in train. Pairs using only their first history item\n* **top_20_new2** - Use cold start users in train. Pairs using only their first history item\n* **top_40_day** - Use only last week data. Forward in time. Clicks/carts/orders to carts/orders. Time decay\n* **top_40_day2** - Use only last week data. Time decay\n* **top_40_less** - Use train users with less than 6 history and test users with less than 3\n* **top_40_more** - Use train users with more than 6 history and test users with more than 3\n* **top_40_less2** - Use item pairs with first item before 2pm. Clicks/carts/orders to carts/orders. Time decay\n* **top_40_more2** - Use item pairs with first item after 2pm. Clicks/carts/orders to carts/orders. Time decay\n\n[1]: https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np, gc, glob\nimport multiprocessing, os, pickle\nfrom collections import Counter\n\nITEM_CT = 20","metadata":{"papermill":{"duration":0.126841,"end_time":"2022-11-03T16:49:27.538248","exception":false,"start_time":"2022-11-03T16:49:27.411407","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-01T22:48:02.893919Z","iopub.execute_input":"2023-02-01T22:48:02.894200Z","iopub.status.idle":"2023-02-01T22:48:03.291453Z","shell.execute_reply.started":"2023-02-01T22:48:02.894174Z","shell.execute_reply":"2023-02-01T22:48:03.290600Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Most Popular Items from test.csv","metadata":{"papermill":{"duration":0.006685,"end_time":"2022-11-03T16:50:19.618761","exception":false,"start_time":"2022-11-03T16:50:19.612076","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# COMPUTED FROM TEST DATA\ntop_orders = [ 986164, 1460571,  329725, 1043508,  332654,  688602,   29735,\n       1495817,  579690, 1022566, 1006198,  471073,  832192,  544144,\n       1825743,  836852,  147526, 1236775,  166037, 1030009, 1609228,\n        508883,  923948, 1462420,  892871,  554660, 1457846,  258353,\n       1734475,  480314,  145332,  108125, 1286213, 1336175, 1359971,\n        137514,  714524,  558573,  172856,  585186,  352192, 1176975,\n       1146575,  954951, 1496287,  823143, 1699089,   25964, 1257293,\n        399315, 1441266, 1196256, 1294924, 1603001, 1274545,  414968,\n       1581568,  247240, 1116095,  383437,  530377,  272744, 1445562,\n        269257,  791627, 1140985, 1708326,  631899,  670066,  122983,\n        223273,  165160,  881286, 1768724,  868327, 1604220,  406358,\n       1722991, 1568011, 1025795, 1647563,  835431, 1531805,  714968,\n        500609, 1217083, 1668343, 1159757, 1610239, 1647157, 1264313,\n       1798916,  423558,  752652,  184976, 1255910, 1413049,  801774,\n        615566, 1034578]\ntop_carts = [ 485256,   33343, 1460571,  986164,  554660,  660655, 1116095,\n        152547, 1022566,  544144,  832192,  579690,  329725, 1043508,\n       1006198,  558573,  471073,  332654,  688602,   29735,  508883,\n        258353, 1736857, 1462420,  166037, 1609228, 1778843,  108125,\n       1495817, 1604220, 1825743, 1562705,  147526,  836852, 1286213,\n         25964, 1236775,  923948, 1281615, 1257293,  917587,  835431,\n       1439409,  892871,  125957,  122983, 1097061, 1449873, 1568011,\n       1030727, 1146575, 1731920,  326904, 1196256,  714524, 1768724,\n        480314, 1800674, 1662401, 1359971,  455191,  496180,  145332,\n        616283, 1708326, 1294924, 1270528,  944778, 1223508,  881286,\n        165160,  272744,  670066,  868327, 1734475,  137514,  172856,\n       1122221,  442293, 1685214,  823143, 1413049, 1722991, 1647157,\n        406358, 1733943,  700995, 1025795,  754412,  530377,  102416,\n        184976, 1445562, 1565495, 1019736, 1274545, 1083665,  667563,\n       1264313,  563117]\ntop_clicks = [1460571,  485256,  108125,  986164, 1551213,  754412,  554660,\n        832192,  579690,   33343, 1006198,  688602,   29735,  329725,\n        184976, 1019736,  496180,  861401,  944778,  659399, 1043508,\n       1022566,  811371, 1604220,  836852,  471073,  819288, 1264313,\n        508883, 1751274,  620545,  959208,  717965,  332654, 1731920,\n        544144,  147526, 1116095, 1294924,  102345, 1645990, 1497089,\n        558573,   95488, 1196256,  199409, 1110150, 1146575, 1236775,\n        137514, 1030009,  435253, 1800674,  881286, 1609228, 1286213,\n        337471,  670066,  831165, 1685214, 1673641,  909449, 1260564,\n       1099100,  995962,  612920, 1647563, 1462420, 1741695, 1281615,\n       1603001, 1722991,  442293,  206735, 1219503,  166037,  799923,\n       1469891,  557072, 1156699,  111891, 1624436, 1782099, 1639229,\n        530377, 1197632, 1140985,  152547,  247240, 1449873, 1825743,\n        901817, 1420240, 1733943,  542343,  680375,  406358,  147278,\n       1627951,  836707]","metadata":{"execution":{"iopub.status.busy":"2023-02-01T22:48:03.292772Z","iopub.execute_input":"2023-02-01T22:48:03.293150Z","iopub.status.idle":"2023-02-01T22:48:03.316116Z","shell.execute_reply.started":"2023-02-01T22:48:03.293121Z","shell.execute_reply":"2023-02-01T22:48:03.315136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Covisit Matrices\nThese covisit matrices were created using GPU RAPIDS cuDF in another notebook and described in the intro section.","metadata":{}},{"cell_type":"code","source":"%%time\nPATH = '/kaggle/input/otto-covisit-matrix/'\ntop_20 = pickle.load(open(PATH+'top_40_aids_v104.pkl', 'rb')) \ntop_20b = pickle.load(open(PATH+'top_40_aids_v23.pkl', 'rb')) \ntop_20c = pickle.load(open(PATH+'top_80_aids_v24.pkl', 'rb')) \ntop_20d = pickle.load(open(PATH+'top_30_aids_v28.pkl', 'rb')) \ntop_20e = pickle.load(open(PATH+'top_80_aids_v130.pkl', 'rb')) \ntop_20f = pickle.load(open(PATH+'top_80_aids_v132.pkl', 'rb')) \n\ntop_20_test2 = pickle.load(open(PATH+'top_40_aids_v34.pkl', 'rb'))\ntop_20_new2 = pickle.load(open(PATH+'top_40_aids_v801_0.pkl', 'rb'))\ntop_40_day = pickle.load(open(PATH+'top_40_aids_v162_d0_0_LB.pkl', 'rb'))\ntop_40_day2 = pickle.load(open(PATH+'top_80_aids_v163_d0_0_LB.pkl', 'rb'))","metadata":{"execution":{"iopub.status.busy":"2023-02-01T22:48:52.754782Z","iopub.execute_input":"2023-02-01T22:48:52.755096Z","iopub.status.idle":"2023-02-01T22:52:16.329915Z","shell.execute_reply.started":"2023-02-01T22:48:52.755068Z","shell.execute_reply":"2023-02-01T22:52:16.328720Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Rules to Suggest Clicks/Carts/Orders","metadata":{}},{"cell_type":"code","source":"import itertools\n\ntype_weight_multipliers = {0: 1, 1: 6, 2: 3}\n\ndef suggest_aids(df):\n    \n    #aids=df.aid.tolist()\n    #types = df.type.tolist()\n    \n    session = df[0]\n    aids = df[1]\n    types = df[2]\n    tss = df[3]\n    ds = df[4]\n    ds2 = df[6]\n    #days = df[7]\n    \n    top_day = top_40_day2\n    \n    #### ALL UNIQUE ITEMS IN USER HISTORY\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    \n    #### LAST ITEM FROM EACH REAL SESSION IN USER HISTORY\n    #df2 = df.sort_values('ts',ascending=False).drop_duplicates('d')\n    #aids2 = df2.aid.tolist()\n    #unique_aids3 = list(dict.fromkeys(aids2[::-1] )) \n    unique_aids3 = list(dict.fromkeys( [f for i, f in enumerate(aids) if ds2[i] == 1][::-1] ))\n    \n    #### ALL ITEMS FROM LAST REAL SESSION IN USER HISTORY\n    #mx = df.d.max()\n    #aids2 = df.loc[df.d==mx].aid.tolist()\n    #unique_aids4 = list(dict.fromkeys(aids2[::-1] ))\n    mx = np.max(ds)\n    unique_aids4 = list(dict.fromkeys( [f for i, f in enumerate(aids) if ds[i] == mx][::-1] ))\n    \n    #### ALL ITEMS FROM USER HISTORY MOST RECENT 24 HOURS\n    #aids2 = df.loc[df.ts >= mx - 60*60*24].aid.tolist()\n    #unique_aids6 = list(dict.fromkeys(aids2[::-1] ))  \n    mx = np.max(tss)\n    unique_aids6 = list(dict.fromkeys( [f for i, f in enumerate(aids) if tss[i] >= mx - 60*60*24 ][::-1] ))\n    \n    #### ALL UNIQUE CARTS/ORDERS IN USER HISTORY\n    #df = df.loc[ df['type'].isin([1,2]) ]\n    #unique_buys = list(dict.fromkeys( df.aid.tolist()[::-1] ))\n    unique_buys = list(dict.fromkeys( [f for i, f in enumerate(aids) if types[i] in [1, 2]][::-1] ))\n    \n    ln = len(unique_aids)\n \n    if len(unique_aids)>=15:\n        weights=np.logspace(0.1,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        aids3 = list(itertools.chain(*[top_20c[aid][:20] for aid in unique_aids[:2] if aid in top_20c]))\n        for i,aid in enumerate(aids3):\n            aids_temp[aid] += 0.6\n        aids3 = list(itertools.chain(*[top_20b[aid][:15] for aid in unique_aids3 if aid in top_20b]))\n        for i,aid in enumerate(aids3):\n            aids_temp[aid] += 0.3\n        aids3 = list(itertools.chain(*[top_20_test2[aid][:20] for aid in unique_aids[:2] if aid in top_20_test2]))\n        for i,aid in enumerate(aids3):\n            aids_temp[aid] += 0.6\n                \n        result = [k for k,v in aids_temp.most_common(ITEM_CT)]\n        return session, (result + top_clicks[:ITEM_CT-len(result)])[:ITEM_CT]\n        #return sorted_aids \n    \n    aids_temp = Counter() \n    \n    # NEW\n    MM = 2\n    aids2 = list(itertools.chain(*[top_day[aid][:10*MM] for aid in unique_aids6 if aid in top_day]))     \n    for i,aid in enumerate( aids2 ):\n        aids_temp[aid] += 1  \n    \n    weights3 = [2,2] + [1]*28 \n    if len(unique_aids)==1:\n        aids5 = list(itertools.chain(*[top_20_new2[aid][:30] for aid in unique_aids[-1:] if aid in top_20_new2]))\n        w5 = weights3* int(len(aids5)//30)\n        for aid,w in zip(aids5,w5):\n            aids_temp[aid] += w\n    \n    #aids2 = list(itertools.chain(*[top_20[aid][:20] for aid in unique_aids if aid in top_20]))\n    #for i,aid in enumerate(aids2):\n    #    m = 0.1 + 0.9*(ln-(i//20))/ln\n    #    aids_temp[aid] += m\n    #    if i%20==0: aids_temp[aid] += m\n            \n    # FROM GIBA\n    for i, a in enumerate(unique_aids):\n        w0 = np.max([1 - (0.35 * i), 0.001]) #Weight aid order starting from the last one. \n        if a in top_20:\n            for j, aj in enumerate(top_20[a]):\n                w1 = np.max([1 - (0.005 * j), 0.01]) #Weight the candidate aid from the dict\n                aids_temp[aj] += (w0*w1)\n            \n    aids3 = list(itertools.chain(*[top_20b[aid][:20] for aid in unique_aids[:2] if aid in top_20b]))\n    for i,aid in enumerate(aids3):\n        aids_temp[aid] += 1\n        if i%20==0: aids_temp[aid] += 1\n            \n    aids3 = list(itertools.chain(*[top_20_test2[aid][:20] for aid in unique_aids[:2] if aid in top_20_test2]))\n    for i,aid in enumerate(aids3):\n        aids_temp[aid] += 1\n        if i%20==0: aids_temp[aid] += 1\n            \n    aids4 = list(itertools.chain(*[top_20f[aid][:10] for aid in unique_aids4 if aid in top_20f]))\n    for i,aid in enumerate(aids4):\n        w = i//10\n        aids_temp[aid] += 1 -w*0.1\n        if i%10==0: aids_temp[aid] += 1 -w*0.1\n            \n    aids5 = list(itertools.chain(*[top_20e[aid][:20] for aid in unique_aids3 if aid in top_20e]))\n    for i,aid in enumerate(aids5):\n        aids_temp[aid] += 1\n        if i%20==0: aids_temp[aid] += 1\n    top_aids2 = [k for k,v in aids_temp.most_common(1) if k not in unique_aids]\n    \n    aids3 = list(itertools.chain(*[top_20c[aid][:10] for aid in top_aids2 if aid in top_20c]))\n    for i,aid in enumerate(aids3):\n        aids_temp[aid] += 1\n        if i%10==0: aids_temp[aid] += 1\n    top_aids2 = [k for k,v in aids_temp.most_common(ITEM_CT) if k not in unique_aids]\n    \n    result = unique_aids + top_aids2[:ITEM_CT - len(unique_aids)]\n    return session, (result + top_clicks[:ITEM_CT-len(result)])[:ITEM_CT]\n\ndef suggest_orders(df):\n    \n    #aids = df.aid.tolist()\n    #types = df.type.tolist()\n    \n    session = df[0]\n    aids = df[1]\n    aids9 = aids.copy()\n    types = df[2]\n    tss = df[3]\n    ds = df[4]\n    ds1 = df[5]\n    ds2 = df[6]\n    days = df[7]\n    \n    top_day = top_40_day\n    click_aids = click_df[session][:10]\n    \n    #### ALL UNIQUE ITEMS IN USER HISTORY\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    \n    #### ALL ITEMS FROM LAST REAL SESSION IN USER HISTORY\n    #mx = df.d.max()\n    #aids2 = df.loc[df.d==mx].aid.tolist()\n    #unique_aids4 = list(dict.fromkeys(aids2[::-1] )) \n    mx = np.max(ds)\n    unique_aids4 = list(dict.fromkeys( [f for i, f in enumerate(aids) if ds[i] == mx][::-1] ))\n    \n    #### ALL ITEMS FROM USER HISTORY MOST RECENT 30 MINUTES\n    #mx = df.ts.max()\n    #aids2 = df.loc[df.ts >= mx - 60*60/2].aid.tolist()\n    #unique_aids5 = list(dict.fromkeys(aids2[::-1] )) \n    mx = np.max(tss)\n    unique_aids5 = list(dict.fromkeys( [f for i, f in enumerate(aids) if tss[i] >= mx - 60*60/2 ][::-1] ))\n    \n    #### ALL ITEMS FROM USER HISTORY MOST RECENT 24 HOURS\n    #aids2 = df.loc[df.ts >= mx - 60*60*24].aid.tolist()\n    #unique_aids6 = list(dict.fromkeys(aids2[::-1] ))  \n    unique_aids6 = list(dict.fromkeys( [f for i, f in enumerate(aids) if tss[i] >= mx - 60*60*24 ][::-1] ))\n    \n    #### FIRST ITEM FROM EACH REAL SESSION IN USER HISTORY\n    #df2 = df.drop_duplicates('d')\n    #aids2 = df2.aid.tolist()\n    #unique_aids2 = list(dict.fromkeys(aids2[::-1] )) #first of each session\n    unique_aids2 = list(dict.fromkeys( [f for i, f in enumerate(aids) if ds1[i] == 1][::-1] ))\n    \n    #### LAST ITEM FROM EACH REAL SESSION IN USER HISTORY\n    #df2 = df.sort_values('ts',ascending=False).drop_duplicates('d')\n    #aids2 = df2.aid.tolist()\n    #unique_aids3 = list(dict.fromkeys(aids2 )) #last of each session\n    unique_aids3 = list(dict.fromkeys( [f for i, f in enumerate(aids) if ds2[i] == 1][::-1] ))\n    \n    #### ALL UNIQUE CARTS/ORDERS IN USER HISTORY\n    #df = df.loc[ df['type'].isin([1,2]) ]\n    #unique_buys = list(dict.fromkeys( df.aid.tolist()[::-1] ))\n    unique_buys = list(dict.fromkeys( [f for i, f in enumerate(aids) if types[i] in [1, 2]][::-1] ))\n    \n    if len(unique_aids)>=20:\n\n        weights=np.logspace(0.5,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        for aid in unique_aids2: \n            aids_temp[aid] += 0.5\n        for aid in unique_aids3: \n            aids_temp[aid] += 0.5\n            \n        aids3 = list(itertools.chain(*[top_20_buy2buy[aid][:40] for aid in unique_buys if aid in top_20_buy2buy]))\n        for i,aid in enumerate(aids3):\n            aids_temp[aid] += 0.05\n            if i%40==0: aids_temp[aid] += 0.05\n        aids3 = list(itertools.chain(*[top_20_buy2buy2[aid][:40] for aid in unique_buys if aid in top_20_buy2buy2]))\n        for i,aid in enumerate(aids3):\n            aids_temp[aid] += 0.1\n            if i%40==0: aids_temp[aid] += 0.1\n                \n        aids4 = list(itertools.chain(*[top_20_test[aid][:40] for aid in unique_aids if aid in top_20_test]))\n        for i,aid in enumerate(aids4):\n            aids_temp[aid] += 0.05\n            if i%40==0: aids_temp[aid] += 0.05\n        aids5 = list(itertools.chain(*[top_20c[aid][:20] for aid in unique_aids[:1] if aid in top_20c]))\n        for i,aid in enumerate(aids5):\n            aids_temp[aid] += 0.05\n            if i%20==0: aids_temp[aid] += 0.05\n        aids6 = list(itertools.chain(*[top_20d[aid][:20] for aid in unique_buys[:1] if aid in top_20d]))\n        for i,aid in enumerate(aids6):\n            aids_temp[aid] += 0.05\n            if i%20==0: aids_temp[aid] += 0.05\n                \n        aids7 = list(itertools.chain(*[top_20b[aid][:5] for aid in unique_aids3 if aid in top_20b]))\n        for i,aid in enumerate(aids7):\n            aids_temp[aid] += 0.25\n            if i%5==0: aids_temp[aid] += 0.25\n        aids7 = list(itertools.chain(*[top_20b[aid][:5] for aid in unique_aids2 if aid in top_20b]))\n        for i,aid in enumerate(aids7):\n            aids_temp[aid] += 0.125\n            if i%5==0: aids_temp[aid] += 0.125\n                \n                \n        aids4 = list(itertools.chain(*[top_day[aid][:20] for aid in unique_aids6 if aid in top_day]))\n        for i,aid in enumerate(aids4):\n            aids_temp[aid] += 0.05\n            if i%20==0: aids_temp[aid] += 0.05\n        aids4 = list(itertools.chain(*[top_20_test[aid][:10] for aid in click_aids if aid in top_20_test]))\n        for i,aid in enumerate(aids4):\n            aids_temp[aid] += 0.05\n            if i%10==0: aids_temp[aid] += 0.05\n        aids4 = list(itertools.chain(*[top_20_buy[aid][:10] for aid in click_aids if aid in top_20_buy]))\n        for i,aid in enumerate(aids4):\n            aids_temp[aid] += 0.05\n            if i%10==0: aids_temp[aid] += 0.05\n        for aid in click_aids:\n            aids_temp[aid] += 0.05\n            \n        MM = 3\n        aids4 = list(itertools.chain(*[top_40_more[aid][:MM] for aid in unique_aids if aid in top_40_more]))\n        for i,aid in enumerate(aids4):\n            aids_temp[aid] += 0.03\n            if i%MM==0: aids_temp[aid] += 0.03\n        aids4 = list(itertools.chain(*[top_40_more2[aid][:MM] for aid in unique_aids if aid in top_40_more2]))\n        for i,aid in enumerate(aids4):\n            aids_temp[aid] += 0.03\n            if i%MM==0: aids_temp[aid] += 0.03\n        aids4 = list(itertools.chain(*[top_40_less[aid][:MM] for aid in unique_aids if aid in top_40_less]))\n        for i,aid in enumerate(aids4):\n            aids_temp[aid] += 0.03\n            if i%MM==0: aids_temp[aid] += 0.03\n        aids4 = list(itertools.chain(*[top_40_less2[aid][:MM] for aid in unique_aids if aid in top_40_less2]))\n        for i,aid in enumerate(aids4):\n            aids_temp[aid] += 0.03\n            if i%MM==0: aids_temp[aid] += 0.03\n                \n        result = [k for k,v in aids_temp.most_common(ITEM_CT)]\n        return session, (result + top_orders[:ITEM_CT-len(result)])[:ITEM_CT]\n        #return sorted_aids \n    \n    weights = [2,2] + [1]*8 \n    weights2 = [2,2] + [1]*53 \n    weights3 = [2,2] + [1]*18 \n    weights4 = [2,2] + [1]*38 \n    weights5 = [2] + [1]*4 \n    \n    ln = len(unique_aids)\n    \n    aids_temp = Counter() \n    aids2 = list(itertools.chain(*[top_20_orders[aid][:10] for aid in unique_aids if aid in top_20_orders]))\n    w2 = weights* int(len(aids2)//10)\n    aids3 = list(itertools.chain(*[top_20_buy2buy[aid][:10] for aid in unique_buys if aid in top_20_buy2buy]))\n    w3 = weights* int(len(aids3)//10)\n    aids4 = list(itertools.chain(*[top_20_test[aid][:10] for aid in unique_aids if aid in top_20_test]))\n    w4 = weights* int(len(aids4)//10)\n    aids5 = list(itertools.chain(*[top_20_buy2buy2[aid][:10] for aid in unique_buys if aid in top_20_buy2buy2]))\n    w5 = weights* int(len(aids5)//10)\n    for i,(aid,w) in enumerate(zip(aids2,w2)):\n        m = 0.25 + 0.75*(ln-(i//10))/ln\n        aids_temp[aid] += w*m\n    for i,(aid,w) in enumerate(zip(aids3,w3)):\n        aids_temp[aid] += w/2\n    for i,(aid,w) in enumerate(zip(aids4,w4)):\n        m = 0.25 + 0.75*(ln-(i//10))/ln\n        aids_temp[aid] += w*m\n    for i,(aid,w) in enumerate(zip(aids5,w5)):\n        aids_temp[aid] += w/2\n        \n    MM = 5\n    top_40_use = top_40_more\n    aids2 = list(itertools.chain(*[top_40_use[aid][:MM] for aid in unique_aids if aid in top_40_use]))\n    w2 = weights5* int(len(aids2)//(MM))\n    for i,(aid,w) in enumerate(zip(aids2,w2)):\n        m = 0.25 + 0.75*(ln-(i//(MM)))/ln\n        aids_temp[aid] += w*m\n    MM = 5\n    top_40_use = top_40_less\n    aids2 = list(itertools.chain(*[top_40_use[aid][:MM] for aid in unique_aids if aid in top_40_use]))\n    w2 = weights5* int(len(aids2)//(MM))\n    for i,(aid,w) in enumerate(zip(aids2,w2)):\n        m = 0.25 + 0.75*(ln-(i//(MM)))/ln\n        aids_temp[aid] += w*m\n    MM = 5\n    top_40_use = top_40_more2\n    aids2 = list(itertools.chain(*[top_40_use[aid][:MM] for aid in unique_aids if aid in top_40_use]))\n    w2 = weights5* int(len(aids2)//(MM))\n    for i,(aid,w) in enumerate(zip(aids2,w2)):\n        m = 0.25 + 0.75*(ln-(i//(MM)))/ln\n        aids_temp[aid] += w*m\n    MM = 5\n    top_40_use = top_40_less2\n    aids2 = list(itertools.chain(*[top_40_use[aid][:MM] for aid in unique_aids if aid in top_40_use]))\n    w2 = weights5* int(len(aids2)//(MM))\n    for i,(aid,w) in enumerate(zip(aids2,w2)):\n        m = 0.25 + 0.75*(ln-(i//(MM)))/ln\n        aids_temp[aid] += w*m\n        \n    MM = 2\n    aids2 = list(itertools.chain(*[top_day[aid][:10*MM] for aid in unique_aids6 if aid in top_day]))\n    w2 = weights3* int(len(aids2)//(10*MM))        \n    for i,(aid,w) in enumerate(zip(aids2,w2)):\n        #m = 0.25 + 0.75*(ln-(i//(10*MM)))/ln\n        aids_temp[aid] += 1 #w*m  \n        \n    ln0 = len(click_aids)\n    aids4 = list(itertools.chain(*[top_20_test[aid][:10] for aid in click_aids if aid in top_20_test]))\n    w4 = weights* int(len(aids4)//(10))\n    for i,(aid,w) in enumerate(zip(aids4,w4)):\n        m = 0.25 + 0.75*(ln0-(i//(10)))/ln0\n        aids_temp[aid] += w*m\n    aids4 = list(itertools.chain(*[top_20_buy[aid][:10] for aid in click_aids if aid in top_20_buy]))\n    w4 = weights* int(len(aids4)//(10))\n    for i,(aid,w) in enumerate(zip(aids4,w4)):\n        m = 0.25 + 0.75*(ln0-(i//(10)))/ln0\n        aids_temp[aid] += w*m\n    for aid in click_aids:\n        aids_temp[aid] += 1\n        \n        \n    aids5 = list(itertools.chain(*[top_20c[aid][:55] for aid in unique_aids[:1] if aid in top_20c]))\n    w5 = weights2* int(len(aids5)//55)\n    for aid,w in zip(aids5,w5):\n        aids_temp[aid] += w\n        \n    if len(unique_aids)==1:\n        aids5 = list(itertools.chain(*[top_20_new2[aid][:20] for aid in unique_aids[-1:] if aid in top_20_new2]))\n        w5 = weights3* int(len(aids5)//20)\n        for aid,w in zip(aids5,w5):\n            aids_temp[aid] += w\n        aids5 = list(itertools.chain(*[top_20_new[aid][:20] for aid in unique_aids[-1:] if aid in top_20_new]))\n        w5 = weights3* int(len(aids5)//20)\n        for aid,w in zip(aids5,w5):\n            aids_temp[aid] += w\n        \n    aids5 = list(itertools.chain(*[top_20d[aid][:20] for aid in unique_buys[:1] if aid in top_20d]))\n    w5 = weights3* int(len(aids5)//20)\n    for aid,w in zip(aids5,w5):\n        aids_temp[aid] += w\n        \n    ln2 = len(unique_aids5)\n    aids5 = list(itertools.chain(*[top_20_buy[aid][:20] for aid in unique_aids5 if aid in top_20_buy]))\n    w5 = weights3* int(len(aids5)//20)\n    for aid,w in zip(aids5,w5):\n        aids_temp[aid] += 2*w/ln2\n        \n    aids4 = list(itertools.chain(*[top_20f[aid][:5] for aid in unique_aids4 if aid in top_20f]))\n    for i,aid in enumerate(aids4):\n        w = i//5\n        aids_temp[aid] += 1/2 -w*0.05\n        if i%5==0: aids_temp[aid] += 1/2 -w*0.05\n    aids5 = list(itertools.chain(*[top_20e[aid][:55] for aid in unique_aids3 if aid in top_20e]))\n    w5 = weights2* int(len(aids5)//55)\n    for i,(aid,w) in enumerate(zip(aids5,w5)):\n        w2 = i//55\n        aids_temp[aid] += w -w2*0.1\n    aids5 = list(itertools.chain(*[top_20e[aid][:10] for aid in unique_aids2 if aid in top_20e]))\n    w5 = weights* int(len(aids5)//10)\n    for i,(aid,w) in enumerate(zip(aids5,w5)):\n        w2 = i//10\n        aids_temp[aid] += w/2. -w2*0.05\n                    \n    sorted_aids = [k for k,v in aids_temp.most_common(ITEM_CT) if k not in unique_aids]\n    \n    result = unique_aids + sorted_aids[:ITEM_CT - len(unique_aids)]\n    return session, (result + top_orders[:ITEM_CT-len(result)])[:ITEM_CT]","metadata":{"papermill":{"duration":0.019192,"end_time":"2022-11-03T16:50:23.026501","exception":false,"start_time":"2022-11-03T16:50:23.007309","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-02-01T22:53:19.732774Z","iopub.execute_input":"2023-02-01T22:53:19.733106Z","iopub.status.idle":"2023-02-01T22:53:19.841391Z","shell.execute_reply.started":"2023-02-01T22:53:19.733078Z","shell.execute_reply":"2023-02-01T22:53:19.840449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Use Parallel Processing\nThe following fast code to run my `suggest clicks` and `suggest carts/orders` quickly is from Aldparis notebook [here][1]\n\n[1]: https://www.kaggle.com/code/adaubas/otto-fast-handcrafted-model-recall-20","metadata":{}},{"cell_type":"code","source":"import psutil\nN_CORES = psutil.cpu_count()     \nprint(f\"N Cores : {N_CORES}\")\nfrom multiprocessing import Pool","metadata":{"execution":{"iopub.status.busy":"2023-02-01T22:53:19.844631Z","iopub.execute_input":"2023-02-01T22:53:19.844956Z","iopub.status.idle":"2023-02-01T22:53:19.856634Z","shell.execute_reply.started":"2023-02-01T22:53:19.844926Z","shell.execute_reply":"2023-02-01T22:53:19.855607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"N_CORES = N_CORES//2\ndef df_parallelize_run(func, t_split):\n    \n    num_cores = np.min([N_CORES, len(t_split)])\n    pool = Pool(num_cores)\n    df = pool.map(func, t_split)\n    pool.close()\n    pool.join()\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2023-02-01T22:53:19.857681Z","iopub.execute_input":"2023-02-01T22:53:19.857931Z","iopub.status.idle":"2023-02-01T22:53:19.864892Z","shell.execute_reply.started":"2023-02-01T22:53:19.857908Z","shell.execute_reply":"2023-02-01T22:53:19.863939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load Kaggle's Test Data\nWe converted Kaggle's test data into lists for each session. Code by Aldparis showing how to do this is [here][1]\n\n[1]: https://www.kaggle.com/code/adaubas/otto-prepare-valid-test","metadata":{}},{"cell_type":"code","source":"%%time\nPIECES = 10\ntest_bysession_list = []\nfor PART in range(PIECES):\n    with open(f'{PATH}test_group_tolist_{PART}_1.pkl', 'rb') as f:\n        test_bysession_list.extend(pickle.load(f))\nprint(len(test_bysession_list))","metadata":{"execution":{"iopub.status.busy":"2023-02-01T22:53:19.866017Z","iopub.execute_input":"2023-02-01T22:53:19.866273Z","iopub.status.idle":"2023-02-01T22:54:34.369930Z","shell.execute_reply.started":"2023-02-01T22:53:19.866250Z","shell.execute_reply":"2023-02-01T22:54:34.369096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Suggest Clicks","metadata":{}},{"cell_type":"code","source":"%%time\ntemp = df_parallelize_run(suggest_aids, test_bysession_list)\nval_clicks = pd.Series([f[1]  for f in temp], index=[f[0] for f in temp])","metadata":{"execution":{"iopub.status.busy":"2023-02-01T22:54:34.371218Z","iopub.execute_input":"2023-02-01T22:54:34.371553Z","iopub.status.idle":"2023-02-01T22:56:20.602515Z","shell.execute_reply.started":"2023-02-01T22:54:34.371523Z","shell.execute_reply":"2023-02-01T22:56:20.601349Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load More Covisit Matrices\nThese covisit matrices were created using GPU RAPIDS cuDF in another notebook and described in the intro section.","metadata":{}},{"cell_type":"code","source":"%%time\ndel top_20\nclick_df = val_clicks.to_dict()\ntop_20_orders = pickle.load(open(PATH+'top_40_orders_carts_v118.pkl', 'rb')) \ntop_20_carts = top_20_orders \ntop_20_buy2buy = pickle.load(open(PATH+'top_40_buy2buy_v16.pkl', 'rb'))    \ntop_20_buy2buy2 = pickle.load(open(PATH+'top_40_buy2buy_v119.pkl', 'rb')) \ntop_20_test = pickle.load(open(PATH+'top_40_aids_v35.pkl', 'rb'))\ntop_20_buy = pickle.load(open(PATH+'top_20_aids_v37.pkl', 'rb')) \ntop_20_new = pickle.load(open(PATH+'top_40_aids_v800_0.pkl', 'rb'))  \n\ntop_40_less = pickle.load(open(PATH+'top_40_aids_v178_0.pkl', 'rb'))   \ntop_40_more = pickle.load(open(PATH+'top_40_aids_v179_0.pkl', 'rb'))\ntop_40_less2 = pickle.load(open(PATH+'top_40_aids_v180_0.pkl', 'rb'))   \ntop_40_more2 = pickle.load(open(PATH+'top_40_aids_v181_0.pkl', 'rb'))","metadata":{"execution":{"iopub.status.busy":"2023-02-01T22:56:50.458438Z","iopub.execute_input":"2023-02-01T22:56:50.458761Z","iopub.status.idle":"2023-02-01T22:59:30.216579Z","shell.execute_reply.started":"2023-02-01T22:56:50.458736Z","shell.execute_reply":"2023-02-01T22:59:30.215256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Suggest Carts/Orders","metadata":{}},{"cell_type":"code","source":"%%time\ntemp = df_parallelize_run(suggest_orders, test_bysession_list)\nval_buys = pd.Series([f[1]  for f in temp], index=[f[0] for f in temp])","metadata":{"execution":{"iopub.status.busy":"2023-02-01T23:00:00.673608Z","iopub.execute_input":"2023-02-01T23:00:00.673917Z","iopub.status.idle":"2023-02-01T23:01:48.058981Z","shell.execute_reply.started":"2023-02-01T23:00:00.673891Z","shell.execute_reply":"2023-02-01T23:01:48.057834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create Submission File","metadata":{}},{"cell_type":"code","source":"clicks_pred_df = pd.DataFrame(val_clicks.add_suffix(\"_clicks\"), columns=[\"labels\"]).reset_index()\norders_pred_df = pd.DataFrame(val_buys.add_suffix(\"_orders\"), columns=[\"labels\"]).reset_index()\ncarts_pred_df = pd.DataFrame(val_buys.add_suffix(\"_carts\"), columns=[\"labels\"]).reset_index()","metadata":{"execution":{"iopub.status.busy":"2023-02-01T23:01:48.060658Z","iopub.execute_input":"2023-02-01T23:01:48.060961Z","iopub.status.idle":"2023-02-01T23:01:52.991611Z","shell.execute_reply.started":"2023-02-01T23:01:48.060932Z","shell.execute_reply":"2023-02-01T23:01:52.990698Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\npred_df = pd.concat(\n    [clicks_pred_df, orders_pred_df, carts_pred_df]\n)\npred_df.columns = [\"session_type\", \"labels\"]\npred_df[\"labels\"] = pred_df.labels.apply(lambda x: \" \".join(map(str,x)))\npred_df.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2023-02-01T23:01:52.992803Z","iopub.execute_input":"2023-02-01T23:01:52.993087Z","iopub.status.idle":"2023-02-01T23:03:33.858121Z","shell.execute_reply.started":"2023-02-01T23:01:52.993062Z","shell.execute_reply":"2023-02-01T23:03:33.857104Z"},"trusted":true},"execution_count":null,"outputs":[]}]}