{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"papermill":{"default_parameters":{},"duration":2157.914579,"end_time":"2022-11-16T20:47:30.276036","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2022-11-16T20:11:32.361457","version":"2.3.4"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":38760,"databundleVersionId":4493939,"sourceType":"competition"},{"sourceId":4436180,"sourceType":"datasetVersion","datasetId":2597726}],"dockerImageVersionId":30823,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"845f91ce","cell_type":"markdown","source":"# Candidate ReRank Model using Handcrafted Rules\nIn this notebook, we present a \"candidate rerank\" model using handcrafted rules. We can improve this model by engineering features, merging them unto items and users, and training a reranker model (such as XGB) to choose our final 20. Furthermore to tune and improve this notebook, we should build a local CV scheme to experiment new logic and/or models.\n\nUPDATE: I published a notebook to compute validation score [here][10] using Radek's scheme described [here][11].\n\nNote in this competition, a \"session\" actually means a unique \"user\". So our task is to predict what each of the `1,671,803` test \"users\" (i.e. \"sessions\") will do in the future. For each test \"user\" (i.e. \"session\") we must predict what they will `click`, `cart`, and `order` during the remainder of the week long test period.\n\n### Step 1 - Generate Candidates\nFor each test user, we generate possible choices, i.e. candidates. In this notebook, we generate candidates from 5 sources:\n* User history of clicks, carts, orders\n* Most popular 20 clicks, carts, orders during test week\n* Co-visitation matrix of click/cart/order to cart/order with type weighting\n* Co-visitation matrix of cart/order to cart/order called buy2buy\n* Co-visitation matrix of click/cart/order to clicks with time weighting\n\n### Step 2 - ReRank and Choose 20\nGiven the list of candidates, we must select 20 to be our predictions. In this notebook, we do this with a set of handcrafted rules. We can improve our predictions by training an XGBoost model to select for us. Our handcrafted rules give priority to:\n* Most recent previously visited items\n* Items previously visited multiple times\n* Items previously in cart or order\n* Co-visitation matrix of cart/order to cart/order\n* Current popular items\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/c_r_model.png)\n  \n# Credits\nWe thank many Kagglers who have shared ideas. We use co-visitation matrix idea from Vladimir [here][1]. We use groupby sort logic from Sinan in comment section [here][4]. We use duplicate prediction removal logic from Radek [here][5]. We use multiple visit logic from Pietro [here][2]. We use type weighting logic from Ingvaras [here][3]. We use leaky test data from my previous notebook [here][4]. And some ideas may have originated from Tawara [here][6] and KJ [here][7]. We use Colum2131's parquets [here][8]. Above image is from Ravi's discussion about candidate rerank models [here][9]\n\n[1]: https://www.kaggle.com/code/vslaykovsky/co-visitation-matrix\n[2]: https://www.kaggle.com/code/pietromaldini1/multiple-clicks-vs-latest-items\n[3]: https://www.kaggle.com/code/ingvarasgalinskas/item-type-vs-multiple-clicks-vs-latest-items\n[4]: https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost\n[5]: https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\n[6]: https://www.kaggle.com/code/ttahara/otto-mors-aid-frequency-baseline\n[7]: https://www.kaggle.com/code/whitelily/co-occurrence-baseline\n[8]: https://www.kaggle.com/datasets/columbia2131/otto-chunk-data-inparquet-format\n[9]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\n[10]: https://www.kaggle.com/cdeotte/compute-validation-score-cv-564\n[11]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991","metadata":{"papermill":{"duration":0.005621,"end_time":"2022-11-16T20:11:40.088445","exception":false,"start_time":"2022-11-16T20:11:40.082824","status":"completed"},"tags":[]}},{"id":"ed81da6c","cell_type":"markdown","source":"# Notes\nBelow are notes about versions:\n* **Version 1 LB 0.573** Uses popular ideas from public notebooks and adds additional co-visitation matrices and additional logic. Has CV `0.563`. See validation notebook version 2 [here][1].\n* **Version 2 LB 573** Refactor logic for `suggest_buys(df)` to make it clear how new co-visitation matrices are reranking the candidates by adding to candidate weights. Also new logic boosts CV by `+0.0003`. Also LB is slightly better too. See validation notebook version 3 [here][1]\n* **Version 3** is the same as version 2 but 1.5x faster co-visitation matrix computation!\n* **Version 4 LB 575** Use top20 for clicks and top15 for carts and buys (instead of top40 and top40). This boosts CV `+0.0015` hooray! New CV is `0.5647`. See validation version 5 [here][1]\n* **Version 5** is the same as version 4 but 2x faster co-visitation matrix computation! (and 3x faster than version 1)\n* **Version 6** Stay tuned for more versions...\n\n[1]: https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-564","metadata":{"papermill":{"duration":0.004,"end_time":"2022-11-16T20:11:40.096933","exception":false,"start_time":"2022-11-16T20:11:40.092933","status":"completed"},"tags":[]}},{"id":"a5804811","cell_type":"markdown","source":"# Step 1 - Candidate Generation with RAPIDS\nFor candidate generation, we build three co-visitation matrices. One computes the popularity of cart/order given a user's previous click/cart/order. We apply type weighting to this matrix. One computes the popularity of cart/order given a user's previous cart/order. We call this \"buy2buy\" matrix. One computes the popularity of clicks given a user previously click/cart/order.  We apply time weighting to this matrix. We will use RAPIDS cuDF GPU to compute these matrices quickly!","metadata":{"papermill":{"duration":0.004049,"end_time":"2022-11-16T20:11:40.105223","exception":false,"start_time":"2022-11-16T20:11:40.101174","status":"completed"},"tags":[]}},{"id":"ec630cb1","cell_type":"code","source":"VER = 5\n\nimport pandas as pd, numpy as np\nfrom tqdm.notebook import tqdm\nimport os, sys, pickle, glob, gc\nfrom collections import Counter\nimport cudf, itertools\nprint('We will use RAPIDS version',cudf.__version__)","metadata":{"papermill":{"duration":2.742174,"end_time":"2022-11-16T20:11:42.851593","exception":false,"start_time":"2022-11-16T20:11:40.109419","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"eaef1cad","cell_type":"markdown","source":"## Compute Three Co-visitation Matrices with RAPIDS\nWe will compute 3 co-visitation matrices using RAPIDS cuDF on GPU. This is 30x faster than using Pandas CPU like other public notebooks! For maximum speed, set the variable `DISK_PIECES` to the smallest number possible based on the GPU you are using without incurring memory errors. If you run this code offline with 32GB GPU ram, then you can use `DISK_PIECES = 1` and compute each co-visitation matrix in almost 1 minute! Kaggle's GPU only has 16GB ram, so we use `DISK_PIECES = 4` and it takes an amazing 3 minutes each! Below are some of the tricks to speed up computation\n* Use RAPIDS cuDF GPU instead of Pandas CPU\n* Read disk once and save in CPU RAM for later GPU multiple use\n* Process largest amount of data possible on GPU at one time\n* Merge data in two stages. Multiple small to single medium. Multiple medium to single large.\n* Write result as parquet instead of dictionary","metadata":{"papermill":{"duration":0.004177,"end_time":"2022-11-16T20:11:42.861237","exception":false,"start_time":"2022-11-16T20:11:42.857060","status":"completed"},"tags":[]}},{"id":"872cc3f0","cell_type":"code","source":"%%time\n# CACHE FUNCTIONS\ndef read_file(f):\n    return cudf.DataFrame( data_cache[f] )\ndef read_file_to_cache(f):\n    df = pd.read_parquet(f)\n    df.ts = (df.ts/1000).astype('int32')\n    df['type'] = df['type'].map(type_labels).astype('int8')\n    return df\n\n# CACHE THE DATA ON CPU BEFORE PROCESSING ON GPU\ndata_cache = {}\ntype_labels = {'clicks':0, 'carts':1, 'orders':2}\nfiles = glob.glob('../input/otto-chunk-data-inparquet-format/*_parquet/*')\nfor f in files: data_cache[f] = read_file_to_cache(f)\n\n# CHUNK PARAMETERS\nREAD_CT = 5\nCHUNK = int( np.ceil( len(files)/6 ))\nprint(f'We will process {len(files)} files, in groups of {READ_CT} and chunks of {CHUNK}.')","metadata":{"_kg_hide-input":true,"papermill":{"duration":67.050747,"end_time":"2022-11-16T20:12:49.916304","exception":false,"start_time":"2022-11-16T20:11:42.865557","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"71ca004b","cell_type":"markdown","source":"## 1) \"Carts Orders\" Co-visitation Matrix - Type Weighted","metadata":{"papermill":{"duration":0.004383,"end_time":"2022-11-16T20:12:49.925362","exception":false,"start_time":"2022-11-16T20:12:49.920979","status":"completed"},"tags":[]}},{"id":"a425ec8b","cell_type":"code","source":"%%time\ntype_weight = {0:1, 1:6, 2:3}\n\n# USE SMALLEST DISK_PIECES POSSIBLE WITHOUT MEMORY ERROR\nDISK_PIECES = 4\nSIZE = 1.86e6/DISK_PIECES\n\n# COMPUTE IN PARTS FOR MEMORY MANGEMENT\nfor PART in range(DISK_PIECES):\n    print()\n    print('### DISK PART',PART+1)\n    \n    # MERGE IS FASTEST PROCESSING CHUNKS WITHIN CHUNKS\n    # => OUTER CHUNKS\n    for j in range(6):\n        a = j*CHUNK\n        b = min( (j+1)*CHUNK, len(files) )\n        print(f'Processing files {a} thru {b-1} in groups of {READ_CT}...')\n        \n        # => INNER CHUNKS\n        for k in range(a,b,READ_CT):\n            # READ FILE\n            df = [read_file(files[k])]\n            for i in range(1,READ_CT): \n                if k+i<b: df.append( read_file(files[k+i]) )\n            df = cudf.concat(df,ignore_index=True,axis=0)\n            df = df.sort_values(['session','ts'],ascending=[True,False])\n            # USE TAIL OF SESSION\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n<30].drop('n',axis=1)\n            # CREATE PAIRS\n            df = df.merge(df,on='session')\n            df = df.loc[ ((df.ts_x - df.ts_y).abs()< 24 * 60 * 60) & (df.aid_x != df.aid_y) ]\n            # MEMORY MANAGEMENT COMPUTE IN PARTS\n            df = df.loc[(df.aid_x >= PART*SIZE)&(df.aid_x < (PART+1)*SIZE)]\n            # ASSIGN WEIGHTS\n            df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n            df['wgt'] = df.type_y.map(type_weight)\n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n            # COMBINE INNER CHUNKS\n            if k==a: tmp2 = df\n            else: tmp2 = tmp2.add(df, fill_value=0)\n            print(k,', ',end='')\n        print()\n        # COMBINE OUTER CHUNKS\n        if a==0: tmp = tmp2\n        else: tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    # CONVERT MATRIX TO DICTIONARY\n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n    # SAVE TOP 40\n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    tmp = tmp.loc[tmp.n<15].drop('n',axis=1)\n    # SAVE PART TO DISK (convert to pandas first uses less memory)\n    tmp.to_pandas().to_parquet(f'top_15_carts_orders_v{VER}_{PART}.pqt')","metadata":{"papermill":{"duration":198.636943,"end_time":"2022-11-16T20:16:08.566733","exception":false,"start_time":"2022-11-16T20:12:49.929790","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"bc3be050","cell_type":"markdown","source":"## 2) \"Buy2Buy\" Co-visitation Matrix","metadata":{"papermill":{"duration":0.010798,"end_time":"2022-11-16T20:16:08.588901","exception":false,"start_time":"2022-11-16T20:16:08.578103","status":"completed"},"tags":[]}},{"id":"3ef6857b","cell_type":"code","source":"%%time\n# USE SMALLEST DISK_PIECES POSSIBLE WITHOUT MEMORY ERROR\nDISK_PIECES = 1\nSIZE = 1.86e6/DISK_PIECES\n\n# COMPUTE IN PARTS FOR MEMORY MANGEMENT\nfor PART in range(DISK_PIECES):\n    print()\n    print('### DISK PART',PART+1)\n    \n    # MERGE IS FASTEST PROCESSING CHUNKS WITHIN CHUNKS\n    # => OUTER CHUNKS\n    for j in range(6):\n        a = j*CHUNK\n        b = min( (j+1)*CHUNK, len(files) )\n        print(f'Processing files {a} thru {b-1} in groups of {READ_CT}...')\n        \n        # => INNER CHUNKS\n        for k in range(a,b,READ_CT):\n            # READ FILE\n            df = [read_file(files[k])]\n            for i in range(1,READ_CT): \n                if k+i<b: df.append( read_file(files[k+i]) )\n            df = cudf.concat(df,ignore_index=True,axis=0)\n            df = df.loc[df['type'].isin([1,2])] # ONLY WANT CARTS AND ORDERS\n            df = df.sort_values(['session','ts'],ascending=[True,False])\n            # USE TAIL OF SESSION\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n<30].drop('n',axis=1)\n            # CREATE PAIRS\n            df = df.merge(df,on='session')\n            df = df.loc[ ((df.ts_x - df.ts_y).abs()< 14 * 24 * 60 * 60) & (df.aid_x != df.aid_y) ] # 14 DAYS\n            # MEMORY MANAGEMENT COMPUTE IN PARTS\n            df = df.loc[(df.aid_x >= PART*SIZE)&(df.aid_x < (PART+1)*SIZE)]\n            # ASSIGN WEIGHTS\n            df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n            df['wgt'] = 1\n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n            # COMBINE INNER CHUNKS\n            if k==a: tmp2 = df\n            else: tmp2 = tmp2.add(df, fill_value=0)\n            print(k,', ',end='')\n        print()\n        # COMBINE OUTER CHUNKS\n        if a==0: tmp = tmp2\n        else: tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    # CONVERT MATRIX TO DICTIONARY\n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n    # SAVE TOP 40\n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    tmp = tmp.loc[tmp.n<15].drop('n',axis=1)\n    # SAVE PART TO DISK (convert to pandas first uses less memory)\n    tmp.to_pandas().to_parquet(f'top_15_buy2buy_v{VER}_{PART}.pqt')","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"papermill":{"duration":30.590555,"end_time":"2022-11-16T20:16:39.190470","exception":false,"start_time":"2022-11-16T20:16:08.599915","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"0226064c","cell_type":"markdown","source":"## 3) \"Clicks\" Co-visitation Matrix - Time Weighted","metadata":{"papermill":{"duration":0.019951,"end_time":"2022-11-16T20:16:39.231177","exception":false,"start_time":"2022-11-16T20:16:39.211226","status":"completed"},"tags":[]}},{"id":"f02f777a","cell_type":"code","source":"%%time\n# USE SMALLEST DISK_PIECES POSSIBLE WITHOUT MEMORY ERROR\nDISK_PIECES = 4\nSIZE = 1.86e6/DISK_PIECES\n\n# COMPUTE IN PARTS FOR MEMORY MANGEMENT\nfor PART in range(DISK_PIECES):\n    print()\n    print('### DISK PART',PART+1)\n    \n    # MERGE IS FASTEST PROCESSING CHUNKS WITHIN CHUNKS\n    # => OUTER CHUNKS\n    for j in range(6):\n        a = j*CHUNK\n        b = min( (j+1)*CHUNK, len(files) )\n        print(f'Processing files {a} thru {b-1} in groups of {READ_CT}...')\n        \n        # => INNER CHUNKS\n        for k in range(a,b,READ_CT):\n            # READ FILE\n            df = [read_file(files[k])]\n            for i in range(1,READ_CT): \n                if k+i<b: df.append( read_file(files[k+i]) )\n            df = cudf.concat(df,ignore_index=True,axis=0)\n            df = df.sort_values(['session','ts'],ascending=[True,False])\n            # USE TAIL OF SESSION\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n<30].drop('n',axis=1)\n            # CREATE PAIRS\n            df = df.merge(df,on='session')\n            df = df.loc[ ((df.ts_x - df.ts_y).abs()< 24 * 60 * 60) & (df.aid_x != df.aid_y) ]\n            # MEMORY MANAGEMENT COMPUTE IN PARTS\n            df = df.loc[(df.aid_x >= PART*SIZE)&(df.aid_x < (PART+1)*SIZE)]\n            # ASSIGN WEIGHTS\n            df = df[['session', 'aid_x', 'aid_y','ts_x']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n            df['wgt'] = 1 + 3*(df.ts_x - 1659304800)/(1662328791-1659304800)\n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n            # COMBINE INNER CHUNKS\n            if k==a: tmp2 = df\n            else: tmp2 = tmp2.add(df, fill_value=0)\n            print(k,', ',end='')\n        print()\n        # COMBINE OUTER CHUNKS\n        if a==0: tmp = tmp2\n        else: tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    # CONVERT MATRIX TO DICTIONARY\n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n    # SAVE TOP 40\n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    tmp = tmp.loc[tmp.n<20].drop('n',axis=1)\n    # SAVE PART TO DISK (convert to pandas first uses less memory)\n    tmp.to_pandas().to_parquet(f'top_20_clicks_v{VER}_{PART}.pqt')","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"papermill":{"duration":195.048085,"end_time":"2022-11-16T20:19:54.341238","exception":false,"start_time":"2022-11-16T20:16:39.293153","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"c5d77570","cell_type":"code","source":"# FREE MEMORY\ndel data_cache, tmp\n_ = gc.collect()","metadata":{"papermill":{"duration":0.179767,"end_time":"2022-11-16T20:19:54.540593","exception":false,"start_time":"2022-11-16T20:19:54.360826","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"2c9d36d3","cell_type":"markdown","source":"# Step 2 - ReRank (choose 20) using handcrafted rules\nFor description of the handcrafted rules, read this notebook's intro.","metadata":{"papermill":{"duration":0.019537,"end_time":"2022-11-16T20:19:54.580031","exception":false,"start_time":"2022-11-16T20:19:54.560494","status":"completed"},"tags":[]}},{"id":"41132bc2","cell_type":"code","source":"def load_test():    \n    dfs = []\n    for e, chunk_file in enumerate(glob.glob('../input/otto-chunk-data-inparquet-format/test_parquet/*')):\n        chunk = pd.read_parquet(chunk_file)\n        chunk.ts = (chunk.ts/1000).astype('int32')\n        chunk['type'] = chunk['type'].map(type_labels).astype('int8')\n        dfs.append(chunk)\n    return pd.concat(dfs).reset_index(drop=True) #.astype({\"ts\": \"datetime64[ms]\"})\n\ntest_df = load_test()\nprint('Test data has shape',test_df.shape)\ntest_df.head()","metadata":{"papermill":{"duration":1.280046,"end_time":"2022-11-16T20:19:55.879501","exception":false,"start_time":"2022-11-16T20:19:54.599455","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"0b675630","cell_type":"code","source":"%%time\ndef pqt_to_dict(df):\n    return df.groupby('aid_x').aid_y.apply(list).to_dict()\n# LOAD THREE CO-VISITATION MATRICES\ntop_20_clicks = pqt_to_dict( pd.read_parquet(f'top_20_clicks_v{VER}_0.pqt') )\nfor k in range(1,DISK_PIECES): \n    top_20_clicks.update( pqt_to_dict( pd.read_parquet(f'top_20_clicks_v{VER}_{k}.pqt') ) )\ntop_20_buys = pqt_to_dict( pd.read_parquet(f'top_15_carts_orders_v{VER}_0.pqt') )\nfor k in range(1,DISK_PIECES): \n    top_20_buys.update( pqt_to_dict( pd.read_parquet(f'top_15_carts_orders_v{VER}_{k}.pqt') ) )\ntop_20_buy2buy = pqt_to_dict( pd.read_parquet(f'top_15_buy2buy_v{VER}_0.pqt') )\n\n# TOP CLICKS AND ORDERS IN TEST\ntop_clicks = test_df.loc[test_df['type']=='clicks','aid'].value_counts().index.values[:20]\ntop_orders = test_df.loc[test_df['type']=='orders','aid'].value_counts().index.values[:20]\n\nprint('Here are size of our 3 co-visitation matrices:')\nprint( len( top_20_clicks ), len( top_20_buy2buy ), len( top_20_buys ) )","metadata":{"papermill":{"duration":97.397608,"end_time":"2022-11-16T20:21:33.297915","exception":false,"start_time":"2022-11-16T20:19:55.900307","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"b107fa16","cell_type":"code","source":"#type_weight_multipliers = {'clicks': 1, 'carts': 6, 'orders': 3}\ntype_weight_multipliers = {0: 1, 1: 6, 2: 3}\n\ndef suggest_clicks(df):\n    # USER HISTORY AIDS AND TYPES\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    # RERANK CANDIDATES USING WEIGHTS\n    if len(unique_aids)>=20:\n        weights=np.logspace(0.1,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        # RERANK BASED ON REPEAT ITEMS AND TYPE OF ITEMS\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    # USE \"CLICKS\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(20) if aid2 not in unique_aids]    \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    # USE TOP20 TEST CLICKS\n    return result + list(top_clicks)[:20-len(result)]\n\ndef suggest_buys(df):\n    # USER HISTORY AIDS AND TYPES\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    # UNIQUE AIDS AND UNIQUE BUYS\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    df = df.loc[(df['type']==1)|(df['type']==2)]\n    unique_buys = list(dict.fromkeys( df.aid.tolist()[::-1] ))\n    # RERANK CANDIDATES USING WEIGHTS\n    if len(unique_aids)>=20:\n        weights=np.logspace(0.5,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        # RERANK BASED ON REPEAT ITEMS AND TYPE OF ITEMS\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        # RERANK CANDIDATES USING \"BUY2BUY\" CO-VISITATION MATRIX\n        aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n        for aid in aids3: aids_temp[aid] += 0.1\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    # USE \"CART ORDER\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_aids if aid in top_20_buys]))\n    # USE \"BUY2BUY\" CO-VISITATION MATRIX\n    aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2+aids3).most_common(20) if aid2 not in unique_aids] \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    # USE TOP20 TEST ORDERS\n    return result + list(top_orders)[:20-len(result)]","metadata":{"papermill":{"duration":0.069514,"end_time":"2022-11-16T20:21:33.387033","exception":false,"start_time":"2022-11-16T20:21:33.317519","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"dc4879e8","cell_type":"markdown","source":"# Create Submission CSV\nInferring test data with Pandas groupby is slow. We need to accelerate the following code.","metadata":{"papermill":{"duration":0.02308,"end_time":"2022-11-16T20:21:33.436050","exception":false,"start_time":"2022-11-16T20:21:33.412970","status":"completed"},"tags":[]}},{"id":"2e7d7d65","cell_type":"code","source":"%%time\npred_df_clicks = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(\n    lambda x: suggest_clicks(x)\n)\n\npred_df_buys = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(\n    lambda x: suggest_buys(x)\n)","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","papermill":{"duration":1512.63605,"end_time":"2022-11-16T20:46:46.091490","exception":false,"start_time":"2022-11-16T20:21:33.455440","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"8b93a757","cell_type":"code","source":"clicks_pred_df = pd.DataFrame(pred_df_clicks.add_suffix(\"_clicks\"), columns=[\"labels\"]).reset_index()\norders_pred_df = pd.DataFrame(pred_df_buys.add_suffix(\"_orders\"), columns=[\"labels\"]).reset_index()\ncarts_pred_df = pd.DataFrame(pred_df_buys.add_suffix(\"_carts\"), columns=[\"labels\"]).reset_index()","metadata":{"papermill":{"duration":2.987022,"end_time":"2022-11-16T20:46:49.098252","exception":false,"start_time":"2022-11-16T20:46:46.111230","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"a5c87606","cell_type":"code","source":"pred_df = pd.concat([clicks_pred_df, orders_pred_df, carts_pred_df])\npred_df.columns = [\"session_type\", \"labels\"]\npred_df[\"labels\"] = pred_df.labels.apply(lambda x: \" \".join(map(str,x)))\npred_df.to_csv(\"submission.csv\", index=False)\npred_df.head()","metadata":{"papermill":{"duration":37.989649,"end_time":"2022-11-16T20:47:27.107993","exception":false,"start_time":"2022-11-16T20:46:49.118344","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null},{"id":"9bd4c95b-1157-4cff-b228-9bde82205576","cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"4cd840de-4a9d-48ff-a85f-2b84e1cd7514","cell_type":"markdown","source":"# New methods","metadata":{}},{"id":"7cd0e32e-9a66-4310-9c5a-07586bd5b33e","cell_type":"markdown","source":"## Imports, Configuration, and Data Caching","metadata":{}},{"id":"c8cb4334-2602-477b-99de-47b8eaa7ee89","cell_type":"code","source":"import pandas as pd, numpy as np\nfrom tqdm.notebook import tqdm\nimport os, sys, pickle, glob, gc\nfrom collections import Counter\nimport cudf, itertools\n\nprint('We will use RAPIDS version', cudf.__version__)\n\n# Define Version\nVER = 5\n\n# Read/Cache functions\ndef read_file_to_cache(f, type_labels):\n    df = pd.read_parquet(f)\n    df.ts = (df.ts/1000).astype('int32')\n    df['type'] = df['type'].map(type_labels).astype('int8')\n    return df\n\ndef read_file(f, data_cache):\n    return cudf.DataFrame(data_cache[f])\n\n# Prepare to cache data on CPU\ndata_cache = {}\ntype_labels = {'clicks':0, 'carts':1, 'orders':2}\nfiles = glob.glob('../input/otto-chunk-data-inparquet-format/*_parquet/*')\n\nfor f in files:\n    data_cache[f] = read_file_to_cache(f, type_labels)\n\nprint(f\"Total files found: {len(files)}\")\n\n# Additional parameters\nREAD_CT = 5\nCHUNK = int(np.ceil(len(files)/6))\nprint(f\"We will process {len(files)} files, in groups of {READ_CT} and chunks of {CHUNK}.\")\n","metadata":{"trusted":true,"scrolled":true,"execution":{"iopub.status.busy":"2024-12-25T21:08:10.626837Z","iopub.execute_input":"2024-12-25T21:08:10.627155Z","iopub.status.idle":"2024-12-25T21:08:38.600709Z","shell.execute_reply.started":"2024-12-25T21:08:10.627126Z","shell.execute_reply":"2024-12-25T21:08:38.599761Z"}},"outputs":[],"execution_count":null},{"id":"c66a8930-dc53-4d78-a7fd-cf9f646f0b5a","cell_type":"markdown","source":"## Build “Carts & Orders” Co-visitation Matrix (Type-Weighted)","metadata":{}},{"id":"0e1d1f5d-c944-4493-9e07-a8ac8f1e6c13","cell_type":"code","source":"%%time\ntype_weight = {0:1, 1:6, 2:3}\n\nDISK_PIECES = 4\nSIZE = 1.86e6 / DISK_PIECES\n\nfor PART in range(DISK_PIECES):\n    print(f\"\\n### DISK PART {PART+1}\")\n    \n    for j in range(6):\n        a = j*CHUNK\n        b = min((j+1)*CHUNK, len(files))\n        print(f'Processing files {a} thru {b-1} in groups of {READ_CT}...')\n        \n        for k in range(a, b, READ_CT):\n            # Read chunk of files\n            df = [read_file(files[k], data_cache)]\n            for i in range(1, READ_CT):\n                if k+i < b:\n                    df.append(read_file(files[k+i], data_cache))\n            df = cudf.concat(df, ignore_index=True, axis=0)\n            \n            # Sort and take tail of session\n            df = df.sort_values(['session','ts'], ascending=[True,False])\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n < 30].drop('n', axis=1)\n            \n            # Create pairs\n            df = df.merge(df, on='session')\n            df = df.loc[((df.ts_x - df.ts_y).abs() < 24*60*60) & (df.aid_x != df.aid_y)]\n            \n            # Memory management: part-specific\n            df = df.loc[(df.aid_x >= PART*SIZE) & (df.aid_x < (PART+1)*SIZE)]\n            \n            # Assign weights\n            df = df[['session', 'aid_x', 'aid_y', 'type_y']].drop_duplicates(['session','aid_x','aid_y'])\n            df['wgt'] = df.type_y.map(type_weight)\n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            \n            # Sum up\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n            if k == a:\n                tmp2 = df\n            else:\n                tmp2 = tmp2.add(df, fill_value=0)\n            print(k, ', ', end='')\n        \n        print()\n        if j == 0:\n            tmp = tmp2\n        else:\n            tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    \n    # Convert matrix to dictionary format\n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'], ascending=[True,False])\n    \n    # Keep top 15\n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    tmp = tmp.loc[tmp.n < 15].drop('n', axis=1)\n    \n    # Save part to disk\n    tmp.to_pandas().to_parquet(f'top_15_carts_orders_v{VER}_{PART}.pqt')\n","metadata":{"trusted":true,"scrolled":true,"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"id":"4a529b05-b51b-4d91-8607-210cc5a9531c","cell_type":"markdown","source":"## Build “Buy2Buy” and “Clicks” Co-visitation Matrices","metadata":{}},{"id":"ce0bafbc-a800-4af4-b266-51ab7ffa84bf","cell_type":"code","source":"%%time\n# Buy2Buy\nDISK_PIECES_B2B = 1\nSIZE_B2B = 1.86e6 / DISK_PIECES_B2B\n\nfor PART in range(DISK_PIECES_B2B):\n    print(f\"\\n### DISK PART Buy2Buy {PART+1}\")\n    for j in range(6):\n        a = j*CHUNK\n        b = min((j+1)*CHUNK, len(files))\n        print(f'Processing files {a} thru {b-1} in groups of {READ_CT}...')\n        \n        for k in range(a, b, READ_CT):\n            df = [read_file(files[k], data_cache)]\n            for i in range(1, READ_CT):\n                if k+i < b:\n                    df.append(read_file(files[k+i], data_cache))\n            df = cudf.concat(df, ignore_index=True, axis=0)\n            \n            # Only carts and orders\n            df = df.loc[df['type'].isin([1,2])]\n            df = df.sort_values(['session','ts'], ascending=[True,False])\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n < 30].drop('n', axis=1)\n            \n            # Merge pairs (14-day window)\n            df = df.merge(df, on='session')\n            df = df.loc[((df.ts_x - df.ts_y).abs() < 14*24*60*60) & (df.aid_x != df.aid_y)]\n            \n            # Part-specific\n            df = df.loc[(df.aid_x >= PART*SIZE_B2B) & (df.aid_x < (PART+1)*SIZE_B2B)]\n            \n            # Assign weight = 1\n            df = df[['session','aid_x','aid_y','type_y']].drop_duplicates(['session','aid_x','aid_y'])\n            df['wgt'] = 1\n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            \n            # Sum up\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n            if k == a:\n                tmp2 = df\n            else:\n                tmp2 = tmp2.add(df, fill_value=0)\n            print(k, ', ', end='')\n        \n        print()\n        if j == 0:\n            tmp = tmp2\n        else:\n            tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    \n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'], ascending=[True,False])\n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    tmp = tmp.loc[tmp.n < 15].drop('n', axis=1)\n    \n    tmp.to_pandas().to_parquet(f'top_15_buy2buy_v{VER}_{PART}.pqt')\n\n\n# Clicks (Time-Weighted)\nDISK_PIECES_CLICK = 4\nSIZE_CLICK = 1.86e6 / DISK_PIECES_CLICK\n\nfor PART in range(DISK_PIECES_CLICK):\n    print(f\"\\n### DISK PART Clicks {PART+1}\")\n    for j in range(6):\n        a = j*CHUNK\n        b = min((j+1)*CHUNK, len(files))\n        print(f'Processing files {a} thru {b-1} in groups of {READ_CT}...')\n        \n        for k in range(a, b, READ_CT):\n            df = [read_file(files[k], data_cache)]\n            for i in range(1, READ_CT):\n                if k+i < b:\n                    df.append(read_file(files[k+i], data_cache))\n            df = cudf.concat(df, ignore_index=True, axis=0)\n            \n            # Sort and take tail\n            df = df.sort_values(['session','ts'], ascending=[True,False])\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n < 30].drop('n', axis=1)\n            \n            # Create pairs\n            df = df.merge(df, on='session')\n            df = df.loc[((df.ts_x - df.ts_y).abs() < 24*60*60) & (df.aid_x != df.aid_y)]\n            \n            # Part-specific\n            df = df.loc[(df.aid_x >= PART*SIZE_CLICK) & (df.aid_x < (PART+1)*SIZE_CLICK)]\n            \n            # Time-based weights\n            df = df[['session','aid_x','aid_y','ts_x']].drop_duplicates(['session','aid_x','aid_y'])\n            df['wgt'] = 1 + 3*(df.ts_x - 1659304800)/(1662328791-1659304800)\n            df.wgt = df.wgt.astype('float32')\n            \n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n            if k == a:\n                tmp2 = df\n            else:\n                tmp2 = tmp2.add(df, fill_value=0)\n            print(k, ', ', end='')\n        \n        print()\n        if j == 0:\n            tmp = tmp2\n        else:\n            tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    \n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'], ascending=[True,False])\n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    tmp = tmp.loc[tmp.n < 20].drop('n', axis=1)\n    \n    tmp.to_pandas().to_parquet(f'top_20_clicks_v{VER}_{PART}.pqt')\n","metadata":{"trusted":true,"scrolled":true,"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"id":"2a3b53e1-f313-4a54-81a6-b594b578cf63","cell_type":"markdown","source":"## Re-Rank (Choose 20) and Make Predictions","metadata":{}},{"id":"55a1983b-9cc9-46dc-b7f3-ce17a966d4c4","cell_type":"code","source":"import glob\n\n# Helper function to load test data\ndef load_test(type_labels):\n    dfs = []\n    for chunk_file in glob.glob('../input/otto-chunk-data-inparquet-format/test_parquet/*'):\n        chunk = pd.read_parquet(chunk_file)\n        chunk.ts = (chunk.ts/1000).astype('int32')\n        chunk['type'] = chunk['type'].map(type_labels).astype('int8')\n        dfs.append(chunk)\n    return pd.concat(dfs).reset_index(drop=True)\n\ntest_df = load_test(type_labels)\nprint('Test data has shape', test_df.shape)\n\n# Dictionary conversion\ndef pqt_to_dict(df):\n    return df.groupby('aid_x').aid_y.apply(list).to_dict()\n\n# Load Co-visitation Matrices\ntop_20_clicks = {}\nDISK_PIECES_CLICK = 4\nfor k in range(DISK_PIECES_CLICK):\n    tmp_dict = pqt_to_dict(pd.read_parquet(f'top_20_clicks_v{VER}_{k}.pqt'))\n    top_20_clicks.update(tmp_dict)\n\ntop_20_buys = {}\nDISK_PIECES = 4\nfor k in range(DISK_PIECES):\n    tmp_dict = pqt_to_dict(pd.read_parquet(f'top_15_carts_orders_v{VER}_{k}.pqt'))\n    top_20_buys.update(tmp_dict)\n\ntop_20_buy2buy = {}\nDISK_PIECES_B2B = 1\nfor k in range(DISK_PIECES_B2B):\n    tmp_dict = pqt_to_dict(pd.read_parquet(f'top_15_buy2buy_v{VER}_{k}.pqt'))\n    top_20_buy2buy.update(tmp_dict)\n\n# Top items from test\ntop_clicks = test_df.loc[test_df['type']==0, 'aid'].value_counts().index.values[:20]\ntop_orders = test_df.loc[test_df['type']==2, 'aid'].value_counts().index.values[:20]\n\nprint('Sizes of co-visitation dictionaries:')\nprint(len(top_20_clicks), len(top_20_buy2buy), len(top_20_buys))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-25T21:08:38.601768Z","iopub.execute_input":"2024-12-25T21:08:38.602229Z"}},"outputs":[],"execution_count":null},{"id":"f8e0816f-cb3b-4291-8105-020533cf0e59","cell_type":"code","source":"# Weighted multipliers\ntype_weight_multipliers = {0:1, 1:6, 2:3}\n\nimport itertools\ndef suggest_clicks(df):\n    aids = df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1]))\n    \n    if len(unique_aids) >= 20:\n        weights = np.logspace(0.1,1,len(aids),base=2, endpoint=True) - 1\n        aids_temp = Counter()\n        for aid, w, t in zip(aids, weights, types):\n            aids_temp[aid] += w * type_weight_multipliers[t]\n        return [k for k, v in aids_temp.most_common(20)]\n    \n    # Extend with co-visitation\n    aids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(20) if aid2 not in unique_aids]\n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    # Fill with top clicks\n    return result + list(top_clicks)[:20-len(result)]\n\ndef suggest_buys(df):\n    aids = df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1]))\n    \n    # Only carts+orders from session\n    df_buys = df.loc[(df['type']==1)|(df['type']==2)]\n    unique_buys = list(dict.fromkeys(df_buys.aid.tolist()[::-1]))\n    \n    if len(unique_aids) >= 20:\n        weights = np.logspace(0.5,1,len(aids),base=2, endpoint=True) - 1\n        aids_temp = Counter()\n        for aid, w, t in zip(aids, weights, types):\n            aids_temp[aid] += w * type_weight_multipliers[t]\n        # Slight bonus for buy2buy neighbors\n        aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n        for aid in aids3:\n            aids_temp[aid] += 0.1\n        return [k for k, v in aids_temp.most_common(20)]\n    \n    # If not enough unique aids, use co-vis\n    aids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_aids if aid in top_20_buys]))\n    aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2 + aids3).most_common(20) if aid2 not in unique_aids]\n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    return result + list(top_orders)[:20 - len(result)]","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"d1c4177a-f1be-4aeb-afc2-43ee42bcfa2b","cell_type":"code","source":"# Generate predictions\npred_df_clicks = test_df.sort_values([\"session\", \"ts\"]).groupby(\"session\").apply(suggest_clicks)\npred_df_buys   = test_df.sort_values([\"session\", \"ts\"]).groupby(\"session\").apply(suggest_buys)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"e6eeb069-c5e1-48f2-ae48-00f8299da778","cell_type":"markdown","source":"## Create and Save Submission","metadata":{}},{"id":"dcf9285c-752a-45fc-9e6b-3f2bf5664ae5","cell_type":"code","source":"# Convert predictions to dataframes\nclicks_pred_df = pd.DataFrame(pred_df_clicks.add_suffix(\"_clicks\"), columns=[\"labels\"]).reset_index()\norders_pred_df = pd.DataFrame(pred_df_buys.add_suffix(\"_orders\"), columns=[\"labels\"]).reset_index()\ncarts_pred_df  = pd.DataFrame(pred_df_buys.add_suffix(\"_carts\"), columns=[\"labels\"]).reset_index()\n\n# Concatenate submissions\npred_df = pd.concat([clicks_pred_df, orders_pred_df, carts_pred_df])\npred_df.columns = [\"session_type\", \"labels\"]\npred_df[\"labels\"] = pred_df.labels.apply(lambda x: \" \".join(map(str,x)))\n\n# Save to CSV\npred_df.to_csv(\"submission.csv\", index=False)\npred_df.head(10)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"73e6ace5-b243-4c8b-b66f-20cee48cbde1","cell_type":"markdown","source":"## Install","metadata":{}},{"id":"162a4ee2-80b1-412b-9594-418089f75ab9","cell_type":"code","source":"!pip install xgboost --quiet\n!pip install networkx --quiet","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"4d30c153-eece-4936-ae35-facbad8c4791","cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom tqdm.notebook import tqdm\nimport os, sys, glob, gc\nfrom collections import Counter\nimport cudf, itertools\n\n# XGBoost for the ML model\n\nimport xgboost as xgb\n\n# For graph/embedding example\n\nimport networkx as nx\n\n# We'll assume these exist from your earlier code:\n# - data_cache (dict of DataFrames from parquet)\n# - files (list of parquet file paths)\n# - top_20_clicks, top_20_buys, top_20_buy2buy (dicts from co-visitation)\n# - test_df (DataFrame with test data)\n# - type_labels\n# - type_weight_multipliers = {0:1, 1:6, 2:3}\n\nVER = 5  # version ID for references\nprint(\"Setup complete.\")","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"7e512b12-2796-43fc-a46a-145178f2b38a","cell_type":"markdown","source":"## Build a Training Dataset for XGBoost","metadata":{}},{"id":"3ba6416b-dfcc-43ae-be3e-fb3a4a5b0d2d","cell_type":"code","source":"%%time\n\ndef build_training_data(\n    data_cache, \n    top_20_clicks, \n    top_20_buys, \n    top_20_buy2buy, \n    sample_size=200_000\n):\n    \"\"\"\n    Example function to create a labeled dataset for XGBoost.\n    - We'll pick some random subset of sessions from data_cache (or all, if memory allows).\n    - We'll label the last item as positive (y=1).\n    - We'll sample some 'negative' items from co-vis candidates or random draws.\n    - We'll create simple numeric features for the pair (session, item).\n    \"\"\"\n    \n    # 1) Concatenate a subset of your training parquet data\n    #    For demonstration, we'll just pick the first N files or random selection. \n    #    You can change this logic to read more data or do a local CV approach.\n    \n    file_subset = list(data_cache.keys())  # all files\n    np.random.shuffle(file_subset)\n    chosen_files = file_subset[:5]  # pick only 5 for demonstration\n    df_list = []\n    for f in chosen_files:\n        df_list.append(data_cache[f])\n    train_df = pd.concat(df_list, ignore_index=True)\n    \n    # Convert ts and type if needed, or assume already done\n    # train_df.ts = (train_df.ts/1000).astype('int32')\n    # train_df['type'] = train_df['type'].map(type_labels).astype('int8')\n    \n    # Let's reduce to a random sample of sessions for memory\n    all_sessions = train_df['session'].unique()\n    chosen_sessions = np.random.choice(all_sessions, size=min(sample_size, len(all_sessions)), replace=False)\n    train_df = train_df[train_df.session.isin(chosen_sessions)]\n    \n    # 2) Build positives (y=1): For each session, the last item is positive\n    train_df = train_df.sort_values(['session','ts'], ascending=[True,True])\n    train_df['rank'] = train_df.groupby('session').cumcount(ascending=False)\n    \n    # The \"last item\" in ascending order will have the highest rank,\n    # or we can just pick the final row in each session:\n    train_df['max_rank'] = train_df.groupby('session')['rank'].transform('max')\n    # Label if rank == max_rank => last item in session\n    train_df['label'] = (train_df['rank'] == train_df['max_rank']).astype('int8')\n    \n    # Keep just needed columns\n    train_df = train_df[['session','aid','type','ts','label']]\n    \n    # 3) Negative sampling: For sessions with label=1, we pick the same session and attach some 'candidate' items.\n    # For demonstration, let's gather each session's last item as positive, \n    # and sample up to 5 negative items from top_20_* or random.\n    \n    out_rows = []\n    grouped = train_df.groupby('session')\n    \n    for session_id, grp in tqdm(grouped, total=len(grouped)):\n        grp = grp.sort_values('ts')\n        \n        # Identify the positive item\n        pos_df = grp[grp['label'] == 1]\n        if len(pos_df) == 0:\n            continue\n        pos_aid = pos_df.iloc[0]['aid']  # last item\n        pos_ts  = pos_df.iloc[0]['ts']\n        pos_type= pos_df.iloc[0]['type']\n        \n        # Collect features for the positive pair\n        out_rows.append({\n            'session': session_id,\n            'aid': pos_aid,\n            'label': 1,\n            'ts': pos_ts,\n            'type': pos_type\n        })\n        \n        # Now sample negative items from the co-vis or random\n        # We'll gather a few aids from top_20_clicks + top_20_buys, if available\n        cand_items = []\n        # Use the session's last item as a key to fetch top co-vis neighbors\n        if pos_aid in top_20_clicks:\n            cand_items += top_20_clicks[pos_aid]\n        if pos_aid in top_20_buys:\n            cand_items += top_20_buys[pos_aid]\n        if pos_aid in top_20_buy2buy:\n            cand_items += top_20_buy2buy[pos_aid]\n        \n        cand_items = list(set(cand_items))  # unique\n        np.random.shuffle(cand_items)\n        cand_items = cand_items[:5]  # limit to 5 negative samples\n        \n        for neg_aid in cand_items:\n            out_rows.append({\n                'session': session_id,\n                'aid': neg_aid,\n                'label': 0,\n                'ts': pos_ts,  # we reuse the time from the last event\n                'type': 0     # guess type=click or 0 for negative example\n            })\n    \n    # Build final DataFrame\n    train_data = pd.DataFrame(out_rows)\n    return train_data\n\n\n# Generate train_data\ntrain_data = build_training_data(\n    data_cache=data_cache,\n    top_20_clicks=top_20_clicks,\n    top_20_buys=top_20_buys,\n    top_20_buy2buy=top_20_buy2buy,\n    sample_size=200_000\n)\nprint(\"Training dataset shape:\", train_data.shape)\ntrain_data.head(10)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"8c9fca59-7cac-49b3-b69e-19909138588a","cell_type":"markdown","source":"## Feature Engineering for XGBoost","metadata":{}},{"id":"0f9e7300-7d22-4049-831d-93ce0145813b","cell_type":"code","source":"%%time\n\ndef compute_features(df):\n    \"\"\"\n    Example function to compute simple numeric features for each row in df.\n    Here each row is (session, aid, ts, type, label).\n    We'll add:\n      - hour of day (just an example)\n      - co-vis stats (like # of neighbors in top_20_* dict)\n      - etc.\n    \"\"\"\n    # Convert timestamp to hour-of-day as an example\n    df['hour'] = (df['ts'] // 3600) % 24  # simplistic\n    df['type_weight'] = df['type'].map({0:1, 1:6, 2:3}).fillna(0).astype('int8')\n    \n    # Example co-vis stats: how many neighbors does this item have in each matrix?\n    # (We store the length of the list from top_20_clicks, etc.)\n    click_len = []\n    buy_len = []\n    buy2buy_len = []\n    for aid in df['aid']:\n        click_len.append(len(top_20_clicks[aid]) if aid in top_20_clicks else 0)\n        buy_len.append(len(top_20_buys[aid]) if aid in top_20_buys else 0)\n        buy2buy_len.append(len(top_20_buy2buy[aid]) if aid in top_20_buy2buy else 0)\n    df['click_neighbors'] = click_len\n    df['buy_neighbors'] = buy_len\n    df['buy2buy_neighbors'] = buy2buy_len\n    \n    # Potentially we can add session-level features if needed\n    # e.g. how many events are in the session? \n    session_counts = df.groupby('session').session.transform('count')\n    df['session_event_count'] = session_counts\n    \n    return df\n\ntrain_data_feats = compute_features(train_data)\ntrain_data_feats.head(10)\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"69ea25d2-fd21-4c34-8a4a-afdfef75a7d0","cell_type":"markdown","source":"## Train XGBoost","metadata":{}},{"id":"26bf0992-e5f9-4849-ba0a-3371e6268666","cell_type":"code","source":"%%time\n\n# Prepare data for XGBoost\nfeature_cols = [\n    'hour',\n    'type_weight',\n    'click_neighbors',\n    'buy_neighbors',\n    'buy2buy_neighbors',\n    'session_event_count'\n]\ntarget_col = 'label'\n\nX = train_data_feats[feature_cols]\ny = train_data_feats[target_col]\n\n# Convert to DMatrix\ndtrain = xgb.DMatrix(X, label=y)\n\n# Simple XGBoost parameters\nparams = {\n    'objective': 'binary:logistic',\n    'eval_metric': 'auc',\n    'max_depth': 6,\n    'eta': 0.1\n}\n\nmodel = xgb.train(params, dtrain, num_boost_round=50)\nprint(\"XGBoost training complete!\")\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"6faf934c-38a8-42f7-85db-7a0851fd011f","cell_type":"markdown","source":"## Generate Test Predictions Using Co-vis + XGBoost + Embeddings","metadata":{}},{"id":"e42747ad-3226-4c16-b3f4-690e895458a2","cell_type":"code","source":"%%time\n\n##########################\n# Candidate Generation\n##########################\ndef suggest_candidates(row, top_20_clicks, top_20_buys, top_20_buy2buy, top_k=40):\n    \"\"\"\n    Example candidate generator: \n    - Takes a single session's DataFrame (`row`),\n    - Gathers unique aids from this session,\n    - Extends with co-vis neighbors from the 3 dicts,\n    - Returns up to top_k unique candidates.\n    \"\"\"\n    aids = row.aid.tolist()\n    # Reverse-unique the aids\n    unique_aids = list(dict.fromkeys(aids[::-1]))\n    \n    # Combine neighbors from all 3 co-vis dictionaries\n    candidates = set()\n    for aid in unique_aids:\n        if aid in top_20_clicks:\n            candidates.update(top_20_clicks[aid])\n        if aid in top_20_buys:\n            candidates.update(top_20_buys[aid])\n        if aid in top_20_buy2buy:\n            candidates.update(top_20_buy2buy[aid])\n    \n    # Include the session's actual items\n    candidates.update(unique_aids)\n    \n    # Return up to top_k\n    candidates = list(candidates)[:top_k]  \n    return candidates\n\n\n##########################\n# Build Test Features\n##########################\ndef build_test_features(\n    session_id, \n    candidate_items, \n    session_df, \n    top_20_clicks, \n    top_20_buys, \n    top_20_buy2buy\n):\n    \"\"\"\n    Build feature set for (session_id, candidate_item).\n    Features:\n      - hour of last event\n      - type_weight (based on last event in the session)\n      - # of neighbors in top_20_clicks, top_20_buys, top_20_buy2buy\n      - session_event_count (length of session)\n    \"\"\"\n    # Sort session by ts to get the last event\n    session_df = session_df.sort_values('ts')\n    last_ts = session_df.iloc[-1]['ts']\n    session_event_count = len(session_df)\n    \n    # Last event's type\n    last_type = session_df.iloc[-1]['type']\n    \n    feature_rows = []\n    for aid in candidate_items:\n        row_dict = {}\n        row_dict['session'] = session_id\n        row_dict['aid'] = aid\n        \n        # Basic features\n        row_dict['hour'] = (last_ts // 3600) % 24\n        row_dict['type_weight'] = type_weight_multipliers.get(last_type, 0)\n        row_dict['click_neighbors'] = len(top_20_clicks[aid]) if aid in top_20_clicks else 0\n        row_dict['buy_neighbors']   = len(top_20_buys[aid])   if aid in top_20_buys   else 0\n        row_dict['buy2buy_neighbors'] = len(top_20_buy2buy[aid]) if aid in top_20_buy2buy else 0\n        row_dict['session_event_count'] = session_event_count\n        \n        feature_rows.append(row_dict)\n    \n    return pd.DataFrame(feature_rows)\n\n\n##########################\n# Scoring Test Sessions\n##########################\ntest_sessions = test_df['session'].unique()\nfinal_preds = []\n\n# Define the columns that your XGBoost model expects\nfeature_cols = [\n    'hour', \n    'type_weight', \n    'click_neighbors', \n    'buy_neighbors', \n    'buy2buy_neighbors', \n    'session_event_count'\n]\n\nfor session_id in tqdm(test_sessions, desc='Scoring test'):\n    # Extract the DataFrame subset for this session\n    session_data = test_df[test_df.session == session_id]\n    \n    # 1) Candidate generation\n    candidates = suggest_candidates(session_data, top_20_clicks, top_20_buys, top_20_buy2buy, top_k=40)\n    \n    # 2) Build feature DataFrame\n    feats_df = build_test_features(session_id, candidates, session_data, \n                                   top_20_clicks, top_20_buys, top_20_buy2buy)\n    \n    # 3) XGBoost predict\n    dtest = xgb.DMatrix(feats_df[feature_cols])\n    preds = model.predict(dtest)\n    \n    feats_df['score'] = preds\n    \n    # 4) Sort by predicted probability\n    feats_df = feats_df.sort_values('score', ascending=False)\n    \n    # 5) Take top 20\n    top_20 = feats_df['aid'].values[:20]\n    \n    # Save final\n    final_preds.append({\n        'session': session_id,\n        'labels': top_20\n    })\n\n##########################\n# Create Submission\n##########################\npred_df_xgb = pd.DataFrame(final_preds)\npred_df_xgb['labels'] = pred_df_xgb['labels'].apply(lambda x: ' '.join(map(str, x)))\npred_df_xgb['session_type'] = pred_df_xgb['session'].astype(str) + '_orders'  # or '_clicks', etc.\n\n# If you want separate lines for clicks/orders/carts, replicate logic or adapt as needed\nsubmission_xgb = pred_df_xgb[['session_type','labels']].reset_index(drop=True)\nsubmission_xgb.to_csv('submission_xgb.csv', index=False)\nsubmission_xgb.head(10)\n","metadata":{"trusted":true,"scrolled":true,"jupyter":{"source_hidden":true}},"outputs":[],"execution_count":null},{"id":"7b83b659-4f13-41a7-abc9-708a8f9abd1c","cell_type":"code","source":"from tqdm import tqdm\n\ndef convert_covisit_to_cudf(covisit_dict):\n    \"\"\"\n    Convert a co-visitation dictionary to a cuDF DataFrame for GPU operations.\n    \"\"\"\n    rows = [(k, v) for k, neighbors in tqdm(covisit_dict.items(), desc='Converting co-visitation dict', total=len(covisit_dict)) for v in neighbors]\n    return cudf.DataFrame(rows, columns=['aid_x', 'aid_y'])\n\n# Convert co-vis dictionaries\ntop_20_clicks_df = convert_covisit_to_cudf(top_20_clicks)\ntop_20_buys_df = convert_covisit_to_cudf(top_20_buys)\ntop_20_buy2buy_df = convert_covisit_to_cudf(top_20_buy2buy)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"9736fbc8-1e53-4187-80f0-5269a0b20b5b","cell_type":"code","source":"import gc\nimport numpy as np\nimport cudf\nfrom tqdm import tqdm\nimport xgboost as xgb\n\n#############################\n# 1) Generate Candidates\n#############################\ndef generate_candidates_in_batches(\n    test_df,\n    top_20_clicks_df,\n    top_20_buys_df,\n    top_20_buy2buy_df,\n    top_k=40,\n    batch_size=100_000\n):\n    \"\"\"\n    Generate candidate items in smaller batches to avoid out-of-memory errors.\n    \n    Returns a cuDF DataFrame with columns including:\n      - 'session'\n      - 'candidate_aid' (the neighbor item)\n      - possibly other columns from merges\n    \"\"\"\n    candidates_list = []\n    total_len = len(test_df)\n    total_batches = int(np.ceil(total_len / batch_size))\n    \n    for i in tqdm(range(total_batches), desc=\"Generating candidates in batches\"):\n        start_idx = i * batch_size\n        end_idx = min((i + 1) * batch_size, total_len)\n        \n        batch = test_df.iloc[start_idx:end_idx]\n\n        # Merge this batch with each co-vis DataFrame\n        candidates_clicks = batch.merge(top_20_clicks_df, left_on='aid', right_on='aid_x', how='left')\n        candidates_buys = batch.merge(top_20_buys_df, left_on='aid', right_on='aid_x', how='left')\n        candidates_buy2buy = batch.merge(top_20_buy2buy_df, left_on='aid', right_on='aid_x', how='left')\n        \n        # Concatenate partial merges\n        batch_candidates = cudf.concat([candidates_clicks, candidates_buys, candidates_buy2buy], ignore_index=True)\n        \n        # Remove duplicates (session, aid_y)\n        batch_candidates = batch_candidates.drop_duplicates(subset=['session', 'aid_y'])\n        \n        # IMPORTANT: rename here so that subsequent code references 'candidate_aid'\n        batch_candidates = batch_candidates.rename(columns={'aid_y': 'candidate_aid'})\n        \n        # Keep only top_k neighbors per session\n        batch_candidates['rank'] = batch_candidates.groupby('session').candidate_aid.cumcount()\n        batch_candidates = batch_candidates[batch_candidates['rank'] < top_k].drop('rank', axis=1)\n        \n        candidates_list.append(batch_candidates)\n        \n        # Free memory\n        del batch, candidates_clicks, candidates_buys, candidates_buy2buy, batch_candidates\n        gc.collect()\n    \n    # Combine all batches into one DataFrame\n    candidates = cudf.concat(candidates_list, ignore_index=True)\n    return candidates\n\n\n#############################\n# 2) Build Features\n#############################\ndef build_features_in_batches(\n    candidates,\n    test_df,\n    top_20_clicks_df,\n    top_20_buys_df,\n    top_20_buy2buy_df,\n    type_weight_multipliers,\n    batch_size=100_000\n):\n    features_list = []\n    total_len = len(candidates)\n    total_batches = int(np.ceil(total_len / batch_size))\n    \n    # Precompute session-level features\n    session_features = test_df[['session', 'ts', 'type']] \\\n        .groupby('session') \\\n        .max() \\\n        .reset_index()\n    session_features['hour'] = (session_features['ts'] // 3600) % 24\n    session_features['type_weight'] = session_features['type'].map(type_weight_multipliers).fillna(0)\n    \n    # ALSO: compute how many events in each session => session_event_count\n    session_size = test_df.groupby('session').size().reset_index().rename(columns={0: 'session_event_count'})\n    # Merge this into session_features\n    session_features = session_features.merge(session_size, on='session', how='left')\n    # Now session_features has columns: [session, ts, type, hour, type_weight, session_event_count]\n\n    for i in range(total_batches):\n        start_idx = i * batch_size\n        end_idx = min((i + 1) * batch_size, total_len)\n        \n        batch = candidates.iloc[start_idx:end_idx]\n        \n        # Merge session-level features\n        batch_features = batch.merge(session_features, on='session', how='left')\n        \n        # Count neighbor size for each candidate\n        neighbor_clicks = (\n            batch.merge(top_20_clicks_df, left_on='candidate_aid', right_on='aid_x', how='left')\n            .groupby('candidate_aid')\n            .size()\n        )\n        neighbor_buys = (\n            batch.merge(top_20_buys_df, left_on='candidate_aid', right_on='aid_x', how='left')\n            .groupby('candidate_aid')\n            .size()\n        )\n        neighbor_buy2buy = (\n            batch.merge(top_20_buy2buy_df, left_on='candidate_aid', right_on='aid_x', how='left')\n            .groupby('candidate_aid')\n            .size()\n        )\n        \n        batch_features['click_neighbors'] = batch_features['candidate_aid'].map(neighbor_clicks).fillna(0)\n        batch_features['buy_neighbors'] = batch_features['candidate_aid'].map(neighbor_buys).fillna(0)\n        batch_features['buy2buy_neighbors'] = batch_features['candidate_aid'].map(neighbor_buy2buy).fillna(0)\n        \n        features_list.append(batch_features)\n    \n    return cudf.concat(features_list, ignore_index=True)\n\n#############################\n# 3) Batch Predict\n#############################\ndef batch_predict(\n    test_df,\n    model,\n    top_20_clicks_df,\n    top_20_buys_df,\n    top_20_buy2buy_df,\n    type_weight_multipliers,\n    feature_cols,\n    top_k=40,\n    batch_size=100_000\n):\n    \"\"\"\n    Full pipeline for generating candidates in batches, building features, then scoring with XGBoost.\n    \n    :param test_df: cuDF DataFrame with columns ['session', 'aid', 'ts', 'type']\n    :param model:    Trained XGBoost model\n    :param top_20_clicks_df, top_20_buys_df, top_20_buy2buy_df: co-vis cuDF DataFrames with columns [aid_x, aid_y]\n    :param type_weight_multipliers: dict for weighting last event type\n    :param feature_cols: list of columns used by XGBoost\n    :param top_k: max neighbors per session item\n    :param batch_size: chunk size to avoid out-of-memory errors\n    :return: submission DataFrame with columns ['session_type','labels']\n    \"\"\"\n    print(\"Step 1: Generating candidates\")\n    candidates = generate_candidates_in_batches(\n        test_df,\n        top_20_clicks_df,\n        top_20_buys_df,\n        top_20_buy2buy_df,\n        top_k=top_k,\n        batch_size=batch_size\n    )\n    \n    print(\"Step 2: Building features\")\n    features = build_features_in_batches(\n        candidates,\n        test_df,\n        top_20_clicks_df,\n        top_20_buys_df,\n        top_20_buy2buy_df,\n        type_weight_multipliers,\n        batch_size=batch_size\n    )\n    \n    print(\"Step 3: Predicting scores\")\n    features_list = []\n    total_len = len(features)\n    total_batches = int(np.ceil(total_len / batch_size))\n    \n    for i in tqdm(range(total_batches), desc=\"Predicting in batches\"):\n        start_idx = i * batch_size\n        end_idx = min((i + 1) * batch_size, total_len)\n        \n        batch = features.iloc[start_idx:end_idx]\n        \n        dtest = xgb.DMatrix(batch[feature_cols].to_pandas())\n        batch['score'] = model.predict(dtest)\n        features_list.append(batch)\n        \n        del batch, dtest\n        gc.collect()\n    \n    # Combine predictions\n    features = cudf.concat(features_list, ignore_index=True)\n    \n    print(\"Step 4: Sorting top predictions per session\")\n    features = features.sort_values(['session', 'score'], ascending=[True, False])\n    \n    # Keep top 20\n    top_20 = features.groupby('session').head(20)\n    \n    # Format final submission\n    final_preds = top_20.groupby('session')['candidate_aid'].apply(\n        lambda x: ' '.join(map(str, x.to_pandas()))\n    ).reset_index()\n    \n    final_preds.columns = ['session', 'labels']\n    final_preds['session_type'] = final_preds['session'].astype(str) + '_orders'\n    \n    return final_preds[['session_type', 'labels']]","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"c0e3678a-a5dc-44ea-90d9-5563d60d51d5","cell_type":"code","source":"# Try a smaller batch_size, e.g. 50_000, 20_000, 10_000, etc.\nsubmission_xgb = batch_predict(\n    test_df=test_df_cudf,\n    model=model,\n    top_20_clicks_df=top_20_clicks_df,\n    top_20_buys_df=top_20_buys_df,\n    top_20_buy2buy_df=top_20_buy2buy_df,\n    type_weight_multipliers={0:1,1:6,2:3},\n    feature_cols=[\n        'hour','type_weight','click_neighbors',\n        'buy_neighbors','buy2buy_neighbors','session_event_count'\n    ],\n    top_k=40,\n    batch_size=10_000  # or even smaller\n)\n\nsubmission_xgb.to_csv('submission_xgb.csv', index=False)\nsubmission_xgb.head(10)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"id":"4c4b3df2-6c99-4de2-969c-cece38d86858","cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}