{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":38760,"databundleVersionId":4493939,"sourceType":"competition"},{"sourceId":4436180,"sourceType":"datasetVersion","datasetId":2597726}],"dockerImageVersionId":30302,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Candidate ReRank Model using Handcrafted Rules\nIn this notebook, we present a \"candidate rerank\" model using handcrafted rules. We can improve this model by engineering features, merging them unto items and users, and training a reranker model (such as XGB) to choose our final 20. Furthermore to tune and improve this notebook, we should build a local CV scheme to experiment new logic and/or models.\n\nUPDATE: I published a notebook to compute validation score [here][10] using Radek's scheme described [here][11].\n\nNote in this competition, a \"session\" actually means a unique \"user\". So our task is to predict what each of the `1,671,803` test \"users\" (i.e. \"sessions\") will do in the future. For each test \"user\" (i.e. \"session\") we must predict what they will `click`, `cart`, and `order` during the remainder of the week long test period.\n\n### Step 1 - Generate Candidates\nFor each test user, we generate possible choices, i.e. candidates. In this notebook, we generate candidates from 5 sources:\n* User history of clicks, carts, orders\n* Most popular 20 clicks, carts, orders during test week\n* Co-visitation matrix of click/cart/order to cart/order with type weighting\n* Co-visitation matrix of cart/order to cart/order called buy2buy\n* Co-visitation matrix of click/cart/order to clicks with time weighting\n\n### Step 2 - ReRank and Choose 20\nGiven the list of candidates, we must select 20 to be our predictions. In this notebook, we do this with a set of handcrafted rules. We can improve our predictions by training an XGBoost model to select for us. Our handcrafted rules give priority to:\n* Most recent previously visited items\n* Items previously visited multiple times\n* Items previously in cart or order\n* Co-visitation matrix of cart/order to cart/order\n* Current popular items\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/c_r_model.png)\n  \n# Credits\nWe thank many Kagglers who have shared ideas. We use co-visitation matrix idea from Vladimir [here][1]. We use groupby sort logic from Sinan in comment section [here][4]. We use duplicate prediction removal logic from Radek [here][5]. We use multiple visit logic from Pietro [here][2]. We use type weighting logic from Ingvaras [here][3]. We use leaky test data from my previous notebook [here][4]. And some ideas may have originated from Tawara [here][6] and KJ [here][7]. We use Colum2131's parquets [here][8]. Above image is from Ravi's discussion about candidate rerank models [here][9]\n\n[1]: https://www.kaggle.com/code/vslaykovsky/co-visitation-matrix\n[2]: https://www.kaggle.com/code/pietromaldini1/multiple-clicks-vs-latest-items\n[3]: https://www.kaggle.com/code/ingvarasgalinskas/item-type-vs-multiple-clicks-vs-latest-items\n[4]: https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost\n[5]: https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\n[6]: https://www.kaggle.com/code/ttahara/otto-mors-aid-frequency-baseline\n[7]: https://www.kaggle.com/code/whitelily/co-occurrence-baseline\n[8]: https://www.kaggle.com/datasets/columbia2131/otto-chunk-data-inparquet-format\n[9]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\n[10]: https://www.kaggle.com/cdeotte/compute-validation-score-cv-564\n[11]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991","metadata":{"papermill":{"duration":0.005198,"end_time":"2022-11-10T16:03:20.966987","exception":false,"start_time":"2022-11-10T16:03:20.961789","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# Notes\nBelow are notes about versions:\n* **Version 1 LB 0.573** Uses popular ideas from public notebooks and adds additional co-visitation matrices and additional logic. Has CV `0.563`. See validation notebook version 2 [here][1].\n* **Version 2 LB 573** Refactor logic for `suggest_buys(df)` to make it clear how new co-visitation matrices are reranking the candidates by adding to candidate weights. Also new logic boosts CV by `+0.0003`. Also LB is slightly better too. See validation notebook version 3 [here][1]\n* **Version 3** is the same as version 2 but 1.5x faster co-visitation matrix computation!\n* **Version 4 LB 575** Use top20 for clicks and top15 for carts and buys (instead of top40 and top40). This boosts CV `+0.0015` hooray! New CV is `0.5647`. See validation version 5 [here][1]\n* **Version 5** is the same as version 4 but 2x faster co-visitation matrix computation! (and 3x faster than version 1)\n* **Version 6** Stay tuned for more versions...\n\n[1]: https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-564","metadata":{}},{"cell_type":"markdown","source":"# Step 1 - Candidate Generation with RAPIDS\nFor candidate generation, we build three co-visitation matrices. One computes the popularity of cart/order given a user's previous click/cart/order. We apply type weighting to this matrix. One computes the popularity of cart/order given a user's previous cart/order. We call this \"buy2buy\" matrix. One computes the popularity of clicks given a user previously click/cart/order.  We apply time weighting to this matrix. We will use RAPIDS cuDF GPU to compute these matrices quickly!","metadata":{"papermill":{"duration":0.00373,"end_time":"2022-11-10T16:03:20.9748","exception":false,"start_time":"2022-11-10T16:03:20.97107","status":"completed"},"tags":[]}},{"cell_type":"code","source":"VER = 5\n\nimport pandas as pd, numpy as np\nfrom tqdm.notebook import tqdm\nimport os, sys, pickle, glob, gc\nfrom collections import Counter\nimport cudf, itertools\nprint('We will use RAPIDS version',cudf.__version__)","metadata":{"papermill":{"duration":3.036143,"end_time":"2022-11-10T16:03:24.014816","exception":false,"start_time":"2022-11-10T16:03:20.978673","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-10-26T20:31:35.835282Z","iopub.execute_input":"2025-10-26T20:31:35.835611Z","iopub.status.idle":"2025-10-26T20:31:36.665669Z","shell.execute_reply.started":"2025-10-26T20:31:35.835585Z","shell.execute_reply":"2025-10-26T20:31:36.664724Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Compute Three Co-visitation Matrices with RAPIDS\nWe will compute 3 co-visitation matrices using RAPIDS cuDF on GPU. This is 30x faster than using Pandas CPU like other public notebooks! For maximum speed, set the variable `DISK_PIECES` to the smallest number possible based on the GPU you are using without incurring memory errors. If you run this code offline with 32GB GPU ram, then you can use `DISK_PIECES = 1` and compute each co-visitation matrix in almost 1 minute! Kaggle's GPU only has 16GB ram, so we use `DISK_PIECES = 4` and it takes an amazing 3 minutes each! Below are some of the tricks to speed up computation\n* Use RAPIDS cuDF GPU instead of Pandas CPU\n* Read disk once and save in CPU RAM for later GPU multiple use\n* Process largest amount of data possible on GPU at one time\n* Merge data in two stages. Multiple small to single medium. Multiple medium to single large.\n* Write result as parquet instead of dictionary","metadata":{"papermill":{"duration":0.00424,"end_time":"2022-11-10T16:03:24.023816","exception":false,"start_time":"2022-11-10T16:03:24.019576","status":"completed"},"tags":[]}},{"cell_type":"code","source":"%%time\n# CACHE FUNCTIONS\ndef read_file(f):\n    return cudf.DataFrame( data_cache[f] )\ndef read_file_to_cache(f):\n    df = pd.read_parquet(f)\n    df.ts = (df.ts/1000).astype('int32')\n    df['type'] = df['type'].map(type_labels).astype('int8')\n    return df\n\n# CACHE THE DATA ON CPU BEFORE PROCESSING ON GPU\ndata_cache = {}\ntype_labels = {'clicks':0, 'carts':1, 'orders':2}\nfiles = glob.glob('../input/otto-chunk-data-inparquet-format/*_parquet/*')\nfor f in files: data_cache[f] = read_file_to_cache(f)\n\n# CHUNK PARAMETERS\nREAD_CT = 5\nCHUNK = int( np.ceil( len(files)/6 ))\nprint(f'We will process {len(files)} files, in groups of {READ_CT} and chunks of {CHUNK}.')","metadata":{"papermill":{"duration":0.063943,"end_time":"2022-11-10T16:03:24.091816","exception":false,"start_time":"2022-11-10T16:03:24.027873","status":"completed"},"tags":[],"_kg_hide-input":true,"execution":{"iopub.status.busy":"2025-10-26T20:31:36.667604Z","iopub.execute_input":"2025-10-26T20:31:36.668403Z","iopub.status.idle":"2025-10-26T20:32:25.824135Z","shell.execute_reply.started":"2025-10-26T20:31:36.668366Z","shell.execute_reply":"2025-10-26T20:32:25.822942Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 1) \"Carts Orders\" Co-visitation Matrix - Type Weighted","metadata":{"papermill":{"duration":0.004089,"end_time":"2022-11-10T16:03:24.100502","exception":false,"start_time":"2022-11-10T16:03:24.096413","status":"completed"},"tags":[]}},{"cell_type":"code","source":"%%time\ntype_weight = {0:1, 1:6, 2:3}\n\n# USE SMALLEST DISK_PIECES POSSIBLE WITHOUT MEMORY ERROR\nDISK_PIECES = 4\nSIZE = 1.86e6/DISK_PIECES\n\n# COMPUTE IN PARTS FOR MEMORY MANGEMENT\nfor PART in range(DISK_PIECES):\n    print()\n    print('### DISK PART',PART+1)\n    \n    # MERGE IS FASTEST PROCESSING CHUNKS WITHIN CHUNKS\n    # => OUTER CHUNKS\n    for j in range(6):\n        a = j*CHUNK\n        b = min( (j+1)*CHUNK, len(files) )\n        print(f'Processing files {a} thru {b-1} in groups of {READ_CT}...')\n        \n        # => INNER CHUNKS\n        for k in range(a,b,READ_CT):\n            # READ FILE\n            df = [read_file(files[k])]\n            for i in range(1,READ_CT): \n                if k+i<b: df.append( read_file(files[k+i]) )\n            df = cudf.concat(df,ignore_index=True,axis=0)\n            df = df.sort_values(['session','ts'],ascending=[True,False])\n            # USE TAIL OF SESSION\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n<30].drop('n',axis=1)\n            # CREATE PAIRS\n            df = df.merge(df,on='session')\n            df = df.loc[ ((df.ts_x - df.ts_y).abs()< 24 * 60 * 60) & (df.aid_x != df.aid_y) ]\n            # MEMORY MANAGEMENT COMPUTE IN PARTS\n            df = df.loc[(df.aid_x >= PART*SIZE)&(df.aid_x < (PART+1)*SIZE)]\n            # ASSIGN WEIGHTS\n            df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n            df['wgt'] = df.type_y.map(type_weight)\n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n            # COMBINE INNER CHUNKS\n            if k==a: tmp2 = df\n            else: tmp2 = tmp2.add(df, fill_value=0)\n            print(k,', ',end='')\n        print()\n        # COMBINE OUTER CHUNKS\n        if a==0: tmp = tmp2\n        else: tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    # CONVERT MATRIX TO DICTIONARY\n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n    # SAVE TOP 40\n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    tmp = tmp.loc[tmp.n<15].drop('n',axis=1)\n    # SAVE PART TO DISK (convert to pandas first uses less memory)\n    tmp.to_pandas().to_parquet(f'top_15_carts_orders_v{VER}_{PART}.pqt')","metadata":{"papermill":{"duration":566.561189,"end_time":"2022-11-10T16:12:50.666123","exception":false,"start_time":"2022-11-10T16:03:24.104934","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-10-26T20:32:25.82601Z","iopub.execute_input":"2025-10-26T20:32:25.826435Z","iopub.status.idle":"2025-10-26T20:35:39.617328Z","shell.execute_reply.started":"2025-10-26T20:32:25.826395Z","shell.execute_reply":"2025-10-26T20:35:39.615905Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2) \"Buy2Buy\" Co-visitation Matrix","metadata":{"papermill":{"duration":0.03219,"end_time":"2022-11-10T16:12:50.730634","exception":false,"start_time":"2022-11-10T16:12:50.698444","status":"completed"},"tags":[]}},{"cell_type":"code","source":"%%time\n# USE SMALLEST DISK_PIECES POSSIBLE WITHOUT MEMORY ERROR\nDISK_PIECES = 1\nSIZE = 1.86e6/DISK_PIECES\n\n# COMPUTE IN PARTS FOR MEMORY MANGEMENT\nfor PART in range(DISK_PIECES):\n    print()\n    print('### DISK PART',PART+1)\n    \n    # MERGE IS FASTEST PROCESSING CHUNKS WITHIN CHUNKS\n    # => OUTER CHUNKS\n    for j in range(6):\n        a = j*CHUNK\n        b = min( (j+1)*CHUNK, len(files) )\n        print(f'Processing files {a} thru {b-1} in groups of {READ_CT}...')\n        \n        # => INNER CHUNKS\n        for k in range(a,b,READ_CT):\n            # READ FILE\n            df = [read_file(files[k])]\n            for i in range(1,READ_CT): \n                if k+i<b: df.append( read_file(files[k+i]) )\n            df = cudf.concat(df,ignore_index=True,axis=0)\n            df = df.loc[df['type'].isin([1,2])] # ONLY WANT CARTS AND ORDERS\n            df = df.sort_values(['session','ts'],ascending=[True,False])\n            # USE TAIL OF SESSION\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n<30].drop('n',axis=1)\n            # CREATE PAIRS\n            df = df.merge(df,on='session')\n            df = df.loc[ ((df.ts_x - df.ts_y).abs()< 14 * 24 * 60 * 60) & (df.aid_x != df.aid_y) ] # 14 DAYS\n            # MEMORY MANAGEMENT COMPUTE IN PARTS\n            df = df.loc[(df.aid_x >= PART*SIZE)&(df.aid_x < (PART+1)*SIZE)]\n            # ASSIGN WEIGHTS\n            df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n            df['wgt'] = 1\n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n            # COMBINE INNER CHUNKS\n            if k==a: tmp2 = df\n            else: tmp2 = tmp2.add(df, fill_value=0)\n            print(k,', ',end='')\n        print()\n        # COMBINE OUTER CHUNKS\n        if a==0: tmp = tmp2\n        else: tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    # CONVERT MATRIX TO DICTIONARY\n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n    # SAVE TOP 40\n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    tmp = tmp.loc[tmp.n<15].drop('n',axis=1)\n    # SAVE PART TO DISK (convert to pandas first uses less memory)\n    tmp.to_pandas().to_parquet(f'top_15_buy2buy_v{VER}_{PART}.pqt')","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"papermill":{"duration":113.735315,"end_time":"2022-11-10T16:14:44.498182","exception":false,"start_time":"2022-11-10T16:12:50.762867","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2025-10-26T20:35:39.619572Z","iopub.execute_input":"2025-10-26T20:35:39.619848Z","iopub.status.idle":"2025-10-26T20:36:09.428522Z","shell.execute_reply.started":"2025-10-26T20:35:39.619821Z","shell.execute_reply":"2025-10-26T20:36:09.427526Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3) \"Clicks\" Co-visitation Matrix - Time Weighted","metadata":{"papermill":{"duration":0.04526,"end_time":"2022-11-10T16:14:44.58589","exception":false,"start_time":"2022-11-10T16:14:44.54063","status":"completed"},"tags":[]}},{"cell_type":"code","source":"%%time\n# USE SMALLEST DISK_PIECES POSSIBLE WITHOUT MEMORY ERROR\nDISK_PIECES = 4\nSIZE = 1.86e6/DISK_PIECES\n\n# COMPUTE IN PARTS FOR MEMORY MANGEMENT\nfor PART in range(DISK_PIECES):\n    print()\n    print('### DISK PART',PART+1)\n    \n    # MERGE IS FASTEST PROCESSING CHUNKS WITHIN CHUNKS\n    # => OUTER CHUNKS\n    for j in range(6):\n        a = j*CHUNK\n        b = min( (j+1)*CHUNK, len(files) )\n        print(f'Processing files {a} thru {b-1} in groups of {READ_CT}...')\n        \n        # => INNER CHUNKS\n        for k in range(a,b,READ_CT):\n            # READ FILE\n            df = [read_file(files[k])]\n            for i in range(1,READ_CT): \n                if k+i<b: df.append( read_file(files[k+i]) )\n            df = cudf.concat(df,ignore_index=True,axis=0)\n            df = df.sort_values(['session','ts'],ascending=[True,False])\n            # USE TAIL OF SESSION\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n<30].drop('n',axis=1)\n            # CREATE PAIRS\n            df = df.merge(df,on='session')\n            df = df.loc[ ((df.ts_x - df.ts_y).abs()< 24 * 60 * 60) & (df.aid_x != df.aid_y) ]\n            # MEMORY MANAGEMENT COMPUTE IN PARTS\n            df = df.loc[(df.aid_x >= PART*SIZE)&(df.aid_x < (PART+1)*SIZE)]\n            # ASSIGN WEIGHTS\n            df = df[['session', 'aid_x', 'aid_y','ts_x']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n            df['wgt'] = 1 + 3*(df.ts_x - 1659304800)/(1662328791-1659304800)\n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n            # COMBINE INNER CHUNKS\n            if k==a: tmp2 = df\n            else: tmp2 = tmp2.add(df, fill_value=0)\n            print(k,', ',end='')\n        print()\n        # COMBINE OUTER CHUNKS\n        if a==0: tmp = tmp2\n        else: tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    # CONVERT MATRIX TO DICTIONARY\n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n    # SAVE TOP 40\n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    tmp = tmp.loc[tmp.n<20].drop('n',axis=1)\n    # SAVE PART TO DISK (convert to pandas first uses less memory)\n    tmp.to_pandas().to_parquet(f'top_20_clicks_v{VER}_{PART}.pqt')","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"papermill":{"duration":null,"end_time":null,"exception":false,"start_time":"2022-11-10T16:14:44.629032","status":"running"},"tags":[],"execution":{"iopub.status.busy":"2025-10-26T20:36:09.430247Z","iopub.execute_input":"2025-10-26T20:36:09.430803Z","iopub.status.idle":"2025-10-26T20:39:19.950911Z","shell.execute_reply.started":"2025-10-26T20:36:09.430763Z","shell.execute_reply":"2025-10-26T20:39:19.949955Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# FREE MEMORY\ndel data_cache, tmp\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2025-10-26T20:39:19.952261Z","iopub.execute_input":"2025-10-26T20:39:19.952609Z","iopub.status.idle":"2025-10-26T20:39:20.116055Z","shell.execute_reply.started":"2025-10-26T20:39:19.952573Z","shell.execute_reply":"2025-10-26T20:39:20.115165Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Step 2 - ReRank (choose 20) using handcrafted rules\nFor description of the handcrafted rules, read this notebook's intro.","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[]}},{"cell_type":"code","source":"def load_test():    \n    dfs = []\n    for e, chunk_file in enumerate(glob.glob('../input/otto-chunk-data-inparquet-format/test_parquet/*')):\n        chunk = pd.read_parquet(chunk_file)\n        chunk.ts = (chunk.ts/1000).astype('int32')\n        chunk['type'] = chunk['type'].map(type_labels).astype('int8')\n        dfs.append(chunk)\n    return pd.concat(dfs).reset_index(drop=True) #.astype({\"ts\": \"datetime64[ms]\"})\n\ntest_df = load_test()\nprint('Test data has shape',test_df.shape)\ntest_df.head()","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[],"execution":{"iopub.status.busy":"2025-10-26T20:39:20.117222Z","iopub.execute_input":"2025-10-26T20:39:20.117457Z","iopub.status.idle":"2025-10-26T20:39:21.271849Z","shell.execute_reply.started":"2025-10-26T20:39:20.117437Z","shell.execute_reply":"2025-10-26T20:39:21.270895Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time\ndef pqt_to_dict(df):\n    return df.groupby('aid_x').aid_y.apply(list).to_dict()\n# LOAD THREE CO-VISITATION MATRICES\ntop_20_clicks = pqt_to_dict( pd.read_parquet(f'top_20_clicks_v{VER}_0.pqt') )\nfor k in range(1,DISK_PIECES): \n    top_20_clicks.update( pqt_to_dict( pd.read_parquet(f'top_20_clicks_v{VER}_{k}.pqt') ) )\ntop_20_buys = pqt_to_dict( pd.read_parquet(f'top_15_carts_orders_v{VER}_0.pqt') )\nfor k in range(1,DISK_PIECES): \n    top_20_buys.update( pqt_to_dict( pd.read_parquet(f'top_15_carts_orders_v{VER}_{k}.pqt') ) )\ntop_20_buy2buy = pqt_to_dict( pd.read_parquet(f'top_15_buy2buy_v{VER}_0.pqt') )\n\n# TOP CLICKS AND ORDERS IN TEST\ntop_clicks = test_df.loc[test_df['type']=='clicks','aid'].value_counts().index.values[:20]\ntop_orders = test_df.loc[test_df['type']=='orders','aid'].value_counts().index.values[:20]\n\nprint('Here are size of our 3 co-visitation matrices:')\nprint( len( top_20_clicks ), len( top_20_buy2buy ), len( top_20_buys ) )","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[],"execution":{"iopub.status.busy":"2025-10-26T20:39:21.272921Z","iopub.execute_input":"2025-10-26T20:39:21.273223Z","iopub.status.idle":"2025-10-26T20:40:41.166561Z","shell.execute_reply.started":"2025-10-26T20:39:21.273198Z","shell.execute_reply":"2025-10-26T20:40:41.165565Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#type_weight_multipliers = {'clicks': 1, 'carts': 6, 'orders': 3}\ntype_weight_multipliers = {0: 1, 1: 6, 2: 3}\n\ndef suggest_clicks(df):\n    # USER HISTORY AIDS AND TYPES\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    # RERANK CANDIDATES USING WEIGHTS\n    if len(unique_aids)>=20:\n        weights=np.logspace(0.1,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        # RERANK BASED ON REPEAT ITEMS AND TYPE OF ITEMS\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    # USE \"CLICKS\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(20) if aid2 not in unique_aids]    \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    # USE TOP20 TEST CLICKS\n    return result + list(top_clicks)[:20-len(result)]\n\ndef suggest_buys(df):\n    # USER HISTORY AIDS AND TYPES\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    # UNIQUE AIDS AND UNIQUE BUYS\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    df = df.loc[(df['type']==1)|(df['type']==2)]\n    unique_buys = list(dict.fromkeys( df.aid.tolist()[::-1] ))\n    # RERANK CANDIDATES USING WEIGHTS\n    if len(unique_aids)>=20:\n        weights=np.logspace(0.5,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        # RERANK BASED ON REPEAT ITEMS AND TYPE OF ITEMS\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        # RERANK CANDIDATES USING \"BUY2BUY\" CO-VISITATION MATRIX\n        aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n        for aid in aids3: aids_temp[aid] += 0.1\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    # USE \"CART ORDER\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_aids if aid in top_20_buys]))\n    # USE \"BUY2BUY\" CO-VISITATION MATRIX\n    aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2+aids3).most_common(20) if aid2 not in unique_aids] \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    # USE TOP20 TEST ORDERS\n    return result + list(top_orders)[:20-len(result)]","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[],"execution":{"iopub.status.busy":"2025-10-26T20:40:41.168076Z","iopub.execute_input":"2025-10-26T20:40:41.168494Z","iopub.status.idle":"2025-10-26T20:40:41.214294Z","shell.execute_reply.started":"2025-10-26T20:40:41.168456Z","shell.execute_reply":"2025-10-26T20:40:41.213329Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# —— 保存基线提交，避免被后续 Notebook 覆盖 ——\nimport shutil, glob, os, pandas as pd\n\n# 1) 备份规则版提交\nif os.path.exists(\"submission.csv\"):\n    shutil.copy(\"submission.csv\", \"submission_rules.csv\")\n\n# 2) 列出将要复用的共现矩阵产物（都在 /kaggle/working 下，会随版本提交成为 Outputs）\nartifacts = sorted(glob.glob(\"top_20_clicks_v*.pqt\")) \\\n         +  sorted(glob.glob(\"top_15_carts_orders_v*.pqt\")) \\\n         +  sorted(glob.glob(\"top_15_buy2buy_v*.pqt\"))\nprint(\"Artifacts to reuse:\", len(artifacts))\npd.Series(artifacts).to_csv(\"artifact_list.txt\", index=False, header=False)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-26T20:40:41.217908Z","iopub.execute_input":"2025-10-26T20:40:41.218461Z","iopub.status.idle":"2025-10-26T20:40:41.231144Z","shell.execute_reply.started":"2025-10-26T20:40:41.218435Z","shell.execute_reply":"2025-10-26T20:40:41.23026Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# XGboost","metadata":{}},{"cell_type":"markdown","source":"## Cell 1 — 备份当前规则提交 + 安装兼容版 XGBoost + 修正热门统计","metadata":{}},{"cell_type":"code","source":"# 备份当前规则提交，避免被覆盖\nimport os, shutil\nif os.path.exists(\"submission.csv\"):\n    shutil.copy(\"submission.csv\", \"submission_rules.csv\")\n    print(\"Backed up baseline to submission_rules.csv\")\n\n# 安装与 Py3.7 兼容的 xgboost 版本\n!pip -q install xgboost==1.6.2\nimport xgboost as xgb, numpy as np, pandas as pd\nprint(\"xgboost version:\", xgb.__version__)\n\n# 你的 test_df 里 'type' 已经映射成 int8(0/1/2)，\n# 这里修正之前用字符串统计热门导致取空的问题\ntop_clicks = test_df.loc[test_df['type']==0, 'aid'].value_counts().index.values[:20]\ntop_orders = test_df.loc[test_df['type']==2, 'aid'].value_counts().index.values[:20]\ntop_clicks_set, top_orders_set = set(top_clicks), set(top_orders)\nprint(\"Top clicks/orders ready.\", len(top_clicks), len(top_orders))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-26T20:41:32.222958Z","iopub.execute_input":"2025-10-26T20:41:32.223406Z","iopub.status.idle":"2025-10-26T20:41:40.532038Z","shell.execute_reply.started":"2025-10-26T20:41:32.223361Z","shell.execute_reply":"2025-10-26T20:41:40.530884Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Cell 2 — 读取带权重 & 名次的共现矩阵（直接用你刚写出的 parquet）","metadata":{}},{"cell_type":"code","source":"import glob\n\ndef pqt_to_weighted_and_rank(paths):\n    from collections import defaultdict\n    weights, ranks = defaultdict(dict), defaultdict(dict)\n    for p in paths:\n        df = pd.read_parquet(p, columns=['aid_x','aid_y','wgt'])\n        df = df.sort_values(['aid_x','wgt'], ascending=[True, False])\n        df['rank'] = df.groupby('aid_x').cumcount().astype('int32')\n        for row in df.itertuples(index=False):\n            # row.aid_x, row.aid_y, row.wgt, row.rank 都是数值字段，不会与方法名冲突\n            ax, ay = int(row.aid_x), int(row.aid_y)\n            weights[ax][ay] = float(row.wgt)\n            ranks[ax][ay]   = int(row.rank)\n    return dict(weights), dict(ranks)\n\n\n# 自动发现当前工作目录下的共现产物（不用依赖 DISK_PIECES）\nCLICK_PQTS = sorted(glob.glob(\"top_20_clicks_v*_*.pqt\"))\nBUYS_PQTS  = sorted(glob.glob(\"top_15_carts_orders_v*_*.pqt\"))\nB2B_PQTS   = sorted(glob.glob(\"top_15_buy2buy_v*_*.pqt\"))\nprint(\"Found:\", len(CLICK_PQTS), \"clicks;\", len(BUYS_PQTS), \"carts_orders;\", len(B2B_PQTS), \"buy2buy\")\n\nclick_w, click_r = pqt_to_weighted_and_rank(CLICK_PQTS)\nbuys_w,  buys_r  = pqt_to_weighted_and_rank(BUYS_PQTS)\nb2b_w,   b2b_r   = pqt_to_weighted_and_rank(B2B_PQTS)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-26T20:43:01.716871Z","iopub.execute_input":"2025-10-26T20:43:01.717715Z","iopub.status.idle":"2025-10-26T20:45:22.633673Z","shell.execute_reply.started":"2025-10-26T20:43:01.717682Z","shell.execute_reply":"2025-10-26T20:45:22.632893Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Cell 3 — 候选与特征（复用你已有逻辑，最小实现）","metadata":{}},{"cell_type":"code","source":"from collections import Counter\nimport itertools\n\nTYPE_W = {0:1, 1:6, 2:3}\n\ndef build_session_seeds(df_hist):\n    aids  = df_hist.aid.tolist()\n    types = df_hist.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1]))\n    unique_buys = list(dict.fromkeys(df_hist[df_hist['type'].isin([1,2])].aid.tolist()[::-1]))\n    return aids, types, unique_aids, unique_buys\n\ndef recent_type_score(aids, types, lo=0.1, hi=1.0):\n    w = np.logspace(lo, hi, len(aids), base=2, endpoint=True) - 1.0\n    sc = Counter()\n    for aid, ww, t in zip(aids, w, types):\n        sc[aid] += float(ww) * TYPE_W[int(t)]\n    return sc\n\ndef aggregate_neighbors(unique_aids, unique_buys):\n    agg_click, agg_buys, agg_b2b = Counter(), Counter(), Counter()\n    for a in unique_aids:\n        for b,w in click_w.get(a, {}).items(): agg_click[b] += w\n        for b,w in buys_w.get(a,  {}).items(): agg_buys[b]  += w\n    for a in unique_buys:\n        for b,w in b2b_w.get(a,   {}).items(): agg_b2b[b]   += w\n    return agg_click, agg_buys, agg_b2b\n\ndef build_candidates(df_hist, k_click=120, k_buys=120):\n    aids, types, unique_aids, unique_buys = build_session_seeds(df_hist)\n    agg_click, agg_buys, agg_b2b = aggregate_neighbors(unique_aids, unique_buys)\n    cand_clicks = set(unique_aids) | set([b for b,_ in agg_click.most_common(k_click)])\n    cand_buys   = set(unique_aids) | set([b for b,_ in (agg_buys+agg_b2b).most_common(k_buys)])\n    return (aids, types, unique_aids, unique_buys, agg_click, agg_buys, agg_b2b, cand_clicks, cand_buys)\n\ndef best_rank_over_seeds(c, seeds, ranks, default=9999):\n    best = default\n    for s in seeds:\n        r = ranks.get(s, {}).get(c, default)\n        if r < best: best = r\n    return best\n\ndef make_features_for_head(df_hist, head, candidates,\n                           aids, types, unique_aids, unique_buys,\n                           agg_click, agg_buys, agg_b2b):\n    rec_sc  = recent_type_score(aids, types)\n    sesslen = len(aids)\n    last_ts = int(df_hist.ts.max()) if sesslen else 0\n    rows, keys = [], []\n    for c in candidates:\n        in_hist = 1 if c in rec_sc else 0\n        w_click = float(agg_click.get(c, 0.0))\n        w_buys  = float(agg_buys.get(c, 0.0))\n        w_b2b   = float(agg_buys.get(c, 0.0))  # 也可以用 agg_buys+agg_b2b 合并后的\n        r_click = best_rank_over_seeds(c, unique_aids, click_r)\n        r_buys  = best_rank_over_seeds(c, unique_aids, buys_r)\n        r_b2b   = best_rank_over_seeds(c, unique_buys, b2b_r)\n        f_best_click = 1.0/(1.0 + r_click)\n        f_best_buys  = 1.0/(1.0 + r_buys)\n        f_best_b2b   = 1.0/(1.0 + r_b2b)\n        f_rec_sc = float(rec_sc.get(c, 0.0))\n        if c in df_hist.aid.values:\n            gap = last_ts - int(df_hist[df_hist.aid==c].ts.max())\n        else:\n            gap = 10*24*3600\n        is_top_click = int(c in top_clicks_set)\n        is_top_order = int(c in top_orders_set)\n        if head=='clicks':\n            rows.append([in_hist, f_rec_sc, w_click, f_best_click,\n                         w_buys, f_best_buys, f_best_b2b,\n                         sesslen, gap, is_top_click, is_top_order])\n        else:\n            rows.append([in_hist, f_rec_sc, w_buys, f_best_buys,\n                         w_click, f_best_click, f_best_b2b,\n                         sesslen, gap, is_top_click, is_top_order])\n        keys.append(c)\n    X = pd.DataFrame(rows, columns=[\n        'in_hist','rec_rule',\n        'w_main','best_main',\n        'w_aux1','best_aux1',\n        'best_b2b',  # 简化：只用名次特征，足够稳定\n        'sess_len','gap',\n        'is_top_click','is_top_order'\n    ])\n    return X, keys\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-26T20:45:41.710836Z","iopub.execute_input":"2025-10-26T20:45:41.711494Z","iopub.status.idle":"2025-10-26T20:45:42.00692Z","shell.execute_reply.started":"2025-10-26T20:45:41.711454Z","shell.execute_reply":"2025-10-26T20:45:42.006042Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Cell 4 — 轻量训练集（会话最后 24h 当“未来”）& 训练模型","metadata":{}},{"cell_type":"code","source":"from tqdm import tqdm\nimport glob, random\n\ndef build_train_rows_from_session(df_sess):\n    if df_sess.empty: return [], []\n    split_ts = int(df_sess.ts.max()) - 24*60*60\n    hist = df_sess[df_sess.ts <  split_ts]\n    fut  = df_sess[df_sess.ts >= split_ts]\n    if hist.empty or fut.empty: return [], []\n    (aids, types, unique_aids, unique_buys,\n     agg_click, agg_buys, agg_b2b,\n     cand_clicks, cand_buys) = build_candidates(hist)\n    labs_clicks = set(fut.loc[fut['type']==0,'aid'].tolist())\n    labs_buys   = set(fut.loc[fut['type'].isin([1,2]),'aid'].tolist())\n    Xc, kc = make_features_for_head(hist, 'clicks', cand_clicks,\n                                    aids, types, unique_aids, unique_buys,\n                                    agg_click, agg_buys, agg_b2b)\n    yc = [1 if a in labs_clicks else 0 for a in kc]\n    Xb, kb = make_features_for_head(hist, 'buys', cand_buys,\n                                    aids, types, unique_aids, unique_buys,\n                                    agg_click, agg_buys, agg_b2b)\n    yb = [1 if a in labs_buys else 0 for a in kb]\n    return [(Xc, yc)], [(Xb, yb)]\n\ndef build_training_dataset(train_glob=\"../input/otto-chunk-data-inparquet-format/*_parquet/*\",\n                           max_sessions=120000, seed=42):\n    files = glob.glob(train_glob)\n    random.seed(seed); random.shuffle(files)\n    Xc_list, yc_list, Xb_list, yb_list = [], [], [], []\n    seen = 0\n    with tqdm(total=max_sessions, desc=\"Build train set\") as pbar:\n        for f in files:\n            df = pd.read_parquet(f)\n            df['ts'] = (df['ts']/1000).astype('int32') if df['ts'].max()>1e10 else df['ts'].astype('int32')\n            if df['type'].dtype=='O':\n                df['type'] = df['type'].map({'clicks':0,'carts':1,'orders':2}).astype('int8')\n            for _, g in df.groupby('session'):\n                xc, xb = build_train_rows_from_session(g)\n                for Xc,yc in xc: Xc_list.append(Xc); yc_list += yc\n                for Xb,yb in xb: Xb_list.append(Xb); yb_list += yb\n                seen += 1; pbar.update(1)\n                if seen >= max_sessions: break\n            if seen >= max_sessions: break\n    Xc = pd.concat(Xc_list, ignore_index=True) if Xc_list else pd.DataFrame()\n    Xb = pd.concat(Xb_list, ignore_index=True) if Xb_list else pd.DataFrame()\n    return Xc, np.array(yc_list, np.int8), Xb, np.array(yb_list, np.int8)\n\nXc, yc, Xb, yb = build_training_dataset(max_sessions=100000)  # 可调小试跑：比如 50_000\nprint(\"Clicks train:\", Xc.shape, \"pos rate:\", float(yc.mean()) if len(yc) else None)\nprint(\"Buys   train:\", Xb.shape, \"pos rate:\", float(yb.mean()) if len(yb) else None)\n\nclf_clicks = xgb.XGBClassifier(\n    n_estimators=300, max_depth=7, learning_rate=0.08,\n    subsample=0.9, colsample_bytree=0.9,\n    reg_alpha=1e-3, reg_lambda=1.0,\n    tree_method='hist', random_state=2025, n_jobs=-1\n)\nclf_buys = xgb.XGBClassifier(\n    n_estimators=300, max_depth=7, learning_rate=0.08,\n    subsample=0.9, colsample_bytree=0.9,\n    reg_alpha=1e-3, reg_lambda=1.0,\n    tree_method='hist', random_state=2025, n_jobs=-1\n)\n\nif (not Xc.empty) and (len(np.unique(yc))>1):\n    clf_clicks.fit(Xc, yc)\nif (not Xb.empty) and (len(np.unique(yb))>1):\n    clf_buys.fit(Xb, yb)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-26T20:45:45.55043Z","iopub.execute_input":"2025-10-26T20:45:45.551259Z","iopub.status.idle":"2025-10-26T21:08:31.221211Z","shell.execute_reply.started":"2025-10-26T20:45:45.551221Z","shell.execute_reply":"2025-10-26T21:08:31.220168Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Cell 5 — ML 排序函数（失败则自动回退到你原先的规则）","metadata":{}},{"cell_type":"code","source":"def rank_and_take20(model, X, keys, fallback_sorted):\n    if model is None or X.empty:\n        return fallback_sorted[:20]\n    s = model.predict_proba(X)[:,1] if hasattr(model, \"predict_proba\") else model.predict(X)\n    order = np.argsort(-s)\n    ranked = [keys[i] for i in order]\n    out, used = [], set()\n    for a in ranked + fallback_sorted:  # 兜底不丢\n        if a not in used:\n            out.append(a); used.add(a)\n        if len(out)==20: break\n    return out\n\ndef suggest_clicks_ml(df):\n    (aids, types, unique_aids, unique_buys,\n     agg_click, agg_buys, agg_b2b,\n     cand_clicks, _) = build_candidates(df)\n    Xc, kc = make_features_for_head(df, 'clicks', cand_clicks,\n                                    aids, types, unique_aids, unique_buys,\n                                    agg_click, agg_buys, agg_b2b)\n    fallback = suggest_clicks(df)  # 复用你现有的规则函数\n    return rank_and_take20(clf_clicks if 'clf_clicks' in globals() else None, Xc, kc, fallback)\n\ndef suggest_buys_ml(df):\n    (aids, types, unique_aids, unique_buys,\n     agg_click, agg_buys, agg_b2b,\n     _, cand_buys) = build_candidates(df)\n    Xb, kb = make_features_for_head(df, 'buys', cand_buys,\n                                    aids, types, unique_aids, unique_buys,\n                                    agg_click, agg_buys, agg_b2b)\n    fallback = suggest_buys(df)\n    return rank_and_take20(clf_buys if 'clf_buys' in globals() else None, Xb, kb, fallback)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-26T21:09:07.745863Z","iopub.execute_input":"2025-10-26T21:09:07.746355Z","iopub.status.idle":"2025-10-26T21:09:07.755568Z","shell.execute_reply.started":"2025-10-26T21:09:07.746323Z","shell.execute_reply":"2025-10-26T21:09:07.754702Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Cell 6 — 批量特征生成 + 一次性模型推理","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom collections import defaultdict\nimport time, psutil\n\ndef _rank_take20_with_fallback(scores_keys_per_sess, fallback_per_sess):\n    \"\"\"\n    scores_keys_per_sess: dict[session] -> list[(score, aid)]\n    fallback_per_sess   : dict[session] -> list[aid]  (规则版候选顺序)\n    return: dict[session] -> list[20 aids]\n    \"\"\"\n    out = {}\n    for sid, pairs in scores_keys_per_sess.items():\n        pairs.sort(key=lambda x: -x[0])  # score 降序\n        used, res = set(), []\n        for _, aid in pairs:\n            if aid not in used:\n                res.append(aid); used.add(aid)\n            if len(res) == 20: break\n        # 兜底：补满 20\n        if len(res) < 20:\n            for aid in fallback_per_sess.get(sid, []):\n                if aid not in used:\n                    res.append(aid); used.add(aid)\n                if len(res) == 20: break\n        out[sid] = res\n    return out\n\ndef _batch_predict_head(head, df_sorted, batch_sessions=120000):\n    \"\"\"\n    head: 'clicks' or 'buys'\n    df_sorted: test_df 已按 [\"session\",\"ts\"] 排好\n    batch_sessions: 每批处理 session 数\n    \"\"\"\n    assert head in (\"clicks\", \"buys\")\n    model = clf_clicks if head == \"clicks\" else clf_buys\n    use_model = (model is not None) and hasattr(model, \"predict_proba\")\n\n    preds_all = {}\n    sessions = df_sorted[\"session\"].drop_duplicates().values\n    N = len(sessions)\n\n    print(f\"\\n🚀 Start rerank inference for [{head}] on {N:,} sessions, batch size {batch_sessions}\")\n    t0_total = time.time()\n\n    for start in range(0, N, batch_sessions):\n        end = min(N, start + batch_sessions)\n        sess_batch = set(sessions[start:end])\n        batch_id = start // batch_sessions + 1\n\n        t0 = time.time()\n        print(f\"\\n🧩 Batch {batch_id}: sessions {start:,} ~ {end-1:,} \"\n              f\"({end-start:,} total) | {head.upper()} head\")\n\n        g = df_sorted[df_sorted[\"session\"].isin(sess_batch)].groupby(\"session\", sort=False)\n\n        X_rows, key_rows, fallback = [], [], {}\n\n        # Step 1: 构建候选 + 特征\n        for sid, df_sess in g:\n            (aids, types, unique_aids, unique_buys,\n             agg_click, agg_buys, agg_b2b,\n             cand_clicks, cand_buys) = build_candidates(df_sess)\n\n            if head == \"clicks\":\n                X, keys = make_features_for_head(\n                    df_sess, \"clicks\", cand_clicks,\n                    aids, types, unique_aids, unique_buys,\n                    agg_click, agg_buys, agg_b2b\n                )\n                fallback[sid] = suggest_clicks(df_sess)\n            else:\n                X, keys = make_features_for_head(\n                    df_sess, \"buys\", cand_buys,\n                    aids, types, unique_aids, unique_buys,\n                    agg_click, agg_buys, agg_b2b\n                )\n                fallback[sid] = suggest_buys(df_sess)\n\n            if not X.empty:\n                X_rows.append(X)\n                key_rows.extend([(sid, a) for a in keys])\n\n        # Step 2: 模型批量推理\n        if not X_rows:\n            print(\"⚠️  No features in this batch, skipping.\")\n            continue\n\n        X_all = pd.concat(X_rows, ignore_index=True)\n        if use_model:\n            print(f\"🔮  Running model on {len(X_all):,} feature rows …\")\n            scores = model.predict_proba(X_all)[:, 1]\n        else:\n            print(\"⚙️  Model not found, using fallback only.\")\n            scores = np.zeros(len(X_all), dtype=np.float32)\n\n        # Step 3: 汇总结果\n        scores_per_sess = defaultdict(list)\n        for (sid, aid), sc in zip(key_rows, scores):\n            scores_per_sess[sid].append((float(sc), int(aid)))\n\n        batch_pred = _rank_take20_with_fallback(scores_per_sess, fallback)\n        preds_all.update(batch_pred)\n\n        # Step 4: 打印批次耗时与内存\n        used_mem = psutil.virtual_memory().used / 1024**3\n        print(f\"✅ Batch {batch_id} done in {time.time()-t0:.1f}s | \"\n              f\"Mem used: {used_mem:.2f} GB\")\n\n        # 手动清理\n        del X_rows, key_rows, X_all, scores_per_sess, batch_pred\n\n    print(f\"\\n🏁 Finished [{head}] rerank: total {time.time()-t0_total:.1f}s\")\n    return preds_all\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-27T00:24:11.772706Z","iopub.execute_input":"2025-10-27T00:24:11.773305Z","iopub.status.idle":"2025-10-27T00:24:11.790536Z","shell.execute_reply.started":"2025-10-27T00:24:11.773266Z","shell.execute_reply":"2025-10-27T00:24:11.789559Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 分别批量预测 clicks/buys → 组装提交（一次性输出）","metadata":{}},{"cell_type":"code","source":"%%time\n# 先对 test_df 排序一次即可\ntest_sorted = test_df.sort_values([\"session\",\"ts\"], kind=\"mergesort\")\n\n# 分别批量推理两个头\nclicks_top20 = _batch_predict_head(\"clicks\", test_sorted, batch_sessions=120000)\nbuys_top20   = _batch_predict_head(\"buys\",   test_sorted, batch_sessions=120000)\n\n# 组装为提交 DataFrame\ndef _to_submission_block(d, suffix):\n    # d: dict[session] -> list[20 aids]\n    s = pd.Series({f\"{sid}_{suffix}\": \" \".join(map(str, aids)) for sid, aids in d.items()})\n    df = s.rename_axis(\"session_type\").reset_index(name=\"labels\")\n    return df\n\nclicks_pred_df = _to_submission_block(clicks_top20, \"clicks\")\norders_pred_df = _to_submission_block(buys_top20,   \"orders\")\ncarts_pred_df  = _to_submission_block(buys_top20,   \"carts\")\n\npred_df_ml = pd.concat([clicks_pred_df, orders_pred_df, carts_pred_df], ignore_index=True)\npred_df_ml.to_csv(\"submission_ml.csv\", index=False)\n\n# 快速自检\nok_ratio = (pred_df_ml[\"labels\"].str.split().map(len) == 20).mean()\nprint(\"submission_ml.csv ready. 20-per-line ratio:\", f\"{ok_ratio:.3f}\")\npred_df_ml.head()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-27T00:24:14.99414Z","iopub.execute_input":"2025-10-27T00:24:14.994793Z","execution_failed":"2025-10-27T04:33:35.273Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Cell 7 — 快速自检（每行 20 个、无空值）","metadata":{}},{"cell_type":"code","source":"# 简单检查\nn_rows = pred_df_ml.shape[0]\nok_len = (pred_df_ml['labels'].str.split().map(len) == 20).mean()\nprint(f\"Rows: {n_rows}, 20-per-line ratio: {ok_len:.3f}\")\n# 两份提交都在：submission_rules.csv（基线） + submission_ml.csv（ML）\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-10-26T20:41:45.34597Z","iopub.status.idle":"2025-10-26T20:41:45.346737Z","shell.execute_reply.started":"2025-10-26T20:41:45.346472Z","shell.execute_reply":"2025-10-26T20:41:45.346504Z"}},"outputs":[],"execution_count":null}]}