{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":38760,"databundleVersionId":4493939,"sourceType":"competition"},{"sourceId":4436180,"sourceType":"datasetVersion","datasetId":2597726}],"dockerImageVersionId":30302,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Modèle de réévaluation des candidats utilisant des règles artisanales (Handcrafted Rules)\nDans ce notebook, on présente un modèle de \"rerank Candidate\" en utilisant des règles faites à la main. On peut améliorer ce modèle en créant des caractéristiques, en les fusionnant avec les articles et les utilisateurs, et en formant un modèle de rerank (comme XGB) pour choisir nos 20 finaux. De plus, pour ajuster et améliorer ce notebook, on devrait construire un schéma de CV local pour expérimenter une nouvelle logique et/ou des modèles.\n\nDans cette compétition, une \"session\" signifie en réalité un \"utilisateur\" unique. Notre tâche est de prédire ce que chacun des 1 671 803 \"utilisateurs\" de test (c'est-à-dire \"sessions\") fera à l'avenir. Pour chaque \"utilisateur\" de test (c'est-à-dire \"session\"), nous devons prédire ce qu'ils vont cliquer, mettre dans le panier et commander pendant le reste de la période de test d'une semaine.\n\n### Étape 1 - Générer des candidats\nPour chaque utilisateur de test, nous générons des choix possibles, c'est-à-dire des candidats. Dans ce Notebook, nous générons des candidats à partir de 5 sources :\n* Historique de l'utilisateur clics, paniers, commandes\n* Les 20 clics, paniers, commandes les plus populaires de la semaine de test\n* Matrice de co-visite des clics/paniers/commandes vers panier/commande avec pondération par type\n* Matrice de co-visite des paniers/commandes vers panier/commande appelée buy2buy\n* Matrice de co-visite des clics/paniers/commandes vers les clics avec pondération temporelle\n\n### Étape 2 - Reclassement et Choix de 20\nÉtant donné la liste des candidats, nous devons en sélectionner 20 pour nos prédictions. Dans ce Notebook, nous le faisons avec un ensemble de règles artisanales. Nous pouvons améliorer nos prédictions en entraînant un modèle XGBoost pour sélectionner à notre place. Nos règles artisanales donnent la priorité à:\n    * Articles récemment consultés\n    * Articles consultés à plusieurs reprises auparavant\n    * Articles précédemment dans le panier ou la commande\n    * Matrice de co-visite du panier/commande au panier/commande\n    * Articles populaires actuellement\n\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/c_r_model.png)\n  \n","metadata":{"papermill":{"duration":0.005198,"end_time":"2022-11-10T16:03:20.966987","exception":false,"start_time":"2022-11-10T16:03:20.961789","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"### Étape 1 - Génération de candidats avec RAPIDS\nPour la génération de candidats, nous construisons trois matrices de co-visite. L'une calcule la popularité du panier/commande étant donné un clic/panier/commande précédent de l'utilisateur. Nous appliquons une pondération par type à cette matrice. Une autre calcule la popularité du panier/commande étant donné un panier/commande précédent de l'utilisateur. Nous appelons cela la matrice \"bytoby\". Une autre calcule la popularité des clics étant donné un clic/panier/commande précédent de l'utilisateur. Nous appliquons une pondération temporelle à cette matrice. Nous utiliserons RAPIDS cuDF GPU pour calculer rapidement ces matrices!","metadata":{"papermill":{"duration":0.00373,"end_time":"2022-11-10T16:03:20.9748","exception":false,"start_time":"2022-11-10T16:03:20.97107","status":"completed"},"tags":[]}},{"cell_type":"code","source":"VER = 5\n\nimport pandas as pd, numpy as np\nfrom tqdm.notebook import tqdm\nimport os, sys, pickle, glob, gc\nfrom collections import Counter\nimport cudf, itertools\nprint('We will use RAPIDS version',cudf.__version__)","metadata":{"papermill":{"duration":3.036143,"end_time":"2022-11-10T16:03:24.014816","exception":false,"start_time":"2022-11-10T16:03:20.978673","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-18T10:11:02.78558Z","iopub.execute_input":"2024-05-18T10:11:02.786333Z","iopub.status.idle":"2024-05-18T10:11:02.79278Z","shell.execute_reply.started":"2024-05-18T10:11:02.786294Z","shell.execute_reply":"2024-05-18T10:11:02.791735Z"},"trusted":true},"execution_count":4,"outputs":[{"name":"stdout","text":"We will use RAPIDS version 21.10.01\n","output_type":"stream"}]},{"cell_type":"markdown","source":"## Calculer trois matrices de co-visite avec RAPIDS\nNous allons calculer 3 matrices de co-visitation en utilisant RAPIDS cuDF sur GPU. C'est 30 fois plus rapide que d'utiliser Pandas CPU comme dans d'autres notebooks publics ! Pour une vitesse maximale, définissez la variable `DISK_PIECES` au plus petit nombre possible en fonction du GPU que vous utilisez sans provoquer d'erreurs mémoire. Si vous exécutez ce code hors ligne avec 32 Go de RAM GPU, vous pouvez utiliser `DISK_PIECES = 1` et calculer chaque matrice de co-visitation en presque 1 minute ! Le GPU de Kaggle a seulement 16 Go de RAM, donc nous utilisons `DISK_PIECES = 4` et cela prend incroyablement 3 minutes chacune ! Voici quelques astuces pour accélérer le calcul\n* Utilisez RAPIDS cuDF GPU au lieu de Pandas CPU\n* Lisez le disque une fois et enregistrez dans la RAM du CPU pour une utilisation ultérieure sur le GPU\n* Traitez la plus grande quantité de données possible sur le GPU en une seule fois\n* Fusionnez les données en deux étapes. De multiples petites à une moyenne unique. De multiples moyennes à une grande unique.\n* Enregistrez le résultat en tant que parquet au lieu de dictionnaire","metadata":{"papermill":{"duration":0.00424,"end_time":"2022-11-10T16:03:24.023816","exception":false,"start_time":"2022-11-10T16:03:24.019576","status":"completed"},"tags":[]}},{"cell_type":"code","source":"%%time\n# CACHE FUNCTIONS\ndef read_file(f):\n    return cudf.DataFrame( data_cache[f] )\ndef read_file_to_cache(f):\n    df = pd.read_parquet(f)\n    df.ts = (df.ts/1000).astype('int32')\n    df['type'] = df['type'].map(type_labels).astype('int8')\n    return df\n\n# CACHE THE DATA ON CPU BEFORE PROCESSING ON GPU\ndata_cache = {}\ntype_labels = {'clicks':0, 'carts':1, 'orders':2}\nfiles = glob.glob('/kaggle/input/otto-chunk-data-inparquet-format/*_parquet/*')\nfor f in files: data_cache[f] = read_file_to_cache(f)\n\n# CHUNK PARAMETERS\nREAD_CT = 5\nCHUNK = int( np.ceil( len(files)/6 ))\nprint(f'We will process {len(files)} files, in groups of {READ_CT} and chunks of {CHUNK}.')","metadata":{"papermill":{"duration":0.063943,"end_time":"2022-11-10T16:03:24.091816","exception":false,"start_time":"2022-11-10T16:03:24.027873","status":"completed"},"tags":[],"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-05-18T10:11:10.701455Z","iopub.execute_input":"2024-05-18T10:11:10.702712Z","iopub.status.idle":"2024-05-18T10:11:53.62881Z","shell.execute_reply.started":"2024-05-18T10:11:10.70266Z","shell.execute_reply":"2024-05-18T10:11:53.627649Z"},"trusted":true},"execution_count":5,"outputs":[{"name":"stdout","text":"We will process 146 files, in groups of 5 and chunks of 25.\nCPU times: user 58.2 s, sys: 8.04 s, total: 1min 6s\nWall time: 42.9 s\n","output_type":"stream"}]},{"cell_type":"markdown","source":"## 1) \"Carts Orders\" Matrice de co-visite - Pondération par type","metadata":{"papermill":{"duration":0.004089,"end_time":"2022-11-10T16:03:24.100502","exception":false,"start_time":"2022-11-10T16:03:24.096413","status":"completed"},"tags":[]}},{"cell_type":"code","source":"%%time\ntype_weight = {0:1, 1:6, 2:3}\n\n# USE SMALLEST DISK_PIECES POSSIBLE WITHOUT MEMORY ERROR\nDISK_PIECES = 4\nSIZE = 1.86e6/DISK_PIECES\n\n# COMPUTE IN PARTS FOR MEMORY MANGEMENT\nfor PART in range(DISK_PIECES):\n    print()\n    print('### DISK PART',PART+1)\n    \n    # MERGE IS FASTEST PROCESSING CHUNKS WITHIN CHUNKS\n    # => OUTER CHUNKS\n    for j in range(6):\n        a = j*CHUNK\n        b = min( (j+1)*CHUNK, len(files) )\n        print(f'Processing files {a} thru {b-1} in groups of {READ_CT}...')\n        \n        # => INNER CHUNKS\n        for k in range(a,b,READ_CT):\n            # READ FILE\n            df = [read_file(files[k])]\n            for i in range(1,READ_CT): \n                if k+i<b: df.append( read_file(files[k+i]) )\n            df = cudf.concat(df,ignore_index=True,axis=0)\n            df = df.sort_values(['session','ts'],ascending=[True,False])\n            # USE TAIL OF SESSION\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n<30].drop('n',axis=1)\n            # CREATE PAIRS\n            df = df.merge(df,on='session')\n            df = df.loc[ ((df.ts_x - df.ts_y).abs()< 24 * 60 * 60) & (df.aid_x != df.aid_y) ]\n            # MEMORY MANAGEMENT COMPUTE IN PARTS\n            df = df.loc[(df.aid_x >= PART*SIZE)&(df.aid_x < (PART+1)*SIZE)]\n            # ASSIGN WEIGHTS\n            df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n            df['wgt'] = df.type_y.map(type_weight)\n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n            # COMBINE INNER CHUNKS\n            if k==a: tmp2 = df\n            else: tmp2 = tmp2.add(df, fill_value=0)\n            print(k,', ',end='')\n        print()\n        # COMBINE OUTER CHUNKS\n        if a==0: tmp = tmp2\n        else: tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    # CONVERT MATRIX TO DICTIONARY\n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n    # SAVE TOP 40\n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    tmp = tmp.loc[tmp.n<15].drop('n',axis=1)\n    # SAVE PART TO DISK (convert to pandas first uses less memory)\n    tmp.to_pandas().to_parquet(f'top_15_carts_orders_v{VER}_{PART}.pqt')","metadata":{"papermill":{"duration":566.561189,"end_time":"2022-11-10T16:12:50.666123","exception":false,"start_time":"2022-11-10T16:03:24.104934","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-18T10:12:07.848467Z","iopub.execute_input":"2024-05-18T10:12:07.848844Z","iopub.status.idle":"2024-05-18T10:15:22.640369Z","shell.execute_reply.started":"2024-05-18T10:12:07.848814Z","shell.execute_reply":"2024-05-18T10:15:22.639231Z"},"trusted":true},"execution_count":6,"outputs":[{"name":"stdout","text":"\n### DISK PART 1\nProcessing files 0 thru 24 in groups of 5...\n0 , 5 , 10 , 15 , 20 , \nProcessing files 25 thru 49 in groups of 5...\n25 , 30 , 35 , 40 , 45 , \nProcessing files 50 thru 74 in groups of 5...\n50 , 55 , 60 , 65 , 70 , \nProcessing files 75 thru 99 in groups of 5...\n75 , 80 , 85 , 90 , 95 , \nProcessing files 100 thru 124 in groups of 5...\n100 , 105 , 110 , 115 , 120 , \nProcessing files 125 thru 145 in groups of 5...\n125 , 130 , 135 , 140 , 145 , \n\n### DISK PART 2\nProcessing files 0 thru 24 in groups of 5...\n0 , 5 , 10 , 15 , 20 , \nProcessing files 25 thru 49 in groups of 5...\n25 , 30 , 35 , 40 , 45 , \nProcessing files 50 thru 74 in groups of 5...\n50 , 55 , 60 , 65 , 70 , \nProcessing files 75 thru 99 in groups of 5...\n75 , 80 , 85 , 90 , 95 , \nProcessing files 100 thru 124 in groups of 5...\n100 , 105 , 110 , 115 , 120 , \nProcessing files 125 thru 145 in groups of 5...\n125 , 130 , 135 , 140 , 145 , \n\n### DISK PART 3\nProcessing files 0 thru 24 in groups of 5...\n0 , 5 , 10 , 15 , 20 , \nProcessing files 25 thru 49 in groups of 5...\n25 , 30 , 35 , 40 , 45 , \nProcessing files 50 thru 74 in groups of 5...\n50 , 55 , 60 , 65 , 70 , \nProcessing files 75 thru 99 in groups of 5...\n75 , 80 , 85 , 90 , 95 , \nProcessing files 100 thru 124 in groups of 5...\n100 , 105 , 110 , 115 , 120 , \nProcessing files 125 thru 145 in groups of 5...\n125 , 130 , 135 , 140 , 145 , \n\n### DISK PART 4\nProcessing files 0 thru 24 in groups of 5...\n0 , 5 , 10 , 15 , 20 , \nProcessing files 25 thru 49 in groups of 5...\n25 , 30 , 35 , 40 , 45 , \nProcessing files 50 thru 74 in groups of 5...\n50 , 55 , 60 , 65 , 70 , \nProcessing files 75 thru 99 in groups of 5...\n75 , 80 , 85 , 90 , 95 , \nProcessing files 100 thru 124 in groups of 5...\n100 , 105 , 110 , 115 , 120 , \nProcessing files 125 thru 145 in groups of 5...\n125 , 130 , 135 , 140 , 145 , \nCPU times: user 2min 9s, sys: 1min 5s, total: 3min 14s\nWall time: 3min 14s\n","output_type":"stream"}]},{"cell_type":"markdown","source":"## 2) \"Buy2Buy\" Co-visitation Matrix","metadata":{"papermill":{"duration":0.03219,"end_time":"2022-11-10T16:12:50.730634","exception":false,"start_time":"2022-11-10T16:12:50.698444","status":"completed"},"tags":[]}},{"cell_type":"code","source":"%%time\n# USE SMALLEST DISK_PIECES POSSIBLE WITHOUT MEMORY ERROR\nDISK_PIECES = 1\nSIZE = 1.86e6/DISK_PIECES\n\n# COMPUTE IN PARTS FOR MEMORY MANGEMENT\nfor PART in range(DISK_PIECES):\n    print()\n    print('### DISK PART',PART+1)\n    \n    # MERGE IS FASTEST PROCESSING CHUNKS WITHIN CHUNKS\n    # => OUTER CHUNKS\n    for j in range(6):\n        a = j*CHUNK\n        b = min( (j+1)*CHUNK, len(files) )\n        print(f'Processing files {a} thru {b-1} in groups of {READ_CT}...')\n        \n        # => INNER CHUNKS\n        for k in range(a,b,READ_CT):\n            # READ FILE\n            df = [read_file(files[k])]\n            for i in range(1,READ_CT): \n                if k+i<b: df.append( read_file(files[k+i]) )\n            df = cudf.concat(df,ignore_index=True,axis=0)\n            df = df.loc[df['type'].isin([1,2])] # ONLY WANT CARTS AND ORDERS\n            df = df.sort_values(['session','ts'],ascending=[True,False])\n            # USE TAIL OF SESSION\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n<30].drop('n',axis=1)\n            # CREATE PAIRS\n            df = df.merge(df,on='session')\n            df = df.loc[ ((df.ts_x - df.ts_y).abs()< 14 * 24 * 60 * 60) & (df.aid_x != df.aid_y) ] # 14 DAYS\n            # MEMORY MANAGEMENT COMPUTE IN PARTS\n            df = df.loc[(df.aid_x >= PART*SIZE)&(df.aid_x < (PART+1)*SIZE)]\n            # ASSIGN WEIGHTS\n            df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n            df['wgt'] = 1\n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n            # COMBINE INNER CHUNKS\n            if k==a: tmp2 = df\n            else: tmp2 = tmp2.add(df, fill_value=0)\n            print(k,', ',end='')\n        print()\n        # COMBINE OUTER CHUNKS\n        if a==0: tmp = tmp2\n        else: tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    # CONVERT MATRIX TO DICTIONARY\n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n    # SAVE TOP 40\n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    tmp = tmp.loc[tmp.n<15].drop('n',axis=1)\n    # SAVE PART TO DISK (convert to pandas first uses less memory)\n    tmp.to_pandas().to_parquet(f'top_15_buy2buy_v{VER}_{PART}.pqt')","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"papermill":{"duration":113.735315,"end_time":"2022-11-10T16:14:44.498182","exception":false,"start_time":"2022-11-10T16:12:50.762867","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-05-18T10:15:30.658376Z","iopub.execute_input":"2024-05-18T10:15:30.658776Z","iopub.status.idle":"2024-05-18T10:16:01.313569Z","shell.execute_reply.started":"2024-05-18T10:15:30.65874Z","shell.execute_reply":"2024-05-18T10:16:01.312552Z"},"trusted":true},"execution_count":7,"outputs":[{"name":"stdout","text":"\n### DISK PART 1\nProcessing files 0 thru 24 in groups of 5...\n0 , 5 , ","output_type":"stream"},{"name":"stderr","text":"/opt/conda/lib/python3.7/site-packages/cudf/core/frame.py:2600: UserWarning: When using a sequence of booleans for `ascending`, `na_position` flag is not yet supported and defaults to treating nulls as greater than all numbers\n  \"When using a sequence of booleans for `ascending`, \"\n","output_type":"stream"},{"name":"stdout","text":"10 , 15 , 20 , \nProcessing files 25 thru 49 in groups of 5...\n25 , 30 , 35 , 40 , 45 , \nProcessing files 50 thru 74 in groups of 5...\n50 , 55 , 60 , 65 , 70 , \nProcessing files 75 thru 99 in groups of 5...\n75 , 80 , 85 , 90 , 95 , \nProcessing files 100 thru 124 in groups of 5...\n100 , 105 , 110 , 115 , 120 , \nProcessing files 125 thru 145 in groups of 5...\n125 , 130 , 135 , 140 , 145 , \nCPU times: user 21.8 s, sys: 8.49 s, total: 30.3 s\nWall time: 30.6 s\n","output_type":"stream"}]},{"cell_type":"markdown","source":"## 3) \"Clicks\" Co-visitation Matrix - Time Weighted","metadata":{"papermill":{"duration":0.04526,"end_time":"2022-11-10T16:14:44.58589","exception":false,"start_time":"2022-11-10T16:14:44.54063","status":"completed"},"tags":[]}},{"cell_type":"code","source":"%%time\n# USE SMALLEST DISK_PIECES POSSIBLE WITHOUT MEMORY ERROR\nDISK_PIECES = 4\nSIZE = 1.86e6/DISK_PIECES\n\n# COMPUTE IN PARTS FOR MEMORY MANGEMENT\nfor PART in range(DISK_PIECES):\n    print()\n    print('### DISK PART',PART+1)\n    \n    # MERGE IS FASTEST PROCESSING CHUNKS WITHIN CHUNKS\n    # => OUTER CHUNKS\n    for j in range(6):\n        a = j*CHUNK\n        b = min( (j+1)*CHUNK, len(files) )\n        print(f'Processing files {a} thru {b-1} in groups of {READ_CT}...')\n        \n        # => INNER CHUNKS\n        for k in range(a,b,READ_CT):\n            # READ FILE\n            df = [read_file(files[k])]\n            for i in range(1,READ_CT): \n                if k+i<b: df.append( read_file(files[k+i]) )\n            df = cudf.concat(df,ignore_index=True,axis=0)\n            df = df.sort_values(['session','ts'],ascending=[True,False])\n            # USE TAIL OF SESSION\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n<30].drop('n',axis=1)\n            # CREATE PAIRS\n            df = df.merge(df,on='session')\n            df = df.loc[ ((df.ts_x - df.ts_y).abs()< 24 * 60 * 60) & (df.aid_x != df.aid_y) ]\n            # MEMORY MANAGEMENT COMPUTE IN PARTS\n            df = df.loc[(df.aid_x >= PART*SIZE)&(df.aid_x < (PART+1)*SIZE)]\n            # ASSIGN WEIGHTS\n            df = df[['session', 'aid_x', 'aid_y','ts_x']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n            df['wgt'] = 1 + 3*(df.ts_x - 1659304800)/(1662328791-1659304800)\n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n            # COMBINE INNER CHUNKS\n            if k==a: tmp2 = df\n            else: tmp2 = tmp2.add(df, fill_value=0)\n            print(k,', ',end='')\n        print()\n        # COMBINE OUTER CHUNKS\n        if a==0: tmp = tmp2\n        else: tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    # CONVERT MATRIX TO DICTIONARY\n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n    # SAVE TOP 40\n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    tmp = tmp.loc[tmp.n<20].drop('n',axis=1)\n    # SAVE PART TO DISK (convert to pandas first uses less memory)\n    tmp.to_pandas().to_parquet(f'top_20_clicks_v{VER}_{PART}.pqt')","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"papermill":{"duration":null,"end_time":null,"exception":false,"start_time":"2022-11-10T16:14:44.629032","status":"running"},"tags":[],"execution":{"iopub.status.busy":"2024-05-18T10:16:30.297575Z","iopub.execute_input":"2024-05-18T10:16:30.298129Z","iopub.status.idle":"2024-05-18T10:19:43.65278Z","shell.execute_reply.started":"2024-05-18T10:16:30.298081Z","shell.execute_reply":"2024-05-18T10:19:43.651707Z"},"trusted":true},"execution_count":8,"outputs":[{"name":"stdout","text":"\n### DISK PART 1\nProcessing files 0 thru 24 in groups of 5...\n0 , 5 , 10 , 15 , 20 , \nProcessing files 25 thru 49 in groups of 5...\n25 , 30 , 35 , 40 , 45 , \nProcessing files 50 thru 74 in groups of 5...\n50 , 55 , 60 , 65 , 70 , \nProcessing files 75 thru 99 in groups of 5...\n75 , 80 , 85 , 90 , 95 , \nProcessing files 100 thru 124 in groups of 5...\n100 , 105 , 110 , 115 , 120 , \nProcessing files 125 thru 145 in groups of 5...\n125 , 130 , 135 , 140 , 145 , \n\n### DISK PART 2\nProcessing files 0 thru 24 in groups of 5...\n0 , 5 , 10 , 15 , 20 , \nProcessing files 25 thru 49 in groups of 5...\n25 , 30 , 35 , 40 , 45 , \nProcessing files 50 thru 74 in groups of 5...\n50 , 55 , 60 , 65 , 70 , \nProcessing files 75 thru 99 in groups of 5...\n75 , 80 , 85 , 90 , 95 , \nProcessing files 100 thru 124 in groups of 5...\n100 , 105 , 110 , 115 , 120 , \nProcessing files 125 thru 145 in groups of 5...\n125 , 130 , 135 , 140 , 145 , \n\n### DISK PART 3\nProcessing files 0 thru 24 in groups of 5...\n0 , 5 , 10 , 15 , 20 , \nProcessing files 25 thru 49 in groups of 5...\n25 , 30 , 35 , 40 , 45 , \nProcessing files 50 thru 74 in groups of 5...\n50 , 55 , 60 , 65 , 70 , \nProcessing files 75 thru 99 in groups of 5...\n75 , 80 , 85 , 90 , 95 , \nProcessing files 100 thru 124 in groups of 5...\n100 , 105 , 110 , 115 , 120 , \nProcessing files 125 thru 145 in groups of 5...\n125 , 130 , 135 , 140 , 145 , \n\n### DISK PART 4\nProcessing files 0 thru 24 in groups of 5...\n0 , 5 , 10 , 15 , 20 , \nProcessing files 25 thru 49 in groups of 5...\n25 , 30 , 35 , 40 , 45 , \nProcessing files 50 thru 74 in groups of 5...\n50 , 55 , 60 , 65 , 70 , \nProcessing files 75 thru 99 in groups of 5...\n75 , 80 , 85 , 90 , 95 , \nProcessing files 100 thru 124 in groups of 5...\n100 , 105 , 110 , 115 , 120 , \nProcessing files 125 thru 145 in groups of 5...\n125 , 130 , 135 , 140 , 145 , \nCPU times: user 2min 9s, sys: 1min 3s, total: 3min 13s\nWall time: 3min 13s\n","output_type":"stream"}]},{"cell_type":"code","source":"# FREE MEMORY\ndel data_cache, tmp\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2024-05-18T10:19:50.115883Z","iopub.execute_input":"2024-05-18T10:19:50.116784Z","iopub.status.idle":"2024-05-18T10:19:50.30165Z","shell.execute_reply.started":"2024-05-18T10:19:50.116735Z","shell.execute_reply":"2024-05-18T10:19:50.3002Z"},"trusted":true},"execution_count":9,"outputs":[]},{"cell_type":"markdown","source":"# Step 2 - ReRank (choose 20) using handcrafted rules\nFor description of the handcrafted rules, read this notebook's intro.","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[]}},{"cell_type":"code","source":"def load_test():    \n    dfs = []\n    for e, chunk_file in enumerate(glob.glob('../input/otto-chunk-data-inparquet-format/test_parquet/*')):\n        chunk = pd.read_parquet(chunk_file)\n        chunk.ts = (chunk.ts/1000).astype('int32')\n        chunk['type'] = chunk['type'].map(type_labels).astype('int8')\n        dfs.append(chunk)\n    return pd.concat(dfs).reset_index(drop=True) #.astype({\"ts\": \"datetime64[ms]\"})\n\ntest_df = load_test()\nprint('Test data has shape',test_df.shape)\ntest_df.head()","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[],"execution":{"iopub.status.busy":"2024-05-18T10:19:58.878007Z","iopub.execute_input":"2024-05-18T10:19:58.878415Z","iopub.status.idle":"2024-05-18T10:20:00.408543Z","shell.execute_reply.started":"2024-05-18T10:19:58.878378Z","shell.execute_reply":"2024-05-18T10:20:00.407519Z"},"trusted":true},"execution_count":10,"outputs":[{"name":"stdout","text":"Test data has shape (6928123, 4)\n","output_type":"stream"},{"execution_count":10,"output_type":"execute_result","data":{"text/plain":"    session     aid          ts  type\n0  13099779  245308  1661795832     0\n1  13099779  245308  1661795862     1\n2  13099779  972319  1661795888     0\n3  13099779  972319  1661795898     1\n4  13099779  245308  1661795907     0","text/html":"<div>\n<style scoped>\n    .dataframe tbody tr th:only-of-type {\n        vertical-align: middle;\n    }\n\n    .dataframe tbody tr th {\n        vertical-align: top;\n    }\n\n    .dataframe thead th {\n        text-align: right;\n    }\n</style>\n<table border=\"1\" class=\"dataframe\">\n  <thead>\n    <tr style=\"text-align: right;\">\n      <th></th>\n      <th>session</th>\n      <th>aid</th>\n      <th>ts</th>\n      <th>type</th>\n    </tr>\n  </thead>\n  <tbody>\n    <tr>\n      <th>0</th>\n      <td>13099779</td>\n      <td>245308</td>\n      <td>1661795832</td>\n      <td>0</td>\n    </tr>\n    <tr>\n      <th>1</th>\n      <td>13099779</td>\n      <td>245308</td>\n      <td>1661795862</td>\n      <td>1</td>\n    </tr>\n    <tr>\n      <th>2</th>\n      <td>13099779</td>\n      <td>972319</td>\n      <td>1661795888</td>\n      <td>0</td>\n    </tr>\n    <tr>\n      <th>3</th>\n      <td>13099779</td>\n      <td>972319</td>\n      <td>1661795898</td>\n      <td>1</td>\n    </tr>\n    <tr>\n      <th>4</th>\n      <td>13099779</td>\n      <td>245308</td>\n      <td>1661795907</td>\n      <td>0</td>\n    </tr>\n  </tbody>\n</table>\n</div>"},"metadata":{}}]},{"cell_type":"code","source":"%%time\ndef pqt_to_dict(df):\n    return df.groupby('aid_x').aid_y.apply(list).to_dict()\n# LOAD THREE CO-VISITATION MATRICES\ntop_20_clicks = pqt_to_dict( pd.read_parquet(f'top_20_clicks_v{VER}_0.pqt') )\nfor k in range(1,DISK_PIECES): \n    top_20_clicks.update( pqt_to_dict( pd.read_parquet(f'top_20_clicks_v{VER}_{k}.pqt') ) )\ntop_20_buys = pqt_to_dict( pd.read_parquet(f'top_15_carts_orders_v{VER}_0.pqt') )\nfor k in range(1,DISK_PIECES): \n    top_20_buys.update( pqt_to_dict( pd.read_parquet(f'top_15_carts_orders_v{VER}_{k}.pqt') ) )\ntop_20_buy2buy = pqt_to_dict( pd.read_parquet(f'top_15_buy2buy_v{VER}_0.pqt') )\n\n# TOP CLICKS AND ORDERS IN TEST\ntop_clicks = test_df.loc[test_df['type']=='clicks','aid'].value_counts().index.values[:20]\ntop_orders = test_df.loc[test_df['type']=='orders','aid'].value_counts().index.values[:20]\n\nprint('Here are size of our 3 co-visitation matrices:')\nprint( len( top_20_clicks ), len( top_20_buy2buy ), len( top_20_buys ) )","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[],"execution":{"iopub.status.busy":"2024-05-18T10:20:07.592485Z","iopub.execute_input":"2024-05-18T10:20:07.593212Z","iopub.status.idle":"2024-05-18T10:22:18.72431Z","shell.execute_reply.started":"2024-05-18T10:20:07.593165Z","shell.execute_reply":"2024-05-18T10:22:18.723308Z"},"trusted":true},"execution_count":11,"outputs":[{"name":"stdout","text":"Here are size of our 3 co-visitation matrices:\n1837166 1168768 1837166\nCPU times: user 2min 11s, sys: 4.88 s, total: 2min 16s\nWall time: 2min 11s\n","output_type":"stream"}]},{"cell_type":"code","source":"#type_weight_multipliers = {'clicks': 1, 'carts': 6, 'orders': 3}\ntype_weight_multipliers = {0: 1, 1: 6, 2: 3}\n\ndef suggest_clicks(df):\n    # USER HISTORY AIDS AND TYPES\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    # RERANK CANDIDATES USING WEIGHTS\n    if len(unique_aids)>=20:\n        weights=np.logspace(0.1,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        # RERANK BASED ON REPEAT ITEMS AND TYPE OF ITEMS\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    # USE \"CLICKS\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(20) if aid2 not in unique_aids]    \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    # USE TOP20 TEST CLICKS\n    return result + list(top_clicks)[:20-len(result)]\n\ndef suggest_buys(df):\n    # USER HISTORY AIDS AND TYPES\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    # UNIQUE AIDS AND UNIQUE BUYS\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    df = df.loc[(df['type']==1)|(df['type']==2)]\n    unique_buys = list(dict.fromkeys( df.aid.tolist()[::-1] ))\n    # RERANK CANDIDATES USING WEIGHTS\n    if len(unique_aids)>=20:\n        weights=np.logspace(0.5,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        # RERANK BASED ON REPEAT ITEMS AND TYPE OF ITEMS\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        # RERANK CANDIDATES USING \"BUY2BUY\" CO-VISITATION MATRIX\n        aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n        for aid in aids3: aids_temp[aid] += 0.1\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    # USE \"CART ORDER\" CO-VISITATION MATRIX\n    aids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_aids if aid in top_20_buys]))\n    # USE \"BUY2BUY\" CO-VISITATION MATRIX\n    aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2+aids3).most_common(20) if aid2 not in unique_aids] \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    # USE TOP20 TEST ORDERS\n    return result + list(top_orders)[:20-len(result)]","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[],"execution":{"iopub.status.busy":"2024-05-18T10:22:29.977437Z","iopub.execute_input":"2024-05-18T10:22:29.978204Z","iopub.status.idle":"2024-05-18T10:22:30.028859Z","shell.execute_reply.started":"2024-05-18T10:22:29.978166Z","shell.execute_reply":"2024-05-18T10:22:30.027784Z"},"trusted":true},"execution_count":12,"outputs":[]},{"cell_type":"markdown","source":"# Create Submission CSV\nInferring test data with Pandas groupby is slow. We need to accelerate the following code.","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[]}},{"cell_type":"code","source":"%%time\npred_df_clicks = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(\n    lambda x: suggest_clicks(x)\n)\n\npred_df_buys = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(\n    lambda x: suggest_buys(x)\n)","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[],"execution":{"iopub.status.busy":"2024-05-18T10:29:50.404782Z","iopub.execute_input":"2024-05-18T10:29:50.405896Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clicks_pred_df = pd.DataFrame(pred_df_clicks.add_suffix(\"_clicks\"), columns=[\"labels\"]).reset_index()\norders_pred_df = pd.DataFrame(pred_df_buys.add_suffix(\"_orders\"), columns=[\"labels\"]).reset_index()\ncarts_pred_df = pd.DataFrame(pred_df_buys.add_suffix(\"_carts\"), columns=[\"labels\"]).reset_index()","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_df = pd.concat([clicks_pred_df, orders_pred_df, carts_pred_df])\npred_df.columns = [\"session_type\", \"labels\"]\npred_df[\"labels\"] = pred_df.labels.apply(lambda x: \" \".join(map(str,x)))\npred_df.to_csv(\"submission.csv\", index=False)\npred_df.head()","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[],"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Calculer le score de validation","metadata":{}},{"cell_type":"code","source":"# FREE MEMORY\ndel pred_df_clicks, pred_df_buys, clicks_pred_df, orders_pred_df, carts_pred_df\ndel top_20_clicks, top_20_buy2buy, top_20_buys, top_clicks, top_orders, test_df\n_ = gc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# COMPUTE METRIC\nscore = 0\nweights = {'clicks': 0.10, 'carts': 0.30, 'orders': 0.60}\nfor t in ['clicks','carts','orders']:\n    sub = pred_df.loc[pred_df.session_type.str.contains(t)].copy()\n    sub['session'] = sub.session_type.apply(lambda x: int(x.split('_')[0]))\n    sub.labels = sub.labels.apply(lambda x: [int(i) for i in x.split(' ')[:20]])\n    test_labels = pd.read_parquet('../input/otto-validation/test_labels.parquet')\n    test_labels = test_labels.loc[test_labels['type']==t]\n    test_labels = test_labels.merge(sub, how='left', on=['session'])\n    test_labels['hits'] = test_labels.apply(lambda df: len(set(df.ground_truth).intersection(set(df.labels))), axis=1)\n    test_labels['gt_count'] = test_labels.ground_truth.str.len().clip(0,20)\n    recall = test_labels['hits'].sum() / test_labels['gt_count'].sum()\n    score += weights[t]*recall\n    print(f'{t} recall =',recall)\n    \nprint('=============')\nprint('Overall Recall =',score)\nprint('=============')","metadata":{},"execution_count":null,"outputs":[]}]}