{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        #print(os.path.join(dirname, filename))\n        continue","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-12-22T16:51:38.791078Z","iopub.execute_input":"2022-12-22T16:51:38.791846Z","iopub.status.idle":"2022-12-22T16:51:38.803319Z","shell.execute_reply.started":"2022-12-22T16:51:38.791806Z","shell.execute_reply":"2022-12-22T16:51:38.802282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Credits\n\nWe thank many Kagglers who have shared ideas. \n* The notebook is from [CHRIS DEOTTE](https://www.kaggle.com/code/cdeotte/compute-validation-score-cv-565/notebook). \n* We use co-visitation matrix idea from Vladimir here. \n* We use groupby sort logic from Sinan in comment section here. \n* We use duplicate prediction removal logic from Radek here. \n* We use multiple visit logic from Pietro here. \n* We use type weighting logic from Ingvaras here. \n* We use leaky test data from my previous notebook here.\n\nAnd some ideas may have originated from Tawara here and KJ here. We use Colum2131's parquets here. Above image is from Ravi's discussion about candidate rerank models here","metadata":{}},{"cell_type":"markdown","source":"### Step 1 - Generate Candidates\n\nFor each test user, we generate possible choices, i.e. candidates. In this notebook, we generate candidates from 5 sources:\n\n* User history of clicks, carts, orders\n* Most popular 20 clicks, carts, orders during test week\n* Co-visitation matrix of click/cart/order to cart/order with type weighting\n* Co-visitation matrix of cart/order to cart/order called buy2buy\n* Co-visitation matrix of click/cart/order to clicks with time weighting\n\n### Step 2 - ReRank and Choose 20\n\nGiven the list of candidates, we must select 20 to be our predictions. In this notebook, we do this with a set of handcrafted rules. We can improve our predictions by training an XGBoost model to select for us. Our handcrafted rules give priority to:\n\n* Most recent previously visited items\n* Items previously visited multiple times\n* Items previously in cart or order\n* Co-visitation matrix of cart/order to cart/order\n* Current popular items","metadata":{}},{"cell_type":"markdown","source":"# 1. Generate Candidates (Recall)","metadata":{}},{"cell_type":"code","source":"VER = 1\nfrom tqdm.notebook import tqdm\nimport os, sys, pickle, glob, gc\nfrom collections import Counter\nimport cudf, itertools\nprint('We will use RAPIDS version',cudf.__version__)","metadata":{"execution":{"iopub.status.busy":"2022-12-22T16:51:43.979752Z","iopub.execute_input":"2022-12-22T16:51:43.980142Z","iopub.status.idle":"2022-12-22T16:51:46.576695Z","shell.execute_reply.started":"2022-12-22T16:51:43.980108Z","shell.execute_reply":"2022-12-22T16:51:46.575757Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Compute Three Co-visitation Matrices with RAPIDS\n\nWe will compute 3 co-visitation matrices using RAPIDS cuDF on GPU. This is 30x faster than using Pandas CPU like other public notebooks! For maximum speed, set the variable DISK_PIECES to the smallest number possible based on the GPU you are using without incurring memory errors. If you run this code offline with 32GB GPU ram, then you can use DISK_PIECES = 1 and compute each co-visitation matrix in almost 1 minute! Kaggle's GPU only has 16GB ram, so we use DISK_PIECES = 4 and it takes an amazing 3 minutes each! Below are some of the tricks to speed up computation\n\n* Use RAPIDS cuDF GPU instead of Pandas CPU\n* Read disk once and save in CPU RAM for later GPU multiple use\n* Process largest amount of data possible on GPU at one time\n* Merge data in two stages. Multiple small to single medium. Multiple medium to single large.\n* Write result as parquet instead of dictionary","metadata":{}},{"cell_type":"code","source":"%%time\n# CACHE FUNCTIONS\ndef read_file(f):\n    return cudf.DataFrame( data_cache[f] )\ndef read_file_to_cache(f):\n    df = pd.read_parquet(f)\n    df.ts = (df.ts/1000).astype('int32')\n    df['type'] = df['type'].map(type_labels).astype('int8')\n    return df\n\n# CACHE THE DATA ON CPU BEFORE PROCESSING ON GPU\ndata_cache = {} # 120个key:value，key是地址str，value是从function read_file_to_cache返回的df\ntype_labels = {'clicks':0, 'carts':1, 'orders':2}\nfiles = glob.glob('../input/otto-validation/*_parquet/*') # 变量files只是一个list，含有地址的str，而且顺序打乱了。len(files) = 120, train=100, test=20\nfor f in files: data_cache[f] = read_file_to_cache(f)\n\n# CHUNK PARAMETERS\nREAD_CT = 5\nCHUNK = int( np.ceil( len(files)/6 )) # 20\nprint(f'We will process {len(files)} files, in groups of {READ_CT} and chunks of {CHUNK}.')","metadata":{"execution":{"iopub.status.busy":"2022-12-22T16:51:54.634410Z","iopub.execute_input":"2022-12-22T16:51:54.634763Z","iopub.status.idle":"2022-12-22T16:52:35.124138Z","shell.execute_reply.started":"2022-12-22T16:51:54.634732Z","shell.execute_reply":"2022-12-22T16:52:35.123087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### \"Carts Orders\" Co-visitation Matrix的理念：\n\n1. 每个用户互动最近的30个item\n2. 两两merge, 去除24小时以外的物品，并去重，生成item1，item2的pair\n3. 按照type给item1赋权重，添加到cart的商品得分最高，order其次\n4. 去掉session属性，合并所有的item pair，形成unique的pair\n4. groupby item1，数出来权重分最高的15个item2，从而获得item1 - 15个item2\n\n### \"Buy_2_Buy\" Co-visitation Matrix的理念：\n\n1. 只保留type为1和2的event\n2. 每个用户互动最近的30个item\n2. 两两merge, 去除24小时以外的物品，并去重，生成item1，item2的pair\n3. 权重就是出现次数，都是1\n4. 去掉session属性，合并所有的item pair，形成unique的pair\n4. groupby item1，数出来权重分最高的15个item2，从而获得item1 - 15个item2\n\n### click_2_click的理念：\n\n1. 每个用户互动最近的30个item\n2. 两两merge, 去除24小时以外的物品，并去重，生成item1，item2的pair\n3. 按照时间顺序给item1赋权重，越近期的商品得分越高\n4. 去掉session属性，合并所有的item pair，形成unique的pair\n4. groupby item1，数出来权重分最高的20个item2，从而获得item1 - 20个item2\n","metadata":{}},{"cell_type":"markdown","source":"### 1) \"Carts Orders\" Co-visitation Matrix - Type Weighted","metadata":{}},{"cell_type":"markdown","source":"#### For one chunk，test the code","metadata":{}},{"cell_type":"code","source":"df = [read_file(files[0])]\nfor i in range(1,5): \n    df.append( read_file(files[i]) )\ndf = cudf.concat(df,ignore_index=True,axis=0)\n# df = df.sort_values(['session','ts'],ascending=[True,False])\n#  # USE TAIL OF SESSION\n# df = df.reset_index(drop=True)\n# df['n'] = df.groupby('session').cumcount()\n# df = df.loc[df.n<30].drop('n',axis=1)\n# # CREATE PAIRS\n# df = df.merge(df,on='session')\n# df = df.loc[ ((df.ts_x - df.ts_y).abs()< 24 * 60 * 60) & (df.aid_x != df.aid_y) ]\n# # ASSIGN WEIGHTS\n# type_weight = {0:1, 1:6, 2:3}\n# df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n# df['wgt'] = df.type_y.map(type_weight)\n# df = df[['aid_x','aid_y','wgt']]\n# df.wgt = df.wgt.astype('float32')\n# df = df.groupby(['aid_x','aid_y']).wgt.sum()\n# df = df.reset_index()\n# df = df.sort_values(['aid_x','wgt'],ascending=[True,False])\n# df = df.reset_index(drop=True)\n# df['n'] = df.groupby('aid_x').aid_y.cumcount()\ndf","metadata":{"execution":{"iopub.status.busy":"2022-12-22T16:52:51.323371Z","iopub.execute_input":"2022-12-22T16:52:51.323720Z","iopub.status.idle":"2022-12-22T16:52:53.211304Z","shell.execute_reply.started":"2022-12-22T16:52:51.323691Z","shell.execute_reply":"2022-12-22T16:52:53.210266Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### For the whole dataset","metadata":{}},{"cell_type":"code","source":"%%time\ntype_weight = {0:1, 1:6, 2:3}\n\n# USE SMALLEST DISK_PIECES POSSIBLE WITHOUT MEMORY ERROR\nDISK_PIECES = 4\nSIZE = 1.86e6/DISK_PIECES\n\n# COMPUTE IN PARTS FOR MEMORY MANGEMENT\nfor PART in range(DISK_PIECES):\n    print()\n    print('### DISK PART',PART+1)\n    \n    # MERGE IS FASTEST PROCESSING CHUNKS WITHIN CHUNKS\n    # => OUTER CHUNKS\n    for j in range(6):\n        a = j*CHUNK\n        b = min( (j+1)*CHUNK, len(files) )\n        print(f'Processing files {a} thru {b-1} in groups of {READ_CT}...')\n        \n        # => INNER CHUNKS\n        for k in range(a,b,READ_CT):\n            # READ FILE\n            df = [read_file(files[k])]\n            for i in range(1,READ_CT): \n                if k+i<b: df.append( read_file(files[k+i]) )\n            df = cudf.concat(df,ignore_index=True,axis=0)\n            df = df.sort_values(['session','ts'],ascending=[True,False])\n            # USE TAIL OF SESSION\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n<30].drop('n',axis=1)\n            # CREATE PAIRS\n            df = df.merge(df,on='session')\n            df = df.loc[ ((df.ts_x - df.ts_y).abs()< 24 * 60 * 60) & (df.aid_x != df.aid_y) ]\n            # MEMORY MANAGEMENT COMPUTE IN PARTS\n            df = df.loc[(df.aid_x >= PART*SIZE)&(df.aid_x < (PART+1)*SIZE)]\n            # ASSIGN WEIGHTS\n            df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n            df['wgt'] = df.type_y.map(type_weight)\n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n            # COMBINE INNER CHUNKS\n            if k==a: tmp2 = df\n            else: tmp2 = tmp2.add(df, fill_value=0)\n            print(k,', ',end='')\n        print()\n        # COMBINE OUTER CHUNKS\n        if a==0: tmp = tmp2\n        else: tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    # CONVERT MATRIX TO DICTIONARY\n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n    # SAVE TOP 40\n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    tmp = tmp.loc[tmp.n<15].drop('n',axis=1)\n    # SAVE PART TO DISK (convert to pandas first uses less memory)\n    tmp.to_pandas().to_parquet(f'top_15_carts_orders_v{VER}_{PART}.pqt')","metadata":{"execution":{"iopub.status.busy":"2022-12-22T16:52:58.254950Z","iopub.execute_input":"2022-12-22T16:52:58.255660Z","iopub.status.idle":"2022-12-22T16:55:07.482160Z","shell.execute_reply.started":"2022-12-22T16:52:58.255623Z","shell.execute_reply":"2022-12-22T16:55:07.481086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2) \"Buy2Buy\" Co-visitation Matrix","metadata":{}},{"cell_type":"markdown","source":"#### For one chunk, test the code","metadata":{}},{"cell_type":"code","source":"df = [read_file(files[0])]\nfor i in range(1,5): \n    df.append( read_file(files[i]) )\ndf = cudf.concat(df,ignore_index=True,axis=0)\ndf = df.loc[df['type'].isin([1,2])] # ONLY WANT CARTS AND ORDERS\ndf = df.sort_values(['session','ts'],ascending=[True,False])\n# USE TAIL OF SESSION\ndf = df.reset_index(drop=True)\ndf['n'] = df.groupby('session').cumcount()\ndf = df.loc[df.n<30].drop('n',axis=1)\n# CREATE PAIRS\ndf = df.merge(df,on='session')\ndf = df.loc[ ((df.ts_x - df.ts_y).abs()< 14 * 24 * 60 * 60) & (df.aid_x != df.aid_y) ] # 14 DAYS\n# ASSIGN WEIGHTS\ndf = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\ndf['wgt'] = 1\ndf = df[['aid_x','aid_y','wgt']]\ndf.wgt = df.wgt.astype('float32')\ndf = df.groupby(['aid_x','aid_y']).wgt.sum()\ndf = df.reset_index()\ndf = df.sort_values(['aid_x','wgt'],ascending=[True,False])\ndf = df.reset_index(drop=True)\ndf['n'] = df.groupby('aid_x').aid_y.cumcount()\ndf","metadata":{"execution":{"iopub.status.busy":"2022-12-16T16:41:27.276008Z","iopub.execute_input":"2022-12-16T16:41:27.277052Z","iopub.status.idle":"2022-12-16T16:41:27.467259Z","shell.execute_reply.started":"2022-12-16T16:41:27.277011Z","shell.execute_reply":"2022-12-16T16:41:27.466187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# USE SMALLEST DISK_PIECES POSSIBLE WITHOUT MEMORY ERROR\nDISK_PIECES = 1\nSIZE = 1.86e6/DISK_PIECES\n\n# COMPUTE IN PARTS FOR MEMORY MANGEMENT\nfor PART in range(DISK_PIECES):\n    print()\n    print('### DISK PART',PART+1)\n    \n    # MERGE IS FASTEST PROCESSING CHUNKS WITHIN CHUNKS\n    # => OUTER CHUNKS\n    for j in range(6):\n        a = j*CHUNK\n        b = min( (j+1)*CHUNK, len(files) )\n        print(f'Processing files {a} thru {b-1} in groups of {READ_CT}...')\n        \n        # => INNER CHUNKS\n        for k in range(a,b,READ_CT):\n            # READ FILE\n            df = [read_file(files[k])]\n            for i in range(1,READ_CT): \n                if k+i<b: df.append( read_file(files[k+i]) )\n            df = cudf.concat(df,ignore_index=True,axis=0)\n            df = df.loc[df['type'].isin([1,2])] # ONLY WANT CARTS AND ORDERS\n            df = df.sort_values(['session','ts'],ascending=[True,False])\n            # USE TAIL OF SESSION\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n<30].drop('n',axis=1)\n            # CREATE PAIRS\n            df = df.merge(df,on='session')\n            df = df.loc[ ((df.ts_x - df.ts_y).abs()< 14 * 24 * 60 * 60) & (df.aid_x != df.aid_y) ] # 14 DAYS\n            # MEMORY MANAGEMENT COMPUTE IN PARTS\n            df = df.loc[(df.aid_x >= PART*SIZE)&(df.aid_x < (PART+1)*SIZE)]\n            # ASSIGN WEIGHTS\n            df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n            df['wgt'] = 1\n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n            # COMBINE INNER CHUNKS\n            if k==a: tmp2 = df\n            else: tmp2 = tmp2.add(df, fill_value=0)\n            print(k,', ',end='')\n        print()\n        # COMBINE OUTER CHUNKS\n        if a==0: tmp = tmp2\n        else: tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    # CONVERT MATRIX TO DICTIONARY\n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n    # SAVE TOP 15\n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    tmp = tmp.loc[tmp.n<15].drop('n',axis=1)\n    # SAVE PART TO DISK (convert to pandas first uses less memory)\n    tmp.to_pandas().to_parquet(f'top_15_buy2buy_v{VER}_{PART}.pqt')","metadata":{"execution":{"iopub.status.busy":"2022-12-22T16:56:37.280289Z","iopub.execute_input":"2022-12-22T16:56:37.281295Z","iopub.status.idle":"2022-12-22T16:56:56.587449Z","shell.execute_reply.started":"2022-12-22T16:56:37.281244Z","shell.execute_reply":"2022-12-22T16:56:56.586156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 3) \"Clicks\" Co-visitation Matrix - Time Weighted","metadata":{}},{"cell_type":"code","source":"# => OUTER CHUNKS\nfor j in range(6):\n    a = j*CHUNK # a = 0, 20, 40, 60, 80, 100\n    b = min( (j+1)*CHUNK, len(files) ) # b = 20, 40, 60, 80, 100, 120\n    print(f'Processing files {a} thru {b-1} in groups of {READ_CT}...')\n\n    # => INNER CHUNKS\n    for k in range(a,b,READ_CT): # k = 0, 5, 10, 15\n        # READ FILE\n        #df = [read_file(files[k])]\n        for i in range(1,READ_CT): \n            #if k+i<b: df.append( read_file(files[k+i]) )\n            print(k, i, k + i, b)","metadata":{"execution":{"iopub.status.busy":"2022-12-16T05:34:47.683058Z","iopub.execute_input":"2022-12-16T05:34:47.683494Z","iopub.status.idle":"2022-12-16T05:34:47.699155Z","shell.execute_reply.started":"2022-12-16T05:34:47.683450Z","shell.execute_reply":"2022-12-16T05:34:47.697626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = [read_file(files[0])]\nfor i in range(1,5): \n    df.append( read_file(files[i]) )\ndf = cudf.concat(df,ignore_index=True,axis=0)\ndf = df.sort_values(['session','ts'],ascending=[True,False])\n# USE TAIL OF SESSION\ndf = df.reset_index(drop=True)\ndf['n'] = df.groupby('session').cumcount()\ndf = df.loc[df.n<30].drop('n',axis=1)\ndf = df.merge(df,on='session')\ndf = df.loc[ ((df.ts_x - df.ts_y).abs()< 24 * 60 * 60) & (df.aid_x != df.aid_y) ]\ndf = df[['session', 'aid_x', 'aid_y','ts_x']].drop_duplicates(['session', 'aid_x', 'aid_y'])\ndf['wgt'] = 1 + 3*(df.ts_x - 1659304800)/(1662328791-1659304800)\ndf = df[['aid_x','aid_y','wgt']]\ndf.wgt = df.wgt.astype('float32')\ndf = df.groupby(['aid_x','aid_y']).wgt.sum()\ndf = df.reset_index()\ndf = df.sort_values(['aid_x','wgt'],ascending=[True,False])\ndf['n'] = df.groupby('aid_x').aid_y.cumcount()\n#df[(df.aid_y == 1295096) | (df.aid_x == 1855594)]\ndf = df.loc[df.n<20].drop('n',axis=1)\n#df[df.aid_y == 881881]\ndf","metadata":{"execution":{"iopub.status.busy":"2022-12-16T05:50:33.721043Z","iopub.execute_input":"2022-12-16T05:50:33.722128Z","iopub.status.idle":"2022-12-16T05:50:34.323147Z","shell.execute_reply.started":"2022-12-16T05:50:33.722083Z","shell.execute_reply":"2022-12-16T05:50:34.322166Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DISK_PIECES = 4\nSIZE = 1.86e6/DISK_PIECES\nSIZE","metadata":{"execution":{"iopub.status.busy":"2022-12-16T03:07:26.625792Z","iopub.execute_input":"2022-12-16T03:07:26.626424Z","iopub.status.idle":"2022-12-16T03:07:26.635067Z","shell.execute_reply.started":"2022-12-16T03:07:26.626374Z","shell.execute_reply":"2022-12-16T03:07:26.633884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# USE SMALLEST DISK_PIECES POSSIBLE WITHOUT MEMORY ERROR\nDISK_PIECES = 4\nSIZE = 1.86e6/DISK_PIECES\n\n# COMPUTE IN PARTS FOR MEMORY MANGEMENT\nfor PART in range(DISK_PIECES):\n    print()\n    print('### DISK PART',PART+1)\n    \n    # MERGE IS FASTEST PROCESSING CHUNKS WITHIN CHUNKS\n    # => OUTER CHUNKS\n    for j in range(6):\n        a = j*CHUNK # a = 0, 20, 40, 60, 80, 100\n        b = min( (j+1)*CHUNK, len(files) ) # b = 20, 40, 60, 80, 100, 120\n        print(f'Processing files {a} thru {b-1} in groups of {READ_CT}...')\n        \n        # => INNER CHUNKS\n        for k in range(a,b,READ_CT): # k = 0, 5, 10, 15\n            # READ FILE 这一条和下面的小型for loop组合，把5个file的数据组合起来。\n            df = [read_file(files[k])]\n            for i in range(1,READ_CT): \n                if k+i<b: df.append( read_file(files[k+i]) )\n            df = cudf.concat(df,ignore_index=True,axis=0)\n            # 对session升序，时间降序\n            df = df.sort_values(['session','ts'],ascending=[True,False])\n            # USE TAIL OF SESSION\n            df = df.reset_index(drop=True)\n            # 统计用户交互行为的累计值。cumcount不是做sum，而是数出现的个数，然后填到每个出现值得后面，如一个session出现4次，则cumcount就会在session出现的行后面填0，1，2，3\n            df['n'] = df.groupby('session').cumcount()\n            # 只保留每个用户最近30次互动\n            df = df.loc[df.n<30].drop('n',axis=1)\n            # CREATE PAIRS\n            # 单一session得到的行数是物品数的平方，如有4个item，得到16行，有8个item，得到64行。相当于每个item都和所有item join一遍\n            df = df.merge(df,on='session')\n            # 这一步就是所谓的co-visitation，在self join之后，去除一天之外的数据以及item self join的数据\n            # 但仍然无法避免重复，比如item1有两次点击0和一次购物车1，它和item2就会有三次join，还不包括item2作为右侧的列和item1也要交三遍\n            df = df.loc[ ((df.ts_x - df.ts_y).abs()< 24 * 60 * 60) & (df.aid_x != df.aid_y) ]\n            # MEMORY MANAGEMENT COMPUTE IN PARTS\n            # 不是按照session或行来分块，而是按照aid来分\n            df = df.loc[(df.aid_x >= PART*SIZE)&(df.aid_x < (PART+1)*SIZE)]\n            # ASSIGN WEIGHTS\n            # 这一步是对于merge之后，去除多余的重复值。经此之后，不会出现x和y都是重叠的问题。但还是会出现x=item1，y=item2和x=item2,y=item1的情况\n            df = df[['session', 'aid_x', 'aid_y','ts_x']].drop_duplicates(['session', 'aid_x', 'aid_y']) \n            # 1659304800 是所有数据中最久的一条，即train set中的第一条\n            # 1662328791 是otto数据(非validation)中test集的最后一条\n            # 权限完全和时间顺序有关，越近则分数最多\n            df['wgt'] = 1 + 3*(df.ts_x - 1659304800)/(1662328791-1659304800)\n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            # 自定义算法的又一个规则：把item pair的权重求和\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n            # COMBINE INNER CHUNKS\n            if k==a: tmp2 = df\n            else: tmp2 = tmp2.add(df, fill_value=0)\n            print(k,', ',end='')\n        print()\n        # COMBINE OUTER CHUNKS\n        if a==0: tmp = tmp2\n        else: tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    # CONVERT MATRIX TO DICTIONARY\n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n    \"\"\"\n    | idx  |aid_x| aid_y |  wgt\n    1310622|  2  |881881 |2.954518\n    4969642|  2  |258090 |2.954518\n    5422461|  2  |929227 |2.954518\n    1999530|  3  |1180285|42.731621\n    ...\n    6505207|1855594|643674|3.011645\n    6733875|1855594|849737|3.011645\n    \"\"\"\n    # SAVE TOP 40\n    tmp = tmp.reset_index(drop=True)\n    # 本意是针对每个aid_x，数出来多少个y和它对应，然后按照时间由近到远排序，留下前20个aid_y。但结果不对，如下\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    \"\"\"\n    \taid_x\taid_y\twgt\tn\n1310622\t2\t881881\t2.954518\t697\n2385962\t258090\t881881\t2.954673\t67\n5290558\t929227\t881881\t2.954766\t35\n2520188\t1101201\t881881\t2.954891\t54\n如果要验证的话，可以在此打断点\n    \"\"\"\n    tmp = tmp.loc[tmp.n<20].drop('n',axis=1)\n    # SAVE PART TO DISK (convert to pandas first uses less memory)\n    tmp.to_pandas().to_parquet(f'top_20_clicks_v{VER}_{PART}.pqt')","metadata":{"execution":{"iopub.status.busy":"2022-12-22T16:58:43.809371Z","iopub.execute_input":"2022-12-22T16:58:43.809740Z","iopub.status.idle":"2022-12-22T17:00:51.787093Z","shell.execute_reply.started":"2022-12-22T16:58:43.809698Z","shell.execute_reply":"2022-12-22T17:00:51.785928Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# FREE MEMORY\ndel data_cache, tmp\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-12-22T17:02:23.691865Z","iopub.execute_input":"2022-12-22T17:02:23.692471Z","iopub.status.idle":"2022-12-22T17:02:23.847864Z","shell.execute_reply.started":"2022-12-22T17:02:23.692431Z","shell.execute_reply":"2022-12-22T17:02:23.845948Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 2 - ReRank (choose 20) using handcrafted rules\n\nFor description of the handcrafted rules, read this notebook's intro","metadata":{}},{"cell_type":"markdown","source":"top_20_clicks（1,812,132）: dict结构，key是sessionID或aid，value是推荐的20个aid\n\ntop_20_buy2buy(1,055,146): 来源是`top_15_buy2buy`。\n\ntop_20_buys (1,812,132): 来源是`top_15_carts_orders`.\n\n### suggest_clicks 功能\n以每个session为基准，先把这个session访问过的aid提取出，按照时间倒序（时间上越近越排在list前面）生成unique的item id列表\n* 如果unique aid大于20个：\n        给所有item赋权，包括时间上和types，时间越近权重越多，交互次数越多权重越多。\n        提取出权重最大的20个item id\n* 如果unique aid不到20个：\n        从top_20_clicks中找到unique item对应的top 20个物品，按照出现次数排序\n        截取这些物品，和unique item补全到20个商品\n        如果还凑不满20个，用全体商品中的top 20个来补全\n        \n返回的结果是每个session得到一个20个item list\n\n### suggest_buys 功能\n\n以每个session为基准，先把这个session访问过的aid提取出，按照时间倒序（时间上越近越排在list前面）生成unique的item id列表\n\n同时准备一个unique buys列表，如果有session有过将某item放入cart或order的记录，则这些item就是unique buys，否则为空\n* 如果unique aid大于20个：\n        给所有item赋权，包括时间上和types，时间越近权重越多，交互次数越多权重越多。\n        再从top_20_buy2buy中取出unique buys对应的20个商品。如果aids_temp中有物品在这些商品中，则+0.1分\n        提取出权重最大的20个item id，如果session没有cart或order的物品，则返回其unique aid。\n* 如果unique aid不到20个：\n        从top_20_buys中找到unique item对应的top 20个物品，按照出现次数排序\n        从top_20_buy2buy中找到unique buy对应的top 20个物品，按照出现次数排序\n        截取这两大类物品，和unique item补全到20个商品\n        如果还凑不满20个，用全体order中的top 20个来补全\n\n返回的结果是每个session得到一个20个item list","metadata":{}},{"cell_type":"code","source":"def load_test():    \n    dfs = []\n    for e, chunk_file in enumerate(glob.glob('../input/otto-validation/test_parquet/*')):\n        chunk = pd.read_parquet(chunk_file)\n        chunk.ts = (chunk.ts/1000).astype('int32')\n        chunk['type'] = chunk['type'].map(type_labels).astype('int8')\n        dfs.append(chunk)\n    return pd.concat(dfs).reset_index(drop=True) #.astype({\"ts\": \"datetime64[ms]\"})\n\ntest_df = load_test()\nprint('Test data has shape',test_df.shape)\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-22T17:02:27.306923Z","iopub.execute_input":"2022-12-22T17:02:27.307602Z","iopub.status.idle":"2022-12-22T17:02:28.602513Z","shell.execute_reply.started":"2022-12-22T17:02:27.307566Z","shell.execute_reply":"2022-12-22T17:02:28.601533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"top_clicks = test_df.loc[test_df['type']==0,'aid'].value_counts().index.values[:20]\ntop_clicks","metadata":{"execution":{"iopub.status.busy":"2022-12-22T19:53:18.143265Z","iopub.execute_input":"2022-12-22T19:53:18.143840Z","iopub.status.idle":"2022-12-22T19:53:18.442005Z","shell.execute_reply.started":"2022-12-22T19:53:18.143805Z","shell.execute_reply":"2022-12-22T19:53:18.440967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# LOAD THREE CO-VISITATION MATRICES\ndef pqt_to_dict(df):\n    return df.groupby('aid_x').aid_y.apply(list).to_dict()\n\ntop_20_clicks = pqt_to_dict( pd.read_parquet(f'top_20_clicks_v{VER}_0.pqt') )\nfor k in range(1,DISK_PIECES): \n    top_20_clicks.update( pqt_to_dict( pd.read_parquet(f'top_20_clicks_v{VER}_{k}.pqt') ) )\ntop_20_buys = pqt_to_dict( pd.read_parquet(f'top_15_carts_orders_v{VER}_0.pqt') )\nfor k in range(1,DISK_PIECES): \n    top_20_buys.update( pqt_to_dict( pd.read_parquet(f'top_15_carts_orders_v{VER}_{k}.pqt') ) )\ntop_20_buy2buy = pqt_to_dict( pd.read_parquet(f'top_15_buy2buy_v{VER}_0.pqt') )\n\n# TOP CLICKS AND ORDERS IN TEST, should be mapped clicks to 0, and orders to 1\ntop_clicks = test_df.loc[test_df['type']=='clicks','aid'].value_counts().index.values[:20]\ntop_orders = test_df.loc[test_df['type']=='orders','aid'].value_counts().index.values[:20]\n\nprint('Here are size of our 3 co-visitation matrices:')\nprint( len( top_20_clicks ), len( top_20_buy2buy ), len( top_20_buys ) )","metadata":{"execution":{"iopub.status.busy":"2022-12-22T17:02:36.308465Z","iopub.execute_input":"2022-12-22T17:02:36.308815Z","iopub.status.idle":"2022-12-22T17:04:10.269762Z","shell.execute_reply.started":"2022-12-22T17:02:36.308784Z","shell.execute_reply":"2022-12-22T17:04:10.268600Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df_sample = test_df.iloc[:30]\ntest_df_sample","metadata":{"execution":{"iopub.status.busy":"2022-12-22T20:09:49.261394Z","iopub.execute_input":"2022-12-22T20:09:49.261747Z","iopub.status.idle":"2022-12-22T20:09:49.277153Z","shell.execute_reply.started":"2022-12-22T20:09:49.261715Z","shell.execute_reply":"2022-12-22T20:09:49.276065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"top_20_clicks[619488]","metadata":{"execution":{"iopub.status.busy":"2022-12-22T19:27:43.891008Z","iopub.execute_input":"2022-12-22T19:27:43.891372Z","iopub.status.idle":"2022-12-22T19:27:43.900643Z","shell.execute_reply.started":"2022-12-22T19:27:43.891342Z","shell.execute_reply":"2022-12-22T19:27:43.899668Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"top_clicks","metadata":{"execution":{"iopub.status.busy":"2022-12-22T19:53:36.542093Z","iopub.execute_input":"2022-12-22T19:53:36.542973Z","iopub.status.idle":"2022-12-22T19:53:36.549468Z","shell.execute_reply.started":"2022-12-22T19:53:36.542923Z","shell.execute_reply":"2022-12-22T19:53:36.548498Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#type_weight_multipliers = {'clicks': 1, 'carts': 6, 'orders': 3}\ntype_weight_multipliers = {0: 1, 1: 6, 2: 3}\n\ndef suggest_clicks(df):\n    \"\"\"\n    以每个session为基准，先把这个session访问过的aid提取出，按照时间倒序（时间上越近越排在list前面）生成unique的item id列表\n    如果unique aid大于20个：\n        给所有item赋权，包括时间上和types，时间越近权重越多，交互次数越多权重越多。\n        提取出权重最大的20个item id\n    如果unique aid不到20个：\n        从top_20_clicks中找到unique item对应的top 20个物品，按照出现次数排序\n        截取这些物品，和unique item补全到20个商品\n        如果还凑不满20个，用全体商品中的top 20个来补全\n    返回的结果是每个session得到一个20个item list\n    \"\"\"\n    # USE USER HISTORY AIDS AND TYPES\n    # 把每一个session的aid和对应的type生成list\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    # aids[::-1]每一个session的aid倒序，倒序的作用是把时间最近的访问物品放到最前面\n    # dict.fromkeys(aids[::-1])以aids中的每一件item倒序排列，并以之为key生成dict：{242544: None, 619488: None, 579241: None, 700554: None}\n    # 上面list(dict.fromkeys)的作用是生成unique的item id\n    unique_aids = list(dict.fromkeys(aids[::-1] ))  #[242544, 619488, 579241, 700554]\n    # RERANK CANDIDATES USING WEIGHTS\n    if len(unique_aids)>=20:\n        # logspace的意义和linespace的定义类似，是以固定的距离分配值。以len(aids)为N，\n        # 即面对着相同的item数量，如25个，将分配25个介于0.4-1的点，而且只要item数量相同，这些点的值就相同，即weight相同\n        # 在时间上越接近的item，被赋的权重越多，越趋近1.\n        weights=np.logspace(0.1,1,len(aids),base=2, endpoint=True)-1\n        # 从collection中引入的Counter\n        aids_temp = Counter() \n        # RERANK BASED ON REPEAT ITEMS AND TYPE OF ITEMS\n        for aid,w,t in zip(aids,weights,types): \n            # 权重是累加的，所以没用unique aid，而是用的原始的aid，出现次数越多，得分越多\n            # Counter({619488: 2.122883676176261, 242544: 1.0, 579241: 0.624504792712471, 700554: 0.41421356237309515})\n            aids_temp[aid] += w * type_weight_multipliers[t]\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    # USE \"CLICKS\" CO-VISITATION MATRIX\n    # 每一个item id都在top_20_clicks对应着20个item，所以如果unique_aids <20, aids是这些unique item对应的20个商品的总和\n    # 即 n 是unique item的数量，aids2的list得到的是 n * 20个商品\n    aids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    # RERANK CANDIDATES\n    # 把这些n * 20的商品按照出现频率统计，只保留最高的前20个item\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(20) if aid2 not in unique_aids]   \n    # 补齐unique item直到20个\n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    # USE TOP20 TEST CLICKS\n    # 如果实在凑不到20个，就用全体数据中点击量最高的20个补充\n    return result + list(top_clicks)[:20-len(result)]\n\ndef suggest_buys(df):\n    \"\"\"\n    以每个session为基准，先把这个session访问过的aid提取出，按照时间倒序（时间上越近越排在list前面）生成unique的item id列表\n    同时准备一个unique buys列表，如果有session有过将某item放入cart或order的记录，则这些item就是unique buys，否则为空\n    如果unique aid大于20个：\n        给所有item赋权，包括时间上和types，时间越近权重越多，交互次数越多权重越多。\n        再从top_20_buy2buy中取出unique buys对应的20个商品。如果aids_temp中有物品在这些商品中，则+0.1分\n        提取出权重最大的20个item id，如果session没有cart或order的物品，则返回其unique aid。\n    如果unique aid不到20个：\n        从top_20_buys中找到unique item对应的top 20个物品，按照出现次数排序\n        从top_20_buy2buy中找到unique buy对应的top 20个物品，按照出现次数排序\n        截取这两大类物品，和unique item补全到20个商品\n        如果还凑不满20个，用全体order中的top 20个来补全\n    返回的结果是每个session得到一个20个item list\n    \"\"\"\n    # USE USER HISTORY AIDS AND TYPES\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    # UNIQUE AIDS AND UNIQUE BUYS\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    # 并不是所有session都有这个df，只有少量的有过cart和order的session才有\n    df = df.loc[(df['type']==1)|(df['type']==2)]\n    # unique buys是以上面的df为基础\n    unique_buys = list(dict.fromkeys( df.aid.tolist()[::-1] ))\n    # RERANK CANDIDATES USING WEIGHTS\n    if len(unique_aids)>=20:\n        weights=np.logspace(0.5,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        # RERANK BASED ON REPEAT ITEMS AND TYPE OF ITEMS\n        for aid,w,t in zip(aids,weights,types):\n            aids_temp[aid] += w * type_weight_multipliers[t]\n        # RERANK CANDIDATES USING \"BUY2BUY\" CO-VISITATION MATRIX\n        # 从top_20_buy2buy中找到unique buys中每件item对应的20个item\n        # 如果没有unique buy，此值为空\n        aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n        # 给counter中（即aids_temp)的每一个在aids3中出现的aid加0.1的权重\n        # 而且如果没在aids_temp中出现过（即没在test_df的aid中），也新增至aids_temp中，并加上0.1的权重\n        for aid in aids3: aids_temp[aid] += 0.1\n        ## 如果session没有cart和order，此时sorted_aids是其unique aid\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    # USE \"CART ORDER\" CO-VISITATION MATRIX\n    # 从top_20_buys中找到unique item对应的top 20个物品，按照出现次数排序\n    aids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_aids if aid in top_20_buys]))\n    # USE \"BUY2BUY\" CO-VISITATION MATRIX\n    # 从top_20_buy2buy中找到unique buy对应的top 20个物品，按照出现次数排序\n    aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n    # RERANK CANDIDATES\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2+aids3).most_common(20) if aid2 not in unique_aids] \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    # USE TOP20 TEST ORDERS\n    return result + list(top_orders)[:20-len(result)]","metadata":{"execution":{"iopub.status.busy":"2022-12-16T13:35:02.712759Z","iopub.execute_input":"2022-12-16T13:35:02.713158Z","iopub.status.idle":"2022-12-16T13:35:02.728511Z","shell.execute_reply.started":"2022-12-16T13:35:02.713125Z","shell.execute_reply":"2022-12-16T13:35:02.727325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create Submission CSV\n\nInferring test data with Pandas groupby is slow. We need to accelerate the following code.","metadata":{}},{"cell_type":"markdown","source":"For sample test data","metadata":{}},{"cell_type":"code","source":"type_weight_multipliers = {0: 1, 1: 6, 2: 3}\ndef clicks(df):\n    print(\"start\")\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    print(unique_aids)\n    if len(unique_aids)>=20:\n        weights=np.logspace(0.5,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter()\n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        print(aids_temp)\n        print(sorted_aids)\n    aids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    #print(aids2)\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(20) if aid2 not in unique_aids]\n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    print(result)\n    \ntest_df_sample.sort_values([\"session\", 'ts']).groupby([\"session\"]).apply(lambda x : clicks(x))","metadata":{"execution":{"iopub.status.busy":"2022-12-22T19:47:49.537705Z","iopub.execute_input":"2022-12-22T19:47:49.538075Z","iopub.status.idle":"2022-12-22T19:47:49.560184Z","shell.execute_reply.started":"2022-12-22T19:47:49.538039Z","shell.execute_reply":"2022-12-22T19:47:49.559052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_list = [171529, 1468470, 1356236, 144290, 657571, 1176975, 480646, 1593125, 1574899, 1793839, 1247728, 1014863, 1394214, 268469, 1734926, 987503, 1782437, 1119232, 58541, 521405, 115404, 1176975, 169239, 1734926, 144290, 520311, 1574899, 171529, 657571, 1073189, 987503, 635996, 260496, 1622692, 1574726, 1247728, 96166, 571090, 621705, 1014863, 169239, 521405, 1782437, 1574899, 1247728, 1363454, 115404, 171529, 1014863, 947372, 1468470, 1434391, 1029598, 1801037, 1509570, 1370415, 1176975, 1775976, 722095, 1074570, 169239, 115404, 171529, 1468470, 1247728, 1782437, 1014863, 1775976, 1176975, 1734926, 144290, 10826, 521405, 722095, 55117, 520311, 823801, 1370415, 260496, 657571]\nfor aid2, cnt in Counter(test_list).most_common(25):\n    print(aid2, cnt)","metadata":{"execution":{"iopub.status.busy":"2022-12-22T20:08:33.472082Z","iopub.execute_input":"2022-12-22T20:08:33.472439Z","iopub.status.idle":"2022-12-22T20:08:33.482880Z","shell.execute_reply.started":"2022-12-22T20:08:33.472406Z","shell.execute_reply":"2022-12-22T20:08:33.481704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def buys(df):\n    print(\"start\")\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    print(unique_aids)\n    df = df.loc[(df['type']==1)|(df['type']==2)]\n    unique_buys = list(dict.fromkeys( df.aid.tolist()[::-1] ))\n    if len(unique_aids)>=2:\n        weights=np.logspace(0.5,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        # RERANK BASED ON REPEAT ITEMS AND TYPE OF ITEMS\n        for aid,w,t in zip(aids,weights,types):\n            aids_temp[aid] += w * type_weight_multipliers[t]\n        print(aids_temp)\n        # RERANK CANDIDATES USING \"BUY2BUY\" CO-VISITATION MATRIX\n        aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n        for aid in aids3: aids_temp[aid] += 0.1\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        print(aids3)\n        print(\"after: \", sorted_aids)\n    \ntest_df_sample.sort_values([\"session\", 'ts']).groupby([\"session\"]).apply(lambda x : buys(x))","metadata":{"execution":{"iopub.status.busy":"2022-12-22T20:56:29.351934Z","iopub.execute_input":"2022-12-22T20:56:29.352641Z","iopub.status.idle":"2022-12-22T20:56:29.381078Z","shell.execute_reply.started":"2022-12-22T20:56:29.352604Z","shell.execute_reply":"2022-12-22T20:56:29.380138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n## 一个session一个session的apply suggest_clicks函数\n\"\"\"\nstart\n    session     aid          ts  type\n0  12089221  700554  1661448002     0\n1  12089221  619488  1661448024     0\n2  12089221  579241  1661449547     0\nstart\n    session     aid          ts  type\n6  12089222  127728  1661448002     0\nstart\n     session      aid          ts  type\n7   12089223  1574899  1661448002     0\n8   12089223  1574899  1661448028     1\n9   12089223    10826  1661448070     0\nstart\n     session     aid          ts  type\n15  12089224  837866  1661448002     0\n\"\"\"\npred_df_clicks = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(\n    lambda x: suggest_clicks(x)\n)\n\npred_df_buys = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(\n    lambda x: suggest_buys(x)\n)","metadata":{"execution":{"iopub.status.busy":"2022-12-16T13:35:07.031717Z","iopub.execute_input":"2022-12-16T13:35:07.032107Z","iopub.status.idle":"2022-12-16T14:02:25.038802Z","shell.execute_reply.started":"2022-12-16T13:35:07.032073Z","shell.execute_reply":"2022-12-16T14:02:25.037737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clicks_pred_df = pd.DataFrame(pred_df_clicks.add_suffix(\"_clicks\"), columns=[\"labels\"]).reset_index()\norders_pred_df = pd.DataFrame(pred_df_buys.add_suffix(\"_orders\"), columns=[\"labels\"]).reset_index()\ncarts_pred_df = pd.DataFrame(pred_df_buys.add_suffix(\"_carts\"), columns=[\"labels\"]).reset_index()","metadata":{"execution":{"iopub.status.busy":"2022-12-16T14:02:55.611478Z","iopub.execute_input":"2022-12-16T14:02:55.612097Z","iopub.status.idle":"2022-12-16T14:02:58.806999Z","shell.execute_reply.started":"2022-12-16T14:02:55.612058Z","shell.execute_reply":"2022-12-16T14:02:58.806007Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_df = pd.concat([clicks_pred_df, orders_pred_df, carts_pred_df])\npred_df.columns = [\"session_type\", \"labels\"]\npred_df[\"labels\"] = pred_df.labels.apply(lambda x: \" \".join(map(str,x)))\npred_df.to_csv(\"validation_preds.csv\", index=False)\npred_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-16T14:03:38.324469Z","iopub.execute_input":"2022-12-16T14:03:38.324932Z","iopub.status.idle":"2022-12-16T14:04:18.118612Z","shell.execute_reply.started":"2022-12-16T14:03:38.324896Z","shell.execute_reply":"2022-12-16T14:04:18.117618Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Compute Validation Score\n\nThis code is from Radek here. It has been modified to use less memory.","metadata":{}},{"cell_type":"code","source":"# FREE MEMORY\ndel pred_df_clicks, pred_df_buys, clicks_pred_df, orders_pred_df, carts_pred_df\ndel top_20_clicks, top_20_buy2buy, top_20_buys, top_clicks, top_orders, test_df\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-12-12T20:56:18.949333Z","iopub.execute_input":"2022-12-12T20:56:18.950302Z","iopub.status.idle":"2022-12-12T20:56:21.756398Z","shell.execute_reply.started":"2022-12-12T20:56:18.950252Z","shell.execute_reply":"2022-12-12T20:56:21.755446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# COMPUTE METRIC\nscore = 0\nweights = {'clicks': 0.10, 'carts': 0.30, 'orders': 0.60}\nfor t in ['clicks','carts','orders']:\n    sub = pred_df.loc[pred_df.session_type.str.contains(t)].copy()\n    sub['session'] = sub.session_type.apply(lambda x: int(x.split('_')[0]))\n    sub.labels = sub.labels.apply(lambda x: [int(i) for i in x.split(' ')[:20]])\n    test_labels = pd.read_parquet('../input/otto-validation/test_labels.parquet')\n    test_labels = test_labels.loc[test_labels['type']==t]\n    test_labels = test_labels.merge(sub, how='left', on=['session'])\n    test_labels['hits'] = test_labels.apply(lambda df: len(set(df.ground_truth).intersection(set(df.labels))), axis=1)\n    test_labels['gt_count'] = test_labels.ground_truth.str.len().clip(0,20)\n    recall = test_labels['hits'].sum() / test_labels['gt_count'].sum()\n    score += weights[t]*recall\n    print(f'{t} recall =',recall)\n    \nprint('=============')\nprint('Overall Recall =',score)\nprint('=============')","metadata":{"execution":{"iopub.status.busy":"2022-12-12T21:03:45.165703Z","iopub.execute_input":"2022-12-12T21:03:45.167252Z","iopub.status.idle":"2022-12-12T21:05:37.555538Z","shell.execute_reply.started":"2022-12-12T21:03:45.167208Z","shell.execute_reply":"2022-12-12T21:05:37.554483Z"},"trusted":true},"execution_count":null,"outputs":[]}]}