{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**What are you trying to do in this notebook?**\n\nTo predict e-commerce clicks, cart additions, and orders. I'll build a multi-objective recommender system based on previous events in a user session.\n\nMy work will help improve the shopping experience for everyone involved. Customers will receive more tailored recommendations while online retailers may increase their sales.\n\n**Why are you trying it?**\n\nTo help online retailers that will select more relevant items from a vast range to recommend to their customers based on their real-time behavior. Improving recommendations will ensure navigating through seemingly endless options is more effortless and engaging for shoppers.\n\nIn this notebook, I'm presenting the \"candidate rerank\" model using handcrafted rules. I can improve this model by engineering features, merging them into items and users, and training a reranker model (such as XGB) to choose my final 20. Furthermore to tune and improve this notebook, I build a local CV scheme to experiment new logic and/or models.","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"execution":{"iopub.status.busy":"2022-12-13T05:57:50.178215Z","iopub.execute_input":"2022-12-13T05:57:50.178597Z","iopub.status.idle":"2022-12-13T05:57:50.192875Z","shell.execute_reply.started":"2022-12-13T05:57:50.178569Z","shell.execute_reply":"2022-12-13T05:57:50.191755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"VER = 5\nfrom tqdm.notebook import tqdm\nimport os, sys, pickle, glob, gc\nfrom collections import Counter\nimport cudf, itertools\nprint('We will use RAPIDS version',cudf.__version__)","metadata":{"execution":{"iopub.status.busy":"2022-12-13T05:57:50.195036Z","iopub.execute_input":"2022-12-13T05:57:50.195805Z","iopub.status.idle":"2022-12-13T05:57:50.202867Z","shell.execute_reply.started":"2022-12-13T05:57:50.195769Z","shell.execute_reply":"2022-12-13T05:57:50.201518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\ndef read_file(f):\n    return cudf.DataFrame( data_cache[f] )\n\ndef read_file_to_cache(f):\n    df = pd.read_parquet(f)\n    df.ts = (df.ts/1000).astype('int32')\n    df['type'] = df['type'].map(type_labels).astype('int8')\n    return df","metadata":{"execution":{"iopub.status.busy":"2022-12-13T05:57:50.204222Z","iopub.execute_input":"2022-12-13T05:57:50.204936Z","iopub.status.idle":"2022-12-13T05:57:50.213565Z","shell.execute_reply.started":"2022-12-13T05:57:50.204903Z","shell.execute_reply":"2022-12-13T05:57:50.212505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_cache = {}\ntype_labels = {'clicks':0, 'carts':1, 'orders':2}\nfiles = glob.glob('../input/otto-chunk-data-inparquet-format/*_parquet/*')\nfor f in files: data_cache[f] = read_file_to_cache(f)","metadata":{"execution":{"iopub.status.busy":"2022-12-13T05:57:50.215943Z","iopub.execute_input":"2022-12-13T05:57:50.216811Z","iopub.status.idle":"2022-12-13T05:58:57.621293Z","shell.execute_reply.started":"2022-12-13T05:57:50.216777Z","shell.execute_reply":"2022-12-13T05:58:57.620318Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"READ_CT = 5\nCHUNK = int( np.ceil( len(files)/6 ))\nprint(f'We will process {len(files)} files, in groups of {READ_CT} and chunks of {CHUNK}.')","metadata":{"execution":{"iopub.status.busy":"2022-12-13T05:58:57.622583Z","iopub.execute_input":"2022-12-13T05:58:57.625112Z","iopub.status.idle":"2022-12-13T05:58:57.631770Z","shell.execute_reply.started":"2022-12-13T05:58:57.625073Z","shell.execute_reply":"2022-12-13T05:58:57.630536Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ntype_weight = {0:1, 1:6, 2:3}\n\nDISK_PIECES = 4\nSIZE = 1.86e6/DISK_PIECES\n\nfor PART in range(DISK_PIECES):\n    print()\n    print('### DISK PART',PART+1)\n    \n    for j in range(6):\n        a = j*CHUNK\n        b = min( (j+1)*CHUNK, len(files) )\n        print(f'Processing files {a} thru {b-1} in groups of {READ_CT}...')\n        \n        for k in range(a,b,READ_CT):\n            df = [read_file(files[k])]\n            for i in range(1,READ_CT): \n                if k+i < b: df.append(read_file(files[k+i]))\n            df = cudf.concat(df,ignore_index = True,axis = 0)\n            df = df.sort_values(['session','ts'],ascending=[True,False])\n            \n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n < 30].drop('n',axis=1)\n            \n            df = df.merge(df,on='session')\n            df = df.loc[ ((df.ts_x - df.ts_y).abs() < 24 * 60 * 60) & (df.aid_x != df.aid_y) ]\n            \n            df = df.loc[(df.aid_x >= PART*SIZE) & (df.aid_x < (PART+1)*SIZE)]\n            \n            df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n            df['wgt'] = df.type_y.map(type_weight)\n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n            \n            if k == a: tmp2 = df\n            else: tmp2 = tmp2.add(df,fill_value = 0)\n            print(k,', ',end='')\n        print()\n        if a==0: tmp = tmp2\n        else: tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    \n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n    \n    tmp = tmp.reset_index(drop = True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    tmp = tmp.loc[tmp.n < 15].drop('n',axis=1)\n    \n    tmp.to_pandas().to_parquet(f'top_15_carts_orders_v{VER}_{PART}.pqt')","metadata":{"execution":{"iopub.status.busy":"2022-12-13T05:58:57.634841Z","iopub.execute_input":"2022-12-13T05:58:57.635670Z","iopub.status.idle":"2022-12-13T06:02:48.620215Z","shell.execute_reply.started":"2022-12-13T05:58:57.635623Z","shell.execute_reply":"2022-12-13T06:02:48.617294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nDISK_PIECES = 1\nSIZE = 1.86e6/DISK_PIECES\n\nfor PART in range(DISK_PIECES):\n    print()\n    print('### DISK PART',PART+1)\n    \n    for j in range(6):\n        a = j*CHUNK\n        b = min( (j+1)*CHUNK, len(files) )\n        print(f'Processing files {a} thru {b-1} in groups of {READ_CT}...')\n        \n        for k in range(a,b,READ_CT):\n            df = [read_file(files[k])]\n            for i in range(1,READ_CT): \n                if k+i<b: df.append( read_file(files[k+i]) )\n            df = cudf.concat(df,ignore_index=True,axis=0)\n            df = df.loc[df['type'].isin([1,2])] \n            df = df.sort_values(['session','ts'],ascending=[True,False])\n            \n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n<30].drop('n',axis=1)\n            \n            df = df.merge(df,on='session')\n            df = df.loc[ ((df.ts_x - df.ts_y).abs()< 14 * 24 * 60 * 60) & (df.aid_x != df.aid_y) ] \n            \n            df = df.loc[(df.aid_x >= PART*SIZE)&(df.aid_x < (PART+1)*SIZE)]\n            \n            df = df[['session', 'aid_x', 'aid_y','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n            df['wgt'] = 1\n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n            if k==a: tmp2 = df\n            else: tmp2 = tmp2.add(df, fill_value=0)\n            print(k,', ',end='')\n        print()\n        \n        if a==0: tmp = tmp2\n        else: tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    \n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n    \n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    tmp = tmp.loc[tmp.n<15].drop('n',axis=1)\n    \n    tmp.to_pandas().to_parquet(f'top_15_buy2buy_v{VER}_{PART}.pqt')","metadata":{"execution":{"iopub.status.busy":"2022-12-13T06:02:48.622109Z","iopub.execute_input":"2022-12-13T06:02:48.622493Z","iopub.status.idle":"2022-12-13T06:03:26.931596Z","shell.execute_reply.started":"2022-12-13T06:02:48.622458Z","shell.execute_reply":"2022-12-13T06:03:26.930440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nDISK_PIECES = 4\nSIZE = 1.86e6/DISK_PIECES\n\nfor PART in range(DISK_PIECES):\n    print()\n    print('### DISK PART',PART+1)\n    \n    for j in range(6):\n        a = j*CHUNK\n        b = min( (j+1)*CHUNK, len(files) )\n        print(f'Processing files {a} thru {b-1} in groups of {READ_CT}...')\n        \n        for k in range(a,b,READ_CT):\n            df = [read_file(files[k])]\n            for i in range(1,READ_CT): \n                if k+i<b: df.append( read_file(files[k+i]) )\n            df = cudf.concat(df,ignore_index=True,axis=0)\n            df = df.sort_values(['session','ts'],ascending=[True,False])\n            \n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n<30].drop('n',axis=1)\n            df = df.merge(df,on='session')\n            df = df.loc[ ((df.ts_x - df.ts_y).abs()< 24 * 60 * 60) & (df.aid_x != df.aid_y) ]\n            \n            df = df.loc[(df.aid_x >= PART*SIZE)&(df.aid_x < (PART+1)*SIZE)]\n            \n            df = df[['session', 'aid_x', 'aid_y','ts_x']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n            df['wgt'] = 1 + 3*(df.ts_x - 1659304800)/(1662328791-1659304800)\n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n            \n            if k==a: tmp2 = df\n            else: tmp2 = tmp2.add(df, fill_value=0)\n            print(k,', ',end='')\n        print()\n        if a==0: tmp = tmp2\n        else: tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    \n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n    \n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    tmp = tmp.loc[tmp.n<20].drop('n',axis=1)\n    \n    tmp.to_pandas().to_parquet(f'top_20_clicks_v{VER}_{PART}.pqt')","metadata":{"execution":{"iopub.status.busy":"2022-12-13T06:03:26.935389Z","iopub.execute_input":"2022-12-13T06:03:26.935678Z","iopub.status.idle":"2022-12-13T06:07:15.930085Z","shell.execute_reply.started":"2022-12-13T06:03:26.935652Z","shell.execute_reply":"2022-12-13T06:07:15.928879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del data_cache, tmp\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-12-13T06:07:15.931888Z","iopub.execute_input":"2022-12-13T06:07:15.934435Z","iopub.status.idle":"2022-12-13T06:07:17.439181Z","shell.execute_reply.started":"2022-12-13T06:07:15.934396Z","shell.execute_reply":"2022-12-13T06:07:17.434244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def load_test():    \n    dfs = []\n    for e, chunk_file in enumerate(glob.glob('../input/otto-chunk-data-inparquet-format/test_parquet/*')):\n        chunk = pd.read_parquet(chunk_file)\n        chunk.ts = (chunk.ts/1000).astype('int32')\n        chunk['type'] = chunk['type'].map(type_labels).astype('int8')\n        dfs.append(chunk)\n    return pd.concat(dfs).reset_index(drop=True) \n\ntest_df = load_test()\nprint('Test data has shape',test_df.shape)\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-13T06:07:17.443286Z","iopub.execute_input":"2022-12-13T06:07:17.446948Z","iopub.status.idle":"2022-12-13T06:07:19.565114Z","shell.execute_reply.started":"2022-12-13T06:07:17.446910Z","shell.execute_reply":"2022-12-13T06:07:19.564104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ndef pqt_to_dict(df):\n    return df.groupby('aid_x').aid_y.apply(list).to_dict()\n\ntop_20_clicks = pqt_to_dict(pd.read_parquet(f'top_20_clicks_v{VER}_0.pqt') )\nfor k in range(1,DISK_PIECES): \n    top_20_clicks.update(pqt_to_dict(pd.read_parquet(f'top_20_clicks_v{VER}_{k}.pqt') ) )\n    \ntop_20_buys = pqt_to_dict( pd.read_parquet(f'top_15_carts_orders_v{VER}_0.pqt') )\nfor k in range(1,DISK_PIECES): \n    top_20_buys.update( pqt_to_dict( pd.read_parquet(f'top_15_carts_orders_v{VER}_{k}.pqt') ) )\n\ntop_20_buy2buy = pqt_to_dict(pd.read_parquet(f'top_15_buy2buy_v{VER}_0.pqt') )\n\nprint('Here are size of our 3 co-visitation matrices:')\nprint( len(top_20_clicks), len(top_20_buy2buy), len(top_20_buys) )","metadata":{"execution":{"iopub.status.busy":"2022-12-13T06:07:19.566644Z","iopub.execute_input":"2022-12-13T06:07:19.567001Z","iopub.status.idle":"2022-12-13T06:08:56.660776Z","shell.execute_reply.started":"2022-12-13T06:07:19.566967Z","shell.execute_reply":"2022-12-13T06:08:56.658950Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"top_clicks = test_df.loc[test_df['type']== 0,'aid'].value_counts().index.values[:20] ","metadata":{"execution":{"iopub.status.busy":"2022-12-13T06:08:56.662264Z","iopub.execute_input":"2022-12-13T06:08:56.662608Z","iopub.status.idle":"2022-12-13T06:08:56.925193Z","shell.execute_reply.started":"2022-12-13T06:08:56.662574Z","shell.execute_reply":"2022-12-13T06:08:56.924219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"top_carts = test_df.loc[test_df['type']== 1,'aid'].value_counts().index.values[:20]","metadata":{"execution":{"iopub.status.busy":"2022-12-13T06:08:56.926714Z","iopub.execute_input":"2022-12-13T06:08:56.927072Z","iopub.status.idle":"2022-12-13T06:08:56.980623Z","shell.execute_reply.started":"2022-12-13T06:08:56.927036Z","shell.execute_reply":"2022-12-13T06:08:56.979684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"top_orders = test_df.loc[test_df['type']== 2,'aid'].value_counts().index.values[:20]","metadata":{"execution":{"iopub.status.busy":"2022-12-13T06:08:56.982191Z","iopub.execute_input":"2022-12-13T06:08:56.982565Z","iopub.status.idle":"2022-12-13T06:08:56.999876Z","shell.execute_reply.started":"2022-12-13T06:08:56.982530Z","shell.execute_reply":"2022-12-13T06:08:56.999009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"type_weight_multipliers = {0: 1, 1: 6, 2: 3}\n\ndef suggest_clicks(df):\n    aids = df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1]))\n    \n    if len(unique_aids) >= 20:\n        weights = np.logspace(0.1,1,len(aids),base = 2, endpoint = True)-1\n        aids_temp = Counter() \n        \n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    \n    aids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    \n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2).most_common(20) if aid2 not in unique_aids]    \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    \n    return result + list(top_clicks)[:20-len(result)]","metadata":{"execution":{"iopub.status.busy":"2022-12-13T06:08:57.001406Z","iopub.execute_input":"2022-12-13T06:08:57.001764Z","iopub.status.idle":"2022-12-13T06:08:57.042524Z","shell.execute_reply.started":"2022-12-13T06:08:57.001730Z","shell.execute_reply":"2022-12-13T06:08:57.041565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def suggest_carts(df):\n    aids = df.aid.tolist()\n    types = df.type.tolist()\n    \n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    df = df.loc[(df['type'] == 0)|(df['type'] == 1)]\n    unique_buys = list(dict.fromkeys(df.aid.tolist()[::-1]))\n    \n    if len(unique_aids) >= 20:\n        weights=np.logspace(0.5,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        \n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        \n        aids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_buys if aid in top_20_buys]))\n        for aid in aids2: aids_temp[aid] += 0.1\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    \n    aids1 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    aids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_aids if aid in top_20_buys]))\n    \n    top_aids2 = [aid2 for aid2, cnt in Counter(aids1+aids2).most_common(20) if aid2 not in unique_aids] \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    \n    return result + list(top_carts)[:20-len(result)]","metadata":{"execution":{"iopub.status.busy":"2022-12-13T06:08:57.044073Z","iopub.execute_input":"2022-12-13T06:08:57.044452Z","iopub.status.idle":"2022-12-13T06:08:57.057479Z","shell.execute_reply.started":"2022-12-13T06:08:57.044414Z","shell.execute_reply":"2022-12-13T06:08:57.056378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def suggest_buys(df):\n    aids = df.aid.tolist()\n    types = df.type.tolist()\n    \n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    df = df.loc[(df['type']==1)|(df['type']==2)]\n    unique_buys = list(dict.fromkeys(df.aid.tolist()[::-1]))\n    \n    if len(unique_aids) >= 20:\n        weights=np.logspace(0.5,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        \n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        \n        aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n        for aid in aids3: aids_temp[aid] += 0.1\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n        \n    aids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_aids if aid in top_20_buys]))\n    aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n    \n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2+aids3).most_common(20) if aid2 not in unique_aids] \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    \n    return result + list(top_orders)[:20-len(result)]","metadata":{"execution":{"iopub.status.busy":"2022-12-13T06:08:57.058793Z","iopub.execute_input":"2022-12-13T06:08:57.059231Z","iopub.status.idle":"2022-12-13T06:08:57.072299Z","shell.execute_reply.started":"2022-12-13T06:08:57.059194Z","shell.execute_reply":"2022-12-13T06:08:57.071342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\npred_df_clicks = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(\n    lambda x: suggest_clicks(x)\n)\n\npred_df_carts = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(\n    lambda x: suggest_carts(x)\n)\n\npred_df_buys = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).apply(\n    lambda x: suggest_buys(x)\n)","metadata":{"execution":{"iopub.status.busy":"2022-12-13T06:08:57.073774Z","iopub.execute_input":"2022-12-13T06:08:57.074111Z","iopub.status.idle":"2022-12-13T06:57:06.496628Z","shell.execute_reply.started":"2022-12-13T06:08:57.074078Z","shell.execute_reply":"2022-12-13T06:57:06.495551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clicks_pred_df = pd.DataFrame(pred_df_clicks.add_suffix(\"_clicks\"), columns=[\"labels\"]).reset_index()\ncarts_pred_df = pd.DataFrame(pred_df_carts.add_suffix(\"_carts\"), columns=[\"labels\"]).reset_index()\norders_pred_df = pd.DataFrame(pred_df_buys.add_suffix(\"_orders\"), columns=[\"labels\"]).reset_index()","metadata":{"execution":{"iopub.status.busy":"2022-12-13T06:57:06.500615Z","iopub.execute_input":"2022-12-13T06:57:06.501624Z","iopub.status.idle":"2022-12-13T06:57:09.448837Z","shell.execute_reply.started":"2022-12-13T06:57:06.501587Z","shell.execute_reply":"2022-12-13T06:57:09.447669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_df = pd.concat([clicks_pred_df, carts_pred_df, orders_pred_df])\npred_df.columns = [\"session_type\", \"labels\"]\npred_df[\"labels\"] = pred_df.labels.apply(lambda x: \" \".join(map(str,x)))\npred_df.to_csv(\"submission.csv\", index=False)\npred_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-13T06:57:09.450444Z","iopub.execute_input":"2022-12-13T06:57:09.451058Z","iopub.status.idle":"2022-12-13T06:57:53.817687Z","shell.execute_reply.started":"2022-12-13T06:57:09.451017Z","shell.execute_reply":"2022-12-13T06:57:53.816212Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Did it work?**\n\nYes, it works. The training data contains full e-commerce session information. For each session in the test data, my task is it to predict the aid values for each session type that occurs after the last timestamp in the test session.\n\nIn other words, the test data contain sessions truncated by timestamp, and I'm predicting what occurs after the point of truncation.\n\n\n**Share your feedback to accomplish the goal!!!**","metadata":{}}]}