{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# OTTOコンペティションの概要\nOTTOはドイツで流行っているeコマースサイト。\n\nレコメンド改善のためのユーザー行動の予測精度を競うコンペティション。\n\n# 予測すること\nユーザーがクリック、またはカートイン、または注文する商品コード（aid）を予測する。\n\n# データの内容\n* session：ユーザー\n* aid：クリック、カートイン、オーダーされた商品コード\n* type：クリック、カートイン、オーダーのうち、どのアクションを行ったか\n* ts：ユーザーがクリック、カートイン、オーダーした日時（UNIXTIME形式）\n\n# Kaggleから与えられるデータ\n\njson形式なのでこのままでは使えない。\n\n<img src=\"https://user-images.githubusercontent.com/101812292/226284383-f802cad4-6ef8-4ea3-a609-78f3d541bd72.jpg\" width=\"350\">\n\n# cudfとparquetファイルで処理を高速化。\n\nデータフレームにできるよう、parquetファイルに変換。\n\nparquetファイルは読み込みと書き出しが速い。\n\n膨大なデータを処理するため、GPUで処理するcudfを使う。\n\nclicks=0,carts=1,orders=2に変換されている。\n\n<img src=\"https://user-images.githubusercontent.com/101812292/226283458-ac039f4f-4426-4a75-bd4d-8d3c233d07d9.jpg\" width=\"400\">\n\n# トレーニングデータの内容\n\n3週間分のユーザーの行動履歴。\n\n<img src=\"https://user-images.githubusercontent.com/101812292/226299762-d3ddcd8d-2046-4549-b986-a1d33b58288b.jpg\" width=\"1000\">\n\n# テストデータの内容\n\n1週間分のユーザーの行動履歴。\n\n### **途中で切り捨てられている**ので、切り捨てられた先のaidsを予測する\n### 1sessionあたり20のaidを予測する\n\n<img src=\"https://user-images.githubusercontent.com/101812292/226299696-f58d7df8-2190-498d-b9db-e18f08347fba.jpg\" width=\"1000\">\n\n# 予測したことの詳細\n### ①トレーニングデータから共起行列を作成する\n\n共起行列は、一緒にクリックされた、もしくはカートイン、オーダーされたaidを2列にしたもの。\n\n図をトレーニングデータの共起とする。\n\n図の場合、テストデータにaid1がある時、aid5,aid7,aid2が予測の候補になる。\n\n<img src=\"https://user-images.githubusercontent.com/101812292/227668753-09bcd89f-addb-4154-bbe7-4ac63eba4cb4.jpg\" width=\"300\">\n\n### ②作成した共起行列に、重みづけルールを作成し、重み順に並び替える\n\n事例の図はオーダーに6、カートに3、クリックに1の重みづけルールを作成して、重み順に並び替える。\n\nテストデータでaid1があれば、aid7が予測aidの最有力候補。\n\nテストデータでaid2があれば、aid8が予測aidの最有力候補。\n\n<img src=\"https://user-images.githubusercontent.com/101812292/227668924-647966e4-5d0c-4df9-9bb5-de8541adba00.jpg\" width=\"400\">\n\n### ③重みづけルールを複数作成し、足し合わせて最終的な20のaidを選択する。\n\n<img src=\"https://user-images.githubusercontent.com/101812292/227465463-1b05fccd-04a5-4d12-b6cf-552b1071935e.jpg\" width=\"600\">\n\n\n\n# 作成した重みづけルール\n### ①Pytorchを利用した行列分解による重みづけ\n\n[このノート](https://www.kaggle.com/code/tashiget/pytorch)により、aidの並び順から埋め込みベクトルを出力し、2次元の座標にする。\n\n座標の距離を計算し、近い順に重みづけをする。\n\n<img src=\"https://user-images.githubusercontent.com/101812292/228710328-172179e0-67e4-4cc5-a259-79360326b242.jpg\" width=\"1200\">\n\n#### ※コンペ終了後で気づいたけど、座標間の距離でなく、内積を出して大きい順に並べた方が精度高そう\n\n### 行列分解とは？\n\naidの特徴を出して、aid同士がどのくらい類似性があるかがわかる。\n\nPytorchで与える情報はaidの並び順のみで、特徴ベクトルを出すことができる。（図ではx座標とy座標の2次元の特徴ベクトルを出した。）\n\n※クリック、カート、オーダーの情報はいらない。\n\n特徴ベクトルの内積が大きいほど、aidは類似している。\n\n### ②typeによる重みづけ　クリック1、カート3、オーダー6\n\nクリック＜カートイン＜オーダーの順で重要。\n\n<img src=\"https://user-images.githubusercontent.com/101812292/227668924-647966e4-5d0c-4df9-9bb5-de8541adba00.jpg\" width=\"400\">\n\n### ③TF、IDF、BM25による重みづけルールを追加した。検索のランキングに使われる重みづけ\n#### TF値\n\n頻繁に出現するaidであるほど重みが高い\n\n→たくさん出てくるaidは重要度が高い。\n\n<img src=\"https://user-images.githubusercontent.com/101812292/227671817-691b9722-22ee-4cf7-ab33-f64cdff404f0.jpg\" width=\"800\">\n\n#### IDF値\n\n複数のaidをまたいでいないほど重みが高い。\n\n<img src=\"https://user-images.githubusercontent.com/101812292/227671630-9d11f017-1904-47ed-ab35-2b5b117f3e0a.jpg\" width=\"800\">\n\n→例えば、コーラを買った人にレコメンドすることを考える。\n\n誰でも買うような水ではなく、メロンソーダの方が重要度が高くなる。\n\n<img src=\"https://user-images.githubusercontent.com/101812292/229458242-43c86bd7-a70f-4dbf-b4fe-57a2e7079688.jpg\" width=\"800\">\n\n#### BM25\n\naidの種類が少ないsessionに重みをつける\n\n→なんでもかんでも商品を買う人よりも、少しの商品を買う人の方がレコメンドの参考にできる。\n\n<img src=\"https://user-images.githubusercontent.com/101812292/227670663-a0b2ce4a-259e-432e-9ea6-520e9f177747.jpg\" width=\"800\">\n\n### ④重みづけルールの全てに、直近であるほど重みがつくようにした。\n\n新しいほど重要度を高くした。\n\n商材のトレンドも加味できる。\n\n<img src=\"https://user-images.githubusercontent.com/101812292/227669807-704a54d7-9ec3-4ede-aec4-5041cb6ca7fc.jpg\" width=\"400\">\n\n\n\n# このコンペで重要なことと感想\n### 共起行列のaidをいかに精度高く並び替えるか\n\n正解する可能性の高い順番に商品をレコメンドしたい。\n\n精度の高い重みづけルールを複数作り、アンサンブルすることで予測精度が上がると考えた。\n\n### いろんなルールを試した中で、直近情報の凄さを感じた。\n\n全ての重みづけルールに、直近情報の重づけルールを掛け合わせてスコアが改善した。\n\n機械学習において直近情報は貴重な情報であることを体感することができた。\n\n### レコメンド系コンペなので、LBのスコアを信じた。\n\n提出は140回したが、OTTOはLBがなかなか上がらなく、苦しかった。\n\nただ、レコメンド系コンペティションはshakeが少ないことを過去コンペから確認していたため、\n\nLBのスコアを信じることにした。\n\n結果、5のみshake downしたが、ブロンズを獲得することができた！\n\n#### 参考ノート\n[このノート](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575)をベースにした。ありがとう！","metadata":{}},{"cell_type":"code","source":"import pandas as pd, numpy as np\nfrom tqdm.notebook import tqdm\nimport os, sys, pickle, glob, gc, itertools\nfrom collections import Counter\nfrom tqdm.notebook import tqdm\nimport cudf\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"papermill":{"duration":3.036143,"end_time":"2022-11-10T16:03:24.014816","exception":false,"start_time":"2022-11-10T16:03:20.978673","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2023-03-29T09:48:57.717029Z","iopub.execute_input":"2023-03-29T09:48:57.717625Z","iopub.status.idle":"2023-03-29T09:49:01.002261Z","shell.execute_reply.started":"2023-03-29T09:48:57.717543Z","shell.execute_reply":"2023-03-29T09:49:01.001261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def read_file(f):\n    return cudf.DataFrame( data_cache[f] )\ndef read_pd_file(f):\n    return pd.DataFrame( data_cache[f] )\ndef read_file_to_cache(f):\n    df = pd.read_parquet(f)\n    #tsはカラム名、intに変換\n    df.ts = (df.ts/1000).astype('int32')\n    #クリック0,カート1,オーダー2の変換\n    df['type'] = df['type'].map(type_labels).astype('int8')\n    return df\n\ndata_cache = {}\ntype_labels = {'clicks':0, 'carts':1, 'orders':2}\nfiles = glob.glob('../input/otto-chunk-data-inparquet-format/*_parquet/*')\n\nVER=5\n\nREAD_CT = 5\nCHUNK = int( np.ceil( len(files)/6 ))\n\ntype_weight = {0: 1, 1: 3, 2: 6}","metadata":{"papermill":{"duration":0.063943,"end_time":"2022-11-10T16:03:24.091816","exception":false,"start_time":"2022-11-10T16:03:24.027873","status":"completed"},"tags":[],"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-03-29T09:49:01.007785Z","iopub.execute_input":"2023-03-29T09:49:01.010493Z","iopub.status.idle":"2023-03-29T09:49:01.094722Z","shell.execute_reply.started":"2023-03-29T09:49:01.010452Z","shell.execute_reply":"2023-03-29T09:49:01.093568Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 2次元に変換したaidごとの特徴ベクトルをx,y座標にプロットした\n# 特徴ベクトルの作成は別ノート\nvector=cudf.read_parquet(\"/kaggle/input/aid-vector/embe_2.pqt\")\n\nx_vector=vector.copy()\nx_vector.columns=[\"aid_x\",\"x_1\",\"y_1\"]\n\ny_vector=vector.copy()\ny_vector.columns=[\"aid_y\",\"x_2\",\"y_2\"]","metadata":{"execution":{"iopub.status.busy":"2023-03-29T09:49:43.716211Z","iopub.execute_input":"2023-03-29T09:49:43.716628Z","iopub.status.idle":"2023-03-29T09:49:46.153475Z","shell.execute_reply.started":"2023-03-29T09:49:43.716593Z","shell.execute_reply":"2023-03-29T09:49:46.152430Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"x_vector","metadata":{"execution":{"iopub.status.busy":"2023-03-29T09:49:55.772250Z","iopub.execute_input":"2023-03-29T09:49:55.772635Z","iopub.status.idle":"2023-03-29T09:49:55.844951Z","shell.execute_reply.started":"2023-03-29T09:49:55.772603Z","shell.execute_reply":"2023-03-29T09:49:55.843249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# このセルでは特徴ベクトルをx,y座標にプロットしたデータフレームを用いて、aidごとの距離を計算し\n# 距離が近い順に重みづけをする。\n\n# メモリエラーを避けるため、5回ループを回して実行する。\nDISK_PIECES = 5\nSIZE = 1.86e6/DISK_PIECES\n\nfor PART in tqdm(range(DISK_PIECES),total=DISK_PIECES):\n    for j in tqdm(range(6),total=6):\n        a = j*CHUNK\n        b = min( (j+1)*CHUNK, len(files) )\n        for k in range(a,b,READ_CT):\n            df = [read_file(files[k])]\n            for i in range(1,READ_CT): \n                if k+i<b: df.append( read_file(files[k+i]) )\n            df = cudf.concat(df,ignore_index=True,axis=0)\n            \n            df = df.sort_values(['session','ts'],ascending=[True,False])\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n<40].drop('n',axis=1)\n            df = df.merge(df,on='session')\n            df = df.loc[ ((df.ts_x - df.ts_y).abs()< 24 * 60 * 60) & (df.aid_x != df.aid_y) ]\n            df = df.loc[(df.aid_x >= PART*SIZE)&(df.aid_x < (PART+1)*SIZE)]\n            df = df.sort_values(['session','type_y'],ascending=[True,False])\n            df = df[['session', 'aid_x', 'aid_y','ts_x',\"type_y\"]].drop_duplicates(['session', 'aid_x', 'aid_y'])\n            \n            # typeによる重みづけも掛け合わせる\n            # 直近であるほど重みづけも掛け合わせる\n            df['wgt'] = df.type_y.map(type_weight)\n            df['wgt'] = df[\"wgt\"]*(1 + 3*(df.ts_x - 1659304800)/(1662328791-1659304800))\n    \n            # ここからaidごとの距離を計算する\n            # 本来は内積を計算した方が良さそう\n            df=df.merge(x_vector,on=\"aid_x\",how=\"left\")\n            df=df.merge(y_vector,on=\"aid_y\",how=\"left\")\n            df[\"distance\"]=np.sqrt((df[\"x_2\"]-df[\"x_1\"])**2+(df[\"y_2\"]-df[\"y_1\"])**2)\n            \n            # 距離が近いほど重みを高くしたいので、1-正規化する。\n            df[\"dist_reverse\"]=1 - ((df[\"distance\"]-df[\"distance\"].min()) / (df[\"distance\"].max() - df[\"distance\"].min()))\n            df[\"wgt\"]=df[\"wgt\"]*df[\"dist_reverse\"]\n        \n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n            if k==a: tmp2 = df\n            else: tmp2 = tmp2.add(df, fill_value=0)\n        if a==0: tmp = tmp2\n        else: tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    tmp = tmp.loc[tmp.n<20].drop('n',axis=1)\n    tmp.to_pandas().to_parquet(f'dist_1_3_6_{PART}.pqt')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# このセルではtypeによるクリック1、カート3、オーダー6の重みづけを行う。\n# type_weight = {0: 1, 1: 3, 2: 6}\nDISK_PIECES = 4\nSIZE = 1.86e6/DISK_PIECES\n\nfor PART in tqdm(range(DISK_PIECES),total=DISK_PIECES):\n    for j in tqdm(range(6),total=6):\n        a = j*CHUNK\n        b = min( (j+1)*CHUNK, len(files) )\n        for k in range(a,b,READ_CT):\n            # READ FILE\n            df = [read_file(files[k])]\n            for i in range(1,READ_CT): \n                if k+i<b: df.append( read_file(files[k+i]) )\n            df = cudf.concat(df,ignore_index=True,axis=0)\n            \n            df = df.sort_values(['session','ts'],ascending=[True,False])\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n<30].drop('n',axis=1)\n            df = df.merge(df,on='session')\n            df = df.loc[ ((df.ts_x - df.ts_y).abs()< 24 * 60 * 60) & (df.aid_x != df.aid_y) ]\n            df = df.loc[(df.aid_x >= PART*SIZE)&(df.aid_x < (PART+1)*SIZE)]\n            df = df.sort_values(['session','type_y'],ascending=[True,False])\n            df = df[['session', 'aid_x', 'aid_y','ts_x',\"type_y\"]].drop_duplicates(['session', 'aid_x', 'aid_y'])\n            \n            # type x 直近情報\n            #1662328791はmaxに近い、1ループ目でのmaxは1662328764\n            #1659304800はmin（1セッション目の1つ目のts）\n            df['wgt'] = df.type_y.map(type_weight)\n            df['wgt'] = df[\"wgt\"]*(1 + 3*(df.ts_x - 1659304800)/(1662328791-1659304800))\n        \n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n            if k==a: tmp2 = df\n            else: tmp2 = tmp2.add(df, fill_value=0)\n        if a==0: tmp = tmp2\n        else: tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    tmp = tmp.loc[tmp.n<15].drop('n',axis=1)\n    tmp.to_pandas().to_parquet(f'n15_log_1_3_6_{PART}.pqt')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#このセルではクリックaidの予測に使う共起行列を作成する\n# BM25を使う\nTFDISK = 5\nSIZE = 1.86e6/TFDISK\n\n# BM25に利用する係数\nk1=2\nbb=0.75\n\n# メモリエラーを避けるため、5回ループを回して処理する\nfor PART in tqdm(range(TFDISK),total=TFDISK):\n    for j in tqdm(range(6),total=6):\n        a = j*CHUNK\n        b = min( (j+1)*CHUNK, len(files) )\n        for k in range(a,b,READ_CT):\n            df = [read_file(files[k])]\n            for i in range(1,READ_CT): \n                if k+i<b: \n                    df.append( read_file(files[k+i]) )\n            df = cudf.concat(df,ignore_index=True,axis=0)\n            df = df.sort_values(['session','ts'],ascending=[True,False])\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n<40]\n            df=df.drop('n',axis=1)\n            \n            # NDL値、BM25の計算に使う\n            aid_nunique=df.drop_duplicates([\"session\",\"aid\"])\n            aid_nunique=aid_nunique.groupby(\"session\")[\"aid\"].nunique()\n            aid_nunique=aid_nunique.reset_index()\n            aid_nunique.columns=[\"session\",\"nunique\"]\n            aid_nunique[\"NDL\"] = aid_nunique[\"nunique\"] / aid_nunique[\"nunique\"].mean()\n            \n            df = df.merge(df,on='session')\n            df = df.loc[ ((df.ts_x - df.ts_y).abs()< 24 * 60 * 60) & (df.aid_x != df.aid_y) ]\n            df = df.loc[(df.aid_x >= PART*SIZE)&(df.aid_x < (PART+1)*SIZE)]\n            \n            # TF値\n            # aid_xのセッションごとの母数を出す\n            df_size=df[[\"session\",\"aid_x\"]]\n            df_size=df_size.reset_index(drop=True)\n            df_size[\"x_size\"]=df_size.groupby([\"session\",\"aid_x\"]).cumcount()+1\n            x_size=df_size.groupby([\"session\",\"aid_x\"],as_index=False).max()\n            df_size=df_size[[\"session\",\"aid_x\"]].merge(x_size,on=[\"session\",\"aid_x\"])\n            df_size=df_size.drop_duplicates()\n            df_tf=df[[\"session\",\"aid_x\",\"aid_y\"]]\n            #aid_yの出現回数を数える\n            df_tf=df_tf.reset_index(drop=True)\n            df_tf[\"y_cum\"]=df_tf.groupby([\"session\",\"aid_x\",\"aid_y\"]).cumcount()+1\n            \n            #aid_yの出現回数のmaxを出す\n            df_tf=df_tf.groupby([\"session\",\"aid_x\",\"aid_y\"],as_index=False).max()\n            \n            df_tf=df_tf.merge(df_size,on=[\"session\",\"aid_x\"])\n            df_tf[\"TF\"]=df_tf[\"y_cum\"]/df_tf[\"x_size\"]\n            df_tf=df_tf[[\"session\",\"aid_x\",\"aid_y\",\"TF\"]]\n\n            # IDF値\n            # session別で行列を作成後、aid_xをユニークにする\n            # あるaid_yがaid_xのユニークをいくつ跨いでいるか\n            df_idf=df_tf[[\"aid_x\",\"aid_y\"]].drop_duplicates()\n            df_idf[\"x_nunique\"]=df_idf[\"aid_x\"].nunique()\n            df_idf=df_idf.reset_index(drop=True)\n            df_idf[\"y_cum\"]=df_idf.groupby(\"aid_y\").cumcount()+1\n            y_cum=df_idf.groupby(\"aid_y\")[\"y_cum\"].max()\n            y_cum=y_cum.reset_index()\n            df_idf=df_idf[[\"aid_x\",\"aid_y\",\"x_nunique\"]].merge(y_cum,on=\"aid_y\")\n            df_idf=df_idf.sort_values([\"aid_x\",\"aid_y\"])\n            df_idf[\"IDF\"]=np.log(df_idf[\"x_nunique\"]/df_idf[\"y_cum\"])+1\n            df_idf=df_idf[[\"aid_x\",\"aid_y\",\"IDF\"]].reset_index(drop=True)\n\n            #TFIDF値\n            df_tfidf=df_tf.merge(df_idf,on=[\"aid_x\",\"aid_y\"])\n            df_tfidf[\"TFIDF\"]=df_tfidf[\"TF\"]*df_tfidf[\"IDF\"]\n            \n            df = df.sort_values(['session','type_y'],ascending=[True,False])\n            df = df[['session', 'aid_x', 'aid_y','ts_x','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n            df['wgt'] = 0\n            \n            df = df.merge(df_tfidf,on=[\"session\",\"aid_x\",\"aid_y\"])  \n            \n            # BM25\n            df=df.merge(aid_nunique,on=\"session\")\n            df[\"BM25\"]=(df[\"TF\"]*df[\"IDF\"]*(k1+1)) / (k1*(1-bb+(bb*df[\"NDL\"]))+df[\"TF\"])\n            \n            # wgtにBM25を入れる\n            df[\"wgt\"] += df[\"BM25\"]\n            \n            # 直近情報の重みづけルールを掛け合わせる\n            df['wgt'] = df[\"wgt\"]*(1 + 3*(df.ts_x - 1659304800)/(1662328791-1659304800))\n            \n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n\n            if k==a: \n                tmp2 = df\n            else: \n                tmp2 = tmp2.add(df, fill_value=0)\n        if a==0: \n            tmp = tmp2\n        else: \n            tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    tmp = tmp.reset_index()\n    \n    # 重みの高い順に並び替える\n    tmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    \n    # 上位20の共起aidを保存する\n    tmp = tmp.loc[tmp.n<20].drop('n',axis=1)\n    tmp.to_pandas().to_parquet(f'n25_bm25_log_{PART}.pqt')","metadata":{"execution":{"iopub.status.busy":"2023-01-30T05:31:48.548796Z","iopub.execute_input":"2023-01-30T05:31:48.549229Z","iopub.status.idle":"2023-01-30T05:31:48.558518Z","shell.execute_reply.started":"2023-01-30T05:31:48.54919Z","shell.execute_reply":"2023-01-30T05:31:48.557444Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# このセルではカートとオーダーの予測aidのための共起行列を作成する。\n# TFIDF値を使う\nTFDISK = 1\nSIZE = 1.86e6/TFDISK\n\n# aidをカートとオーダーに絞り込むので、ループは1回でもメモリエラーを起こさない\nfor PART in tqdm(range(TFDISK),total=TFDISK):\n    for j in tqdm(range(6),total=6):\n        a = j*CHUNK\n        b = min( (j+1)*CHUNK, len(files) )\n        for k in range(a,b,READ_CT):\n            df = [read_file(files[k])]\n            for i in range(1,READ_CT): \n                if k+i<b: \n                    df.append( read_file(files[k+i]) )\n            df = cudf.concat(df,ignore_index=True,axis=0)\n            \n            # aidをカートとオーダーのみに絞りこむ\n            df = df.loc[df['type'].isin([1,2])] \n            df = df.sort_values(['session','ts'],ascending=[True,False])\n            df = df.reset_index(drop=True)\n            df['n'] = df.groupby('session').cumcount()\n            df = df.loc[df.n<40]\n            df=df.drop('n',axis=1)\n            \n            # TF値\n            # aid_xのセッションごとの母数を出す\n            df_size=df[[\"session\",\"aid_x\"]]\n            df_size=df_size.reset_index(drop=True)\n            df_size[\"x_size\"]=df_size.groupby([\"session\",\"aid_x\"]).cumcount()+1\n            x_size=df_size.groupby([\"session\",\"aid_x\"],as_index=False).max()\n            df_size=df_size[[\"session\",\"aid_x\"]].merge(x_size,on=[\"session\",\"aid_x\"])\n            df_size=df_size.drop_duplicates()\n            df_tf=df[[\"session\",\"aid_x\",\"aid_y\"]]\n            #aid_yの出現回数を数える\n            df_tf=df_tf.reset_index(drop=True)\n            df_tf[\"y_cum\"]=df_tf.groupby([\"session\",\"aid_x\",\"aid_y\"]).cumcount()+1\n            \n            #aid_yの出現回数のmaxを出す\n            df_tf=df_tf.groupby([\"session\",\"aid_x\",\"aid_y\"],as_index=False).max()\n            \n            df_tf=df_tf.merge(df_size,on=[\"session\",\"aid_x\"])\n            df_tf[\"TF\"]=df_tf[\"y_cum\"]/df_tf[\"x_size\"]\n            df_tf=df_tf[[\"session\",\"aid_x\",\"aid_y\",\"TF\"]]\n\n            # IDF値\n            # session別で行列を作成後、aid_xをユニークにする\n            # あるaid_yがaid_xのユニークをいくつ跨いでいるか\n            df_idf=df_tf[[\"aid_x\",\"aid_y\"]].drop_duplicates()\n            df_idf[\"x_nunique\"]=df_idf[\"aid_x\"].nunique()\n            df_idf=df_idf.reset_index(drop=True)\n            df_idf[\"y_cum\"]=df_idf.groupby(\"aid_y\").cumcount()+1\n            y_cum=df_idf.groupby(\"aid_y\")[\"y_cum\"].max()\n            y_cum=y_cum.reset_index()\n            df_idf=df_idf[[\"aid_x\",\"aid_y\",\"x_nunique\"]].merge(y_cum,on=\"aid_y\")\n            df_idf=df_idf.sort_values([\"aid_x\",\"aid_y\"])\n            df_idf[\"IDF\"]=np.log(df_idf[\"x_nunique\"]/df_idf[\"y_cum\"])+1\n            df_idf=df_idf[[\"aid_x\",\"aid_y\",\"IDF\"]].reset_index(drop=True)\n\n            #TFIDF\n            df_tfidf=df_tf.merge(df_idf,on=[\"aid_x\",\"aid_y\"])\n            df_tfidf[\"TFIDF\"]=df_tfidf[\"TF\"]*df_tfidf[\"IDF\"]\n            \n            df = df.sort_values(['session','type_y'],ascending=[True,False])\n            df = df[['session', 'aid_x', 'aid_y','ts_x','type_y']].drop_duplicates(['session', 'aid_x', 'aid_y'])\n            df['wgt'] = 0\n            \n            df = df.merge(df_tfidf,on=[\"session\",\"aid_x\",\"aid_y\"])  \n            \n            # 重みにTFIDFを追加する\n            df[\"wgt\"] += df[\"TFIDF\"]\n            \n            # 直近情報の重みづけルールを掛け合わせる\n            df['wgt'] = df[\"wgt\"]*(1 + 3*(df.ts_x - 1659304800)/(1662328791-1659304800))\n            \n            df = df[['aid_x','aid_y','wgt']]\n            df.wgt = df.wgt.astype('float32')\n            df = df.groupby(['aid_x','aid_y']).wgt.sum()\n\n            if k==a: \n                tmp2 = df\n            else: \n                tmp2 = tmp2.add(df, fill_value=0)\n        if a==0: \n            tmp = tmp2\n        else: \n            tmp = tmp.add(tmp2, fill_value=0)\n        del tmp2, df\n        gc.collect()\n    tmp = tmp.reset_index()\n    tmp = tmp.sort_values(['aid_x','wgt'],ascending=[True,False])\n    tmp = tmp.reset_index(drop=True)\n    tmp['n'] = tmp.groupby('aid_x').aid_y.cumcount()\n    tmp = tmp.loc[tmp.n<15].drop('n',axis=1)\n    tmp.to_pandas().to_parquet(f'n25_tfidf_log_buy2buy_all.pqt')","metadata":{"execution":{"iopub.status.busy":"2023-01-30T05:31:48.560757Z","iopub.execute_input":"2023-01-30T05:31:48.561896Z","iopub.status.idle":"2023-01-30T05:31:48.579386Z","shell.execute_reply.started":"2023-01-30T05:31:48.561815Z","shell.execute_reply":"2023-01-30T05:31:48.578267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def load_test():    \n    dfs = []\n    for e, chunk_file in enumerate(glob.glob('../input/otto-chunk-data-inparquet-format/test_parquet/*')):\n        chunk = pd.read_parquet(chunk_file)\n        chunk.ts = (chunk.ts/1000).astype('int32')\n        chunk['type'] = chunk['type'].map(type_labels).astype('int8')\n        dfs.append(chunk)\n    return pd.concat(dfs).reset_index(drop=True)\n\ntest_df = load_test()","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[],"execution":{"iopub.status.busy":"2023-01-30T05:31:48.580966Z","iopub.execute_input":"2023-01-30T05:31:48.581587Z","iopub.status.idle":"2023-01-30T05:31:51.546894Z","shell.execute_reply.started":"2023-01-30T05:31:48.581552Z","shell.execute_reply":"2023-01-30T05:31:51.545752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def pqt_to_dict(df):\n    return df.groupby('aid_x').aid_y.apply(list).to_dict()\n\n# TFIDF x 直近情報の重みづけ\ntfidf_buy2buy20 = pqt_to_dict( pd.read_parquet(f'/kaggle/input/log-tfidf-bm25/tfidf_log_buy2buy_0.pqt') )\n\n# クリック1、カート3、オーダー6 x 直近情報の重みづけ\ntop_20_buys = pqt_to_dict( pd.read_parquet(f'/kaggle/input/distmerge/distmerge_1_3_6.pqt') )\n\n# カートと　オーダーだけに絞り込み、カート3、オーダー6の重みづけ x 直近情報の重みづけ\ntop_20_buy2buy = pqt_to_dict( pd.read_parquet(f'/kaggle/input/distmerge/dist_buy2buy_all.pqt') )\n\n# 行列分解 x type x 直近情報の重みづけ\ndist_only_buy2buy=pqt_to_dict(pd.read_parquet(\"/kaggle/input/distmerge/distonly_log_buy2buy_all.pqt\"))\n\ntop_orders=test_df.loc[test_df[\"type\"]==2][\"aid\"]\ntop_orders=top_orders.value_counts().index.values[:20]","metadata":{"execution":{"iopub.status.busy":"2023-01-30T05:31:51.548211Z","iopub.execute_input":"2023-01-30T05:31:51.548536Z","iopub.status.idle":"2023-01-30T05:34:00.411279Z","shell.execute_reply.started":"2023-01-30T05:31:51.548507Z","shell.execute_reply":"2023-01-30T05:34:00.409893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# カートとオーダーの予測aidを作成する。\ntype_weight_multipliers = {0: 1, 1: 3, 2: 6}\ndef suggest_buys(df):\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    df = df.loc[(df['type']==1)|(df['type']==2)]\n    unique_buys = list(dict.fromkeys( df.aid.tolist()[::-1] ))\n    \n    ordersdf=df.loc[(df['type']==2)]\n    unique_orders = list(dict.fromkeys( ordersdf.aid.tolist()[::-1] ))\n    \n    if len(unique_aids)>=20:\n        \n        weights=np.logspace(0.5,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n        for aid in aids3: aids_temp[aid] += 0.1\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    \n    # 全ての重みづけルールをアンサンブルする\n    aids2 = list(itertools.chain(*[top_20_buys[aid] for aid in unique_aids if aid in top_20_buys]))\n    aids3 = list(itertools.chain(*[top_20_buy2buy[aid] for aid in unique_buys if aid in top_20_buy2buy]))\n    aids4 = list(itertools.chain(*[tfidf_buy2buy20[aid] for aid in unique_aids if aid in tfidf_buy2buy20]))\n    aids5=list(itertools.chain(*[dist_only_buy2buy[aid] for aid in unique_aids if aid in dist_only_buy2buy]))\n    \n    # 最も出現したaidから上位20個を選択する。\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2+aids3+aids4+aids5).most_common(20) if aid2 not in unique_aids] \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    return result + [i for i in top_orders if i not in result][:20-len(result)]","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[],"execution":{"iopub.status.busy":"2023-01-30T05:34:00.412899Z","iopub.execute_input":"2023-01-30T05:34:00.413596Z","iopub.status.idle":"2023-01-30T05:34:00.46206Z","shell.execute_reply.started":"2023-01-30T05:34:00.41354Z","shell.execute_reply":"2023-01-30T05:34:00.460747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tqdm.pandas()\npred_df_buys = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).progress_apply(lambda x: suggest_buys(x))","metadata":{"execution":{"iopub.status.busy":"2023-01-30T05:34:00.463388Z","iopub.execute_input":"2023-01-30T05:34:00.463721Z","iopub.status.idle":"2023-01-30T06:22:06.00728Z","shell.execute_reply.started":"2023-01-30T05:34:00.463692Z","shell.execute_reply":"2023-01-30T06:22:05.964852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del top_20_buys,top_20_buy2buy","metadata":{"execution":{"iopub.status.busy":"2023-01-30T06:22:06.041771Z","iopub.execute_input":"2023-01-30T06:22:06.042217Z","iopub.status.idle":"2023-01-30T06:22:07.084265Z","shell.execute_reply.started":"2023-01-30T06:22:06.042181Z","shell.execute_reply":"2023-01-30T06:22:07.083189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def pqt_to_dict(df):\n    return df.groupby('aid_x').aid_y.apply(list).to_dict()\n\n# クリック１、カート3、オーダー6 x 直近情報の重みづけ\ntop_20_clicks = pqt_to_dict( pd.read_parquet(f'/kaggle/input/distmerge/distmerge_1_3_6.pqt') )\n\n# BM25 x 直近情報の重みづけ\nbm25 = pqt_to_dict( pd.read_parquet(f'/kaggle/input/log-tfidf-bm25/bm25_log_all.pqt') )\n\n# 行列分解 x type x 直近情報の重みづけ\ndist_only=pqt_to_dict(pd.read_parquet(\"/kaggle/input/distmerge/distonly_log_all.pqt\"))\n\ntop_clicks=test_df.loc[test_df[\"type\"]==0][\"aid\"]\ntop_clicks=top_clicks.value_counts().index.values[:20]","metadata":{"execution":{"iopub.status.busy":"2023-01-30T06:22:07.088842Z","iopub.execute_input":"2023-01-30T06:22:07.089244Z","iopub.status.idle":"2023-01-30T06:24:41.080416Z","shell.execute_reply.started":"2023-01-30T06:22:07.089211Z","shell.execute_reply":"2023-01-30T06:24:41.0791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# クリックの予測aidを作成する\ntype_weight_multipliers = {0: 1, 1: 3, 2: 6}\n\ndef suggest_clicks(df):\n    aids=df.aid.tolist()\n    types = df.type.tolist()\n    unique_aids = list(dict.fromkeys(aids[::-1] ))\n    if len(unique_aids)>=20:\n        weights=np.logspace(0.1,1,len(aids),base=2, endpoint=True)-1\n        aids_temp = Counter() \n        for aid,w,t in zip(aids,weights,types): \n            aids_temp[aid] += w * type_weight_multipliers[t]\n        sorted_aids = [k for k,v in aids_temp.most_common(20)]\n        return sorted_aids\n    \n    # 全ての重みづけルールをアンサンブルする\n    aids2 = list(itertools.chain(*[top_20_clicks[aid] for aid in unique_aids if aid in top_20_clicks]))\n    aids4 = list(itertools.chain(*[bm25[aid] for aid in unique_aids if aid in bm25]))\n    aids5=list(itertools.chain(*[dist_only[aid] for aid in unique_aids if aid in dist_only]))\n    \n    # 最も出現したaidから上位20個を選択する。\n    top_aids2 = [aid2 for aid2, cnt in Counter(aids2+aids4+aids5).most_common(20) if aid2 not in unique_aids] \n    result = unique_aids + top_aids2[:20 - len(unique_aids)]\n    return result + [i for i in top_clicks if i not in result][:20-len(result)]","metadata":{"execution":{"iopub.status.busy":"2023-01-30T06:24:41.081953Z","iopub.execute_input":"2023-01-30T06:24:41.082391Z","iopub.status.idle":"2023-01-30T06:24:41.147419Z","shell.execute_reply.started":"2023-01-30T06:24:41.082359Z","shell.execute_reply":"2023-01-30T06:24:41.146339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tqdm.pandas()\npred_df_clicks = test_df.sort_values([\"session\", \"ts\"]).groupby([\"session\"]).progress_apply(lambda x: suggest_clicks(x))","metadata":{"execution":{"iopub.status.busy":"2023-01-30T06:24:41.148682Z","iopub.execute_input":"2023-01-30T06:24:41.14961Z","iopub.status.idle":"2023-01-30T06:32:39.893191Z","shell.execute_reply.started":"2023-01-30T06:24:41.149567Z","shell.execute_reply":"2023-01-30T06:32:39.892005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del top_20_clicks","metadata":{"execution":{"iopub.status.busy":"2023-01-30T06:32:39.894609Z","iopub.execute_input":"2023-01-30T06:32:39.895001Z","iopub.status.idle":"2023-01-30T06:32:40.670637Z","shell.execute_reply.started":"2023-01-30T06:32:39.894959Z","shell.execute_reply":"2023-01-30T06:32:40.669361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clicks_pred_df = pd.DataFrame(pred_df_clicks.add_suffix(\"_clicks\"), columns=[\"labels\"]).reset_index()\ncarts_pred_df = pd.DataFrame(pred_df_buys.add_suffix(\"_carts\"), columns=[\"labels\"]).reset_index()\norders_pred_df = pd.DataFrame(pred_df_buys.add_suffix(\"_orders\"), columns=[\"labels\"]).reset_index()","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[],"execution":{"iopub.status.busy":"2023-01-30T06:32:40.672753Z","iopub.execute_input":"2023-01-30T06:32:40.673231Z","iopub.status.idle":"2023-01-30T06:32:43.864414Z","shell.execute_reply.started":"2023-01-30T06:32:40.673187Z","shell.execute_reply":"2023-01-30T06:32:43.863124Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"carts_pred_df","metadata":{"execution":{"iopub.status.busy":"2023-01-30T06:32:43.866064Z","iopub.execute_input":"2023-01-30T06:32:43.866378Z","iopub.status.idle":"2023-01-30T06:32:43.891703Z","shell.execute_reply.started":"2023-01-30T06:32:43.866349Z","shell.execute_reply":"2023-01-30T06:32:43.890419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"orders_pred_df","metadata":{"execution":{"iopub.status.busy":"2023-01-30T06:32:43.894247Z","iopub.execute_input":"2023-01-30T06:32:43.895083Z","iopub.status.idle":"2023-01-30T06:32:43.913786Z","shell.execute_reply.started":"2023-01-30T06:32:43.895034Z","shell.execute_reply":"2023-01-30T06:32:43.912572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clicks_pred_df","metadata":{"execution":{"iopub.status.busy":"2023-01-30T06:32:43.915562Z","iopub.execute_input":"2023-01-30T06:32:43.915988Z","iopub.status.idle":"2023-01-30T06:32:43.937527Z","shell.execute_reply.started":"2023-01-30T06:32:43.915946Z","shell.execute_reply":"2023-01-30T06:32:43.936216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 予測aidを全て繋ぎ合わせて提出データにする\npred_df = pd.concat([clicks_pred_df, carts_pred_df,orders_pred_df])\npred_df.columns = [\"session_type\", \"labels\"]\npred_df[\"labels\"] = pred_df.labels.progress_apply(lambda x: \" \".join(map(str,x)))\npred_df.to_csv(\"submit_clicks_log_type.csv\", index=False)\npred_df.head(),pred_df.shape","metadata":{"papermill":{"duration":null,"end_time":null,"exception":null,"start_time":null,"status":"pending"},"tags":[],"execution":{"iopub.status.busy":"2023-01-30T06:32:43.941087Z","iopub.execute_input":"2023-01-30T06:32:43.941444Z","iopub.status.idle":"2023-01-30T06:33:46.470134Z","shell.execute_reply.started":"2023-01-30T06:32:43.941412Z","shell.execute_reply":"2023-01-30T06:33:46.468877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}