{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Prediction Notebook\n\n## This notebook is a work in progress. I'll update it when I'll get some free time. If you are intereseted in something in particular leave a comment and I'll try to answer as soon as possible.\n\n\nIn this notebook I use the models trained to make inference on the test set.\n\nMy last submission was made of 3 LGBM classifier one for each interaction type.\n\nThe prediction on the test set would take too much memory so the prediction is done in 10 folds.\n\nThe LGBM classifiers combine the score of many models and features to make predictions.\n\nThe model I use uses candidates from 3 groups of models:\nThe models used in the first group are the 3 candidate selectors from an improved version of Co-visitation matrix starting from this notebook  [notebook](https://www.kaggle.com/code/carnozhao/otto-fast-cpu-end-to-end-pipeline) by [carnozhao](https://www.kaggle.com/carnozhao), a tuned version of Matrix Factorization from the [notebook](https://www.kaggle.com/code/cpmpml/matrix-factorization-with-gpu) by [cpmpml](https://www.kaggle.com/cpmpml), a slightly changed version of W2V model form the [notebook](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission) by [radek1](https://www.kaggle.com/radek1). In addition to these models I tried also to reverse the order of each session for Co-visitation.\nThe second group uses a set of candidate selector based purely on the history of each session, to each item we give a weight and we sum the weight for each item in the session, I used different weighting for each type of interaction but also a weighting based on the temporal distance from the last seen item in the session.\nThe last group in the end was only a single model using RecBole, in particular I used SRGNN model using the code from this [notebook](https://www.kaggle.com/code/yamsam/recbole-gru4rec-sample-code) by [yamsam](https://www.kaggle.com/yamsam) as a starting point.\n\nThe features used are fairly easy ones, for each candidate how many times it already appeared in the session for each type of interaction and the sum of this 3 counts, the length of the session (number of items) and the temporal length of the session (just now I'm thinking that I forgot to also add the count of unique items seen in the session).\n\nOther features I used are Target Encodings trained on the week before the test set and target encoding trained on the test week ( for training the models I used the Target encodings calculated on the week before the validation week and encoding calculated on the validation week itself).\n\n# Everything has been trained on kaggle!!!\nI have over 50 Notebooks combining model validation, model training and inference. \n\nI struggled a lot for some parts, an example is SRGNN: Inference would take more than 24h on CPU and 10h on GPU. My solution was training on GPU (2h) saving the model and making predictions on 7 different notebooks each for a split of the dataset, in this way the time to get the result was around 4-5h and the GPU time used was minimal.\n\nI'll try polishing as much as I can and share to the community as many notebook as possible.\n\nThis was a hard but fun competition, I hope everyone enjoyed it. \n\nLet's keep kaggling!","metadata":{}},{"cell_type":"code","source":"splits=10\nDEBUG = False","metadata":{"execution":{"iopub.status.busy":"2023-01-27T07:56:20.133437Z","iopub.execute_input":"2023-01-27T07:56:20.134711Z","iopub.status.idle":"2023-01-27T07:56:36.048679Z","shell.execute_reply.started":"2023-01-27T07:56:20.134575Z","shell.execute_reply":"2023-01-27T07:56:36.047134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd \n\nfrom functools import reduce\nimport os\nimport gc\nfrom sklearn.model_selection import train_test_split\nimport lightgbm\nimport logging\nimport sys\n\n\nfrom tqdm.auto import tqdm\n\nFORMAT = '%(asctime)s - %(levelname)s - %(message)s'\n\n\ntry:\n    logger\nexcept:\n    logger=logging.getLogger(\"experiment_logger\")\n    formatter = logging.Formatter(FORMAT)\n    consoleHandler = logging.StreamHandler(sys.stdout)\n    consoleHandler.setFormatter(formatter)\n    logger.addHandler(consoleHandler)\n    logger.setLevel(logging.DEBUG)\n    logger.debug(\"debugging\")","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-01-27T07:56:36.050871Z","iopub.execute_input":"2023-01-27T07:56:36.051777Z","iopub.status.idle":"2023-01-27T07:56:37.794159Z","shell.execute_reply.started":"2023-01-27T07:56:36.051733Z","shell.execute_reply":"2023-01-27T07:56:37.792792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_sessions_all=pd.read_parquet(\"/kaggle/input/otto-full-optimized-memory-footprint/test.parquet\")\nsessions_full=test_sessions_all[\"session\"].drop_duplicates().reset_index(drop=True)\nif DEBUG:\n    sessions_full=sessions_full[-10:]","metadata":{"execution":{"iopub.status.busy":"2023-01-27T07:56:37.795867Z","iopub.execute_input":"2023-01-27T07:56:37.796235Z","iopub.status.idle":"2023-01-27T07:56:38.886232Z","shell.execute_reply.started":"2023-01-27T07:56:37.796201Z","shell.execute_reply":"2023-01-27T07:56:38.885061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Combine score of candidates for each candidate selector with various features\ndef prepare_split(split,splits):\n    assert split<splits\n    \n\n    sessions=sessions_full[(len(sessions_full)*split)//splits:(len(sessions_full)*(split+1))//splits]\n\n    test_sessions=test_sessions_all.query(\"session in @sessions\")\n    count_click=test_sessions.loc[test_sessions[\"type\"]==\"clicks\",[\"session\",\"aid\",\"type\"]].groupby([\"session\",\"aid\"]).agg(count_click=(\"type\",\"count\")).reset_index()\n    count_cart=test_sessions.loc[test_sessions[\"type\"]==\"carts\",[\"session\",\"aid\",\"type\"]].groupby([\"session\",\"aid\"]).agg(count_cart=(\"type\",\"count\")).reset_index()\n    count_order=test_sessions.loc[test_sessions[\"type\"]==\"orders\",[\"session\",\"aid\",\"type\"]].groupby([\"session\",\"aid\"]).agg(count_order=(\"type\",\"count\")).reset_index()\n    count_overall=test_sessions.groupby([\"session\",\"aid\"]).agg(count_overall=(\"type\",\"count\")).reset_index()\n\n\n    length=test_sessions.groupby([\"session\"]).agg(length=(\"type\",\"count\")).reset_index()\n    duration=test_sessions.groupby([\"session\"]).agg(max_time=(\"ts\",\"max\"),min_time=(\"ts\",\"min\")).reset_index()\n    duration[\"duration\"]=duration[\"max_time\"]-duration[\"min_time\"]\n    del duration[\"max_time\"]\n    del duration[\"min_time\"]\n    session_aid_feats=count_click.merge(count_cart,on=[\"session\",\"aid\"],how=\"left\")\n    del count_cart,count_click\n    gc.collect()\n    session_aid_feats=session_aid_feats.merge(count_order,on=[\"session\",\"aid\"],how=\"left\")\n    del count_order\n    gc.collect()\n    session_aid_feats=session_aid_feats.merge(count_overall,on=[\"session\",\"aid\"],how=\"left\")\n    del count_overall\n    gc.collect()\n\n    session_feats=length.merge(duration,on=\"session\",how=\"left\")\n    del duration,length\n    gc.collect()\n    \n\n    dfs=[]\n    for file in [file for file in os.listdir(\"/kaggle/input/test-convert-format\") if \"aid\" not in file and \".feather\"  in file]:\n        df = pd.read_feather(f\"/kaggle/input/test-convert-format/{file}\")\n        df=df.query(\"session in @sessions\")\n        score_col=[col for col in list(df.columns) if \"score\" in col][0]\n        #df[score_col]=df[score_col].astype(\"float16\")\n        dfs.append(df)\n\n    dfs1 = reduce(lambda  left,right: pd.merge(left,right,on=['aid',\"session\"],\n                                                how='outer'),dfs[1:],dfs[0])\n    dfs=[]\n    for file in [file for file in os.listdir(\"/kaggle/input/test-convert-format-part-2\") if \"aid\" not in file and \".feather\"  in file]:\n        df = pd.read_feather(f\"/kaggle/input/test-convert-format-part-2/{file}\")\n        df=df.query(\"session in @sessions\")\n        score_col=[col for col in list(df.columns) if \"score\" in col][0]\n        #df[score_col]=df[score_col].astype(\"float16\")\n        dfs.append(df)\n\n    dfs2 = reduce(lambda  left,right: pd.merge(left,right,on=['aid',\"session\"],\n                                                how='outer'), dfs[1:],dfs[0])\n    dfs=[]\n    for file in [file for file in os.listdir(\"/kaggle/input/convert-format-test-recbole\") if \"aid\" not in file and \".feather\"  in file]:\n        df = pd.read_feather(f\"/kaggle/input/convert-format-test-recbole/{file}\")\n        df=df.query(\"session in @sessions\")\n        score_col=[col for col in list(df.columns) if \"score\" in col][0]\n        #df[score_col]=df[score_col].astype(\"float16\")\n        dfs.append(df)\n        \n    logger.debug(\"loaded scores\")\n    #dfs3 = reduce(lambda  left,right: pd.merge(left,right,on=['aid',\"session\"],\n    #                                            how='outer'), tqdm(dfs[1:]),dfs[0])\n    dfs3 = dfs[0]\n\n    df = dfs1.merge(dfs3,on=[\"aid\",\"session\"],how=\"outer\")\n    df = df.merge(dfs2,on=[\"aid\",\"session\"],how=\"outer\")\n\n    #df = dfs1.merge(dfs3,on=[\"aid\",\"session\"],how=\"outer\")\n    del dfs,dfs1,dfs2,dfs3\n\n    gc.collect()\n    df.fillna(-1,inplace=True)\n    \n    TE=pd.read_feather(\"/kaggle/input/test-te-preparation/TE_valid.feather\")\n    TE=TE.fillna(-1)\n    \n    df=df.merge(session_aid_feats,on=[\"session\",\"aid\"],how=\"left\")\n    df=df.merge(session_feats,on=\"session\",how=\"left\")\n    df=df.merge(TE,on=\"aid\",how=\"left\")\n    df.fillna(0,inplace=True)\n    df.head()\n\n    del TE,session_aid_feats,session_feats\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2023-01-27T07:56:38.888647Z","iopub.execute_input":"2023-01-27T07:56:38.889009Z","iopub.status.idle":"2023-01-27T07:56:38.914303Z","shell.execute_reply.started":"2023-01-27T07:56:38.888977Z","shell.execute_reply":"2023-01-27T07:56:38.913132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_predictions(split):\n    #click\n    features=np.load(\"/kaggle/input/train-lgbm-classifier-click/features.npy\")\n    model_click= lightgbm.Booster(model_file='/kaggle/input/train-lgbm-classifier-click/model_click_classifier.txt')\n    split[\"click_pred\"]=model_click.predict(split[list(features)])\n    #cart\n    features=np.load(\"/kaggle/input/train-lgbm-classifier-cart/features.npy\")\n    model_cart= lightgbm.Booster(model_file='/kaggle/input/train-lgbm-classifier-cart/model_cart_classifier.txt')\n    split[\"cart_pred\"]=model_cart.predict(split[list(features)])\n    #order\n    features=np.load(\"/kaggle/input/train-lgbm-classifier-orders/features.npy\")\n    model_order= lightgbm.Booster(model_file='/kaggle/input/train-lgbm-classifier-orders/model_order_classifier.txt')\n    split[\"order_pred\"]=model_order.predict(split[list(features)])","metadata":{"execution":{"iopub.status.busy":"2023-01-27T07:56:38.915855Z","iopub.execute_input":"2023-01-27T07:56:38.916335Z","iopub.status.idle":"2023-01-27T07:56:38.932176Z","shell.execute_reply.started":"2023-01-27T07:56:38.916277Z","shell.execute_reply":"2023-01-27T07:56:38.930889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def score_to_csv(df,name,n):\n    df[\"labels\"]=df[\"labels\"].astype(\"int32\").astype(\"string\")\n    df=df.sort_values(['session_type', 'score'], ascending=False).groupby('session_type')[\"labels\"].agg(list).reset_index()\n    df[\"labels\"]=df[\"labels\"].str[:20].str.join(\" \")\n    df[\"session_type\"]=df[\"session_type\"].astype(\"int32\").astype(str)+f\"_{name}\"\n    df.to_csv(f\"pred_{name}_part_{n}.csv\",index=False)\n    print(df.head(1))\n    \ndef preds_to_file(df,n):\n    \n    \n    score_to_csv(df[[\"session\",\"aid\",\"click_pred\"]].rename(columns={\"click_pred\":\"score\",\"session\":\"session_type\",\"aid\":\"labels\"}),\"clicks\",n)\n    score_to_csv(df[[\"session\",\"aid\",\"cart_pred\"]].rename(columns={\"cart_pred\":\"score\",\"session\":\"session_type\",\"aid\":\"labels\"}),\"carts\",n)\n    score_to_csv(df[[\"session\",\"aid\",\"order_pred\"]].rename(columns={\"order_pred\":\"score\",\"session\":\"session_type\",\"aid\":\"labels\"}),\"orders\",n)","metadata":{"execution":{"iopub.status.busy":"2023-01-27T07:56:38.933785Z","iopub.execute_input":"2023-01-27T07:56:38.934212Z","iopub.status.idle":"2023-01-27T07:56:38.951809Z","shell.execute_reply.started":"2023-01-27T07:56:38.934171Z","shell.execute_reply":"2023-01-27T07:56:38.950676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nfor split in tqdm(range(splits)):\n    df=prepare_split(split,splits)\n    logger.debug(\"data ready\")\n    get_predictions(df)\n    logger.debug(\"done predictions\")\n    df=df[[\"session\",\"aid\",\"click_pred\",\"cart_pred\",\"order_pred\"]]\n    gc.collect()\n    preds_to_file(df,split)\n    logger.debug(\"prepared submission partial file\")\n    del df\n    gc.collect()\n    ","metadata":{"execution":{"iopub.status.busy":"2023-01-27T07:56:38.953715Z","iopub.execute_input":"2023-01-27T07:56:38.954081Z","iopub.status.idle":"2023-01-27T07:58:04.864000Z","shell.execute_reply.started":"2023-01-27T07:56:38.954048Z","shell.execute_reply":"2023-01-27T07:58:04.859152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df=pd.concat([pd.read_csv(f\"pred_clicks_part_{split}.csv\") for  split in range(splits)]\n             +[pd.read_csv(f\"pred_carts_part_{split}.csv\") for  split in range(splits)]\n             +[pd.read_csv(f\"pred_orders_part_{split}.csv\") for  split in range(splits)])\n#df[\"session_type\"]=df[\"session_type\"].str.replace(\".0\",\"\")\ndf.to_csv(\"submission.csv\",index=False)","metadata":{"execution":{"iopub.status.busy":"2023-01-27T07:58:04.864900Z","iopub.status.idle":"2023-01-27T07:58:04.865359Z","shell.execute_reply.started":"2023-01-27T07:58:04.865116Z","shell.execute_reply":"2023-01-27T07:58:04.865134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-27T07:58:04.869545Z","iopub.status.idle":"2023-01-27T07:58:04.870056Z","shell.execute_reply.started":"2023-01-27T07:58:04.869840Z","shell.execute_reply":"2023-01-27T07:58:04.869863Z"},"trusted":true},"execution_count":null,"outputs":[]}]}