{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"***OTTO – Multi-Objective Recommender System***","metadata":{}},{"cell_type":"markdown","source":"_`Task overview`_\n\n_The aim of this competition is to predict e-commerce clicks, cart additions, and orders. You'll build a multi-objective recommender system based on previous events in a user session._\n\n_Current recommender systems consist of various models with different approaches, ranging from simple matrix factorization to a transformer-type deep neural network. However, no single model exists that can simultaneously optimize multiple objectives. In this competition, you’ll build a single entry to predict click-through, add-to-cart, and conversion rates based on previous same-session events._","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-11-16T03:55:36.488582Z","iopub.execute_input":"2022-11-16T03:55:36.488995Z","iopub.status.idle":"2022-11-16T03:55:36.517614Z","shell.execute_reply.started":"2022-11-16T03:55:36.488905Z","shell.execute_reply":"2022-11-16T03:55:36.516748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nfrom pathlib import Path\nimport os\nimport random\nimport numpy as np\nimport json\nfrom datetime import timedelta\nfrom collections import Counter\nfrom tqdm.notebook import tqdm\nfrom heapq import nlargest\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nsns.set_theme()\n\nimport warnings\nwarnings.filterwarnings('ignore')\n\nimport math","metadata":{"execution":{"iopub.status.busy":"2022-11-16T03:55:43.513315Z","iopub.execute_input":"2022-11-16T03:55:43.513706Z","iopub.status.idle":"2022-11-16T03:55:44.099600Z","shell.execute_reply.started":"2022-11-16T03:55:43.513671Z","shell.execute_reply":"2022-11-16T03:55:44.098729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Data_Path = Path('../input/otto-recommender-system')\nTrain_Path = Data_Path/'train.jsonl'\nTest_Path = Data_Path/'test.jsonl'\nSample_sub_Path = Path('../input/otto-recommender-system/sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-11-16T03:55:52.682855Z","iopub.execute_input":"2022-11-16T03:55:52.683687Z","iopub.status.idle":"2022-11-16T03:55:52.688844Z","shell.execute_reply.started":"2022-11-16T03:55:52.683641Z","shell.execute_reply":"2022-11-16T03:55:52.687822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"with open(Train_Path, 'r') as f:\n    print(f\"We have {len(f.readlines()):,} lines in the training data\")","metadata":{"execution":{"iopub.status.busy":"2022-11-16T03:58:21.865468Z","iopub.execute_input":"2022-11-16T03:58:21.865858Z","iopub.status.idle":"2022-11-16T03:59:17.232770Z","shell.execute_reply.started":"2022-11-16T03:58:21.865818Z","shell.execute_reply":"2022-11-16T03:59:17.231624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_size = 150000\n\nchunks = pd.read_json(Train_Path, lines=True, chunksize = sample_size)\n\nfor c in chunks:\n    sample_train_df = c\n    break","metadata":{"execution":{"iopub.status.busy":"2022-11-16T03:59:40.863360Z","iopub.execute_input":"2022-11-16T03:59:40.864540Z","iopub.status.idle":"2022-11-16T03:59:51.013329Z","shell.execute_reply.started":"2022-11-16T03:59:40.864492Z","shell.execute_reply":"2022-11-16T03:59:51.012417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_train_df.set_index('session', drop=True, inplace=True)\nsample_train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-11-16T04:00:28.019175Z","iopub.execute_input":"2022-11-16T04:00:28.020017Z","iopub.status.idle":"2022-11-16T04:00:28.125886Z","shell.execute_reply.started":"2022-11-16T04:00:28.019966Z","shell.execute_reply":"2022-11-16T04:00:28.125044Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Data structure***\n\n_`session` - the unique session id. Each session contains a list of time ordered events._\n\n_`events` - the time ordered sequence of events in the session. Each event contains 3 pieces of information:_\n\n- _`aid` - the article id (product code) of the associated event_\n\n- _`ts` - the Unix timestamp of the event (Unix time is the number of microseconds that have elapsed since 00:00:00 UTC on 1 January 1970)_\n\n- _`type` - the event type, i.e., whether a product was clicked (`clicks`), added to the user's cart (`carts`), or ordered during the session (`orders`)_","metadata":{}},{"cell_type":"code","source":"example_session = sample_train_df.iloc[0].item()\nprint(f'This session was {len(example_session)} actions long \\n')\nprint(f'The first action in the session: \\n {example_session[0]} \\n')\n\ntime_elapsed = example_session[-1][\"ts\"] - example_session[0][\"ts\"]\nprint(f'The first session elapsed: {str(timedelta(microseconds=time_elapsed))} \\n')\n\naction_counts = {}\nfor action in example_session:\n    action_counts[action['type']] = action_counts.get(action['type'], 0) + 1  \nprint(f'The first session contains the following frequency of actions: {action_counts}')","metadata":{"execution":{"iopub.status.busy":"2022-11-16T04:01:08.783918Z","iopub.execute_input":"2022-11-16T04:01:08.784365Z","iopub.status.idle":"2022-11-16T04:01:08.793351Z","shell.execute_reply.started":"2022-11-16T04:01:08.784332Z","shell.execute_reply":"2022-11-16T04:01:08.792068Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"_`Intital EDA`_","metadata":{}},{"cell_type":"code","source":"action_counts_list, article_id_counts_list, session_length_time_list, session_length_action_list = ([] for i in range(4))\noverall_action_counts = {}\noverall_article_id_counts = {}\n\nfor i, row in tqdm(sample_train_df.iterrows(), total=len(sample_train_df)):\n    \n    actions = row['events']\n    \n    action_counts = {}\n    article_id_counts = {}\n    for action in actions:\n        action_counts[action['type']] = action_counts.get(action['type'], 0) + 1\n        article_id_counts[action['aid']] = article_id_counts.get(action['aid'], 0) + 1\n        overall_action_counts[action['type']] = overall_action_counts.get(action['type'], 0) + 1\n        overall_article_id_counts[action['aid']] = overall_article_id_counts.get(action['aid'], 0) + 1\n        \n    session_length_time = actions[-1]['ts'] - actions[0]['ts']\n    \n    action_counts_list.append(action_counts)\n    article_id_counts_list.append(article_id_counts)\n    session_length_time_list.append(session_length_time)\n    session_length_action_list.append(len(actions))\n    \nsample_train_df['action_counts'] = action_counts_list\nsample_train_df['article_id_counts'] = article_id_counts_list\nsample_train_df['session_length_unix'] = session_length_time_list\nsample_train_df['session_length_minutes'] = sample_train_df['session_length_unix']*1.66667e-8 \nsample_train_df['session_length_action'] = session_length_action_list","metadata":{"execution":{"iopub.status.busy":"2022-11-16T04:01:30.774955Z","iopub.execute_input":"2022-11-16T04:01:30.775898Z","iopub.status.idle":"2022-11-16T04:01:55.943769Z","shell.execute_reply.started":"2022-11-16T04:01:30.775860Z","shell.execute_reply":"2022-11-16T04:01:55.942794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"total_actions = sum(overall_action_counts.values())\n\nplt.figure(figsize=(8,6))\nsns.barplot(x=list(overall_action_counts.keys()), y=[i/total_actions for i in overall_action_counts.values()]);\nplt.title(f'Action frequency', fontsize=12)\nplt.ylabel('Count', fontsize=12)\nplt.xlabel('Category', fontsize=12)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-11-16T04:02:02.528733Z","iopub.execute_input":"2022-11-16T04:02:02.529145Z","iopub.status.idle":"2022-11-16T04:02:02.745166Z","shell.execute_reply.started":"2022-11-16T04:02:02.529108Z","shell.execute_reply":"2022-11-16T04:02:02.744353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1,2, figsize=(24, 10))\n\np = sns.distplot(sample_train_df['session_length_action'], color=\"y\", bins= 70, ax=ax[0], kde=False)\np.set_xlabel(\"Number of actions\", fontsize = 16)\np.set_ylabel(\"Density\", fontsize = 16)\np.set_title(\"Distribution of the number of actions taken in each session\", fontsize = 14)\np.axvline(sample_train_df['session_length_action'].mean(), color='r', linestyle='--', label=\"Mean\")\n\np = sns.distplot(sample_train_df['session_length_minutes'], color=\"b\", bins= 70, ax=ax[1], kde=False)\np.set_xlabel(\"Minutes\", fontsize = 16)\np.set_ylabel(\"Density\", fontsize = 16)\np.set_title(\"Length of each session\", fontsize = 16);","metadata":{"execution":{"iopub.status.busy":"2022-11-16T04:02:11.183563Z","iopub.execute_input":"2022-11-16T04:02:11.184798Z","iopub.status.idle":"2022-11-16T04:02:11.954081Z","shell.execute_reply.started":"2022-11-16T04:02:11.184717Z","shell.execute_reply":"2022-11-16T04:02:11.953026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'{round(len(sample_train_df[sample_train_df[\"session_length_action\"]<10])/len(sample_train_df),3)*100}% of the sessions had less than 10 actions')","metadata":{"execution":{"iopub.status.busy":"2022-11-16T04:02:21.708306Z","iopub.execute_input":"2022-11-16T04:02:21.708730Z","iopub.status.idle":"2022-11-16T04:02:21.775119Z","shell.execute_reply.started":"2022-11-16T04:02:21.708700Z","shell.execute_reply":"2022-11-16T04:02:21.773901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"article_id_freq = list(overall_article_id_counts.values())\ncut_off = [i for i in article_id_freq if i<30]\n\nplt.figure(figsize=(8,6))\nsns.distplot(cut_off, bins=30, kde=False);\nplt.title(f'Article ID frequency', fontsize=12)\nplt.ylabel('Count', fontsize=12)\nplt.xlabel('Article', fontsize=12);","metadata":{"execution":{"iopub.status.busy":"2022-11-16T04:02:32.719853Z","iopub.execute_input":"2022-11-16T04:02:32.720306Z","iopub.status.idle":"2022-11-16T04:02:33.152214Z","shell.execute_reply.started":"2022-11-16T04:02:32.720270Z","shell.execute_reply":"2022-11-16T04:02:33.151011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'Frequency of most common articles: {sorted(list(overall_article_id_counts.values()))[-5:]} \\n')\nres = nlargest(5, overall_article_id_counts, key = overall_article_id_counts.get)\nprint(f'IDs for those common articles: {res}')","metadata":{"execution":{"iopub.status.busy":"2022-11-16T04:02:40.394642Z","iopub.execute_input":"2022-11-16T04:02:40.395072Z","iopub.status.idle":"2022-11-16T04:02:40.766201Z","shell.execute_reply.started":"2022-11-16T04:02:40.395035Z","shell.execute_reply":"2022-11-16T04:02:40.765100Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"_`Baseline`_\n\n_The test data contains truncated session data similar to that of the training data. The task is to predict the next aid clicked after the session truncation, as well as the the remaining aids that are added to carts and orders; you may predict up to 20 values for each session type_\n\n_Submissions are evaluated on Recall each action type, and the three recall values are weight-averaged: {'clicks': 0.10, 'carts': 0.30, 'orders': 0.60} . It is important to get the 'orders' predictions correct as they carry most of the weigthing_\n\n_For each session in the test data, your task it to predict the aid values for each type that occur after the last timestamp ts the test session. In other words, the test data contains sessions truncated by timestamp, and you are to predict what occurs after the point of truncation._\n\n_For clicks there is only a single ground truth value for each session, which is the next aid clicked during the session (although you can still predict up to 20 aid values). The ground truth for carts and orders contains all aid values that were added to a cart and ordered respectively during the session._\n\n_Each session and type combination should appear on its own session_type row in the submission (3 rows per session), and predictions should be space delimited. This can be seen in the sample_test_df below._","metadata":{}},{"cell_type":"code","source":"with open(Test_Path, 'r') as f:\n    print(f\"We have {len(f.readlines()):,} lines in the test data\")","metadata":{"execution":{"iopub.status.busy":"2022-11-16T04:03:20.323287Z","iopub.execute_input":"2022-11-16T04:03:20.324387Z","iopub.status.idle":"2022-11-16T04:03:25.314314Z","shell.execute_reply.started":"2022-11-16T04:03:20.324310Z","shell.execute_reply":"2022-11-16T04:03:25.313131Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_size = 150\n\nchunks = pd.read_json(Test_Path, lines=True, chunksize = sample_size)\n\nfor c in chunks:\n    sample_test_df = c\n    break","metadata":{"execution":{"iopub.status.busy":"2022-11-16T04:03:39.288531Z","iopub.execute_input":"2022-11-16T04:03:39.288921Z","iopub.status.idle":"2022-11-16T04:03:39.304498Z","shell.execute_reply.started":"2022-11-16T04:03:39.288889Z","shell.execute_reply":"2022-11-16T04:03:39.303370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_test_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-11-16T04:03:48.093831Z","iopub.execute_input":"2022-11-16T04:03:48.094730Z","iopub.status.idle":"2022-11-16T04:03:48.136578Z","shell.execute_reply.started":"2022-11-16T04:03:48.094687Z","shell.execute_reply":"2022-11-16T04:03:48.135689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"_Below shows a sample submission. For each session in the test set there is a prediction (labels). This predicts what articles will be next interacted with in that session. For each session there are three actions (clicks, carts, orders), predictions are made for all three actions._","metadata":{}},{"cell_type":"code","source":"sample_submission = pd.read_csv(Sample_sub_Path)\nsample_submission.head()","metadata":{"execution":{"iopub.status.busy":"2022-11-16T04:04:05.673432Z","iopub.execute_input":"2022-11-16T04:04:05.673873Z","iopub.status.idle":"2022-11-16T04:04:11.390173Z","shell.execute_reply.started":"2022-11-16T04:04:05.673834Z","shell.execute_reply":"2022-11-16T04:04:11.389076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"_Lets find the most common article for each different type of action._","metadata":{}},{"cell_type":"code","source":"sample_size = 150000\n\nchunks = pd.read_json(Train_Path, lines=True, chunksize = sample_size)\n\nclicks_article_list = []\ncarts_article_list = []\norders_article_list = []\n\nfor e, c in enumerate(chunks):\n    \n    if e > 2:\n        break\n    \n    sample_train_df = c\n    \n    for i, row in c.iterrows():\n        actions = row['events']\n        for action in actions:\n            if action['type'] == 'clicks':\n                clicks_article_list.append(action['aid'])\n            elif action['type'] == 'carts':\n                carts_article_list.append(action['aid'])\n            else:\n                orders_article_list.append(action['aid'])","metadata":{"execution":{"iopub.status.busy":"2022-11-16T04:04:47.809067Z","iopub.execute_input":"2022-11-16T04:04:47.810240Z","iopub.status.idle":"2022-11-16T04:06:08.355992Z","shell.execute_reply.started":"2022-11-16T04:04:47.810195Z","shell.execute_reply":"2022-11-16T04:06:08.355063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"article_click_freq = Counter(clicks_article_list)\narticle_carts_freq = Counter(carts_article_list)\narticle_order_freq = Counter(orders_article_list)","metadata":{"execution":{"iopub.status.busy":"2022-11-16T04:06:08.358256Z","iopub.execute_input":"2022-11-16T04:06:08.359193Z","iopub.status.idle":"2022-11-16T04:06:16.152765Z","shell.execute_reply.started":"2022-11-16T04:06:08.359136Z","shell.execute_reply":"2022-11-16T04:06:16.151463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"top_click_article = nlargest(20, article_click_freq, key = article_click_freq.get)\ntop_carts_article = nlargest(20, article_carts_freq, key = article_carts_freq.get)\ntop_order_article = nlargest(20, article_order_freq, key = article_order_freq.get) ","metadata":{"execution":{"iopub.status.busy":"2022-11-16T04:06:33.003961Z","iopub.execute_input":"2022-11-16T04:06:33.005193Z","iopub.status.idle":"2022-11-16T04:06:33.674738Z","shell.execute_reply.started":"2022-11-16T04:06:33.005151Z","shell.execute_reply":"2022-11-16T04:06:33.673534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"frequent_articles = {'clicks': top_click_article, 'carts':top_carts_article, 'order':top_order_article}","metadata":{"execution":{"iopub.status.busy":"2022-11-16T04:06:53.647864Z","iopub.execute_input":"2022-11-16T04:06:53.648302Z","iopub.status.idle":"2022-11-16T04:06:53.653473Z","shell.execute_reply.started":"2022-11-16T04:06:53.648267Z","shell.execute_reply":"2022-11-16T04:06:53.652220Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for action in ['clicks', 'carts', 'order']:\n    print(f'Most frequent articles for {action}: {frequent_articles[action][-5:]}')","metadata":{"execution":{"iopub.status.busy":"2022-11-16T04:07:03.303430Z","iopub.execute_input":"2022-11-16T04:07:03.303803Z","iopub.status.idle":"2022-11-16T04:07:03.309771Z","shell.execute_reply.started":"2022-11-16T04:07:03.303764Z","shell.execute_reply":"2022-11-16T04:07:03.308500Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"_There is some overlap but the articles do change for the different actions!_\n\n_This baseline will use the fact that people will often interact with articles they have previouslt interacted with. The prediction will consist of the top 20 most frequent articles in the session. If there are less than 20 articles in the session the prediction will be padded with the most frequent articles in the training data as found above._","metadata":{}},{"cell_type":"code","source":"test_data = pd.read_json(Test_Path, lines=True, chunksize=1000)\n\npreds = []\n\nfor chunk in tqdm(test_data, total=1671):\n    \n    for i, row in chunk.iterrows():\n        actions = row['events']\n        article_id_list = []\n        for action in actions:\n            article_id_list.append(action['aid'])\n            \n        article_freq = Counter(article_id_list)\n        top_articles = nlargest(20, article_freq, key = article_freq.get)\n        \n        padding_size = -(20 - len(top_articles))\n        for action in ['clicks', 'carts', 'order']:\n            top_articles = top_articles + frequent_articles[action][padding_size:]\n            preds.append(\" \".join([str(id) for id in top_articles]))","metadata":{"execution":{"iopub.status.busy":"2022-11-16T04:07:46.638704Z","iopub.execute_input":"2022-11-16T04:07:46.639139Z","iopub.status.idle":"2022-11-16T04:11:44.995200Z","shell.execute_reply.started":"2022-11-16T04:07:46.639102Z","shell.execute_reply":"2022-11-16T04:11:44.994070Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submission['labels'] = preds","metadata":{"execution":{"iopub.status.busy":"2022-11-16T04:11:53.102977Z","iopub.execute_input":"2022-11-16T04:11:53.103817Z","iopub.status.idle":"2022-11-16T04:11:53.914891Z","shell.execute_reply.started":"2022-11-16T04:11:53.103777Z","shell.execute_reply":"2022-11-16T04:11:53.913934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sample_submission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-11-16T04:12:02.408240Z","iopub.execute_input":"2022-11-16T04:12:02.409385Z","iopub.status.idle":"2022-11-16T04:12:34.861496Z","shell.execute_reply.started":"2022-11-16T04:12:02.409345Z","shell.execute_reply":"2022-11-16T04:12:34.859913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Thank You***","metadata":{}}]}