{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"},"kaggle":{"accelerator":"none","dataSources":[{"sourceType":"competition","sourceId":38760,"databundleVersionId":4493939}],"dockerImageVersionId":30301,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style=\"font-family:verdana;\"> <center>🛍 OTTO – Multi-Objective Recommender System - Getting Started 🧑‍💻</center> </h1>\n\n***","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"<div style=\"color:white;\n           display:fill;\n           border-radius:5px;\n           background-color:#0daae3;\n           font-size:110%;\n           font-family:Verdana;\n           letter-spacing:0.5px\">\n        <p style=\"padding: 10px;\n              color:white;\">\n            Hopefully this notebook will give you a basic understanding of the task and data involved in this competition. Please give an upvote if you find it useful 👍\n        </p>\n    </div>\n    \n<div align = 'center'><img src= \"https://livewin.net/wp-content/uploads/2021/02/Ecommerce.png\" alt =\"Shop\" style='width: 1000px;height 500px'>","metadata":{"execution":{"iopub.status.busy":"2022-11-01T22:06:01.156463Z","iopub.execute_input":"2022-11-01T22:06:01.15692Z","iopub.status.idle":"2022-11-01T22:06:01.16978Z","shell.execute_reply.started":"2022-11-01T22:06:01.15683Z","shell.execute_reply":"2022-11-01T22:06:01.167973Z"}}},{"cell_type":"markdown","source":"### <span style=\"font-family:verdana; word-spacing:1.5px;\"> Contents:\n[Load in the data ⏳](#first-bullet)\n    \n[Data Structure 🗂](#second-bullet)\n    \n[Intital EDA 📊](#third-bullet)\n    \n[Baseline 📈](#fourth-bullet)\n    \n[Where to go next 🚀](#fith-bullet)","metadata":{}},{"cell_type":"markdown","source":"### <span style=\"font-family:verdana; word-spacing:1.5px;\">  Task overview\n    \n<span style=\"font-family:verdana; word-spacing:1.5px;\">  The aim of this competition is to predict e-commerce <span style=\"color:#159364;\">clicks, cart additions, and orders</span>. You'll build a multi-objective recommender system based on previous events in a user session.\n    \n<span style=\"font-family:verdana; word-spacing:1.5px;\"> Current recommender systems consist of various models with different approaches, ranging from simple matrix factorization to a transformer-type deep neural network. However, no single model exists that can simultaneously optimize multiple objectives. In this competition, you’ll build a single entry to predict click-through, add-to-cart, and conversion rates based on previous same-session events.","metadata":{}},{"cell_type":"markdown","source":"### <span style=\"font-family:verdana; word-spacing:1.5px;\">   Imports / setup 🚚","metadata":{}},{"cell_type":"code","source":"### Imports ###\n\nimport pandas as pd\nfrom pathlib import Path\nimport os\nimport random\nimport numpy as np\nimport json\nfrom datetime import timedelta\nfrom collections import Counter\nfrom tqdm.notebook import tqdm\nfrom heapq import nlargest\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nsns.set_theme()\n\nimport warnings\nwarnings.filterwarnings('ignore')\n","metadata":{"execution":{"iopub.status.busy":"2022-12-20T18:59:45.791654Z","iopub.execute_input":"2022-12-20T18:59:45.792055Z","iopub.status.idle":"2022-12-20T18:59:45.800906Z","shell.execute_reply.started":"2022-12-20T18:59:45.792022Z","shell.execute_reply":"2022-12-20T18:59:45.799618Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"### Paths ###\n\nDATA_PATH = Path('../input/otto-recommender-system')\nTRAIN_PATH = DATA_PATH/'train.jsonl'\nTEST_PATH = DATA_PATH/'test.jsonl'\nSAMPLE_SUB_PATH = Path('../input/otto-recommender-system/sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-12-20T18:59:45.948533Z","iopub.execute_input":"2022-12-20T18:59:45.948962Z","iopub.status.idle":"2022-12-20T18:59:45.954779Z","shell.execute_reply.started":"2022-12-20T18:59:45.948928Z","shell.execute_reply":"2022-12-20T18:59:45.953351Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <span style=\"font-family:verdana; word-spacing:1.5px;\">   Load in the data ⏳","metadata":{}},{"cell_type":"code","source":"# Lets check how many lines the training data has!\n\nwith open(TRAIN_PATH, 'r') as f:\n    print(f\"We have {len(f.readlines()):,} lines in the training data\")","metadata":{"execution":{"iopub.status.busy":"2022-12-20T18:59:46.293516Z","iopub.execute_input":"2022-12-20T18:59:46.293952Z","iopub.status.idle":"2022-12-20T19:02:13.570024Z","shell.execute_reply.started":"2022-12-20T18:59:46.29392Z","shell.execute_reply":"2022-12-20T19:02:13.56814Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Load in a sample to a pandas df\n\nsample_size = 150000\n\nchunks = pd.read_json(TRAIN_PATH, lines=True, chunksize = sample_size)\n\nfor c in chunks:\n    sample_train_df = c\n    break","metadata":{"execution":{"iopub.status.busy":"2022-12-20T19:02:13.573018Z","iopub.execute_input":"2022-12-20T19:02:13.574342Z","iopub.status.idle":"2022-12-20T19:02:53.561741Z","shell.execute_reply.started":"2022-12-20T19:02:13.574295Z","shell.execute_reply":"2022-12-20T19:02:53.557177Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sample_train_df.set_index('session', drop=True, inplace=True)\nsample_train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-20T19:02:53.566661Z","iopub.execute_input":"2022-12-20T19:02:53.567263Z","iopub.status.idle":"2022-12-20T19:02:53.671751Z","shell.execute_reply.started":"2022-12-20T19:02:53.567211Z","shell.execute_reply":"2022-12-20T19:02:53.670439Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <span style=\"font-family:verdana; word-spacing:1.5px;\">   Data structure 🗂\n    \n<span style=\"font-family:verdana; word-spacing:1.5px;\">  `session` - the unique session id. Each session contains a list of time ordered events.\n\n<span style=\"font-family:verdana; word-spacing:1.5px;\">  `events` - the time ordered sequence of events in the session. Each event contains 3 pieces of information:\n\n- <span style=\"font-family:verdana; word-spacing:1.5px;\">  `aid` - the article id (product code) of the associated event\n    \n- <span style=\"font-family:verdana; word-spacing:1.5px;\">  `ts` - the Unix timestamp of the event (Unix time is the number of **milliseconds** that have elapsed since 00:00:00 UTC on 1 January 1970) (Thanks to Junji Takeshima https://www.kaggle.com/junjitakeshima for the correction)\n\n- <span style=\"font-family:verdana; word-spacing:1.5px;\">  `type` - the event type, i.e., whether a product was clicked (`clicks`), added to the user's cart (`carts`), or ordered during the session (`orders`)\n","metadata":{"execution":{"iopub.status.busy":"2022-11-01T23:14:19.492394Z","iopub.execute_input":"2022-11-01T23:14:19.492818Z","iopub.status.idle":"2022-11-01T23:14:19.49865Z","shell.execute_reply.started":"2022-11-01T23:14:19.492782Z","shell.execute_reply":"2022-11-01T23:14:19.496998Z"}}},{"cell_type":"code","source":"# Let's look at an example session and print out some basic info\n\n# Sample the first session in the df\nexample_session = sample_train_df.iloc[0].item()\nprint(f'This session was {len(example_session)} actions long \\n')\nprint(f'The first action in the session: \\n {example_session[0]} \\n')\n\n# Time of session\ntime_elapsed = example_session[-1][\"ts\"] - example_session[0][\"ts\"]\n# The timestamp is in milliseconds since 00:00:00 UTC on 1 January 1970\nprint(f'The first session elapsed: {str(timedelta(milliseconds=time_elapsed))} \\n')\n\n# Count the frequency of actions within the session\naction_counts = {}\nfor action in example_session:\n    action_counts[action['type']] = action_counts.get(action['type'], 0) + 1  \nprint(f'The first session contains the following frequency of actions: {action_counts}')","metadata":{"execution":{"iopub.status.busy":"2022-12-20T19:02:53.676743Z","iopub.execute_input":"2022-12-20T19:02:53.677156Z","iopub.status.idle":"2022-12-20T19:02:53.689077Z","shell.execute_reply.started":"2022-12-20T19:02:53.677126Z","shell.execute_reply":"2022-12-20T19:02:53.687312Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <span style=\"font-family:verdana; word-spacing:1.5px;\"> Intital EDA 📊","metadata":{"execution":{"iopub.status.busy":"2022-11-01T23:34:37.808224Z","iopub.execute_input":"2022-11-01T23:34:37.808672Z","iopub.status.idle":"2022-11-01T23:34:37.816409Z","shell.execute_reply.started":"2022-11-01T23:34:37.808638Z","shell.execute_reply":"2022-11-01T23:34:37.814557Z"}}},{"cell_type":"code","source":"### Extract information from each session and add it to the df ###\n\naction_counts_list, article_id_counts_list, session_length_time_list, session_length_action_list = ([] for i in range(4))\noverall_action_counts = {}\noverall_article_id_counts = {}\n\nfor i, row in tqdm(sample_train_df.iterrows(), total=len(sample_train_df)):\n    \n    actions = row['events']\n    \n    # Get the frequency of actions and article_ids\n    action_counts = {}\n    article_id_counts = {}\n    for action in actions:\n        action_counts[action['type']] = action_counts.get(action['type'], 0) + 1\n        article_id_counts[action['aid']] = article_id_counts.get(action['aid'], 0) + 1\n        overall_action_counts[action['type']] = overall_action_counts.get(action['type'], 0) + 1\n        overall_article_id_counts[action['aid']] = overall_article_id_counts.get(action['aid'], 0) + 1\n        \n    # Get the length of the session\n    session_length_time = actions[-1]['ts'] - actions[0]['ts']\n    \n    # Add to list\n    action_counts_list.append(action_counts)\n    article_id_counts_list.append(article_id_counts)\n    session_length_time_list.append(session_length_time)\n    session_length_action_list.append(len(actions))\n    \nsample_train_df['action_counts'] = action_counts_list\nsample_train_df['article_id_counts'] = article_id_counts_list\nsample_train_df['session_length_unix'] = session_length_time_list\nsample_train_df['session_length_hours'] = sample_train_df['session_length_unix']*2.77778e-7  # Convert to hours\nsample_train_df['session_length_action'] = session_length_action_list","metadata":{"execution":{"iopub.status.busy":"2022-12-20T19:02:53.691479Z","iopub.execute_input":"2022-12-20T19:02:53.692896Z","iopub.status.idle":"2022-12-20T19:03:24.438483Z","shell.execute_reply.started":"2022-12-20T19:02:53.69284Z","shell.execute_reply":"2022-12-20T19:03:24.43683Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"### Actions ###\n\ntotal_actions = sum(overall_action_counts.values())\n\nplt.figure(figsize=(8,6))\nsns.barplot(x=list(overall_action_counts.keys()), y=[i/total_actions for i in overall_action_counts.values()]);\nplt.title(f'Action frequency', fontsize=12)\nplt.ylabel('Count', fontsize=12)\nplt.xlabel('Category', fontsize=12)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-20T19:03:24.440951Z","iopub.execute_input":"2022-12-20T19:03:24.441562Z","iopub.status.idle":"2022-12-20T19:03:24.618959Z","shell.execute_reply.started":"2022-12-20T19:03:24.441507Z","shell.execute_reply":"2022-12-20T19:03:24.617913Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"fig, ax = plt.subplots(1,2, figsize=(24, 10))\n\np = sns.distplot(sample_train_df['session_length_action'], color=\"y\", bins= 70, ax=ax[0], kde=False)\np.set_xlabel(\"Number of actions\", fontsize = 16)\np.set_ylabel(\"Density\", fontsize = 16)\np.set_title(\"Distribution of the number of actions taken in each session\", fontsize = 14)\np.axvline(sample_train_df['session_length_action'].mean(), color='r', linestyle='--', label=\"Mean\")\n\np = sns.distplot(sample_train_df['session_length_hours'], color=\"b\", bins= 70, ax=ax[1], kde=False)\np.set_xlabel(\"Hours\", fontsize = 16)\np.set_ylabel(\"Density\", fontsize = 16)\np.set_title(\"Length of each session\", fontsize = 16);","metadata":{"execution":{"iopub.status.busy":"2022-12-20T19:03:24.620838Z","iopub.execute_input":"2022-12-20T19:03:24.621635Z","iopub.status.idle":"2022-12-20T19:03:25.431035Z","shell.execute_reply.started":"2022-12-20T19:03:24.621588Z","shell.execute_reply":"2022-12-20T19:03:25.429583Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Something seems a bit odd with the minutes plot. All the sessions are capped at 650 hours - this needs looking into .. 🤔","metadata":{}},{"cell_type":"code","source":"print(f'{round(len(sample_train_df[sample_train_df[\"session_length_action\"]<10])/len(sample_train_df),3)*100}% of the sessions had less than 10 actions')","metadata":{"execution":{"iopub.status.busy":"2022-12-20T19:03:25.432925Z","iopub.execute_input":"2022-12-20T19:03:25.433307Z","iopub.status.idle":"2022-12-20T19:03:25.542081Z","shell.execute_reply.started":"2022-12-20T19:03:25.433272Z","shell.execute_reply":"2022-12-20T19:03:25.540843Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"article_id_freq = list(overall_article_id_counts.values())\ncut_off = [i for i in article_id_freq if i<30]\n\nplt.figure(figsize=(8,6))\nsns.distplot(cut_off, bins=30, kde=False);\nplt.title(f'Article ID frequency', fontsize=12)\nplt.ylabel('Count', fontsize=12)\nplt.xlabel('Article', fontsize=12);","metadata":{"execution":{"iopub.status.busy":"2022-12-20T19:03:25.543984Z","iopub.execute_input":"2022-12-20T19:03:25.544745Z","iopub.status.idle":"2022-12-20T19:03:26.022734Z","shell.execute_reply.started":"2022-12-20T19:03:25.544698Z","shell.execute_reply":"2022-12-20T19:03:26.021385Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"As we can see from the plot above the vast majority of atricles have a very small number of actions relating to them. There are some exceptions..","metadata":{"execution":{"iopub.status.busy":"2022-11-02T00:44:03.271921Z","iopub.execute_input":"2022-11-02T00:44:03.272395Z","iopub.status.idle":"2022-11-02T00:44:03.565989Z","shell.execute_reply.started":"2022-11-02T00:44:03.272355Z","shell.execute_reply":"2022-11-02T00:44:03.564369Z"}}},{"cell_type":"code","source":"### Look at the most interacted with articles ###\nprint(f'Frequency of most common articles: {sorted(list(overall_article_id_counts.values()))[-5:]} \\n')\nres = nlargest(5, overall_article_id_counts, key = overall_article_id_counts.get)\nprint(f'IDs for those common articles: {res}')","metadata":{"execution":{"iopub.status.busy":"2022-12-20T19:03:26.028404Z","iopub.execute_input":"2022-12-20T19:03:26.02881Z","iopub.status.idle":"2022-12-20T19:03:26.550971Z","shell.execute_reply.started":"2022-12-20T19:03:26.028776Z","shell.execute_reply":"2022-12-20T19:03:26.54964Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <span style=\"font-family:verdana; word-spacing:1.5px;\">   Baseline 📈\n    \n<span style=\"font-family:verdana; word-spacing:1.5px;\"> The `test data` contains truncated session data similar to that of the training data. The task is to <span style=\"color:#159364;\"> predict the next aid clicked </span> after the session truncation, as well as the the remaining aids that are added to carts and orders; you may predict up to 20 values for each session type\n    \n<span style=\"font-family:verdana; word-spacing:1.5px;\"> Submissions are evaluated on <span style=\"color:#159364;\"> Recall </span> each action type, and the three recall values are weight-averaged: <span style=\"color:#159364;\"> {'clicks': 0.10, 'carts': 0.30, 'orders': 0.60} </span>. It is important to get the 'orders' predictions correct as they carry most of the weigthing :)\n    \n<span style=\"font-family:verdana; word-spacing:1.5px;\"> For each <span style=\"color:#159364;\"> session</span> in the test data, your task it to predict the <span style=\"color:#159364;\"> aid</span> values for each <span style=\"color:#159364;\"> type</span> that occur after the last timestamp ts the test session. In other words, the test data contains sessions truncated by timestamp, and you are to predict what occurs after the point of truncation.\n\n<span style=\"font-family:verdana; word-spacing:1.5px;\"> For <span style=\"color:#159364;\"> clicks </span>there is only a single ground truth value for each session, which is the next <span style=\"color:#159364;\"> aid</span> clicked during the session (although you can still predict up to 20 aid values). The ground truth for <span style=\"color:#159364;\"> carts</span> and <span style=\"color:#159364;\"> orders</span> contains all <span style=\"color:#159364;\"> aid </span>values that were added to a cart and ordered respectively during the session.\n    \nEach session and type combination should appear on its own session_type row in the submission (3 rows per session), and predictions should be space delimited. This can be seen in the `sample_test_df` below..","metadata":{}},{"cell_type":"code","source":"with open(TEST_PATH, 'r') as f:\n    print(f\"We have {len(f.readlines()):,} lines in the test data\")","metadata":{"execution":{"iopub.status.busy":"2022-12-20T19:03:26.552847Z","iopub.execute_input":"2022-12-20T19:03:26.553229Z","iopub.status.idle":"2022-12-20T19:03:32.669482Z","shell.execute_reply.started":"2022-12-20T19:03:26.553177Z","shell.execute_reply":"2022-12-20T19:03:32.667811Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Load in a sample to a pandas df\n\nsample_size = 150\n\nchunks = pd.read_json(TEST_PATH, lines=True, chunksize = sample_size)\n\nfor c in chunks:\n    sample_test_df = c\n    break","metadata":{"execution":{"iopub.status.busy":"2022-12-20T19:03:32.671576Z","iopub.execute_input":"2022-12-20T19:03:32.672074Z","iopub.status.idle":"2022-12-20T19:03:32.690794Z","shell.execute_reply.started":"2022-12-20T19:03:32.672028Z","shell.execute_reply":"2022-12-20T19:03:32.689491Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sample_test_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-20T19:03:32.692653Z","iopub.execute_input":"2022-12-20T19:03:32.693088Z","iopub.status.idle":"2022-12-20T19:03:32.738323Z","shell.execute_reply.started":"2022-12-20T19:03:32.693049Z","shell.execute_reply":"2022-12-20T19:03:32.737065Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<span style=\"font-family:verdana; word-spacing:1.5px;\"> Below shows a sample submission. For each `session` in the test set there is a prediction (`labels`). This predicts what articles will be next interacted with in that session. For each session there are three actions (clicks, carts, orders), predictions are made for all three actions.","metadata":{}},{"cell_type":"code","source":"sample_submission = pd.read_csv(SAMPLE_SUB_PATH)\nsample_submission.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-20T19:03:32.740026Z","iopub.execute_input":"2022-12-20T19:03:32.740499Z","iopub.status.idle":"2022-12-20T19:03:40.478827Z","shell.execute_reply.started":"2022-12-20T19:03:32.740463Z","shell.execute_reply":"2022-12-20T19:03:40.477092Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<span style=\"font-family:verdana; word-spacing:1.5px;\"> Lets find the most common article for each different type of action.","metadata":{}},{"cell_type":"code","source":"sample_size = 150000\n\nchunks = pd.read_json(TRAIN_PATH, lines=True, chunksize = sample_size)\n\nclicks_article_list = []\ncarts_article_list = []\norders_article_list = []\n\nfor e, c in enumerate(chunks):\n    \n    # Save time by not using all the data\n    if e > 2:\n        break\n    \n    sample_train_df = c\n    \n    for i, row in c.iterrows():\n        actions = row['events']\n        for action in actions:\n            if action['type'] == 'clicks':\n                clicks_article_list.append(action['aid'])\n            elif action['type'] == 'carts':\n                carts_article_list.append(action['aid'])\n            else:\n                orders_article_list.append(action['aid'])\n    ","metadata":{"execution":{"iopub.status.busy":"2022-12-20T19:03:40.481332Z","iopub.execute_input":"2022-12-20T19:03:40.481737Z","iopub.status.idle":"2022-12-20T19:05:30.51631Z","shell.execute_reply.started":"2022-12-20T19:03:40.481703Z","shell.execute_reply":"2022-12-20T19:05:30.514706Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create dictionaries with articles and their frequencies\narticle_click_freq = Counter(clicks_article_list)\narticle_carts_freq = Counter(carts_article_list)\narticle_order_freq = Counter(orders_article_list)","metadata":{"execution":{"iopub.status.busy":"2022-12-20T19:05:30.521891Z","iopub.execute_input":"2022-12-20T19:05:30.522933Z","iopub.status.idle":"2022-12-20T19:05:42.928857Z","shell.execute_reply.started":"2022-12-20T19:05:30.522877Z","shell.execute_reply":"2022-12-20T19:05:42.927219Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Get the 20 most frequent articles for each action\ntop_click_article = nlargest(20, article_click_freq, key = article_click_freq.get)\ntop_carts_article = nlargest(20, article_carts_freq, key = article_carts_freq.get)\ntop_order_article = nlargest(20, article_order_freq, key = article_order_freq.get) ","metadata":{"execution":{"iopub.status.busy":"2022-12-20T19:05:42.930826Z","iopub.execute_input":"2022-12-20T19:05:42.931378Z","iopub.status.idle":"2022-12-20T19:05:44.012512Z","shell.execute_reply.started":"2022-12-20T19:05:42.931333Z","shell.execute_reply":"2022-12-20T19:05:44.011032Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Create a dict with this info\nfrequent_articles = {'clicks': top_click_article, 'carts':top_carts_article, 'order':top_order_article}","metadata":{"execution":{"iopub.status.busy":"2022-12-20T19:05:44.014472Z","iopub.execute_input":"2022-12-20T19:05:44.014895Z","iopub.status.idle":"2022-12-20T19:05:44.021655Z","shell.execute_reply.started":"2022-12-20T19:05:44.014856Z","shell.execute_reply":"2022-12-20T19:05:44.020033Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for action in ['clicks', 'carts', 'order']:\n    print(f'Most frequent articles for {action}: {frequent_articles[action][:5]}') # Correction by @danielliao 🙏","metadata":{"execution":{"iopub.status.busy":"2022-12-20T19:05:44.024052Z","iopub.execute_input":"2022-12-20T19:05:44.025143Z","iopub.status.idle":"2022-12-20T19:05:44.035063Z","shell.execute_reply.started":"2022-12-20T19:05:44.025095Z","shell.execute_reply":"2022-12-20T19:05:44.034062Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<span style=\"font-family:verdana; word-spacing:1.5px;\"> There is some overlap but the articles do change for the different actions!\n    \n<span style=\"font-family:verdana; word-spacing:1.5px;\"> This baseline will use the fact that people will often interact with articles they have previouslt interacted with. The prediction will consist of the top 20 most frequent articles in the session. If there are less than 20 articles in the session the prediction will be padded with the most frequent articles in the training data as found above.","metadata":{}},{"cell_type":"code","source":"test_data = pd.read_json(TEST_PATH, lines=True, chunksize=1000)\n\npreds = []\n\nfor chunk in tqdm(test_data, total=1671):\n    \n    for i, row in chunk.iterrows():\n        actions = row['events']\n        article_id_list = []\n        for action in actions:\n            article_id_list.append(action['aid'])\n            \n        # Get 20 most common article ID for the session\n        article_freq = Counter(article_id_list)\n        top_articles = nlargest(20, article_freq, key = article_freq.get)\n        \n        # Pad with most popular items in training\n        padding_size = (20 - len(top_articles)) # Correction by @danielliao 🙏\n        for action in ['clicks', 'carts', 'order']:\n            top_articles_added = top_articles + frequent_articles[action][:padding_size] # Correction by @danielliao 🙏\n            preds.append(\" \".join([str(id) for id in top_articles_added]))","metadata":{"execution":{"iopub.status.busy":"2022-12-20T19:10:31.646599Z","iopub.execute_input":"2022-12-20T19:10:31.647936Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Predict the 20 most common atricles for each test session\nsample_submission['labels'] = preds","metadata":{"execution":{"iopub.status.busy":"2022-12-20T18:57:19.32662Z","iopub.execute_input":"2022-12-20T18:57:19.327094Z","iopub.status.idle":"2022-12-20T18:57:20.274746Z","shell.execute_reply.started":"2022-12-20T18:57:19.327047Z","shell.execute_reply":"2022-12-20T18:57:20.27323Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sample_submission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-11-02T11:00:24.080451Z","iopub.execute_input":"2022-11-02T11:00:24.08139Z","iopub.status.idle":"2022-11-02T11:00:59.182417Z","shell.execute_reply.started":"2022-11-02T11:00:24.081354Z","shell.execute_reply":"2022-11-02T11:00:59.181337Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Where to go next 🚀\n\n- At the moment we pad every prediction with the same most frequently occuring articles. The model could be improved if we looked at which articles co-occured frequently in the same session ([KJ](https://www.kaggle.com/code/whitelily/co-occurrence-baseline) and [VLADIMIR SLAYKOVSKIY](https://www.kaggle.com/code/vslaykovsky/co-visitation-matrix) have started looking at this).\n\n- While shopping there is a common sequence of events: click -> cart -> order. Currently we look at which articles are most common for each action but surely the most likely item to be ordered is one already in the cart 🤔\n","metadata":{}},{"cell_type":"markdown","source":"# WIP :)","metadata":{}},{"cell_type":"code","source":"","metadata":{},"outputs":[],"execution_count":null}]}