{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## OTTO – Multi-Objective Recommender System","metadata":{}},{"cell_type":"markdown","source":"## 1. Setup","metadata":{}},{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings('ignore')\nimport pathlib\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-11-12T06:52:37.380516Z","iopub.execute_input":"2022-11-12T06:52:37.380866Z","iopub.status.idle":"2022-11-12T06:52:37.386311Z","shell.execute_reply.started":"2022-11-12T06:52:37.380838Z","shell.execute_reply":"2022-11-12T06:52:37.385490Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"competition_dataset_directory = pathlib.Path('../input/otto-recommender-system')\npickled_dataset_directory = pathlib.Path('../input/otto-multi-objective-recommender-system-pickle')\n\ndf_train = pd.read_pickle(pickled_dataset_directory / 'train.pkl')\ndf_test = pd.read_pickle(pickled_dataset_directory / 'test.pkl')\n\nprint(f'Training Shape: {df_train.shape} - Memory Usage: {df_train.memory_usage().sum() / 1024 ** 2:.2f} MB')\nprint(f'Test Shape: {df_test.shape} - Memory Usage: {df_test.memory_usage().sum() / 1024 ** 2:.2f} MB')","metadata":{"execution":{"iopub.status.busy":"2022-11-12T06:52:37.715921Z","iopub.execute_input":"2022-11-12T06:52:37.716274Z","iopub.status.idle":"2022-11-12T06:53:12.734936Z","shell.execute_reply.started":"2022-11-12T06:52:37.716245Z","shell.execute_reply":"2022-11-12T06:53:12.733521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2. Introduction\n\nThis is a large dataset of e-commerce sessions. Sessions consist of user events such as clicking to a product, adding a product to cart or ordering a product. Each session belongs to a unique user and duration of sessions are different from each other.\n\n* `session`: unique ID of the user session\n* `aid`: unique ID of the product\n* `ts`: timestamp of the event\n* `type`: category of the event\n\nIt is a multi-objective ranking problem since there are 3 categories (click, cart, order) to recommend for each session. Recommendations for each category are evaluated on recall@20 and scores are multiplied with their corresponding weights.\n\n* **216.7M** events in training set and **6.9M** events in test set\n* **12.8M** sessions in training set and **1.6M** sessions in test set\n* **1.8M** products in training set and **783K** products in test set\n* **194.7M** clicks in training set and **6.2M** clicks in test set\n* **16.8M** carts in training set and **570K** carts in test set\n* **5M** orders in training set and **65K** orders in test set","metadata":{}},{"cell_type":"code","source":"train_events = df_train.shape[0]\ntest_events = df_test.shape[0]\nprint(f'Number of Events - Training: {train_events} | Test: {test_events}')\n\ntrain_unique_sessions = df_train['session'].unique()\ntest_unique_sessions = df_test['session'].unique()\nprint(f'Number of Unique Sessions - Training: {len(train_unique_sessions)} | Test: {len(test_unique_sessions)}')\ndel train_unique_sessions, test_unique_sessions\n\ntrain_unique_aids = df_train['aid'].unique()\ntest_unique_aids = df_test['aid'].unique()\noverlapping_aids = set(train_unique_aids).intersection(set(test_unique_aids))\nprint(f'Number of Unique Products - Training: {len(train_unique_aids)} | Test: {len(test_unique_aids)} - ({len(overlapping_aids)} Overlapping Products)')\ndel train_unique_aids, test_unique_aids, overlapping_aids\n\ntrain_clicks = df_train[df_train['type'] == 0].shape[0]\ntest_clicks = df_test[df_test['type'] == 0].shape[0]\nprint(f'Number of Clicks - Training: {train_clicks} | Test: {test_clicks}')\n\ntrain_carts = df_train[df_train['type'] == 1].shape[0]\ntest_carts = df_test[df_test['type'] == 1].shape[0]\nprint(f'Number of Carts - Training: {train_carts} | Test: {test_carts}')\n\ntrain_orders = df_train[df_train['type'] == 2].shape[0]\ntest_orders = df_test[df_test['type'] == 2].shape[0]\nprint(f'Number of Orders - Training: {train_orders} | Test: {test_orders}')","metadata":{"execution":{"iopub.status.busy":"2022-11-05T05:27:42.639682Z","iopub.execute_input":"2022-11-05T05:27:42.640397Z","iopub.status.idle":"2022-11-05T05:28:05.488052Z","shell.execute_reply.started":"2022-11-05T05:27:42.640351Z","shell.execute_reply":"2022-11-05T05:28:05.486096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_start = df_train['ts'].min().strftime('%Y.%m.%d %X')\ntrain_end = df_train['ts'].max().strftime('%Y.%m.%d %X')\ntest_start = df_test['ts'].min().strftime('%Y.%m.%d %X')\ntest_end = df_test['ts'].max().strftime('%Y.%m.%d %X')\nprint(f'Events Time Range\\nTraining: {train_start} - {train_end}\\nTest: {test_start} - {test_end}')","metadata":{"execution":{"iopub.status.busy":"2022-11-07T16:29:15.096560Z","iopub.execute_input":"2022-11-07T16:29:15.097007Z","iopub.status.idle":"2022-11-07T16:29:16.572333Z","shell.execute_reply.started":"2022-11-07T16:29:15.096975Z","shell.execute_reply":"2022-11-07T16:29:16.571388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Training and test set are split by time. Training set is 4 weeks of user events and test set is the following week after training set. There aren't any overlapping events between training and test set because some of the training set events are trimmed in order to prevent leakage. Proportions of event types are similar in training and test set which can be seen from the visualization below.","metadata":{}},{"cell_type":"code","source":"def visualize_categorical_feature_distribution(df, feature, path=None):\n\n    \"\"\"\n    Visualize distribution of given categorical column in given dataframe\n    \n    Parameters\n    ----------\n    df: pandas.DataFrame of shape (n_samples, 4)\n        Dataframe with session, aid, ts, type columns\n        \n    feature: str\n        Name of the categorical feature\n    \n    path: path-like str or None\n        Path of the output file or None (if path is None, plot is displayed with selected backend)\n    \"\"\"\n\n    fig, ax = plt.subplots(figsize=(24, df[feature].value_counts().shape[0] + 4), dpi=100)\n    sns.barplot(\n        y=df[feature].value_counts().values,\n        x=df[feature].value_counts().index,\n        color='tab:blue',\n        ax=ax\n    )\n    ax.set_xlabel('')\n    ax.set_ylabel('')\n    ax.set_xticklabels([\n        f'{x} ({value_count:,})' for value_count, x in zip(\n            df[feature].value_counts().values,\n            df[feature].value_counts().index\n        )\n    ])\n    ax.tick_params(axis='x', labelsize=15, pad=10)\n    ax.tick_params(axis='y', labelsize=15, pad=10)\n    ax.set_title(f'Value Counts {feature}', size=20, pad=15)\n\n    if path is None:\n        plt.show()\n    else:\n        plt.savefig(path)\n        plt.close(fig)\n\n\nvisualize_categorical_feature_distribution(df=df_train, feature='type')\nvisualize_categorical_feature_distribution(df=df_test, feature='type')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-11-04T14:02:12.024212Z","iopub.execute_input":"2022-11-04T14:02:12.024783Z","iopub.status.idle":"2022-11-04T14:02:20.070331Z","shell.execute_reply.started":"2022-11-04T14:02:12.024722Z","shell.execute_reply":"2022-11-04T14:02:20.068868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3. Users and Products\n\nAs we already know, there are 12.8M unique users (sessions) in training set and 1.6m unique sessions in test set. However, their statistics are also different between training and test set. In training set, mean session duration is 4 times longer than test set. Shortest session in training set is 2 because it wouldn't be possible to create ground-truth on single event sessions. Longest session in training set is 500 which looks like a threshold for trimming extremely long sessions.\n\nDiscrepancy in session statistics may not be a problem because training set is 4 weeks while test set is only a single week. Mean session duration being 4 times longer in training set is also an artifact of that.\n\nDistributions of session durations can be seen from the visualization below. Both densities are on log scale for interpretability.","metadata":{}},{"cell_type":"code","source":"def visualize_continuous_feature_distribution(df_train, df_test, feature, path=None):\n\n    \"\"\"\n    Visualize distribution of given continuous column in given dataframe\n    \n    Parameters\n    ----------\n    df_train: pandas.DataFrame of shape (n_samples, 4)\n        Training dataframe with session, aid, ts, type columns\n        \n    df_test: pandas.DataFrame of shape (n_samples, 4)\n        Test dataframe with session, aid, ts, type columns\n        \n    feature: str\n        Name of the continuous feature\n    \n    path: path-like str or None\n        Path of the output file or None (if path is None, plot is displayed with selected backend)\n    \"\"\"\n    \n    fig, ax = plt.subplots(figsize=(24, 6), dpi=100)\n    sns.kdeplot(df_train[feature], label='train', fill=True, log_scale=True, ax=ax)\n    sns.kdeplot(df_test[feature], label='test', fill=True, log_scale=True, ax=ax)\n    ax.tick_params(axis='x', labelsize=12.5)\n    ax.tick_params(axis='y', labelsize=12.5)\n    ax.set_xlabel('')\n    ax.set_ylabel('')\n    ax.legend(prop={'size': 15})\n    title = f'''\n    {feature}\n    Mean - Train: {df_train[feature].mean():.2f} |  Test: {df_test[feature].mean():.2f}\n    Median - Train: {df_train[feature].median():.2f} |  Test: {df_test[feature].median():.2f}\n    Std - Train: {df_train[feature].std():.2f} |  Test: {df_test[feature].std():.2f}\n    Min - Train: {df_train[feature].min():.2f} |  Test: {df_test[feature].min():.2f}\n    Max - Train: {df_train[feature].max():.2f} |  Test: {df_test[feature].max():.2f}\n    '''\n    ax.set_title(title, size=20, pad=15)\n    \n    if path is None:\n        plt.show()\n    else:\n        plt.savefig(path)\n        plt.close(fig)\n\n    \ndf_train_session_counts = df_train.groupby('session')[['session']].count()\ndf_test_session_counts = df_test.groupby('session')[['session']].count()\nvisualize_continuous_feature_distribution(df_train=df_train_session_counts, df_test=df_test_session_counts, feature='session')\ndel df_train_session_counts, df_test_session_counts","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-11-07T16:57:14.948573Z","iopub.execute_input":"2022-11-07T16:57:14.949087Z","iopub.status.idle":"2022-11-07T16:58:23.970797Z","shell.execute_reply.started":"2022-11-07T16:57:14.949045Z","shell.execute_reply":"2022-11-07T16:58:23.969397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 1.8M unique products (aids) in training set and 783K unique aids in test set. aids in test set are a subset of aids in training set so there aren't any unseen aid in test set. Distribution of aid occurences is different in training and test set as well because of the previously mentioned reason. Training set aid occurences distribution is more normal while test set aid occurences distribution is multi-modal which could be an artifact of session truncation.\n\nDistributions of aid occurences can be seen from the visualization below. Both densities are on log scale for interpretability.","metadata":{}},{"cell_type":"code","source":"df_train_aid_counts = df_train.groupby('aid')[['aid']].count()\ndf_test_aid_counts = df_test.groupby('aid')[['aid']].count()\nvisualize_continuous_feature_distribution(df_train=df_train_aid_counts, df_test=df_test_aid_counts, feature='aid')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-11-07T17:01:44.884351Z","iopub.execute_input":"2022-11-07T17:01:44.885514Z","iopub.status.idle":"2022-11-07T17:02:13.444652Z","shell.execute_reply.started":"2022-11-07T17:01:44.885467Z","shell.execute_reply":"2022-11-07T17:02:13.443341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_aid_counts = df_train_aid_counts.rename(columns={'aid': 'count'}).reset_index()\ndf_test_aid_counts = df_test_aid_counts.rename(columns={'aid': 'count'}).reset_index()\ndf_all_aid_counts = pd.concat((df_train_aid_counts, df_test_aid_counts), axis=0, ignore_index=True).reset_index(drop=True)\ndf_all_aid_counts = df_all_aid_counts.groupby('aid')['count'].sum().reset_index()\n\ndf_train_aid_counts.sort_values(by='count', ascending=False, inplace=True)\ndf_test_aid_counts.sort_values(by='count', ascending=False, inplace=True)\ndf_all_aid_counts.sort_values(by='count', ascending=False, inplace=True)\n\ntrain_20_most_frequent_aids = df_train_aid_counts.set_index('aid').head(20).to_dict()['count']\ntest_20_most_frequent_aids = df_test_aid_counts.set_index('aid').head(20).to_dict()['count']\nall_20_most_frequent_aids = df_all_aid_counts.set_index('aid').head(20).to_dict()['count']\n\ndel df_train_aid_counts, df_test_aid_counts, df_all_aid_counts","metadata":{"execution":{"iopub.status.busy":"2022-11-07T17:05:20.647151Z","iopub.execute_input":"2022-11-07T17:05:20.647672Z","iopub.status.idle":"2022-11-07T17:05:21.882890Z","shell.execute_reply.started":"2022-11-07T17:05:20.647629Z","shell.execute_reply":"2022-11-07T17:05:21.882012Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"aid frequencies in training and test are very similar. There are minor differences between them since test set is almost 20 times smaller than training set. Computing aid frequencies on concatenated training and test set would be more reliable if aid frequencies are utilized in any way.\n\nTop 20 most frequent aids are visualized for training, test and training + test sets below. X axis is the frequency and y axis labels are aid and its frequency inside parentheses.","metadata":{}},{"cell_type":"code","source":"def visualize_aid_frequencies(aid_frequencies, title, path=None):\n    \n    \"\"\"\n    Visualize aids and their frequencies in given dictionary\n    \n    Parameters\n    ----------\n    aid_frequencies: dict\n        Dictionary of aids and their frequencies\n        \n    title: str\n        Title of the plot\n    \n    path: path-like str or None\n        Path of the output file or None (if path is None, plot is displayed with selected backend)\n    \"\"\"\n    \n    fig, ax = plt.subplots(figsize=(24, 20), dpi=100)\n    ax.barh(range(len(aid_frequencies)), aid_frequencies.values(), align='center')\n    ax.set_xlabel('')\n    ax.set_ylabel('')\n    ax.set_yticks(range(len(aid_frequencies)))\n    ax.set_yticklabels([f'{x} ({value_count:,})' for x, value_count in aid_frequencies.items()])\n    ax.tick_params(axis='x', labelsize=15, pad=10)\n    ax.tick_params(axis='y', labelsize=15, pad=10)\n    ax.set_title(title, size=20, pad=15)\n    plt.gca().invert_yaxis()\n\n    if path is None:\n        plt.show()\n    else:\n        plt.savefig(path)\n        plt.close(fig)\n\n\nvisualize_aid_frequencies(\n    aid_frequencies=train_20_most_frequent_aids,\n    title='Top 20 Most Frequent aids in Training Set'\n)\nvisualize_aid_frequencies(\n    aid_frequencies=test_20_most_frequent_aids,\n    title='Top 20 Most Frequent aids in Test Set'\n)\nvisualize_aid_frequencies(\n    aid_frequencies=all_20_most_frequent_aids,\n    title='Top 20 Most Frequent aids in Training + Test Set'\n)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-11-07T17:08:25.505982Z","iopub.execute_input":"2022-11-07T17:08:25.506433Z","iopub.status.idle":"2022-11-07T17:08:27.167671Z","shell.execute_reply.started":"2022-11-07T17:08:25.506394Z","shell.execute_reply":"2022-11-07T17:08:27.166451Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"aid frequencies are also different between clicks, carts and orders. global aid frequencies are closer to click aid frequencies because clicks dominate the event types.","metadata":{}},{"cell_type":"code","source":"df_train_aids_by_types  = df_train.groupby(['type', 'aid'])[['aid']].count()\ndf_test_aids_by_types  = df_test.groupby(['type', 'aid'])[['aid']].count()\ndf_train_click_aid_counts = df_train_aids_by_types.loc[0].rename(columns={'aid': 'count'}).reset_index()\ndf_train_cart_aid_counts = df_train_aids_by_types.loc[1].rename(columns={'aid': 'count'}).reset_index()\ndf_train_order_aid_counts = df_train_aids_by_types.loc[2].rename(columns={'aid': 'count'}).reset_index()\ndf_test_click_aid_counts = df_test_aids_by_types.loc[0].rename(columns={'aid': 'count'}).reset_index()\ndf_test_cart_aid_counts = df_test_aids_by_types.loc[1].rename(columns={'aid': 'count'}).reset_index()\ndf_test_order_aid_counts = df_test_aids_by_types.loc[2].rename(columns={'aid': 'count'}).reset_index()\ndf_all_click_aid_counts = pd.concat((df_train_click_aid_counts, df_test_click_aid_counts), axis=0, ignore_index=True).reset_index(drop=True)\ndf_all_click_aid_counts = df_all_click_aid_counts.groupby('aid')['count'].sum().reset_index()\ndf_all_cart_aid_counts = pd.concat((df_train_cart_aid_counts, df_test_cart_aid_counts), axis=0, ignore_index=True).reset_index(drop=True)\ndf_all_cart_aid_counts = df_all_cart_aid_counts.groupby('aid')['count'].sum().reset_index()\ndf_all_order_aid_counts = pd.concat((df_train_order_aid_counts, df_test_order_aid_counts), axis=0, ignore_index=True).reset_index(drop=True)\ndf_all_order_aid_counts = df_all_order_aid_counts.groupby('aid')['count'].sum().reset_index()\n\ndf_train_click_aid_counts.sort_values(by='count', ascending=False, inplace=True)\ndf_test_click_aid_counts.sort_values(by='count', ascending=False, inplace=True)\ndf_all_click_aid_counts.sort_values(by='count', ascending=False, inplace=True)\n\ndf_train_cart_aid_counts.sort_values(by='count', ascending=False, inplace=True)\ndf_test_cart_aid_counts.sort_values(by='count', ascending=False, inplace=True)\ndf_all_cart_aid_counts.sort_values(by='count', ascending=False, inplace=True)\n\ndf_train_order_aid_counts.sort_values(by='count', ascending=False, inplace=True)\ndf_test_order_aid_counts.sort_values(by='count', ascending=False, inplace=True)\ndf_all_order_aid_counts.sort_values(by='count', ascending=False, inplace=True)\n\ntrain_20_most_frequent_click_aids = df_train_click_aid_counts.set_index('aid').head(20).to_dict()['count']\ntest_20_most_frequent_click_aids = df_test_click_aid_counts.set_index('aid').head(20).to_dict()['count']\nall_20_most_frequent_click_aids = df_all_click_aid_counts.set_index('aid').head(20).to_dict()['count']\n\ntrain_20_most_frequent_cart_aids = df_train_cart_aid_counts.set_index('aid').head(20).to_dict()['count']\ntest_20_most_frequent_cart_aids = df_test_cart_aid_counts.set_index('aid').head(20).to_dict()['count']\nall_20_most_frequent_cart_aids = df_all_cart_aid_counts.set_index('aid').head(20).to_dict()['count']\n\ntrain_20_most_frequent_order_aids = df_train_order_aid_counts.set_index('aid').head(20).to_dict()['count']\ntest_20_most_frequent_order_aids = df_test_order_aid_counts.set_index('aid').head(20).to_dict()['count']\nall_20_most_frequent_order_aids = df_all_order_aid_counts.set_index('aid').head(20).to_dict()['count']","metadata":{"execution":{"iopub.status.busy":"2022-11-07T17:20:02.656206Z","iopub.execute_input":"2022-11-07T17:20:02.656702Z","iopub.status.idle":"2022-11-07T17:20:04.718244Z","shell.execute_reply.started":"2022-11-07T17:20:02.656664Z","shell.execute_reply":"2022-11-07T17:20:04.717176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Top 20 most frequent click aids are visualized for training, test and training + test sets below. X axis is the frequency and y axis labels are aid and its frequency inside parentheses.","metadata":{}},{"cell_type":"code","source":"visualize_aid_frequencies(\n    aid_frequencies=train_20_most_frequent_click_aids,\n    title='Top 20 Most Frequent click aids in Training Set'\n)\nvisualize_aid_frequencies(\n    aid_frequencies=test_20_most_frequent_click_aids,\n    title='Top 20 Most Frequent click aids in Test Set'\n)\nvisualize_aid_frequencies(\n    aid_frequencies=all_20_most_frequent_click_aids,\n    title='Top 20 Most Frequent click aids in Training + Test Set'\n)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-11-07T17:20:48.817638Z","iopub.execute_input":"2022-11-07T17:20:48.818047Z","iopub.status.idle":"2022-11-07T17:20:50.489585Z","shell.execute_reply.started":"2022-11-07T17:20:48.818014Z","shell.execute_reply":"2022-11-07T17:20:50.488207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Top 20 most frequent cart aids are visualized for training, test and training + test sets below. X axis is the frequency and y axis labels are aid and its frequency inside parentheses.","metadata":{}},{"cell_type":"code","source":"visualize_aid_frequencies(\n    aid_frequencies=train_20_most_frequent_cart_aids,\n    title='Top 20 Most Frequent cart aids in Training Set'\n)\nvisualize_aid_frequencies(\n    aid_frequencies=test_20_most_frequent_cart_aids,\n    title='Top 20 Most Frequent cart aids in Test Set'\n)\nvisualize_aid_frequencies(\n    aid_frequencies=all_20_most_frequent_cart_aids,\n    title='Top 20 Most Frequent cart aids in Training + Test Set'\n)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T17:22:41.971931Z","iopub.execute_input":"2022-11-07T17:22:41.972548Z","iopub.status.idle":"2022-11-07T17:22:43.654121Z","shell.execute_reply.started":"2022-11-07T17:22:41.972501Z","shell.execute_reply":"2022-11-07T17:22:43.652918Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Top 20 most frequent order aids are visualized for training, test and training + test sets below. X axis is the frequency and y axis labels are aid and its frequency inside parentheses.","metadata":{}},{"cell_type":"code","source":"visualize_aid_frequencies(\n    aid_frequencies=train_20_most_frequent_order_aids,\n    title='Top 20 Most Frequent order aids in Training Set'\n)\nvisualize_aid_frequencies(\n    aid_frequencies=test_20_most_frequent_order_aids,\n    title='Top 20 Most Frequent order aids in Test Set'\n)\nvisualize_aid_frequencies(\n    aid_frequencies=all_20_most_frequent_order_aids,\n    title='Top 20 Most Frequent order aids in Training + Test Set'\n)","metadata":{"execution":{"iopub.status.busy":"2022-11-07T17:23:14.776572Z","iopub.execute_input":"2022-11-07T17:23:14.777048Z","iopub.status.idle":"2022-11-07T17:23:16.363707Z","shell.execute_reply.started":"2022-11-07T17:23:14.777014Z","shell.execute_reply":"2022-11-07T17:23:16.362608Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4. Events\n\nIn training set 93% of sessions are clicks, 16% of sessions are carts and 10% of sessions are orders on average. Those numbers don't add up to 100 because they are averages calculated on all sessions. They have 12, 10 and 8 standard deviations respectively. In test set those numbers are little bit higher for averages and standard deviations because sessions are truncated randomly. Events are likely over/under represented in test sessions.\n\n(0 -> clicks, 1 -> carts, 2 -> orders)","metadata":{}},{"cell_type":"code","source":"train_session_type_groupby = df_train.groupby(['session', 'type'])\ndf_train_sessions_types = train_session_type_groupby[['session']].count().rename(columns={'session': 'count'}).reset_index()\ndf_train_sessions_types['total'] = df_train_sessions_types.groupby('session')['count'].transform('sum')\ndf_train_sessions_types.sort_values(by=['total', 'session', 'type'], ascending=False, inplace=True)\ndf_train_sessions_types['rate'] = df_train_sessions_types['count'] / df_train_sessions_types['total']\ndf_train_sessions_type_rate_means = df_train_sessions_types.groupby('type')['rate'].mean().to_dict()\ndf_train_sessions_type_rate_stds = df_train_sessions_types.groupby('type')['rate'].std().to_dict()\n\ntest_session_type_groupby = df_test.groupby(['session', 'type'])\ndf_test_sessions_types = test_session_type_groupby[['session']].count().rename(columns={'session': 'count'}).reset_index()\ndf_test_sessions_types['total'] = df_test_sessions_types.groupby('session')['count'].transform('sum')\ndf_test_sessions_types.sort_values(by=['total', 'session', 'type'], ascending=False, inplace=True)\ndf_test_sessions_types['rate'] = df_test_sessions_types['count'] / df_test_sessions_types['total']\ndf_test_sessions_type_rate_means = df_test_sessions_types.groupby('type')['rate'].mean().to_dict()\ndf_test_sessions_type_rate_stds = df_test_sessions_types.groupby('type')['rate'].std().to_dict()\n\nprint(\nf'''\nSession Event Type Percentages on Average\nTraining: 0: {df_train_sessions_type_rate_means[0]:.4f}(±{df_train_sessions_type_rate_stds[0]:.4f}) | 1: {df_train_sessions_type_rate_means[1]:.4f}(±{df_train_sessions_type_rate_stds[1]:.4f}) | 2: {df_train_sessions_type_rate_means[2]:.4f}(±{df_train_sessions_type_rate_stds[2]:.4f})\nTest: 0: {df_test_sessions_type_rate_means[0]:.4f}(±{df_test_sessions_type_rate_stds[0]:.4f}) | 1: {df_test_sessions_type_rate_means[1]:.4f}(±{df_test_sessions_type_rate_stds[1]:.4f}) | 2: {df_test_sessions_type_rate_means[2]:.4f}(±{df_test_sessions_type_rate_stds[2]:.4f})\n'''\n)\ndel df_train_sessions_type_rate_means, df_train_sessions_type_rate_stds, df_test_sessions_type_rate_means, df_test_sessions_type_rate_stds","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-11-06T08:02:56.142319Z","iopub.execute_input":"2022-11-06T08:02:56.142708Z","iopub.status.idle":"2022-11-06T08:03:57.230773Z","shell.execute_reply.started":"2022-11-06T08:02:56.142677Z","shell.execute_reply":"2022-11-06T08:03:57.229616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Some of the sessions might not have any clicks, carts or orders but the statistics below doesn't include their missing event types. The highlight is, some of the sessions in test set can only have clicks, carts or orders, but training set sessions are more heterogeneous. Homogeneous sessions in training set are click-only sessions. This discrepancy could be a problem during modelling.","metadata":{}},{"cell_type":"code","source":"df_train_sessions_type_rate_mins = df_train_sessions_types.groupby('type')['rate'].min().to_dict()\ndf_train_sessions_type_rate_maxs = df_train_sessions_types.groupby('type')['rate'].max().to_dict()\ndf_test_sessions_type_rate_mins = df_test_sessions_types.groupby('type')['rate'].min().to_dict()\ndf_test_sessions_type_rate_maxs = df_test_sessions_types.groupby('type')['rate'].max().to_dict()\n\nprint(\nf'''\nSession Event Type Least and Most Percentages\nTraining: 0: {df_train_sessions_type_rate_mins[0]:.4f}-{df_train_sessions_type_rate_maxs[0]:.4f} | 1: {df_train_sessions_type_rate_mins[1]:.4f}-{df_train_sessions_type_rate_maxs[1]:.4f} | 2: {df_train_sessions_type_rate_mins[2]:.4f}-{df_train_sessions_type_rate_maxs[2]:.4f}\nTest: 0: {df_test_sessions_type_rate_mins[0]:.4f}-{df_test_sessions_type_rate_maxs[0]:.4f} | 1: {df_test_sessions_type_rate_mins[1]:.4f}-{df_test_sessions_type_rate_maxs[1]:.4f} | 2: {df_test_sessions_type_rate_mins[2]:.4f}-{df_test_sessions_type_rate_maxs[2]:.4f}\n'''\n)\ndel df_train_sessions_type_rate_mins, df_train_sessions_type_rate_maxs, df_test_sessions_type_rate_mins, df_test_sessions_type_rate_maxs","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-11-06T08:03:57.233187Z","iopub.execute_input":"2022-11-06T08:03:57.233697Z","iopub.status.idle":"2022-11-06T08:03:58.054183Z","shell.execute_reply.started":"2022-11-06T08:03:57.233649Z","shell.execute_reply":"2022-11-06T08:03:58.052999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Another important thing to consider is how sessions start and end. There are two interesting things here. First, sessions are supposed to start with clicks but very few of them start with carts or orders. Those sessions are less than 1% and they might be truncated from their beginning so they could fit into the selected time frame. The other thing is, sessions in test set are less likely to end with an order because of the truncation. Truncated timesteps in test set are more likely order events.  ","metadata":{}},{"cell_type":"code","source":"df_train_session_firsts = (df_train.groupby('session')['type'].first().value_counts() / df_train['session'].nunique()).to_dict()\ndf_train_session_lasts = (df_train.groupby('session')['type'].last().value_counts() / df_train['session'].nunique()).to_dict()\ndf_test_session_firsts = (df_test.groupby('session')['type'].first().value_counts() / df_test['session'].nunique()).to_dict()\ndf_test_session_lasts = (df_test.groupby('session')['type'].last().value_counts() / df_test['session'].nunique()).to_dict()\n\nprint(\nf'''\nSession Event Type First and Last Percentages\nTraining: 0: {df_train_session_firsts[0]:.4f}-{df_train_session_lasts[0]:.4f} | 1: {df_train_session_firsts[1]:.4f}-{df_train_session_lasts[1]:.4f} | 2: {df_train_session_firsts[2]:.4f}-{df_train_session_lasts[2]:.4f}\nTest: 0: {df_test_session_firsts[0]:.4f}-{df_test_session_lasts[0]:.4f} | 1: {df_test_session_firsts[1]:.4f}-{df_test_session_lasts[1]:.4f} | 2: {df_test_session_firsts[2]:.4f}-{df_test_session_lasts[2]:.4f}\n'''\n)\ndel df_train_session_firsts, df_train_session_lasts, df_test_session_firsts, df_test_session_lasts","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-11-06T08:28:53.315365Z","iopub.execute_input":"2022-11-06T08:28:53.315875Z","iopub.status.idle":"2022-11-06T08:29:18.785329Z","shell.execute_reply.started":"2022-11-06T08:28:53.315838Z","shell.execute_reply":"2022-11-06T08:29:18.784061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A natural sequence is clicking to a product, adding it to the cart and ordering it. It can be seen from the visualization of session 3 below that cart events lead to order events multiple times. However, this doesn't apply to all sessions.","metadata":{}},{"cell_type":"code","source":"def visualize_session(df, session, path=None):\n    \n    \"\"\"\n    Visualize given session in given dataframe as a sequence\n    \n    Parameters\n    ----------\n    df: pandas.DataFrame of shape (n_samples, 4)\n        Training or test dataframe with session, aid, ts, type columns\n\n    feature: int\n        Unique ID of the session\n    \n    path: path-like str or None\n        Path of the output file or None (if path is None, plot is displayed with selected backend)\n    \"\"\"\n    \n    df_session = df.loc[df['session'] == session, :]\n    \n    fig, ax = plt.subplots(figsize=(24, 6))\n    ax.plot(df_session.set_index('ts')['type'], 'o-')\n    ax.set_yticks(range(3), ['Click (0)', 'Cart (1)', 'Order (2)'])\n    ax.tick_params(axis='x', labelsize=12.5)\n    ax.tick_params(axis='y', labelsize=12.5)\n    ax.set_xlabel('Timestamps', fontsize=15, labelpad=12.5)\n    ax.set_ylabel('Events', fontsize=15, labelpad=12.5)\n    title = f'''\n    Session: {session}\n    Clicks: {(df_session['type'] == 0).sum()} |  Carts: {(df_session['type'] == 1).sum()} | Orders: {(df_session['type'] == 2).sum()}\n    '''\n    ax.set_title(title, size=20, pad=15)\n    \n    if path is None:\n        plt.show()\n    else:\n        plt.savefig(path)\n        plt.close(fig)\n\n\nvisualize_session(df=df_train, session=3)","metadata":{"execution":{"iopub.status.busy":"2022-11-06T07:58:55.878446Z","iopub.execute_input":"2022-11-06T07:58:55.878922Z","iopub.status.idle":"2022-11-06T07:58:56.342409Z","shell.execute_reply.started":"2022-11-06T07:58:55.878890Z","shell.execute_reply":"2022-11-06T07:58:56.341243Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In training set, session 39 starts with cart event and session 747 starts with order event which are anomalies mentioned before.","metadata":{}},{"cell_type":"code","source":"visualize_session(df=df_train, session=39)\nvisualize_session(df=df_train, session=747)","metadata":{"execution":{"iopub.status.busy":"2022-11-06T11:59:55.638980Z","iopub.execute_input":"2022-11-06T11:59:55.639414Z","iopub.status.idle":"2022-11-06T11:59:56.609059Z","shell.execute_reply.started":"2022-11-06T11:59:55.639370Z","shell.execute_reply":"2022-11-06T11:59:56.608023Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 5. Ground-truth\n\nGround-truth of both training and test set can be created as long as a session has more than 1 timesteps. Ground-truth of clicks and other events are created differently. Ground-truth click of a timestep is the next click in that session. Ground-truth of carts or orders of a timestep is the collection of all next unique carts or orders in that session.","metadata":{}},{"cell_type":"code","source":"def get_labels(aids, event_types):\n    \n    \"\"\"\n    Create ground-truth labels from given session aids and event types\n    \n    Parameters\n    ----------\n    aids: pandas.Series of shape (n_events)\n        Session aids\n\n    event_types: pandas.Series of shape (n_events)\n        Session event types\n        \n    Returns\n    -------\n    labels: list of shape (n_events)\n        Ground-truth labels\n    \"\"\"\n    \n    previous_click = None\n    previous_carts = set()\n    previous_orders = set()\n    labels = []\n\n    for aid, event_type in zip(reversed(aids.values), reversed(event_types.values)):\n        \n        label = {}\n        \n        if event_type == 0:\n            previous_click = aid\n        elif event_type == 1:\n            previous_carts.add(aid)\n        elif event_type == 2:\n            previous_orders.add(aid)\n            \n        label[0] = previous_click \n        label[1] = previous_carts.copy() if len(previous_carts) > 0 else np.nan\n        label[2] = previous_orders.copy() if len(previous_orders) > 0 else np.nan\n        labels.append(label)\n        \n    labels = labels[:-1][::-1]\n    labels.append({0: np.nan, 1: np.nan, 2: np.nan})\n    \n    return labels\n","metadata":{"execution":{"iopub.status.busy":"2022-11-12T06:54:49.332347Z","iopub.execute_input":"2022-11-12T06:54:49.333701Z","iopub.status.idle":"2022-11-12T06:54:49.348083Z","shell.execute_reply.started":"2022-11-12T06:54:49.333654Z","shell.execute_reply":"2022-11-12T06:54:49.346367Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Labels of session 747 are extracted with the function above. It can be seen that after a cart or order event, corresponding aid is dropped from the labels.","metadata":{}},{"cell_type":"code","source":"df_session747 = df_train.loc[df_train['session'] == 747, :]\nsession747_labels = get_labels(aids=df_session747['aid'], event_types=df_session747['type'])\ndf_session747.loc[:, 'label'] = session747_labels\ndf_session747","metadata":{"execution":{"iopub.status.busy":"2022-11-12T06:54:50.909331Z","iopub.execute_input":"2022-11-12T06:54:50.909742Z","iopub.status.idle":"2022-11-12T06:54:51.992960Z","shell.execute_reply.started":"2022-11-12T06:54:50.909710Z","shell.execute_reply":"2022-11-12T06:54:51.991740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 6. Evaluation\n\n**Precision (positive predictive value)** and **recall (true positive rate)** are most commonly used binary and multi-class classification metrics. Precision formula is TP / (TP + FP) and recall formula is TP / (TP + FN). Predicting less false positives is important for improving precision and predicting more true positives is important for recall. There is always a trade-off between those metrics so harmonic mean of them (F1 score) is also used conjunctly with them.\n\nThose metrics are also used in information retrieval problems. For this problem, predictions are evaluated on recall@20 for each event type and 3 recall values are multiplied with their specified weights. Event type weights are; 0.1 for clicks, 0.3 for carts and 0.6 for orders. Recall makes more sense in this context because false positives aren't important as much as false negatives.\n\n![recall@20](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/recall.png)\n\nSessions in test set are truncated by their last events and aids of those truncated events are expected to be predicted. Regardless of the truncated event's type, 20 predictions can be made for clicks, carts and orders. Recalls are calculated separately on those 20 predictions and weighted based on the event type. Finally, 3 recall values are averaged.\n\nFor example; let's assume that session 747 is truncated at index 59605 and we are trying to predict aids of that event. ","metadata":{}},{"cell_type":"code","source":"df_session747.loc[:59605]","metadata":{"execution":{"iopub.status.busy":"2022-11-12T07:44:37.470753Z","iopub.execute_input":"2022-11-12T07:44:37.471152Z","iopub.status.idle":"2022-11-12T07:44:37.494188Z","shell.execute_reply.started":"2022-11-12T07:44:37.471117Z","shell.execute_reply":"2022-11-12T07:44:37.492522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since there is only one ground-truth for clicks, recall can be either 0 or 1. This can be evaluated by simply using an `in` operation. If ground-truth click aid is in click predictions, then recall is 1. Otherwise, recall is 0. After multiplying the click recall with 0.1, it can be either 0 or 0.1.","metadata":{}},{"cell_type":"code","source":"def click_recall(y_true, y_pred):\n    \n    \"\"\"\n    Calculate recall for clicks on ground-truth and predictions\n    \n    Parameters\n    ----------\n    y_true: int\n        Ground-truth click aid\n\n    y_pred: array-like of shape (n_aids) (1 <= n_aids <= 20)\n        Prediction click aids\n        \n    Returns\n    -------\n    recall: int\n        Recall calculated on ground-truth and predictions\n    \"\"\"\n    \n    recall = int(y_true in y_pred)\n    return recall\n\n\nsession747_click_recall = click_recall(df_session747.loc[59605, 'label'][0] , df_session747.loc[:59605, 'aid'].unique())\nprint(f'Click Recall of Session 747: {session747_click_recall}')","metadata":{"execution":{"iopub.status.busy":"2022-11-12T08:34:57.530118Z","iopub.execute_input":"2022-11-12T08:34:57.530792Z","iopub.status.idle":"2022-11-12T08:34:57.541187Z","shell.execute_reply.started":"2022-11-12T08:34:57.530745Z","shell.execute_reply":"2022-11-12T08:34:57.539584Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For cart and order event types, there are multiple ground-truth aids so recall can be any value between 0 and 1. Recall of those event types can be calculated with set operations. True positives are the intersection of ground-truth and prediction aids and false negatives are ground-truth - prediction aids.\n\nIn some rare cases, number of ground-truth aids can be greater than 20. In order to prevent calculation errors for those cases, denominator should be min(20, n_aids).","metadata":{}},{"cell_type":"code","source":"def cart_order_recall(y_true, y_pred):\n    \n    \"\"\"\n    Calculate recall for carts/orders on ground-truth and predictions\n    \n    Parameters\n    ----------\n    y_true: array-like of shape (n_aids) (1 <= n_aids)\n        Ground-truth cart/order aids\n\n    y_pred: array-like of shape (n_aids) (1 <= n_aids <= 20)\n        Prediction cart/order aids\n        \n    Returns\n    -------\n    recall: int\n        Recall calculated on ground-truth and predictions\n    \"\"\"\n    \n    y_true = set(y_true)\n    y_pred = set(y_pred)\n    \n    tp = len(y_true.intersection(y_pred))\n    fn = len(y_true - y_pred)\n    \n    recall = tp / min(20, (tp + fn))\n    return recall\n\nsession747_cart_recall = cart_order_recall(df_session747.loc[59605, 'label'][1] , df_session747.loc[:59605, 'aid'].unique())\nprint(f'Cart Recall of Session 747: {session747_cart_recall}')\nsession747_order_recall = cart_order_recall(df_session747.loc[59605, 'label'][2] , df_session747.loc[:59605, 'aid'].unique())\nprint(f'Order Recall of Session 747: {session747_order_recall}')","metadata":{"execution":{"iopub.status.busy":"2022-11-12T08:26:57.962702Z","iopub.execute_input":"2022-11-12T08:26:57.963141Z","iopub.status.idle":"2022-11-12T08:26:57.972485Z","shell.execute_reply.started":"2022-11-12T08:26:57.963105Z","shell.execute_reply":"2022-11-12T08:26:57.971249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After multiplying event type recalls with their weights, recall of session 747 becomes 0.7. Recalls are calculated for each session's truncated last event and averaged which would be the final score.","metadata":{}},{"cell_type":"code","source":"session_747_recall = (session747_click_recall * 0.1) + (session747_cart_recall * 0.3) + (session747_order_recall * 0.6)\nprint(f'Recall of Session 747: {session_747_recall}')","metadata":{"execution":{"iopub.status.busy":"2022-11-12T08:33:08.108221Z","iopub.execute_input":"2022-11-12T08:33:08.108636Z","iopub.status.idle":"2022-11-12T08:33:08.115785Z","shell.execute_reply.started":"2022-11-12T08:33:08.108603Z","shell.execute_reply":"2022-11-12T08:33:08.113878Z"},"trusted":true},"execution_count":null,"outputs":[]}]}