{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<img src=\"https://i.imgur.com/MLND8w0.png\">\n\n<center><h1> I was warned </h1></center>\n<center><h2> - but didn't listen 💧 - </h2></center>\n\n> 📌 **Scope**: The [goal](https://www.kaggle.com/competitions/otto-recommender-system/data) of this competition is to predict e-commerce *clicks, cart additions, and orders*. You'll build a **multi-objective recommender system** based on previous events in a user session.\n\nIf you still are confused about the scope, I asked a [dumb question](https://www.kaggle.com/competitions/otto-recommender-system/discussion/367154) for myself too, so maybe this one will clear the air a bit more.\n\n**Dataset**: The data I'm using is from [Radek's dataset here](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint).\n\n**Additional Note**: As the data is 200 mil rows +, I did the preprocessing locally on my computer, so I will be also using my own dataset, but the original is still from Radek. :)\n\n### ○ Libraries","metadata":{}},{"cell_type":"code","source":"# General Libraries\nimport os\nimport re\nimport gc\nimport wandb\nimport random\nimport math\nfrom tqdm import tqdm\nfrom pprint import pprint\nfrom time import time\nfrom datetime import datetime\nimport itertools\nimport warnings\nimport pandas as pd\nimport numpy as np\n\n# For the Visuals\nimport seaborn as sns\nimport matplotlib as mpl\nfrom matplotlib import cm\nimport matplotlib.patches as patches\nimport matplotlib.pyplot as plt\nimport matplotlib.image as mpimg\nfrom matplotlib.offsetbox import AnnotationBbox, OffsetImage\nfrom matplotlib.colors import ListedColormap, LinearSegmentedColormap\nfrom matplotlib.patches import Rectangle\nfrom IPython.display import display_html\nplt.rcParams.update({'font.size': 16})\n\n# RAPIDS\nimport cudf\nimport cupy\n\n# Environment check\nwarnings.filterwarnings(\"ignore\")\nos.environ[\"WANDB_SILENT\"] = \"true\"\nCONFIG = {'competition': 'Otto', '_wandb_kernel': 'aot'}\n\n# Custom colors\nclass clr:\n    S = '\\033[1m' + '\\033[91m'\n    E = '\\033[0m'\n    \nmy_colors = [\"#f3afc2\", \"#a86c4a\", \"#7f5c10\", \"#d79a7b\", \n             \"#ab883e\", \"#7a7300\", \"#004c00\"]\nCMAP1 = ListedColormap(my_colors)\n\nprint(clr.S+\"Notebook Color Schemes:\"+clr.E)\nsns.palplot(sns.color_palette(my_colors))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:41:42.760923Z","iopub.execute_input":"2022-12-01T11:41:42.761486Z","iopub.status.idle":"2022-12-01T11:41:48.674038Z","shell.execute_reply.started":"2022-12-01T11:41:42.761326Z","shell.execute_reply":"2022-12-01T11:41:48.672440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 🐝 W&B Fork & Run\n\nIn order to run this notebook you will need to input your own **secret API key** within the `! wandb login $secret_value_0` line. \n\n🐝**How do you get your own API key?**\n\nSuper simple! Go to **https://wandb.ai/site** -> Login -> Click on your profile in the top right corner -> Settings -> Scroll down to API keys -> copy your very own key (for more info check [this amazing notebook for ML Experiment Tracking on Kaggle](https://www.kaggle.com/ayuraj/experiment-tracking-with-weights-and-biases)).\n\n<center><img src=\"https://i.imgur.com/fFccmoS.png\" width=500></center>","metadata":{}},{"cell_type":"code","source":"# 🐝 Secrets\nfrom kaggle_secrets import UserSecretsClient\nuser_secrets = UserSecretsClient()\nsecret_value_0 = user_secrets.get_secret(\"wandb\")\n\n! wandb login $secret_value_0","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:41:50.649799Z","iopub.execute_input":"2022-12-01T11:41:50.650467Z","iopub.status.idle":"2022-12-01T11:41:54.383483Z","shell.execute_reply.started":"2022-12-01T11:41:50.650409Z","shell.execute_reply":"2022-12-01T11:41:54.380935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### ○ Helpers","metadata":{}},{"cell_type":"code","source":"# === General Functions ===\n\ndef show_values_on_bars(axs, h_v=\"v\", space=0.4):\n    '''Plots the value at the end of the a seaborn barplot.\n    axs: the ax of the plot\n    h_v: weather or not the barplot is vertical/ horizontal'''\n    \n    def _show_on_single_plot(ax):\n        if h_v == \"v\":\n            for p in ax.patches:\n                _x = p.get_x() + p.get_width() / 2\n                _y = p.get_y() + p.get_height()\n                value = int(p.get_height())\n                ax.text(_x, _y, format(value, ','), ha=\"center\") \n        elif h_v == \"h\":\n            for p in ax.patches:\n                _x = p.get_x() + p.get_width() + float(space)\n                _y = p.get_y() + p.get_height()\n                value = int(p.get_width())\n                ax.text(_x, _y, format(value, ','), ha=\"left\")\n\n    if isinstance(axs, np.ndarray):\n        for idx, ax in np.ndenumerate(axs):\n            _show_on_single_plot(ax)\n    else:\n        _show_on_single_plot(axs)\n        \n        \n# === 🐝 W&B ===\ndef save_dataset_artifact(run_name, artifact_name, path, data_type=\"dataset\"):\n    '''Saves dataset to W&B Artifactory.\n    run_name: name of the experiment\n    artifact_name: under what name should the dataset be stored\n    path: path to the dataset'''\n    \n    run = wandb.init(project='Otto', \n                     name=run_name, \n                     config=CONFIG)\n    artifact = wandb.Artifact(name=artifact_name, \n                              type=data_type)\n    artifact.add_file(path)\n\n    wandb.log_artifact(artifact)\n    wandb.finish()\n    print(\"Artifact has been saved successfully.\")\n    \n    \ndef create_wandb_plot(x_data=None, y_data=None, x_name=None, y_name=None, title=None, log=None, plot=\"line\"):\n    '''Create and save lineplot/barplot in W&B Environment.\n    x_data & y_data: Pandas Series containing x & y data\n    x_name & y_name: strings containing axis names\n    title: title of the graph\n    log: string containing name of log'''\n    \n    data = [[label, val] for (label, val) in zip(x_data, y_data)]\n    table = wandb.Table(data=data, columns = [x_name, y_name])\n    \n    if plot == \"line\":\n        wandb.log({log : wandb.plot.line(table, x_name, y_name, title=title)})\n    elif plot == \"bar\":\n        wandb.log({log : wandb.plot.bar(table, x_name, y_name, title=title)})\n    elif plot == \"scatter\":\n        wandb.log({log : wandb.plot.scatter(table, x_name, y_name, title=title)})\n        \n        \ndef create_wandb_hist(x_data=None, x_name=None, title=None, log=None):\n    '''Create and save histogram in W&B Environment.\n    x_data: Pandas Series containing x values\n    x_name: strings containing axis name\n    title: title of the graph\n    log: string containing name of log'''\n    \n    data = [[x] for x in x_data]\n    table = wandb.Table(data=data, columns=[x_name])\n    wandb.log({log : wandb.plot.histogram(table, x_name, title=title)})","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true},"execution":{"iopub.status.busy":"2022-12-01T11:41:54.387038Z","iopub.execute_input":"2022-12-01T11:41:54.387549Z","iopub.status.idle":"2022-12-01T11:41:54.413249Z","shell.execute_reply.started":"2022-12-01T11:41:54.387502Z","shell.execute_reply":"2022-12-01T11:41:54.409196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 🐝 Bonus: Cover Photo\nrun = wandb.init(project='Otto', name='CoverPhoto', config=CONFIG)\ncover = plt.imread(\"../input/otto-helper-data/recsys_cover.png\")\nwandb.log({\"cover\": wandb.Image(cover)})\nwandb.finish()","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:41:54.416077Z","iopub.execute_input":"2022-12-01T11:41:54.416710Z","iopub.status.idle":"2022-12-01T11:43:01.611723Z","shell.execute_reply.started":"2022-12-01T11:41:54.416659Z","shell.execute_reply":"2022-12-01T11:43:01.610304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 🐝 Experiment\nrun = wandb.init(project='Otto', name='base_info', config=CONFIG)","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:43:01.615514Z","iopub.execute_input":"2022-12-01T11:43:01.616405Z","iopub.status.idle":"2022-12-01T11:43:11.184147Z","shell.execute_reply.started":"2022-12-01T11:43:01.616320Z","shell.execute_reply":"2022-12-01T11:43:11.182583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 216 mil rows in train data\ntrain = cudf.read_parquet(\"/kaggle/input/otto-full-optimized-memory-footprint/train.parquet\")\ntest = cudf.read_parquet(\"/kaggle/input/otto-full-optimized-memory-footprint/test.parquet\")\n\nprint(clr.S+\"Train Shape:\"+clr.E, train.shape)\nprint(clr.S+\"Test Shape:\"+clr.E, test.shape)","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:43:11.191409Z","iopub.execute_input":"2022-12-01T11:43:11.196124Z","iopub.status.idle":"2022-12-01T11:43:35.665905Z","shell.execute_reply.started":"2022-12-01T11:43:11.196046Z","shell.execute_reply":"2022-12-01T11:43:35.664467Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wandb.log({\"train_len\":train.shape[0],\n           \"test_len\": test.shape[0]})","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:43:35.672901Z","iopub.execute_input":"2022-12-01T11:43:35.677960Z","iopub.status.idle":"2022-12-01T11:43:38.219033Z","shell.execute_reply.started":"2022-12-01T11:43:35.677888Z","shell.execute_reply":"2022-12-01T11:43:38.217639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:43:38.222398Z","iopub.execute_input":"2022-12-01T11:43:38.228998Z","iopub.status.idle":"2022-12-01T11:43:40.064663Z","shell.execute_reply.started":"2022-12-01T11:43:38.228926Z","shell.execute_reply":"2022-12-01T11:43:40.063217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:43:40.071613Z","iopub.execute_input":"2022-12-01T11:43:40.076562Z","iopub.status.idle":"2022-12-01T11:43:41.803312Z","shell.execute_reply.started":"2022-12-01T11:43:40.076472Z","shell.execute_reply":"2022-12-01T11:43:41.801851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wandb.finish()","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:43:41.810662Z","iopub.execute_input":"2022-12-01T11:43:41.814780Z","iopub.status.idle":"2022-12-01T11:44:34.046211Z","shell.execute_reply.started":"2022-12-01T11:43:41.814695Z","shell.execute_reply":"2022-12-01T11:44:34.044448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1.Sessions x Events\n\n> 📌**Note**: There is no intersection in the sessions from `train` and `test` datasets. All sessions in the `test` dataset are unique (and truncated in time).\n\n*FYI, as I've seen in Discussions, a session is a user. So each unique session is actually a unique user - meaning there are no 2 sessions from 1 person in time.*","metadata":{}},{"cell_type":"code","source":"# 🐝 Experiment\nrun = wandb.init(project='Otto', name='session_look', config=CONFIG)","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:44:34.052852Z","iopub.execute_input":"2022-12-01T11:44:34.053298Z","iopub.status.idle":"2022-12-01T11:44:44.505607Z","shell.execute_reply.started":"2022-12-01T11:44:34.053260Z","shell.execute_reply":"2022-12-01T11:44:44.504022Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Intersection between sessions in train vs test = 0\n# which means all users in test data are NEW\nset(train[\"session\"].to_pandas()).intersection(set(test[\"session\"].to_pandas()))","metadata":{"execution":{"iopub.status.busy":"2022-11-25T19:33:54.684167Z","iopub.execute_input":"2022-11-25T19:33:54.686869Z","iopub.status.idle":"2022-11-25T19:34:42.054045Z","shell.execute_reply.started":"2022-11-25T19:33:54.686822Z","shell.execute_reply":"2022-11-25T19:34:42.053004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_describe(x):\n    '''Get .describe() info.'''\n    return {\"mean\": x.mean(),\n            \"min\": x.min(),\n            \"max\": x.max(),\n            \"count\": len(x)}\n\nsess_tr = train[\"session\"].value_counts().values\nsess_te = test[\"session\"].value_counts().values","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:44:44.514101Z","iopub.execute_input":"2022-12-01T11:44:44.520469Z","iopub.status.idle":"2022-12-01T11:44:46.873602Z","shell.execute_reply.started":"2022-12-01T11:44:44.520378Z","shell.execute_reply":"2022-12-01T11:44:46.871273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(clr.S+\"=== Train ===\"+clr.E)\nprint(clr.S+\"Event Number Summary per Session\"+clr.E)\nprint(get_describe(sess_tr.get()))\nwandb.log(get_describe(sess_tr.get()))\n\nprint(clr.S+\"=== Test ===\"+clr.E)\nprint(clr.S+\"Event Number Summary per Session\"+clr.E)\nprint(get_describe(sess_te.get()))","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:32:26.559882Z","iopub.execute_input":"2022-12-01T11:32:26.560455Z","iopub.status.idle":"2022-12-01T11:32:28.847263Z","shell.execute_reply.started":"2022-12-01T11:32:26.560376Z","shell.execute_reply":"2022-12-01T11:32:28.840866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(clr.S+\"- Percent of Sessions in Terms of Event Numbers -\"+clr.E)\n\nprint(clr.S+\"=== Train ===\"+clr.E)\nprint(\"<10 events:\", sum(sess_tr<10)/get_describe(sess_tr)[\"count\"], \"\\n\"+\n      \"<50 events:\", sum(sess_tr<50)/get_describe(sess_tr)[\"count\"], \"\\n\"+\n      \"<150 events:\", sum(sess_tr<150)/get_describe(sess_tr)[\"count\"], \"\\n\"+\n      \"150 events or more:\", sum(sess_tr>=150)/get_describe(sess_tr)[\"count\"])\n\nprint(clr.S+\"=== Test ===\"+clr.E)\nprint(\"<10 events:\", sum(sess_te<10)/get_describe(sess_te)[\"count\"], \"\\n\"+\n      \"<50 events:\", sum(sess_te<50)/get_describe(sess_te)[\"count\"], \"\\n\"+\n      \"<150 events:\", sum(sess_te<150)/get_describe(sess_te)[\"count\"], \"\\n\"+\n      \"150 events or more:\", sum(sess_te>=150)/get_describe(sess_te)[\"count\"])","metadata":{"execution":{"iopub.status.busy":"2022-11-24T16:16:44.795922Z","iopub.execute_input":"2022-11-24T16:16:44.796298Z","iopub.status.idle":"2022-11-24T16:46:09.885102Z","shell.execute_reply.started":"2022-11-24T16:16:44.796266Z","shell.execute_reply":"2022-11-24T16:46:09.883996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Train Overview**:\n* There are 12,899,779 unique sessions (a session ~ a user) in the `train` data\n* The *average* number of events (clicks, added to cart, ordered) is ~17\n* ~63% of sessions have *less than 10 events*\n* ~92% of sessions have *less than 50 events*\n* ~98% of sessions have *less than 150 events*\n* ~2% of sessions have *very high engagement* aka events (more than 150 per session)","metadata":{}},{"cell_type":"code","source":"sess = sess_tr[::60].copy()\nf, (a0, a1) = plt.subplots(2, 1, gridspec_kw={'height_ratios': [3, 1]}, figsize=(24, 15))\nsns.distplot(sess.get(), rug=True, hist=False, \n#              bins=10,\n             rug_kws={\"color\": my_colors[5]},\n             kde_kws={\"color\": my_colors[5], \"lw\": 5, \"alpha\": 0.7},\n#              hist_kws={\"histtype\": \"step\", \"linewidth\": 3, \"alpha\": 1, \"color\": my_colors[5]},\n             ax=a0)\n\na0.axvline(x=10, ls=\":\", lw=2, color=\"#707B7C\")\na0.text(x=-40, y=0.075, s=\"63% of data\", size=17, color=\"#707B7C\", weight=\"bold\")\na0.text(x=-30, y=0.072, s=\"is here\", size=17, color=\"#707B7C\",weight=\"bold\")\na0.axvline(x=50, ls=\"--\", lw=2, color=\"#424949\")\na0.text(x=-30, y=0.05, s=\"92% of data is here\", size=17, color=\"#424949\", weight=\"bold\")\na0.axvline(x=150, ls=\"-\", lw=2, color=\"black\")\na0.text(x=-15, y=0.03, s=\"98%    of    data    is    here\", size=17, weight=\"bold\")\na0.text(x=250, y=0.01, s=\"2% of session have 150+ events.\", size=17,weight=\"bold\")\n\nsns.boxplot(x=sess.get(), ax=a1, notch=False, showcaps=True, \n            flierprops={\"marker\": \"x\"},\n            boxprops={\"facecolor\": my_colors[5]},\n            medianprops={\"color\": my_colors[0]},)\n\nplt.suptitle(\"TRAIN: Distribution of Number of Events per Session\", weight=\"bold\", size=25)\nsns.despine(right=True, top=True, left=True);","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:44:46.875958Z","iopub.execute_input":"2022-12-01T11:44:46.876594Z","iopub.status.idle":"2022-12-01T11:45:01.928836Z","shell.execute_reply.started":"2022-12-01T11:44:46.876543Z","shell.execute_reply":"2022-12-01T11:45:01.927097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"create_wandb_hist(x_data=sess.tolist(), \n                  x_name=\"Number of Events / Session\",\n                  title=\"TRAIN: Distribution of Number of Events per Session\",\n                  log=\"hist_sess_train\")","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:45:01.930840Z","iopub.execute_input":"2022-12-01T11:45:01.931381Z","iopub.status.idle":"2022-12-01T11:45:19.989376Z","shell.execute_reply.started":"2022-12-01T11:45:01.931306Z","shell.execute_reply":"2022-12-01T11:45:19.987780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Test Overview**:\n* There are 1,671,803 unique sessions (a session ~ a user)\n* The session length is much smaller, with an average of 4 events (~3 times smaller than for `train`)\n* There are some longer sessions that should not be disregarded? ***","metadata":{}},{"cell_type":"code","source":"sess = sess_te[::20].copy()\nf, (a0, a1) = plt.subplots(2, 1, gridspec_kw={'height_ratios': [3, 1]}, figsize=(24, 15))\nsns.distplot(sess.get(), rug=True, hist=False, \n#              bins=10,\n             rug_kws={\"color\": my_colors[5]},\n             kde_kws={\"color\": my_colors[5], \"lw\": 5, \"alpha\": 0.7},\n#              hist_kws={\"histtype\": \"step\", \"linewidth\": 3, \"alpha\": 1, \"color\": my_colors[5]},\n             ax=a0)\n\na0.axvline(x=50, ls=\"--\", lw=2, color=\"black\")\na0.text(x=-22, y=0.23, s=\"99% of data is here\", size=17, color=\"black\", weight=\"bold\")\na0.text(x=60, y=0.175, s=\"Actually, 90%+ session have less than 10 events\", size=17, \n        color=\"black\", weight=\"bold\")\na0.text(x=60, y=0.165, s=\"(that need to be predicted)\", size=17, color=\"black\", weight=\"bold\")\n\nsns.boxplot(x=sess.get(), ax=a1, notch=False, showcaps=True, \n            flierprops={\"marker\": \"x\"},\n            boxprops={\"facecolor\": my_colors[5]},\n            medianprops={\"color\": my_colors[0]},)\n\nplt.suptitle(\"TEST: Distribution of Number of Events per Session\", weight=\"bold\", size=25)\nsns.despine(right=True, top=True, left=True);","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-12-01T11:45:19.997776Z","iopub.execute_input":"2022-12-01T11:45:20.002107Z","iopub.status.idle":"2022-12-01T11:45:25.691592Z","shell.execute_reply.started":"2022-12-01T11:45:20.002036Z","shell.execute_reply":"2022-12-01T11:45:25.690147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"create_wandb_hist(x_data=sess.tolist(), \n                  x_name=\"Number of Events / Session\",\n                  title=\"TEST: Distribution of Number of Events per Session\",\n                  log=\"hist_sess_test\")","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:45:25.698786Z","iopub.execute_input":"2022-12-01T11:45:25.702918Z","iopub.status.idle":"2022-12-01T11:45:33.522809Z","shell.execute_reply.started":"2022-12-01T11:45:25.702845Z","shell.execute_reply":"2022-12-01T11:45:33.521239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wandb.finish()","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:45:33.530013Z","iopub.execute_input":"2022-12-01T11:45:33.534019Z","iopub.status.idle":"2022-12-01T11:46:45.157004Z","shell.execute_reply.started":"2022-12-01T11:45:33.533948Z","shell.execute_reply":"2022-12-01T11:46:45.155128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Types of Events\n\n📍 **3 Types of Events** (encoded numericaly for memory convenience): \n* 0: clicks\n* 1: carts (added to cart)\n* 2: orders","metadata":{}},{"cell_type":"code","source":"# 🐝 Experiment\nrun = wandb.init(project='Otto', name='event_look', config=CONFIG)","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:46:45.159098Z","iopub.execute_input":"2022-12-01T11:46:45.159642Z","iopub.status.idle":"2022-12-01T11:46:54.537131Z","shell.execute_reply.started":"2022-12-01T11:46:45.159592Z","shell.execute_reply":"2022-12-01T11:46:54.535549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# See percentage of events as clicks, carts, orders\ntypes = train[\"type\"].value_counts().reset_index()\ntypes.columns = [\"type\", \"train\"]\ntypes[\"test\"] = test[\"type\"].value_counts().values\n\ntotal_tr = types[\"train\"].sum()\ntotal_te = types[\"test\"].sum()\ntypes[\"train\"] = (types[\"train\"]/total_tr)*100\ntypes[\"test\"] = (types[\"test\"]/total_te)*100\n\ntypes = types.melt(id_vars=[\"type\"], \n                   var_name=\"Data\", \n                   value_name=\"Count\")\n\ntypes","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:46:54.545069Z","iopub.execute_input":"2022-12-01T11:46:54.549166Z","iopub.status.idle":"2022-12-01T11:46:58.064608Z","shell.execute_reply.started":"2022-12-01T11:46:54.549087Z","shell.execute_reply":"2022-12-01T11:46:58.063123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# See how many sessions contain distinct events\ndistincts = train.groupby(\"session\")[\"type\"].nunique()\\\n                .reset_index()[\"type\"].value_counts().reset_index()\ndistincts.columns = [\"unique_events\", \"train\"]\ndistincts[\"test\"] = test.groupby(\"session\")[\"type\"].nunique()\\\n                        .reset_index()[\"type\"].value_counts().values\n\ntotal_tr = distincts[\"train\"].sum()\ntotal_te = distincts[\"test\"].sum()\ndistincts[\"train\"] = (distincts[\"train\"]/total_tr)*100\ndistincts[\"test\"] = (distincts[\"test\"]/total_te)*100\n\ndistincts = distincts.sort_values(\"unique_events\").reset_index(drop=True)\n\ndistincts = distincts.melt(id_vars=[\"unique_events\"], \n                           var_name=\"Data\", \n                           value_name=\"Count\")\n\ndistincts","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:46:58.072768Z","iopub.execute_input":"2022-12-01T11:46:58.077495Z","iopub.status.idle":"2022-12-01T11:47:00.774777Z","shell.execute_reply.started":"2022-12-01T11:46:58.077421Z","shell.execute_reply":"2022-12-01T11:47:00.773322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Event Types Analysis**:\n* `train` and `test` have around the same distribution\n* 90% of events are clicks, ~7% are \"add to carts\" and 2% are actual \"orders\". There is also a *very low percentage* of orders in `test` set.\n* unique event count:\n    * for train: \n        * 70% of sessions have only 1 event type\n        * 17% of sessions have 2 event types\n        * 12% of sessions have all 3 event types\n    * for test:\n        * 85% of sessions have only 1 event type (much higher than for train)\n        * 12% of sessions have 2 event types\n        * 1% of sessions have all 3 event types (much lower than for train - as number of orders in test is also much lower)\n* *Big difference between train and test","metadata":{}},{"cell_type":"code","source":"fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(24, 12))\naxs = [ax1, ax2]\nfig.suptitle('Event Types Analysis', \n             weight=\"bold\", size=25)\n\nsns.barplot(data=types.to_pandas(), x=\"Data\", y=\"Count\", hue=\"type\", ax=ax1,\n            palette=[my_colors[2], my_colors[3], my_colors[0]])\nshow_values_on_bars(ax1, h_v=\"v\", space=0.4)\nax1.set_title('Event Frequency per Type', weight=\"bold\", size=20)\n\nsns.barplot(data=distincts.to_pandas(), x=\"Data\", y=\"Count\", hue=\"unique_events\", ax=ax2,\n            palette=[my_colors[6], my_colors[5], my_colors[4]])\nshow_values_on_bars(ax2, h_v=\"v\", space=0.4)\nax2.set_title('Unique Event Count', weight=\"bold\", size=20)\n\nfor ax in axs:\n    ax.set_yticks([])\n    ax.set_xlabel(\"\")\n    ax.set_ylabel(\"Frequency (%)\", size = 18, weight=\"bold\")\n    \nsns.despine(right=True, top=True, left=True);","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:47:00.781889Z","iopub.execute_input":"2022-12-01T11:47:00.785924Z","iopub.status.idle":"2022-12-01T11:47:03.411862Z","shell.execute_reply.started":"2022-12-01T11:47:00.785854Z","shell.execute_reply":"2022-12-01T11:47:03.410476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now I wanted to look on session level (not overall) how do the users behave.\n* 9 mil sessions did only **1 type of activity** (either click, cart or order).\n* 2.2 mil session did **2 types of events**\n* only 1.5 sessions have **all 3 activities**","metadata":{}},{"cell_type":"code","source":"session_type = train.groupby([\"session\"])[\"type\"].nunique().reset_index()\nsession_type.columns = [\"session\", \"types_count\"]","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:47:03.420206Z","iopub.execute_input":"2022-12-01T11:47:03.424186Z","iopub.status.idle":"2022-12-01T11:47:05.942772Z","shell.execute_reply.started":"2022-12-01T11:47:03.424116Z","shell.execute_reply":"2022-12-01T11:47:05.941314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(20, 7))\n\nfigure = sns.barplot(data=session_type[\"types_count\"].value_counts().reset_index().to_pandas(),\n                     x=\"index\", y=\"types_count\", palette=my_colors)\nshow_values_on_bars(figure, h_v=\"v\", space=0.4)\nplt.title('[train] Sessions with only 1, 2 or 3 events', weight=\"bold\", size=20)\n\nplt.xlabel(\"Number of Unique Events\", size = 18, weight=\"bold\")\nplt.ylabel(\"Count\", size = 18, weight=\"bold\")\n    \nsns.despine(right=True, top=True, left=True);","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:47:05.950213Z","iopub.execute_input":"2022-12-01T11:47:05.954304Z","iopub.status.idle":"2022-12-01T11:47:08.156842Z","shell.execute_reply.started":"2022-12-01T11:47:05.954219Z","shell.execute_reply":"2022-12-01T11:47:08.155308Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"create_wandb_plot(x_data=session_type[\"types_count\"].value_counts().reset_index().to_pandas()[\"index\"],\n                  y_data=session_type[\"types_count\"].value_counts().reset_index().to_pandas()[\"types_count\"],\n                  x_name=\"Number of Unique Events\", y_name=\"Count\",\n                  title=\"[train] Sessions with only 1, 2 or 3 events\",\n                  log=\"bar_sess_type\", plot=\"bar\")","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:47:08.164935Z","iopub.execute_input":"2022-12-01T11:47:08.169034Z","iopub.status.idle":"2022-12-01T11:47:10.748702Z","shell.execute_reply.started":"2022-12-01T11:47:08.168963Z","shell.execute_reply":"2022-12-01T11:47:10.747205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Q**: Is a pattern between clicks carts orders? Is an item after a certain amount of clicks added to cart?\n\n> **Note**: unfortunatelly here I get OOM error when trying to filter the `train` dataset on only the sessions that contain at least 1 order in them. Hence, back to local work it is! The result (3 rows) of information that I got locally I just pasted after the function `get_order_behavior()` that I used on my machine.","metadata":{}},{"cell_type":"code","source":"# First, let's select only sessions that have at least 1 order in them\nsess_list = train[train[\"type\"]==2][\"session\"].unique().values.tolist()\nprint(clr.S+\"No. of Session with at least 1 order in them:\"+clr.E, len(sess_list))","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:47:10.755850Z","iopub.execute_input":"2022-12-01T11:47:10.759879Z","iopub.status.idle":"2022-12-01T11:47:12.859850Z","shell.execute_reply.started":"2022-12-01T11:47:10.759807Z","shell.execute_reply":"2022-12-01T11:47:12.858463Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# List with all aids that are present in orders\naid_list = train[train[\"type\"]==2][\"aid\"].unique().values.tolist()\nprint(clr.S+\"No. of Aids that are present in orders:\"+clr.E, len(sess_list))","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:47:12.866770Z","iopub.execute_input":"2022-12-01T11:47:12.872727Z","iopub.status.idle":"2022-12-01T11:47:14.752864Z","shell.execute_reply.started":"2022-12-01T11:47:12.872653Z","shell.execute_reply":"2022-12-01T11:47:14.751437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_order_behavior(train):\n    \n    # Select only sessions that contain at least 1 order in them\n    # and only AIDs (products) that appear in an order\n    data = train[(train[\"session\"].isin(sess_list)) & (train[\"aid\"].isin(aid_list))]\\\n                        .reset_index(drop=True)\n\n    # List of sessions that have more than 50 events\n    # as we established all these sessions are outliers and might\n    # hence we want to EXCLUDE them to not mess up our averages afterwards\n    outliers = train[\"session\"].value_counts().reset_index()\n    outliers = outliers[outliers[\"session\"]>=50][\"index\"].unique().tolist()\n    data = data[~data[\"session\"].isin(outliers)]\n\n    # On average how many clicks/ how many carts does it take to make an order?\n    final = data.groupby([\"session\", \"aid\", \"type\"])[\"night\"].count().reset_index().\\\n                groupby([\"aid\", \"type\"])[\"night\"].mean().reset_index().\\\n                groupby([\"type\"])[\"night\"].mean()\n    \n    return final\n\n# Function ran locally (Kaggle env out of memory error) - result can be seen below\nfinal = pd.DataFrame({\"type\": [0, 1, 2],\n                      \"average\": [1.573084, 1.167270, 1.051570]})","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:47:14.759700Z","iopub.execute_input":"2022-12-01T11:47:14.764151Z","iopub.status.idle":"2022-12-01T11:47:16.378254Z","shell.execute_reply.started":"2022-12-01T11:47:14.764077Z","shell.execute_reply":"2022-12-01T11:47:16.376854Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, what was the question again?\n\n**Q**: Is a pattern between clicks carts orders? Is an item after a certain amount of clicks added to cart?\n\n**A**: So, what I wanted to check with my question was if users click multiple times before they order. E.g. if you want a speciffic aid, do you go and click on it 3 times to \"make sure\" you like it before you order? The answer is ... kinda no? It takes 1.57 clicks (so I would say some users click twice, some only once) for a user to actually purchase the item. \n\n*(keeping in mind I have *removed* all sessions that had 50+ events, as these were outliers that would have messed up with my average)*","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(20, 7))\n\nfigure = sns.barplot(data=final,\n                     x=\"type\", y=\"average\", palette=my_colors[3:])\n# show_values_on_bars(figure, h_v=\"v\", space=0.4)\nplt.title('[train] Average Number of Events Until Order', weight=\"bold\", size=20)\n\nplt.xlabel(\"Event Type (clicks, carts, orders)\", size = 18, weight=\"bold\")\nplt.ylabel(\"Average\", size = 18, weight=\"bold\")\n    \nsns.despine(right=True, top=True, left=True);","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:47:16.385264Z","iopub.execute_input":"2022-12-01T11:47:16.389641Z","iopub.status.idle":"2022-12-01T11:47:18.661856Z","shell.execute_reply.started":"2022-12-01T11:47:16.389568Z","shell.execute_reply":"2022-12-01T11:47:18.659856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"create_wandb_plot(x_data=final[\"type\"],\n                  y_data=final[\"average\"],\n                  x_name=\"Event Type (clicks, carts, orders)\", y_name=\"Average\",\n                  title=\"[train] Average Number of Events Until Order\",\n                  log=\"bar_until_order\", plot=\"bar\")","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:47:18.675638Z","iopub.execute_input":"2022-12-01T11:47:18.682815Z","iopub.status.idle":"2022-12-01T11:47:21.016754Z","shell.execute_reply.started":"2022-12-01T11:47:18.682726Z","shell.execute_reply":"2022-12-01T11:47:21.013630Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> 📍 **Bonus**: As stated before, I have excluded the outliers from the data, so they don't mess up my average. However, locally, out of fun, I wanted to see what it would happen if I would keep the outliers. Would my numbers change completely? And the answer is no, the number of clicks changes only slighlty, from 1.57 to 1.70. When looking closely to the data, **there is rarely the case when an `aid` repeats more than 2-3 times per session**. This is actually very interesting (as my behavior for example is usually to recheck a product multiple times, go back and forth before buying).","metadata":{}},{"cell_type":"code","source":"wandb.finish()","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:47:21.025706Z","iopub.execute_input":"2022-12-01T11:47:21.036804Z","iopub.status.idle":"2022-12-01T11:49:07.792094Z","shell.execute_reply.started":"2022-12-01T11:47:21.036726Z","shell.execute_reply":"2022-12-01T11:49:07.790324Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del sess, train, test, session_type, aid_list\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:49:07.797188Z","iopub.execute_input":"2022-12-01T11:49:07.798320Z","iopub.status.idle":"2022-12-01T11:49:08.207245Z","shell.execute_reply.started":"2022-12-01T11:49:07.798260Z","shell.execute_reply":"2022-12-01T11:49:08.205918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. Time\n\n**Some of my questions:**\n* What is the timeline of sessions for train and test?\n* Is there less or more activity during the night?\n* How does the number of sessions group over time? (aka are they more condensed in certain periods or they are equaly distributed?)\n\n📍**Keep in mind**: as the Kaggle kernel does not have enough compute power and memory to perform the following tasks (converting Unix timestamp to datetime, getting the month, day, hour etc. to perform further analytics), I have done the following `get_datetime_info()` function locally and saved the data in [my dataset](https://www.kaggle.com/datasets/andradaolteanu/otto-helper-data). I've used `parallel_apply()` to compute them faster, this and `RAPIDS` library have been of real help for the huge amount of data we're handling in this competition.\n\n<center><img src=\"https://i.imgur.com/qF8C6sQ.png\" width=800></center>","metadata":{}},{"cell_type":"code","source":"# 🐝 Experiment\nrun = wandb.init(project='Otto', name='time_look', config=CONFIG)","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:49:08.214417Z","iopub.execute_input":"2022-12-01T11:49:08.218406Z","iopub.status.idle":"2022-12-01T11:49:18.536550Z","shell.execute_reply.started":"2022-12-01T11:49:08.218320Z","shell.execute_reply":"2022-12-01T11:49:18.534797Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> 📍 **Note** - do NOT divide by /1000 if you're using Radek's dataset, as he already did that. :) Otherwise you'll get data starting from 1970 and you'll wonder (like I did) if we went back in time.","metadata":{}},{"cell_type":"code","source":"# Simple function\ndef get_time(x):\n    '''Convert from Unix to Datetime.'''\n    return datetime.utcfromtimestamp(x).strftime('%Y-%m-%d %H:%M:%S')\n\n# With parallel\nfrom pandarallel import pandarallel\npandarallel.initialize(progress_bar=True)\n\ndef get_time_parallel(row):\n    '''Convert from Unix to Datetime.'''\n    return datetime.utcfromtimestamp(row.ts).strftime('%Y-%m-%d %H:%M:%S')\n\ndef get_night(x):\n    night = [21, 22, 23, 0, 1, 2, 3, 4, 5]\n    \n    if x in night:\n        return 1\n    else:\n        return 0","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:49:18.544179Z","iopub.execute_input":"2022-12-01T11:49:18.548391Z","iopub.status.idle":"2022-12-01T11:49:20.782862Z","shell.execute_reply.started":"2022-12-01T11:49:18.548278Z","shell.execute_reply":"2022-12-01T11:49:20.781427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_datetime_info(df):\n    df[\"datetime\"] = cudf.Series(df.to_pandas().parallel_apply(get_time_parallel, axis=1))\n\n    df[\"datetime\"] = cudf.to_datetime(df[\"datetime\"])\n\n    # Get day to retrieve activity\n    df[\"day\"] = df[\"datetime\"].dt.day\n    df[\"month\"] = df[\"datetime\"].dt.month\n    df[\"weekday\"] = df[\"datetime\"].dt.weekday\n    df[\"hour\"] = df[\"datetime\"].dt.hour\n    \n    df[\"night\"] = cudf.Series(df[\"hour\"].to_pandas().apply(lambda x: get_night(x)))\n    \n    df.drop(columns=[\"datetime\"], inplace=True)\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:49:20.790065Z","iopub.execute_input":"2022-12-01T11:49:20.794046Z","iopub.status.idle":"2022-12-01T11:49:24.636107Z","shell.execute_reply.started":"2022-12-01T11:49:20.793972Z","shell.execute_reply":"2022-12-01T11:49:24.634739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Read in prepped data - made locally using the above function\ntrain = cudf.read_parquet(\"/kaggle/input/otto-helper-data/train_prep.parquet\")\ntest = cudf.read_parquet(\"/kaggle/input/otto-helper-data/test_prep.parquet\")\n\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:49:24.643105Z","iopub.execute_input":"2022-12-01T11:49:24.646959Z","iopub.status.idle":"2022-12-01T11:49:56.514859Z","shell.execute_reply.started":"2022-12-01T11:49:24.646888Z","shell.execute_reply":"2022-12-01T11:49:56.513442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Q**: What is the **timeline** of sessions for train and test?\n\n**A**: The dataset starts from 31 July 2022 midnight and goes for 1 month, until 28 Aug 2022 for the training part. For the test part we need to predict 1 week in the future, from 28 Aug (where we left off during train) until 4 Sep 2022.","metadata":{}},{"cell_type":"code","source":"print(clr.S+\"Start Train:\"+clr.E, get_time(train[\"ts\"].min()), \"\\n\"+\n      clr.S+\"Finish Train:\"+clr.E, get_time(train[\"ts\"].max()), \"\\n\"+\n      clr.S+\"Start Test:\"+clr.E, get_time(test[\"ts\"].min()), \"\\n\"+\n      clr.S+\"Finish Test:\"+clr.E, get_time(test[\"ts\"].max()))","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:49:56.521843Z","iopub.execute_input":"2022-12-01T11:49:56.525940Z","iopub.status.idle":"2022-12-01T11:49:58.625070Z","shell.execute_reply.started":"2022-12-01T11:49:56.525864Z","shell.execute_reply":"2022-12-01T11:49:58.622957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Q**: How does the **number of sessions group over time**? (aka are they more condensed in certain periods or they are equaly distributed?)\n\n**A**: Looks like Sundays are the most active days of the month, with the highest activity happening then! There is somewhat of a pattern of \"lazy\" Mondays and a decend in activity until Thursday, with Friday beginning to surge again with the peak on Sunday.\n\n<center><img src=\"https://i.imgur.com/VOJMlcu.jpg\" width=800></center>","metadata":{}},{"cell_type":"code","source":"activity1 = train[\"day\"].value_counts().reset_index().to_pandas()\nactivity2 = test[\"day\"].value_counts().reset_index().to_pandas()\n\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize=(24, 12), gridspec_kw={'width_ratios': [3, 1]})\naxs = [ax1, ax2]\nfig.suptitle('Timeline Activity', \n             weight=\"bold\", size=25)\n\nsns.barplot(data=activity1, x=\"index\", y=\"day\", lw=3, color=my_colors[1], \n            order=[31, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15,\n                   16, 17, 18, 19, 20, 21, 22, 23, 23, 25, 26, 27, 28], ax=ax1)\n# show_values_on_bars(ax1, h_v=\"v\", space=0.4)\nax1.set_title('Train Activity (measure in Events)', weight=\"bold\", size=20)\n\nsns.barplot(data=activity2, x=\"index\", y=\"day\", lw=3, color=my_colors[5],\n            order = [28, 29, 30, 1, 2, 3, 4])\n# show_values_on_bars(ax2, h_v=\"v\", space=0.4)\nax2.set_title('Test Activity (measure in Events)', weight=\"bold\", size=20)\n\nfor ax in axs:\n#     ax.set_yticks([])\n    ax.set_xlabel(\"Day\", size = 18, weight=\"bold\")\n    ax.set_ylabel(\"Frequency\", size = 18, weight=\"bold\")\n    ax.set_ylim(activity1[\"day\"].min(), activity1[\"day\"].max())\n    \nsns.despine(right=True, top=True, left=True);","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:49:58.632077Z","iopub.execute_input":"2022-12-01T11:49:58.636005Z","iopub.status.idle":"2022-12-01T11:50:03.507390Z","shell.execute_reply.started":"2022-12-01T11:49:58.635935Z","shell.execute_reply":"2022-12-01T11:50:03.505831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"create_wandb_plot(x_data=activity1[\"index\"],\n                  y_data=activity1[\"day\"],\n                  x_name=\"Day\", y_name=\"Frequency\",\n                  title=\"Train Activity (measure in Events)\",\n                  log=\"bar_tr_activity\", plot=\"bar\")\n\ncreate_wandb_plot(x_data=activity2[\"index\"],\n                  y_data=activity2[\"day\"],\n                  x_name=\"Day\", y_name=\"Frequency\",\n                  title=\"Test Activity (measure in Events)\",\n                  log=\"bar_te_activity\", plot=\"bar\")","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:50:03.517421Z","iopub.execute_input":"2022-12-01T11:50:03.522774Z","iopub.status.idle":"2022-12-01T11:50:06.052918Z","shell.execute_reply.started":"2022-12-01T11:50:03.522701Z","shell.execute_reply":"2022-12-01T11:50:06.051333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Q**: But what about only for orders? How does the month activity looks like?\n\n**A**: They look about the same, with an outlier on 9th August (a Tuesday).","metadata":{}},{"cell_type":"code","source":"activity1 = train[train[\"type\"]==2][\"day\"].value_counts().reset_index().to_pandas()\nactivity2 = train[train[\"type\"]==2][\"day\"].value_counts().reset_index().to_pandas()\n\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize=(24, 12), gridspec_kw={'width_ratios': [3, 1]})\naxs = [ax1, ax2]\nfig.suptitle('Timeline Activity (Orders Only)', \n             weight=\"bold\", size=25)\n\nsns.barplot(data=activity1, x=\"index\", y=\"day\", lw=3, color=my_colors[1], \n            order=[31, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15,\n                   16, 17, 18, 19, 20, 21, 22, 23, 23, 25, 26, 27, 28], ax=ax1)\n# show_values_on_bars(ax1, h_v=\"v\", space=0.4)\nax1.set_title('Train Activity (measure in Events)', weight=\"bold\", size=20)\n\nsns.barplot(data=activity2, x=\"index\", y=\"day\", lw=3, color=my_colors[5],\n            order = [28, 29, 30, 1, 2, 3, 4])\n# show_values_on_bars(ax2, h_v=\"v\", space=0.4)\nax2.set_title('Test Activity (measure in Events)', weight=\"bold\", size=20)\n\nfor ax in axs:\n#     ax.set_yticks([])\n    ax.set_xlabel(\"Day\", size = 18, weight=\"bold\")\n    ax.set_ylabel(\"Frequency\", size = 18, weight=\"bold\")\n    ax.set_ylim(activity1[\"day\"].min(), activity1[\"day\"].max())\n    \nsns.despine(right=True, top=True, left=True);","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:50:06.054913Z","iopub.execute_input":"2022-12-01T11:50:06.055445Z","iopub.status.idle":"2022-12-01T11:50:09.422176Z","shell.execute_reply.started":"2022-12-01T11:50:06.055393Z","shell.execute_reply":"2022-12-01T11:50:09.420872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Q**: Is there **less or more activity during the night**? (aka is there a pattern in events number that correlates to the time of day?)\n\n**A**: Distributions are extremely similar. There is a steady increase throughout the day, with the lowest activity in the morning and noon, and the highest in the afternoon (between 5PM and 7PM).","metadata":{}},{"cell_type":"code","source":"hours1 = train.groupby([\"day\", \"hour\", \"night\"])[\"session\"].count()\\\n            .reset_index().groupby([\"hour\", \"night\"])[\"session\"]\\\n            .mean().reset_index().to_pandas()\nhours2 = test.groupby([\"day\", \"hour\", \"night\"])[\"session\"].count()\\\n            .reset_index().groupby([\"hour\", \"night\"])[\"session\"]\\\n            .mean().reset_index().to_pandas()\n\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize=(24, 12), gridspec_kw={'width_ratios': [2, 2]})\naxs = [ax1, ax2]\nfig.suptitle('Hour Activity', \n             weight=\"bold\", size=25)\n\nsns.barplot(data=hours1, x=\"hour\", y=\"session\", lw=3, color=my_colors[1], ax=ax1)\n# show_values_on_bars(ax1, h_v=\"v\", space=0.4)\nax1.set_title('Train Activity (measure in Events)', weight=\"bold\", size=20)\n\nsns.barplot(data=hours2, x=\"hour\", y=\"session\", lw=3, color=my_colors[5], ax=ax2)\n# show_values_on_bars(ax2, h_v=\"v\", space=0.4)\nax2.set_title('Test Activity (measure in Events)', weight=\"bold\", size=20)\n\nfor ax in axs:\n#     ax.set_yticks([])\n    ax.set_xlabel(\"Hour\", size = 18, weight=\"bold\")\n    ax.set_ylabel(\"Frequency\", size = 18, weight=\"bold\")\n    \nsns.despine(right=True, top=True, left=True);","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:50:09.423938Z","iopub.execute_input":"2022-12-01T11:50:09.425054Z","iopub.status.idle":"2022-12-01T11:50:13.296859Z","shell.execute_reply.started":"2022-12-01T11:50:09.425010Z","shell.execute_reply":"2022-12-01T11:50:13.295552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"create_wandb_plot(x_data=hours1[\"hour\"],\n                  y_data=hours1[\"session\"],\n                  x_name=\"Hour\", y_name=\"Frequency\",\n                  title=\"Train (Hour) Activity (measure in Events)\",\n                  log=\"bar_tr_hour\", plot=\"bar\")\n\ncreate_wandb_plot(x_data=hours2[\"hour\"],\n                  y_data=hours2[\"session\"],\n                  x_name=\"Hour\", y_name=\"Frequency\",\n                  title=\"Test (Hour) Activity (measure in Events)\",\n                  log=\"bar_te_hour\", plot=\"bar\")","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:50:13.298547Z","iopub.execute_input":"2022-12-01T11:50:13.299041Z","iopub.status.idle":"2022-12-01T11:50:16.482250Z","shell.execute_reply.started":"2022-12-01T11:50:13.298986Z","shell.execute_reply":"2022-12-01T11:50:16.480189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Q**: Do the users (sessions) have **registered activity** on 1 hour? 1 Day? Whole month? How does their activity look like?\n\n**A**: Looks like ... it varies!\n* ~45% of sessions have registered activity in less than 1 day. Which means more than half of the sessions have activity on 1 or multiple days!\n* for the rest it looks like it's an even 10%:\n    * ~10% of sessions have registered activity for 5->10 days\n    * ~10% of sessions have registered activity for 10->15 days\n    * ~10% of sessions have registered activity for 15->20 days\n    * ~10% of sessions have registered activity for 20->25 days\n* we have some strong outliers too, with ~4% of sessions having 25+ days of registered activity","metadata":{}},{"cell_type":"code","source":"# Get min and max time for each session\nsess_time = train.groupby(\"session\").aggregate({\"ts\": [\"min\", \"max\"]}).reset_index()\nsess_time.columns = [\"session\", \"min_ts\", \"max_ts\"]\n\n# Get time difference\nsess_time[\"diff\"] = cudf.to_datetime(sess_time[\"max_ts\"], unit=\"s\") - \\\n                    cudf.to_datetime(sess_time[\"min_ts\"], unit=\"s\")\n# Convert difference in days\n# days are floats as I also added the hours to them\nsess_time[\"diff\"] = sess_time[\"diff\"].dt.components[\"hours\"]/24 + sess_time[\"diff\"].dt.days","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:50:16.484106Z","iopub.execute_input":"2022-12-01T11:50:16.484816Z","iopub.status.idle":"2022-12-01T11:50:18.548124Z","shell.execute_reply.started":"2022-12-01T11:50:16.484761Z","shell.execute_reply":"2022-12-01T11:50:18.546740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sess = sess_time[\"diff\"][::60].copy()\nf, (a0, a1) = plt.subplots(2, 1, gridspec_kw={'height_ratios': [3, 1]}, figsize=(24, 15))\nsns.distplot(sess.to_pandas(), rug=True, hist=False, \n#              bins=10,\n             rug_kws={\"color\": my_colors[5]},\n             kde_kws={\"color\": my_colors[5], \"lw\": 5, \"alpha\": 0.7},\n#              hist_kws={\"histtype\": \"step\", \"linewidth\": 3, \"alpha\": 1, \"color\": my_colors[5]},\n             ax=a0)\n\na0.axvline(x=3, ls=\":\", lw=2, color=\"black\")\na0.text(x=-2, y=0.25, s=\"50% of data is here\", size=17, color=\"black\", weight=\"bold\")\na0.text(x=10, y=0.05, s=\"rest of 50% of data is spread evenly here\", size=17, color=\"black\", weight=\"bold\")\n\nsns.boxplot(x=sess.to_pandas(), ax=a1, notch=False, showcaps=True, \n            flierprops={\"marker\": \"x\"},\n            boxprops={\"facecolor\": my_colors[5]},\n            medianprops={\"color\": my_colors[0]},)\n\nplt.suptitle(\"TRAIN: Distribution of Number of Days a Session took place\", weight=\"bold\", size=25)\nsns.despine(right=True, top=True, left=True);","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:50:18.550127Z","iopub.execute_input":"2022-12-01T11:50:18.550656Z","iopub.status.idle":"2022-12-01T11:50:25.956551Z","shell.execute_reply.started":"2022-12-01T11:50:18.550604Z","shell.execute_reply":"2022-12-01T11:50:25.953067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Q**: We looked at an user overall activity (from the first event to the last). Now I wanna see for some of them **how is the intensity**? Do they have more events (clicks, carts, orders) at the beginning and fade away? Or are they constant?\n\n**A**: Honestly, I do not have an e xact answer here. Majority of sessions seen below have a spike of activity in the first day and then another one in the last day (looks like they looked once at the beginning and then again - maybe even ordered - in the last day). However it's not always tha case, other sessions have a latent activity over the entire period, so here I cannot conclude something very ... speciffic.","metadata":{}},{"cell_type":"code","source":"def get_session_activity(sess_example):\n    # Filter only needed session\n    dt = train[train[\"session\"] == sess_example][\"day\"].value_counts().reset_index()\n    dt.columns=[\"day\", \"activity_count\"]\n\n    # Append 0 for any days that are missing\n    # So the lineplot is accurate\n    for k in range(1, 30):\n        if k not in dt[\"day\"]:\n            dt = dt.append({\"day\":k, \"activity_count\":0}, ignore_index=True)\n\n    # Plot\n    plt.figure(figsize=(20, 8))\n    plt.title(f\"Session {sess_example}: activity over days\", weight=\"bold\", size=25)\n    sns.lineplot(data=dt.to_pandas(), x=\"day\", y=\"activity_count\",\n                 color=my_colors[1], lw=10)\n    plt.xlabel(\"Day\", size = 18, weight=\"bold\")\n    plt.ylabel(\"Activity Count (in events)\", size = 18, weight=\"bold\")\n    sns.despine(right=True, top=True, left=True)\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:52:48.768562Z","iopub.execute_input":"2022-12-01T11:52:48.769032Z","iopub.status.idle":"2022-12-01T11:52:50.377120Z","shell.execute_reply.started":"2022-12-01T11:52:48.768994Z","shell.execute_reply":"2022-12-01T11:52:50.375466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Low Activity - <= 5 days","metadata":{}},{"cell_type":"code","source":"print(clr.S+\"Session that had activity for 4 days:\"+clr.E, \"\\n\")\nget_session_activity(sess_example=sess_time[sess_time[\"diff\"]==4][\"session\"].values.tolist()[0])","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:54:01.580166Z","iopub.execute_input":"2022-12-01T11:54:01.581595Z","iopub.status.idle":"2022-12-01T11:54:04.014984Z","shell.execute_reply.started":"2022-12-01T11:54:01.581541Z","shell.execute_reply":"2022-12-01T11:54:04.013486Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(clr.S+\"Session that had activity for 5 days:\"+clr.E, \"\\n\")\nget_session_activity(sess_example=sess_time[sess_time[\"diff\"]==5][\"session\"].values.tolist()[0])","metadata":{"execution":{"iopub.status.busy":"2022-12-01T11:53:21.272175Z","iopub.execute_input":"2022-12-01T11:53:21.272760Z","iopub.status.idle":"2022-12-01T11:53:23.726561Z","shell.execute_reply.started":"2022-12-01T11:53:21.272712Z","shell.execute_reply":"2022-12-01T11:53:23.725134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Mediun Activity - 5 -> 15 days","metadata":{}},{"cell_type":"code","source":"print(clr.S+\"Session that had activity for 7 days:\"+clr.E, \"\\n\")\nget_session_activity(sess_example=sess_time[sess_time[\"diff\"]==7][\"session\"].values.tolist()[0])","metadata":{"execution":{"iopub.status.busy":"2022-12-01T12:00:21.303520Z","iopub.execute_input":"2022-12-01T12:00:21.304456Z","iopub.status.idle":"2022-12-01T12:00:23.892709Z","shell.execute_reply.started":"2022-12-01T12:00:21.304411Z","shell.execute_reply":"2022-12-01T12:00:23.890785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(clr.S+\"Session that had activity for 10 days:\"+clr.E, \"\\n\")\nget_session_activity(sess_example=sess_time[sess_time[\"diff\"]==10][\"session\"].values.tolist()[0])","metadata":{"execution":{"iopub.status.busy":"2022-12-01T12:00:27.483830Z","iopub.execute_input":"2022-12-01T12:00:27.484273Z","iopub.status.idle":"2022-12-01T12:00:30.677934Z","shell.execute_reply.started":"2022-12-01T12:00:27.484235Z","shell.execute_reply":"2022-12-01T12:00:30.676543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### High Activity - 15+ days","metadata":{}},{"cell_type":"code","source":"print(clr.S+\"Activity for a session that had activity for 25 days\"+clr.E)\nget_session_activity(sess_example=sess_time[sess_time[\"diff\"]==25][\"session\"].values.tolist()[1])","metadata":{"execution":{"iopub.status.busy":"2022-12-01T12:00:56.107831Z","iopub.execute_input":"2022-12-01T12:00:56.109024Z","iopub.status.idle":"2022-12-01T12:00:59.294849Z","shell.execute_reply.started":"2022-12-01T12:00:56.108966Z","shell.execute_reply":"2022-12-01T12:00:59.293468Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(clr.S+\"Activity for a session that had activity for 26 days\"+clr.E)\nget_session_activity(sess_example=sess_time[sess_time[\"diff\"]==26][\"session\"].values.tolist()[1])","metadata":{"execution":{"iopub.status.busy":"2022-12-01T12:01:09.439704Z","iopub.execute_input":"2022-12-01T12:01:09.440211Z","iopub.status.idle":"2022-12-01T12:01:12.122051Z","shell.execute_reply.started":"2022-12-01T12:01:09.440169Z","shell.execute_reply":"2022-12-01T12:01:12.120666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del sess_time\ngc.collect()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"wandb.finish()","metadata":{"execution":{"iopub.status.busy":"2022-11-27T13:47:46.694515Z","iopub.execute_input":"2022-11-27T13:47:46.694869Z","iopub.status.idle":"2022-11-27T13:48:56.021768Z","shell.execute_reply.started":"2022-11-27T13:47:46.694833Z","shell.execute_reply":"2022-11-27T13:48:56.020754Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Item Analysis (aid)\n\n* `train`: 1,855,603 unique products\n* `test`: 783,486 unique products\n* there is no product in the test set that doesn't appear at least once in the train set\n\n<center><img src=\"https://i.imgur.com/VQLrSGL.jpg\" width=800></center>","metadata":{}},{"cell_type":"code","source":"# 🐝 Experiment\nrun = wandb.init(project='Otto', name='product_look', config=CONFIG)","metadata":{"execution":{"iopub.status.busy":"2022-11-27T13:48:56.023400Z","iopub.execute_input":"2022-11-27T13:48:56.023777Z","iopub.status.idle":"2022-11-27T13:49:06.134340Z","shell.execute_reply.started":"2022-11-27T13:48:56.023737Z","shell.execute_reply":"2022-11-27T13:49:06.133190Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Products Overview\nprint(clr.S+\"Train - Unique products:\"+clr.E, train[\"aid\"].nunique())\nprint(clr.S+\"Test - Unique products:\"+clr.E, test[\"aid\"].nunique())\nprint(clr.S+\"Products Found in Test & Train:\"+clr.E,\n     len(set(test[\"aid\"].unique().to_pandas())\\\n         .intersection(set(train[\"aid\"].unique().to_pandas()))))\nprint(clr.S+\"Products in Test but not in Train:\"+clr.E, \n      test[\"aid\"].nunique() - len(set(test[\"aid\"].unique().to_pandas())\\\n                                     .intersection(set(train[\"aid\"].unique().to_pandas()))))\n\nwandb.log({\"new_in_test\": test[\"aid\"].nunique() - len(set(test[\"aid\"].unique().to_pandas())\\\n                                     .intersection(set(train[\"aid\"].unique().to_pandas())))})","metadata":{"execution":{"iopub.status.busy":"2022-11-27T13:49:06.140109Z","iopub.execute_input":"2022-11-27T13:49:06.142582Z","iopub.status.idle":"2022-11-27T13:49:11.830897Z","shell.execute_reply.started":"2022-11-27T13:49:06.142535Z","shell.execute_reply":"2022-11-27T13:49:11.829872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Q**: Do products repeat in a same patterns, or are there products that have more **interaction** than others?\n\n**A**: Vast majority (62% in `train`) of products are seen **less than 30 times** within the dataset. There are however some pretty strong outliers.","metadata":{}},{"cell_type":"code","source":"sess = train[\"aid\"].value_counts().values\nf, (a0, a1) = plt.subplots(2, 1, gridspec_kw={'height_ratios': [3, 1]}, figsize=(24, 15))\nsns.distplot(sess[::20].get(), rug=True, hist=False, \n#              bins=10,\n             rug_kws={\"color\": my_colors[5]},\n             kde_kws={\"color\": my_colors[5], \"lw\": 5, \"alpha\": 0.7},\n#              hist_kws={\"histtype\": \"step\", \"linewidth\": 3, \"alpha\": 1, \"color\": my_colors[5]},\n             ax=a0)\n\na0.axvline(x=1300, ls=\"--\", lw=2, color=\"black\")\na0.text(x=1500, y=0.00013, s=\"62% of products repeat less than 30 times\",\n        size=17, color=\"black\", weight=\"bold\")\n\nsns.boxplot(x=sess.get(), ax=a1, notch=False, showcaps=True, \n            flierprops={\"marker\": \"x\"},\n            boxprops={\"facecolor\": my_colors[5]},\n            medianprops={\"color\": my_colors[0]},)\n\nplt.suptitle(\"TRAIN: Distribution of Number of Apparitions of Products\", weight=\"bold\", size=25)\nsns.despine(right=True, top=True, left=True);","metadata":{"execution":{"iopub.status.busy":"2022-11-27T13:49:11.835871Z","iopub.execute_input":"2022-11-27T13:49:11.838270Z","iopub.status.idle":"2022-11-27T13:49:19.964161Z","shell.execute_reply.started":"2022-11-27T13:49:11.838228Z","shell.execute_reply":"2022-11-27T13:49:19.963228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"create_wandb_hist(x_data=sess[::20].tolist(), \n                  x_name=\"Number of Apparitions / Product\",\n                  title=\"TRAIN: Distribution of Number of Apparitions of Products\",\n                  log=\"hist_product_tr\")","metadata":{"execution":{"iopub.status.busy":"2022-11-27T13:49:19.968851Z","iopub.execute_input":"2022-11-27T13:49:19.971136Z","iopub.status.idle":"2022-11-27T13:49:27.380022Z","shell.execute_reply.started":"2022-11-27T13:49:19.971097Z","shell.execute_reply":"2022-11-27T13:49:27.379074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sess = test[\"aid\"].value_counts().values\nf, (a0, a1) = plt.subplots(2, 1, gridspec_kw={'height_ratios': [3, 1]}, figsize=(24, 15))\nsns.distplot(sess[::10].get(), rug=True, hist=False, \n#              bins=10,\n             rug_kws={\"color\": my_colors[5]},\n             kde_kws={\"color\": my_colors[5], \"lw\": 5, \"alpha\": 0.7},\n#              hist_kws={\"histtype\": \"step\", \"linewidth\": 3, \"alpha\": 1, \"color\": my_colors[5]},\n             ax=a0)\n\na0.axvline(x=200, ls=\"--\", lw=2, color=\"black\")\na0.text(x=210, y=0.00178, s=\"95% of products repeat less than 30 times\",\n        size=17, color=\"black\", weight=\"bold\")\n\nsns.boxplot(x=sess.get(), ax=a1, notch=False, showcaps=True, \n            flierprops={\"marker\": \"x\"},\n            boxprops={\"facecolor\": my_colors[5]},\n            medianprops={\"color\": my_colors[0]},)\n\nplt.suptitle(\"TEST: Distribution of Number of Apparitions of Products\", weight=\"bold\", size=25)\nsns.despine(right=True, top=True, left=True);","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-11-27T13:49:27.384688Z","iopub.execute_input":"2022-11-27T13:49:27.386978Z","iopub.status.idle":"2022-11-27T13:49:32.407531Z","shell.execute_reply.started":"2022-11-27T13:49:27.386939Z","shell.execute_reply":"2022-11-27T13:49:32.406538Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"create_wandb_hist(x_data=sess[::10].tolist(), \n                  x_name=\"Number of Apparitions / Product\",\n                  title=\"TEST: Distribution of Number of Apparitions of Products\",\n                  log=\"hist_product_te\")","metadata":{"execution":{"iopub.status.busy":"2022-11-27T13:49:32.412197Z","iopub.execute_input":"2022-11-27T13:49:32.414454Z","iopub.status.idle":"2022-11-27T13:49:40.018858Z","shell.execute_reply.started":"2022-11-27T13:49:32.414414Z","shell.execute_reply":"2022-11-27T13:49:40.017809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Q**: Do the products with most events have **same interactions** in train and test? In other words, do the products that are most viewed in train also have a similar pattern in test data? What about on events alone?\n\n**A**: On first glance (top 25 products) they do. 9/25 top products (with most interactions) appear in Train and Test (looking at the overall engagement). For \"add to cart\" the similarity is even stronger, with 14/25 aids being in boths Train and Test.","metadata":{}},{"cell_type":"code","source":"# Top products to show\nN = 25","metadata":{"execution":{"iopub.status.busy":"2022-11-27T13:49:40.023808Z","iopub.execute_input":"2022-11-27T13:49:40.026190Z","iopub.status.idle":"2022-11-27T13:49:42.404883Z","shell.execute_reply.started":"2022-11-27T13:49:40.026150Z","shell.execute_reply":"2022-11-27T13:49:42.403821Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dt_tr = train[\"aid\"].value_counts().head(N).reset_index()\ndt_tr.columns = [\"aid\", \"count\"]\ndt_tr[\"aid\"] = dt_tr[\"aid\"].astype(str)\n\ndt_te = test[\"aid\"].value_counts().head(N).reset_index()\ndt_te.columns = [\"aid\", \"count\"]\ndt_te[\"aid\"] = dt_te[\"aid\"].astype(str)\n\nprint(clr.S+f\"Top products in both Train and Test:\"+clr.E, \n      set(dt_tr[\"aid\"].unique().to_pandas())\\\n            .intersection(set(dt_te[\"aid\"].unique().to_pandas())))\n\n# Plot\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize=(24, 12))\naxs = [ax1, ax2]\nfig.suptitle('Most frequented products (overall)', \n             weight=\"bold\", size=25)\n\nsns.barplot(data=dt_tr.to_pandas(), x=\"count\", y=\"aid\", ax=ax1,\n            palette = sns.blend_palette(colors = [my_colors[6], my_colors[4]], n_colors=N))\nshow_values_on_bars(ax1, h_v=\"h\", space=0.4)\nax1.set_title('TRAIN', weight=\"bold\", size=20)\n\nsns.barplot(data=dt_te.to_pandas(), x=\"count\", y=\"aid\", ax=ax2,\n            palette = sns.blend_palette(colors = [my_colors[1], my_colors[3]], n_colors=N))\nshow_values_on_bars(ax2, h_v=\"h\", space=0.4)\nax2.set_title('TEST', weight=\"bold\", size=20)\n\nfor ax in axs:\n#     ax.set_yticks([])\n    ax.set_xlabel(\"Count\", size = 18, weight=\"bold\")\n    ax.set_ylabel(\"AID\", size = 18, weight=\"bold\")\n    \nsns.despine(right=True, top=True, left=True);","metadata":{"execution":{"iopub.status.busy":"2022-11-27T13:49:42.406174Z","iopub.execute_input":"2022-11-27T13:49:42.406710Z","iopub.status.idle":"2022-11-27T13:49:46.089010Z","shell.execute_reply.started":"2022-11-27T13:49:42.406663Z","shell.execute_reply":"2022-11-27T13:49:46.088064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"type_value = 0","metadata":{"execution":{"iopub.status.busy":"2022-11-27T13:49:46.090573Z","iopub.execute_input":"2022-11-27T13:49:46.090954Z","iopub.status.idle":"2022-11-27T13:49:48.479999Z","shell.execute_reply.started":"2022-11-27T13:49:46.090916Z","shell.execute_reply":"2022-11-27T13:49:48.478847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dt_tr = train[train[\"type\"]==type_value][\"aid\"].value_counts().head(N).reset_index()\ndt_tr.columns = [\"aid\", \"count\"]\ndt_tr[\"aid\"] = dt_tr[\"aid\"].astype(str)\n\ndt_te = test[test[\"type\"]==type_value][\"aid\"].value_counts().head(N).reset_index()\ndt_te.columns = [\"aid\", \"count\"]\ndt_te[\"aid\"] = dt_te[\"aid\"].astype(str)\n\nprint(clr.S+f\"Top products in both Train and Test:\"+clr.E, \n      set(dt_tr[\"aid\"].unique().to_pandas())\\\n            .intersection(set(dt_te[\"aid\"].unique().to_pandas())))\n\n# Plot\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize=(24, 12))\naxs = [ax1, ax2]\nfig.suptitle(f'Most frequented products (CLICKS)', \n             weight=\"bold\", size=25)\n\nsns.barplot(data=dt_tr.to_pandas(), x=\"count\", y=\"aid\", ax=ax1,\n            palette = sns.blend_palette(colors = [my_colors[6], my_colors[4]], n_colors=N))\nshow_values_on_bars(ax1, h_v=\"h\", space=0.4)\nax1.set_title('TRAIN', weight=\"bold\", size=20)\n\nsns.barplot(data=dt_te.to_pandas(), x=\"count\", y=\"aid\", ax=ax2,\n            palette = sns.blend_palette(colors = [my_colors[1], my_colors[3]], n_colors=N))\nshow_values_on_bars(ax2, h_v=\"h\", space=0.4)\nax2.set_title('TEST', weight=\"bold\", size=20)\n\nfor ax in axs:\n#     ax.set_yticks([])\n    ax.set_xlabel(\"Count\", size = 18, weight=\"bold\")\n    ax.set_ylabel(\"AID\", size = 18, weight=\"bold\")\n    \nsns.despine(right=True, top=True, left=True);","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-11-27T13:49:48.481524Z","iopub.execute_input":"2022-11-27T13:49:48.482453Z","iopub.status.idle":"2022-11-27T13:49:52.439743Z","shell.execute_reply.started":"2022-11-27T13:49:48.482378Z","shell.execute_reply":"2022-11-27T13:49:52.438699Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"type_value = 1","metadata":{"execution":{"iopub.status.busy":"2022-11-27T13:49:52.441162Z","iopub.execute_input":"2022-11-27T13:49:52.441715Z","iopub.status.idle":"2022-11-27T13:49:55.731501Z","shell.execute_reply.started":"2022-11-27T13:49:52.441675Z","shell.execute_reply":"2022-11-27T13:49:55.730579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dt_tr = train[train[\"type\"]==type_value][\"aid\"].value_counts().head(N).reset_index()\ndt_tr.columns = [\"aid\", \"count\"]\ndt_tr[\"aid\"] = dt_tr[\"aid\"].astype(str)\n\ndt_te = test[test[\"type\"]==type_value][\"aid\"].value_counts().head(N).reset_index()\ndt_te.columns = [\"aid\", \"count\"]\ndt_te[\"aid\"] = dt_te[\"aid\"].astype(str)\n\nprint(clr.S+f\"Top products in both Train and Test:\"+clr.E, \n      set(dt_tr[\"aid\"].unique().to_pandas())\\\n            .intersection(set(dt_te[\"aid\"].unique().to_pandas())))\n\n# Plot\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize=(24, 12))\naxs = [ax1, ax2]\nfig.suptitle(f'Most frequented products (CARTS)', \n             weight=\"bold\", size=25)\n\nsns.barplot(data=dt_tr.to_pandas(), x=\"count\", y=\"aid\", ax=ax1,\n            palette = sns.blend_palette(colors = [my_colors[6], my_colors[4]], n_colors=N))\nshow_values_on_bars(ax1, h_v=\"h\", space=0.4)\nax1.set_title('TRAIN', weight=\"bold\", size=20)\n\nsns.barplot(data=dt_te.to_pandas(), x=\"count\", y=\"aid\", ax=ax2,\n            palette = sns.blend_palette(colors = [my_colors[1], my_colors[3]], n_colors=N))\nshow_values_on_bars(ax2, h_v=\"h\", space=0.4)\nax2.set_title('TEST', weight=\"bold\", size=20)\n\nfor ax in axs:\n#     ax.set_yticks([])\n    ax.set_xlabel(\"Count\", size = 18, weight=\"bold\")\n    ax.set_ylabel(\"AID\", size = 18, weight=\"bold\")\n    \nsns.despine(right=True, top=True, left=True);","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-11-27T13:49:55.732773Z","iopub.execute_input":"2022-11-27T13:49:55.733124Z","iopub.status.idle":"2022-11-27T13:49:59.304989Z","shell.execute_reply.started":"2022-11-27T13:49:55.733089Z","shell.execute_reply":"2022-11-27T13:49:59.304050Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"type_value = 2","metadata":{"execution":{"iopub.status.busy":"2022-11-27T13:49:59.306543Z","iopub.execute_input":"2022-11-27T13:49:59.306925Z","iopub.status.idle":"2022-11-27T13:50:01.813523Z","shell.execute_reply.started":"2022-11-27T13:49:59.306887Z","shell.execute_reply":"2022-11-27T13:50:01.812562Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dt_tr = train[train[\"type\"]==type_value][\"aid\"].value_counts().head(N).reset_index()\ndt_tr.columns = [\"aid\", \"count\"]\ndt_tr[\"aid\"] = dt_tr[\"aid\"].astype(str)\n\ndt_te = test[test[\"type\"]==type_value][\"aid\"].value_counts().head(N).reset_index()\ndt_te.columns = [\"aid\", \"count\"]\ndt_te[\"aid\"] = dt_te[\"aid\"].astype(str)\n\nprint(clr.S+f\"Top products in both Train and Test:\"+clr.E, \n          set(dt_tr[\"aid\"].unique().to_pandas())\\\n                .intersection(set(dt_te[\"aid\"].unique().to_pandas())))\n\n# Plot\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize=(24, 12))\naxs = [ax1, ax2]\nfig.suptitle(f'Most frequented products (ORDERS)', \n             weight=\"bold\", size=25)\n\nsns.barplot(data=dt_tr.to_pandas(), x=\"count\", y=\"aid\", ax=ax1,\n            palette = sns.blend_palette(colors = [my_colors[6], my_colors[4]], n_colors=N))\nshow_values_on_bars(ax1, h_v=\"h\", space=0.4)\nax1.set_title('TRAIN', weight=\"bold\", size=20)\n\nsns.barplot(data=dt_te.to_pandas(), x=\"count\", y=\"aid\", ax=ax2,\n            palette = sns.blend_palette(colors = [my_colors[1], my_colors[3]], n_colors=N))\nshow_values_on_bars(ax2, h_v=\"h\", space=0.4)\nax2.set_title('TEST', weight=\"bold\", size=20)\n\nfor ax in axs:\n#     ax.set_yticks([])\n    ax.set_xlabel(\"Count\", size = 18, weight=\"bold\")\n    ax.set_ylabel(\"AID\", size = 18, weight=\"bold\")\n    \nsns.despine(right=True, top=True, left=True);","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-11-27T13:50:01.815009Z","iopub.execute_input":"2022-11-27T13:50:01.815345Z","iopub.status.idle":"2022-11-27T13:50:05.666777Z","shell.execute_reply.started":"2022-11-27T13:50:01.815309Z","shell.execute_reply":"2022-11-27T13:50:05.665679Z"},"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"📍 So there are quite a few products that appear in both Train and Test when looking at top 25 prods with most event types:\n* CLICKS: {'108125', '554660', '832192', '184976', '29735', '986164', '1460571', '485256'}\n* CARTS: {'33343', '152547', '554660', '832192', '29735', '258353', '986164', '544144', '1022566', '332654', '1462420', '1460571', '166037', '485256'}\n* ORDERS: {'923948', '29735', '1022566', '986164', '544144', '332654', '166037'}\n\n**Q**: What about first 1000 products? What about their pairing? Can we find some rules?","metadata":{}},{"cell_type":"code","source":"wandb.finish()","metadata":{"execution":{"iopub.status.busy":"2022-11-27T13:50:05.668508Z","iopub.execute_input":"2022-11-27T13:50:05.668999Z","iopub.status.idle":"2022-11-27T13:51:15.412560Z","shell.execute_reply.started":"2022-11-27T13:50:05.668940Z","shell.execute_reply":"2022-11-27T13:51:15.411316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5. Baseline - Predict Top Clicks/Carts/Orders\n\n📍 **Note**: We can predict up to 20 products for each type (we won't be penalized for predicting more than necessary). So, for a dummy baseline, I will predict for each session and type most common aids in each type.","metadata":{}},{"cell_type":"code","source":"# Get top N products for each type for train\ngroups = train.groupby([\"aid\", \"type\"])[\"night\"].count().reset_index()\\\n        .sort_values([\"type\", \"night\"], ascending=False).reset_index(drop=True)\ngroups.columns = [\"aid\", \"type\", \"count\"]\n\nN = 60\nclicks = groups[groups[\"type\"]==0].head(N)[\"aid\"].values.tolist()\ncarts = groups[groups[\"type\"]==1].head(N)[\"aid\"].values.tolist()\norders = groups[groups[\"type\"]==2].head(N)[\"aid\"].values.tolist()\n\nprint(clr.S+\"(Clicks) Top Aids:\"+clr.E, clicks)\nprint(clr.S+\"(Carts) Top Aids:\"+clr.E, carts)\nprint(clr.S+\"(Orders) Top Aids:\"+clr.E, orders)","metadata":{"execution":{"iopub.status.busy":"2022-11-27T19:17:39.907799Z","iopub.execute_input":"2022-11-27T19:17:39.908530Z","iopub.status.idle":"2022-11-27T19:17:40.340104Z","shell.execute_reply.started":"2022-11-27T19:17:39.908492Z","shell.execute_reply":"2022-11-27T19:17:40.338960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Sample Submission\nss = cudf.read_csv(\"/kaggle/input/otto-recommender-system/sample_submission.csv\")\nss.head()","metadata":{"execution":{"iopub.status.busy":"2022-11-27T19:17:40.343270Z","iopub.execute_input":"2022-11-27T19:17:40.343653Z","iopub.status.idle":"2022-11-27T19:17:42.105246Z","shell.execute_reply.started":"2022-11-27T19:17:40.343617Z","shell.execute_reply":"2022-11-27T19:17:42.104293Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Submission 1: Top 20 products (repeating)\n\n📍**Note**: For the first submission I will just predict top 20 products, without looking at the user's activity before (hence, if session 9999 in test clicked on product 001 and product 001 is also in top 20, I will still consider it).\n\n*LB Score: 0.006 - EXTREMELY low :), not good*","metadata":{}},{"cell_type":"code","source":"def get_top_20_repeats(clicks, carts, orders):\n    # Make them strings\n    # like we're shown in sample submission\n    clicks_1 = \" \".join(str(v) for v in clicks[:20])\n    carts_1 = \" \".join(str(v) for v in carts[:20])\n    orders_1 = \" \".join(str(v) for v in orders[:20])\n\n    labels = []\n\n    for k in tqdm(range(len(ss))):\n        sess_type = ss.iloc[k, :][\"session_type\"][k]\n\n        if \"clicks\" in sess_type:\n            labels.append(clicks_1)\n        elif \"carts\" in sess_type:\n            labels.append(carts_1)\n        else:\n            labels.append(orders_1)\n            \n    return labels\n\n# Processed this locally\n# takes me 20 seconds vs Kaggle env ~20 mins\nlabels_top20 = np.load(\"/kaggle/input/otto-helper-data/top_20_repeats.npy\")","metadata":{"execution":{"iopub.status.busy":"2022-11-27T19:18:42.384368Z","iopub.execute_input":"2022-11-27T19:18:42.384727Z","iopub.status.idle":"2022-11-27T19:19:01.104465Z","shell.execute_reply.started":"2022-11-27T19:18:42.384678Z","shell.execute_reply":"2022-11-27T19:19:01.103423Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Submission 2: Last 20 AIDs\n\nI also wanted to combine [Radek's](https://www.kaggle.com/code/radek1/last-20-aids) last 20 aids solution with my *top 20* labels and see where that leads me. His LB score is **0.464** with that baseline only.","metadata":{}},{"cell_type":"code","source":"# Radek's Solution\n# taking all aids per session in the test set and predicting them as next steps\ntest_sess_aids = test.to_pandas().groupby('session')['aid'].apply(lambda x: list(x)[-20:])","metadata":{"execution":{"iopub.status.busy":"2022-11-27T19:24:44.443377Z","iopub.execute_input":"2022-11-27T19:24:44.443782Z","iopub.status.idle":"2022-11-27T19:25:21.182105Z","shell.execute_reply.started":"2022-11-27T19:24:44.443744Z","shell.execute_reply":"2022-11-27T19:25:21.181035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"session_types = ['clicks', 'carts', 'orders']\nlabels_radek = []\n\nfor _, aids in tqdm(test_sess_aids.iteritems()):\n    for types in session_types:\n        labels_radek.append(' '.join([str(a) for a in aids]))","metadata":{"execution":{"iopub.status.busy":"2022-11-27T19:25:48.891022Z","iopub.execute_input":"2022-11-27T19:25:48.891380Z","iopub.status.idle":"2022-11-27T19:25:56.351767Z","shell.execute_reply.started":"2022-11-27T19:25:48.891350Z","shell.execute_reply":"2022-11-27T19:25:56.350757Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Radek's + top20\n# Combining these 2, as Radek's only has in many cases less than 20 total aids predicted\nnew_labels = []\n\nfor l1, l2 in tqdm(zip(labels_radek, labels_top20)):\n    new_labels.append(\" \".join((l1 + \" \" + l2).split(\" \")[:20]))","metadata":{"execution":{"iopub.status.busy":"2022-11-27T19:27:00.198839Z","iopub.execute_input":"2022-11-27T19:27:00.199295Z","iopub.status.idle":"2022-11-27T19:27:13.484645Z","shell.execute_reply.started":"2022-11-27T19:27:00.199255Z","shell.execute_reply":"2022-11-27T19:27:13.483392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Clean memory\ndel labels_radek, labels_top20, test_sess_aids\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-11-27T19:27:21.823529Z","iopub.execute_input":"2022-11-27T19:27:21.823893Z","iopub.status.idle":"2022-11-27T19:27:34.379403Z","shell.execute_reply.started":"2022-11-27T19:27:21.823861Z","shell.execute_reply":"2022-11-27T19:27:34.378087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Make the submission\nss[\"labels\"] = new_labels\nss.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-11-27T19:28:05.049318Z","iopub.execute_input":"2022-11-27T19:28:05.049675Z","iopub.status.idle":"2022-11-27T19:28:08.844023Z","shell.execute_reply.started":"2022-11-27T19:28:05.049646Z","shell.execute_reply":"2022-11-27T19:28:08.835253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This leads to a very slight increase from **0.464** to **0.465**. :) Next Notebook will explore more various (and not dummy) approaches.","metadata":{}},{"cell_type":"code","source":"# 🐝 Save Artifacts\nsave_dataset_artifact(run_name=\"save_train_prep\", \n                      artifact_name=\"train_prep\",\n                      path=\"/kaggle/input/otto-helper-data/train_prep.parquet\",\n                      data_type=\"dataset\")\nsave_dataset_artifact(run_name=\"save_test_prep\", \n                      artifact_name=\"test_prep\",\n                      path=\"/kaggle/input/otto-helper-data/test_prep.parquet\",\n                      data_type=\"dataset\")","metadata":{"execution":{"iopub.status.busy":"2022-11-27T19:30:03.706109Z","iopub.execute_input":"2022-11-27T19:30:03.706479Z","iopub.status.idle":"2022-11-27T19:32:00.418778Z","shell.execute_reply.started":"2022-11-27T19:30:03.706446Z","shell.execute_reply":"2022-11-27T19:32:00.417550Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# <center><video src=\".mp4\" width=800 controls></center>","metadata":{"execution":{"iopub.status.busy":"2022-11-27T12:32:15.307437Z","iopub.execute_input":"2022-11-27T12:32:15.307804Z","iopub.status.idle":"2022-11-27T12:32:15.314043Z","shell.execute_reply.started":"2022-11-27T12:32:15.307768Z","shell.execute_reply":"2022-11-27T12:32:15.312967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 🐝 [W&B Dashboard](https://wandb.ai/andrada/Otto?workspace=user-andrada)\n    \n<center><img src=\"https://i.imgur.com/0alEevD.png\"></center>\n\n------\n\n<center><img src=\"https://i.imgur.com/knxTRkO.png\"></center>\n\n### My Specs\n\n* 🖥 Z8 G4 Workstation\n* 💾 2 CPUs & 96GB Memory\n* 🎮 2x NVIDIA A6000\n* 💻 Zbook Studio G7 on the go","metadata":{}}]}