{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# BERT + Nearest Neighbors - H&M Product Recommendation\nThis notebook follows the industry standard pipeline *CRISP* for *Data Mining*.\n\nIt goes from the understanding of the business requirements, through data analysis and data preparation, and finally modeling and evaluation. \n\nDelve right into the table of contents below!","metadata":{}},{"cell_type":"markdown","source":"<a id=\"contents\"></a>\n# Contents\n[Libraries](#libs)\n1. [Business understanding](#business)\n2. [Data understanding](#du)\n\n    2.1. [Articles](#du-articles)\n    \n    2.2. [Customers](#customers)\n    \n    2.3. [Transactions](#transactions)\n    \n    2.4. [Conclusion](#du-conclusion)\n\n3. [Data preparation](#dp)\n\n    3.1. [Merge chosen textual columns](#dp-1)\n    \n    3.2. [Embed texts into vectors using BERT](#dp-2)\n    \n    3.3. [Add the vectors into the transactions](#dp-3)\n    \n    3.4. [Calculate the average transaction vector for each customer](#dp-4)\n    \n    3.5. [Add new column for bought articles](#dp-5)\n    \n    3.6. [Conclusion](#dp-conclusion)\n\n4. [Modeling](#modeling)\n5. [Evaluation](#eval)\n6. [System Usage](#usage)\n7. [Submission](#submission)","metadata":{}},{"cell_type":"markdown","source":"<a id=\"libs\"></a>\n# Libraries","metadata":{}},{"cell_type":"code","source":"import os\nimport re\nfrom typing import List, Union, Any\nfrom dataclasses import dataclass\n\nimport numpy as np\nimport pandas as pd\nfrom torch import nn\nfrom tqdm.notebook import tqdm\nimport matplotlib.pyplot as plt\nfrom sklearn.neighbors import NearestNeighbors\nfrom transformers import AutoTokenizer, AutoModel","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-05-24T20:50:51.722062Z","iopub.execute_input":"2023-05-24T20:50:51.722462Z","iopub.status.idle":"2023-05-24T20:50:58.634199Z","shell.execute_reply.started":"2023-05-24T20:50:51.722428Z","shell.execute_reply":"2023-05-24T20:50:58.633133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"business\"></a>\n# 1. Business understanding\n[Back to Contents](#contents)\n\n### Objectives\nThe objective of this notebook is to create a *recommendation system* for the customers of *H&M*. As it is stated in the dataset description, the online store offers an extensive selection of products to browse through, hence, that might be overwhelming for the customers, and ultimately, they might just give up on searching. Also, solving this task might have positive implications for the environment - more right choices, less emission from transportation because of returns.\n\n### Resources\nThe resources that we have are:\n- `images/` - images of all products\n- `articles.csv` - metadata for each product/article\n- `customers.csv` - metadata for each customer\n- `transactions_train.csv` - product/article transactions made by the customers; *the product-customer entries are not unique*, therefore there might be multiple transaction for the same product by the same customer\n- `sample_submission.csv` - sample submission file in the correct format; all `customer_ids` in it will be used for predictions, some of the customers' IDs might not be in `transactions_train.csv`.\n\n### Goals\nIn addition to the objectives, we should define what success looks like after the whole process is finished (in technical terms). As described in the [competition evaluation page](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/overview/evaluation), the *Mean Average Precision* (*MAP@12*) metric will be used for evaluation. The higher it is, the better. Also, the missing `customer_ids` in `transactions_train.csv` will be excluded from the scoring.\n\nAdditionally, we might manually test if the model recommendations make sense.\n\nOne more thing to note is that there is no penalty for making all $12$ predictions for a customer, even though the customer has made an order for less than $12$. It is even advantageous to make $12$ predictions.\n\n### Project plan\nThe project will be separated into $5$ more steps:\n- Data understanding - going through each file, exploring its fields and what can be used for our purpose. Here, I'll mostly use the packages `pandas`, `numpy`, `matplotlib`.\n- Data preparation - after I've decided what part of the data will actually be used, I'll  follow the steps for its preparation. The used packages here will be the same as the ones above, with (maybe) the addition of `PyTorch`.\n- Modeling - at this step I'll take the *Natural Language Processing* (*NLP*) route with *BERT* and *Nearest Neighbors* search algorithm. This would require `PyTorch`, `HuggingFace`'s `transformers` and `sklearn`. Initially, I had an idea to use a *Convolutional Autoencoder* and train it on the images, but chose the prior option because it seems that the *Computer Vision* path would take a lot more time and the textual data is quite descriptive.\n- Evaluation - here I will manually test the model, I'll evaluate whether the recommendations are meaningful based on the transactions of certain users. The actual evaluation will be made after submission.\n- System Usage - at this stage I'll describe how to use the defined Recommendation System.\n- Submission - this is the part in which I generate a DataFrame of submission-like format.","metadata":{}},{"cell_type":"markdown","source":"# 2. Data understanding\n[Back to Contents](#contents)\n\nI'll be exploring each file (file group) separately. The first one is `articles.csv`.","metadata":{}},{"cell_type":"code","source":"# Defining the base paths.\nBASE_IN_PATH = \"/kaggle/input/h-and-m-personalized-fashion-recommendations\"\nBASE_OUT_PATH = \"/kaggle/working\"","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:50:58.643603Z","iopub.execute_input":"2023-05-24T20:50:58.643888Z","iopub.status.idle":"2023-05-24T20:50:58.649266Z","shell.execute_reply.started":"2023-05-24T20:50:58.643861Z","shell.execute_reply":"2023-05-24T20:50:58.648161Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"du-articles\"></a>\n## 2.1. Articles\n[Back to Contents](#contents)","metadata":{}},{"cell_type":"markdown","source":"Let's take a look at the dataframe of `articles.csv`. \n\nAt each stage of the analysis I might add a section that looks like this:\n> Sample section\n\nIt will be used to denote what changes I am going to make in the *Data preparation* step.","metadata":{}},{"cell_type":"code","source":"articles_df = pd.read_csv(os.path.join(BASE_IN_PATH, \"articles.csv\"))\nprint(f\"Num. articles: {len(articles_df)}\")\nprint(f\"Unique articles: {len(articles_df.article_id.unique().tolist())}\")\nprint(f\"Columns: {list(articles_df.columns)}\")\narticles_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:50:58.650515Z","iopub.execute_input":"2023-05-24T20:50:58.650841Z","iopub.status.idle":"2023-05-24T20:50:59.780078Z","shell.execute_reply.started":"2023-05-24T20:50:58.650813Z","shell.execute_reply":"2023-05-24T20:50:59.779112Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Analyzing the impact of missing values\n\nWe should find out whether there is missing data. If there is, we should fill it in or remove the rows without data.\n\nAs we can see, there are only $416$ missing descriptions. This means that if we don't want to miss out on these entries, while using an NLP model on these articles, we might need to include some of the other categorical columns.","metadata":{}},{"cell_type":"code","source":"articles_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:50:59.781697Z","iopub.execute_input":"2023-05-24T20:50:59.782023Z","iopub.status.idle":"2023-05-24T20:51:00.241627Z","shell.execute_reply.started":"2023-05-24T20:50:59.781995Z","shell.execute_reply":"2023-05-24T20:51:00.240302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Informatively, let's see the distribution of the types of products (`index_name`s) that don't have descriptions.","metadata":{}},{"cell_type":"code","source":"index_names = articles_df[\"index_name\"].unique().tolist()\n\n# Defining two rows of subplots.\nfig, ax = plt.subplots(2, 1, figsize=(10, 5))\nfig.suptitle(\"Missing description distribution vs full dataset distribution\")\n# Disabling the x axis labels for the upper plot.\nx_axis = ax[0].axes.get_xaxis()\nx_axis.set_visible(False)\n\n# Plotting the histogram of the entries with Missing descriptions.\narticles_df[\n    articles_df[\"detail_desc\"].isnull()\n][\"index_name\"].value_counts().loc[index_names].plot(kind=\"bar\", ax=ax[0])\n# Plotting the histogram of the index_names in the whole dataset.\narticles_df[\"index_name\"].value_counts().loc[index_names].plot(kind=\"bar\", ax=ax[1])","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:51:00.243101Z","iopub.execute_input":"2023-05-24T20:51:00.243481Z","iopub.status.idle":"2023-05-24T20:51:00.789266Z","shell.execute_reply.started":"2023-05-24T20:51:00.243432Z","shell.execute_reply":"2023-05-24T20:51:00.788182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If we then compare the two histograms above, we can see that there are not too many discrepancies in regards to the counts of all `index_name`s. Since the maximum number of missing entries per `index_name` is approx. $90$, in comparison to the even smallest `index_name` (like `Sport` - approx. $2500$ entries), the missing descriptions are insignificant.\n\n> I've decided to use the entries without descriptions, even though they are such a small subset of the whole dataset. That's because in the real world there might be products without descriptions that might be the right fit for the customers. Therefore, the other categorical features might be enough.","metadata":{}},{"cell_type":"markdown","source":"### Sequence analysis in textual columns\nHere, I'll figure out what are the lengths of the sequences (product names and descriptions). Naively, I'll split the descriptions based on a single whitespace.\nAlso, I'll try to filter out some of the columns, based on redundant information.\n\nThis should be done in the preprocessing step:\n> Even though there are some cases in which the text might have a brand name in it, I should lowercase the text fields in the dataframe, since the capitalization in our context won't matter that much. The benefit would be smaller a vocabulary.","metadata":{}},{"cell_type":"code","source":"# Let's start by filling the NaN values in `detail_desc` with an empty string.\narticles_df[\"detail_desc\"] = articles_df[\"detail_desc\"].fillna(\"\")","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:51:00.790770Z","iopub.execute_input":"2023-05-24T20:51:00.791120Z","iopub.status.idle":"2023-05-24T20:51:00.822477Z","shell.execute_reply.started":"2023-05-24T20:51:00.791091Z","shell.execute_reply":"2023-05-24T20:51:00.821395Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_text_size(text: str) -> int:\n    # Substituting multiple whitespaces with one.\n    text = re.sub(r\"\\s+\", \" \", text)\n    # Removing surrounding whitespace.\n    text = text.strip()\n    \n    return len(text.split())\n\n\narticles_df[\"detail_desc_len\"] = articles_df[\"detail_desc\"].apply(get_text_size)\n\nplt.title(\"Description length distribution\")\narticles_df[\"detail_desc_len\"].hist()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:51:00.823983Z","iopub.execute_input":"2023-05-24T20:51:00.824327Z","iopub.status.idle":"2023-05-24T20:51:02.959359Z","shell.execute_reply.started":"2023-05-24T20:51:00.824299Z","shell.execute_reply":"2023-05-24T20:51:02.958312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, most of the articles have descriptions of length around $20$. I'll look into it more in the next section - *Data preparation*.\n\n> There, I'll merge all text columns into one and use it for vector embedding.","metadata":{}},{"cell_type":"code","source":"# Dropping the previously created column. We wouldn't need it anymore.\narticles_df = articles_df.drop(\"detail_desc_len\", axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:51:02.963678Z","iopub.execute_input":"2023-05-24T20:51:02.964004Z","iopub.status.idle":"2023-05-24T20:51:02.992355Z","shell.execute_reply.started":"2023-05-24T20:51:02.963975Z","shell.execute_reply":"2023-05-24T20:51:02.991413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"text_cols = [\n    \"prod_name\",\n    \"product_type_name\",\n    \"product_group_name\",\n    \"graphical_appearance_name\",\n    \"colour_group_name\",\n    \"perceived_colour_value_name\",\n    \"perceived_colour_master_name\",\n    \"department_name\",\n    \"index_name\",\n    \"index_group_name\",\n    \"section_name\",\n    \"garment_group_name\",\n]\n\narticles_df[\"prod_name_len\"] = articles_df[\"prod_name\"].apply(get_text_size)\n\nplt.title(\"Product name length distribution\")\narticles_df[\"prod_name_len\"].hist()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:51:02.994038Z","iopub.execute_input":"2023-05-24T20:51:02.994502Z","iopub.status.idle":"2023-05-24T20:51:03.755078Z","shell.execute_reply.started":"2023-05-24T20:51:02.994467Z","shell.execute_reply":"2023-05-24T20:51:03.753877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles_df = articles_df.drop(\"prod_name_len\", axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:51:03.756337Z","iopub.execute_input":"2023-05-24T20:51:03.756693Z","iopub.status.idle":"2023-05-24T20:51:03.792978Z","shell.execute_reply.started":"2023-05-24T20:51:03.756655Z","shell.execute_reply":"2023-05-24T20:51:03.791588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The product lengths are small (range from $1$ to approx. $7$).\n\nLet's also print some of the values in the other text columns to see what information we might miss if we exclude some of them.","metadata":{}},{"cell_type":"code","source":"for col in text_cols:\n    # Taking a random sample of 10 entries for each text column.\n    print(f\"{col}:\", articles_df.sample(n=10, random_state=1)[col].tolist())\n    print()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:51:03.794368Z","iopub.execute_input":"2023-05-24T20:51:03.794748Z","iopub.status.idle":"2023-05-24T20:51:03.882168Z","shell.execute_reply.started":"2023-05-24T20:51:03.794715Z","shell.execute_reply":"2023-05-24T20:51:03.880928Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After comparing the information of each column to the others, I conclude that I'll use these textual columns (based on mostly unique information):\n- `prod_name`\n- `product_type_name`\n- `product_group_name`\n- `graphical_appearance_name`\n- `colour_group_name`\n- `department_name`\n- `index_name`\n- `detail_desc`\n\nThe data from the other columns can be derived from these ones.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"customers\"></a>\n## 2.2. Customers\n[Back to Contents](#contents)","metadata":{}},{"cell_type":"code","source":"customers_df = pd.read_csv(os.path.join(BASE_IN_PATH, \"customers.csv\"))\nprint(f\"Num. customers: {len(customers_df)}\")\nprint(f\"Unique customers: {len(customers_df.customer_id.unique().tolist())}\")\nprint(f\"Columns: {list(customers_df.columns)}\")\ncustomers_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:51:03.884575Z","iopub.execute_input":"2023-05-24T20:51:03.885554Z","iopub.status.idle":"2023-05-24T20:51:10.559404Z","shell.execute_reply.started":"2023-05-24T20:51:03.885492Z","shell.execute_reply":"2023-05-24T20:51:10.558054Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Analyzing the impact of missing values\n\nWe can see that only two columns don't have missing values `customer_id` and `postal_code`. I guess these were the fields that were mandatory for the customer to fill in.","metadata":{}},{"cell_type":"code","source":"customers_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:51:10.561047Z","iopub.execute_input":"2023-05-24T20:51:10.561821Z","iopub.status.idle":"2023-05-24T20:51:12.251876Z","shell.execute_reply.started":"2023-05-24T20:51:10.561789Z","shell.execute_reply":"2023-05-24T20:51:12.250929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Customer activity","metadata":{}},{"cell_type":"code","source":"# Analyzing the activity of the users.\ncustomers_df[\"Active\"] = customers_df[\"Active\"].fillna(0)\n\nplt.title(\"Inactive/Active (0/1) users\")\ncustomers_df[\"Active\"].hist()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:51:12.253141Z","iopub.execute_input":"2023-05-24T20:51:12.253552Z","iopub.status.idle":"2023-05-24T20:51:12.598328Z","shell.execute_reply.started":"2023-05-24T20:51:12.253497Z","shell.execute_reply":"2023-05-24T20:51:12.596868Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are a lot of inactive users. That shouldn't bother us though! The data from them is still useful.","metadata":{"execution":{"iopub.status.busy":"2023-05-24T06:01:48.067882Z","iopub.execute_input":"2023-05-24T06:01:48.068363Z","iopub.status.idle":"2023-05-24T06:01:48.075434Z","shell.execute_reply.started":"2023-05-24T06:01:48.068330Z","shell.execute_reply":"2023-05-24T06:01:48.074324Z"}}},{"cell_type":"markdown","source":"### Customer age","metadata":{}},{"cell_type":"code","source":"# Figuring out the age distribution of the customers.\nplt.title(\"Customers age distribution\")\ncustomers_df[\"age\"].hist()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:51:12.599878Z","iopub.execute_input":"2023-05-24T20:51:12.601014Z","iopub.status.idle":"2023-05-24T20:51:12.927902Z","shell.execute_reply.started":"2023-05-24T20:51:12.600972Z","shell.execute_reply":"2023-05-24T20:51:12.926735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Of course, the age distribution is scewed towards the younger folks. The interesting part here for me is that there is a slight peek between $40$ and $60$. Maybe when people reach the middle age, they want to gratify themselves more with garment. ","metadata":{}},{"cell_type":"markdown","source":"<a id=\"transactions\"></a>\n## 2.3. Transactions\n[Back to Contents](#contents)","metadata":{}},{"cell_type":"code","source":"transactions_df = pd.read_csv(os.path.join(BASE_IN_PATH, \"transactions_train.csv\"))\nprint(f\"Num. transactions: {len(transactions_df)}\")\nprint(f\"Columns: {list(transactions_df.columns)}\")\ntransactions_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:51:12.929067Z","iopub.execute_input":"2023-05-24T20:51:12.929378Z","iopub.status.idle":"2023-05-24T20:52:24.083875Z","shell.execute_reply.started":"2023-05-24T20:51:12.929351Z","shell.execute_reply":"2023-05-24T20:52:24.082796Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Exploring whether there are any missing values in `transactions.csv`.\ntransactions_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:52:24.085322Z","iopub.execute_input":"2023-05-24T20:52:24.085843Z","iopub.status.idle":"2023-05-24T20:52:43.629212Z","shell.execute_reply.started":"2023-05-24T20:52:24.085802Z","shell.execute_reply":"2023-05-24T20:52:43.628078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Reviewing the price column and its distribution.\ntransactions_df[\"price\"].describe(\n    percentiles=[0.1, 0.25, 0.5, 0.75, 0.9]\n)","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:52:43.630608Z","iopub.execute_input":"2023-05-24T20:52:43.630895Z","iopub.status.idle":"2023-05-24T20:52:45.010475Z","shell.execute_reply.started":"2023-05-24T20:52:43.630869Z","shell.execute_reply":"2023-05-24T20:52:45.009400Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, there are some outliers on boths sides (minimum and maximum). We can conclude this by observing the $90$% percentile and the max. value $\\approx 6 \\times 10^{-1}$, and the $10$% percentile and the min. value $\\approx 1.7 \\times 10^{-5}$.","metadata":{}},{"cell_type":"markdown","source":"Here, I might've delved deeper into the relationships of the fields from `articles.csv` and`customers.csv`, based on the transactions in `transactions_train.csv`, but I decided to dive right into the Data preparation and Modeling, since I took the NLP route and I've already inspected the needed text columns.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"du-conclusion\"></a>\n## 2.4. Conclusion\n[Back to Contents](#contents)\n\nBased on the analysis in this section, I'll preprocess the data and then it will be passed to a model. The preprocessing pipeline will include:\n\n- Fill in the NaN values of `detail_desc` with an empty string (I've already done it)\n- Merge *the chosen* textual columns into one called `text`\n- Lowercase the new column `text`\n- Using *BERT*, embed the values in `text` into vectors\n- Left join the transactions dataframe with the articles dataframe (aquire the new vector column)\n- Calculate the average transaction vector for each customer and add it to `customers_df` (or create a new DataFrame)\n- Add a new column called `bought_articles` in `customers_df`, in which all article IDs of the bough articles for each customer will be saved\n\nLet's continue with the data preparation now!","metadata":{}},{"cell_type":"markdown","source":"<a id=\"dp\"></a>\n# 3. Data preparation\n[Back to Contents](#contents)","metadata":{}},{"cell_type":"markdown","source":"<a id=\"dp-1\"></a>\n## 3.1. Merge chosen textual columns","metadata":{}},{"cell_type":"code","source":"# These were the selected textual columns to be merged in the previous section.\ntext_cols = [\n    \"prod_name\",\n    \"product_type_name\",\n    \"product_group_name\",\n    \"graphical_appearance_name\",\n    \"colour_group_name\",\n    \"department_name\",\n    \"index_name\",\n    \"detail_desc\",\n]\n\n\ndef merge_text_columns(row, columns):\n    texts = []\n    \n    # Looping through the columns except for `detail_desc`.\n    # It will be appended with a '-' separator.\n    for col in columns[:-1]:\n        texts.append(row[col])\n        \n    texts = \", \".join(texts)\n    texts = \" - \".join([texts, row[columns[-1]]])\n    \n    return texts\n\narticles_df[\"text\"] = articles_df.apply(lambda row: merge_text_columns(row, text_cols), axis=1)\narticles_df[\"text\"].head()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:52:45.011897Z","iopub.execute_input":"2023-05-24T20:52:45.012322Z","iopub.status.idle":"2023-05-24T20:52:50.539321Z","shell.execute_reply.started":"2023-05-24T20:52:45.012282Z","shell.execute_reply":"2023-05-24T20:52:50.538272Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lowercase the newly created `text` column.\narticles_df[\"text\"] = articles_df[\"text\"].apply(lambda text: text.lower())","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:52:50.540849Z","iopub.execute_input":"2023-05-24T20:52:50.541249Z","iopub.status.idle":"2023-05-24T20:52:50.643974Z","shell.execute_reply.started":"2023-05-24T20:52:50.541211Z","shell.execute_reply":"2023-05-24T20:52:50.642788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"dp-2\"></a>\n## 3.2. Embed texts into vectors using BERT\n[Back to Contents](#contents)\n\nI'll define a class with which all texts will be embedded as vectors of size $768$.\n\nIt will be done by taking the $1^{st}$ vector of the output, corresponding to the \\[CLS\\] token. It has information about the context of all the data coming from the input, and it should be a good representation.","metadata":{}},{"cell_type":"code","source":"# The fraction of the articles that we are going to embed. I use a subset of the whole dataset\n# because I want to speed up the whole process. A larger subset might also be used, but the preprocessing\n# will take a lot more time.\nEMBED_FRAC = 0.1\n# If this is set to True, the `EMBED_FRAC` fraction of the dataset will be shuffled randomly.\nRANDOMNESS = False\n\n# Maximum length of a tokenized sequence. I chose these values based on the histograms above.\n# BERT uses a subword tokenizer, but still, a lot of samples have much less than 60 words.\nMAX_LEN = 60","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:52:50.645430Z","iopub.execute_input":"2023-05-24T20:52:50.646557Z","iopub.status.idle":"2023-05-24T20:52:50.652563Z","shell.execute_reply.started":"2023-05-24T20:52:50.646495Z","shell.execute_reply":"2023-05-24T20:52:50.651418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class BertVectorizer:\n    \n    def __init__(self):\n        self._model_id = \"bert-base-uncased\"\n        self._tokenizer = AutoTokenizer.from_pretrained(self._model_id)\n        self._base_model = AutoModel.from_pretrained(self._model_id)\n        \n    def embed(self, texts: List[str], max_length=60) -> np.ndarray:\n        \"\"\"Embed `text` into a vector of size 768.\n        Args:\n            text (List[str]): Input text.\n            max_length (int): The maximum length of a text in `texts`. Defaults to 60.\n        \n        Returns:\n            numpy.ndarray: The vector representation of `text`.\n        \"\"\"\n        # Since the input size vary, I pad or truncate, based on the lengths.\n        inputs = self._tokenizer(\n            texts, \n            max_length=max_length, \n            padding=\"max_length\",\n            truncation=True,\n            return_tensors=\"pt\"\n        )\n        # Getting the output tensor of the model. It is of shape (batch_size, seq_len, embedding_size).\n        # Then I get only the vectors for each [CLS] corresponding to each input in `text`.\n        embedding = self._base_model(**inputs).last_hidden_state[:, 0, :].detach()\n        # `output` shape: (batch_size, embedding_size)\n        \n        return embedding.numpy()\n\n\nbert = BertVectorizer()\nembedding = bert.embed(articles_df[\"text\"][:10].tolist(), max_length=MAX_LEN)\nprint(f\"Vector shape: {embedding.shape}\")","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:52:50.654028Z","iopub.execute_input":"2023-05-24T20:52:50.654879Z","iopub.status.idle":"2023-05-24T20:53:09.065312Z","shell.execute_reply.started":"2023-05-24T20:52:50.654846Z","shell.execute_reply":"2023-05-24T20:53:09.064368Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# I could use multithreading here. This method is too slow.\ndef create_embeddings(dataframe: pd.DataFrame, vectorizer: nn.Module, batch_size=5) -> pd.DataFrame:\n    vectors = []\n    \n    for i in tqdm(range(0, len(dataframe), batch_size)):\n        curr_df = dataframe.iloc[i:i + batch_size]\n        vectors.extend(vectorizer.embed(curr_df[\"text\"].tolist()))\n\n    dataframe[\"embedding\"] = vectors\n        \n    return dataframe\n\n\nif RANDOMNESS:\n    print(\"Shuffling the articles dataframe...\")\n    articles_sample = articles_df.sample(frac=EMBED_FRAC, random_state=1)\nelse:\n    articles_sample = articles_df.iloc[:int(EMBED_FRAC * len(articles_df))]\n        \nembedded_articles = create_embeddings(articles_sample, vectorizer=bert, batch_size=100)\nembedded_articles[\"embedding\"]","metadata":{"execution":{"iopub.status.busy":"2023-05-24T20:53:09.066964Z","iopub.execute_input":"2023-05-24T20:53:09.067308Z","iopub.status.idle":"2023-05-24T21:07:59.472577Z","shell.execute_reply.started":"2023-05-24T20:53:09.067277Z","shell.execute_reply":"2023-05-24T21:07:59.471907Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Num. embedded articles:\", len(embedded_articles))","metadata":{"execution":{"iopub.status.busy":"2023-05-24T21:07:59.473910Z","iopub.execute_input":"2023-05-24T21:07:59.474504Z","iopub.status.idle":"2023-05-24T21:07:59.479601Z","shell.execute_reply.started":"2023-05-24T21:07:59.474474Z","shell.execute_reply":"2023-05-24T21:07:59.478684Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"dp-3\"></a>\n## 3.3. Add the vectors into the transactions\n[Back to Contents](#contents)","metadata":{}},{"cell_type":"code","source":"# Getting the subset of `transactions_train.csv` which has these particular article IDs.\nsample_transaction_df = transactions_df[\n    transactions_df[\"article_id\"].isin(\n        embedded_articles[\"article_id\"].tolist()\n    )\n]\nprint(\"Num. of transactions with these article IDs:\", len(sample_transaction_df))\nsample_transaction_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T21:07:59.480925Z","iopub.execute_input":"2023-05-24T21:07:59.481529Z","iopub.status.idle":"2023-05-24T21:08:00.715188Z","shell.execute_reply.started":"2023-05-24T21:07:59.481487Z","shell.execute_reply":"2023-05-24T21:08:00.714182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Executing a left join on the transactions with the embedded articles.\n# This maps the embedding vectors to each transaction.\nembedded_transactions = sample_transaction_df.merge(\n    embedded_articles, \n    how=\"left\", left_on=\"article_id\", right_on=\"article_id\"\n)[[\n    \"customer_id\",\n    \"article_id\",\n    \"price\",\n    \"embedding\"\n]]\nembedded_transactions.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T21:08:00.722023Z","iopub.execute_input":"2023-05-24T21:08:00.722357Z","iopub.status.idle":"2023-05-24T21:08:05.632197Z","shell.execute_reply.started":"2023-05-24T21:08:00.722327Z","shell.execute_reply":"2023-05-24T21:08:05.631257Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"dp-4\"></a>\n## 3.4. Calculate the average transaction vector for each customer\n[Back to Contents](#contents)\n\nHere, for each customer ID, I'll be calculating the average of all embeddings.\n\nIf there are $n$ vectors $\\textbf{x}_i$ for a customer in `embedded_transactions`, and each vector is of shape $\\mathbb{R}^{768}$, the vector describing the customer is defined as:\n\n$$ \\textbf{x}_{cust} = \\frac{1}{n} \\sum_{\\textbf{x}}^{n} x_i $$","metadata":{}},{"cell_type":"code","source":"embedded_customer_ids = embedded_transactions[\"customer_id\"].unique().tolist()\nprint(f\"Num. Customers in the embedded transactions: {len(embedded_customer_ids)}\")","metadata":{"execution":{"iopub.status.busy":"2023-05-24T21:08:05.633650Z","iopub.execute_input":"2023-05-24T21:08:05.634540Z","iopub.status.idle":"2023-05-24T21:08:06.840083Z","shell.execute_reply.started":"2023-05-24T21:08:05.634488Z","shell.execute_reply":"2023-05-24T21:08:06.838999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customer_embeddings = embedded_transactions.groupby([\"customer_id\"])[\"embedding\"].apply(\n    lambda emb: emb.mean()\n).reset_index()\ncustomer_embeddings.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T21:08:06.841322Z","iopub.execute_input":"2023-05-24T21:08:06.841753Z","iopub.status.idle":"2023-05-24T21:09:59.587105Z","shell.execute_reply.started":"2023-05-24T21:08:06.841707Z","shell.execute_reply":"2023-05-24T21:09:59.585594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"dp-5\"></a>\n## 3.5. Add new column for bought articles\n[Back to Contents](#contents)\n\nAdding a new column to the DataFrame `customer_embeddings`, which will be a list of *all bought article IDs* for each Customer.","metadata":{}},{"cell_type":"code","source":"embedded_transactions.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T21:09:59.589646Z","iopub.execute_input":"2023-05-24T21:09:59.590375Z","iopub.status.idle":"2023-05-24T21:09:59.610442Z","shell.execute_reply.started":"2023-05-24T21:09:59.590320Z","shell.execute_reply":"2023-05-24T21:09:59.609411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Group the article IDs based on the customer ID.\nembedded_transactions[\"article_id\"] = embedded_transactions[\"article_id\"].astype(str)\nbought_articles = embedded_transactions.groupby([\"customer_id\"]).agg({\n    \"article_id\": \",\".join\n})\nbought_articles.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T21:09:59.611589Z","iopub.execute_input":"2023-05-24T21:09:59.611884Z","iopub.status.idle":"2023-05-24T21:10:28.510014Z","shell.execute_reply.started":"2023-05-24T21:09:59.611858Z","shell.execute_reply":"2023-05-24T21:10:28.508988Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Add the list of bought articles into the `customer_embeddings` DataFrame.\n# Here, it doesn't matter if it is inner, left or right, since the customer IDs are the same.\ncustomer_embeddings= customer_embeddings.merge(bought_articles, how=\"left\", left_on=\"customer_id\", right_on=\"customer_id\")\ncustomer_embeddings.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T21:10:28.511501Z","iopub.execute_input":"2023-05-24T21:10:28.512391Z","iopub.status.idle":"2023-05-24T21:10:29.585852Z","shell.execute_reply.started":"2023-05-24T21:10:28.512335Z","shell.execute_reply":"2023-05-24T21:10:29.584805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"dp-conclusion\"></a>\n## 3.6. Conclusion\n[Back to Contents](#contents)","metadata":{}},{"cell_type":"markdown","source":"Finally, the preprocessing step is finished. The data is ready to be used for inference.\n\nThe important results of these preprocessing steps are these *two* DataFrames:\n- `embedded_articles` - consists of the articles and their vector representations\n- `customer_embeddings` - consists of all customers and their vector representations (aggregated by the average of their transactional vectors)\n\nLet's print them one more time.","metadata":{}},{"cell_type":"code","source":"embedded_articles = embedded_articles[[\n    \"article_id\",\n    \"embedding\"\n]]\n\nembedded_articles.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T21:10:29.587175Z","iopub.execute_input":"2023-05-24T21:10:29.587484Z","iopub.status.idle":"2023-05-24T21:10:29.610646Z","shell.execute_reply.started":"2023-05-24T21:10:29.587457Z","shell.execute_reply":"2023-05-24T21:10:29.609537Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"customer_embeddings.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T21:10:29.612182Z","iopub.execute_input":"2023-05-24T21:10:29.612582Z","iopub.status.idle":"2023-05-24T21:10:29.636732Z","shell.execute_reply.started":"2023-05-24T21:10:29.612544Z","shell.execute_reply":"2023-05-24T21:10:29.635486Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"modeling\"></a>\n# 4. Modeling\n[Back to Contents](#contents)\n\nIn regards to the modeling, as it can be seen from the previous section, we used *BERT* for the embedding of the texts (multiple string columns) into vectors. \n\nThe remaining objective, which has not been tackled yet, is to actually create a system that uses these vectors and recommends products to the Customers. For this, I'll use the algorithm **Nearest Neighbors** implemented in `sklearn`.\n\nBased on the *Euclidean distance* between the vectors, the Nearest Neighbors algorithm finds top $k$ closest vectors. \n\nThe model will fit the vectors from `embedded_articles` DataFrame. Then, a customer ID will be passed to it. Based on the vector representing the certain user (in `customer_embeddings`), the model will predict top $k$ neighbors, hence, it will recomment top $k$ articles.","metadata":{}},{"cell_type":"code","source":"@dataclass\nclass SystemMetadata:\n    articles_metadata: pd.DataFrame\n    customers_metadata: pd.DataFrame\n\n\nclass ArticleRecommender:\n    \"\"\"Recommendation system for H&M products. Based on previous purchases it\n    suggests new products that the customers might like.\n    \n    Args:\n        metadata (SystemMetadata): Dataclass consisting of vectors describing each article and each customer.   \n    \"\"\"\n    \n    def __init__(self, metadata: SystemMetadata):\n        self._customers_metadata = metadata.customers_metadata\n        self._articles_metadata = metadata.articles_metadata\n        self._articles_metadata[\"article_id\"] = self._articles_metadata[\"article_id\"].astype(str)\n        \n        self._model = NearestNeighbors(n_neighbors=12)\n    \n    def recommend(self, customer_id: str, topk: int = 12) -> List[str]:\n        \"\"\"Recommends `topk` articles based on `customer_id`'s previous purchases.\n        \n        Args:\n            customer_id (str): ID of the customer to which you want to recommend new products.\n            topk (int): Denotes how many suggestions to make. They are ordered (top K) suggestions. Defaults to 12.\n        \n        Returns:\n            List[str]: List of article IDs.\n        \"\"\"\n        \n        # Creating deep copies, since I don't want to alter the original DataFrames.\n        # Also, when we call `recommend()` multiple times, each time we want to have\n        # all the metadata.\n        articles_metadata = self._articles_metadata.copy(deep=True)\n        customers_metadata = self._customers_metadata.copy(deep=True)\n        \n        # Getting the already purchased articles. We want to suggest new things to our Customers, right?\n        customer_purchases = self._get_customer_field_value(\n            customer_id=customer_id,\n            field_name=\"article_id\"\n        ).split(\",\")\n        \n        # Get the DataFrame IDs of the articles that were already purchased by this customer.\n        # Then, remove these entries from the DataFrames.\n        article_df_ids = self._articles_metadata[\n            self._articles_metadata[\"article_id\"].isin(customer_purchases)\n        ].index.tolist()\n        articles_metadata.drop(article_df_ids, inplace=True)\n        customers_metadata.drop(article_df_ids, inplace=True)\n        \n        train_embeddings = self._col2numpy(\n            column=articles_metadata[\"embedding\"].tolist()\n        )\n        \n        # Fitting the model on the article vectors.\n        self._model.fit(train_embeddings)\n        \n        # Getting the vector of the Customer with ID `customer_id`.\n        customer_embedding = self._get_customer_field_value(\n            customer_id, field_name=\"embedding\"\n        )\n        customer_embedding = np.expand_dims(customer_embedding, 0)\n        # Here `customer_embedding` is a NumPy array with shape (1, 768).\n        \n        # Making a prediction.\n        predictions = self._model.kneighbors(\n            customer_embedding, \n            n_neighbors=topk,\n            return_distance=False\n        )[0]\n        \n        # Returning the respective article IDs, based on the predicted indices.\n        return articles_metadata.iloc[\n            predictions.tolist()\n        ][\"article_id\"].tolist()\n        \n    def _col2numpy(self, column: List[np.ndarray]) -> np.ndarray:\n        # Stacking the list of NumPy arrays on the row axis.\n        array = np.stack(column, axis=0)\n        \n        return array\n    \n    def _get_customer_field_value(self, customer_id: str, field_name: str) -> Any:\n        return self._customers_metadata[\n             self._customers_metadata[\"customer_id\"] == customer_id\n        ][field_name].tolist()[0]\n\n\n# Selecting an arbitrary customer.\ncustomer_id = customer_embeddings[\"customer_id\"][42]\n    \nmetadata = SystemMetadata(\n    articles_metadata=embedded_articles,\n    customers_metadata=customer_embeddings\n)\n# Making a recommendation\narticle_recommender = ArticleRecommender(metadata)\nrecommended_articles = article_recommender.recommend(\n    customer_id=customer_id\n)\nprint(f\"Recommended articles for customer with ID '{customer_id}':\\n{recommended_articles}\")","metadata":{"execution":{"iopub.status.busy":"2023-05-24T21:10:29.639278Z","iopub.execute_input":"2023-05-24T21:10:29.639670Z","iopub.status.idle":"2023-05-24T21:10:30.574832Z","shell.execute_reply.started":"2023-05-24T21:10:29.639619Z","shell.execute_reply":"2023-05-24T21:10:30.573730Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"eval\"></a>\n# 5. Evaluation\n[Back to Contents](#contents)","metadata":{}},{"cell_type":"markdown","source":"### Manual evaluation\n\nLet's explore the purchases of the certain Customer and the recommendations. \nHere I validate whether the suggestions for the Customer make sense.\n\nSince there is no test set (there is only a sample submission, with random predictions), the only way in which we can actually evaluate our model using a metric, is by submission. Therefore, I can only manually evaluate it before the submission.\n\nIn a real-world scenario, we might include more people into the manual evaluation. \nFor example, we might use our Customers for evaluation by:\n- Click-Through Rate - percentage of users, who clicked on recommended items\n- Conversion Rate - how many users made a purchase on a recommended item\n- Dwell Time - the time users interact with recommended items\n\nSimilarly, we can use A/B testing - splitting Customers into different groups and comparing performance of different recommendation systems.","metadata":{}},{"cell_type":"code","source":"customer_purchases = customer_embeddings[\n    customer_embeddings[\"customer_id\"] == customer_id\n][\"article_id\"].astype(str).str.split(\",\").tolist()[0]\n\narticles_df[\"article_id\"] = articles_df[\"article_id\"].astype(str)\nprint(\"Customer's choices:\")\narticles_df[articles_df[\"article_id\"].isin(customer_purchases)]","metadata":{"execution":{"iopub.status.busy":"2023-05-24T21:10:30.576215Z","iopub.execute_input":"2023-05-24T21:10:30.576570Z","iopub.status.idle":"2023-05-24T21:10:30.915257Z","shell.execute_reply.started":"2023-05-24T21:10:30.576516Z","shell.execute_reply":"2023-05-24T21:10:30.914270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's take a look at the top 5 recommendations.\nprint(\"Our system's recommendations:\")\narticles_df[articles_df[\"article_id\"].isin(recommended_articles[:5])]","metadata":{"execution":{"iopub.status.busy":"2023-05-24T21:10:30.916548Z","iopub.execute_input":"2023-05-24T21:10:30.916854Z","iopub.status.idle":"2023-05-24T21:10:30.955718Z","shell.execute_reply.started":"2023-05-24T21:10:30.916827Z","shell.execute_reply":"2023-05-24T21:10:30.954935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> As we can see, the results are quite good. The recommendation system suggests different types of dresses and skirts, which is really close to what the certain customer would expect.\nYou can test it out with other customer IDs.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"usage\"></a>\n# 6. System Usage\n[Back to Contents](#contents)\n\nFor the usage of the recommendation system, you can take a look at the [Modeling](#modeling) section.\n\nBasically, we start with the definition of the metadata. It consists of two *DataFrames* - the first one `article_metadata` has all articles and their vector representation, and `customers_metadata` has all customers and their vector representation, based on previous usage.\n```python\nmetadata = SystemMetadata(\n    articles_metadata=embedded_articles,\n    customers_metadata=customer_embeddings\n)\n```\nThen, the next step would be from the perspective of the end-user (given that there is no other API infront of this class):\n```python\narticle_recommender = ArticleRecommender(metadata)\nrecommended_articles = article_recommender.recommend(\n    customer_id=customer_id\n)\n```\nHere, we recommend $k$ articles (by default it is $12$) to a customer with ID `customer_id`.\n\n*In a real-world scenario, we might use a Vector DB for the storage and querying of the vectors.*","metadata":{}},{"cell_type":"markdown","source":"<a id=\"submission\"></a>\n# 7. Submission\n[Back to Contents](#contents)\n\nLet's take a look at the sample submission. I should create the same structure using all customer IDs.","metadata":{}},{"cell_type":"code","source":"sample_submission_df = pd.read_csv(os.path.join(BASE_IN_PATH, \"sample_submission.csv\"))\nsample_submission_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T21:10:30.956925Z","iopub.execute_input":"2023-05-24T21:10:30.957254Z","iopub.status.idle":"2023-05-24T21:10:36.239045Z","shell.execute_reply.started":"2023-05-24T21:10:30.957224Z","shell.execute_reply":"2023-05-24T21:10:36.237970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def generate_submission(system: ArticleRecommender, customer_ids: List[str]) -> pd.DataFrame:\n    recommendations = []\n    progressbar = tqdm(customer_ids)\n    \n    for i, customer_id in enumerate(progressbar):\n        progressbar.set_description(f\"Customer {i + 1}/{len(customer_ids)}\")\n        current_recommendations = system.recommend(\n            customer_id=customer_id\n        )\n        recommendations.append(\" \".join(current_recommendations))\n        \n    return pd.DataFrame.from_dict({\n        \"customer_id\": customer_ids,\n        \"prediction\": recommendations,\n    })\n\n\n# Generating a submission for a small subset of all Customers, just as an example.\nsubmission_df = generate_submission(\n    system=article_recommender,\n    customer_ids=customer_embeddings[\"customer_id\"].tolist()[:100]\n)\nsubmission_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-05-24T21:10:36.243241Z","iopub.execute_input":"2023-05-24T21:10:36.243594Z","iopub.status.idle":"2023-05-24T21:11:32.437688Z","shell.execute_reply.started":"2023-05-24T21:10:36.243564Z","shell.execute_reply":"2023-05-24T21:11:32.436600Z"},"trusted":true},"execution_count":null,"outputs":[]}]}