{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":31254,"databundleVersionId":3103714,"sourceType":"competition"}],"dockerImageVersionId":31012,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# H&M Fashion Recommendations with Product2Vec\n\nUniversity of Colorado Boulder\n\nDTSA5511 Week 6 - Final Project\n\n## 1. Introduction\n\nShopping online has created the opportunity to bring fashion to the customer through personalized digital ad experiences. Developing recommendation algorithms based on customer transaction data for H&M addresses this need to bridge the gap between product and online shoppers. The goal of this project is to predict which article each customer will purchase in a 7-day period following the training data timeframe.\n\nFor this project, I implemented a Product2Vec approach, which adapts the popular Word2Vec natural language processing technique to the task of product recommendation. This method treats a customer's purchase history as a sentence and individual products as words, allowing the model to learn embeddings of products as vector representations based on their co-occurrence patterns in customer purchase histories.\n\nThe overarching purpose of this project is to demonstrate mastery of unsupervised learning technique. Product2Vec approach is particularly well-suited for fashion recommendations because it captures latent relationships between products without requiring explicit attributes as well as naturally modelling the sequntial nature of cutomer purchase much like the sequential nature of words in a sentence. Additionally, it should be able to substitute product relationships.\n\nThe MAP@12 evalutation score of 0.002 this implementation achieved provides a foundation for understanding embedding-based recommendation systems.\n\n## 2. Load Modules and Data\n\nThe H&M dataset contains three primary data sources including over 31 million rows of transactions, more than 105 thousand items, and over 1.3 million users. Each transaction record includes the customer ID, pruchase data, article ID, and price.","metadata":{"_uuid":"78331593-995a-486d-afa6-b1e3182d09fa","_cell_guid":"d8665ee5-ec91-42c8-a41e-14cab36cbf78","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"!pip install -q umap-learn","metadata":{"_uuid":"8bdbee41-16fa-45a4-8a33-dcbb389c283f","_cell_guid":"8c2a2e5f-06a9-4a50-976a-ddc49af8c602","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:49.99178Z","iopub.status.idle":"2025-04-28T06:29:49.992063Z","shell.execute_reply.started":"2025-04-28T06:29:49.991943Z","shell.execute_reply":"2025-04-28T06:29:49.991956Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport umap\nimport random\nimport os\nimport pickle\nimport logging\nimport gc\nfrom pathlib import Path\nfrom concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor\nfrom datetime import datetime, timedelta\nfrom tqdm import tqdm\nfrom collections import Counter\nfrom gensim.models import Word2Vec\nfrom sklearn.manifold import TSNE\nfrom PIL import Image\n\nlogging.basicConfig(format='%(asctime)s : %(levelname)s : %(message)s', level=logging.INFO)","metadata":{"_uuid":"0616666e-3bbe-4566-aed2-0737ea4d7acf","_cell_guid":"1bde8185-ac26-4f60-985d-91a5174513be","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:49.993185Z","iopub.status.idle":"2025-04-28T06:29:49.9936Z","shell.execute_reply.started":"2025-04-28T06:29:49.993389Z","shell.execute_reply":"2025-04-28T06:29:49.993428Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"FULL_RUN = True\nTRAIN_NEW_MODEL = False\nVIZ = False\nEVAL = False \nSUBMIT_ISSUE = True\nif FULL_RUN:\n    SAMPLE_SIZE = None\n    MODEL_PATH = '/kaggle/working/product2vec_model.pkl'  \n    EMBED_PATH = '/kaggle/working/embeddings.pkl'  \nelse:\n    SAMPLE_SIZE = 250000\n    MODEL_PATH = '/kaggle/working/product2vec_model_small.pkl'\n    EMBED_PATH = '/kaggle/working/embeddings_small.pkl'\n    \nTRAIN_SEQUENCES_PATH = '/kaggle/working/train_sequences.pkl'\nHISTORIES_PATH = '/kaggle/working/histories.pkl'\nSEQUENCES_PATH = '/kaggle/working/sequences.pkl'\nGROUND_TRUTH = '/kaggle/working/saved_ground_truth.pkl'\nIMG_PATH = '/kaggle/input/h-and-m-personalized-fashion-recommendations/images'\nSAMPLE_SUBMISSIONS = '/kaggle/input/h-and-m-personalized-fashion-recommendations/sample_submission.csv'\nSUBMISSIONS = '/kaggle/working/submission.csv'","metadata":{"_uuid":"2702f48e-f08c-4138-9077-46dbc6f88c8a","_cell_guid":"c8012b83-d7cb-4e5f-9971-c1391ec6568d","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:49.995185Z","iopub.status.idle":"2025-04-28T06:29:49.995526Z","shell.execute_reply.started":"2025-04-28T06:29:49.995368Z","shell.execute_reply":"2025-04-28T06:29:49.995379Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Load in data:","metadata":{"_uuid":"4a778531-e63e-45b6-9e43-5a5b9e901d23","_cell_guid":"a4b5b522-3148-41a8-87ad-8f0935617baa","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"def load_data(full_run=FULL_RUN, chunk_size=1000000):\n    print(f\"Loading data with {'full' if full_run else 'small'} dataset...\")\n    \n    txs_path = '/kaggle/input/h-and-m-personalized-fashion-recommendations/transactions_train.csv'\n    cxs_path = '/kaggle/input/h-and-m-personalized-fashion-recommendations/customers.csv'\n    axs_path = '/kaggle/input/h-and-m-personalized-fashion-recommendations/articles.csv'\n    \n    articles = pd.read_csv(axs_path, dtype={'article_id': str})\n    customers = pd.read_csv(cxs_path, dtype={'customer_id': str, 'age': 'float32'})\n    customers.fillna(0, inplace=True)\n    \n    if full_run:\n\n        # Read transactions in chunks for efficiency\n        chunks = pd.read_csv(\n            txs_path,\n            dtype={'article_id': str, 'customer_id': str, 'price': 'float32'},\n            parse_dates=['t_dat'],\n            chunksize=chunk_size,\n            usecols=['t_dat', 'customer_id', 'article_id']  # Load only needed columns\n        )\n        \n        transactions = pd.concat([chunk for chunk in tqdm(chunks, desc=\"Loading transactions\")])\n   \n    else:\n        \n        # Sample a smaller dataset\n        transactions = pd.read_csv(\n            txs_path,\n            dtype={'article_id': str, 'customer_id': str, 'price': 'float32'},\n            parse_dates=['t_dat'],\n            nrows=500000,\n            usecols=['t_dat', 'customer_id', 'article_id']\n        )\n        \n        # filter for recent transactions for recency\n        recent_date = transactions['t_dat'].max() - pd.Timedelta(days=30)\n        transactions = transactions[transactions['t_dat'] >= recent_date]\n    \n    print(f\"Transactions: {transactions.shape}, Articles: {articles.shape}, Customers: {customers.shape}\")\n    return transactions, articles, customers","metadata":{"_uuid":"da409b21-333f-4c24-8736-6dbbbdb3fbb8","_cell_guid":"3748799e-bf73-4831-91ef-422f24ccd2ba","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-28T06:29:49.996778Z","iopub.status.idle":"2025-04-28T06:29:49.99706Z","shell.execute_reply.started":"2025-04-28T06:29:49.996934Z","shell.execute_reply":"2025-04-28T06:29:49.996946Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"transactions, articles, customers = load_data(FULL_RUN)\ngc.collect()","metadata":{"_uuid":"d2245b18-a300-4c3d-883d-b7b18239fa48","_cell_guid":"3b24eb2e-862a-47ba-b6da-9308f6ded228","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-28T06:29:49.998033Z","iopub.status.idle":"2025-04-28T06:29:49.99847Z","shell.execute_reply.started":"2025-04-28T06:29:49.998278Z","shell.execute_reply":"2025-04-28T06:29:49.998294Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3. Data Exploration\n\n### 3.1 Cursory View\n\nExplore transactions data:","metadata":{"_uuid":"74eb2f47-db24-45e0-bca0-0fa6b706723b","_cell_guid":"6c200c4e-7d7d-4e16-8c42-d3579638d340","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"print(\"Transactions dataset info:\")\nprint(f\"Shape: {transactions.shape}\")\nprint(f\"Columns: {transactions.columns.tolist()}\")\nprint(\"\\nSample transactions:\")\ntransactions.head()","metadata":{"_uuid":"c6310847-4d9f-4519-93c0-0255b2336fc6","_cell_guid":"18824569-70cb-42db-b6c1-50293c16289c","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:49.99923Z","iopub.status.idle":"2025-04-28T06:29:49.999531Z","shell.execute_reply.started":"2025-04-28T06:29:49.999375Z","shell.execute_reply":"2025-04-28T06:29:49.999386Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Explore articles of clothing:","metadata":{"_uuid":"b1b2216b-299f-4743-bc9e-4f2586e244b0","_cell_guid":"3d2ccd40-9a9a-49c7-9166-b6e2606e5979","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# Explore articles data\nprint(\"Articles dataset info:\")\nprint(f\"Shape: {articles.shape}\")\nprint(f\"Columns: {articles.columns.tolist()}\")\nprint(\"\\nSample articles:\")\narticles.head()","metadata":{"_uuid":"d9381e39-1418-454c-bdc5-2d9079df792c","_cell_guid":"74a885b1-7e48-4f12-b125-35449ce2c78e","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.000926Z","iopub.status.idle":"2025-04-28T06:29:50.001255Z","shell.execute_reply.started":"2025-04-28T06:29:50.001087Z","shell.execute_reply":"2025-04-28T06:29:50.001102Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Explore customer data:","metadata":{"_uuid":"3ace536a-a4dd-47ad-94b3-200256d8a345","_cell_guid":"cead85d2-8675-4b58-96f0-1fd8220e5c11","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# Explore customers data\nprint(\"Customers dataset info:\")\nprint(f\"Shape: {customers.shape}\")\nprint(f\"Columns: {customers.columns.tolist()}\")\nprint(\"\\nSample customers:\")\ncustomers.head()","metadata":{"_uuid":"60b48e3a-6200-4281-a74e-3ae6d5625dcd","_cell_guid":"6504d601-9890-4632-9fcf-615b5135dcf2","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.002788Z","iopub.status.idle":"2025-04-28T06:29:50.003184Z","shell.execute_reply.started":"2025-04-28T06:29:50.002985Z","shell.execute_reply":"2025-04-28T06:29:50.003012Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 3.2 Transactions Analysis\n\nTake a look at purchases by day, week, and month. This is important to identify trends in customer behavior. This may be useful for providing recommendations relevant to the season or current week. Based on the plots below we can see peaks in online shopping during summer months, particularly in June. This aligns with summer sales, change of weather, and end of school-year shopping activities. On the daily chart it is evident that the highest daily spike happens around black friday. These temporal patterns suggest that recommendation strategies might benefits from incorporating season context as fashion preferences vary considerably throughout the year.","metadata":{"_uuid":"47c3b698-58c2-4c95-8bf6-8e39b3e0896b","_cell_guid":"833c7931-0254-4ee5-8c81-ddfd4167b2da","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"if VIZ:\n    print(f\"Date range: {transactions['t_dat'].min()} to {transactions['t_dat'].max()}\")\n    \n    transactions['week'] = transactions['t_dat'].dt.to_period('W')\n    transactions['month'] = transactions['t_dat'].dt.to_period('M')\n    \n    daily_purchases = transactions.groupby('t_dat').size()\n    weekly_purchases = transactions.groupby('week').size()\n    monthly_purchases = transactions.groupby('month').size()\n    \n    # Plot weekly and monthly trends\n    fig, (ax1, ax2, ax3) = plt.subplots(3, 1, figsize=(15, 10))\n    \n    daily_purchases.plot(ax=ax1)\n    ax1.set_title('Daily Purchase Volume')\n    ax1.set_ylabel('Number of Purchases')\n    ax1.set_xlabel('Day')\n    ax1.grid(True)\n    \n    weekly_purchases.plot(ax=ax2)\n    ax2.set_title('Weekly Purchase Volume')\n    ax2.set_ylabel('Number of Purchases')\n    ax2.set_xlabel('Week')\n    ax2.grid(True)\n    \n    monthly_purchases.plot(ax=ax3)\n    ax3.set_title('Monthly Purchase Volume')\n    ax3.set_ylabel('Number of Purchases')\n    ax3.set_xlabel('Month')\n    ax3.grid(True)\n    \n    plt.tight_layout()\n    plt.show()","metadata":{"_uuid":"817c5628-4928-4ea9-9389-e2b85873b21a","_cell_guid":"8b05f336-37bc-4780-ad63-5ad82bd321a4","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.004725Z","iopub.status.idle":"2025-04-28T06:29:50.005133Z","shell.execute_reply.started":"2025-04-28T06:29:50.004927Z","shell.execute_reply":"2025-04-28T06:29:50.004944Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 3.3 Customer Analysis\n\nLook at customer purchase frequency on linear and logular scale so we can see the long tail of the distribution. Every non-outlier customer has fewer than 1250 purchases, which is evident on the log scale. There is a large frequency around zero purchases. It is possible to use this insight when building the model, which is likely to work best for repeat customers. With so many customers purchasing so few items, the recommendation power is limited.","metadata":{"_uuid":"0dd154c3-6e5f-4cea-9114-50a40585e53a","_cell_guid":"c6f5ab2c-70f6-4f4d-a95f-950a7fea603b","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"if VIZ:\n    # Customer purchase frequency\n    customer_purchase_counts = transactions.groupby('customer_id').size()\n    \n    print(f\"Average purchases per customer: {customer_purchase_counts.mean()}\")\n    print(f\"Median purchases per customer: {customer_purchase_counts.median()}\")\n    print(f\"Min purchases per customer: {customer_purchase_counts.min()}\")\n    print(f\"Max purchases per customer: {customer_purchase_counts.max()}\")\n    \n    # lin distribution\n    plt.figure(figsize=(12, 6))\n    customer_purchase_counts.hist(bins=50)\n    plt.title('Distribution of Purchases per Customer')\n    plt.xlabel('Number of Purchases')\n    plt.ylabel('Number of Customers')\n    plt.grid(True)\n    plt.tight_layout()\n    plt.show()\n    \n    # log distribution\n    plt.figure(figsize=(12, 6))\n    customer_purchase_counts.hist(bins=50, log=True)\n    plt.title('Distribution of Purchases per Customer (Log Scale)')\n    plt.xlabel('Number of Purchases')\n    plt.ylabel('Number of Customers (Log Scale)')\n    plt.grid(True)\n    plt.tight_layout()\n    plt.show()","metadata":{"_uuid":"0092599d-e6bd-46d1-b5a3-02cf6708bd04","_cell_guid":"6bb7d11f-682b-4b20-90a6-001e2f13f097","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.006139Z","iopub.status.idle":"2025-04-28T06:29:50.006529Z","shell.execute_reply.started":"2025-04-28T06:29:50.006323Z","shell.execute_reply":"2025-04-28T06:29:50.006339Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 3.4 Product Analysis\n\nNow, we can take a look at product popularity by item and by category. The most popular item is a pair of Jade HW Skinny Denim TRS, and the most populr category is trousers. This is useful information to compare to the recommendations made by the model.","metadata":{"_uuid":"bd1b6164-e1d4-4e5c-8f75-c2332a07420c","_cell_guid":"f61f0c31-1b94-47c8-ab32-3cdc652ab6a5","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"product_popularity = transactions['article_id'].value_counts()\n\nprint(f\"Total unique products purchased: {len(product_popularity)}\")\nprint(f\"Average purchases per product: {product_popularity.mean()}\")\nprint(f\"Median purchases per product: {product_popularity.median()}\")\nprint(f\"Min purchases per product: {product_popularity.min()}\")\nprint(f\"Max purchases per product: {product_popularity.max()}\")","metadata":{"_uuid":"516f9ede-d2bc-491d-a3b3-5781540c9c95","_cell_guid":"ed5f3182-e23a-4a04-87cc-4a68455a82b8","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.007819Z","iopub.status.idle":"2025-04-28T06:29:50.008094Z","shell.execute_reply.started":"2025-04-28T06:29:50.007964Z","shell.execute_reply":"2025-04-28T06:29:50.007974Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if VIZ:\n    # map article_id to product name\n    article_names = articles[['article_id', 'prod_name']].set_index('article_id')['prod_name']\n    \n    # replace article_ids with product names\n    product_popularity_named = product_popularity.copy()\n    product_popularity_named.index = [article_names.get(id, id) for id in product_popularity.index]\n    \n    most_popular_article_id = product_popularity.index[0]\n    most_popular_article_name = product_popularity_named.index[0]\n    \n    plt.figure(figsize=(12, 6))\n    product_popularity_named.head(50).plot(kind='bar')\n    plt.title('Top 50 Most Popular Products')\n    plt.xlabel('Product Name')\n    plt.ylabel('Purchase Count')\n    plt.grid(True, axis='y')\n    plt.tight_layout()\n    plt.show()","metadata":{"_uuid":"8833a01a-1b0f-4873-97ed-3e708aee6d88","_cell_guid":"62767746-fe09-4f30-97d7-8ef646b57a58","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.008844Z","iopub.status.idle":"2025-04-28T06:29:50.009165Z","shell.execute_reply.started":"2025-04-28T06:29:50.009033Z","shell.execute_reply":"2025-04-28T06:29:50.009046Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def display_item(article_id, title=None, figsize=(6, 6)):\n\n    product_name = article_id\n    \n    subfolder = article_id[:3]\n    item_path = os.path.join(IMG_PATH, subfolder, f\"{article_id}.jpg\")\n    \n    plt.figure(figsize=figsize)\n    img = Image.open(item_path)\n    plt.imshow(img)\n    \n    plt.title(title, fontsize=14)\n    \n    plt.axis('off')\n    plt.tight_layout()\n    plt.show()","metadata":{"_uuid":"6ed5ec67-f1ac-4b03-8733-0427f6b92c53","_cell_guid":"7ef70089-e408-4e44-b474-86578618807a","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-28T06:29:50.009692Z","iopub.status.idle":"2025-04-28T06:29:50.009968Z","shell.execute_reply.started":"2025-04-28T06:29:50.009819Z","shell.execute_reply":"2025-04-28T06:29:50.009828Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if VIZ:\n    display_item(most_popular_article_id, f'Most Popular Item: {most_popular_article_name}')","metadata":{"_uuid":"35e929db-e1d8-4d02-9f59-bc67a1365546","_cell_guid":"3650ea2e-bb5f-4e86-b9fa-7ffe9db6e215","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-28T06:29:50.011646Z","iopub.status.idle":"2025-04-28T06:29:50.011968Z","shell.execute_reply.started":"2025-04-28T06:29:50.011839Z","shell.execute_reply":"2025-04-28T06:29:50.011851Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if VIZ:\n    transactions_with_products = transactions.merge(articles, on='article_id', how='left')\n    \n    product_type_counts = transactions_with_products['product_type_name'].value_counts()\n    product_group_counts = transactions_with_products['product_group_name'].value_counts()\n    \n    plt.figure(figsize=(14, 8))\n    product_type_counts.head(20).plot(kind='barh')\n    plt.title('Top 20 Product Types')\n    plt.xlabel('Purchase Count')\n    plt.ylabel('Product Type')\n    plt.grid(True, axis='x')\n    plt.tight_layout()\n    plt.show()","metadata":{"_uuid":"4b688a7f-e064-4fb9-bb3a-df0fcc1b2daa","_cell_guid":"73711aae-5f1b-4fd1-9147-8a52242d1d52","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.012893Z","iopub.status.idle":"2025-04-28T06:29:50.013196Z","shell.execute_reply.started":"2025-04-28T06:29:50.013069Z","shell.execute_reply":"2025-04-28T06:29:50.013084Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4. Data Preprocessing\n\nThe transaction data was preprocessed to create pruchase sequences for each cutomer. This involved sorting transactions by customer and timestamp, grouping transactions by customer, converting each customer's purchases into ordered sequences, and limiting sequence length to the most recent 50 purchases. \n\nIn early iterations, out of memory errors were common, so I implemented parallel processing with chunking for sequence creation, making the preprocessing pipeline scalable to the full 31 million plus transactions.","metadata":{"_uuid":"9017b135-3693-4ac4-adbd-e58baec4257c","_cell_guid":"741fb921-6000-414f-bbe7-bd1c20fe5842","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"def process_group_item(item, min_purchases=2):\n    \n    customer_id, group = item\n    if len(group) >= min_purchases:\n        return customer_id, group['article_id'].tolist()[-50:]\n    return None\n\ndef create_sequences(transactions, min_purchases=2, cache_path=SEQUENCES_PATH, return_dict=False, chunk_size=5000000):\n    \n    cache_file = Path(cache_path)\n    if cache_file.exists():\n        print(\"Loading sequences from cache...\")\n        with open(cache_file, 'rb') as f:\n            return pickle.load(f)\n    \n    print(\"Creating transaction sequences...\")\n    \n    sequences = {} if return_dict else []\n\n    # chunk it up for efficiency\n    for start_idx in tqdm(range(0, len(transactions), chunk_size), desc=\"Processing chunks\"):\n        chunk = transactions[start_idx:start_idx + chunk_size][['customer_id', 't_dat', 'article_id']]\n        chunk_sorted = chunk.sort_values(['customer_id', 't_dat'])\n        \n        groups = chunk_sorted.groupby('customer_id')\n        groups_list = list(groups)  # Convert to list to get length\n        \n        with ProcessPoolExecutor(max_workers=4) as executor:\n            results = executor.map(process_group_item, groups_list, [min_purchases] * len(groups_list))\n            for result in results:\n                if result is not None:\n                    cid, seq = result\n                    if return_dict:\n                        sequences[cid] = seq\n                    else:\n                        sequences.append(seq)\n        \n        del chunk, chunk_sorted, groups_list, results\n        gc.collect()\n    \n    print(f\"Created {len(sequences)} sequences\")\n    \n    with open(cache_file, 'wb') as f:\n        pickle.dump(sequences, f)\n    \n    return sequences","metadata":{"_uuid":"332f79fe-fb58-441b-aabb-09e747b1bac1","_cell_guid":"acca5a0e-d8c7-4711-92fc-a29e819039a2","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.014316Z","iopub.status.idle":"2025-04-28T06:29:50.014589Z","shell.execute_reply.started":"2025-04-28T06:29:50.014468Z","shell.execute_reply":"2025-04-28T06:29:50.014479Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"For evaluation, I created time-based train, validation, and test splits. The training data had all transactions before the validation period. The validation data was limited to 7 days before the test period, and the test data had the final 7 days. This approach mimics a real-wprld scenario where we use historical data to predict future purchases, respecting the temporal and trendy nature of fashion recommendations.","metadata":{"_uuid":"d4f93e91-bc33-4a41-a7f6-b888f25be3e7","_cell_guid":"8d766f80-6d79-4881-a113-402cb04802a2","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"def create_time_splits(transactions, test_days=7, validation_days=7):\n\n    print(\"Creating time splits...\")\n    \n    # set split dates\n    max_date = transactions['t_dat'].max()\n    test_start_date = max_date - timedelta(days=test_days-1)\n    validation_start_date = test_start_date - timedelta(days=validation_days)\n    \n    # split data\n    train = transactions[transactions['t_dat'] < validation_start_date]\n    validation = transactions[(transactions['t_dat'] >= validation_start_date) & \n                              (transactions['t_dat'] < test_start_date)]\n    test = transactions[transactions['t_dat'] >= test_start_date]\n    \n    print(f\"Train period: until {validation_start_date - timedelta(days=1)}\")\n    print(f\"Validation period: {validation_start_date} to {test_start_date - timedelta(days=1)}\")\n    print(f\"Test period: {test_start_date} to {max_date}\")\n    \n    print(f\"Train shape: {train.shape}\")\n    print(f\"Validation shape: {validation.shape}\")\n    print(f\"Test shape: {test.shape}\")\n    \n    return {\n        'train': train,\n        'validation': validation,\n        'test': test\n    }","metadata":{"_uuid":"73a14973-8017-4b4e-bf39-7fc33ed9f634","_cell_guid":"57094ffc-27f0-43e2-96b2-937dbd2c3867","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.015943Z","iopub.status.idle":"2025-04-28T06:29:50.016302Z","shell.execute_reply.started":"2025-04-28T06:29:50.016123Z","shell.execute_reply":"2025-04-28T06:29:50.016138Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def create_ground_truth(transactions, customer_ids_set, start_date, days=7):\n\n    print(\"Creating ground truth...\")\n    \n    end_date = start_date + timedelta(days=days-1)\n    \n    # filter transactions for the evaluation period\n    eval_transactions = transactions[(transactions['t_dat'] >= start_date) & \n                                    (transactions['t_dat'] <= end_date)]\n    \n    # using set for faster lookups for customers we actually care about\n    eval_transactions = eval_transactions[eval_transactions['customer_id'].isin(customer_ids_set)]\n    \n    # ground truth dictionary\n    ground_truth = {}\n    \n    for customer_id, group in eval_transactions.groupby('customer_id'):\n        # get unique purchased articles\n        purchased_articles = group['article_id'].unique().tolist()\n        ground_truth[customer_id] = purchased_articles\n    \n    print(f\"Created ground truth for {len(ground_truth)} customers\")\n    return ground_truth","metadata":{"_uuid":"02cd2c9b-8dca-4abb-a009-393323a81fda","_cell_guid":"7dcb9460-6ef2-4cc1-9e88-cead6d3bf364","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-28T06:29:50.017393Z","iopub.status.idle":"2025-04-28T06:29:50.017787Z","shell.execute_reply.started":"2025-04-28T06:29:50.017591Z","shell.execute_reply":"2025-04-28T06:29:50.017607Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5. Product2Vec Model Implementation\n\nThe core of this project is the Product2Vec model that leverages Word2Vec to learn product embeddings from purchase sequences. The model architecture below has key components, starting with the embedding learning using skip-gram to predict context products given a target product. Each product is represented as a dense 100-dimensional vector. The default window size of 5 captures local purchase patterns. Tunable hyper parameters include vector size, windows, min count of products, skip gram, epochs, and hierarchical softmax. The fine-tuning in this notebook is limited, and will be considered part of future work as much of my effort was spent reducing run time and memory management.\n\nThe generation process involves retrieving a customer's purchase history, aggregaring the embedding of purchased products with recency bias, finding products with embeddings similar to the aggregate vector, excluding already purchased items (how annoying is it to be recommended a product after you purchase it), and recommending the top n most similar products.","metadata":{"_uuid":"9fbd48ed-141b-4a7c-bb93-9720ef88621e","_cell_guid":"0f34107b-0a15-4878-93f2-88757877426d","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"# Product2Vec model for learning product embeddings from purchase sequences using word2vec\nclass Product2Vec:\n\n    def __init__(self, vector_size=100, window=5, min_count=5, \n                 sg=1, workers=4, epochs=10, hs=1):\n\n        self.vector_size = vector_size\n        self.window = window\n        self.min_count = min_count\n        self.sg = sg\n        self.workers = workers\n        self.epochs = epochs\n        self.hs = hs\n        self.model = None\n        self.item_embeddings = None\n        \n    def fit(self, sequences):\n\n        print(f\"Training Product2Vec with {len(sequences)} sequences...\")\n\n        # filter rare items\n        item_counts = Counter(item for seq in sequences for item in seq)\n        sequences = [[item for item in seq if item_counts[item] >= self.min_count] for seq in sequences]\n        sequences = [seq for seq in sequences if seq]\n        \n        # train\n        self.model = Word2Vec(\n            sentences=sequences,\n            vector_size=self.vector_size,\n            window=self.window,\n            min_count=self.min_count,\n            sg=self.sg,\n            workers=self.workers,\n            epochs=self.epochs,\n            hs=self.hs\n        )\n        \n        self.item_embeddings = {item: self.model.wv[item] for item in self.model.wv.index_to_key}\n        print(f\"Model trained with {len(self.item_embeddings)} items\")\n        return self\n        \n    def get_embedding(self, item_id):\n        return self.model.wv[item_id]\n    \n    def get_similar_items(self, item_id, top_n=12): # playing with top_n, was 10\n        return self.model.wv.most_similar(item_id, topn=top_n)\n\n    \n    def get_recommendations(self, history, top_n=12, strategy='mean', \n                           recency_bias=True, exclude_history=True):\n\n        # filter history items that are in the vocabulary\n        valid_history = [item for item in history if item in self.model.wv.key_to_index]\n        \n        if not valid_history:\n            return []\n        \n        # apply recency bias if enabled\n        if recency_bias and len(valid_history) > 1:\n            # more recent items get higher weights\n            weights = np.linspace(0.5, 1.0, len(valid_history))\n        else:\n            weights = np.ones(len(valid_history))\n        \n        # get embeddings for history items\n        history_embeddings = [self.get_embedding(item) for item in valid_history]\n        \n        # combine embeddings based on strategy\n        if strategy == 'mean':\n            user_vector = np.mean(history_embeddings, axis=0)\n        elif strategy == 'weighted_mean':\n            user_vector = np.average(history_embeddings, axis=0, weights=weights)\n        elif strategy == 'recent':\n            # use only the most recent item\n            user_vector = history_embeddings[-1]\n        else:\n            raise ValueError(f\"Unknown strategy: {strategy}\")\n        \n        # find similar items to the user vector\n        items_to_exclude = set(valid_history) if exclude_history else set()\n        similar_items = self._find_similar_items(user_vector, top_n, items_to_exclude)\n        \n        return [item_id for item_id, _ in similar_items]\n    \n    def _find_similar_items(self, vector, top_n=12, exclude_items=None):\n\n        if exclude_items is None:\n            exclude_items = set()\n        \n        # norm the query vector\n        norm_vector = vector / np.linalg.norm(vector)\n        \n        # get all item vectors\n        all_items = [(item, self.model.wv[item]) for item in self.model.wv.index_to_key \n                    if item not in exclude_items]\n        \n        # get similarities\n        similarities = [(item, np.dot(norm_vector, vec / np.linalg.norm(vec))) \n                       for item, vec in all_items]\n        \n        # sort similarities descendingly \n        similarities.sort(key=lambda x: x[1], reverse=True)\n        \n        return similarities[:top_n]\n    \n    def save(self, model_path, embeddings_path):\n        self.model.save(model_path)\n        with open(embeddings_path, 'wb') as f:\n            pickle.dump(self.item_embeddings, f)\n    \n    @classmethod # access class and load model\n    def load(cls, model_path, embeddings_path=None):\n\n        # create a new instance\n        instance = cls()\n        \n        # load the Word2Vec model\n        instance.model = Word2Vec.load(model_path)\n        \n        # load embeddings if path provided, otherwise extract from model\n        if embeddings_path and os.path.exists(embeddings_path):\n            with open(embeddings_path, 'rb') as f:\n                instance.item_embeddings = pickle.load(f)\n        else:\n            instance.item_embeddings = {item: instance.model.wv[item] \n                                     for item in instance.model.wv.index_to_key}\n        \n        return instance","metadata":{"_uuid":"b8f8cc0b-bab9-4898-aee7-68e49a361d91","_cell_guid":"4b2bc0be-615a-4bd1-bfa1-06fbc7a4ba02","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.018726Z","iopub.status.idle":"2025-04-28T06:29:50.019046Z","shell.execute_reply.started":"2025-04-28T06:29:50.018883Z","shell.execute_reply":"2025-04-28T06:29:50.018899Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def generate_user_recommendations(model, user_histories, popular_items, top_n=12, \n                                strategy='weighted_mean', recency_bias=True):\n    \n    print(f\"Generating recommendations for {len(user_histories)} users...\")\n    \n    recommendations = {}\n    \n    for user_id, history in tqdm(user_histories.items()):\n        recs = model.get_recommendations(\n            history=history,\n            top_n=top_n,\n            strategy=strategy,\n            recency_bias=recency_bias\n        )\n        \n        # If not enough recommendations, pad with popular items\n        if len(recs) < top_n:\n            remaining_popular = [item for item in popular_items if item not in recs]\n            needed = top_n - len(recs)\n            recs.extend(remaining_popular[:needed])\n                \n        recommendations[user_id] = recs[:top_n]\n    \n    print(f\"Generated recommendations for {len(recommendations)} users\")\n    return recommendations","metadata":{"_uuid":"8bb440a1-3063-4540-a77c-59ece53449bb","_cell_guid":"fd9cedca-4e33-4aa9-becf-bcfa45257644","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.020999Z","iopub.status.idle":"2025-04-28T06:29:50.021441Z","shell.execute_reply.started":"2025-04-28T06:29:50.021215Z","shell.execute_reply":"2025-04-28T06:29:50.021233Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 6. Model Training","metadata":{"_uuid":"b7b8a472-51d2-4b39-a62c-28f63f48fbe4","_cell_guid":"d990aac6-b4db-4b9a-a318-8ecdc1acb488","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"gc.collect()","metadata":{"_uuid":"a99bbca3-9c24-4a26-b38c-2dd047f23d5a","_cell_guid":"898ed924-5edb-49c7-aef9-4c9319491f5b","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-28T06:29:50.02279Z","iopub.status.idle":"2025-04-28T06:29:50.023171Z","shell.execute_reply.started":"2025-04-28T06:29:50.022989Z","shell.execute_reply":"2025-04-28T06:29:50.023005Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"popular_items = transactions['article_id'].value_counts().head(12).index.tolist()  # Optimize with head()\ngc.collect()\n\nbase_params = {\n    'vector_size': 128,\n    'window': 3,\n    'min_count': 10,\n    'sg': 1,\n    'workers': 4,\n    'epochs': 15,\n    'hs': 0\n}\n\nif TRAIN_NEW_MODEL:\n    \n    train_sequences = create_sequences(\n        transactions,\n        min_purchases=2,\n        cache_path=TRAIN_SEQUENCES_PATH,\n        return_dict=False,\n        chunk_size=5000000\n    )\n    \n    customer_histories = create_sequences(\n        transactions,\n        min_purchases=1,\n        cache_path=HISTORIES_PATH,\n        return_dict=True,\n        chunk_size=5000000\n    )\n    \n    base_model = Product2Vec(**base_params)\n    base_model.fit(train_sequences)\n    base_model.save('/kaggle/working/product2vec_model.pkl', '/kaggle/working/embeddings.pkl')\n    del train_sequences\n    \n    gc.collect()\n        \nelse:\n    \n    base_model = Product2Vec.load('/kaggle/working/product2vec_model.pkl', '/kaggle/working/embeddings.pkl')\n    \n    customer_histories = create_sequences(\n        transactions,\n        min_purchases=1,\n        cache_path=HISTORIES_PATH,\n        return_dict=True,\n        chunk_size=5000000\n    )\n    \n    gc.collect()","metadata":{"_uuid":"2323620e-7c2d-4803-a43a-f8c4d3c82365","_cell_guid":"48aad6b8-017b-4514-9994-e8a49f5c94fb","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.023812Z","iopub.status.idle":"2025-04-28T06:29:50.02418Z","shell.execute_reply.started":"2025-04-28T06:29:50.023997Z","shell.execute_reply":"2025-04-28T06:29:50.024013Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 7. Model Evaluation\n\nThe moddel was evaluated using the Mean Average Precision at 12 (MAP@12) metric, which rewards models for placing relevant items higher in the recommendation list. For each customer, the ground truth consists of the articles they actually purchased in the 7-day validation period. Some optimization efforts includes batch processing of users, precomputation of embeddings, parallel computation, and early filtering of invalid histories.","metadata":{"_uuid":"00a4c24f-899f-4d25-b595-ffdb3b545417","_cell_guid":"6732af73-12b0-41e4-8b12-fac1fdd5810b","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"if EVAL:\n    # Create time-based splits\n    splits = create_time_splits(transactions, test_days=7, validation_days=14)\n    train_df = splits['train']\n    validation_df = splits['validation']\n    test_df = splits['test']","metadata":{"_uuid":"0b84afd5-f4d2-4b03-bde2-296f9dbef2a6","_cell_guid":"8e141a01-551e-4efd-8999-172b743d6905","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.025297Z","iopub.status.idle":"2025-04-28T06:29:50.025682Z","shell.execute_reply.started":"2025-04-28T06:29:50.025498Z","shell.execute_reply":"2025-04-28T06:29:50.025514Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if EVAL:\n    if TRAIN_NEW_MODEL:\n        \n        customer_ids_set = set(customer_histories.keys())\n        \n        validation_ground_truth = create_ground_truth(\n            validation_df, \n            customer_ids_set, \n            validation_df['t_dat'].min(),\n            days=7\n        )\n        \n        with open(GROUND_TRUTH, 'wb') as f:\n            pickle.dump(validation_ground_truth, f)\n            \n    else:\n        \n        with open(GROUND_TRUTH, 'rb') as f:\n            validation_ground_truth = pickle.load(f)\n        \n        print('Loaded ground truth from save file')","metadata":{"_uuid":"2920dec2-17e0-4587-881e-a51b9264d287","_cell_guid":"a1b4760a-32b4-4f7a-8376-4d104bd3e1f5","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.026757Z","iopub.status.idle":"2025-04-28T06:29:50.027016Z","shell.execute_reply.started":"2025-04-28T06:29:50.026899Z","shell.execute_reply":"2025-04-28T06:29:50.02691Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def map_at_k_kaggle(actual, predicted, k=12):\n    \n    ap_at_k = []\n    for act, pred in zip(actual, predicted):\n        \n        if len(act) > 0:\n            pred_k = pred[:k]\n            precision_sum = 0\n            num_correct = 0\n            \n            for i, item in enumerate(pred_k):\n                if item in act:\n                    num_correct += 1\n                    precision_sum += num_correct / (i + 1)\n            \n            ap = precision_sum / min(len(act), k) if len(act) > 0 else 0\n            ap_at_k.append(ap)\n            \n    return np.mean(ap_at_k) if ap_at_k else 0","metadata":{"_uuid":"685f1bb6-ef41-4883-995a-cb3d2483f2cc","_cell_guid":"8c385721-2b18-4147-97c0-80c864b75bc2","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.02852Z","iopub.status.idle":"2025-04-28T06:29:50.028897Z","shell.execute_reply.started":"2025-04-28T06:29:50.028711Z","shell.execute_reply":"2025-04-28T06:29:50.028728Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def evaluate_model(model, user_histories, ground_truth, top_n=12, strategy='weighted_mean', \n                  recency_bias=True, n_workers=4):\n    \n    print(f\"Evaluating model with strategy '{strategy}'...\")\n    \n    # precompute embedding matrix\n    vocab_items = list(model.model.wv.index_to_key)\n    embedding_matrix = np.array([model.model.wv[item] for item in vocab_items])\n    \n    # filter valid users to save time\n    eval_users = [uid for uid in ground_truth.keys() if uid in user_histories and user_histories[uid]]\n    print(f\"Evaluating for {len(eval_users)} users\")\n    \n    def process_batch(user_batch):\n        batch_recs = []\n        batch_actual = []\n        \n        for user_id in user_batch:\n            history = user_histories[user_id]\n            actual = ground_truth[user_id]\n            \n            # filter valid history \n            valid_history = [item for item in history if item in model.model.wv.key_to_index]\n            if not valid_history:\n                continue\n                \n            # compute user vector\n            history_embeddings = np.array([model.model.wv[item] for item in valid_history])\n            weights = np.linspace(0.5, 2.0, len(valid_history)) if recency_bias else np.ones(len(valid_history))\n            user_vector = np.average(history_embeddings, axis=0, weights=weights)\n            \n            # Batch similarity computation\n            norm_user = user_vector / np.linalg.norm(user_vector)\n            norm_embeddings = embedding_matrix / np.linalg.norm(embedding_matrix, axis=1, keepdims=True)\n            similarities = np.dot(norm_embeddings, norm_user)\n            \n            # exlude history items\n            valid_indices = [i for i, item in enumerate(vocab_items) if item not in set(valid_history)]\n            similarities = similarities[valid_indices]\n            valid_items = [vocab_items[i] for i in valid_indices]\n            \n            # Top-N recs\n            top_indices = np.argsort(similarities)[::-1][:top_n]\n            recs = [valid_items[i] for i in top_indices]\n            \n            batch_recs.append(recs)\n            batch_actual.append(actual)\n        \n        return batch_recs, batch_actual\n    \n    batch_size = 1000\n    actual_lists = []\n    pred_lists = []\n    \n    with ThreadPoolExecutor(max_workers=n_workers) as executor:\n        \n        futures = [\n            executor.submit(process_batch, eval_users[i:i + batch_size])\n            for i in range(0, len(eval_users), batch_size)\n        ]\n        \n        for future in tqdm(futures, desc=\"Evaluating batches\"):\n            recs, actual = future.result()\n            pred_lists.extend(recs)\n            actual_lists.extend(actual)\n    \n    map_score = map_at_k_kaggle(actual_lists, pred_lists, k=top_n)\n    \n    # coverage metrics\n    all_items = set(item for items in ground_truth.values() for item in items)\n    recommended_items = set(item for items in pred_lists for item in items)\n    item_coverage = len(recommended_items & all_items) / len(all_items) if all_items else 0\n    customer_coverage = len(pred_lists) / len(ground_truth) if ground_truth else 0\n    \n    return {\n        'map@12': map_score,\n        'item_coverage': item_coverage,\n        'customer_coverage': customer_coverage\n    }","metadata":{"_uuid":"49035fc4-0333-40a8-be7d-9511705612ea","_cell_guid":"bf6592aa-b880-4eaf-9341-beba89e353bd","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-28T06:29:50.030309Z","iopub.status.idle":"2025-04-28T06:29:50.030708Z","shell.execute_reply.started":"2025-04-28T06:29:50.030519Z","shell.execute_reply":"2025-04-28T06:29:50.030535Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The results of evaluation shows that the MAP@12 is 0.002, the item coverage is 0.564, and the customer coverage is 0.723.\n\nWhile the MAP@12 score might seem low, it's actually within the expected range for challenging recommendation tasks like fashion, several factors contribute to this. First, fashion is inherently difficult to predict because it changes with seasons and is usually biased by novelty (perhaps something that can be implemented in future iterations). The sparsity of the dataset also contributes to this since most customers only purchase a tiny fraction of the available items, and not every customer was active.\n\nThe item and customer coverage metrics are more optimistic. They show that the model recommends a substantial portion of relevant items and can generate recommendations for most customers.","metadata":{"_uuid":"dc61e3c9-e267-44ac-ba56-c81c4f9bcca4","_cell_guid":"9bc15240-ba72-49f3-9517-b508104f3634","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"if EVAL:\n    # evaluate model\n    results = evaluate_model(base_model, customer_histories, validation_ground_truth)\n    print(f\"MAP@12: {results['map@12']}\")\n    print(f\"Item coverage: {results['item_coverage']}\")\n    print(f\"Customer coverage: {results['customer_coverage']}\")","metadata":{"_uuid":"ea61c066-7c9b-482d-a6fa-5210bcfd8635","_cell_guid":"217f06db-52f7-44e6-9e6c-f7cf66e961eb","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.032224Z","iopub.status.idle":"2025-04-28T06:29:50.032574Z","shell.execute_reply.started":"2025-04-28T06:29:50.032433Z","shell.execute_reply":"2025-04-28T06:29:50.032451Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 8. Model Analysis and Visualizations","metadata":{"_uuid":"6465ea8b-d15f-4030-9c8e-26ecd76f1a2e","_cell_guid":"94457b12-fae7-449a-a4eb-de3de6de029f","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"Example recommendations for random customers demonstrated that the model recommends products that are stylistically coherent with past purchases. For instance, a customer who purchased casual tops and jeans received recommendations for similar casual items, suggesting the embeddings capture meaningful style patterns. Using the eyeball method, the model seems to be working well.","metadata":{"_uuid":"173d0b94-cf8c-4ed8-97f6-08e576628d24","_cell_guid":"19294172-2602-4015-9686-1b8534834fd9","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"def display_multiple_items(article_ids, articles_df=None, title=\"Product Recommendations\"):\n    \n    n_items = min(len(article_ids), 12)  \n    cols = int(np.ceil(np.sqrt(n_items)))  \n    rows = int(np.ceil(n_items / cols))    \n    \n    fig, axes = plt.subplots(rows, cols, figsize=(15, 10), squeeze=False)\n    fig.suptitle(title, fontsize=16)\n    \n    axes_flat = axes.flatten()\n    \n    for i, article_id in enumerate(article_ids[:12]):\n        if i >= n_items:\n            break\n            \n        # get subfolder from article_id\n        subfolder = article_id[:3]\n        \n        # build the full path correctly\n        full_path = f\"{IMG_PATH}/{subfolder}/{article_id}.jpg\"\n            \n        if os.path.exists(full_path):\n            img = Image.open(full_path)\n            axes_flat[i].imshow(img)\n            axes_flat[i].set_title(f\"Article: {article_id}\", fontsize=10)\n        \n        else:\n            print(f'Image not found for article {article_id}')\n            axes_flat[i].set_facecolor('lightgray')\n            axes_flat[i].text(0.5, 0.5, f'No image\\n{article_id}', \n                         horizontalalignment='center', \n                         verticalalignment='center')\n        \n        axes_flat[i].axis('off')\n    \n    # hide unused subplots\n    for j in range(n_items, len(axes_flat)):\n        axes_flat[j].set_visible(False)\n    \n    plt.tight_layout()\n    plt.subplots_adjust(top=0.9)\n    plt.show()","metadata":{"_uuid":"973636f3-bd8f-4316-91b3-ccbedd811d04","_cell_guid":"aee0b589-2538-480d-959a-b007a570b7bb","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-28T06:29:50.033783Z","iopub.status.idle":"2025-04-28T06:29:50.034062Z","shell.execute_reply.started":"2025-04-28T06:29:50.033939Z","shell.execute_reply":"2025-04-28T06:29:50.03395Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def show_customer_recommendations(customer_id, model, transactions, articles_df=None):\n\n    # get customer purchase history\n    history = transactions[transactions['customer_id'] == customer_id]['article_id'].tolist()\n    \n    # generaet recommendations\n    recommendations = model.get_recommendations(\n        history=history,\n        top_n=12,\n        strategy='weighted_mean',\n        recency_bias=True\n    )\n    \n    # Display purchase history\n    print(f\"\\nCustomer purchase history: {len(history)} items\")\n    if len(history) == 1:\n        display_item(history[0], title=f\"Customer Purchase History {customer_id}\")\n    elif len(history) == 0:\n        print(f'no purchase history for {customer_id}')\n    else:\n        display_multiple_items(history[:12], articles_df=articles_df, title=f\"{customer_id} Purchase History \")\n    \n    # show recommendations\n    print(f\"\\nDisplaying {len(recommendations)} recommendations for customer {customer_id}\")\n    if len(recommendations) == 1:\n        display_item(recommendations, title=f\"Recommendations for Customer {customer_id}\")\n    elif len(recommendations) == 0:\n        print(f'no recommendations history for {customer_id}')\n    else:\n        display_multiple_items(recommendations, articles_df=articles_df, title=f\"Recommendations for {customer_id}\")\n    \n    return recommendations","metadata":{"_uuid":"03200c60-e365-4b0c-adfa-36e37bcaa210","_cell_guid":"1131cb2b-2d66-490a-813d-5a8c0d9a12b4","trusted":true,"collapsed":false,"execution":{"iopub.status.busy":"2025-04-28T06:29:50.035496Z","iopub.status.idle":"2025-04-28T06:29:50.03583Z","shell.execute_reply.started":"2025-04-28T06:29:50.035692Z","shell.execute_reply":"2025-04-28T06:29:50.035707Z"},"jupyter":{"outputs_hidden":false}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def visualize_embeddings(model, articles_df, dim_reduction='tsne', n_items=1000, \n                        color_by='product_group_name', figsize=(12, 10)):\n\n    # get embeddings for visualization\n    vocab_items = list(model.model.wv.key_to_index.keys())\n    \n    # dim reduction options\n    if dim_reduction == 'tsne':\n        reducer = TSNE(n_components=2, random_state=19)\n    elif dim_reduction == 'umap':\n        reducer = umap.UMAP(random_state=19)\n    \n    # sample items if there are too many\n    if len(vocab_items) > n_items:\n        sample_items = random.sample(vocab_items, n_items)\n    else:\n        sample_items = vocab_items\n    \n    item_embeddings = np.array([model.model.wv[item] for item in sample_items])\n    \n    reduced_embeddings = reducer.fit_transform(item_embeddings)\n    \n    # create df for plotting\n    plot_df = pd.DataFrame({\n        'x': reduced_embeddings[:, 0],\n        'y': reduced_embeddings[:, 1],\n        'article_id': sample_items\n    })\n    \n    plot_df = plot_df.merge(articles_df, on='article_id', how='left')\n    \n    plt.figure(figsize=figsize)\n    sns.scatterplot(data=plot_df, x='x', y='y', hue=color_by, alpha=0.7, s=20)\n    plt.title(f'Product Embeddings Visualization ({dim_reduction.upper()})')\n    plt.xlabel(f'{dim_reduction.upper()} Dimension 1')\n    plt.ylabel(f'{dim_reduction.upper()} Dimension 2')\n    \n    plt.legend(bbox_to_anchor=(1.05, 1), loc='upper left')\n    plt.tight_layout()\n    \n    return plt.gcf()","metadata":{"_uuid":"c9c5aaf6-59bd-4cb6-b70d-8ee49c7fd75e","_cell_guid":"7cabde3b-ec4d-47d5-9e3d-858cca95ef26","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.037424Z","iopub.status.idle":"2025-04-28T06:29:50.037714Z","shell.execute_reply.started":"2025-04-28T06:29:50.037573Z","shell.execute_reply":"2025-04-28T06:29:50.037584Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# find similar products for an example item\narticles_in_vocab = list(base_model.model.wv.key_to_index.keys())\nexample_article_id = random.choice(articles_in_vocab)\n\nsimilar_items = base_model.get_similar_items(example_article_id, top_n=5)\n\nprint(f\"Similar products to {example_article_id}:\")\n\nfor similar_id, similarity in similar_items:\n    print(f\"  {similar_id} (Similarity: {similarity})\")","metadata":{"_uuid":"d5e7524e-2a80-4688-a3b4-912d118d0952","_cell_guid":"a8afa1c3-6ba8-44c3-9fe6-f24a80bc524a","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.038782Z","iopub.status.idle":"2025-04-28T06:29:50.039179Z","shell.execute_reply.started":"2025-04-28T06:29:50.03899Z","shell.execute_reply":"2025-04-28T06:29:50.039007Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if VIZ:\n    similar_item_ids = [i for i, x in similar_items] \n    display_item(example_article_id, f'Item {example_article_id}')\n    display_multiple_items(similar_item_ids, title=f\"Items Similar to {example_article_id}\")","metadata":{"_uuid":"5ca3a1cc-c146-434a-a0ed-55f7c58c81d5","_cell_guid":"ae0a697e-e96a-4017-982c-3dc69e32eed8","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.040171Z","iopub.status.idle":"2025-04-28T06:29:50.040571Z","shell.execute_reply.started":"2025-04-28T06:29:50.040361Z","shell.execute_reply":"2025-04-28T06:29:50.040377Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if VIZ:\n    # generate recommendations for a user\n    example_customer_id = transactions[ transactions['article_id'] == example_article_id]['customer_id'].iloc[1]\n    history = transactions[transactions['customer_id'] == example_customer_id]['article_id'].tolist()\n    \n    recommendations = base_model.get_recommendations(\n        history=history,\n        top_n=12,\n        strategy='weighted_mean',\n        recency_bias=True\n    )\n    \n    print(f\"Recommendations for customer {example_customer_id}:\")\n    print(recommendations)\n    print('\\n')\n    \n    display_multiple_items(recommendations, title=\"Product Recommendations\")","metadata":{"_uuid":"528e4d24-c989-4d39-9d2e-06747acb845e","_cell_guid":"19bd592a-462c-4ba3-8cf9-a5a93ab597a8","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.046942Z","iopub.status.idle":"2025-04-28T06:29:50.047339Z","shell.execute_reply.started":"2025-04-28T06:29:50.047173Z","shell.execute_reply":"2025-04-28T06:29:50.047188Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if VIZ:\n    example_customer_id2 = transactions['customer_id'].iloc[np.random.randint(1000)]\n    show_customer_recommendations(example_customer_id2, base_model, transactions, articles_df=articles)","metadata":{"_uuid":"e34b526f-fe0e-4add-8cf9-5d2e138b808c","_cell_guid":"e015e43c-2720-4f92-8d28-6a8f9fb998bc","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.048655Z","iopub.status.idle":"2025-04-28T06:29:50.049366Z","shell.execute_reply.started":"2025-04-28T06:29:50.049162Z","shell.execute_reply":"2025-04-28T06:29:50.04918Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Visualize embeddings using t-SNE:","metadata":{"_uuid":"4bd4ff53-1b09-4ad3-ae99-5f6a86ff0366","_cell_guid":"8bd94fb6-4989-42e7-8e3b-c63527ce158e","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"if VIZ:\n    tsne_fig = visualize_embeddings(\n        base_model, \n        articles, \n        dim_reduction='tsne', \n        n_items=2000, \n        color_by='product_group_name',\n        figsize=(15, 12)\n    )","metadata":{"_uuid":"919a4f03-11dd-4f56-815d-4a5058ab6b4f","_cell_guid":"72fc796b-bb26-43e1-a000-c63552822820","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.050336Z","iopub.status.idle":"2025-04-28T06:29:50.050734Z","shell.execute_reply.started":"2025-04-28T06:29:50.050545Z","shell.execute_reply":"2025-04-28T06:29:50.050562Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Visualize embeddings using UMAP:","metadata":{"_uuid":"b650deb2-36eb-49ef-979a-94da070a564f","_cell_guid":"d674b22e-92f2-4897-8faa-0ff510860347","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"if VIZ:\n    umap_fig = visualize_embeddings(\n        base_model, \n        articles, \n        dim_reduction='umap', \n        n_items=2000, \n        color_by='colour_group_name',\n        figsize=(15, 12)\n    )","metadata":{"_uuid":"98dccf8d-8cd9-4e35-a226-658ad753ab85","_cell_guid":"c3e92717-0ad8-4f12-b16d-adbaf0201eeb","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.052271Z","iopub.status.idle":"2025-04-28T06:29:50.052678Z","shell.execute_reply.started":"2025-04-28T06:29:50.052485Z","shell.execute_reply":"2025-04-28T06:29:50.052507Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 9. Fine-tuning and Optimizations\n\nThe model implementation includes several key optimizations. In order to use memory efficiently on such a large dataset with limited compute power, the data processing was chunked and parallelized using process pool executor and thread pool executor. Garbage collection after memory-intensive operations also proved to spare room in memory. The model hyperparameter tuning focused on the architecture of the model in terms of overall efficiency. For example, hierarchical softmax was implemented to replace negative sampling for efficiency. Also, recency bias in user vector computation was activated to emphasize recent purchases, and batch similarity computation was impemented for faster recommendations. These optimizations were important for making the appraoch scalable to the full dataset while maintaining reasonable computation times.","metadata":{"_uuid":"e28ae3ac-84a1-4282-bf86-de998f82725f","_cell_guid":"d8ac0fbf-06f2-4839-9736-081b723ccea3","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"markdown","source":"## 10. Submission","metadata":{"_uuid":"dbe556f4-0734-4018-8409-e74921d928ed","_cell_guid":"43272cf0-4300-4047-b316-05b69afb32ec","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}},{"cell_type":"code","source":"def create_submission_file(model, transactions, sample_submission_path, output_path, popular_items,\n                          strategy='weighted_mean', recency_bias=True, top_n=12):\n    \n    print(\"Creating submission file...\")\n    \n    sample_submission = pd.read_csv(sample_submission_path)\n    submission_customers = sample_submission['customer_id'].tolist()\n    \n    # load or create customer histories as a dictionary\n    cache_file = Path(HISTORIES_PATH)\n    \n    # check if cache exists and is a dictionary\n    customer_histories = None\n    if cache_file.exists():\n        print(\"Checking cached sequences...\")\n        with open(cache_file, 'rb') as f:\n            customer_histories = pickle.load(f)\n        if not isinstance(customer_histories, dict):\n            print(\"Cached file is not a dictionary, regenerating...\")\n            os.remove(cache_path)  # Delete invalid cache\n            customer_histories = None\n    \n    # generate sequences if no valid cache\n    if customer_histories is None:\n        customer_histories = create_sequences(\n            transactions,\n            min_purchases=1,\n            cache_path=HISTORIES_PATH,\n            return_dict=True,\n            chunk_size=5000000\n        )\n    \n    # precompute embedding matrix\n    vocab_items = list(model.model.wv.index_to_key)\n    embedding_matrix = np.array([model.model.wv[item] for item in vocab_items])\n    \n    def process_batch(customer_batch):\n        batch_recs = {}\n        for customer_id in customer_batch:\n            history = customer_histories.get(customer_id, [])\n            \n            if history:\n                valid_history = [item for item in history if item in model.model.wv.key_to_index]\n                \n                if valid_history:\n                    \n                    # user vector\n                    history_embeddings = np.array([model.model.wv[item] for item in valid_history])\n                    weights = np.linspace(0.5, 1.0, len(valid_history)) if recency_bias else np.ones(len(valid_history))\n                    user_vector = np.average(history_embeddings, axis=0, weights=weights)\n                    \n                    # batch similarity\n                    norm_user = user_vector / np.linalg.norm(user_vector)\n                    norm_embeddings = embedding_matrix / np.linalg.norm(embedding_matrix, axis=1, keepdims=True)\n                    similarities = np.dot(norm_embeddings, norm_user)\n                    \n                    # exclude history\n                    valid_indices = [i for i, item in enumerate(vocab_items) if item not in set(valid_history)]\n                    similarities = similarities[valid_indices]\n                    valid_items = [vocab_items[i] for i in valid_indices]\n                    \n                    # Top-N\n                    top_indices = np.argsort(similarities)[::-1][:top_n]\n                    recs = [valid_items[i] for i in top_indices]\n                    \n                else:\n                    recs = popular_items\n                    \n            else:\n                recs = popular_items\n                \n            batch_recs[customer_id] = recs\n        \n        return batch_recs\n    \n    batch_size = 1000\n    recommendations = {}\n    with ThreadPoolExecutor(max_workers=4) as executor:\n        futures = [\n            executor.submit(process_batch, submission_customers[i:i + batch_size])\n            for i in range(0, len(submission_customers), batch_size)\n        ]\n        for future in tqdm(futures, desc=\"Generating recommendations\"):\n            recommendations.update(future.result())\n    \n    # Create submission\n    submission = pd.DataFrame({\n        'customer_id': submission_customers,\n        'prediction': [' '.join(recommendations.get(cid, popular_items)) for cid in submission_customers]\n    })\n    \n    submission.to_csv(output_path, index=False)\n    print(f\"Submission file saved to {output_path}\")\n    return submission","metadata":{"_uuid":"44a837a0-99d0-4478-a665-96807d7f91a4","_cell_guid":"48eadb4c-b01a-4b60-9bc5-d7bc6adad6e5","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.053389Z","iopub.status.idle":"2025-04-28T06:29:50.053767Z","shell.execute_reply.started":"2025-04-28T06:29:50.053587Z","shell.execute_reply":"2025-04-28T06:29:50.053603Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if SUBMIT_ISSUE:\n    #import shutil\n    #shutil.rmtree('folder_name')\n    import os\n    os.rename('/kaggle/working/submission.csv','/kaggle/working/old_submission.csv')\n    #os.remove('file_path')\n    submission = pd.read_csv('/kaggle/working/old_submission.csv')\n    submission.to_csv(SUBMISSIONS, index=False)\n    ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-04-28T06:32:59.34816Z","iopub.execute_input":"2025-04-28T06:32:59.348617Z","iopub.status.idle":"2025-04-28T06:32:59.353567Z","shell.execute_reply.started":"2025-04-28T06:32:59.34859Z","shell.execute_reply":"2025-04-28T06:32:59.352636Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if TRAIN_NEW_MODEL:\n    submission = create_submission_file(\n        model=base_model,\n        transactions=transactions,\n        sample_submission_path=SAMPLE_SUBMISSIONS,\n        output_path=SUBMISSIONS,\n        popular_items=popular_items\n    )\nelse:\n    submission = pd.read_csv(SUBMISSIONS)","metadata":{"_uuid":"5224a7f3-6b67-4f85-a82e-c95b4986158c","_cell_guid":"8c2d4ad7-2610-48ca-ba97-3ac4e11ceae3","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.055931Z","iopub.status.idle":"2025-04-28T06:29:50.056335Z","shell.execute_reply.started":"2025-04-28T06:29:50.056147Z","shell.execute_reply":"2025-04-28T06:29:50.056164Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# explore submission data\nprint(\"submission dataset info:\")\nprint(f\"Shape: {submission.shape}\")\nprint(f\"Columns: {submission.columns.tolist()}\")\nprint(\"\\nSample customers:\")\nsubmission.head()","metadata":{"_uuid":"ae372c1c-f46f-4d02-916b-89c8e4fd8f5b","_cell_guid":"8713abfa-7eaf-4ffb-bbbb-ab2c3e50194f","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false},"execution":{"iopub.status.busy":"2025-04-28T06:29:50.057795Z","iopub.status.idle":"2025-04-28T06:29:50.058056Z","shell.execute_reply.started":"2025-04-28T06:29:50.057939Z","shell.execute_reply":"2025-04-28T06:29:50.05795Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 11. Conclusion\n\nThe implementation of Product2Vec for H&M fashion recommendations demonstrates the potential of embedding-based approaches for recommendor systems. \n\nSeveral limitations of the current iteration are earmarked for improvement in the future. The model in its current state does not utilize temporal information other than recency bias. Additionally, the model faces a cold-start problem where customers with very few purchases and new products receive lower quality recommendations.\n\nIn future work, I plan to utilize the temporal information more heavily, fine-tune the base hyperparameters like sequence length, window size, skip-gram, and possibly add attention mechanisms. But, as an exercise in unsupervised learning, I believe that this notebook is a solid foundation.","metadata":{"_uuid":"a986ab83-9112-4225-bdd3-b3e1556b09ac","_cell_guid":"ef418b08-e5c8-4c53-9b94-062640a56159","trusted":true,"collapsed":false,"jupyter":{"outputs_hidden":false}}}]}