{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":31254,"databundleVersionId":3103714,"sourceType":"competition"},{"sourceId":3205676,"sourceType":"datasetVersion","datasetId":1945136}],"dockerImageVersionId":30163,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 🛍️ H&M Personalized Fashion Recommendation\n\n## 📌 Project Overview\nThis project builds a **Personalized Fashion Recommendation System** using the **H&M Personalized Fashion Recommendations dataset**.  \nThe goal is to perform customer segmentation using Recency Frequency Monetary (RFM) Model and K-Means clustering to predict and recommend **relevant products** to customers based on their past interactions.\n\n## 📂 Dataset Information\nThe dataset is sourced from [Kaggle's H&M Personalized Fashion Recommendations](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/).  \nIt consists of three main tables:\n\n### 1️⃣ **Transactions (`transactions_train.csv`)**\n- Contains historical purchase data of customers.\n- **Columns:**\n  - `t_dat` → Transaction date\n  - `customer_id` → Unique customer identifier\n  - `article_id` → Product ID\n  - `price` → Price of the product at purchase\n  - `sales_channel_id` → 1 = Online, 2 = Store purchase\n\n### 2️⃣ **Articles (`articles.csv`)**\n- Contains metadata about the fashion items.\n- **Columns:**\n  - `article_id` → Unique product ID\n  - `product_code` → Product category code\n  - `product_type_no` → Product type number\n  - `graphical_appearance_no` → Image style\n  - `perceived_colour_value_id` → Light/Dark color category\n  - `perceived_colour_master_id` → Main color (e.g., red, blue)\n  - `department_no` → Department category\n  - `detail_desc` → Detailed product description\n\n### 3️⃣ **Customers (`customers.csv`)**\n- Contains information about each customer.\n- **Columns:**\n  - `customer_id` → Unique customer identifier\n  - `FN` → Loyalty program membership flag\n  - `Active` → Active status in loyalty program\n  - `club_member_status` → Membership type\n  - `fashion_news_frequency` → Email subscription frequency\n  - `age` → Customer's age\n\n## 🚀 **Project Workflow**\n1. **Exploratory Data Analysis (EDA)**:\n   - Purchase trends over time\n   - Most popular products\n   - Customer behavior insights\n   - RFM and K-Means Clustering\n2. **Building Recommendation Models**:\n   - Recommend Iems Frequently Purchased Together\n3. **Deploying the Model**:\n   - Creating a recommendation function\n   - Generating top-N fashion recommendations per customer","metadata":{}},{"cell_type":"code","source":"%pip install kneed\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:37:37.574036Z","iopub.execute_input":"2025-03-05T01:37:37.574551Z","iopub.status.idle":"2025-03-05T01:37:45.645500Z","shell.execute_reply.started":"2025-03-05T01:37:37.574515Z","shell.execute_reply":"2025-03-05T01:37:45.644611Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport seaborn as sns\nfrom matplotlib import pyplot as plt\nfrom tqdm.notebook import tqdm\nimport cudf, gc\nimport cv2, matplotlib.pyplot as plt\nfrom os.path import exists\nprint('RAPIDS version',cudf.__version__)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:37:45.647289Z","iopub.execute_input":"2025-03-05T01:37:45.647529Z","iopub.status.idle":"2025-03-05T01:37:50.293617Z","shell.execute_reply.started":"2025-03-05T01:37:45.647496Z","shell.execute_reply":"2025-03-05T01:37:50.292964Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Let's load and preview data","metadata":{}},{"cell_type":"code","source":"articles = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/articles.csv\")\ncustomers = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/customers.csv\")\ntransactions = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:37:50.295056Z","iopub.execute_input":"2025-03-05T01:37:50.295526Z","iopub.status.idle":"2025-03-05T01:38:53.483310Z","shell.execute_reply.started":"2025-03-05T01:37:50.295479Z","shell.execute_reply":"2025-03-05T01:38:53.482710Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"articles.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:38:53.484843Z","iopub.execute_input":"2025-03-05T01:38:53.485054Z","iopub.status.idle":"2025-03-05T01:38:53.570564Z","shell.execute_reply.started":"2025-03-05T01:38:53.485028Z","shell.execute_reply":"2025-03-05T01:38:53.569906Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"customers.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:38:53.571439Z","iopub.execute_input":"2025-03-05T01:38:53.571635Z","iopub.status.idle":"2025-03-05T01:38:53.849382Z","shell.execute_reply.started":"2025-03-05T01:38:53.571611Z","shell.execute_reply":"2025-03-05T01:38:53.848697Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"transactions.info()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:38:53.850566Z","iopub.execute_input":"2025-03-05T01:38:53.850839Z","iopub.status.idle":"2025-03-05T01:38:53.859713Z","shell.execute_reply.started":"2025-03-05T01:38:53.850803Z","shell.execute_reply":"2025-03-05T01:38:53.858965Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"We use CUDF to reduce memory usage","metadata":{}},{"cell_type":"code","source":"# LOAD TRANSACTIONS DATAFRAME\ndf = cudf.read_csv('../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv')\nprint('Transactions shape',df.shape)\ndisplay( df.head() )\n\n# REDUCE MEMORY OF DATAFRAME\ndf = df[['customer_id','article_id']]\ndf.customer_id = df.customer_id.str[-16:].str.hex_to_int().astype('int64')\ndf.article_id = df.article_id.astype('int32')\n_ = gc.collect()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:38:53.860714Z","iopub.execute_input":"2025-03-05T01:38:53.860881Z","iopub.status.idle":"2025-03-05T01:38:59.189252Z","shell.execute_reply.started":"2025-03-05T01:38:53.860860Z","shell.execute_reply":"2025-03-05T01:38:59.188650Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Exploratory Data Analysis","metadata":{}},{"cell_type":"markdown","source":"## 1. Customers Data EDA","metadata":{}},{"cell_type":"markdown","source":"There are no missing values or duplicates in customer dataset","metadata":{}},{"cell_type":"code","source":"# Check for missing values\nprint(\"\\nMissing Values in Customers Dataframe:\")\nprint(customers.isnull().sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:38:59.190322Z","iopub.execute_input":"2025-03-05T01:38:59.190589Z","iopub.status.idle":"2025-03-05T01:38:59.471253Z","shell.execute_reply.started":"2025-03-05T01:38:59.190555Z","shell.execute_reply":"2025-03-05T01:38:59.470520Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"customers.shape[0] - customers['customer_id'].nunique()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:38:59.472208Z","iopub.execute_input":"2025-03-05T01:38:59.472405Z","iopub.status.idle":"2025-03-05T01:39:00.041083Z","shell.execute_reply.started":"2025-03-05T01:38:59.472380Z","shell.execute_reply":"2025-03-05T01:39:00.040356Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The most common age is 21 - 23","metadata":{}},{"cell_type":"code","source":"# Exploratory data analysis for customers\n# Histogram of age\nplt.figure(figsize=(10, 6))\nplt.hist(customers['age'].dropna(), bins=30, color ='brown')  # Drop NA values for histogram\nplt.title('Distribution of Customer Age')\nplt.xlabel('Age')\nplt.ylabel('Frequency')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:39:00.043174Z","iopub.execute_input":"2025-03-05T01:39:00.043377Z","iopub.status.idle":"2025-03-05T01:39:00.328065Z","shell.execute_reply.started":"2025-03-05T01:39:00.043352Z","shell.execute_reply":"2025-03-05T01:39:00.327399Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 2. Articles Data EDA","metadata":{}},{"cell_type":"markdown","source":"Ladieswear accounts for the highest portion of Product Index while Sportwear accounts for the lowest","metadata":{}},{"cell_type":"code","source":"# Precompute counts and sort in descending order\nsorted_articles = articles['index_name'].value_counts().reset_index()\nsorted_articles.columns = ['index_name', 'count']\nsorted_articles = sorted_articles.sort_values(by='count', ascending=False)\n\n# Plot\nf, ax = plt.subplots(figsize=(15, 7))\nax = sns.barplot(data=sorted_articles, y='index_name', x='count', palette=\"flare\")\nax.set_xlabel('Count by Index Name')\nax.set_ylabel('Index Name')\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:39:00.329284Z","iopub.execute_input":"2025-03-05T01:39:00.329806Z","iopub.status.idle":"2025-03-05T01:39:00.576766Z","shell.execute_reply.started":"2025-03-05T01:39:00.329765Z","shell.execute_reply":"2025-03-05T01:39:00.576028Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"There are about 45.8K unique products in this dataset","metadata":{}},{"cell_type":"code","source":"for col in articles.columns:\n    if not 'no' in col and not 'code' in col and not 'id' in col:\n        un_n = articles[col].nunique()\n        print(f'no of unique {col}: {un_n}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:39:00.577682Z","iopub.execute_input":"2025-03-05T01:39:00.577894Z","iopub.status.idle":"2025-03-05T01:39:00.676835Z","shell.execute_reply.started":"2025-03-05T01:39:00.577870Z","shell.execute_reply":"2025-03-05T01:39:00.676142Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3. Transactions Data EDA","metadata":{}},{"cell_type":"markdown","source":"There are no missing values in Transactions data","metadata":{}},{"cell_type":"code","source":"print(\"\\nMissing Values in Transactions Dataframe:\")\nprint(transactions.isnull().sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:39:00.677950Z","iopub.execute_input":"2025-03-05T01:39:00.678213Z","iopub.status.idle":"2025-03-05T01:39:03.478790Z","shell.execute_reply.started":"2025-03-05T01:39:00.678174Z","shell.execute_reply":"2025-03-05T01:39:03.478140Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Transactions data ranged from 2018-09-20 to 2020-09-22","metadata":{}},{"cell_type":"code","source":"# Convert 't_dat' to datetime objects\ntransactions['t_dat'] = pd.to_datetime(transactions['t_dat'])\n\n# Analyze transaction dates\nprint(\"\\nTransaction Date Range:\")\nprint(f\"Minimum Date: {transactions['t_dat'].min()}\")\nprint(f\"Maximum Date: {transactions['t_dat'].max()}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:39:03.480246Z","iopub.execute_input":"2025-03-05T01:39:03.480919Z","iopub.status.idle":"2025-03-05T01:39:07.987605Z","shell.execute_reply.started":"2025-03-05T01:39:03.480876Z","shell.execute_reply":"2025-03-05T01:39:07.986879Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Number of transactions peaked during October, December, and April, possibly from End-of-Season Sale","metadata":{}},{"cell_type":"code","source":"# Visualize transaction counts over time\nplt.figure(figsize=(12, 6))\ntransactions['t_dat'].value_counts().sort_index().plot(color = 'brown')\nplt.title('Number of Transactions Over Time')\nplt.xlabel('Date')\nplt.ylabel('Number of Transactions')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:39:07.988830Z","iopub.execute_input":"2025-03-05T01:39:07.989038Z","iopub.status.idle":"2025-03-05T01:39:08.416195Z","shell.execute_reply.started":"2025-03-05T01:39:07.989013Z","shell.execute_reply":"2025-03-05T01:39:08.415504Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Find top 10 bestselling products\nbestsellers = transactions['article_id'].value_counts().head(10)\ntop_10_ids = bestsellers.index.tolist()\n\n# Get details of top 10 products\ntop_products = articles[articles['article_id'].isin(top_10_ids)]\n\n# Merge with sales count\ntop_products = top_products.merge(\n    pd.DataFrame({'article_id': bestsellers.index, 'sales_count': bestsellers.values}),\n    on='article_id'\n)\n\n# Sort by sales count\ntop_products = top_products.sort_values('sales_count', ascending=False)\n\n# Path to images\nBASE = '../input/h-and-m-personalized-fashion-recommendations/images/'\n\n# Create a figure to display the top 10 products\nplt.figure(figsize=(20, 15))\n\n# Loop through top 10 products\nfor i, (_, product) in enumerate(top_products.iterrows()):\n    article_id = product['article_id']\n    \n    # Format image path\n    img_path = f\"{BASE}0{str(article_id)[:2]}/0{str(article_id)}.jpg\"\n    \n    if exists(img_path):\n        # Read and convert image from BGR to RGB\n        img = cv2.imread(img_path)[:,:,::-1]\n        \n        # Create subplot\n        plt.subplot(2, 5, i+1)\n        \n        # Display product information\n        title = f\"#{i+1}: {product['prod_name']}\\n\"\n        title += f\"Sales: {product['sales_count']}\\n\"\n        title += f\"Color: {product['colour_group_name']}\"\n        \n        plt.title(title, fontsize=12)\n        plt.imshow(img)\n        plt.axis('off')\n    else:\n        plt.subplot(2, 5, i+1)\n        plt.text(0.5, 0.5, f\"Image not found\\nArticle ID: {article_id}\", \n                 horizontalalignment='center', verticalalignment='center')\n        plt.axis('off')\n\nplt.suptitle('Top 10 Bestselling Products', fontsize=20)\nplt.tight_layout()\nplt.subplots_adjust(top=0.9)\nplt.savefig('top_10_bestsellers.png', dpi=300)\nplt.show()\n\n# Display detailed information table\nprint(\"Top 10 Bestselling Products Details:\")\ndisplay_cols = ['article_id', 'prod_name', 'product_type_name', 'colour_group_name', \n                'department_name', 'garment_group_name', 'sales_count']\nprint(top_products[display_cols].to_string(index=False))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T02:00:41.806331Z","iopub.execute_input":"2025-03-05T02:00:41.807017Z","iopub.status.idle":"2025-03-05T02:00:51.334255Z","shell.execute_reply.started":"2025-03-05T02:00:41.806979Z","shell.execute_reply":"2025-03-05T02:00:51.333593Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Customers Segmentation using RFM Model","metadata":{}},{"cell_type":"markdown","source":"Load transactions data with cudf for memory optimization ","metadata":{}},{"cell_type":"code","source":"train = cudf.read_csv('../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv')\ntrain['customer_id'] = train['customer_id'].str[-16:].str.hex_to_int().astype('int64')\ntrain['article_id'] = train.article_id.astype('int32')\ntrain.t_dat = cudf.to_datetime(train.t_dat)\ntrain = train[['t_dat','customer_id','article_id']]\ntrain.to_parquet('train.pqt',index=False)\nprint( train.shape )\ntrain.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:39:08.417120Z","iopub.execute_input":"2025-03-05T01:39:08.417305Z","iopub.status.idle":"2025-03-05T01:39:11.847687Z","shell.execute_reply.started":"2025-03-05T01:39:08.417281Z","shell.execute_reply":"2025-03-05T01:39:11.846989Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Create RFM data","metadata":{}},{"cell_type":"code","source":"import cudf\nimport cupy as cp\n\n# Load the dataset (GPU-accelerated DataFrame)\ntrain = cudf.read_csv('../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv')\n\n# Convert customer_id to integer format\ntrain['customer_id'] = train['customer_id'].str[-16:].str.hex_to_int().astype('int64')\ntrain['article_id'] = train['article_id'].astype('int32')\n\n# Convert date to datetime\ntrain['t_dat'] = cudf.to_datetime(train['t_dat'])\n\n# Compute the max order date in the dataset for Recency calculation\nmax_date = train['t_dat'].max()\n\n# Compute Recency (Days since last purchase)\nrecency_df = train.groupby('customer_id').agg({'t_dat': 'max'}).reset_index()\nrecency_df['recency'] = (max_date - recency_df['t_dat']).dt.days\nrecency_df = recency_df[['customer_id', 'recency']]\n\n# Compute Frequency (Total number of purchases)\nfrequency_df = train.groupby('customer_id').agg({'article_id': 'count'}).reset_index()\nfrequency_df = frequency_df.rename(columns={'article_id': 'frequency'})\n\n# Compute Monetary (Unique items bought)\nmonetary_df = train.groupby('customer_id').agg({'article_id': 'nunique'}).reset_index()\nmonetary_df = monetary_df.rename(columns={'article_id': 'monetary'})\n\n# Merge all RFM metrics\nrfm = recency_df.merge(frequency_df, on='customer_id').merge(monetary_df, on='customer_id')\n\n# Sort by best RFM scores (lowest recency, highest frequency & monetary)\nrfm_sorted = rfm.sort_values(by=['recency', 'frequency', 'monetary'], ascending=[True, False, False])\n\n# Save as Parquet file (GPU-optimized format)\nrfm_sorted.to_parquet('rfm_model.pqt', index=False)\n\n# Display top customers\nprint(rfm_sorted.head(10))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:39:11.848690Z","iopub.execute_input":"2025-03-05T01:39:11.848886Z","iopub.status.idle":"2025-03-05T01:39:15.500090Z","shell.execute_reply.started":"2025-03-05T01:39:11.848861Z","shell.execute_reply":"2025-03-05T01:39:15.499410Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Check the distribution of data","metadata":{}},{"cell_type":"code","source":"# Convert cuDF DataFrame to pandas for plotting\nrfm_pandas = rfm_sorted.to_pandas()\n\n# Set up the figure with subplots\nfig, axes = plt.subplots(1, 3, figsize=(18, 6))\nfig.suptitle('RFM Metrics Distributions', fontsize=16)\n\n# Plot Recency distribution\nsns.histplot(rfm_pandas['recency'], kde=True, ax=axes[0], color='skyblue')\naxes[0].set_title('Recency Distribution')\naxes[0].set_xlabel('Days Since Last Purchase')\naxes[0].set_ylabel('Frequency')\n\n# Plot Frequency distribution\nsns.histplot(rfm_pandas['frequency'], kde=True, ax=axes[1], color='lightgreen')\naxes[1].set_title('Frequency Distribution')\naxes[1].set_xlabel('Number of Purchases')\naxes[1].set_ylabel('Frequency')\n\n# Plot Monetary distribution\nsns.histplot(rfm_pandas['monetary'], kde=True, ax=axes[2], color='salmon')\naxes[2].set_title('Monetary Distribution')\naxes[2].set_xlabel('Number of Unique Items Bought')\naxes[2].set_ylabel('Frequency')\n\nplt.tight_layout(rect=[0, 0, 1, 0.95])\nplt.savefig('rfm_distributions.png', dpi=300)\nplt.show()\n\n# For better visualization of skewed distributions, also create log-transformed plots\nfig, axes = plt.subplots(1, 3, figsize=(18, 6))\nfig.suptitle('Log-Transformed RFM Metrics Distributions', fontsize=16)\n\n# Add 1 before log transform to handle zeros\nsns.histplot(np.log1p(rfm_pandas['recency']), kde=True, ax=axes[0], color='skyblue')\naxes[0].set_title('Log-Transformed Recency')\naxes[0].set_xlabel('Log(Days Since Last Purchase + 1)')\n\nsns.histplot(np.log1p(rfm_pandas['frequency']), kde=True, ax=axes[1], color='lightgreen')\naxes[1].set_title('Log-Transformed Frequency')\naxes[1].set_xlabel('Log(Number of Purchases + 1)')\n\nsns.histplot(np.log1p(rfm_pandas['monetary']), kde=True, ax=axes[2], color='salmon')\naxes[2].set_title('Log-Transformed Monetary')\naxes[2].set_xlabel('Log(Number of Unique Items + 1)')\n\nplt.tight_layout(rect=[0, 0, 1, 0.95])\nplt.savefig('rfm_log_distributions.png', dpi=300)\nplt.show()\n\n# Create a pairplot to visualize relationships between RFM metrics\nsns.pairplot(rfm_pandas[['recency', 'frequency', 'monetary']], diag_kind='kde')\nplt.suptitle('RFM Metrics Relationships', y=1.02, fontsize=16)\nplt.savefig('rfm_pairplot.png', dpi=300)\nplt.show()\n\n# Print summary statistics\nprint(\"RFM Summary Statistics:\")\nprint(rfm_pandas[['recency', 'frequency', 'monetary']].describe())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:39:15.501617Z","iopub.execute_input":"2025-03-05T01:39:15.501875Z","iopub.status.idle":"2025-03-05T01:41:41.710188Z","shell.execute_reply.started":"2025-03-05T01:39:15.501841Z","shell.execute_reply":"2025-03-05T01:41:41.709439Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from scipy.stats import skew\n\n# For each column in rfm_sorted\nfor col in ['recency', 'frequency', 'monetary']:\n    skewness = skew(rfm_sorted[col].values_host)\n    print(f\"Skewness of {col}: {skewness:.5f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:41:41.711216Z","iopub.execute_input":"2025-03-05T01:41:41.711411Z","iopub.status.idle":"2025-03-05T01:41:41.763576Z","shell.execute_reply.started":"2025-03-05T01:41:41.711386Z","shell.execute_reply":"2025-03-05T01:41:41.762820Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Apply log transformation to combat severe positive skewness","metadata":{}},{"cell_type":"code","source":"# For frequency column\nfreq_data = rfm_pandas['frequency'].values\noriginal_freq_skew = skew(freq_data)\n\n# Log transformation for frequency\nfreq_log = np.log1p(freq_data)  # log(x+1) to handle zeros\nlog_freq_skew = skew(freq_log)\n\n# For monetary column\nmon_data = rfm_pandas['monetary'].values\noriginal_mon_skew = skew(mon_data)\n\n# Log transformation for monetary\nmon_log = np.log1p(mon_data)  # log(x+1) to handle zeros\nlog_mon_skew = skew(mon_log)\n\n# Print summary of skewness values before and after log transformation\nprint(\"Skewness before and after log transformation:\")\nprint(f\"Frequency - Original: {original_freq_skew:.5f}, Log-transformed: {log_freq_skew:.5f}\")\nprint(f\"Monetary - Original: {original_mon_skew:.5f}, Log-transformed: {log_mon_skew:.5f}\")\nprint(f\"Frequency - Skewness reduction: {original_freq_skew - log_freq_skew:.5f} ({(1 - log_freq_skew/original_freq_skew)*100:.2f}%)\")\nprint(f\"Monetary - Skewness reduction: {original_mon_skew - log_mon_skew:.5f} ({(1 - log_mon_skew/original_mon_skew)*100:.2f}%)\")\n\n# Add the log-transformed columns to the rfm_pandas dataframe\nrfm_pandas['frequency_log'] = freq_log\nrfm_pandas['monetary_log'] = mon_log\n\n# Show the first few rows of the dataframe with the new columns\nprint(\"\\nDataframe with log-transformed columns:\")\nprint(rfm_pandas[['frequency', 'frequency_log', 'monetary', 'monetary_log']].head())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:41:41.764588Z","iopub.execute_input":"2025-03-05T01:41:41.764807Z","iopub.status.idle":"2025-03-05T01:41:41.877377Z","shell.execute_reply.started":"2025-03-05T01:41:41.764783Z","shell.execute_reply":"2025-03-05T01:41:41.876600Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Since K-Means model is sensitive to feature scale, we will normalize data","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import StandardScaler\n\n# Create a copy of the DataFrame to avoid modifying the original\nrfm_scaled = rfm_pandas.copy()\n\n# Features to scale\nfeatures = ['recency', 'frequency_log', 'monetary_log']\n\n# Initialize the scaler\nscaler = StandardScaler()\n\n# Fit and transform the features\nscaled_features = scaler.fit_transform(rfm_scaled[features])\n\n# Create a DataFrame with the scaled features\nscaled_df = pd.DataFrame(scaled_features, columns=[f'{col}_scaled' for col in features])\n\n# Add the scaled features to the original DataFrame\nrfm_scaled = pd.concat([rfm_scaled, scaled_df], axis=1)\n\n# Display the first few rows of the DataFrame with original and scaled features\nprint(rfm_scaled[['customer_id', 'recency', 'recency_scaled', \n                 'frequency_log', 'frequency_log_scaled',\n                 'monetary_log', 'monetary_log_scaled']].head())\n\n# Verify that scaled features have mean ≈ 0 and std ≈ 1\nprint(\"\\nMean of scaled features:\")\nprint(rfm_scaled[[col for col in rfm_scaled.columns if col.endswith('_scaled')]].mean())\n\nprint(\"\\nStandard deviation of scaled features:\")\nprint(rfm_scaled[[col for col in rfm_scaled.columns if col.endswith('_scaled')]].std())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:41:41.878304Z","iopub.execute_input":"2025-03-05T01:41:41.878499Z","iopub.status.idle":"2025-03-05T01:41:42.364183Z","shell.execute_reply.started":"2025-03-05T01:41:41.878466Z","shell.execute_reply":"2025-03-05T01:41:42.363345Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Clustering with K-Means ","metadata":{}},{"cell_type":"markdown","source":"## First we need to choose the optimal k neighbors ","metadata":{}},{"cell_type":"markdown","source":"From the Elbow Method, we will look for the \"bend\" in the inertia curve where adding more clusters doesn't significantly reduce the within-cluster sum of squares.","metadata":{}},{"cell_type":"code","source":"# Import GPU-accelerated libraries\nimport cudf\nimport cuml\nfrom sklearn.cluster import KMeans\nimport numpy as np\nimport matplotlib.pyplot as plt\nfrom kneed import KneeLocator\n\n# Convert pandas DataFrame to cuDF\nfeatures_gpu = cudf.DataFrame(rfm_scaled[['recency_scaled', 'frequency_log_scaled', 'monetary_log_scaled']])\n\n# Elbow Method with GPU acceleration\ninertia = []\nk_range = range(2, 15)\n\nfor k in k_range:\n    kmeans = cuKMeans(n_clusters=k, random_state=42, n_init=10)\n    kmeans.fit(features_gpu)\n    inertia.append(kmeans.inertia_)\n\n# Use kneed to find the elbow point\nkl = KneeLocator(list(k_range), inertia, curve='convex', direction='decreasing')\nelbow_k = kl.elbow\nprint(f\"Elbow point detected at k = {elbow_k}\")\n\n# Plot the elbow curve\nplt.figure(figsize=(10, 6))\nplt.plot(k_range, inertia, 'bo-')\nplt.xlabel('Number of Clusters (k)')\nplt.ylabel('Inertia')\nplt.title('Elbow Method for Optimal k')\nplt.grid(True)\nplt.savefig('elbow_method.png')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:48:45.092936Z","iopub.execute_input":"2025-03-05T01:48:45.093233Z","iopub.status.idle":"2025-03-05T01:49:28.151029Z","shell.execute_reply.started":"2025-03-05T01:48:45.093203Z","shell.execute_reply":"2025-03-05T01:49:28.150282Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Apply K-Means with the K = 5","metadata":{}},{"cell_type":"code","source":"features_for_clustering = rfm_scaled[['recency_scaled', 'frequency_log_scaled', 'monetary_log_scaled']]\nrecommended_k = elbow_k\n\n# Apply K-means with the recommended k\nfinal_kmeans = KMeans(n_clusters=elbow_k, random_state=42, n_init=10)\nrfm_scaled['cluster'] = final_kmeans.fit_predict(features_for_clustering)\n\n# Analyze the resulting clusters\ncluster_summary = rfm_scaled.groupby('cluster').agg({\n    'recency': 'mean',\n    'frequency': 'mean',\n    'monetary': 'mean',\n    'recency_scaled': 'mean',\n    'frequency_log_scaled': 'mean',\n    'monetary_log_scaled': 'mean',\n    'customer_id': 'count'\n}).rename(columns={'customer_id': 'count'})\n\nprint(\"\\nCluster Summary:\")\nprint(cluster_summary)\n\n# Visualize the clusters\nplt.figure(figsize=(12, 8))\nscatter = plt.scatter(rfm_scaled['frequency_log_scaled'], \n                     rfm_scaled['monetary_log_scaled'],\n                     c=rfm_scaled['cluster'], \n                     cmap='viridis', \n                     alpha=0.5)\nplt.colorbar(scatter, label='Cluster')\nplt.xlabel('Frequency (Log, Scaled)')\nplt.ylabel('Monetary (Log, Scaled)')\nplt.title(f'Customer Segments with K-means (k={recommended_k})')\nplt.grid(True, alpha=0.3)\nplt.savefig('cluster_visualization.png')\nplt.show()\n\n# 3D visualization\nfrom mpl_toolkits.mplot3d import Axes3D\n\nfig = plt.figure(figsize=(12, 10))\nax = fig.add_subplot(111, projection='3d')\nscatter = ax.scatter(rfm_scaled['recency_scaled'], \n                    rfm_scaled['frequency_log_scaled'],\n                    rfm_scaled['monetary_log_scaled'],\n                    c=rfm_scaled['cluster'],\n                    cmap='viridis',\n                    alpha=0.5)\nax.set_xlabel('Recency (Scaled)')\nax.set_ylabel('Frequency (Log, Scaled)')\nax.set_zlabel('Monetary (Log, Scaled)')\nplt.colorbar(scatter, label='Cluster')\nplt.title(f'3D Visualization of Customer Segments (k={recommended_k})')\nplt.savefig('3d_cluster_visualization.png')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T01:49:44.853895Z","iopub.execute_input":"2025-03-05T01:49:44.854171Z","iopub.status.idle":"2025-03-05T01:51:33.003167Z","shell.execute_reply.started":"2025-03-05T01:49:44.854141Z","shell.execute_reply":"2025-03-05T01:51:33.002284Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Interpreting the K-means Clusters (k=5)\n\nBased on the 3D visualization and cluster summary, I can provide an interpretation of your 5 customer segments:\n\n## Cluster 0 (Teal/Blue-Green, Top)\n- **Characteristics**: High recency (1.58 scaled), low frequency (-1.06), low monetary value (-1.05)\n- **Interpretation**: These appear to be **recent but low-value customers** who have made purchases recently but don't buy frequently and spend less than average.\n- **Size**: 234,063 customers\n\n## Cluster 1 (Green, Middle-Right)\n- **Characteristics**: Negative recency (-0.33), slightly below average frequency (-0.91), below average monetary (-0.91)\n- **Interpretation**: These represent **inactive, low-value customers** who haven't purchased recently and have below-average purchase frequency and spending.\n- **Size**: 311,669 customers\n\n## Cluster 2 (Yellow, Bottom-Left)\n- **Characteristics**: Very negative recency (-0.79), highest frequency (1.48), highest monetary (1.49)\n- **Interpretation**: These are your **inactive high-value customers** - they haven't purchased recently but historically had high frequency and spending. These are valuable customers you might want to win back.\n- **Size**: 273,267 customers\n\n## Cluster 3 (Dark Blue, Bottom)\n- **Characteristics**: Positive recency (1.01), average frequency (-0.01), average monetary (-0.02)\n- **Interpretation**: These are **recent average-value customers** who have made purchases recently with average frequency and spending.\n- **Size**: 177,755 customers\n\n## Cluster 4 (Purple, Bottom-Right)\n- **Characteristics**: Negative recency (-0.63), above average frequency (0.35), above average monetary (0.35)\n- **Interpretation**: These represent **moderately inactive, above-average value customers** who haven't purchased very recently but have above-average purchase frequency and spending.\n- **Size**: 365,527 customers (largest segment)\n\n## Business Recommendations:\n1. **Cluster 0**: Convert these recent customers into repeat buyers with loyalty incentives\n2. **Cluster 1**: Consider targeted reactivation campaigns or potentially deprioritize\n3. **Cluster 2**: High-priority win-back campaigns as they were valuable customers\n4. **Cluster 3**: Nurture these customers to increase their purchase frequency\n5. **Cluster 4**: Reactivation campaigns to bring these valuable customers back\n\nThe visualization shows clear separation between segments, indicating your clustering has effectively identified distinct customer groups based on their RFM characteristics.","metadata":{}},{"cell_type":"markdown","source":"| Cluster | Cluster Name | Customer Count | Recommended Actions |\n|---------|--------------|----------------|---------------------|\n| 0 | Recent Low-Value | 234,063 | Offer loyalty programs to increase purchase frequency and value |\n| 1 | Inactive Low-Value | 311,669 | Send re-engagement emails with special discounts or remove from active marketing |\n| 2 | Dormant High-Value | 273,267 | Launch win-back campaigns with personalized offers based on past purchases |\n| 3 | Active Average-Value | 177,755 | Encourage repeat purchases with cross-sell/upsell opportunities |\n| 4 | Semi-Active Above-Average | 365,527 | Send targeted promotions to reactivate and increase purchase frequency |","metadata":{}},{"cell_type":"markdown","source":"# Find Items Purchased Together","metadata":{}},{"cell_type":"code","source":"# LOAD TRANSACTIONS DATAFRAME\ndf = cudf.read_csv('../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv')\nprint('Transactions shape',df.shape)\ndisplay( df.head() )\n\n# REDUCE MEMORY OF DATAFRAME\ndf = df[['customer_id','article_id']]\ndf.customer_id = df.customer_id.str[-16:].str.hex_to_int().astype('int64')\ndf.article_id = df.article_id.astype('int32')\n_ = gc.collect()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T02:05:56.025325Z","iopub.execute_input":"2025-03-05T02:05:56.025606Z","iopub.status.idle":"2025-03-05T02:05:59.213988Z","shell.execute_reply.started":"2025-03-05T02:05:56.025575Z","shell.execute_reply":"2025-03-05T02:05:59.213151Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# FIND ITEMS PURCHASED TOGETHER\nvc = df.article_id.value_counts()\npairs = {}\nfor j,i in enumerate(vc.index.values[1000:1032]):\n    #if j%10==0: print(j,', ',end='')\n    USERS = df.loc[df.article_id==i.item(),'customer_id'].unique()\n    vc2 = df.loc[(df.customer_id.isin(USERS))&(df.article_id!=i.item()),'article_id'].value_counts()\n    pairs[i.item()] = [vc2.index[0], vc2.index[1], vc2.index[2]]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T02:06:06.979870Z","iopub.execute_input":"2025-03-05T02:06:06.980527Z","iopub.status.idle":"2025-03-05T02:06:14.787811Z","shell.execute_reply.started":"2025-03-05T02:06:06.980491Z","shell.execute_reply":"2025-03-05T02:06:14.786980Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Display Item Purchased Together\nWhen customers bought the item in the 1st column below, then those customers also bought the items in the 2nd, 3rd, and 4th column too!","metadata":{}},{"cell_type":"code","source":"items = cudf.read_csv('../input/h-and-m-personalized-fashion-recommendations/articles.csv')\nBASE = '../input/h-and-m-personalized-fashion-recommendations/images/'\n\nfor i,(k,v) in enumerate( pairs.items() ):\n    name1 = BASE+'0'+str(k)[:2]+'/0'+str(k)+'.jpg'\n    name2 = BASE+'0'+str(v[0])[:2]+'/0'+str(v[0])+'.jpg'\n    name3 = BASE+'0'+str(v[1])[:2]+'/0'+str(v[1])+'.jpg'\n    name4 = BASE+'0'+str(v[2])[:2]+'/0'+str(v[2])+'.jpg'\n    if exists(name1) & exists(name2) & exists(name3) & exists(name4):\n        plt.figure(figsize=(20,5))\n        img1 = cv2.imread(name1)[:,:,::-1]\n        img2 = cv2.imread(name2)[:,:,::-1]\n        img3 = cv2.imread(name3)[:,:,::-1]\n        img4 = cv2.imread(name4)[:,:,::-1]\n        plt.subplot(1,4,1)\n        plt.title('When customers buy this',size=18)\n        plt.imshow(img1)\n        plt.subplot(1,4,2)\n        plt.title('They buy this',size=18)\n        plt.imshow(img2)\n        plt.subplot(1,4,3)\n        plt.title('They buy this',size=18)\n        plt.imshow(img3)\n        plt.subplot(1,4,4)\n        plt.title('They buy this',size=18)\n        plt.imshow(img4)\n        plt.show()\n    if i==63: break","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-03-05T02:06:30.309471Z","iopub.execute_input":"2025-03-05T02:06:30.310172Z","iopub.status.idle":"2025-03-05T02:07:05.828502Z","shell.execute_reply.started":"2025-03-05T02:06:30.310136Z","shell.execute_reply":"2025-03-05T02:07:05.827784Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Recommend Items Frequently Purchased Together\nThis notebook demonstrates how recommending items that are frequently purchased together is effective. The current best scoring public notebook [here][1] recommends to customers those customers' last purchases and scores public LB 0.020. In this notebook here, we will begin with that idea and add recommending items that are frequently purchased together with a customers' previous purchaes. This notebook improves the LB and scores LB 0.021. This notebook's strategy is as follows:\n* recommend items previously purchased [idea here][1]\n* recommend items that are bought together with previous purchases [idea here][2]\n* recommend popular items [idea here][1]\n\n[1]: https://www.kaggle.com/hengzheng/time-is-our-best-friend-v2\n[2]: https://www.kaggle.com/cdeotte/customers-who-bought-this-frequently-buy-this","metadata":{}},{"cell_type":"code","source":"train = cudf.read_csv('../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv')\ntrain['customer_id'] = train['customer_id'].str[-16:].str.hex_to_int().astype('int64')\ntrain['article_id'] = train.article_id.astype('int32')\ntrain.t_dat = cudf.to_datetime(train.t_dat)\ntrain = train[['t_dat','customer_id','article_id']]\ntrain.to_parquet('train.pqt',index=False)\nprint( train.shape )\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2025-03-05T02:07:40.606049Z","iopub.execute_input":"2025-03-05T02:07:40.606797Z","iopub.status.idle":"2025-03-05T02:07:44.175572Z","shell.execute_reply.started":"2025-03-05T02:07:40.606757Z","shell.execute_reply":"2025-03-05T02:07:44.174836Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Find Each Customer's Last Week of Purchases\nOur final predictions will have the row order from of our dataframe. Each row of our dataframe will be a prediction. We will create the `predictionstring` later by `train.groupby('customer_id').article_id.sum()`. Since `article_id` is a string, when we groupby sum, it will concatenate all the customer predictions into a single string. It will also create the string in the order of the dataframe. So as we proceed in this notebook, we will order the dataframe how we want our predictions ordered.","metadata":{}},{"cell_type":"code","source":"tmp = train.groupby('customer_id').t_dat.max().reset_index()\ntmp.columns = ['customer_id','max_dat']\ntrain = train.merge(tmp,on=['customer_id'],how='left')\ntrain['diff_dat'] = (train.max_dat - train.t_dat).dt.days\ntrain = train.loc[train['diff_dat']<=6]\nprint('Train shape:',train.shape)","metadata":{"execution":{"iopub.status.busy":"2025-03-05T02:26:19.720079Z","iopub.execute_input":"2025-03-05T02:26:19.720777Z","iopub.status.idle":"2025-03-05T02:26:19.744327Z","shell.execute_reply.started":"2025-03-05T02:26:19.720739Z","shell.execute_reply":"2025-03-05T02:26:19.743714Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# (1) Recommend Most Often Previously Purchased Items\nNote that many operations in cuDF will shuffle the order of the dataframe rows. Therefore we need to sort afterward because we want the most often previously purchased items first. Because this will be the order of our predictons. Since we sort by `ct` and then `t_dat` will will recommend items that have been purchased more frequently first followed by items purchased more recently second.","metadata":{}},{"cell_type":"code","source":"tmp = train.groupby(['customer_id','article_id'])['t_dat'].agg('count').reset_index()\ntmp.columns = ['customer_id','article_id','ct']\ntrain = train.merge(tmp,on=['customer_id','article_id'],how='left')\ntrain = train.sort_values(['ct','t_dat'],ascending=False)\ntrain = train.drop_duplicates(['customer_id','article_id'])\ntrain = train.sort_values(['ct','t_dat'],ascending=False)\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2025-03-05T02:08:27.097570Z","iopub.execute_input":"2025-03-05T02:08:27.098329Z","iopub.status.idle":"2025-03-05T02:08:27.347883Z","shell.execute_reply.started":"2025-03-05T02:08:27.098291Z","shell.execute_reply":"2025-03-05T02:08:27.347191Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# (2) Recommend Items Purchased Together\nIn my notebook [here][1], we compute a dictionary of items frequently purchased together. We will load and use that dictionary below. Note that we use the command `drop_duplicates` so that we don't recommend an item that the user has already bought and we have already recommended above. We will need to use Pandas for some commands because RAPIDS cuDF doesn't have two conveinent commands, (1) create new column from dictionary map of another column (2) groupby aggregate strings sum.\n\nWe concatenate these rows after the rows containing customers' previous purchases. Therefore we will recommend previous items first and then items purchased together second. Note the trick to convert a column of int32 into a prediction string (using groupby agg str sum) is from notebook [here][2]\n\n[1]: https://www.kaggle.com/cdeotte/customers-who-bought-this-frequently-buy-this\n[2]: https://www.kaggle.com/hiroshisakiyama/recommending-items-recently-bought","metadata":{}},{"cell_type":"code","source":"# USE PANDAS TO MAP COLUMN WITH DICTIONARY\nimport pandas as pd, numpy as np\ntrain = train.to_pandas()\npairs = np.load('../input/hmitempairs/pairs_cudf.npy',allow_pickle=True).item()\ntrain['article_id2'] = train.article_id.map(pairs)","metadata":{"execution":{"iopub.status.busy":"2025-03-05T02:08:36.175911Z","iopub.execute_input":"2025-03-05T02:08:36.176172Z","iopub.status.idle":"2025-03-05T02:08:36.657425Z","shell.execute_reply.started":"2025-03-05T02:08:36.176141Z","shell.execute_reply":"2025-03-05T02:08:36.656866Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# RECOMMENDATION OF PAIRED ITEMS\ntrain2 = train[['customer_id','article_id2']].copy()\ntrain2 = train2.loc[train2.article_id2.notnull()]\ntrain2 = train2.drop_duplicates(['customer_id','article_id2'])\ntrain2 = train2.rename({'article_id2':'article_id'},axis=1)","metadata":{"execution":{"iopub.status.busy":"2025-03-05T02:08:39.091550Z","iopub.execute_input":"2025-03-05T02:08:39.092227Z","iopub.status.idle":"2025-03-05T02:08:40.078566Z","shell.execute_reply.started":"2025-03-05T02:08:39.092193Z","shell.execute_reply":"2025-03-05T02:08:40.077971Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# CONCATENATE PAIRED ITEM RECOMMENDATION AFTER PREVIOUS PURCHASED RECOMMENDATIONS\ntrain = train[['customer_id','article_id']]\ntrain = pd.concat([train,train2],axis=0,ignore_index=True)\ntrain.article_id = train.article_id.astype('int32')\ntrain = train.drop_duplicates(['customer_id','article_id'])","metadata":{"execution":{"iopub.status.busy":"2025-03-05T02:08:40.792221Z","iopub.execute_input":"2025-03-05T02:08:40.792921Z","iopub.status.idle":"2025-03-05T02:08:42.663562Z","shell.execute_reply.started":"2025-03-05T02:08:40.792888Z","shell.execute_reply":"2025-03-05T02:08:42.662947Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# CONVERT RECOMMENDATIONS INTO SINGLE STRING\ntrain.article_id = ' 0' + train.article_id.astype('str')\npreds = cudf.DataFrame( train.groupby('customer_id').article_id.sum().reset_index() )\npreds.columns = ['customer_id','prediction']\npreds.head()","metadata":{"execution":{"iopub.status.busy":"2025-03-05T02:08:44.915887Z","iopub.execute_input":"2025-03-05T02:08:44.916605Z","iopub.status.idle":"2025-03-05T02:08:53.162249Z","shell.execute_reply.started":"2025-03-05T02:08:44.916564Z","shell.execute_reply":"2025-03-05T02:08:53.161512Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# (3) Recommend Last Week's Most Popular Items\nAfter recommending previous purchases and items purchased together we will then recommend the 12 most popular items. Therefore if our previous recommendations did not fill up a customer's 12 recommendations, then it will be filled by popular items.","metadata":{}},{"cell_type":"code","source":"train = cudf.read_parquet('train.pqt')\ntrain.t_dat = cudf.to_datetime(train.t_dat)\ntrain = train.loc[train.t_dat >= cudf.to_datetime('2020-09-16')]\ntop12 = ' 0' + ' 0'.join(train.article_id.value_counts().to_pandas().index.astype('str')[:12])\nprint(\"Last week's top 12 popular items:\")\nprint( top12 )","metadata":{"execution":{"iopub.status.busy":"2025-03-05T02:10:24.916199Z","iopub.execute_input":"2025-03-05T02:10:24.916582Z","iopub.status.idle":"2025-03-05T02:10:25.314704Z","shell.execute_reply.started":"2025-03-05T02:10:24.916548Z","shell.execute_reply":"2025-03-05T02:10:25.313971Z"},"trusted":true},"outputs":[],"execution_count":null}]}