{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":31254,"databundleVersionId":3103714,"sourceType":"competition"}],"dockerImageVersionId":30887,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import os\n\n# Define required files\nrequired_files = {\n    \"sample_submission.csv\": None,\n    \"articles.csv\": None,\n    \"transactions_train.csv\": None,\n    \"customers.csv\": None\n}\n\n# Search for files\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        if filename in required_files:  # Check only required files\n            required_files[filename] = os.path.join(dirname, filename)\n        if all(required_files.values()):  # Stop early if all found\n            break\n\n# Print found files\nfor name, path in required_files.items():\n    if path:\n        print(path)\n    else:\n        print(f\"Missing: {name}\")\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-02-14T09:12:30.410612Z","iopub.execute_input":"2025-02-14T09:12:30.410896Z","iopub.status.idle":"2025-02-14T09:16:13.425799Z","shell.execute_reply.started":"2025-02-14T09:12:30.410852Z","shell.execute_reply":"2025-02-14T09:16:13.424978Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Standard library imports\nimport sys\nimport warnings\nimport time\nimport os\nimport copy\nimport gc\nimport re\nimport random\nimport pickle\nfrom collections import Counter\nfrom datetime import datetime, timedelta\nfrom pathlib import Path\nfrom pprint import pprint\n\n# Third-party imports\nimport numpy as np\nimport cupy as cp\nimport pandas as pd\nimport cudf\nfrom cuml.cluster import KMeans\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom tqdm import tqdm\nfrom sklearn import preprocessing\n\n# Suppress warnings for cleaner output\nwarnings.filterwarnings('ignore')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-14T09:16:13.427043Z","iopub.execute_input":"2025-02-14T09:16:13.427313Z","iopub.status.idle":"2025-02-14T09:16:31.588928Z","shell.execute_reply.started":"2025-02-14T09:16:13.427291Z","shell.execute_reply":"2025-02-14T09:16:31.588301Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# CustomerSegmentation Class\n## This class handles customer data preprocessing and clustering.\n\n### Preprocessing Customer Data\n* Encodes categorical variables like fashion_news_frequency and club_member_status with numerical values.\n* Fills missing values with appropriate defaults.\n* Drops irrelevant columns like postal_code.\n### Creating Customer Segments\n* Selects customer features like age, membership status, and activity.\n* Applies feature scaling using StandardScaler (or MinMaxScaler if specified).\n* Clusters customers into groups using KMeans (with n_clusters=12).\n* Group customers with similar characteristics into distinct clusters.\n\nThoughts: To group customers with similar behaviors, I decided to use clustering. I chose K-Means because it's effective for segmentation when the number of clusters is predefined. Before clustering, I needed to preprocess the customer data. This involved handling missing values, encoding categorical variables like 'club_member_status' and 'fashion_news_frequency' into numerical values, and normalizing features so that no single feature dominates the clustering.\n\nImplementation: The CustomerSegmentation class handles preprocessing using mappings for categorical variables and applies standardization to normalize features. K-Means from cuML (a GPU-accelerated library) was used for faster processing, especially important with large datasets.","metadata":{}},{"cell_type":"code","source":"class CustomerSegmentation:\n    def __init__(self, random_state=2025, n_clusters=12):\n        self.random_state = random_state\n        self.n_clusters = n_clusters\n        \n        self.news_frequency_mapping = {\n            np.nan: 0, \n            'None': 0, \n            'NONE': 0,\n            'Monthly': 1, \n            'Regularly': 2\n        }\n        \n        self.member_status_mapping = {\n            np.nan: 0,\n            'PRE-CREATE': 1,\n            'ACTIVE': 2,\n            'LEFT CLUB': -1\n        }\n\n    def preprocess_customer_data(self, customer_df, columns_to_drop=['postal_code']):\n        \"\"\"Preprocess customer data by handling missing values and encoding categorical variables.\"\"\"\n        # Convert to cuDF if input is pandas DataFrame\n        if isinstance(customer_df, pd.DataFrame):\n            processed_df = cudf.from_pandas(customer_df)\n        else:\n            processed_df = customer_df.copy()\n            \n        # Drop specified columns\n        if any(col in processed_df.columns for col in columns_to_drop):\n            processed_df = processed_df.drop(columns=[col for col in columns_to_drop if col in processed_df.columns])\n        \n        # Handle categorical and missing values\n        if 'fashion_news_frequency' in processed_df.columns:\n            processed_df['fashion_news_frequency'] = (\n                processed_df['fashion_news_frequency']\n                .astype('str')  # Convert to string to handle NaN values\n                .replace('NONE', 'None')\n                .map(self.news_frequency_mapping)\n                .fillna(0)\n            )\n            \n        if 'club_member_status' in processed_df.columns:\n            processed_df['club_member_status'] = (\n                processed_df['club_member_status']\n                .astype('str')\n                .map(self.member_status_mapping)\n                .fillna(0)\n            )\n            \n        if 'age' in processed_df.columns:\n            processed_df['age'] = processed_df['age'].fillna(-1)\n            \n        # Handle binary indicators\n        for col in ['FN', 'Active']:\n            if col in processed_df.columns:\n                processed_df[col] = processed_df[col].fillna(0)\n        \n        return processed_df\n\n    def create_customer_segments(self, df, id_columns, feature_columns, normalization='StandardScaler', return_normalized=False):\n        \"\"\"Create customer segments using KMeans clustering.\"\"\"\n        # Ensure all feature columns exist\n        missing_cols = [col for col in feature_columns if col not in df.columns]\n        if missing_cols:\n            raise ValueError(f\"Missing feature columns: {missing_cols}\")\n\n        # Convert features to cupy array for sklearn\n        X = df[feature_columns].to_pandas().values\n        \n        # Apply normalization if specified\n        if normalization == 'StandardScaler':\n            normalizer = preprocessing.StandardScaler()\n            X = normalizer.fit_transform(X)\n        elif normalization == 'minMax':\n            normalizer = preprocessing.MinMaxScaler()\n            X = normalizer.fit_transform(X)\n        \n        print(f'Normalization Method: {normalization}')\n        \n        # Perform clustering\n        kmeans = KMeans(n_clusters=self.n_clusters, random_state=self.random_state)\n        kmeans.fit(X)\n        \n        print(f'Clustering Distortion: {kmeans.inertia_:.2f}')\n        \n\n        # Add predictions to original dataframe\n        cluster_assignments = cudf.Series(kmeans.labels_, name='cluster_id')\n        result_df = df.copy()\n        result_df['cluster_id'] = cluster_assignments\n        \n        if return_normalized:\n            norm_features_df = cudf.DataFrame(X, columns=feature_columns)\n            print(\"\\n=== Normalized Features Summary ===\")\n            print(norm_features_df.describe())\n            return result_df, norm_features_df\n            \n        return result_df\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-14T09:16:31.590065Z","iopub.execute_input":"2025-02-14T09:16:31.590476Z","iopub.status.idle":"2025-02-14T09:16:31.600381Z","shell.execute_reply.started":"2025-02-14T09:16:31.590454Z","shell.execute_reply":"2025-02-14T09:16:31.599567Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# PurchaseAnalyzer Class\n## This class analyzes transaction patterns and generates product recommendations for analyzing purchase.\n\n### Getting Recent Transactions\n* Filters transactions from the last few weeks (September 1–21, 2020).\n* Uses GPU-accelerated cudf for faster processing.\n### Analyzing Cluster Preferences\n* Merges transactions with customer segments.\n* Counts how often each product is purchased in each cluster.\n* Identifies the top N articles for each cluster to understand cluster-level preferences.\n### Generating Recommendations\n* Assigns weights to each purchase based on recency (more recent = higher weight).\n* Calculates a weighted popularity score for each product per customer.\n* Ranks products and selects the top 12 recommendations.\n\nThoughts: To generate relevant recommendations, focusing on recent transactions makes sense as fashion trends change. I filtered transactions from the last three weeks (2020-09-01 to 2020-09-21) to capture the latest purchasing patterns. Then, I calculated each cluster's product preferences by counting article purchases within each cluster.\n\nImplementation: The PurchaseAnalyzer class loads and filters transactions. It merges transaction data with cluster assignments to compute purchase counts per cluster. The analyze_cluster_preferences method identifies top products for each cluster and computes a similarity matrix to understand how clusters overlap in preferences","metadata":{}},{"cell_type":"code","source":"class PurchaseAnalyzer:\n    def __init__(self, input_path, date_range=('2020-09-01', '2020-09-21')):\n        self.input_path = Path(input_path)\n        self.date_range = date_range\n        \n    def get_recent_transactions(self):\n        \"\"\"Load and filter recent transactions using cuDF.\"\"\"\n        transactions_path = self.input_path / 'transactions_train.csv'\n        if not transactions_path.exists():\n            raise FileNotFoundError(f\"Transactions file not found at {transactions_path}\")\n\n        transactions_df = cudf.read_csv(\n            transactions_path,\n            usecols=['t_dat', 'customer_id', 'article_id'],\n            dtype={\n                'article_id': 'int32',\n                't_dat': 'str',\n                'customer_id': 'str'\n            }\n        )\n        \n        # Convert date and filter range\n        transactions_df['t_dat'] = cudf.to_datetime(transactions_df['t_dat'])\n        mask = (transactions_df['t_dat'] >= self.date_range[0]) & (transactions_df['t_dat'] <= self.date_range[1])\n        return transactions_df[mask]\n    \n    def analyze_cluster_preferences(self, recent_transactions, customer_segments, top_n=100):\n        \"\"\"Analyze product preferences for each customer cluster.\"\"\"\n        # Ensure required columns exist\n        required_cols = ['customer_id', 'cluster_id']\n        if not all(col in customer_segments.columns for col in required_cols):\n            raise ValueError(f\"Missing required columns in customer_segments: {required_cols}\")\n\n        # Merge transactions with cluster assignments\n        merged_df = recent_transactions.merge(\n            customer_segments[['customer_id', 'cluster_id']], \n            on='customer_id', \n            how='inner'\n        )\n        \n        # Calculate purchase counts by cluster and article\n        purchase_counts = (\n            merged_df.groupby(['cluster_id', 'article_id'])\n            .size()\n            .reset_index()\n            .rename(columns={0: 'purchase_count'})\n        )\n        \n        purchase_counts = purchase_counts.to_pandas()\n        \n        # Get top products for each cluster\n        cluster_preferences = {}\n        for cluster_id in purchase_counts['cluster_id'].unique():\n            cluster_purchases = purchase_counts[purchase_counts['cluster_id'] == cluster_id]\n            top_products = cluster_purchases.nlargest(top_n, 'purchase_count')['article_id'].tolist()\n            cluster_preferences[cluster_id] = top_products\n            \n        # Calculate similarity matrix\n        similarity_df = pd.DataFrame([cluster_preferences]).T.rename(columns={0: 'top_products'})\n        for cluster_id in similarity_df.index:\n            similarity_df[cluster_id] = [\n                len(set(similarity_df.at[cluster_id, 'top_products']) & \n                    set(similarity_df.at[x, 'top_products'])) / top_n \n                for x in similarity_df.index\n            ]\n            \n        return similarity_df.drop(columns='top_products')\n\n    def generate_recommendations(self, customer_segments, recent_transactions, n_recommendations=12):\n        \"\"\"Generate product recommendations for each customer segment.\"\"\"\n        if len(recent_transactions) == 0:\n            raise ValueError(\"No transactions found in the specified date range\")\n\n        last_date = recent_transactions['t_dat'].max()\n        \n        # Calculate days since last purchase for weighting\n        recent_transactions = recent_transactions.copy()\n        recent_transactions['days_since'] = (\n            last_date - recent_transactions['t_dat']\n        ).dt.days\n        \n        # Convert days_since to cupy array for GPU calculations\n        days_since_array = cp.asarray(recent_transactions['days_since'].values.astype('float32'))\n        \n        # Weight calculation parameters\n        a, b, c, d = 2.5e4, 1.5e5, 2e-1, 1e3\n        \n        # Calculate weights using cupy operations\n        weights = (\n            a / cp.sqrt(days_since_array) + \n            b * cp.exp(-c * days_since_array) - \n            d\n        )\n        \n        # Convert weights back to cudf Series and clip negative values\n        recent_transactions['weight'] = cudf.Series(\n            cp.asnumpy(weights), \n            index=recent_transactions.index\n        ).clip(lower=0)\n        \n        # Calculate weighted purchase counts\n        weighted_counts = (\n            recent_transactions\n            .groupby(['customer_id', 'article_id'])\n            ['weight']\n            .sum()\n            .reset_index()\n        )\n        \n        # Rank products for each customer\n        weighted_counts['rank'] = weighted_counts.groupby('customer_id')['weight'].rank(\n            method='dense',\n            ascending=False\n        )\n        \n        # Filter top N recommendations\n        recommendations = weighted_counts[weighted_counts['rank'] <= n_recommendations]\n        \n        # Clean up GPU memory\n        del days_since_array, weights\n        cp.get_default_memory_pool().free_all_blocks()\n        \n        return recommendations\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-14T09:16:31.601700Z","iopub.execute_input":"2025-02-14T09:16:31.601995Z","iopub.status.idle":"2025-02-14T09:16:31.626304Z","shell.execute_reply.started":"2025-02-14T09:16:31.601965Z","shell.execute_reply":"2025-02-14T09:16:31.625525Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# RecommendationEngine Class\n## This class retrieves recommended articles for a given customer.\n\n### Generating User Recommendations\n* Identifies the cluster the customer belongs to.\n* Finds the most popular articles within that cluster that the customer hasn't purchased yet.\n* GPU-accelerated group operations for efficiency\n\nThought: Recommendations should balance between cluster preferences and individual purchase history. For each user, I combine their cluster's popular items with their own recent purchases, excluding items they've already bought. Including article details (like product type) makes recommendations more interpretable.\n\nImplementation: The RecommendationEngine loads article data and, for a given user, retrieves their cluster's top products. It merges this with article details and excludes already purchased items. The generate_recommendations method weights recent purchases more heavily using an exponential decay to prioritize newer items.","metadata":{}},{"cell_type":"code","source":"class RecommendationEngine:\n    def __init__(self, input_path):\n        self.input_path = Path(input_path)\n        self.articles_df = None\n        self.load_article_data()\n        \n    def load_article_data(self):\n        \"\"\"Load and preprocess article data.\"\"\"\n        articles_path = self.input_path / 'articles.csv'\n        if not articles_path.exists():\n            raise FileNotFoundError(f\"Articles file not found at {articles_path}\")\n            \n        self.articles_df = pd.read_csv(articles_path)\n        \n    def get_user_recommendations(self, customer_id, segmented_customers, recent_transactions, \n                               n_recommendations=12, include_article_details=True):\n        \"\"\"\n        Generate personalized recommendations for a specific user.\n        \n        Args:\n            customer_id (str): The customer ID to generate recommendations for\n            segmented_customers (cudf.DataFrame): DataFrame with customer segments\n            recent_transactions (cudf.DataFrame): Recent transaction data\n            n_recommendations (int): Number of recommendations to generate\n            include_article_details (bool): Whether to include article details in results\n        \"\"\"\n        # Get user's cluster\n        user_cluster = segmented_customers[\n            segmented_customers['customer_id'] == customer_id\n        ]['cluster_id'].iloc[0]\n        \n        # Get cluster-based recommendations\n        cluster_transactions = recent_transactions.merge(\n            segmented_customers[segmented_customers['cluster_id'] == user_cluster][['customer_id']],\n            on='customer_id',\n            how='inner'\n        )\n        \n        # Get user's recent purchases\n        user_purchases = set(\n            recent_transactions[\n                recent_transactions['customer_id'] == customer_id\n            ]['article_id'].to_pandas().tolist()\n        )\n        \n        # Calculate article popularity within cluster\n        article_scores = (\n            cluster_transactions\n            .groupby('article_id')\n            .size()\n            .reset_index(name='popularity_score')\n        )\n        \n        # Convert to pandas for easier processing\n        article_scores = article_scores.to_pandas()\n        \n        # Remove articles user has already purchased\n        article_scores = article_scores[\n            ~article_scores['article_id'].isin(user_purchases)\n        ]\n        \n        # Sort by popularity and get top recommendations\n        recommendations = article_scores.nlargest(n_recommendations, 'popularity_score')\n        \n        if include_article_details and self.articles_df is not None:\n            recommendations = recommendations.merge(\n                self.articles_df,\n                on='article_id',\n                how='left'\n            )\n            \n        return recommendations","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-14T09:32:25.743430Z","iopub.execute_input":"2025-02-14T09:32:25.743800Z","iopub.status.idle":"2025-02-14T09:32:25.751029Z","shell.execute_reply.started":"2025-02-14T09:32:25.743775Z","shell.execute_reply":"2025-02-14T09:32:25.750339Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Main Execution Flow\n## This section runs the entire pipeline.\n\n### Step-by-Step Process\n* Preprocess Customer Data: Clean and encode customer info.\n* Create Segments: Apply clustering to identify customer groups.\n* Analyze Transactions: Identify cluster-level purchase patterns.\n* Generate Recommendations: Get personalized article suggestions for each customer.\n* Visualize Similarity: Show cluster similarity as a heatmap.\n* Example User: Generate and print top recommendations for a sample customer.\n\nThought: To validate the clustering, visualizing cluster similarity helps check if clusters are distinct. A heatmap of the similarity matrix shows how much clusters share top products. Testing recommendations on a sample user ensures the system works end-to-end.\n\nImplementation: Using seaborn, a heatmap visualizes cluster similarities. The main function runs the pipeline and prints sample recommendations, including article details for clarity.","metadata":{}},{"cell_type":"code","source":"def main():\n    INPUT_PATH = Path('../input/h-and-m-personalized-fashion-recommendations')\n    \n    # Initialize classes\n    segmentation = CustomerSegmentation(random_state=2025, n_clusters=12)\n    analyzer = PurchaseAnalyzer(input_path=INPUT_PATH)\n    recommendation_engine = RecommendationEngine(input_path=INPUT_PATH)  # Add recommendation engine\n    \n    try:\n        # Load and preprocess customer data\n        customers_path = INPUT_PATH / 'customers.csv'\n        if not customers_path.exists():\n            raise FileNotFoundError(f\"Customers file not found at {customers_path}\")\n            \n        customers_df = cudf.read_csv(customers_path)\n        processed_customers = segmentation.preprocess_customer_data(customers_df)\n        \n        # Create customer segments\n        feature_cols = ['club_member_status', 'fashion_news_frequency', 'age', 'FN', 'Active']\n        segmented_customers = segmentation.create_customer_segments(\n            processed_customers,\n            id_columns=['customer_id'],\n            feature_columns=feature_cols\n        )\n        \n        # Analyze recent transactions\n        recent_transactions = analyzer.get_recent_transactions()\n        \n        # Generate recommendations\n        recommendations = analyzer.generate_recommendations(\n            segmented_customers,\n            recent_transactions\n        )\n        \n        # Analyze cluster similarities\n        cluster_similarity = analyzer.analyze_cluster_preferences(\n            recent_transactions,\n            segmented_customers\n        )\n        \n        # Visualize results\n        plt.figure(figsize=(10, 6))\n        sns.heatmap(cluster_similarity, annot=True, cbar=False)\n        plt.title('Cluster Purchase Pattern Similarity')\n        plt.show()\n\n         # Generate recommendations for a sample user\n        sample_user_id = segmented_customers['customer_id'].iloc[0]\n        print(f\"\\nGenerating recommendations for user: {sample_user_id}\")\n        \n        user_recommendations = recommendation_engine.get_user_recommendations(\n            sample_user_id,\n            segmented_customers,\n            recent_transactions\n        )\n        \n        print(\"\\nTop Recommended Articles:\")\n        print(user_recommendations[['article_id', 'popularity_score', 'product_type_name']].head())\n        \n    except Exception as e:\n        print(f\"An error occurred: {str(e)}\")\n        raise\n\nif __name__ == \"__main__\":\n    main()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-02-14T09:32:29.338656Z","iopub.execute_input":"2025-02-14T09:32:29.339006Z","iopub.status.idle":"2025-02-14T09:32:33.343112Z","shell.execute_reply.started":"2025-02-14T09:32:29.338977Z","shell.execute_reply":"2025-02-14T09:32:33.342123Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Optimization Considerations**:\n\n- **GPU Acceleration**: Using cuDF and cuML leverages GPU power for faster data processing and model training, crucial for handling large datasets efficiently.\n\n- **Memory Management**: Explicitly freeing GPU memory after operations prevents memory leaks, which is vital in resource-constrained environments like Kaggle.\n\n**Challenges and Solutions**:\n\n- **Data Preprocessing**: Handling missing values and converting categorical data required careful mapping to avoid introducing bias. For example, mapping 'club_member_status' to numerical values reflecting activity levels.\n\n- **Feature Selection**: Choosing the right features (age, news frequency, member status) for clustering was based on domain knowledge, assuming these affect purchasing behavior.\n\n- **Weighting Recent Purchases**: The weight formula (combining inverse square root and exponential decay) was chosen to emphasize recent purchases without completely ignoring older ones.\n\n**Potential Improvements**:\n\n- **Dynamic Time Range**: Instead of a fixed date range, dynamically determine the period based on data recency according to need.\n\n- **Incorporating Article Features**: Using article attributes (e.g., color, season) could enhance recommendations by matching customer preferences more granularly.","metadata":{}}]}