{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### Contents\n\n\n#### 1. What is Customer Segmentation?\n\n > Which customer is more likely similar?\n\n\n\n\n#### 2. How to represent customers?\n\n> The way to represent customer as features\n\n\n\n\n#### 3. How to make a group with similar customers?\n\n> The way to seperate customer groups\n\n","metadata":{}},{"cell_type":"markdown","source":"### [Note]\nThis code is a notebook designed to study customer segmentation tasks only. If you are not familiar with the data of the competition yet, it would be good to look at other EDA Notebooks first.\n\n- Recommendation 1 : https://www.kaggle.com/code/vanguarde/h-m-eda-first-look\n\n- Recommendation 2 : https://www.kaggle.com/code/andradaolteanu/h-m-eda-rapids-and-similarity-recommenders\n\n- Recommendation 3 : https://www.kaggle.com/code/ludovicocuoghi/h-m-sales-and-customers-deep-analysis","metadata":{}},{"cell_type":"markdown","source":"### 1. What is Customer Segmentation?\n\n- Customer segmentation is the practice of dividing a customer base into groups of individuals that are similar in specific ways relevant to marketing, such as age, gender, interests and spending habits.\n\n> Source : https://www.techtarget.com/searchcustomerexperience/definition/customer-segmentation\n\n<br>\n  In order for sales to occur, items must be recommended based on the interests of customers. To do this, it is important to decide what kind of customer information to look at, and to group similar customers together. We will segment customers and find out the characteristics of each customer through machine learning methods.","metadata":{}},{"cell_type":"markdown","source":"#### Import Libraries","metadata":{}},{"cell_type":"code","source":"# data analysis\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# progress bar\nfrom tqdm.notebook import tqdm\n\n# dimensionality reduction (for visualization)\n#from sklearn.decomposition import PCA\nfrom sklearn.manifold import TSNE\n\n# clustering\nfrom sklearn.cluster import KMeans\nfrom sklearn.metrics import silhouette_score","metadata":{"execution":{"iopub.status.busy":"2022-07-24T02:37:11.257680Z","iopub.execute_input":"2022-07-24T02:37:11.258066Z","iopub.status.idle":"2022-07-24T02:37:11.608424Z","shell.execute_reply.started":"2022-07-24T02:37:11.257993Z","shell.execute_reply":"2022-07-24T02:37:11.607505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Data Preparation","metadata":{}},{"cell_type":"code","source":"articles = pd.read_csv('../input/h-and-m-personalized-fashion-recommendations/articles.csv')\ncustomers = pd.read_csv('../input/h-and-m-personalized-fashion-recommendations/customers.csv')\ntransactions = pd.read_csv('../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv')\nprint(articles.shape, customers.shape, transactions.shape)","metadata":{"execution":{"iopub.status.busy":"2022-07-24T02:37:11.610484Z","iopub.execute_input":"2022-07-24T02:37:11.610864Z","iopub.status.idle":"2022-07-24T02:38:43.449999Z","shell.execute_reply.started":"2022-07-24T02:37:11.610808Z","shell.execute_reply":"2022-07-24T02:38:43.448544Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will use only a small portion of the data for the purposes of our analysis. For actual analysis, it is better to use a lot of resources and perform it with the entire data.","metadata":{}},{"cell_type":"code","source":"# choose portion\ntransactions = transactions[:100000]","metadata":{"execution":{"iopub.status.busy":"2022-07-24T02:38:43.451827Z","iopub.execute_input":"2022-07-24T02:38:43.453293Z","iopub.status.idle":"2022-07-24T02:38:43.459023Z","shell.execute_reply.started":"2022-07-24T02:38:43.453223Z","shell.execute_reply":"2022-07-24T02:38:43.457841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To define customer representation, it is necessary to bring related data into one table. Based on some selected transactions, the data is combined into one table.","metadata":{}},{"cell_type":"code","source":"# we just need \"customer_id\", \"article_id\", \"price\" columns.\ntransactions = transactions[[\"customer_id\", \"article_id\", \"price\"]]\ntransactions","metadata":{"execution":{"iopub.status.busy":"2022-07-24T02:38:43.462237Z","iopub.execute_input":"2022-07-24T02:38:43.463496Z","iopub.status.idle":"2022-07-24T02:38:44.162029Z","shell.execute_reply.started":"2022-07-24T02:38:43.463410Z","shell.execute_reply":"2022-07-24T02:38:44.160899Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# we only need \"customer_id\", \"age\"\ncustomers = customers[[\"customer_id\", \"age\"]]\ncustomers","metadata":{"execution":{"iopub.status.busy":"2022-07-24T02:38:44.163530Z","iopub.execute_input":"2022-07-24T02:38:44.163820Z","iopub.status.idle":"2022-07-24T02:38:44.290348Z","shell.execute_reply.started":"2022-07-24T02:38:44.163790Z","shell.execute_reply":"2022-07-24T02:38:44.288787Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# We will use category information from articles, other columns don't need.\narticles = articles[[\"article_id\", \"prod_name\", \"product_type_name\", \"product_group_name\",\n                     \"department_name\", \"index_name\", \"index_group_name\",\n                     \"section_name\", \"garment_group_name\"]]\narticles","metadata":{"execution":{"iopub.status.busy":"2022-07-24T02:38:44.292103Z","iopub.execute_input":"2022-07-24T02:38:44.292452Z","iopub.status.idle":"2022-07-24T02:38:44.342559Z","shell.execute_reply.started":"2022-07-24T02:38:44.292423Z","shell.execute_reply":"2022-07-24T02:38:44.341871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# merge into just one table!\ntemp = pd.merge(transactions, articles, on=\"article_id\", how='inner')\ntemp = pd.merge(temp, customers, on=\"customer_id\", how='inner')\ntemp","metadata":{"execution":{"iopub.status.busy":"2022-07-24T02:38:44.343947Z","iopub.execute_input":"2022-07-24T02:38:44.344237Z","iopub.status.idle":"2022-07-24T02:40:22.714869Z","shell.execute_reply.started":"2022-07-24T02:38:44.344209Z","shell.execute_reply":"2022-07-24T02:40:22.713608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check null values and drop it.\ndisplay(temp[temp.isnull().any(axis=1)])\ntemp = temp.dropna()","metadata":{"execution":{"iopub.status.busy":"2022-07-24T02:40:22.716313Z","iopub.execute_input":"2022-07-24T02:40:22.716651Z","iopub.status.idle":"2022-07-24T02:41:34.112900Z","shell.execute_reply.started":"2022-07-24T02:40:22.716621Z","shell.execute_reply":"2022-07-24T02:41:34.111881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# grab some garbage\nimport gc\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-07-24T02:41:34.114293Z","iopub.execute_input":"2022-07-24T02:41:34.114619Z","iopub.status.idle":"2022-07-24T02:41:34.437980Z","shell.execute_reply.started":"2022-07-24T02:41:34.114589Z","shell.execute_reply":"2022-07-24T02:41:34.436911Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 2. How to represent customers?\n\nThere are many different ways to define the characteristics of a customer.\nWe will use the customer age and the article category info the customer has purchased.\nEach customer will be presented with a numerical value based on which item they purchased.\nThe data represented is defined as a feature vector, after that clustering is performed.\n\n\nThere are many categories we can choose from articles table, we'll start with the highest level first.","metadata":{}},{"cell_type":"code","source":"# find highest level of groups\nfor col in [\"prod_name\", \"product_type_name\", \"product_group_name\",\n            \"department_name\", \"index_name\", \"index_group_name\",\n            \"section_name\", \"garment_group_name\"]:\n    print(f\"{col}\\t>> {temp[col].nunique()} number of unique categories.\")","metadata":{"execution":{"iopub.status.busy":"2022-07-24T02:41:34.442147Z","iopub.execute_input":"2022-07-24T02:41:34.442527Z","iopub.status.idle":"2022-07-24T02:42:05.895668Z","shell.execute_reply.started":"2022-07-24T02:41:34.442496Z","shell.execute_reply":"2022-07-24T02:42:05.894886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# define customer-article matrix\nca_matrix = pd.crosstab(index=temp.customer_id, columns=temp.index_group_name)\nca_matrix","metadata":{"execution":{"iopub.status.busy":"2022-07-24T02:42:05.899577Z","iopub.execute_input":"2022-07-24T02:42:05.899886Z","iopub.status.idle":"2022-07-24T02:42:56.808997Z","shell.execute_reply.started":"2022-07-24T02:42:05.899856Z","shell.execute_reply":"2022-07-24T02:42:56.807910Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Get average prices they've been bought.\nprices = temp.groupby([\"customer_id\"])[\"price\"].mean()\nprices","metadata":{"execution":{"iopub.status.busy":"2022-07-24T02:42:56.810341Z","iopub.execute_input":"2022-07-24T02:42:56.810701Z","iopub.status.idle":"2022-07-24T02:43:16.082916Z","shell.execute_reply.started":"2022-07-24T02:42:56.810671Z","shell.execute_reply":"2022-07-24T02:43:16.081874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# get customer ages\nages = temp.groupby([\"customer_id\"])[\"age\"].mean()\nages","metadata":{"execution":{"iopub.status.busy":"2022-07-24T02:43:16.084170Z","iopub.execute_input":"2022-07-24T02:43:16.084482Z","iopub.status.idle":"2022-07-24T02:43:35.986079Z","shell.execute_reply.started":"2022-07-24T02:43:16.084426Z","shell.execute_reply":"2022-07-24T02:43:35.984359Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# merge them into one table\ntrain = pd.merge(ca_matrix, prices, on=\"customer_id\", how=\"left\")\ntrain = pd.merge(train, ages, on=\"customer_id\", how=\"left\")\ntrain","metadata":{"execution":{"iopub.status.busy":"2022-07-24T02:43:35.988433Z","iopub.execute_input":"2022-07-24T02:43:35.988757Z","iopub.status.idle":"2022-07-24T02:43:39.898738Z","shell.execute_reply.started":"2022-07-24T02:43:35.988728Z","shell.execute_reply":"2022-07-24T02:43:39.896529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"When using distance-based machine learning models, feature scaling must be performed.","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import MinMaxScaler\n\nscaled = MinMaxScaler().fit_transform(train)\nX = pd.DataFrame(data=scaled, columns=train.columns, index=train.index)\nX","metadata":{"execution":{"iopub.status.busy":"2022-07-24T02:43:39.900294Z","iopub.execute_input":"2022-07-24T02:43:39.900636Z","iopub.status.idle":"2022-07-24T02:43:40.004973Z","shell.execute_reply.started":"2022-07-24T02:43:39.900606Z","shell.execute_reply":"2022-07-24T02:43:40.003529Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### 3. How to make a group with similar customers?\n\nUsually, we use insights or statistics from existing research to segment customers.\nWe use machine learning methods to practice how to construct a customer segment.\n\n\nWe use \"clustering\", an unsupervised settings, because we don't know which customer belongs to which customer group.","metadata":{}},{"cell_type":"code","source":"# we segment 28,317 users into 8 groups by using K-means clustering.\n\nmodel = KMeans()\npreds = model.fit_predict(X)  # no needed 'y'\nprint(\"SSE : \", model.inertia_)\nprint(\"Silhouette Score : %.4f\" % silhouette_score(X, preds))","metadata":{"execution":{"iopub.status.busy":"2022-07-24T02:43:40.007046Z","iopub.execute_input":"2022-07-24T02:43:40.007758Z","iopub.status.idle":"2022-07-24T07:23:20.283813Z","shell.execute_reply.started":"2022-07-24T02:43:40.007729Z","shell.execute_reply":"2022-07-24T07:23:20.282530Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To find optimal number of clusters, we can use elbow methods.\n\nOR\n\nwe can use silhouette scores as metric.","metadata":{}},{"cell_type":"code","source":"sses = []\nscores = []\nN = list(range(2, 11))  # check between K=2 to K=10\n\nfor n_clusters in tqdm(N):\n    model = KMeans(n_clusters=n_clusters, random_state=42)\n    preds = model.fit_predict(X)\n    sse = model.inertia_\n    score = silhouette_score(X, preds)\n    sses.append(sse)\n    scores.append(score)","metadata":{"execution":{"iopub.status.busy":"2022-07-24T07:23:20.285575Z","iopub.execute_input":"2022-07-24T07:23:20.285927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Plot the result","metadata":{}},{"cell_type":"code","source":"# SSE\nplt.figure(figsize=(12, 6))\nplt.title(\"SSE in K-means\")\nsns.lineplot(x=N, y=sses)\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"By looking above graph, we derive best **K=4** with elbow methods.","metadata":{}},{"cell_type":"code","source":"# Silhouette\nplt.figure(figsize=(12, 6))\nplt.title(\"Silhouette Scores using K-means\")\nsns.lineplot(x=N, y=scores)\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"By looking above graph, we derive best **K=2** using silhouette scores.","metadata":{}},{"cell_type":"markdown","source":"#### Select best K=4, see the final result!","metadata":{}},{"cell_type":"code","source":"# visualize the result\n# if you want to use PCA, then uncomment PCA related codes.\n# if you want to use full data, try MiniBatchSparsePCA.\nfrom sklearn.decomposition import MiniBatchSparsePCA\n\n#pca = PCA(n_components=2)\n#tsne = TSNE(n_components=2, perplexity=50, random_state=42)\n#svd = TruncatedSVD(n_components=2, )\npca = MiniBatchSparsePCA(n_components=2, batch_size=100000, random_state=42)\n\nX_reduced = pca.fit_transform(X)\n#X_reduced = tsne.fit_transform(X)\npca_df = pd.DataFrame(data=X_reduced,\n                     columns=[f\"PC{i}\" for i in range(X_reduced.shape[1])])\n\n# tsne_df = pd.DataFrame(data=X_reduced,\n#                        columns=[f\"dim{i}\" for i in range(1, X_reduced.shape[1]+1)])\n\n\n# finalize cluster model\n# with sse\nresult_sse = KMeans(n_clusters=4).fit_predict(X)\n\n# with silhouette\n#result_silhouette = KMeans(n_clusters=2).fit_transform(X)\n\npca_df['group'] = result_sse\n#tsne_df['group'] = result_sse\n\nplt.figure(figsize=(8, 6))\nsns.scatterplot(data=pca_df, x=\"PC0\", y=\"PC1\", hue=\"group\")\n#sns.scatterplot(data=tsne_df, x=\"dim1\", y=\"dim2\", hue=\"group\")\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}