{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Our Goal ✨\n1. Similarity Recommendation Based on Node2Vec\n2. Let's Build a Recommendation System Using H&M Dataset!","metadata":{}},{"cell_type":"markdown","source":"Hello, Kagglers!\n\nIn this notebook, we'll try fashion recommendations using the Node2Vec algorithm, which is implemented by applying the simple concept of Graph Neural Network. Let's compare the images of clothes that the user purchased and the recommended clothes to see if the similarity score is reliable!\n\nYou can also find the .ipynb file that can be run all at once on this [GitHub link](https://github.com/H4Y3J1N/Graph-Travel/tree/H4Y3J1N).","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport seaborn as sns\nfrom matplotlib import pyplot as plt\nimport networkx as nx\nfrom gensim.models import Word2Vec\nfrom sklearn.metrics.pairwise import cosine_similarity\nimport matplotlib.image as mpimg\nimport random","metadata":{"execution":{"iopub.status.busy":"2023-07-01T05:56:53.980340Z","iopub.execute_input":"2023-07-01T05:56:53.981015Z","iopub.status.idle":"2023-07-01T05:56:54.000027Z","shell.execute_reply.started":"2023-07-01T05:56:53.980867Z","shell.execute_reply":"2023-07-01T05:56:53.998888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/articles.csv\")\n# customers = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/customers.csv\")\ntransactions = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\")","metadata":{"execution":{"iopub.status.busy":"2023-07-01T05:56:54.002706Z","iopub.execute_input":"2023-07-01T05:56:54.003306Z","iopub.status.idle":"2023-07-01T05:56:57.923385Z","shell.execute_reply.started":"2023-07-01T05:56:54.003251Z","shell.execute_reply":"2023-07-01T05:56:57.921242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In order to depict this dataset as a Graph, we need to retain only the meaningful rows. The goal is to reduce sparsity as much as possible for better performance.\n\nSo, the first thing to do is to calculate the frequency of 'article_id' (product ID) and 'customer_id' (user ID), and consider filtering only users who have many purchase records & products with many purchase records.","metadata":{}},{"cell_type":"code","source":"item_freq = transactions.groupby('article_id')['customer_id'].nunique()\nuser_freq = transactions.groupby('customer_id')['article_id'].nunique()\n\nitems = item_freq[item_freq >= 100].index\nusers = user_freq[user_freq >= 100].index\n\nfiltered_df = transactions[transactions['article_id'].isin(items) & transactions['customer_id'].isin(users)]","metadata":{"execution":{"iopub.status.busy":"2023-07-01T05:56:57.924388Z","iopub.status.idle":"2023-07-01T05:56:57.925328Z","shell.execute_reply.started":"2023-07-01T05:56:57.925001Z","shell.execute_reply":"2023-07-01T05:56:57.925040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next, let's aggregate to reflect the weight we learned in the previous tutorial on the edges. This is a more in-depth application of what we've learned on the journey so far. The weight given to the edge can contain various information values. By assigning appropriate values, we can also reflect in the graph how much the user prefers the item.\n\nLet's take an example. Consider user 0, item A, and item B as nodes. If user 0 purchased item A 10 times and item B once... We could say the preference for item A is higher, right?\n\nThere are two edges connecting these three nodes. The edges that stretch from the user 0 node to each item node. What relation might these edges imply? It's the purchase history! Therefore, we can assign a weight of 10 to the edge connected to item A node and a weight of 1 to the edge connected to item B.","metadata":{}},{"cell_type":"code","source":"freq = filtered_df.groupby(['customer_id', 'article_id']).size().reset_index(name='frequency')\n\nGraphTravel_HM = filtered_df.merge(freq, on=['customer_id', 'article_id'], how='left')\n\nGraphTravel_HM = GraphTravel_HM[GraphTravel_HM['frequency'] >= 10]","metadata":{"execution":{"iopub.status.busy":"2023-07-01T05:56:57.927385Z","iopub.status.idle":"2023-07-01T05:56:57.927929Z","shell.execute_reply.started":"2023-07-01T05:56:57.927633Z","shell.execute_reply":"2023-07-01T05:56:57.927661Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\nNow we have information that allows us to understand a specific user's preference (repeat purchase frequency) for an item. It seems we can assign weight information to the edge using the 'frequency' column!\n\n> Is it okay to use weight and random walk search bias together?\n\nYes! Weights and search bias can be used simultaneously and do not conflict with each other. Each controls a different aspect of the random walk, and using them together allows for more detailed control over how the random walk navigates the graph. To explain each concept a little more,\n\n1. Weight is used to indicate the importance of an edge and is used to determine the probability that the random walk algorithm will follow a particular edge. Edges with high weights have a higher probability of being chosen by the random walk than edges with low weights.\n\n2. Search bias is controlled by the two parameters p and q in Node2Vec. These parameters control the degree to which the random walk prefers to revisit nodes it has visited before or visit new nodes it has not yet visited.\n\nAnd lastly, in order to reflect only meaningful information in the Graph, the dataset was filtered to leave only rows of products purchased more than 10 times by a user. \n\nNow that some preprocessing seems to have been done, shall we take a moment to check the state of the dataset?","metadata":{}},{"cell_type":"code","source":"display(GraphTravel_HM)\n\nprint(\"unique customer_id\" , GraphTravel_HM.customer_id.nunique())\nprint(\"unique article_id\" , GraphTravel_HM.article_id.nunique())","metadata":{"execution":{"iopub.status.busy":"2023-07-01T05:56:57.929494Z","iopub.status.idle":"2023-07-01T05:56:57.930017Z","shell.execute_reply.started":"2023-07-01T05:56:57.929743Z","shell.execute_reply":"2023-07-01T05:56:57.929771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Wow, the original transactions dataframe had as many as 31,788,324 rows. It has significantly reduced. It's quite surprising that there are still around 25,000 left, even though we filtered with quite a high value.\n\nThere are 922 unique users and 1,013 unique items remaining. It seems reasonable to expect quite useful recommendation results from this amount.\n\nOn one hand, I'm curious about the distribution of frequency values. Shall we draw it directly and check?","metadata":{}},{"cell_type":"code","source":"sns.distplot(GraphTravel_HM['frequency'], kde=True, bins=30)\n\nplt.title('Distribution of frequency')\nplt.xlabel('Frequency')\nplt.ylabel('Density')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-01T05:56:57.931571Z","iopub.status.idle":"2023-07-01T05:56:57.932086Z","shell.execute_reply.started":"2023-07-01T05:56:57.931811Z","shell.execute_reply":"2023-07-01T05:56:57.931838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Considering we left only the products with more than 10 in the dataframe, the distribution is quite understandable. Most of the repurchase frequencies are concentrated between 10 and 100.\n\nIt looks like there's a product that's been purchased as many as 550 times... I wonder what this could be? Although EDA is not important to us right now and Erica will ignore her curiosity, those who are curious might find it fun to check for themselves.\n\nNow, let's move to the final step of preprocessing. The values of customer_id were mixed with integers and strings as hash values, weren't they? I'll map these to integers. If it's an integer starting from 0, it's likely not to overlap with the article_id.\n\nAlso, to print the item name along with the recommendation result, I'll create one dictionary as well.","metadata":{}},{"cell_type":"code","source":"unique_customer_ids = GraphTravel_HM['customer_id'].unique()\ncustomer_id_mapping = {id: i for i, id in enumerate(unique_customer_ids)}\nGraphTravel_HM['customer_id'] = GraphTravel_HM['customer_id'].map(customer_id_mapping)\n\nitem_name_mapping = dict(zip(articles['article_id'], articles['prod_name'])) # prod_name","metadata":{"execution":{"iopub.status.busy":"2023-07-01T05:56:57.933931Z","iopub.status.idle":"2023-07-01T05:56:57.934457Z","shell.execute_reply.started":"2023-07-01T05:56:57.934156Z","shell.execute_reply":"2023-07-01T05:56:57.934185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# customer_id_mapping","metadata":{"execution":{"iopub.status.busy":"2023-07-01T05:56:57.936352Z","iopub.status.idle":"2023-07-01T05:56:57.937030Z","shell.execute_reply.started":"2023-07-01T05:56:57.936826Z","shell.execute_reply":"2023-07-01T05:56:57.936847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that the preprocessing for building a recommendation system is complete, let's officially create a Graph.\n\nAfter creating the graph, I will implement and apply the Biased Random Walk and Generate Walk code.\nWhile we're at it, let's proceed straight through the process of extracting embeddings with the Word2Vec model!","metadata":{}},{"cell_type":"code","source":"G = nx.Graph()\n\nfor index, row in GraphTravel_HM.iterrows():\n    G.add_node(row['customer_id'], type='user')\n    G.add_node(row['article_id'], type='item')\n    G.add_edge(row['customer_id'], row['article_id'], weight=row['frequency'])","metadata":{"execution":{"iopub.status.busy":"2023-07-01T05:56:57.938248Z","iopub.status.idle":"2023-07-01T05:56:57.938591Z","shell.execute_reply.started":"2023-07-01T05:56:57.938404Z","shell.execute_reply":"2023-07-01T05:56:57.938421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# biased random walk  \ndef biased_random_walk(G, start_node, walk_length, p=1, q=1):\n    walk = [start_node]\n\n    while len(walk) < walk_length:\n        cur_node = walk[-1]\n        cur_neighbors = list(G.neighbors(cur_node))\n\n        if len(cur_neighbors) > 0:\n            if len(walk) == 1:\n                walk.append(random.choice(cur_neighbors))\n            else:\n                prev_node = walk[-2]\n\n                probability = []\n                for neighbor in cur_neighbors:\n                    if neighbor == prev_node:\n                        # Return parameter \n                        probability.append(1/p)\n                    elif G.has_edge(neighbor, prev_node):\n                        # Stay parameter \n                        probability.append(1)\n                    else:\n                        # In-out parameter \n                        probability.append(1/q)\n\n                probability = np.array(probability)\n                probability = probability / probability.sum()  # normalize\n\n                next_node = np.random.choice(cur_neighbors, p=probability)\n                walk.append(next_node)\n        else:\n            break\n\n    return walk","metadata":{"execution":{"iopub.status.busy":"2023-07-01T05:56:57.939421Z","iopub.status.idle":"2023-07-01T05:56:57.940370Z","shell.execute_reply.started":"2023-07-01T05:56:57.940109Z","shell.execute_reply":"2023-07-01T05:56:57.940152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def generate_walks(G, num_walks, walk_length, p=1, q=1):\n    walks = []\n    nodes = list(G.nodes())\n    for _ in range(num_walks):\n        random.shuffle(nodes)  # to ensure randomness\n        for node in nodes:\n            walk_from_node = biased_random_walk(G, node, walk_length, p, q)\n            walks.append(walk_from_node)\n    return walks","metadata":{"execution":{"iopub.status.busy":"2023-07-01T05:56:57.941327Z","iopub.status.idle":"2023-07-01T05:56:57.942181Z","shell.execute_reply.started":"2023-07-01T05:56:57.941821Z","shell.execute_reply":"2023-07-01T05:56:57.941843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# generate_walks(G, 2, 8, p=0.5, q=0.5)","metadata":{"execution":{"iopub.status.busy":"2023-07-01T05:56:57.943456Z","iopub.status.idle":"2023-07-01T05:56:57.943829Z","shell.execute_reply.started":"2023-07-01T05:56:57.943629Z","shell.execute_reply":"2023-07-01T05:56:57.943647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Random Walk \nwalks = generate_walks(G, num_walks=10, walk_length=20, p=9, q=1)\nfiltered_walks = [walk for walk in walks if len(walk) >= 5]\n\n# to String  (for Word2Vec input)\nwalks = [[str(node) for node in walk] for walk in walks]\n\n# Word2Vec train\nmodel = Word2Vec(walks, vector_size=128, window=5, min_count=0,  hs=1, sg=1, workers=4, epochs=10)\n\n# node embedding extract\nembeddings = {node_id: model.wv[node_id] for node_id in model.wv.index_to_key}","metadata":{"execution":{"iopub.status.busy":"2023-07-01T05:56:57.944745Z","iopub.status.idle":"2023-07-01T05:56:57.945076Z","shell.execute_reply.started":"2023-07-01T05:56:57.944897Z","shell.execute_reply":"2023-07-01T05:56:57.944915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# embeddings","metadata":{"execution":{"iopub.status.busy":"2023-07-01T05:56:57.947090Z","iopub.status.idle":"2023-07-01T05:56:57.947616Z","shell.execute_reply.started":"2023-07-01T05:56:57.947323Z","shell.execute_reply":"2023-07-01T05:56:57.947351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the middle of the code, it might be good to run the generate_walks() method separately or to print out the embeddings to visually check their shape at least once.\n\nThe line of code creating the Random Walk is a point of interest. In this tutorial code, p is set very high as generate_walks(G, num_walks=10, walk_length=20, p=9, q=1). I hope you try changing this number yourself and check how the recommendation result changes!\n\n\n## get_user_embedding \n\nThis function returns the Embedding of the given user Node. Here, embedding refers to a vector that quantifies the user's preferences or behavior patterns. This vector is used in the item recommendation algorithm we are implementing to calculate the similarity between the user and the item.","metadata":{}},{"cell_type":"code","source":"def get_user_embedding(user_id, embeddings):\n    return embeddings[str(user_id)]","metadata":{"execution":{"iopub.status.busy":"2023-07-01T05:56:57.949082Z","iopub.status.idle":"2023-07-01T05:56:57.949922Z","shell.execute_reply.started":"2023-07-01T05:56:57.949676Z","shell.execute_reply":"2023-07-01T05:56:57.949703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## get_rated_items\n\nThis function returns a list of items already rated (or purchased) by the given user. This information will be used to exclude items already rated by the user from the recommendation list.\nOf course, there's no rule that items with a history of purchase should not be recommended, and it could help encourage repurchase... but if the recommendation list becomes an order review list, there's no point in creating embeddings, right?\n\nI will set the purpose of this recommendation system to recommend unseen items. That means I'm going to recommend items that the user hasn't bought yet but might like!","metadata":{}},{"cell_type":"code","source":"def get_rated_items(user_id, df):\n    return set(df[df['customer_id'] == user_id]['article_id'])","metadata":{"execution":{"iopub.status.busy":"2023-07-01T05:56:57.950983Z","iopub.status.idle":"2023-07-01T05:56:57.951346Z","shell.execute_reply.started":"2023-07-01T05:56:57.951155Z","shell.execute_reply":"2023-07-01T05:56:57.951174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## calculate_similarities\n\nThis function calculates the similarity between the given user and all items, and returns this similarity as a list along with the item ID. In other words, it will calculate the cosine similarity between the user embedding and the embedding of each item. Items that the user has already rated are excluded from the similarity calculation.","metadata":{}},{"cell_type":"code","source":"def calculate_similarities(user_id, df, embeddings):\n    rated_items = get_rated_items(user_id, df)\n    user_embedding = get_user_embedding(user_id, embeddings)\n\n    item_similarities = []\n    for item_id in set(df['article_id']):\n        if item_id not in rated_items:  \n            item_embedding = embeddings[str(item_id)]\n            similarity = cosine_similarity([user_embedding], [item_embedding])[0][0]\n            item_similarities.append((item_id, similarity))\n\n    return item_similarities","metadata":{"execution":{"iopub.status.busy":"2023-07-01T05:56:57.952277Z","iopub.status.idle":"2023-07-01T05:56:57.952631Z","shell.execute_reply.started":"2023-07-01T05:56:57.952431Z","shell.execute_reply":"2023-07-01T05:56:57.952449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## show_images\n\nThis is a function that takes a list of items as input, prints out images for each item, and also prints out the similarity score. The item list given as input includes both the item ID and similarity.\n\nSince our currently implemented similarity-based recommendation system doesn't have a separately defined label to judge the accuracy of recommendations, we implemented the tutorial code to be able to see the images together in order to intuitively check the results of the recommendation system!","metadata":{}},{"cell_type":"code","source":"def show_images(items, item_name_mapping, num_items, show_similarity=False):\n    f, ax = plt.subplots(1, num_items, figsize=(20,10))\n    if num_items == 1:\n        ax = [ax]\n    for i, item in enumerate(items):\n        item_id, similarity = item\n        print(f\"- Item {item_id}: {item_name_mapping[item_id]}\", end='')\n        if show_similarity:\n            print(f\" with similarity score: {similarity}\")\n        else:\n            print()\n        img_path = f\"../input/h-and-m-personalized-fashion-recommendations/images/0{str(item_id)[:2]}/0{int(item_id)}.jpg\"\n        try:\n            img = mpimg.imread(img_path)\n            ax[i].imshow(img)\n            ax[i].set_title(f'Item {item_id}')\n            ax[i].set_xticks([], [])\n            ax[i].set_yticks([], [])\n            ax[i].grid(False)\n        except FileNotFoundError:\n            print(f\"Image for item {item_id} not found.\")\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-07-01T05:56:57.953910Z","iopub.status.idle":"2023-07-01T05:56:57.954244Z","shell.execute_reply.started":"2023-07-01T05:56:57.954065Z","shell.execute_reply":"2023-07-01T05:56:57.954083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## recommend_items\n\nThis function places the functions defined above appropriately, and outputs the final recommendation result. Ultimately, it finds items that are close in Embedding distance based on the user's transaction history, in other words, items the user might like, or in other words, the most similar items, and visually outputs the result.\n\nIt means that you can present a recommendation result that can be called personalized recommendation!","metadata":{}},{"cell_type":"code","source":"\ndef recommend_items(user_id, df, embeddings, item_name_mapping, num_items=5):\n    rated_items = get_rated_items(user_id, df)\n    \n    print(f\"User {user_id} has purchased:\")\n    show_images([(item_id, 0) for item_id in list(rated_items)[:5]], item_name_mapping, min(len(rated_items), 5))\n    \n    item_similarities = calculate_similarities(user_id, df, embeddings)\n\n    recommended_items = sorted(item_similarities, key=lambda x: x[1], reverse=True)[:num_items]\n\n    print(f\"\\nRecommended items for user {user_id}:\")\n    show_images(recommended_items, item_name_mapping, num_items, show_similarity=True)\n","metadata":{"execution":{"iopub.status.busy":"2023-07-01T05:56:57.955428Z","iopub.status.idle":"2023-07-01T05:56:57.955786Z","shell.execute_reply.started":"2023-07-01T05:56:57.955589Z","shell.execute_reply":"2023-07-01T05:56:57.955607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I made this code to output the top 5 recommended results.\nTo be able to judge visually whether it was a good recommendation, I arranged not only the similarity score but also the image.\n\n1. It prints out five items that the user has purchased in the past, and\n2. It presents five items with high similarity that are considered to be likable to the user.\n\nIn this way, you can judge whether it is a completely arbitrary recommendation or not, even without an accuracy score or labels, through similarity scores and images. Well, shall we finally activate the long-awaited first recommendation system?","metadata":{}},{"cell_type":"code","source":"# costomer 45's top 5 \nrecommend_items(45, GraphTravel_HM, embeddings, item_name_mapping, num_items=5)","metadata":{"execution":{"iopub.status.busy":"2023-07-01T05:56:57.957065Z","iopub.status.idle":"2023-07-01T05:56:57.957396Z","shell.execute_reply.started":"2023-07-01T05:56:57.957218Z","shell.execute_reply":"2023-07-01T05:56:57.957235Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Oh, quite interesting results came out, didn't they? Of course, it's more of a dream interpretation... The clothes that User 45 has purchased have the following design commonalities:\n\n1. Casual and comfortable clothes\n2. Clothes with shirring (wrinkles)\n3. Preference for both vivid primary colors and achromatic colors\n\nAnd the recommended items also seem to have similar commonalities! Plus, a hoodie is also recommended. It seems like we've made quite a decent recommendation with a simple code.\n\nIt would also be fun to manually print out the recommended results for other users by changing the numbers.","metadata":{}},{"cell_type":"markdown","source":"# If this tutorial was helpful, please Upvote! 🥰","metadata":{}},{"cell_type":"markdown","source":"혹시 한국인인가요? 이 튜토리얼 코드의 한국어 버전이 필요하다면, Graph User Group 사이트에 찾아오세요!\n\nGUG에서는 매주 GNN 기술을 논문과 연계해 공유하는 뉴스레터와, 튜토리얼이 연재되고 있답니다.\n\n✈🎫 Want to know about GUG? [Link](https://www.graphusergroup.com/)","metadata":{}}]}