{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Purpose\nThe purpose of this notebook is to find similar products for any given article_id.\nThis is done by using TF-IDF and cosine similarity approach\n\nGiven any article ID the script finds 10 similar articles.\n\nPlease upvote if you find this helpful!","metadata":{}},{"cell_type":"code","source":"import cv2\nimport numpy as np # linear algebra\nimport os\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.express as px\nfrom os import listdir\nfrom os.path import isfile, join\n\nfrom termcolor import colored\nfrom IPython.display import HTML\nfrom PIL import Image\n\nimport warnings\npd.set_option('display.max_rows', None)\npd.set_option('display.max_columns', None)\npd.set_option('float_format', '{:f}'.format)\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:43:52.176922Z","iopub.execute_input":"2022-03-23T10:43:52.178259Z","iopub.status.idle":"2022-03-23T10:43:54.810294Z","shell.execute_reply.started":"2022-03-23T10:43:52.178081Z","shell.execute_reply":"2022-03-23T10:43:54.809395Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/articles.csv\")\ncustomers = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/customers.csv\")\ntransactions = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:43:57.676762Z","iopub.execute_input":"2022-03-23T10:43:57.677057Z","iopub.status.idle":"2022-03-23T10:45:07.507337Z","shell.execute_reply.started":"2022-03-23T10:43:57.677027Z","shell.execute_reply":"2022-03-23T10:45:07.506120Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:45:07.509067Z","iopub.execute_input":"2022-03-23T10:45:07.509370Z","iopub.status.idle":"2022-03-23T10:45:07.543400Z","shell.execute_reply.started":"2022-03-23T10:45:07.509331Z","shell.execute_reply":"2022-03-23T10:45:07.542408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Let's build a single word column for every article id present","metadata":{}},{"cell_type":"markdown","source":"**<span style=\"color:#023e8a;\"> This table contains all h&m articles with details such as a type of product, a color, a product group and other features.</span>**  \n**<span style=\"color:#023e8a;\"> Article data description: </span>**\n  \n- 105542 rows and 25 columns  \n- No nulls apart from detail_desc  \n- 11 int and 14 obj types are present \n\n> `article_id` **<span style=\"color:#023e8a;\">: A unique identifier of every article.</span>**  \n>  - The primary column  \n  \n> `product_code`, `prod_name` **<span style=\"color:#023e8a;\">: A unique identifier of every product and its name (not the same).</span>** >  - 47224 unique product_code  \n>  - product_code and article id are highly correlated  \n>  - 45875 uniqe prod_name. **Different from product_code**. Which one to use?  \n>  - No specific dominant levels in prod_name  \n  \n> `product_type`, `product_type_name` **<span style=\"color:#023e8a;\">: The group of product_code and its name</span>**  \n>  - 132 unique product types, but 131 unique product names  \n>  - dominant levels are present in both. first 8 form ~80% of total data  \n  \n> `product_group_name` **<span style=\"color:#023e8a;\">: 19 unique values. highly dominant levels are present.</span>**   \n  \n> `graphical_appearance_no`, `graphical_appearance_name` **<span style=\"color:#023e8a;\">: The group of graphics and its name</span>**  \n>  - both has 30 unique values. 1-1 mapping  \n>  - highly dominant levels present  \n  \n> `colour_group_code`, `colour_group_name` **<span style=\"color:#023e8a;\">: The group of color and its name</span>**  \n>  - both 50 unique values. 1-1 mapping  \n>  - mildly dominant levels  \n  \n> `perceived_colour_value_id`, `perceived_colour_value_name`, `perceived_colour_master_id`, `perceived_colour_master_name` **<span style=\"color:#023e8a;\">: The added color info</span>**    \n>  - only 8 levels in both\n  \n> `department_no`, `department_name` **<span style=\"color:#023e8a;\">: A unique identifier of every dep and its name</span>**  \n>  - 299 unique department_no, 250 unique department_name. **Not matching**  \n>  - no dominant levels.   \n  \n> `index_code`, `index_name` **<span style=\"color:#023e8a;\">: A unique identifier of every index and its name</span>**  \n>  - 10 levels in both 1-1 mapping.  \n>  - obviously dominant  \n  \n> `index_group_no`, `index_group_name` **<span style=\"color:#023e8a;\">: A group of indeces and its name</span>**  \n>  - 5 levels in both  \n>  - obviously dominant  \n  \n> `section_no`, `section_name` **<span style=\"color:#023e8a;\">: A unique identifier of every section and its name</span>**  \n>  - 57 in section no, 56 in section name . **Not one-one matching**  \n>  - Non dominant  \n  \n> `garment_group_no`, `garment_group_name` **<span style=\"color:#023e8a;\">: A unique identifier of every garment and its name</span>**  \n>  - 21 in both levels. 1-1 mapping  \n>  - some dominant levels  \n  \n> `detail_desc` **<span style=\"color:#023e8a;\">: Details</span>**  \n>  - All unique descriptions. many are nulls. Not sure how helpful it'll be","metadata":{}},{"cell_type":"markdown","source":"## Most of the columns are paired and hence only the string versions(& not ID) of these columns are taken.\nThese are then processed as shown below","metadata":{}},{"cell_type":"code","source":"articles_sub = articles[['article_id','prod_name','product_type_name','product_group_name','graphical_appearance_name','colour_group_name'\n                         ,'perceived_colour_value_name','perceived_colour_master_name','department_name','index_name','index_group_name'\n                         ,'section_name','garment_group_name','detail_desc']]\narticles_sub.shape","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:45:07.545359Z","iopub.execute_input":"2022-03-23T10:45:07.545694Z","iopub.status.idle":"2022-03-23T10:45:07.571952Z","shell.execute_reply.started":"2022-03-23T10:45:07.545649Z","shell.execute_reply":"2022-03-23T10:45:07.571062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's remove space in all string columns\nfor i in articles_sub.columns[1:]:\n    articles_sub[i] = articles_sub[i].str.replace(\" \",\"\")","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:45:07.574700Z","iopub.execute_input":"2022-03-23T10:45:07.575039Z","iopub.status.idle":"2022-03-23T10:45:08.653212Z","shell.execute_reply.started":"2022-03-23T10:45:07.574997Z","shell.execute_reply":"2022-03-23T10:45:08.652209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Combine all info from columns to a single column separated by space\n\ncols = ['prod_name', 'product_type_name', 'product_group_name',\n       'graphical_appearance_name', 'colour_group_name',\n       'perceived_colour_value_name', 'perceived_colour_master_name',\n       'department_name', 'index_name', 'index_group_name', 'section_name',\n       'garment_group_name', 'detail_desc']\narticles_sub['combined'] = articles_sub[cols].apply(lambda row: ' '.join(row.values.astype(str)), axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:45:08.654672Z","iopub.execute_input":"2022-03-23T10:45:08.654929Z","iopub.status.idle":"2022-03-23T10:45:11.687826Z","shell.execute_reply.started":"2022-03-23T10:45:08.654899Z","shell.execute_reply":"2022-03-23T10:45:11.687169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles_sub.head()","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:45:11.689138Z","iopub.execute_input":"2022-03-23T10:45:11.689412Z","iopub.status.idle":"2022-03-23T10:45:11.710737Z","shell.execute_reply.started":"2022-03-23T10:45:11.689379Z","shell.execute_reply":"2022-03-23T10:45:11.709871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles_final = articles_sub[['article_id','combined']]","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:45:11.712048Z","iopub.execute_input":"2022-03-23T10:45:11.712311Z","iopub.status.idle":"2022-03-23T10:45:11.750485Z","shell.execute_reply.started":"2022-03-23T10:45:11.712280Z","shell.execute_reply":"2022-03-23T10:45:11.749690Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Find related articles of all articles\nGiven an article_id, let's find 10 similar products for it","metadata":{}},{"cell_type":"code","source":"#Only 5000 products are taken because of computational issues\narticles_final = articles_final.loc[:5000]","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:45:11.764645Z","iopub.execute_input":"2022-03-23T10:45:11.765378Z","iopub.status.idle":"2022-03-23T10:45:11.769358Z","shell.execute_reply.started":"2022-03-23T10:45:11.765332Z","shell.execute_reply":"2022-03-23T10:45:11.768588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Import TfIdfVectorizer from scikit-learn\nfrom sklearn.feature_extraction.text import TfidfVectorizer\n\n#Define a TF-IDF Vectorizer Object. Remove all english stop words such as 'the', 'a'\ntfidf = TfidfVectorizer(stop_words='english')\n\n#Replace NaN with an empty string\narticles_final['combined'] = articles_final['combined'].fillna('')\n\n#Construct the required TF-IDF matrix by fitting and transforming the data\ntfidf_matrix = tfidf.fit_transform(articles_final['combined'])\n\n#Output the shape of tfidf_matrix\ntfidf_matrix.shape","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:45:11.770522Z","iopub.execute_input":"2022-03-23T10:45:11.771248Z","iopub.status.idle":"2022-03-23T10:45:12.058080Z","shell.execute_reply.started":"2022-03-23T10:45:11.771176Z","shell.execute_reply":"2022-03-23T10:45:12.057227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import linear_kernel\nfrom sklearn.metrics.pairwise import linear_kernel\n\n# Compute the cosine similarity matrix\ncosine_sim = linear_kernel(tfidf_matrix, tfidf_matrix)","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:45:12.060639Z","iopub.execute_input":"2022-03-23T10:45:12.060887Z","iopub.status.idle":"2022-03-23T10:45:12.714831Z","shell.execute_reply.started":"2022-03-23T10:45:12.060857Z","shell.execute_reply":"2022-03-23T10:45:12.714058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"indices = pd.Series(articles_final.index, index=articles_final['article_id']).drop_duplicates()","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:45:12.716056Z","iopub.execute_input":"2022-03-23T10:45:12.716465Z","iopub.status.idle":"2022-03-23T10:45:12.721580Z","shell.execute_reply.started":"2022-03-23T10:45:12.716421Z","shell.execute_reply":"2022-03-23T10:45:12.720775Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Function that takes in article_id as input and outputs most similar articles\ndef get_recommendations(title, cosine_sim=cosine_sim):\n    # Get the index of the article that matches the title\n    idx = indices[title]\n\n    # Get the pairwsie similarity scores of all articles\n    sim_scores = list(enumerate(cosine_sim[idx]))\n\n    # Sort the articles based on the similarity scores\n    sim_scores = sorted(sim_scores, key=lambda x: x[1], reverse=True)\n\n    # Get the scores of the 10 most similar articles\n    sim_scores = sim_scores[:12]\n\n    # Get the article indices\n    article_indices = [i[0] for i in sim_scores]\n\n    # Return the top 10 most similar articles\n    return articles_final['article_id'].iloc[article_indices]\n","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:45:12.722936Z","iopub.execute_input":"2022-03-23T10:45:12.723885Z","iopub.status.idle":"2022-03-23T10:45:12.732390Z","shell.execute_reply.started":"2022-03-23T10:45:12.723848Z","shell.execute_reply":"2022-03-23T10:45:12.731735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"recom = list(get_recommendations(108775044))\nrecom","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:45:12.733594Z","iopub.execute_input":"2022-03-23T10:45:12.734103Z","iopub.status.idle":"2022-03-23T10:45:12.902283Z","shell.execute_reply.started":"2022-03-23T10:45:12.734069Z","shell.execute_reply":"2022-03-23T10:45:12.901687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualising the predictions","metadata":{}},{"cell_type":"code","source":"def display_articles(article_ids):\n    rows = 4 #len(article_ids)\n    cols = 3\n    image_path = \"/kaggle/input/h-and-m-personalized-fashion-recommendations/images/\"\n    plt.figure(figsize=(2 + 3 * cols, 2 + 4 * rows))\n    for i in range(len(article_ids)):\n\n        article_id = (\"0\" + str(article_ids[i]))[-10:]\n        plt.subplot(rows, cols, i + 1)\n        plt.axis('off')\n        #plt.title(f\"{product_group_name} {article_id[:3]}\\n{article_id}.jpg\")\n        try:\n            image = Image.open(f\"{image_path}{article_id[:3]}/{article_id}.jpg\")\n            plt.imshow(image)\n        except:\n            None","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:47:01.204080Z","iopub.execute_input":"2022-03-23T10:47:01.204658Z","iopub.status.idle":"2022-03-23T10:47:01.212364Z","shell.execute_reply.started":"2022-03-23T10:47:01.204598Z","shell.execute_reply":"2022-03-23T10:47:01.211449Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#First image (top left) is the input article and rest are all recommended similar articles\ndisplay_articles(recom)\n","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:47:03.707444Z","iopub.execute_input":"2022-03-23T10:47:03.707783Z","iopub.status.idle":"2022-03-23T10:47:07.967353Z","shell.execute_reply.started":"2022-03-23T10:47:03.707744Z","shell.execute_reply":"2022-03-23T10:47:07.966270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Let's try with few more articles","metadata":{}},{"cell_type":"code","source":"recom = list(get_recommendations(252298006))\ndisplay_articles(recom)","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:47:16.044007Z","iopub.execute_input":"2022-03-23T10:47:16.044767Z","iopub.status.idle":"2022-03-23T10:47:20.044991Z","shell.execute_reply.started":"2022-03-23T10:47:16.044726Z","shell.execute_reply":"2022-03-23T10:47:20.044108Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"recom = list(get_recommendations(224337008))\ndisplay_articles(recom)","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:47:26.888969Z","iopub.execute_input":"2022-03-23T10:47:26.889633Z","iopub.status.idle":"2022-03-23T10:47:31.760910Z","shell.execute_reply.started":"2022-03-23T10:47:26.889569Z","shell.execute_reply":"2022-03-23T10:47:31.759969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"recom = list(get_recommendations(245348002))\ndisplay_articles(recom)","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:47:44.830787Z","iopub.execute_input":"2022-03-23T10:47:44.831095Z","iopub.status.idle":"2022-03-23T10:47:50.716021Z","shell.execute_reply.started":"2022-03-23T10:47:44.831064Z","shell.execute_reply":"2022-03-23T10:47:50.712345Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"recom = list(get_recommendations(112679048))\ndisplay_articles(recom)","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:47:57.684839Z","iopub.execute_input":"2022-03-23T10:47:57.685673Z","iopub.status.idle":"2022-03-23T10:48:03.580968Z","shell.execute_reply.started":"2022-03-23T10:47:57.685626Z","shell.execute_reply":"2022-03-23T10:48:03.580275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Denim short\nrecom = list(get_recommendations(156289011))\ndisplay_articles(recom)","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:48:09.336330Z","iopub.execute_input":"2022-03-23T10:48:09.336808Z","iopub.status.idle":"2022-03-23T10:48:14.116187Z","shell.execute_reply.started":"2022-03-23T10:48:09.336775Z","shell.execute_reply":"2022-03-23T10:48:14.115139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Greenish Khakhi short\nrecom = list(get_recommendations(212766041))\ndisplay_articles(recom)","metadata":{"execution":{"iopub.status.busy":"2022-03-23T10:48:14.118020Z","iopub.execute_input":"2022-03-23T10:48:14.118811Z","iopub.status.idle":"2022-03-23T10:48:18.474967Z","shell.execute_reply.started":"2022-03-23T10:48:14.118756Z","shell.execute_reply":"2022-03-23T10:48:18.474019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Under development","metadata":{}},{"cell_type":"markdown","source":"# Join articles and transaction details","metadata":{}},{"cell_type":"code","source":"transactions = transactions[transactions.customer_id.isin(['000058a12d5b43e67d225668fa1f8d618c13dc232df0cad8ffe7ad4a1091e318','00007d2de826758b65a93dd24ce629ed66842531df6699338c5570910a014cc2','00083cda041544b2fbb0e0d2905ad17da7cf1007526fb4c73235dccbbc132280'])]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"art_trans = pd.merge(articles_final, transactions, on='article_id')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tot_trans = pd.merge(art_trans,customers,on='customer_id')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tot_trans.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tot_trans.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fin_art = tot_trans[['customer_id','combined']].groupby(['customer_id'])['combined'].apply(' '.join).reset_index()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fin_art","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Customer ID X Article description","metadata":{}},{"cell_type":"code","source":"#Import TfIdfVectorizer from scikit-learn\nfrom sklearn.feature_extraction.text import TfidfVectorizer\n\n#Define a TF-IDF Vectorizer Object. Remove all english stop words such as 'the', 'a'\ntfidf = TfidfVectorizer(stop_words='english')\n\n#Replace NaN with an empty string\nfin_art['combined'] = fin_art['combined'].fillna('')\n\n#Construct the required TF-IDF matrix by fitting and transforming the data\ntfidf_matrix = tfidf.fit_transform(fin_art['combined'])\n\n#Output the shape of tfidf_matrix\ntfidf_matrix.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import linear_kernel\nfrom sklearn.metrics.pairwise import linear_kernel\n\n# Compute the cosine similarity matrix\ncosine_sim = linear_kernel(tfidf_matrix, tfidf_matrix)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cosine_sim.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"indices = pd.Series(fin_art.index, index=fin_art['customer_id']).drop_duplicates()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Function that takes in movie title as input and outputs most similar movies\ndef get_recommendations(title, cosine_sim=cosine_sim):\n    # Get the index of the movie that matches the title\n    idx = indices[title]\n\n    # Get the pairwsie similarity scores of all movies with that movie\n    sim_scores = list(enumerate(cosine_sim[idx]))\n\n    # Sort the movies based on the similarity scores\n    sim_scores = sorted(sim_scores, key=lambda x: x[1], reverse=True)\n\n    # Get the scores of the 10 most similar movies\n    sim_scores = sim_scores[1:11]\n\n    # Get the movie indices\n    movie_indices = [i[0] for i in sim_scores]\n\n    # Return the top 10 most similar movies\n    return fin_art['customer_id'].iloc[movie_indices]\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"get_recommendations('00083cda041544b2fbb0e0d2905ad17da7cf1007526fb4c73235dccbbc132280')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Find related articles of all articles","metadata":{}},{"cell_type":"code","source":"articles_final = articles_final.loc[:500]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Import TfIdfVectorizer from scikit-learn\nfrom sklearn.feature_extraction.text import TfidfVectorizer\n\n#Define a TF-IDF Vectorizer Object. Remove all english stop words such as 'the', 'a'\ntfidf = TfidfVectorizer(stop_words='english')\n\n#Replace NaN with an empty string\narticles_final['combined'] = articles_final['combined'].fillna('')\n\n#Construct the required TF-IDF matrix by fitting and transforming the data\ntfidf_matrix = tfidf.fit_transform(articles_final['combined'])\n\n#Output the shape of tfidf_matrix\ntfidf_matrix.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import linear_kernel\nfrom sklearn.metrics.pairwise import linear_kernel\n\n# Compute the cosine similarity matrix\ncosine_sim = linear_kernel(tfidf_matrix, tfidf_matrix)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"indices = pd.Series(articles_final.index, index=articles_final['article_id']).drop_duplicates()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Function that takes in movie title as input and outputs most similar movies\ndef get_recommendations(title, cosine_sim=cosine_sim):\n    # Get the index of the movie that matches the title\n    idx = indices[title]\n\n    # Get the pairwsie similarity scores of all movies with that movie\n    sim_scores = list(enumerate(cosine_sim[idx]))\n\n    # Sort the movies based on the similarity scores\n    sim_scores = sorted(sim_scores, key=lambda x: x[1], reverse=True)\n\n    # Get the scores of the 10 most similar movies\n    sim_scores = sim_scores[:12]\n\n    # Get the movie indices\n    movie_indices = [i[0] for i in sim_scores]\n\n    # Return the top 10 most similar movies\n    return articles_final['article_id'].iloc[movie_indices]\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"recom = list(get_recommendations(108775044))\nrecom","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualise the predictions","metadata":{}},{"cell_type":"code","source":"def display_articles(article_ids):\n    rows = 4 #len(article_ids)\n    cols = 3\n    image_path = \"/kaggle/input/h-and-m-personalized-fashion-recommendations/images/\"\n    plt.figure(figsize=(2 + 3 * cols, 2 + 4 * rows))\n    for i in range(len(article_ids)):\n\n        article_id = (\"0\" + str(article_ids[i]))[-10:]\n        plt.subplot(rows, cols, i + 1)\n        plt.axis('off')\n        #plt.title(f\"{product_group_name} {article_id[:3]}\\n{article_id}.jpg\")\n        try:\n            image = Image.open(f\"{image_path}{article_id[:3]}/{article_id}.jpg\")\n            plt.imshow(image)\n        except:\n            None","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display_articles(recom)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}