{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Building a Recommendation System Using CNN \n![](https://www.researchgate.net/profile/Andreas_Veit/publication/282181243/figure/fig1/AS:360995122892808@1463079352720/Visualization-of-a-2D-embedding-of-the-style-space-trained-with-strategic-sampling.png)\n","metadata":{}},{"cell_type":"markdown","source":"\n## Introduction\n\nIn this notebook, I will use a CNN Model to create a Fashion Embedding. This information can be used in ML algorithms with higher semantic quality and similarity between Objects. We will use embeddings to identify similar items, this information will be used to recommend similar content in RecSys.\n\n* **Introduction**\n    * What is Embedding ?\n    * How to use Embedding ?\n* **Data Preparation**\n* **Use Pre-Trained Model to Recommendation**\n* Visualization Latent Space of Contents","metadata":{"_uuid":"108a9f6ae089a97c442efb22c7c5fb064aa98016"}},{"cell_type":"markdown","source":"#### Configure VM","metadata":{}},{"cell_type":"code","source":"# !pip install swifter\n# !pip install tensorflow==2.0.0","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2023-02-24T15:01:45.402565Z","iopub.execute_input":"2023-02-24T15:01:45.402874Z","iopub.status.idle":"2023-02-24T15:01:45.407976Z","shell.execute_reply.started":"2023-02-24T15:01:45.402819Z","shell.execute_reply":"2023-02-24T15:01:45.407137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## What is Embedding ?","metadata":{}},{"cell_type":"markdown","source":"An embedding is a relatively low-dimensional space into which you can translate high-dimensional vectors. Embeddings make it easier to do machine learning on large inputs like sparse vectors representing words. Ideally, an embedding captures some of the semantics of the input by placing semantically similar inputs close together in the embedding space. An embedding can be learned and reused across models.","metadata":{}},{"cell_type":"markdown","source":"So a natural language modelling technique like Word Embedding is used to map words or phrases from a vocabulary to a corresponding vector of real numbers. As well as being amenable to processing by learning algorithms, this vector representation has two important and advantageous properties:\n\n* **Dimensionality Reduction** — it is a more efficient representation\n* **Contextual Similarity** — it is a more expressive representation","metadata":{}},{"cell_type":"markdown","source":"![](https://encrypted-tbn0.gstatic.com/images?q=tbn:ANd9GcSiR683wW4f9httU7krJeLcgDQRB3Fmxi4v2SIr8QLSht204cmk&s)","metadata":{}},{"cell_type":"markdown","source":"We can use the Embedding as input of the model, containing a reduced dimensionality but with much semantic information. ","metadata":{}},{"cell_type":"markdown","source":"## Data Preparation\nTo begin this exploratory analysis, first use `matplotlib` to import libraries and define functions for plotting the data.","metadata":{"_uuid":"0a5dd1bb4db47e3313cc9957857eb3f4684dd11f"}},{"cell_type":"code","source":"from mpl_toolkits.mplot3d import Axes3D\nfrom sklearn.preprocessing import StandardScaler\nimport matplotlib.pyplot as plt # plotting\nimport matplotlib.image as mpimg\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os # accessing directory structure","metadata":{"_kg_hide-input":false,"_uuid":"875b42ec5baee5274279d8a7b7a72159f3a586de","execution":{"iopub.status.busy":"2023-02-24T15:01:45.410018Z","iopub.execute_input":"2023-02-24T15:01:45.410481Z","iopub.status.idle":"2023-02-24T15:01:46.565239Z","shell.execute_reply.started":"2023-02-24T15:01:45.410265Z","shell.execute_reply":"2023-02-24T15:01:46.564353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"DATASET_PATH = \"/kaggle/input/h-and-m-personalized-fashion-recommendations/images/\"\nprint(os.listdir(DATASET_PATH))\nDATA_PRODUCT = \"/kaggle/input/store-product/store_product.csv\"","metadata":{"_kg_hide-input":false,"_uuid":"7c96d605282a65cebf83737bbf0a3386c5c3f19e","execution":{"iopub.status.busy":"2023-02-24T15:01:46.567519Z","iopub.execute_input":"2023-02-24T15:01:46.568097Z","iopub.status.idle":"2023-02-24T15:01:46.585733Z","shell.execute_reply.started":"2023-02-24T15:01:46.568043Z","shell.execute_reply":"2023-02-24T15:01:46.584999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.read_csv(DATA_PRODUCT,encoding='unicode_escape')\ndf['image'] = df.apply(lambda row: str(row['images']).replace(\"photos/products/\",\"\"), axis=1)\ndf = df.reset_index(drop=True)\ndf.head(10)","metadata":{"_uuid":"f92cc76567188bdfe9dfa0d720d163f2702ab4fb","execution":{"iopub.status.busy":"2023-02-24T15:01:46.587920Z","iopub.execute_input":"2023-02-24T15:01:46.588359Z","iopub.status.idle":"2023-02-24T15:01:46.701156Z","shell.execute_reply.started":"2023-02-24T15:01:46.588308Z","shell.execute_reply":"2023-02-24T15:01:46.700386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import cv2\ndef plot_figures(figures, nrows = 1, ncols=1,figsize=(8, 8)):\n    \"\"\"Plot a dictionary of figures.\n\n    Parameters\n    ----------\n    figures : <title, figure> dictionary\n    ncols : number of columns of subplots wanted in the display\n    nrows : number of rows of subplots wanted in the figure\n    \"\"\"\n\n    fig, axeslist = plt.subplots(ncols=ncols, nrows=nrows,figsize=figsize)\n    for ind,title in enumerate(figures):\n        axeslist.ravel()[ind].imshow(cv2.cvtColor(figures[title], cv2.COLOR_BGR2RGB))\n        axeslist.ravel()[ind].set_title(title)\n        axeslist.ravel()[ind].set_axis_off()\n    plt.tight_layout() # optional\n    \ndef img_path(img):\n    return DATASET_PATH+img\n\ndef load_image(img, resized_fac = 0.1):\n    img     = cv2.imread(img_path(img))\n    w, h, _ = img.shape\n    resized = cv2.resize(img, (int(h*resized_fac), int(w*resized_fac)), interpolation = cv2.INTER_AREA)\n    return resized","metadata":{"execution":{"iopub.status.busy":"2023-02-24T15:01:46.703387Z","iopub.execute_input":"2023-02-24T15:01:46.703854Z","iopub.status.idle":"2023-02-24T15:01:46.917369Z","shell.execute_reply.started":"2023-02-24T15:01:46.703803Z","shell.execute_reply":"2023-02-24T15:01:46.916674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport numpy as np\n\n# generation of a dictionary of (title, images)\nfigures = {str(row.id): load_image(row.image) for i, row in df.sample(12).iterrows()}\n# plot of the images in a figure, with 2 rows and 3 columns\nplot_figures(figures, 2, 6)","metadata":{"execution":{"iopub.status.busy":"2023-02-24T15:01:46.920984Z","iopub.execute_input":"2023-02-24T15:01:46.921244Z","iopub.status.idle":"2023-02-24T15:01:48.254858Z","shell.execute_reply.started":"2023-02-24T15:01:46.921188Z","shell.execute_reply":"2023-02-24T15:01:48.253960Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The Dataset is made up of different items that can be found in a marketplace. The idea is to use embeddings to search for similarity and find similar items just using the image.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(7,20))\ndf.category_id.value_counts().sort_values().plot(kind='barh')","metadata":{"execution":{"iopub.status.busy":"2023-02-24T17:02:09.951409Z","iopub.execute_input":"2023-02-24T17:02:09.951790Z","iopub.status.idle":"2023-02-24T17:02:10.132601Z","shell.execute_reply.started":"2023-02-24T17:02:09.951726Z","shell.execute_reply":"2023-02-24T17:02:10.131657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Use Pre-Trained Model to Recommendation","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nimport keras\nfrom keras import Model\nfrom keras.applications.resnet50 import ResNet50\nfrom keras.preprocessing import image\nfrom keras.applications.resnet50 import preprocess_input, decode_predictions\nfrom keras.layers import GlobalMaxPooling2D\ntf.__version__","metadata":{"execution":{"iopub.status.busy":"2023-02-24T15:01:48.487336Z","iopub.execute_input":"2023-02-24T15:01:48.487806Z","iopub.status.idle":"2023-02-24T15:01:49.896626Z","shell.execute_reply.started":"2023-02-24T15:01:48.487752Z","shell.execute_reply":"2023-02-24T15:01:49.895768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nprint(os.listdir(\"../input\"))\nimport tensorflow as tf\nprint(tf.__version__)","metadata":{"execution":{"iopub.status.busy":"2023-02-24T16:47:06.632560Z","iopub.execute_input":"2023-02-24T16:47:06.632876Z","iopub.status.idle":"2023-02-24T16:47:06.638557Z","shell.execute_reply.started":"2023-02-24T16:47:06.632823Z","shell.execute_reply":"2023-02-24T16:47:06.637647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Input Shape\nimg_width, img_height, _ = 224, 224, 3 #load_image(df.iloc[0].image).shape\n\n# Pre-Trained Model\nbase_model = ResNet50(weights='imagenet', \n                      include_top=False, \n                      input_shape = (img_width, img_height, 3))\nbase_model.trainable = False\n\n# Add Layer Embedding\nmodel = keras.Sequential([\n    base_model,\n    GlobalMaxPooling2D()\n])\n\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2023-02-24T15:01:49.898255Z","iopub.execute_input":"2023-02-24T15:01:49.898844Z","iopub.status.idle":"2023-02-24T15:02:06.415806Z","shell.execute_reply.started":"2023-02-24T15:01:49.898787Z","shell.execute_reply":"2023-02-24T15:02:06.414985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_embedding(model, img_name):\n    # Reshape\n    img = image.load_img(img_path(img_name), target_size=(img_width, img_height))\n    # img to Array\n    x   = image.img_to_array(img)\n    # Expand Dim (1, w, h)\n    x   = np.expand_dims(x, axis=0)\n    # Pre process Input\n    x   = preprocess_input(x)\n    return model.predict(x).reshape(-1)","metadata":{"execution":{"iopub.status.busy":"2023-02-24T15:02:06.418922Z","iopub.execute_input":"2023-02-24T15:02:06.419165Z","iopub.status.idle":"2023-02-24T15:02:06.425232Z","shell.execute_reply.started":"2023-02-24T15:02:06.419118Z","shell.execute_reply":"2023-02-24T15:02:06.424426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Get item Embedding","metadata":{}},{"cell_type":"code","source":"df.iloc[2].image\nimg_path(df.iloc[20].image)\n# emb = get_embedding(model, img_path(df.iloc[23].image))\nemb = get_embedding(model, df.iloc[23].image)\nemb.shape","metadata":{"execution":{"iopub.status.busy":"2023-02-24T15:02:06.426782Z","iopub.execute_input":"2023-02-24T15:02:06.427353Z","iopub.status.idle":"2023-02-24T15:02:09.607837Z","shell.execute_reply.started":"2023-02-24T15:02:06.427302Z","shell.execute_reply":"2023-02-24T15:02:09.607011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"str_image = df.loc[df['id']==82].image.values[0]\n# print(one_image_path)\n# str_image = one_image_path.image.values[0]\n# print(str_image)\n# one_image_path.image\nimg_array = load_image(str(str_image))\nplt.imshow(cv2.cvtColor(img_array, cv2.COLOR_BGR2RGB))\nprint(img_array.shape)\nprint(emb)","metadata":{"execution":{"iopub.status.busy":"2023-02-24T16:10:44.257576Z","iopub.execute_input":"2023-02-24T16:10:44.257902Z","iopub.status.idle":"2023-02-24T16:10:44.437007Z","shell.execute_reply.started":"2023-02-24T16:10:44.257844Z","shell.execute_reply":"2023-02-24T16:10:44.436242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.shape","metadata":{"execution":{"iopub.status.busy":"2023-02-24T15:02:09.828733Z","iopub.execute_input":"2023-02-24T15:02:09.829387Z","iopub.status.idle":"2023-02-24T15:02:09.836870Z","shell.execute_reply.started":"2023-02-24T15:02:09.829321Z","shell.execute_reply":"2023-02-24T15:02:09.835351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Get Embedding for all itens in dataset","metadata":{}},{"cell_type":"code","source":"%%time\n#import swifter\n\n# Parallel apply\ndf_sample      = df#.sample(10)\nmap_embeddings = df_sample['image'].apply(lambda img: get_embedding(model, img))\ndf_embs        = map_embeddings.apply(pd.Series)\n\nprint(df_embs.shape)\ndf_embs.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-24T15:02:09.838410Z","iopub.execute_input":"2023-02-24T15:02:09.839136Z","iopub.status.idle":"2023-02-24T15:04:46.811934Z","shell.execute_reply.started":"2023-02-24T15:02:09.838916Z","shell.execute_reply":"2023-02-24T15:04:46.811145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Compute Similarity Between Items","metadata":{}},{"cell_type":"markdown","source":"![](http://dataaspirant.com/wp-content/uploads/2015/04/cosine.png)","metadata":{}},{"cell_type":"code","source":"# https://scikit-learn.org/stable/modules/generated/sklearn.metrics.pairwise_distances.html\nfrom sklearn.metrics.pairwise import pairwise_distances\n\n# Calcule DIstance Matriz\ncosine_sim = 1-pairwise_distances(df_embs, metric='cosine')\ncosine_sim[:4, :4]","metadata":{"execution":{"iopub.status.busy":"2023-02-24T15:06:43.846206Z","iopub.execute_input":"2023-02-24T15:06:43.846539Z","iopub.status.idle":"2023-02-24T15:06:44.101520Z","shell.execute_reply.started":"2023-02-24T15:06:43.846475Z","shell.execute_reply":"2023-02-24T15:06:44.100391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Recommender Similar Items","metadata":{}},{"cell_type":"code","source":"indices = pd.Series(range(len(df)), index=df.id)\nindices\n\n# Function that get movie recommendations based on the cosine similarity score of movie genres\ndef get_recommender(idx, df, top_n = 12):\n    sim_idx    = indices[idx]\n    sim_scores = list(enumerate(cosine_sim[sim_idx]))\n    sim_scores = sorted(sim_scores, key=lambda x: x[1], reverse=True)\n    sim_scores = sim_scores[1:top_n+1]\n#     print(sim_scores)\n    idx_rec    = [indices.index[i[0]] for i in sim_scores]\n    idx_sim    = [i[1] for i in sim_scores]\n#     print(idx_rec)\n    \n#     return indices.iloc[idx_rec].index, idx_sim\n    return idx_rec, idx_sim\n\nget_recommender(82, df)","metadata":{"execution":{"iopub.status.busy":"2023-02-24T17:32:51.007630Z","iopub.execute_input":"2023-02-24T17:32:51.007935Z","iopub.status.idle":"2023-02-24T17:32:51.019364Z","shell.execute_reply.started":"2023-02-24T17:32:51.007883Z","shell.execute_reply":"2023-02-24T17:32:51.018601Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"indices.index[0]","metadata":{"execution":{"iopub.status.busy":"2023-02-24T16:34:53.526569Z","iopub.execute_input":"2023-02-24T16:34:53.526870Z","iopub.status.idle":"2023-02-24T16:34:53.533324Z","shell.execute_reply.started":"2023-02-24T16:34:53.526818Z","shell.execute_reply":"2023-02-24T16:34:53.532576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Idx Item to Recommender\nidx_ref = 82\n\n# Recommendations\nidx_rec, idx_sim = get_recommender(idx_ref, df, top_n = 12)\n\n# Plot\n#===================\nstr_image = df.loc[df[\"id\"] == idx_ref].image.values[0]\nplt.imshow(cv2.cvtColor(load_image(str_image), cv2.COLOR_BGR2RGB))\n# generation of a dictionary of (title, images)\n\nfigures = {'id: '+str(i): load_image(df.loc[df[\"id\"] == i].image.values[0]) for i in idx_rec}\n# print(figures)\nplot_figures(figures, 2, 6)","metadata":{"execution":{"iopub.status.busy":"2023-02-24T17:31:04.984716Z","iopub.execute_input":"2023-02-24T17:31:04.985020Z","iopub.status.idle":"2023-02-24T17:31:06.160186Z","shell.execute_reply.started":"2023-02-24T17:31:04.984965Z","shell.execute_reply":"2023-02-24T17:31:06.159473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Create file","metadata":{}},{"cell_type":"code","source":"model.save('model_content_base2.h')\n# model = keras.models.load_model('path/to/location')","metadata":{"execution":{"iopub.status.busy":"2023-02-24T17:45:20.059360Z","iopub.execute_input":"2023-02-24T17:45:20.059729Z","iopub.status.idle":"2023-02-24T17:45:20.311651Z","shell.execute_reply.started":"2023-02-24T17:45:20.059666Z","shell.execute_reply":"2023-02-24T17:45:20.310803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"idx_rec","metadata":{"execution":{"iopub.status.busy":"2023-02-24T17:31:50.500757Z","iopub.execute_input":"2023-02-24T17:31:50.501064Z","iopub.status.idle":"2023-02-24T17:31:50.508595Z","shell.execute_reply.started":"2023-02-24T17:31:50.501008Z","shell.execute_reply":"2023-02-24T17:31:50.507572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# print(df_export)\narr_id_product = df[\"id\"].to_numpy()\nmy_dict_recomendation = {}\nfor i in arr_id_product:\n    recommen_list,tmp2 = get_recommender(i,df)\n    str_id_recom = ' '.join([str(elem) for elem in recommen_list])\n    print(str_id_recom)\n    my_dict_recomendation[i] = str_id_recom","metadata":{"execution":{"iopub.status.busy":"2023-02-24T17:38:19.372962Z","iopub.execute_input":"2023-02-24T17:38:19.373268Z","iopub.status.idle":"2023-02-24T17:38:23.472175Z","shell.execute_reply.started":"2023-02-24T17:38:19.373211Z","shell.execute_reply":"2023-02-24T17:38:23.471442Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(my_dict_recomendation)","metadata":{"execution":{"iopub.status.busy":"2023-02-24T17:39:37.522308Z","iopub.execute_input":"2023-02-24T17:39:37.522632Z","iopub.status.idle":"2023-02-24T17:39:37.531825Z","shell.execute_reply.started":"2023-02-24T17:39:37.522575Z","shell.execute_reply":"2023-02-24T17:39:37.531065Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"with open('my_dict_recomendation.csv', 'w') as f:\n    for key in my_dict_recomendation.keys():\n        f.write(\"%s, %s\\n\" % (key, my_dict_recomendation[key]))","metadata":{"execution":{"iopub.status.busy":"2023-02-24T17:41:31.313850Z","iopub.execute_input":"2023-02-24T17:41:31.314177Z","iopub.status.idle":"2023-02-24T17:41:31.325212Z","shell.execute_reply.started":"2023-02-24T17:41:31.314121Z","shell.execute_reply":"2023-02-24T17:41:31.324441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## new image path not id","metadata":{}},{"cell_type":"code","source":"# img_path = ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Conclusion\n\nConvolutional networks can be used to generate generic embeddings of any content. These embeddings can be used to identify similar items and in a recommendation process.\n\nA big improvement would be to retrain some network layers in a dataset similar to the one that will be used. So the network learns better features for a specific problem.","metadata":{}}]}