{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# H&M Fashion | EDA | H&M Personalized Fashion Recommendations\nHello everyone! In this my new notebook we are going to look through [H&M Personalized Fashion Recommendations](https://www.kaggle.com/c/h-and-m-personalized-fashion-recommendations) Prediction Competition.\n\n#### Acknowledgements 😍\nMy acknowledgments are given to:\n* [GABRIEL PREDA. H&M EDA and Prediction](https://www.kaggle.com/gpreda/h-m-eda-and-prediction)\n\n#### About 👚\nH&M Group is a family of brands and businesses with 53 online markets and approximately 4,850 stores. In this competition, H&M Group invites you to develop product recommendations based on data from previous transactions, as well as from customer and product meta data.","metadata":{}},{"cell_type":"markdown","source":"# 1. Import libraries 📚\nHere we import libraries that will be used.","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nfrom PIL import Image\nfrom tqdm import tqdm\nfrom datetime import datetime\nimport matplotlib.pyplot as plt\nfrom wordcloud import WordCloud, STOPWORDS","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-23T10:26:22.714463Z","iopub.execute_input":"2022-02-23T10:26:22.7148Z","iopub.status.idle":"2022-02-23T10:26:23.722723Z","shell.execute_reply.started":"2022-02-23T10:26:22.71471Z","shell.execute_reply":"2022-02-23T10:26:23.72202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Read data 📖","metadata":{}},{"cell_type":"code","source":"total_folders = 0\ntotal_files = 0\n\nfolder_info = []\nimages_names = []\n\npath = \"../input/h-and-m-personalized-fashion-recommendations\"\n\nfor base, dirs, files in tqdm(os.walk(path)):\n    for directories in dirs:\n        folder_info.append((directories, \n                            len(os.listdir(os.path.join(base, directories)))))\n        total_folders = total_folders + 1\n    \n    for _files in files:\n        total_files = total_files + 1\n        if (len(_files.split(\".jpg\"))==2):\n            images_names.append(_files.split(\".jpg\")[0])","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-02-23T10:26:23.724617Z","iopub.execute_input":"2022-02-23T10:26:23.725009Z","iopub.status.idle":"2022-02-23T10:26:49.658217Z","shell.execute_reply.started":"2022-02-23T10:26:23.724964Z","shell.execute_reply":"2022-02-23T10:26:49.657231Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After that we can check the final reading result:","metadata":{}},{"cell_type":"code","source":"print(f\"• Total number of folders: {total_folders}\")\nprint(f\"• Total number of files: {total_files}\")","metadata":{"execution":{"iopub.status.busy":"2022-02-23T10:26:49.65986Z","iopub.execute_input":"2022-02-23T10:26:49.660416Z","iopub.status.idle":"2022-02-23T10:26:49.667229Z","shell.execute_reply.started":"2022-02-23T10:26:49.660372Z","shell.execute_reply":"2022-02-23T10:26:49.666456Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"folder_info_df = pd.DataFrame(folder_info, \n                              columns=[\"folder\", \n                                       \"files count\"])\n\nfolder_info_df.sort_values([\"files count\"], ascending=False).head().style.set_properties(**{'background-color': 'rgba(184,230,194,.5)'})","metadata":{"execution":{"iopub.status.busy":"2022-02-23T11:27:41.235077Z","iopub.execute_input":"2022-02-23T11:27:41.235972Z","iopub.status.idle":"2022-02-23T11:27:41.253033Z","shell.execute_reply.started":"2022-02-23T11:27:41.235928Z","shell.execute_reply":"2022-02-23T11:27:41.252106Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"articles_df = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/articles.csv\")\ncustomers_df = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/customers.csv\")\nsample_submission_df = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/sample_submission.csv\")\n\ntransactions_train_df = pd.read_csv(\"../input/h-and-m-personalized-fashion-recommendations/transactions_train.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-02-23T10:26:49.711445Z","iopub.execute_input":"2022-02-23T10:26:49.712064Z","iopub.status.idle":"2022-02-23T10:28:12.571557Z","shell.execute_reply.started":"2022-02-23T10:26:49.712023Z","shell.execute_reply":"2022-02-23T10:28:12.570228Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After that, we can check, what we've read:","metadata":{}},{"cell_type":"code","source":"# articles_df\narticles_df.head(3).style.set_properties(**{'background-color': 'rgba(184,230,194,.5)'})","metadata":{"execution":{"iopub.status.busy":"2022-02-23T10:28:12.573585Z","iopub.execute_input":"2022-02-23T10:28:12.573845Z","iopub.status.idle":"2022-02-23T10:28:12.695839Z","shell.execute_reply.started":"2022-02-23T10:28:12.573805Z","shell.execute_reply":"2022-02-23T10:28:12.695058Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# customers_df\ncustomers_df.head(3).style.set_properties(**{'background-color': 'rgba(184,230,194,.5)'})","metadata":{"execution":{"iopub.status.busy":"2022-02-23T10:28:12.697178Z","iopub.execute_input":"2022-02-23T10:28:12.697408Z","iopub.status.idle":"2022-02-23T10:28:12.713523Z","shell.execute_reply.started":"2022-02-23T10:28:12.697379Z","shell.execute_reply":"2022-02-23T10:28:12.712865Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# sample_submission_df\n# prediction in sample submission is a sequence of article ids (max 12 article ids)\nsample_submission_df.head(3).style.set_properties(**{'background-color': 'rgba(184,230,194,.5)'})","metadata":{"execution":{"iopub.status.busy":"2022-02-23T10:28:12.716317Z","iopub.execute_input":"2022-02-23T10:28:12.7167Z","iopub.status.idle":"2022-02-23T10:28:12.733502Z","shell.execute_reply.started":"2022-02-23T10:28:12.716658Z","shell.execute_reply":"2022-02-23T10:28:12.73259Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# transactions_train_df\ntransactions_train_df.head(3).style.set_properties(**{'background-color': 'rgba(184,230,194,.5)'})","metadata":{"execution":{"iopub.status.busy":"2022-02-23T10:28:12.734819Z","iopub.execute_input":"2022-02-23T10:28:12.735134Z","iopub.status.idle":"2022-02-23T10:28:12.754976Z","shell.execute_reply.started":"2022-02-23T10:28:12.735095Z","shell.execute_reply":"2022-02-23T10:28:12.754295Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. Feature engineering 💻\n***Transactions table*** is the train data. It contains customer_id and article_id, which are foreign keys for the customer and articles tables. Also, Transactions also contains sales_channel_id.\n\nHere we can **check missing data, count unique values** etc.","metadata":{}},{"cell_type":"code","source":"def missing_data(data):\n    total = data.isnull().sum().sort_values(ascending = False)\n    percent = (data.isnull().sum()/data.isnull().count()*100).sort_values(ascending = False)\n    return pd.concat([total, percent], axis=1, keys=['Total', 'Percent'])\n\ndef unique_values(data):\n    total = data.count()\n    tt = pd.DataFrame(total)\n    tt.columns = ['Total']\n    uniques = []\n    for col in data.columns:\n        unique = data[col].nunique()\n        uniques.append(unique)\n    tt['Uniques'] = uniques\n    return tt","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-23T10:28:12.756225Z","iopub.execute_input":"2022-02-23T10:28:12.756603Z","iopub.status.idle":"2022-02-23T10:28:12.770857Z","shell.execute_reply.started":"2022-02-23T10:28:12.756566Z","shell.execute_reply":"2022-02-23T10:28:12.769986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here we can look through missing data in different files.\n\nIn the **articles_df**, the only missing data is for the detailed description of the article (0.4% missing data).\n\nIn the **cutomers_df**, only *customer_id* and *postal_code* are completely filled. *Age*, *fashion_news_frequency* have around 1% misssing data, *FN* has 65% missing and *Active* has 66% missing data.\n\nIn the **transactions_train_df**, there is no missing data.","metadata":{}},{"cell_type":"code","source":"missing_data(articles_df).head(7).style.set_properties(**{'background-color': 'rgba(245, 181, 152,.5)'})","metadata":{"execution":{"iopub.status.busy":"2022-02-23T10:28:12.772297Z","iopub.execute_input":"2022-02-23T10:28:12.772534Z","iopub.status.idle":"2022-02-23T10:28:13.286063Z","shell.execute_reply.started":"2022-02-23T10:28:12.772507Z","shell.execute_reply":"2022-02-23T10:28:13.285194Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_data(customers_df).style.set_properties(**{'background-color': 'rgba(245, 245, 152,.5)'})","metadata":{"execution":{"iopub.status.busy":"2022-02-23T10:28:13.287374Z","iopub.execute_input":"2022-02-23T10:28:13.28821Z","iopub.status.idle":"2022-02-23T10:28:15.128235Z","shell.execute_reply.started":"2022-02-23T10:28:13.288161Z","shell.execute_reply":"2022-02-23T10:28:15.127405Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_data(transactions_train_df).style.set_properties(**{'background-color': 'rgba(152, 243, 245,.5)'})","metadata":{"execution":{"iopub.status.busy":"2022-02-23T10:28:15.129514Z","iopub.execute_input":"2022-02-23T10:28:15.130196Z","iopub.status.idle":"2022-02-23T10:28:35.791542Z","shell.execute_reply.started":"2022-02-23T10:28:15.130159Z","shell.execute_reply":"2022-02-23T10:28:35.790648Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After that we can check **unique values**:","metadata":{}},{"cell_type":"code","source":"unique_values(articles_df).head(5).style.set_properties(**{'background-color': 'rgba(145, 178, 227,.5)'})","metadata":{"execution":{"iopub.status.busy":"2022-02-23T10:28:35.792863Z","iopub.execute_input":"2022-02-23T10:28:35.793205Z","iopub.status.idle":"2022-02-23T10:28:36.1368Z","shell.execute_reply.started":"2022-02-23T10:28:35.793172Z","shell.execute_reply":"2022-02-23T10:28:36.135928Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"unique_values(customers_df).head(5).style.set_properties(**{'background-color': 'rgba(130, 126, 230,.5)'})","metadata":{"execution":{"iopub.status.busy":"2022-02-23T10:28:36.138135Z","iopub.execute_input":"2022-02-23T10:28:36.138355Z","iopub.status.idle":"2022-02-23T10:28:38.396741Z","shell.execute_reply.started":"2022-02-23T10:28:36.138328Z","shell.execute_reply":"2022-02-23T10:28:38.395923Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"unique_values(transactions_train_df).head(5).style.set_properties(**{'background-color': 'rgba(188, 126, 230,.5)'})","metadata":{"execution":{"iopub.status.busy":"2022-02-23T10:28:38.397766Z","iopub.execute_input":"2022-02-23T10:28:38.398014Z","iopub.status.idle":"2022-02-23T10:28:57.218695Z","shell.execute_reply.started":"2022-02-23T10:28:38.397986Z","shell.execute_reply":"2022-02-23T10:28:57.217805Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that **not all** the customers in customer data are appearing as having transactions in transaction train data. \n**Moreover,** not all articles are represented in this data. \n\nIt is **interesting** that the number of different prices is quite small, out of 31.7M transactions, and for 1.3M customers, buying 104K different articles. Same for the dates, there are only 734 different dates. ","metadata":{}},{"cell_type":"markdown","source":"# 4. Data visualisations. Articles Data 📊","metadata":{}},{"cell_type":"code","source":"def pie_chart(df, col_values, labels, ax, color, title):\n    n_classes = len(df)\n    explode = (0.1,) * n_classes # explode for 0.1 each slice\n    ax.pie(df[col_values],\n           colors=color, \n           explode=explode,\n           labels=df[labels],\n           shadow=True)\n    ax.set_title(title, fontsize=16)\n    \ndef bar_plot(df, col_x, col_y, ax, color, title):\n    ax.bar(x=df[col_x],\n           height=df[col_y],\n           color=color)\n    ax.set_title(title, fontsize=16) \n    plt.xticks(rotation=90)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-23T10:28:57.220181Z","iopub.execute_input":"2022-02-23T10:28:57.220488Z","iopub.status.idle":"2022-02-23T10:28:57.228923Z","shell.execute_reply.started":"2022-02-23T10:28:57.220448Z","shell.execute_reply":"2022-02-23T10:28:57.228238Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = articles_df.groupby([\"product_group_name\"])[\"product_type_name\"].nunique()\ndf = pd.DataFrame({'Product Group': temp.index,'Product Types': temp.values})\ndf = df.sort_values(['Product Types'], ascending=False)\n\nfig, axes = plt.subplots(nrows=1, ncols=1, figsize=(20,10))\ncolor = plt.cm.autumn(np.linspace(0, 1, len(df)))\n\npie_chart(df,\n          'Product Types', \n          'Product Group',\n          axes, \n          color,  \n          \"Product Types per each Product Group\")  ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-23T10:28:57.230046Z","iopub.execute_input":"2022-02-23T10:28:57.230262Z","iopub.status.idle":"2022-02-23T10:28:57.721547Z","shell.execute_reply.started":"2022-02-23T10:28:57.230236Z","shell.execute_reply":"2022-02-23T10:28:57.720607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = articles_df.groupby([\"product_group_name\"])[\"article_id\"].nunique()\ndf = pd.DataFrame({'Product Group': temp.index,'Articles': temp.values})\ndf = df.sort_values(['Articles'], ascending=False)\n\nfig, axes = plt.subplots(nrows=1, ncols=1, figsize=(20,10))\ncolor = plt.cm.summer(np.linspace(0, 1, len(df)))\n\npie_chart(df,\n          'Articles', \n          'Product Group',\n          axes, \n          color,  \n          \"Product Types per each Product Group\")  ","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-23T10:28:57.722874Z","iopub.execute_input":"2022-02-23T10:28:57.723249Z","iopub.status.idle":"2022-02-23T10:28:58.160685Z","shell.execute_reply.started":"2022-02-23T10:28:57.723218Z","shell.execute_reply":"2022-02-23T10:28:58.159834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = articles_df.groupby([\"index_group_name\"])[\"article_id\"].nunique()\ndf = pd.DataFrame({'Index Group Name': temp.index,'Articles': temp.values})\ndf = df.sort_values(['Articles'], ascending=False)\n\nfig, axes = plt.subplots(nrows=1, ncols=1, figsize=(12,6))\ncolor = plt.cm.spring(np.linspace(0, 1, len(df)))\n\nbar_plot(df,\n         'Index Group Name',\n         'Articles',\n         axes, \n         color, \n         \"Number of Articles per each Index Group Name\")\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-23T10:28:58.161964Z","iopub.execute_input":"2022-02-23T10:28:58.162205Z","iopub.status.idle":"2022-02-23T10:28:58.40784Z","shell.execute_reply.started":"2022-02-23T10:28:58.162175Z","shell.execute_reply":"2022-02-23T10:28:58.40687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = articles_df.groupby([\"product_type_name\"])[\"article_id\"].nunique()\ndf = pd.DataFrame({'Product Type': temp.index,'Articles': temp.values})\ntotal_types = len(df['Product Type'].unique())\ndf = df.sort_values(['Articles'], ascending=False)[0:30]\n\nfig, axes = plt.subplots(nrows=1, ncols=1, figsize=(10,6))\ncolor = plt.cm.cool(np.linspace(0, 1, len(df)))\n\nbar_plot(df,\n         'Product Type',\n         'Articles',\n         axes, \n         color, \n         \"Number of Articles per each Product Type (top 30)\")\n","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-23T10:28:58.409308Z","iopub.execute_input":"2022-02-23T10:28:58.409688Z","iopub.status.idle":"2022-02-23T10:28:58.830565Z","shell.execute_reply.started":"2022-02-23T10:28:58.409642Z","shell.execute_reply":"2022-02-23T10:28:58.829958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = articles_df.groupby([\"department_name\"])[\"article_id\"].nunique()\ndf = pd.DataFrame({'Department Name': temp.index,'Articles': temp.values})\ntotal_depts = len(df['Department Name'].unique())\ndf = df.sort_values(['Articles'], ascending=False).head(20)\n\n\nfig, axes = plt.subplots(nrows=1, ncols=1, figsize=(22,10))\ncolor = plt.cm.pink(np.linspace(0, 1, len(df)))\n\npie_chart(df,\n          'Articles', \n          'Department Name',\n          axes, \n          color,  \n          \"Number of Articles per each Department (top 20)\")  ","metadata":{"execution":{"iopub.status.busy":"2022-02-23T10:40:52.38233Z","iopub.execute_input":"2022-02-23T10:40:52.382669Z","iopub.status.idle":"2022-02-23T10:40:52.832037Z","shell.execute_reply.started":"2022-02-23T10:40:52.382629Z","shell.execute_reply":"2022-02-23T10:40:52.831067Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = articles_df.groupby([\"section_name\"])[\"article_id\"].nunique()\ndf = pd.DataFrame({'Section Name': temp.index,'Articles': temp.values})\ndf = df.sort_values(['Articles'], ascending=False)\n\n\nfig, axes = plt.subplots(nrows=1, ncols=1, figsize=(22,10))\ncolor = plt.cm.PRGn(np.linspace(0, 1, len(df)))\n\npie_chart(df.head(15),\n          'Articles', \n          'Section Name',\n          axes, \n          color,  \n          \"Number of Articles per each Section Name (top 15)\")  \n","metadata":{"execution":{"iopub.status.busy":"2022-02-23T10:41:08.357864Z","iopub.execute_input":"2022-02-23T10:41:08.358184Z","iopub.status.idle":"2022-02-23T10:41:08.767994Z","shell.execute_reply.started":"2022-02-23T10:41:08.358152Z","shell.execute_reply":"2022-02-23T10:41:08.766995Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = articles_df.groupby([\"graphical_appearance_name\"])[\"article_id\"].nunique()\ndf = pd.DataFrame({'Graphical Appearance Name': temp.index,'Articles': temp.values})\ndf = df.sort_values(['Articles'], ascending=False).head(50)\n\nfig, axes = plt.subplots(nrows=1, ncols=1, figsize=(22,10))\ncolor = plt.cm.PiYG(np.linspace(0, 1, len(df)))\n\npie_chart(df.head(15),\n          'Articles', \n          'Graphical Appearance Name',\n          axes, \n          color,  \n          \"Number of Articles per each Graphical Appearance Name (top 15)\")  ","metadata":{"execution":{"iopub.status.busy":"2022-02-23T10:41:06.868058Z","iopub.execute_input":"2022-02-23T10:41:06.868374Z","iopub.status.idle":"2022-02-23T10:41:07.245702Z","shell.execute_reply.started":"2022-02-23T10:41:06.868338Z","shell.execute_reply":"2022-02-23T10:41:07.244723Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"stopwords = set(STOPWORDS)\n\ndef show_wordcloud(data, title = None):\n    wordcloud = WordCloud(background_color='#edf7ee',\n                          stopwords=stopwords,\n                          max_words=400,\n                          max_font_size=40, \n                          scale=5,\n                          colormap=\"spring\",\n                          random_state=1).generate(str(data))\n\n    \n    fig = plt.figure(1, figsize=(10,10))\n    plt.axis('off')\n    if title: \n        fig.suptitle(title, fontsize=14)\n        fig.subplots_adjust(top=2.3)\n\n    plt.imshow(wordcloud)\n    plt.show()\n    \n    \n    \nshow_wordcloud(articles_df[\"prod_name\"], \"Wordcloud from product name\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-23T10:28:59.059999Z","iopub.execute_input":"2022-02-23T10:28:59.06031Z","iopub.status.idle":"2022-02-23T10:28:59.730184Z","shell.execute_reply.started":"2022-02-23T10:28:59.060283Z","shell.execute_reply":"2022-02-23T10:28:59.729288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5. Data visualisations. Customers Data 📊","metadata":{}},{"cell_type":"code","source":"temp = customers_df.groupby([\"age\"])[\"customer_id\"].count()\ndf = pd.DataFrame({'Age': temp.index,'Customers': temp.values})\ndf = df.sort_values(['Age'], ascending=False)\n\n\nfig, axes = plt.subplots(nrows=1, ncols=1, figsize=(10,6))\ncolor = plt.cm.cool(np.linspace(0, 1, len(df)))\n\nbar_plot(df,\n         'Age',\n         'Customers',\n         axes, \n         color, \n         \"Number of Customers per each Age\")","metadata":{"execution":{"iopub.status.busy":"2022-02-23T10:55:09.68515Z","iopub.execute_input":"2022-02-23T10:55:09.686974Z","iopub.status.idle":"2022-02-23T10:55:10.334682Z","shell.execute_reply.started":"2022-02-23T10:55:09.686913Z","shell.execute_reply":"2022-02-23T10:55:10.333758Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = customers_df.groupby([\"fashion_news_frequency\"])[\"customer_id\"].count()\ndf = pd.DataFrame({'Fashion News Frequency': temp.index,'Customers': temp.values})\ndf = df.sort_values(['Customers'], ascending=False)\n\nfig, axes = plt.subplots(nrows=1, ncols=1, figsize=(4,6))\ncolor = plt.cm.seismic(np.linspace(0, 1, len(df)))\n\nbar_plot(df,\n         'Fashion News Frequency',\n         'Customers',\n         axes, \n         color, \n         \"Number of Customers per each Fashion News Frequency\")","metadata":{"execution":{"iopub.status.busy":"2022-02-23T10:57:02.997151Z","iopub.execute_input":"2022-02-23T10:57:02.997436Z","iopub.status.idle":"2022-02-23T10:57:03.532556Z","shell.execute_reply.started":"2022-02-23T10:57:02.997406Z","shell.execute_reply":"2022-02-23T10:57:03.531613Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"temp = customers_df.groupby([\"club_member_status\"])[\"customer_id\"].count()\ndf = pd.DataFrame({'Club Member Status': temp.index,'Customers': temp.values})\ndf = df.sort_values(['Customers'], ascending=False)\n\nfig, axes = plt.subplots(nrows=1, ncols=1, figsize=(22,10))\ncolor = plt.cm.summer(np.linspace(0, 1, len(df)))\n\npie_chart(df.head(15),\n          'Customers', \n          'Club Member Status',\n          axes, \n          color,  \n          \"Number of Customers per each Club Member Status\") ","metadata":{"execution":{"iopub.status.busy":"2022-02-23T11:00:10.499587Z","iopub.execute_input":"2022-02-23T11:00:10.500342Z","iopub.status.idle":"2022-02-23T11:00:11.069015Z","shell.execute_reply.started":"2022-02-23T11:00:10.500291Z","shell.execute_reply":"2022-02-23T11:00:11.067976Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 6. Data visualisations. Transactions Data 📊","metadata":{}},{"cell_type":"code","source":"df = transactions_train_df.sample(100_000)\nfig, ax = plt.subplots(1, 1, figsize=(14, 7))\nsns.kdeplot(np.log(df.loc[df[\"sales_channel_id\"]==1].price.value_counts()),\n           color=\"red\")\nsns.kdeplot(np.log(df.loc[df[\"sales_channel_id\"]==2].price.value_counts()),\n           color=\"blue\")\n\nax.legend(labels=['Sales channel 1', \n                  'Sales channel 2'])\n\nplt.title(\"Logaritmic distribution of price frequency \\\nin transactions, grouped per sales channel (100k sample)\")\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-23T11:03:17.158716Z","iopub.execute_input":"2022-02-23T11:03:17.158996Z","iopub.status.idle":"2022-02-23T11:03:19.285413Z","shell.execute_reply.started":"2022-02-23T11:03:17.158966Z","shell.execute_reply":"2022-02-23T11:03:19.284499Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = transactions_train_df.sample(100_000).groupby([\"t_dat\"])[\"article_id\"].count().reset_index()\ndf[\"t_dat\"] = df[\"t_dat\"].apply(lambda x: datetime.strptime(x, '%Y-%m-%d'))\ndf.columns = [\"Date\", \"Transactions\"]\n\nfig, ax = plt.subplots(1, 1, figsize=(16,6))\nplt.plot(df[\"Date\"], df[\"Transactions\"], color=\"red\")\nplt.xlabel(\"Date\")\nplt.ylabel(\"Transactions\")\nplt.title(f\"Transactions per day (100k sample)\")\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-23T11:04:20.785701Z","iopub.execute_input":"2022-02-23T11:04:20.786026Z","iopub.status.idle":"2022-02-23T11:04:22.799473Z","shell.execute_reply.started":"2022-02-23T11:04:20.785995Z","shell.execute_reply":"2022-02-23T11:04:22.798845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = transactions_train_df.sample(100_000).groupby([\"t_dat\", \"sales_channel_id\"])[\"article_id\"].count().reset_index()\ndf[\"t_dat\"] = df[\"t_dat\"].apply(lambda x: datetime.strptime(x, '%Y-%m-%d'))\n\ndf.columns = [\"Date\", \"Sales Channel Id\", \"Transactions\"]\n\nfig, ax = plt.subplots(1, 1, figsize=(16,6))\ng1 = ax.plot(df.loc[df[\"Sales Channel Id\"]==1, \"Date\"], \n             df.loc[df[\"Sales Channel Id\"]==1, \n                    \"Transactions\"], \n             label=\"Sales Channel 1\", \n             color=\"Blue\")\n\ng2 = ax.plot(df.loc[df[\"Sales Channel Id\"]==2, \"Date\"], \n             df.loc[df[\"Sales Channel Id\"]==2, \n                    \"Transactions\"], \n             label=\"Sales Channel 2\", \n             color=\"Red\")\n\nplt.xlabel(\"Date\")\nplt.ylabel(\"Transactions\")\nax.legend()\nplt.title(f\"Transactions per day, grouped by Sales Channel (100k sample)\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-23T11:05:42.460905Z","iopub.execute_input":"2022-02-23T11:05:42.461221Z","iopub.status.idle":"2022-02-23T11:05:44.454386Z","shell.execute_reply.started":"2022-02-23T11:05:42.461183Z","shell.execute_reply":"2022-02-23T11:05:44.453617Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = transactions_train_df.groupby([\"t_dat\", \"sales_channel_id\"])[\"article_id\"].nunique().reset_index()\ndf[\"t_dat\"] = df[\"t_dat\"].apply(lambda x: datetime.strptime(x, '%Y-%m-%d'))\ndf.columns = [\"Date\", \"Sales Channel Id\", \"Unique Articles\"]\n\nfig, ax = plt.subplots(1, 1, figsize=(16,6))\ng1 = ax.plot(df.loc[df[\"Sales Channel Id\"]==1, \n                    \"Date\"], \n             df.loc[df[\"Sales Channel Id\"]==1, \n                    \"Unique Articles\"], \n             label=\"Sales Channel 1\", \n             color=\"Blue\")\n\ng2 = ax.plot(df.loc[df[\"Sales Channel Id\"]==2, \n                    \"Date\"], \n             df.loc[df[\"Sales Channel Id\"]==2, \n                    \"Unique Articles\"], \n             label=\"Sales Channel 2\", \n             color=\"Orange\")\n\nplt.xlabel(\"Date\")\nplt.ylabel(\"Unique Articles / Day\")\nax.legend()\nplt.title(f\"Unique articles per day, grouped by Sales Channel\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-23T11:06:21.752124Z","iopub.execute_input":"2022-02-23T11:06:21.752432Z","iopub.status.idle":"2022-02-23T11:06:37.48158Z","shell.execute_reply.started":"2022-02-23T11:06:21.75239Z","shell.execute_reply":"2022-02-23T11:06:37.478898Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 7. Data visualisations. Image Data 📊","metadata":{}},{"cell_type":"code","source":"image_name_df = pd.DataFrame(images_names, columns = [\"image_name\"])\nimage_name_df[\"article_id\"] = image_name_df[\"image_name\"].apply(lambda x: int(x[1:]))\nimage_name_df.head().style.set_properties(**{'background-color': 'rgba(184,230,194,.5)'})","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-23T11:28:12.541769Z","iopub.execute_input":"2022-02-23T11:28:12.542139Z","iopub.status.idle":"2022-02-23T11:28:12.651707Z","shell.execute_reply.started":"2022-02-23T11:28:12.542096Z","shell.execute_reply":"2022-02-23T11:28:12.650685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_article_df = articles_df[[\"article_id\", \n                                \"product_code\", \n                                \"product_group_name\", \n                                \"product_type_name\"]].merge(image_name_df, \n                                                            on=[\"article_id\"], \n                                                            how=\"left\")\nimage_article_df.head().style.set_properties(**{'background-color': 'rgba(184,230,194,.5)'})","metadata":{"execution":{"iopub.status.busy":"2022-02-23T11:28:39.131408Z","iopub.execute_input":"2022-02-23T11:28:39.131764Z","iopub.status.idle":"2022-02-23T11:28:39.209477Z","shell.execute_reply.started":"2022-02-23T11:28:39.131727Z","shell.execute_reply":"2022-02-23T11:28:39.208614Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# products without images\narticle_no_image_df = image_article_df.loc[image_article_df.image_name.isna()]\narticle_no_image_df.head().style.set_properties(**{'background-color': 'rgba(184,230,194,.5)'})","metadata":{"execution":{"iopub.status.busy":"2022-02-23T11:29:08.970697Z","iopub.execute_input":"2022-02-23T11:29:08.970997Z","iopub.status.idle":"2022-02-23T11:29:09.008309Z","shell.execute_reply.started":"2022-02-23T11:29:08.970961Z","shell.execute_reply":"2022-02-23T11:29:09.007412Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's plot some image data.","metadata":{}},{"cell_type":"code","source":"def plot_image_samples(image_article_df, product_group_name, cols=1, rows=-1):\n    image_path = \"../input/h-and-m-personalized-fashion-recommendations/images/\"\n    _df = image_article_df.loc[image_article_df.product_group_name==product_group_name]\n    article_ids = _df.article_id.values[0:cols*rows]\n    plt.figure(figsize=(2 + 3 * cols, 2 + 4 * rows))\n    for i in range(cols * rows):\n        article_id = (\"0\" + str(article_ids[i]))[-10:]\n        plt.subplot(rows, cols, i + 1)\n        plt.axis('off')\n        plt.title(f\"{product_group_name} {article_id[:3]}\\n{article_id}.jpg\")\n        image = Image.open(f\"{image_path}{article_id[:3]}/{article_id}.jpg\")\n        plt.imshow(image)\n        \nplot_image_samples(image_article_df, \"Garment Lower body\", 5, 1)\nplot_image_samples(image_article_df, \"Accessories\", 5, 1)\nplot_image_samples(image_article_df, \"Swimwear\", 5, 1)\nplot_image_samples(image_article_df, \"Bags\", 5, 1)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-23T11:34:22.287012Z","iopub.execute_input":"2022-02-23T11:34:22.287299Z","iopub.status.idle":"2022-02-23T11:34:30.245958Z","shell.execute_reply.started":"2022-02-23T11:34:22.28727Z","shell.execute_reply":"2022-02-23T11:34:30.24502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 8. Predictions ☂️\nFor this initial submission, I dediced to follow logic, that was described in article, that is give in anknowledgements:\n\n* if there are articles for a certain client, pick the most recent buys;\n* if there are not articles for a certain client, just pick the most frequently buyed articles.","metadata":{}},{"cell_type":"code","source":"transactions_train_df = transactions_train_df.sort_values([\"customer_id\", \n                                                           \"t_dat\"], \n                                                          ascending=False)\ntransactions_train_df.head().style.set_properties(**{'background-color': 'rgba(184,230,194,.5)'})","metadata":{"execution":{"iopub.status.busy":"2022-02-23T11:38:54.481519Z","iopub.execute_input":"2022-02-23T11:38:54.482237Z","iopub.status.idle":"2022-02-23T11:39:24.153378Z","shell.execute_reply.started":"2022-02-23T11:38:54.482193Z","shell.execute_reply":"2022-02-23T11:39:24.15249Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's capture first what are the most frequent recently bought articles.","metadata":{}},{"cell_type":"code","source":"last_date = transactions_train_df.t_dat.max()\nprint(last_date)\nprint(transactions_train_df.loc[transactions_train_df.t_dat==last_date].shape)\nprint()\n\nmost_frequent_articles = list(transactions_train_df.loc[transactions_train_df.t_dat==last_date].article_id.value_counts()[0:12].index)\nart_list = []\nfor art in most_frequent_articles:\n    art = \"0\"+str(art)\n    art_list.append(art)\nart_str = \" \".join(art_list)\nprint(\"Frequent articles bought recently:\", art_str, end=\"\\n\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-02-23T11:46:15.591334Z","iopub.execute_input":"2022-02-23T11:46:15.591566Z","iopub.status.idle":"2022-02-23T11:46:30.581593Z","shell.execute_reply.started":"2022-02-23T11:46:15.591538Z","shell.execute_reply":"2022-02-23T11:46:30.580653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"agg_df = transactions_train_df.groupby([\"customer_id\"])[\"article_id\"].agg(lambda x: str(x.values[0:12])[1:-1]).reset_index()","metadata":{"execution":{"iopub.status.busy":"2022-02-23T11:46:59.125178Z","iopub.execute_input":"2022-02-23T11:46:59.125976Z","iopub.status.idle":"2022-02-23T11:48:45.714004Z","shell.execute_reply.started":"2022-02-23T11:46:59.125931Z","shell.execute_reply":"2022-02-23T11:48:45.712624Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def padding_articles(x):\n    if x:\n        xl = x.split()\n        x = []\n        for xi in xl:\n            x.append(\"0\"+xi)\n        dimm_x = len(x)\n        if dimm_x < 12:\n            x.extend(art_list[:12-dimm_x])\n        return(\" \".join(x))","metadata":{"execution":{"iopub.status.busy":"2022-02-23T11:48:45.71636Z","iopub.execute_input":"2022-02-23T11:48:45.717355Z","iopub.status.idle":"2022-02-23T11:48:45.724147Z","shell.execute_reply.started":"2022-02-23T11:48:45.717297Z","shell.execute_reply":"2022-02-23T11:48:45.723038Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"agg_df[\"article_id\"] = agg_df[\"article_id\"].apply(lambda x: padding_articles(x))\nprint(\"Aggregated transaction history: \", agg_df.customer_id.nunique())\nprint(\"Submission sample: \", sample_submission_df.customer_id.nunique())","metadata":{"execution":{"iopub.status.busy":"2022-02-23T11:48:50.327372Z","iopub.execute_input":"2022-02-23T11:48:50.327696Z","iopub.status.idle":"2022-02-23T11:48:57.947961Z","shell.execute_reply.started":"2022-02-23T11:48:50.32765Z","shell.execute_reply":"2022-02-23T11:48:57.946755Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We'll replace the values in sample submission with the existent in aggregated transactions data and just let the default one otherwise.","metadata":{}},{"cell_type":"code","source":"sample_submission_df.head().style.set_properties(**{'background-color': 'rgba(184,230,194,.5)'})","metadata":{"execution":{"iopub.status.busy":"2022-02-23T11:48:57.94974Z","iopub.execute_input":"2022-02-23T11:48:57.950095Z","iopub.status.idle":"2022-02-23T11:48:57.964236Z","shell.execute_reply.started":"2022-02-23T11:48:57.950049Z","shell.execute_reply":"2022-02-23T11:48:57.963187Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For the customers with missing articles, we simply replace with most frequent buyed articles in most recent days.","metadata":{}},{"cell_type":"code","source":"submission_df = agg_df.merge(sample_submission_df[[\"customer_id\"]], how=\"right\")\nsubmission_df.columns = [\"customer_id\", \"prediction\"]\nprint(submission_df.shape)\nsubmission_df.head().style.set_properties(**{'background-color': 'rgba(184,230,194,.5)'})","metadata":{"execution":{"iopub.status.busy":"2022-02-23T11:50:55.081462Z","iopub.execute_input":"2022-02-23T11:50:55.081786Z","iopub.status.idle":"2022-02-23T11:50:57.597291Z","shell.execute_reply.started":"2022-02-23T11:50:55.081754Z","shell.execute_reply":"2022-02-23T11:50:57.596382Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Rows with missing data in submission: \", submission_df.loc[submission_df.prediction.isna()].shape[0])","metadata":{"execution":{"iopub.status.busy":"2022-02-23T11:51:17.153778Z","iopub.execute_input":"2022-02-23T11:51:17.154411Z","iopub.status.idle":"2022-02-23T11:51:17.548761Z","shell.execute_reply.started":"2022-02-23T11:51:17.154369Z","shell.execute_reply":"2022-02-23T11:51:17.548196Z"},"_kg_hide-output":false,"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We replace the missing data with the most frequently bought articles, from recent days. We calculated it before.","metadata":{}},{"cell_type":"code","source":"submission_df.loc[submission_df.prediction.isna(), [\"prediction\"]] = art_str\nprint(\"Rows with missing data in submission: \", submission_df.loc[submission_df.prediction.isna()].shape[0])\nsubmission_df.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-02-23T11:53:45.784174Z","iopub.execute_input":"2022-02-23T11:53:45.785027Z","iopub.status.idle":"2022-02-23T11:54:00.178869Z","shell.execute_reply.started":"2022-02-23T11:53:45.784982Z","shell.execute_reply":"2022-02-23T11:54:00.178158Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 9. Conclusion 💖\nThank you for reading my new article! **If you liked it, please, make an upvote 💖**\n\nMy other articles:\n* [House Prices Regression sklearn](https://www.kaggle.com/maricinnamon/house-prices-regression-sklearn)\n* [Harry Potter Movies Dataset | Starter Notebook](https://www.kaggle.com/maricinnamon/harry-potter-movies-dataset-starter-notebook)\n* [Automobile Customer Clustering (K-means & PCA)](https://www.kaggle.com/maricinnamon/automobile-customer-clustering-k-means-pca)","metadata":{}}]}