{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1><center>In-Depth EDA [All Files]+Competition Understanding</center></h1>\n<h3><center>H&M Personalized Fashion Recommendations</center></h3>\n\n<center><img src = \"http://www.hm.com/entrance/assets/bundle/img/HM-Share-Image.jpg\" width = \"750\" height = \"500\"/></center>                                                                          ","metadata":{}},{"cell_type":"markdown","source":"<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='background:orange; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>Contents</center></h2>","metadata":{}},{"cell_type":"markdown","source":"> | S.No       |                   Heading                |\n> | :------------- | :-------------------:                |         \n> |  01 |  [**Competition Overview**](#competition-overview)  |                   \n> |  02 |  [**Libraries**](#libraries)                        |  \n> |  03 |  [**Global Config**](#global-config)                |\n> |  04 |  [**Weights and Biases**](#weights-and-biases)      |\n> |  05 |  [**Load Datasets**](#load-datasets)                |\n> |  06 |  [**Articles EDA**](#articles-eda)  |","metadata":{}},{"cell_type":"markdown","source":"<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h3 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='background:maroon; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>If you find this notebook useful, do give me an upvote, it helps to keep up my motivation. This notebook will be updated frequently so keep checking for furthur developments.</center></h3>","metadata":{}},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"<a id=\"competition-overview\"></a>\n<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='background:orange; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>Competition Overview</center></h2>","metadata":{}},{"cell_type":"markdown","source":"## **<span style=\"color:orange;\">Description</span>**\n  \nIn this competition, H&M Group invites you to develop product recommendations based on data from previous transactions, as well as from customer and product meta data. The available meta data spans from simple data, such as garment type and customer age, to text data from product descriptions, to image data from garment images.\n  \nThere are no preconceptions on what information that may be useful – that is for you to find out. If you want to investigate a categorical data type algorithm, or dive into NLP and image processing deep learning, that is up to you.\n\n---\n\n## **<span style=\"color:orange;\">Evaluation Metric</span>**\n\nSubmissions are evaluated according to the Mean Average Precision @ 12 (MAP@12).\n\n**Notes:**\n\n- You will be making purchase predictions for all customer_id values provided, regardless of whether these customers made purchases in the training data.\n- Customer that did not make any purchase during test period are excluded from the scoring.\n- There is never a penalty for using the full 12 predictions for a customer that ordered fewer than 12 items; thus, it's advantageous to make 12 predictions for each customer.\n\n---","metadata":{}},{"cell_type":"markdown","source":"<a id=\"libraries\"></a>\n<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='background:orange; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>Libraries</center></h2>","metadata":{}},{"cell_type":"code","source":"# Necessities\nimport os\nimport random\n\nimport numpy as np\nimport pandas as pd\n\n#Text Processing\nimport re\nimport nltk\nnltk.download('popular')\n\n# Aesthetics\nfrom termcolor import colored\n\n# Data Visualization\nimport plotly.express as px\n\nimport matplotlib.pyplot as plt\n%matplotlib inline","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"global-config\"></a>\n<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='background:orange; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>Global Config</center></h2>","metadata":{}},{"cell_type":"code","source":"class config:\n    SEED=42\n    DIRECTORY_PATH = \"../input/h-and-m-personalized-fashion-recommendations\"\n    IMAGES_FILE_PATH = os.path.join(DIRECTORY_PATH, 'images')\n    ARTICLES_FILE_PATH = os.path.join(DIRECTORY_PATH, 'articles.csv')\n    CUSTOMERS_FILE_PATH = os.path.join(DIRECTORY_PATH, 'customers.csv')\n    SAMPLE_FILE_PATH = os.path.join(DIRECTORY_PATH, 'sample_submission.csv')\n    TRANSACTIONS_FILE_PATH = os.path.join(DIRECTORY_PATH, 'transactions_train.csv')","metadata":{"execution":{"iopub.status.busy":"2022-02-08T05:08:38.298049Z","iopub.execute_input":"2022-02-08T05:08:38.298349Z","iopub.status.idle":"2022-02-08T05:08:38.304427Z","shell.execute_reply.started":"2022-02-08T05:08:38.298317Z","shell.execute_reply":"2022-02-08T05:08:38.303505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def set_seed(seed=config.SEED):\n    random.seed(seed)\n    os.environ[\"PYTHONHASHSEED\"] = str(seed)\n    np.random.seed(seed)\n    \nset_seed()","metadata":{"execution":{"iopub.status.busy":"2022-02-08T05:04:40.067191Z","iopub.execute_input":"2022-02-08T05:04:40.067846Z","iopub.status.idle":"2022-02-08T05:04:40.072904Z","shell.execute_reply.started":"2022-02-08T05:04:40.067792Z","shell.execute_reply":"2022-02-08T05:04:40.071985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"<a id=\"weights-and-biases\"></a>\n<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='background:orange; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>Weights and Biases</center></h2>","metadata":{}},{"cell_type":"markdown","source":"<center><img src = \"https://i.imgur.com/1sm6x8P.png\" width = \"750\" height = \"500\"/></center>  ","metadata":{}},{"cell_type":"markdown","source":"**Weights & Biases** is the machine learning platform for developers to build better models faster.\n\nYou can use W&B's lightweight, interoperable tools to\n\n- quickly track experiments,\n- version and iterate on datasets,\n- evaluate model performance,\n- reproduce models,\n- visualize results and spot regressions,\n- and share findings with colleagues.\n  \nSet up W&B in 5 minutes, then quickly iterate on your machine learning pipeline with the confidence that your datasets and models are tracked and versioned in a reliable system of record.\n\nIn this notebook I will use Weights and Biases's amazing features to perform wonderful visualizations and logging seamlessly.\n\n---","metadata":{}},{"cell_type":"markdown","source":"<a id=\"load-datasets\"></a>\n<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='background:orange; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>Load Datasets</center></h2>","metadata":{}},{"cell_type":"code","source":"articles = pd.read_csv(config.ARTICLES_FILE_PATH)\ncustomers = pd.read_csv(config.CUSTOMERS_FILE_PATH)\nsample = pd.read_csv(config.SAMPLE_FILE_PATH)\n# transactions = pd.read_csv(config.TRANSACTIONS_FILE_PATH)","metadata":{"execution":{"iopub.status.busy":"2022-02-08T05:09:47.830872Z","iopub.execute_input":"2022-02-08T05:09:47.831123Z","iopub.status.idle":"2022-02-08T05:09:54.630291Z","shell.execute_reply.started":"2022-02-08T05:09:47.831097Z","shell.execute_reply":"2022-02-08T05:09:54.629412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"articles-eda\"></a>\n<div class=\"list-group\" id=\"list-tab\" role=\"tablist\">\n<h2 class=\"list-group-item list-group-item-action active\" data-toggle=\"list\" style='background:orange; border:0; color:white' role=\"tab\" aria-controls=\"home\"><center>Articles EDA</center></h2>","metadata":{}},{"cell_type":"markdown","source":"## **<span style=\"color:orange;\">Basic Tabular Analysis</span>**","metadata":{}},{"cell_type":"markdown","source":"### Head","metadata":{}},{"cell_type":"code","source":"articles.head()","metadata":{"execution":{"iopub.status.busy":"2022-02-08T05:11:56.050269Z","iopub.execute_input":"2022-02-08T05:11:56.050548Z","iopub.status.idle":"2022-02-08T05:11:56.100169Z","shell.execute_reply.started":"2022-02-08T05:11:56.050515Z","shell.execute_reply":"2022-02-08T05:11:56.099368Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Dataset Size","metadata":{}},{"cell_type":"code","source":"print(\"Shape of articles.csv: \", colored(articles.shape, 'yellow'))","metadata":{"execution":{"iopub.status.busy":"2022-02-08T05:13:17.608572Z","iopub.execute_input":"2022-02-08T05:13:17.608872Z","iopub.status.idle":"2022-02-08T05:13:17.615156Z","shell.execute_reply.started":"2022-02-08T05:13:17.608844Z","shell.execute_reply":"2022-02-08T05:13:17.613922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thus, the file has 105,542 rows and 25 columns.","metadata":{}},{"cell_type":"markdown","source":"### Info","metadata":{}},{"cell_type":"code","source":"articles.info()","metadata":{"execution":{"iopub.status.busy":"2022-02-08T05:14:23.563901Z","iopub.execute_input":"2022-02-08T05:14:23.564163Z","iopub.status.idle":"2022-02-08T05:14:23.660288Z","shell.execute_reply.started":"2022-02-08T05:14:23.564135Z","shell.execute_reply":"2022-02-08T05:14:23.659290Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This tells us that there are no null values in the **articles.csv** file.","metadata":{}},{"cell_type":"markdown","source":"### Column wise Unique Values","metadata":{}},{"cell_type":"code","source":"for col in articles.columns:\n    print(col + \":\" + colored(str(len(articles[col].unique())), 'yellow'))","metadata":{"execution":{"iopub.status.busy":"2022-02-08T05:15:56.230830Z","iopub.execute_input":"2022-02-08T05:15:56.231414Z","iopub.status.idle":"2022-02-08T05:15:56.373157Z","shell.execute_reply.started":"2022-02-08T05:15:56.231374Z","shell.execute_reply":"2022-02-08T05:15:56.372421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## **<span style=\"color:orange;\">Unique Value Counts</span>**","metadata":{}},{"cell_type":"markdown","source":"### Product Name","metadata":{}},{"cell_type":"code","source":"articles.prod_name.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-02-08T05:33:28.894274Z","iopub.execute_input":"2022-02-08T05:33:28.895856Z","iopub.status.idle":"2022-02-08T05:33:28.937282Z","shell.execute_reply.started":"2022-02-08T05:33:28.895807Z","shell.execute_reply":"2022-02-08T05:33:28.936658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Product Type Name","metadata":{}},{"cell_type":"code","source":"articles.product_type_name.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-02-08T05:36:58.796387Z","iopub.execute_input":"2022-02-08T05:36:58.796984Z","iopub.status.idle":"2022-02-08T05:36:58.811615Z","shell.execute_reply.started":"2022-02-08T05:36:58.796947Z","shell.execute_reply":"2022-02-08T05:36:58.810771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Product Group Name","metadata":{}},{"cell_type":"code","source":"articles.product_group_name.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-02-08T05:37:17.461851Z","iopub.execute_input":"2022-02-08T05:37:17.462142Z","iopub.status.idle":"2022-02-08T05:37:17.477541Z","shell.execute_reply.started":"2022-02-08T05:37:17.462111Z","shell.execute_reply":"2022-02-08T05:37:17.476481Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Graphical Appearance Name","metadata":{}},{"cell_type":"code","source":"articles.graphical_appearance_name\t.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-02-08T05:38:24.562650Z","iopub.execute_input":"2022-02-08T05:38:24.562953Z","iopub.status.idle":"2022-02-08T05:38:24.577572Z","shell.execute_reply.started":"2022-02-08T05:38:24.562924Z","shell.execute_reply":"2022-02-08T05:38:24.576734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Colour Group Name","metadata":{}},{"cell_type":"code","source":"articles.colour_group_name.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-02-08T05:38:43.229883Z","iopub.execute_input":"2022-02-08T05:38:43.230163Z","iopub.status.idle":"2022-02-08T05:38:43.245585Z","shell.execute_reply.started":"2022-02-08T05:38:43.230127Z","shell.execute_reply":"2022-02-08T05:38:43.244642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}},{"cell_type":"markdown","source":"## **<span style=\"color:orange;\">Detail Description Distributions</span>**","metadata":{}},{"cell_type":"code","source":"articles.detail_desc","metadata":{"execution":{"iopub.status.busy":"2022-02-08T05:42:24.931638Z","iopub.execute_input":"2022-02-08T05:42:24.931986Z","iopub.status.idle":"2022-02-08T05:42:24.940278Z","shell.execute_reply.started":"2022-02-08T05:42:24.931953Z","shell.execute_reply":"2022-02-08T05:42:24.939461Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Preprocessing the text","metadata":{}},{"cell_type":"code","source":"def preprocess_text(text, flg_stemm=False, flg_lemm=True):\n\n    lst_stopwords = nltk.corpus.stopwords.words(\"english\")\n    \n    ## clean (convert to lowercase and remove punctuations and characters and then strip)\n    text = re.sub(r'[^\\w\\s]', '', str(text).lower().strip())\n            \n    ## Tokenize (convert from string to list)\n    lst_text = text.split()\n    ## remove Stopwords\n    if lst_stopwords is not None:\n        lst_text = [word for word in lst_text if word not in \n                    lst_stopwords]\n                \n    ## Stemming (remove -ing, -ly, ...)\n    if flg_stemm == True:\n        ps = nltk.stem.porter.PorterStemmer()\n        lst_text = [ps.stem(word) for word in lst_text]\n                \n    ## Lemmatisation (convert the word into root word)\n    if flg_lemm == True:\n        lem = nltk.stem.wordnet.WordNetLemmatizer()    \n        lst_text = [lem.lemmatize(word) for word in lst_text]\n            \n    ## back to string from list\n    text = \" \".join(lst_text)\n    return text\n","metadata":{"execution":{"iopub.status.busy":"2022-02-08T05:46:50.690809Z","iopub.execute_input":"2022-02-08T05:46:50.691379Z","iopub.status.idle":"2022-02-08T05:46:50.699076Z","shell.execute_reply.started":"2022-02-08T05:46:50.691341Z","shell.execute_reply":"2022-02-08T05:46:50.698226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Clean the text","metadata":{}},{"cell_type":"code","source":"# Clean Detail Description\narticles[\"clean_detail_desc\"] = articles[\"detail_desc\"].apply(\n    lambda x: preprocess_text(x, flg_stemm=False, flg_lemm=True, )\n)","metadata":{"execution":{"iopub.status.busy":"2022-02-08T05:48:37.234985Z","iopub.execute_input":"2022-02-08T05:48:37.235245Z","iopub.status.idle":"2022-02-08T05:49:02.460775Z","shell.execute_reply.started":"2022-02-08T05:48:37.235216Z","shell.execute_reply":"2022-02-08T05:49:02.459935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Generate New Features","metadata":{}},{"cell_type":"code","source":"#Length of Detail Description\narticles['clean_detail_desc_len'] = articles['clean_detail_desc'].apply(lambda x: len(x))\n\n#Word Count\narticles['clean_detail_desc_word_count'] = articles[\"clean_detail_desc\"].apply(lambda x: len(str(x).split(\" \")))\n\n#Character Count\narticles['clean_detail_desc_char_count'] = articles[\"clean_detail_desc\"].apply(lambda x: sum(len(word) for word in str(x).split(\" \")))\n\n#Average Word Length\narticles['clean_detail_desc_avg_word_length'] = articles['clean_detail_desc_char_count'] / articles['clean_detail_desc_word_count']","metadata":{"execution":{"iopub.status.busy":"2022-02-08T05:50:59.084144Z","iopub.execute_input":"2022-02-08T05:50:59.084452Z","iopub.status.idle":"2022-02-08T05:50:59.534522Z","shell.execute_reply.started":"2022-02-08T05:50:59.084421Z","shell.execute_reply":"2022-02-08T05:50:59.533620Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Distribution Plots","metadata":{}},{"cell_type":"code","source":"def plot_distribution(x, title):\n\n    fig = px.histogram(\n    articles, \n    x = x,\n    width = 800,\n    height = 500,\n    title = title\n    )\n    \n    fig.show()","metadata":{"execution":{"iopub.status.busy":"2022-02-08T05:55:43.718243Z","iopub.execute_input":"2022-02-08T05:55:43.718987Z","iopub.status.idle":"2022-02-08T05:55:43.724880Z","shell.execute_reply.started":"2022-02-08T05:55:43.718942Z","shell.execute_reply":"2022-02-08T05:55:43.723880Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_distribution(x = 'clean_detail_desc_len', title = 'Description Length Distribution')","metadata":{"execution":{"iopub.status.busy":"2022-02-08T05:55:55.580498Z","iopub.execute_input":"2022-02-08T05:55:55.581379Z","iopub.status.idle":"2022-02-08T05:55:56.911901Z","shell.execute_reply.started":"2022-02-08T05:55:55.581339Z","shell.execute_reply":"2022-02-08T05:55:56.911065Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_distribution(x = 'clean_detail_desc_word_count', title = 'Word Count Distribution')","metadata":{"execution":{"iopub.status.busy":"2022-02-08T05:59:38.693362Z","iopub.execute_input":"2022-02-08T05:59:38.693647Z","iopub.status.idle":"2022-02-08T05:59:39.012132Z","shell.execute_reply.started":"2022-02-08T05:59:38.693613Z","shell.execute_reply":"2022-02-08T05:59:39.011255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_distribution(x = 'clean_detail_desc_char_count', title = 'Character Count Distribution')","metadata":{"execution":{"iopub.status.busy":"2022-02-08T06:00:10.918914Z","iopub.execute_input":"2022-02-08T06:00:10.919206Z","iopub.status.idle":"2022-02-08T06:00:11.246720Z","shell.execute_reply.started":"2022-02-08T06:00:10.919169Z","shell.execute_reply":"2022-02-08T06:00:11.245615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_distribution(x = 'clean_detail_desc_avg_word_length', title = 'Average Word Length Distribution')","metadata":{"execution":{"iopub.status.busy":"2022-02-08T06:00:30.673923Z","iopub.execute_input":"2022-02-08T06:00:30.674183Z","iopub.status.idle":"2022-02-08T06:00:31.013278Z","shell.execute_reply.started":"2022-02-08T06:00:30.674156Z","shell.execute_reply":"2022-02-08T06:00:31.012134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---","metadata":{}}]}