{"cells": [{"metadata": {"_cell_guid": "a896c938-8be7-4453-90a9-6eefa176c18d", "_uuid": "e5d62013ddf9016dc17f7fd059f915d210981dfa"}, "cell_type": "markdown", "source": ["Hello\n", "This Kernel/Notebook work towards converting the categories_names.csv hierarchy into categories binary hierarchy, the need of a binary hierarchical categories , would make it easier when splitting the data according to the categories hieracrchy **in case you want to approch this problem by using multiple models for the categories in each level.** for example you may create one model to classfiey the class in category level 1, then for each prediction on the level 1 model you choose which model to choose next..."]}, {"metadata": {"_cell_guid": "96ac9e2e-f888-4f06-91c1-6fa41da85c2d", "_uuid": "9091f650bcae6882849b036bfd59cff632b99e98"}, "execution_count": null, "source": ["# This Python 3 environment comes with many helpful analytics libraries installed\n", "# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n", "# For example, here's several helpful packages to load in \n", "\n", "import numpy as np # linear algebra\n", "import pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n", "\n", "# Input data files are available in the \"../input/\" directory.\n", "# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n", "\n", "from subprocess import check_output\n", "print(check_output([\"ls\", \"../input\"]).decode(\"utf8\"))\n", "\n", "# Any results you write to the current directory are saved as output.\n"], "cell_type": "code", "outputs": []}, {"metadata": {"_cell_guid": "8c5b2c19-51cc-44b7-9b28-d2cd5ad94096", "_uuid": "5c7e37bf8683ef07f96c9f8a8a73bab7cd367af5", "collapsed": true}, "execution_count": null, "source": ["categories=pd.read_csv(\"../input/category_names.csv\")"], "cell_type": "code", "outputs": []}, {"metadata": {"_cell_guid": "252137e5-ef5f-44cd-8d72-a849abfb312d", "_uuid": "2d426a991fdad04cf3cafe8406ea4e36525183d4"}, "cell_type": "markdown", "source": ["Let Expolre the name of the columns, types...etc"]}, {"metadata": {"_cell_guid": "23495a5e-dae4-427e-96e7-977ef0b4ce1b", "_uuid": "ecf934f8b4a9c82e5fb12dc7f776004764b199c2"}, "execution_count": null, "source": ["print(categories.info())\n", "print(categories.head(10))"], "cell_type": "code", "outputs": []}, {"metadata": {"_cell_guid": "f7062b3f-da6e-4dbc-a95d-bcefd9aa9853", "_uuid": "0cf169ca47c2d017461cfa739fcf5aa6bbfe139c"}, "cell_type": "markdown", "source": ["Lets check how many categories there are for the level 1 :"]}, {"metadata": {"_cell_guid": "743996ba-cd8e-4cd5-81ae-30edcba0644f", "_uuid": "c29e4c907166191e54cc968c7dbea66f6a7a19b0"}, "execution_count": null, "source": ["cat_lv1=categories[\"category_level1\"].unique()\n", "print(len(cat_lv1))"], "cell_type": "code", "outputs": []}, {"metadata": {"_cell_guid": "f6862f27-ab52-4517-9a60-0c5311ec013e", "_uuid": "f1ec929897e220a38f2f898237c8a2497a59cf50"}, "cell_type": "markdown", "source": ["Important note:  Given that category level 3 is the most lower category level, does each name of the category level 3 corrsponds to a unique category_id ?,\n", "let's check that out:\n"]}, {"metadata": {"_cell_guid": "4df7e56e-8606-4985-9dc6-84be3dd3d321", "_uuid": "0fe701850e2c2fae70520c25a4bb86bed5081531"}, "execution_count": null, "source": ["print(\"Total unique Categories Names (Level 3)\",len(categories.category_level3.unique()))\n", "print(\"Total unique Categories Ids\",len(categories.category_id.unique()))\n", "print(\"Number of total observation/rows\",len(categories))"], "cell_type": "code", "outputs": []}, {"metadata": {"_cell_guid": "c0ccac82-d0b0-4840-9064-536d2959b402", "_uuid": "bd4b585b6c67c637223792b977437bc6b87a7288"}, "cell_type": "markdown", "source": ["As you can observe there are seven duplicated names in category level 3 "]}, {"metadata": {"_cell_guid": "5860a1f6-0aae-464d-85ef-04a363867c8a", "_uuid": "7e7ea0bbcc4c6a46319c7fb11e80ad69941843e4"}, "execution_count": null, "source": ["categories[categories.category_level3.duplicated()]\n"], "cell_type": "code", "outputs": []}, {"metadata": {"_cell_guid": "749c71b7-0d20-40a7-9613-fed84424833f", "_uuid": "3170abf5ed27c9233fffeaff16f24b7e353f5de3", "_kg_hide-output": true}, "cell_type": "markdown", "source": ["Now Let's create the basis of the categories binary dateframe:"]}, {"metadata": {"_cell_guid": "fd195f2b-608e-4f78-be61-570a6345acb2", "_uuid": "7553323cb4c0c6f5869942b9b54fe94216b46907"}, "execution_count": null, "source": ["cat_bin=categories.iloc[:,[0]]\n", "cat_bin[\"cl1\"]=pd.Series()\n", "cat_bin[\"cl2\"]=pd.Series()\n", "cat_bin[\"cl3\"]=pd.Series()"], "cell_type": "code", "outputs": []}, {"metadata": {"_cell_guid": "5824a3bf-1c35-4785-97dd-cab2d1c4bdd3", "_uuid": "7d213ace0fbf2b12a94daa26eda7d8c9ee9d24d7"}, "cell_type": "markdown", "source": ["The transformation algorthim:"]}, {"metadata": {"_cell_guid": "2bc9c3a8-444b-4756-ae11-e46a4d2190c8", "_uuid": "2ff725767aa138f1d057894b20590be025bdb3c5"}, "execution_count": null, "source": ["for i in range(len(cat_lv1)):\n", "    df=categories[categories[\"category_level1\"]==cat_lv1[i]]\n", "    cat_bin.cl1[categories.category_level1==cat_lv1[i]]=i\n", "    cls2=df.category_level2.unique()\n", "    no_cl2=len(cls2)\n", "    for j in range(no_cl2):\n", "        df2=df[df[\"category_level2\"]==cls2[j]]\n", "        cat_bin.cl2[categories.category_level2==cls2[j]]=j\n", "        no_cl3=len(df2)\n", "        cls3=df2.category_level3.unique()\n", "        for k in range(no_cl3):\n", "            cat_bin.cl3[categories.category_level3==cls3[k]]=k\n"], "cell_type": "code", "outputs": []}, {"metadata": {"_cell_guid": "1959a6a1-73bc-4e85-81f1-0a38f7ac8b32", "_uuid": "c59c2a3c87b628c726f701e0a8c58f281b6b9d0d"}, "execution_count": null, "source": ["print(cat_bin.head(10))"], "cell_type": "code", "outputs": []}, {"metadata": {"_cell_guid": "4019ed1e-4950-43ca-ac26-a0b14ec8b6b2", "_uuid": "582e702e5ea1e95d1c984c48952dd53307e6f24f"}, "cell_type": "markdown", "source": ["As You can see Now you have binary hierarchy of the categories names.\n", "But there is a bug in line 11 in the transforming algorthim because there are dulpicated names in categorey level 3, \"You can see the effect of the bug in raw 3\" I have corrected that after checking the duplicated raws of categories that crosspond to cat_bin, here is the solution:"]}, {"metadata": {"_cell_guid": "ef61fca5-0fa0-4827-af3c-cba3f98e628d", "_uuid": "1d7826e7c141067ed3a73d1025583e552a684c52", "collapsed": true}, "execution_count": null, "source": ["cat_bin[\"cl3\"][3]=2\n", "cat_bin[\"cl3\"][1149]=23\n", "cat_bin[\"cl3\"][758]=96\n", "cat_bin[\"cl3\"][3704]=1\n", "cat_bin[\"cl3\"][4018]=21\n", "cat_bin[\"cl3\"][893]=11\n", "cat_bin[\"cl3\"][105]=9"], "cell_type": "code", "outputs": []}, {"metadata": {"_cell_guid": "dcc915b6-0cba-4e4b-a17a-eac8b9f96dc7", "_uuid": "e54a2af86e0344500ac1c846756cc8e4f4870781"}, "cell_type": "markdown", "source": ["Finally create the csv file to use it when needed:"]}, {"metadata": {"_cell_guid": "4ca5d5a1-7dad-4cd8-8629-3156b8613748", "_uuid": "64d80c564d9dad27071aa6432cf5e76fe3978ea0"}, "execution_count": null, "source": ["cat_bin.to_csv('category_binary',index=False)"], "cell_type": "code", "outputs": []}], "nbformat": 4, "metadata": {"language_info": {"codemirror_mode": {"name": "ipython", "version": 3}, "version": "3.6.1", "file_extension": ".py", "name": "python", "mimetype": "text/x-python", "nbconvert_exporter": "python", "pygments_lexer": "ipython3"}, "kernelspec": {"display_name": "Python 3", "language": "python", "name": "python3"}}, "nbformat_minor": 1}