{"metadata": {"kernelspec": {"display_name": "Python 3", "name": "python3", "language": "python"}, "language_info": {"file_extension": ".py", "name": "python", "codemirror_mode": {"version": 3, "name": "ipython"}, "version": "3.6.1", "mimetype": "text/x-python", "nbconvert_exporter": "python", "pygments_lexer": "ipython3"}}, "nbformat": 4, "nbformat_minor": 1, "cells": [{"metadata": {"_uuid": "e5d62013ddf9016dc17f7fd059f915d210981dfa", "_cell_guid": "a896c938-8be7-4453-90a9-6eefa176c18d"}, "source": ["Hello\n", "This Kernel/Notebook work towards converting the categories_names.csv hierarchy into categories binary hierarchy, the need of a binary hierarchical categories , would make it easier when splitting the data according to the categories hieracrchy **in case you want to approch this problem by using multiple models for the categories in each level.** for example you may create one model to classfiey the class in category level 1, then for each prediction on the level 1 model you choose which model to choose next..."], "cell_type": "markdown"}, {"execution_count": null, "metadata": {"_uuid": "9091f650bcae6882849b036bfd59cff632b99e98", "_cell_guid": "96ac9e2e-f888-4f06-91c1-6fa41da85c2d"}, "source": ["# This Python 3 environment comes with many helpful analytics libraries installed\n", "# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n", "# For example, here's several helpful packages to load in \n", "\n", "import numpy as np # linear algebra\n", "import pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n", "\n", "# Input data files are available in the \"../input/\" directory.\n", "# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n", "\n", "from subprocess import check_output\n", "print(check_output([\"ls\", \"../input\"]).decode(\"utf8\"))\n", "\n", "# Any results you write to the current directory are saved as output.\n"], "outputs": [], "cell_type": "code"}, {"execution_count": null, "metadata": {"_uuid": "5c7e37bf8683ef07f96c9f8a8a73bab7cd367af5", "_cell_guid": "8c5b2c19-51cc-44b7-9b28-d2cd5ad94096", "collapsed": true}, "source": ["categories=pd.read_csv(\"../input/category_names.csv\")"], "outputs": [], "cell_type": "code"}, {"metadata": {"_uuid": "2d426a991fdad04cf3cafe8406ea4e36525183d4", "_cell_guid": "252137e5-ef5f-44cd-8d72-a849abfb312d"}, "source": ["Let Expolre the name of the columns, types...etc"], "cell_type": "markdown"}, {"execution_count": null, "metadata": {"_uuid": "ecf934f8b4a9c82e5fb12dc7f776004764b199c2", "_cell_guid": "23495a5e-dae4-427e-96e7-977ef0b4ce1b"}, "source": ["print(categories.info())\n", "print(categories.head(10))"], "outputs": [], "cell_type": "code"}, {"metadata": {"_uuid": "0cf169ca47c2d017461cfa739fcf5aa6bbfe139c", "_cell_guid": "f7062b3f-da6e-4dbc-a95d-bcefd9aa9853"}, "source": ["Lets check how many categories there are for the level 1 :"], "cell_type": "markdown"}, {"execution_count": null, "metadata": {"_uuid": "c29e4c907166191e54cc968c7dbea66f6a7a19b0", "_cell_guid": "743996ba-cd8e-4cd5-81ae-30edcba0644f"}, "source": ["cat_lv1=categories[\"category_level1\"].unique()\n", "print(len(cat_lv1))"], "outputs": [], "cell_type": "code"}, {"metadata": {"_uuid": "f1ec929897e220a38f2f898237c8a2497a59cf50", "_cell_guid": "f6862f27-ab52-4517-9a60-0c5311ec013e"}, "source": ["Important note:  Given that category level 3 is the most lower category level, does each name of the category level 3 corrsponds to a unique category_id ?,\n", "let's check that out:\n"], "cell_type": "markdown"}, {"execution_count": null, "metadata": {"_uuid": "0fe701850e2c2fae70520c25a4bb86bed5081531", "_cell_guid": "4df7e56e-8606-4985-9dc6-84be3dd3d321"}, "source": ["print(\"Total unique Categories Names (Level 3)\",len(categories.category_level3.unique()))\n", "print(\"Total unique Categories Ids\",len(categories.category_id.unique()))\n", "print(\"Number of total observation/rows\",len(categories))"], "outputs": [], "cell_type": "code"}, {"metadata": {"_uuid": "bd4b585b6c67c637223792b977437bc6b87a7288", "_cell_guid": "c0ccac82-d0b0-4840-9064-536d2959b402"}, "source": ["As you can observe there are seven duplicated names in category level 3 "], "cell_type": "markdown"}, {"execution_count": null, "metadata": {"_uuid": "7e7ea0bbcc4c6a46319c7fb11e80ad69941843e4", "_cell_guid": "5860a1f6-0aae-464d-85ef-04a363867c8a"}, "source": ["categories[categories.category_level3.duplicated()]\n"], "outputs": [], "cell_type": "code"}, {"metadata": {"_uuid": "3170abf5ed27c9233fffeaff16f24b7e353f5de3", "_cell_guid": "749c71b7-0d20-40a7-9613-fed84424833f", "_kg_hide-output": true}, "source": ["Now Let's create the basis of the categories binary dateframe:"], "cell_type": "markdown"}, {"execution_count": null, "metadata": {"_uuid": "7553323cb4c0c6f5869942b9b54fe94216b46907", "_cell_guid": "fd195f2b-608e-4f78-be61-570a6345acb2"}, "source": ["cat_bin=categories.iloc[:,[0]]\n", "cat_bin[\"cl1\"]=pd.Series()\n", "cat_bin[\"cl2\"]=pd.Series()\n", "cat_bin[\"cl3\"]=pd.Series()"], "outputs": [], "cell_type": "code"}, {"metadata": {"_uuid": "7d213ace0fbf2b12a94daa26eda7d8c9ee9d24d7", "_cell_guid": "5824a3bf-1c35-4785-97dd-cab2d1c4bdd3"}, "source": ["The transformation algorthim:"], "cell_type": "markdown"}, {"execution_count": null, "metadata": {"_uuid": "2ff725767aa138f1d057894b20590be025bdb3c5", "_cell_guid": "2bc9c3a8-444b-4756-ae11-e46a4d2190c8"}, "source": ["for i in range(len(cat_lv1)):\n", "    df=categories[categories[\"category_level1\"]==cat_lv1[i]]\n", "    cat_bin.cl1[categories.category_level1==cat_lv1[i]]=i\n", "    cls2=df.category_level2.unique()\n", "    no_cl2=len(cls2)\n", "    for j in range(no_cl2):\n", "        df2=df[df[\"category_level2\"]==cls2[j]]\n", "        cat_bin.cl2[categories.category_level2==cls2[j]]=j\n", "        no_cl3=len(df2)\n", "        cls3=df2.category_level3.unique()\n", "        for k in range(no_cl3):\n", "            cat_bin.cl3[categories.category_level3==cls3[k]]=k\n"], "outputs": [], "cell_type": "code"}, {"execution_count": null, "metadata": {"_uuid": "c59c2a3c87b628c726f701e0a8c58f281b6b9d0d", "_cell_guid": "1959a6a1-73bc-4e85-81f1-0a38f7ac8b32"}, "source": ["print(cat_bin.head(10))"], "outputs": [], "cell_type": "code"}, {"metadata": {"_uuid": "582e702e5ea1e95d1c984c48952dd53307e6f24f", "_cell_guid": "4019ed1e-4950-43ca-ac26-a0b14ec8b6b2"}, "source": ["As You can see Now you have binary hierarchy of the categories names.\n", "But there is a bug in line 11 in the transforming algorthim because there are dulpicated names in categorey level 3, \"You can see the effect of the bug in raw 3\" I have corrected that after checking the duplicated raws of categories that crosspond to cat_bin, here is the solution:"], "cell_type": "markdown"}, {"execution_count": null, "metadata": {"_uuid": "1d7826e7c141067ed3a73d1025583e552a684c52", "_cell_guid": "ef61fca5-0fa0-4827-af3c-cba3f98e628d", "collapsed": true}, "source": ["cat_bin[\"cl3\"][3]=2\n", "cat_bin[\"cl3\"][1149]=23\n", "cat_bin[\"cl3\"][758]=96\n", "cat_bin[\"cl3\"][3704]=1\n", "cat_bin[\"cl3\"][4018]=21\n", "cat_bin[\"cl3\"][893]=11\n", "cat_bin[\"cl3\"][105]=9"], "outputs": [], "cell_type": "code"}, {"metadata": {"_uuid": "e54a2af86e0344500ac1c846756cc8e4f4870781", "_cell_guid": "dcc915b6-0cba-4e4b-a17a-eac8b9f96dc7"}, "source": ["Finally create the csv file to use it when needed:"], "cell_type": "markdown"}, {"execution_count": null, "metadata": {"_uuid": "64d80c564d9dad27071aa6432cf5e76fe3978ea0", "_cell_guid": "4ca5d5a1-7dad-4cd8-8629-3156b8613748"}, "source": ["cat_bin.to_csv('category_binary',index=False)"], "outputs": [], "cell_type": "code"}]}