{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"markdown","source":"# Exploring classes and Heirarchy (fixed)\n"},{"metadata":{},"cell_type":"markdown","source":"This is a fork of [very useful notebook](https://www.kaggle.com/thanatoz/understanding-open-image-v5-classes-hierarchy) that has helped me and others understand label classes and how they relate. The only problem was that public dataset anotations have 600 classes while competition uses annotations with class count reduced to 500. Original kernel has incorrectly used labels and class descriptions from public instead of competition dataset and thus found too many classes missing from the class hierarchy. After fixing data sources, there are still classes missing from the hierarchy, but this time only 19, \n\nIf you feel that I am wrong anywhere, feel free to comment below and help improvise this kernel. "},{"metadata":{},"cell_type":"markdown","source":"> These annotation files cover the 500 boxable object classes, and span the 1,743,042 training images where we annotated bounding boxes, object segmentations, and visual relationships, as well as the full validation (41,620 images) and test (125,436 images) sets."},{"metadata":{},"cell_type":"markdown","source":"### Downloading the required files\n\nI am using the annotation files and the class_names files. So download them from the given link."},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true,"_kg_hide-output":true},"cell_type":"code","source":"# Downloading the hierarchy json\n# !wget https://storage.googleapis.com/openimages/2018_04/bbox_labels_600_hierarchy.json # --> old link\n!wget https://storage.googleapis.com/openimages/challenge_2019/challenge-2019-label500-hierarchy.json\n\n# Downloading class names\n!wget https://storage.googleapis.com/openimages/challenge_2019/challenge-2019-classes-description-500.csv\n# used to be https://storage.googleapis.com/openimages/v5/class-descriptions-boxable.csv\n    \n# Downlaoding class-annotations\n!wget https://storage.googleapis.com/openimages/challenge_2019/challenge-2019-train-detection-bbox.csv\n# used to be https://storage.googleapis.com/openimages/2018_04/train/train-annotations-bbox.csv","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import pandas as pd\ncls=pd.read_csv('challenge-2019-classes-description-500.csv', header=None)\nclasses2name={i:j for i,j in zip(cls[0], cls[1])}\nname2classes={j:i for i,j in zip(cls[0], cls[1])}\ncls.tail()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"trn=pd.read_csv('challenge-2019-train-detection-bbox.csv')\ntrn.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Lets now print the classes from the Attached JSON given with the dataset. \nHere you can understand the heirarchy level as per the name indentation."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"import json\nhier = json.load(open('challenge-2019-label500-hierarchy.json','r'))\nlevel1=[]\nlevel2=[]\nlevel3=[]\nfor l2 in hier['Subcategory']:\n    print(classes2name[l2['LabelName']])\n    level3.append(classes2name[l2['LabelName']])\n    try:\n        for j in l2['Subcategory']:\n            print('----> ',classes2name[j['LabelName']])\n            level2.append(classes2name[j['LabelName']])\n            try:\n                for k in j['Subcategory']:\n                    print('\\t----> ',classes2name[k['LabelName']])\n                    level1.append(classes2name[k['LabelName']])\n            except:\n                pass\n    except:\n        pass\n        \nlevel1 = set(level1)\nlevel2 = set(level2)\nlevel3 = set(level3)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Classes count\n\nWe can see that we obtain 3 levels of classes hierarchy. "},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print('Classes count in level 1 is {}'.format(len(level1)))\nprint('Classes count in level 2 is {}'.format(len(level2)))\nprint('Classes count in level 3 is {}'.format(len(level3)))\nprint('Total unique class counts are {}'.format(len(level1)+len(level2)+len(level3)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So there are 483 unique classes that we obtain from the json provided to us. But there could be classes overlap. Lets run a quick check over this."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"print('level1 and level2 overlaps = {}'.format(len(level2&level1)))\nprint('level2 and level3 overlaps = {}'.format(len(level2&level3)))\nprint('level1 and level3 overlaps = {}'.format(len(level3&level1)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"With 2 oveplapping classes, that makes effective number of unique classes in the hierarchy 481."},{"metadata":{},"cell_type":"markdown","source":"### Classes along with counts \n\nThe index of the classes in a list along with the value count in the training dataset."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"res=trn['LabelName'].value_counts()\ntrn_classes=[]\nfor idx, (i,j) in enumerate(zip(res.index, res)):\n    trn_classes.append(classes2name[i])\n    print('{} \\t {} \\t {}'.format(idx+1, classes2name[i], j))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Classes that could be missing\n\nWe have seen that we are provided with 500 classes but we obtained lesser classes from the classes hierarchy JSON. So lets try to print the classes that could be missing."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"all_training_classes = set(list(cls[1]))\nall_json_classes = level1.union(level2).union(level3)\nprint('There are {} classes from JSON file and {} classes from training file'.format(len(all_json_classes), len(all_training_classes)))\nprint('Thus there are {} missing classes'.format(len(all_training_classes)-len(all_json_classes)))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"So the missing classes could be these."},{"metadata":{"_kg_hide-input":true,"trusted":true},"cell_type":"code","source":"all_training_classes-all_json_classes","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.4","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}