{"cells":[{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","collapsed":true,"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":false},"cell_type":"markdown","source":"The dataset for this competition is a subset of NIH dataset. \n\nIt might be useful to use labels provided by NIH to train the model.\n\nI share the list which associates train images to original NIH images so that you can make use of NIH labels. The list is obtrained by calculating the diffs among images."},{"metadata":{"_uuid":"891ab69e6d359842ec86d84659018effcd9ce515"},"cell_type":"markdown","source":"## Update\n\nAs there is a request to add this for stage1 test images, I've uploaded a new csv 'nih_for_test1_images.csv'.\n\n- nih.csv: for train images\n- nih_for_test1_images.csv: for stage1 test images\n\nOn this competition, you may use NIH label for training your model, but I not sure you can use NIH label for testing. Please check the disussion about external data https://www.kaggle.com/c/rsna-pneumonia-detection-challenge/discussion/64345"},{"metadata":{"_uuid":"89cf770e1b722da3ff58af1c0f0c512f6fee5fb0"},"cell_type":"markdown","source":"## NIH dataset\n\nLink to the dataset: https://nihcc.app.box.com/v/ChestXray-NIHCC\n\nNIH labels: https://nihcc.app.box.com/v/ChestXray-NIHCC/file/219760887468"},{"metadata":{"_uuid":"312dc88f583fb4d7ecfa257d1421b12cdd8dc301"},"cell_type":"markdown","source":"## Dataframe to associate patientId to original NIH image and its label"},{"metadata":{"_uuid":"1096c558a13b16eff5d8e176efdc34581c73f4fc"},"cell_type":"markdown","source":"I calculated the pixel diffs between competition's dataset images and NIH images. A image pair with the minimum diff is treated as the same image. \n\nI manually checked hundreds of image pairs and they were all correct.\n\nAfter finding out the original NIH images, I put all labels of this competition and NIH into the dataframe, nih.csv."},{"metadata":{"_uuid":"7f38ddb7f33642d5447b835d8fefad312948b94a"},"cell_type":"markdown","source":"#### for train images"},{"metadata":{"trusted":true,"_uuid":"5550dbb2a7462ed659f197414526e992cd4437df"},"cell_type":"code","source":"import pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\ndf = pd.read_csv('../input/nihcsv/nih.csv')\nprint(df.shape)\ndf.head(10)\n# class2 is the label given in NIH dataset (and imageIndex is filename of the image)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a111bc9ef9e8624797c31aa8d6e3fc04f1020856"},"cell_type":"markdown","source":"#### for stage1 test images"},{"metadata":{"trusted":true,"_uuid":"b9a6f54e25309c3ef763d72146e84bf1eebcc9d1"},"cell_type":"code","source":"import os\nprint(os.listdir('../input/nihcsv'))\n\ntest = pd.read_csv('../input/nihcsv/nih_for_test1_images.csv')\nprint(test.shape)\ntest.head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"784890e5c8a736fbb0db9e9dd00408c85d68e6d7"},"cell_type":"markdown","source":"## NIH label for 'Normal' class\n\nThere are some images labeled 'Nodule' or 'Atelectasis' at NIH but labeled as 'Normal' in this competition. I'm a bit surprised to know there are lots of images with Infiltration at NIH but treated at 'Normal' here."},{"metadata":{"trusted":true,"_uuid":"22a759b15ec8c09668677f39b9e5c6bb5fe16728"},"cell_type":"code","source":"df[df.class1 == 'Normal'].class2.value_counts().to_frame().head(10).plot.bar()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cd8ffb9216087cbedf19307edccd86230c34145e"},"cell_type":"markdown","source":"## NIH label for 'Lung Opacity' class\n\nI was expecting most images are labeled as 'Infiltration' at NIH and that was somewhat correct."},{"metadata":{"trusted":true,"_uuid":"362514945b9d75beba4362c5cc1c7b23fbe4d848"},"cell_type":"code","source":"df[df.class1 == 'Lung Opacity'].class2.value_counts().to_frame().head(10).plot.bar()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fcb8cb35139133fff5dd3a4752304874743e36f0"},"cell_type":"markdown","source":"## NIH label for 'No Lung Opacity / Not Normal' class\n\nA bit surprising to know there are many images labeled as 'No Finding' at NIH."},{"metadata":{"trusted":true,"_uuid":"4fb13dcc20f985d0bbd93aa9f4d9482ff9f29b4b"},"cell_type":"code","source":"df[df.class1 == 'No Lung Opacity / Not Normal'].class2.value_counts().to_frame().head(10).plot.bar()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"086cec2e3992c0af8c9599d75abf20177cbc285e"},"cell_type":"markdown","source":"## Guessing a reason for label inconsistency between this competition and NIH.\n\nI'm not a specialist in this field and I can not say much about the variance among specialists's decisions.\n\nI guess one of the reason is the fact how NIH label is created. https://arxiv.org/pdf/1705.02315.pdf\nNIH label is generated by NLP technique and it might not be as solid as the label in this competition.\n"},{"metadata":{"trusted":true,"_uuid":"282dd46b51d9335ca17b29cb23c4084d459c2d4d"},"cell_type":"code","source":"for class2, count2 in df.class2.value_counts().items():\n\n    if count2 < 100: # ignore small count\n        continue\n\n    print('\\n----- %s -----' % class2)\n\n    _df = df[df.class2 == class2].class1\n    for class1, count1 in _df.value_counts().items():\n        ratio = count1 / _df.count()\n        print('%d (%.2f%%) %s' % (count1, ratio * 100, class1))","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}