{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"In most real-world projects, the data files are not orderly placed in a single folder. Instead, to be more organized they are often clustered under diffrent sub-folders. This is the case for the data provided for [siim-covid19-detection](https://www.kaggle.com/c/siim-covid19-detection) competition.\n\nIn order to handle this project the very first step is to load the images efficiently. In this notebook, two methods are provided for this step.\n\nThis is particulatly important to automate loading large datasets with a few lines of code.\n\nFirst let's import some libraries","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport os\nfrom os import path\nimport itertools","metadata":{"execution":{"iopub.status.busy":"2021-06-21T22:33:43.790563Z","iopub.execute_input":"2021-06-21T22:33:43.791143Z","iopub.status.idle":"2021-06-21T22:33:43.794701Z","shell.execute_reply.started":"2021-06-21T22:33:43.791108Z","shell.execute_reply":"2021-06-21T22:33:43.793958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Method1: Using the files information sheet**:\nIn the [siim-covid19-detection](https://www.kaggle.com/c/siim-covid19-detection) data, the folder and file names are given in the train_image_level.csv file. We can use this information to load all files.\n \nFirst, let's define the paths and load the .csv file.","metadata":{}},{"cell_type":"code","source":"InputPath = \"../input/siim-covid19-detection\"\nTrainPath = f\"{InputPath}/train\"\nTestPath = f\"{InputPath}/test\"\n\ntrain_image_level = pd.read_csv(f\"{InputPath}/train_image_level.csv\")\ntrain_image_level.head(\n)\n","metadata":{"execution":{"iopub.status.busy":"2021-06-21T22:33:48.530579Z","iopub.execute_input":"2021-06-21T22:33:48.531180Z","iopub.status.idle":"2021-06-21T22:33:48.576900Z","shell.execute_reply.started":"2021-06-21T22:33:48.531147Z","shell.execute_reply":"2021-06-21T22:33:48.575149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#let's change the id to ImageID, and remove the '-image' from the end of each element\ntrain_image_level.rename({\"id\" : \"ImageID\"}, axis = 1, inplace = True)\ntrain_image_level[\"ImageID\"] = train_image_level[\"ImageID\"].apply(lambda x:f'{x[:-6]}')\n\ntrain_image_level.head()","metadata":{"execution":{"iopub.status.busy":"2021-06-21T22:33:49.827344Z","iopub.execute_input":"2021-06-21T22:33:49.828009Z","iopub.status.idle":"2021-06-21T22:33:49.843836Z","shell.execute_reply.started":"2021-06-21T22:33:49.827969Z","shell.execute_reply":"2021-06-21T22:33:49.842184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In order to save time, let's pick a sample of 5 images to read for now. You can later apply the same method to read all files.","metadata":{}},{"cell_type":"code","source":"sample_df = train_image_level.head(5).reset_index(drop = True)\nprint(sample_df)","metadata":{"execution":{"iopub.status.busy":"2021-06-21T22:33:51.972258Z","iopub.execute_input":"2021-06-21T22:33:51.972582Z","iopub.status.idle":"2021-06-21T22:33:51.980448Z","shell.execute_reply.started":"2021-06-21T22:33:51.972554Z","shell.execute_reply":"2021-06-21T22:33:51.979788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i, rows in sample_df.iterrows():\n    dir = os.listdir(TrainPath + \"/\" + rows[\"StudyInstanceUID\"])\n    for k in dir:\n        ImagePath_1 = TrainPath + \"/\" + rows[\"StudyInstanceUID\"] + \"/\"+ k + \"/\" + rows[\"ImageID\"] + \".dcm\"\n        if path.exists(ImagePath_1):\n            print(ImagePath_1)\n            break","metadata":{"execution":{"iopub.status.busy":"2021-06-21T22:33:52.830879Z","iopub.execute_input":"2021-06-21T22:33:52.831355Z","iopub.status.idle":"2021-06-21T22:33:52.851490Z","shell.execute_reply.started":"2021-06-21T22:33:52.831309Z","shell.execute_reply":"2021-06-21T22:33:52.850830Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Method2: Using Pythom method walk()**:\nThe [walk()](https://www.tutorialspoint.com/python/os_walk.htm) method generates the file names in a directory tree.\n\nMore information can be found [here](https://docs.python.org/3/library/os.html).","metadata":{}},{"cell_type":"code","source":"AllTrainFiles=os.walk(TrainPath)\nfor root, dirs, files in itertools.islice(AllTrainFiles,15):\n    for name in files:\n        if name[-4:]=='.dcm':\n            ImagePath_2 = os.path.join(root, name)\n            print(ImagePath_2)","metadata":{"execution":{"iopub.status.busy":"2021-06-21T22:33:55.093702Z","iopub.execute_input":"2021-06-21T22:33:55.094662Z","iopub.status.idle":"2021-06-21T22:33:55.124898Z","shell.execute_reply.started":"2021-06-21T22:33:55.094609Z","shell.execute_reply":"2021-06-21T22:33:55.123674Z"},"trusted":true},"execution_count":null,"outputs":[]}]}