{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Goal \n### Identifying and localizing COVID-19 abnormalities on chest radiographs","metadata":{}},{"cell_type":"markdown","source":"## Reference  \n* [YOLO ref1]( https://www.kaggle.com/ayuraj/train-covid-19-detection-using-yolov5)  \n* [YOLO ref2](https://www.kaggle.com/h053473666/siim-cov19-yolov5-train#YOLOv5)  \n* [simple-tutorial](https://www.kaggle.com/yujiariyasu/catch-up-on-positive-samples-plot-submission-csv?scriptVersionId=63394385)  ","metadata":{}},{"cell_type":"markdown","source":"## The Domain Knowledge\n#### ★Ground glass opacties\nground glass opacities (GGOs, for short) indicate abnormalities in the lungs. \"Ground glass opacities [are] a pattern that can be seen when the lungs are sick,\" says Dr. Cortopassi. She adds that, while normal lung CT scans appear black, an abnormal chest CT with GGOs will show lighter-colored or gray patches.\n\n#### ★Opacity(不透明度)\nthe degree of transparenet(x-ray image)\n\n#### 1. Typical Appearance  \nCommonly reported imaging features of greater specificity for COVID-19 pneumonia.\n#### 2. Atypical Appearance  \nUncommonly or not reported features of COVID-19 pneumonia.\n#### 3. Indeterminate Appearance(不確定)   \nNonspecific imaging features of COVID-19 pneumonia.\n#### 4. Negative for Pneumonia(陰性） \n\n#### boxes\nbounding boxes in easily-readable dictionary format\n\n#### DICOM format\nAny DICOM medical image consists of two parts—a header and the actual image itself. The header consists of data that describes the image, the most important being patient data.This includes the patient’s demographic information such as the patient’s name, age, gender, and date of birth.Hy\n(https://theaisummer.com/medical-image-coordinates/)","metadata":{}},{"cell_type":"markdown","source":"# Data","metadata":{}},{"cell_type":"markdown","source":"* train_study_level.csv - the train study-level metadata, with one row for each study, including correct labels.\n* train_image_level.csv - the train image-level metadata, with one row for each image, including both correct labels and any bounding boxes in a dictionary format.  \nSome images in both test and train have multiple bounding boxes.\n* sample_submission.csv - a sample submission file containing all image- and study-level IDs.\n* train folder - comprises 6,334 chest scans in DICOM format, stored in paths with the form study/series/image\n* test folder - The hidden test dataset is of roughly the same scale as the training dataset.","metadata":{}},{"cell_type":"markdown","source":"## EDA","metadata":{}},{"cell_type":"markdown","source":"The Process of EDA\n1. Data visualization\n2. Feature select\n3. Feature engineering\n4. fill in missing value","metadata":{}},{"cell_type":"markdown","source":"### set up W&B\n* save learning parameter\n* vizualiztion image file","metadata":{}},{"cell_type":"code","source":"import wandb\nimport os\n# os._Environは環境変数名keyと値valueが対になったマップ型オブジェクト\n# print(os.environ)\n# wandbとのAPI接続を暗号化する\n#!wandb login $api_key","metadata":{"execution":{"iopub.status.busy":"2021-08-05T02:49:03.686792Z","iopub.execute_input":"2021-08-05T02:49:03.687172Z","iopub.status.idle":"2021-08-05T02:49:04.439997Z","shell.execute_reply.started":"2021-08-05T02:49:03.687136Z","shell.execute_reply":"2021-08-05T02:49:04.43875Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"\"\"from kaggle_secrets import UserSecretsClient\nuser_secrets = UserSecretsClient()\napi_key = user_secrets.get_secret(\"edc8af4e0cd1f3bba30aaea945348675bb6346de\")\n\nos.environ[\"WANDB_SILENT\"] = \"true\"\nCONFIG = {'competition': 'siim-fisabio-rsna', '_wandb_kernel': 'ruch'}\"\"\"","metadata":{"execution":{"iopub.status.busy":"2021-08-05T02:49:04.442192Z","iopub.execute_input":"2021-08-05T02:49:04.44263Z","iopub.status.idle":"2021-08-05T02:49:04.44926Z","shell.execute_reply.started":"2021-08-05T02:49:04.442579Z","shell.execute_reply":"2021-08-05T02:49:04.448272Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Libarary","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport pandas_profiling\nimport cv2\nimport pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport pydicom \nimport random\nimport albumentations as A\nfrom sklearn.model_selection import train_test_split\nfrom fastai.vision.all import *\nfrom fastai.medical.imaging import *\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut","metadata":{"execution":{"iopub.status.busy":"2021-08-06T01:29:28.948107Z","iopub.execute_input":"2021-08-06T01:29:28.948464Z","iopub.status.idle":"2021-08-06T01:29:34.423152Z","shell.execute_reply.started":"2021-08-06T01:29:28.948428Z","shell.execute_reply":"2021-08-06T01:29:34.422231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Look at train_study_level.csv / train_image_level.csv","metadata":{}},{"cell_type":"code","source":"%cd input","metadata":{"execution":{"iopub.status.busy":"2021-08-06T00:13:58.593188Z","iopub.execute_input":"2021-08-06T00:13:58.593610Z","iopub.status.idle":"2021-08-06T00:13:58.601388Z","shell.execute_reply.started":"2021-08-06T00:13:58.593528Z","shell.execute_reply":"2021-08-06T00:13:58.600078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_study_df = pd.read_csv(\"../input/siim-covid19-detection/train_study_level.csv\")\ntrain_image_data = pd.read_csv(\"../input/siim-covid19-detection/train_image_level.csv\")","metadata":{"execution":{"iopub.status.busy":"2021-08-06T00:14:01.619158Z","iopub.execute_input":"2021-08-06T00:14:01.619679Z","iopub.status.idle":"2021-08-06T00:14:01.667087Z","shell.execute_reply.started":"2021-08-06T00:14:01.619646Z","shell.execute_reply":"2021-08-06T00:14:01.666042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_study_df.shape)\ntrain_study_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-08-06T00:14:03.754764Z","iopub.execute_input":"2021-08-06T00:14:03.755139Z","iopub.status.idle":"2021-08-06T00:14:03.769762Z","shell.execute_reply.started":"2021-08-06T00:14:03.755109Z","shell.execute_reply":"2021-08-06T00:14:03.768171Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Show Distribution \ntrain_study_level.csv ","metadata":{}},{"cell_type":"code","source":"study_class = [\"Negative for Pneumonia\", \"Typical Appearance\",\"Indeterminate Appearance\", \"Atypical Appearance\"]\nplt.figure(figsize = (10,5))\nplt.bar([1,2,3,4], train_study_df[study_class].values.sum(axis=0))\nplt.xticks([1,2,3,4],study_class)\nplt.ylabel('Frequency')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-08-06T00:14:14.254634Z","iopub.execute_input":"2021-08-06T00:14:14.255012Z","iopub.status.idle":"2021-08-06T00:14:14.671876Z","shell.execute_reply.started":"2021-08-06T00:14:14.254960Z","shell.execute_reply":"2021-08-06T00:14:14.670522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_image_data.head()","metadata":{"execution":{"iopub.status.busy":"2021-08-06T00:14:15.084077Z","iopub.execute_input":"2021-08-06T00:14:15.084413Z","iopub.status.idle":"2021-08-06T00:14:15.097642Z","shell.execute_reply.started":"2021-08-06T00:14:15.084385Z","shell.execute_reply":"2021-08-06T00:14:15.096594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have our bounding box labels provided in the label column. The format is as follows:  \n[class ID] [confidence score] [bounding box]","metadata":{}},{"cell_type":"markdown","source":"#### look at the distribution of opacity vs none:","metadata":{}},{"cell_type":"code","source":"train_image_data['split_label'] = train_image_data.label.apply(lambda x: [x.split()[offs:offs+6] for offs in range(0, len(x.split()), 6)])\n# show the split_label\ntrain_image_data['split_label'][:5]","metadata":{"execution":{"iopub.status.busy":"2021-08-05T02:49:08.176638Z","iopub.execute_input":"2021-08-05T02:49:08.177005Z","iopub.status.idle":"2021-08-05T02:49:08.214643Z","shell.execute_reply.started":"2021-08-05T02:49:08.176969Z","shell.execute_reply":"2021-08-05T02:49:08.213629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_image_data['split_label'].values[0]","metadata":{"execution":{"iopub.status.busy":"2021-08-05T02:49:09.761755Z","iopub.execute_input":"2021-08-05T02:49:09.76216Z","iopub.status.idle":"2021-08-05T02:49:09.769762Z","shell.execute_reply.started":"2021-08-05T02:49:09.762126Z","shell.execute_reply":"2021-08-05T02:49:09.768744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"classes_freq = []\nfor i in range(len(train_image_data)):\n    for j in train_image_data.iloc[i].split_label: classes_freq.append(j[0])\nplt.hist(classes_freq)\nplt.ylabel('Frequency')","metadata":{"execution":{"iopub.status.busy":"2021-08-05T02:49:09.928454Z","iopub.execute_input":"2021-08-05T02:49:09.928823Z","iopub.status.idle":"2021-08-05T02:49:10.808932Z","shell.execute_reply.started":"2021-08-05T02:49:09.928792Z","shell.execute_reply":"2021-08-05T02:49:10.808093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### target label distribution","metadata":{}},{"cell_type":"code","source":"image_data_path = \"../input/siim-covid19-detection/train\"\n\n# show image\n#cv2.imshow(\"image_data\")","metadata":{"execution":{"iopub.status.busy":"2021-08-05T02:49:10.810232Z","iopub.execute_input":"2021-08-05T02:49:10.810685Z","iopub.status.idle":"2021-08-05T02:49:10.814319Z","shell.execute_reply.started":"2021-08-05T02:49:10.810631Z","shell.execute_reply":"2021-08-05T02:49:10.813613Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### look at the images \n","metadata":{}},{"cell_type":"code","source":"# pixel_arrayプロパティを用いることで画像データがNumPyのndarrayとして取得する\n\ndef dicom2array(path, voi_lut=True, fix_monochrome=True):\n    dicom = pydicom.read_file(path)\n    # transform raw DICOM data to \"human-friendly\" view\n    if voi_lut:\n        data = apply_voi_lut(dicom.pixel_array, dicom)\n    else:\n        data = dicom.pixel_array\n    if fix_monochrome and dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = np.amax(data) - data\n    data = data - np.min(data)\n    data = data / np.max(data)\n    data = (data * 255).astype(np.uint8)\n    return data\n\n\ndef plot_img(img, size=(7, 7),is_rgb=True, title=\"\", cmap='grap'):\n    plt.figure(figsize=size)\n    plt.imshow(img, cmap=camp)\n    plt.suptitle(title)\n    plt.show()\n    \ndef plot_imgs(imgs, cols=4, size=7,is_rgb=True ,title=\"\", cmap='gray', img_size=(500, 500)):\n    rows = len(imgs) // cols + 1\n    fig = plt.figure(figsize=(cols*size, rows*size))\n    for i, img in enumerate(imgs):\n        if img_size is not None:\n            img = cv2.resize(img, img_size)\n        fig.add_subplot(rows, cols, i+1)\n        plt.imshow(img, cmap=cmap)\n    plt.suptitle(title)\n    plt.show","metadata":{"execution":{"iopub.status.busy":"2021-08-05T02:49:11.794287Z","iopub.execute_input":"2021-08-05T02:49:11.794847Z","iopub.status.idle":"2021-08-05T02:49:11.807712Z","shell.execute_reply.started":"2021-08-05T02:49:11.794797Z","shell.execute_reply":"2021-08-05T02:49:11.806531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset_path = Path('../input/siim-covid19-detection')\ndicom_paths = get_dicom_files(dataset_path/'train')\nimgs = [dicom2array(path) for path in dicom_paths[:4]]\nplot_imgs(imgs)","metadata":{"execution":{"iopub.status.busy":"2021-08-05T02:52:46.120289Z","iopub.execute_input":"2021-08-05T02:52:46.120644Z","iopub.status.idle":"2021-08-05T02:52:46.132916Z","shell.execute_reply.started":"2021-08-05T02:52:46.120615Z","shell.execute_reply":"2021-08-05T02:52:46.132184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Let's actually look at how many images are available per study:","metadata":{}},{"cell_type":"code","source":"num_images_per_study = []\nfor i in (dataset_path/'train').ls():\n    num_images_per_study.append(len(get_dicom_files(i)))\n    if len(get_dicom_files(i)) > 5:\n        print(f'Study {i} had {len(get_dicom_files(i))} images')","metadata":{"execution":{"iopub.status.busy":"2021-08-05T02:49:18.758814Z","iopub.execute_input":"2021-08-05T02:49:18.759234Z","iopub.status.idle":"2021-08-05T02:49:29.999532Z","shell.execute_reply.started":"2021-08-05T02:49:18.759204Z","shell.execute_reply":"2021-08-05T02:49:29.99839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.hist(num_images_per_study)","metadata":{"execution":{"iopub.status.busy":"2021-08-05T02:49:30.001401Z","iopub.execute_input":"2021-08-05T02:49:30.001749Z","iopub.status.idle":"2021-08-05T02:49:30.230569Z","shell.execute_reply.started":"2021-08-05T02:49:30.001714Z","shell.execute_reply":"2021-08-05T02:49:30.22946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### look at image appled boundig box","metadata":{}},{"cell_type":"code","source":"# 一致するファイルを抽出\ndef image_path(row):\n    study_path = dataset_path/'train'/row.StudyInstanceUID\n    for i in get_dicom_files(study_path):\n        # 拡張子なしのファイル名の文字列はstem属性で取得\n        if row.id.split('_')[0] == i.stem: return i\n    \n\ntrain_image_data['image_path'] = train_image_data.apply(image_path, axis=1)","metadata":{"execution":{"iopub.status.busy":"2021-08-05T02:49:32.456591Z","iopub.execute_input":"2021-08-05T02:49:32.456941Z","iopub.status.idle":"2021-08-05T02:49:39.386715Z","shell.execute_reply.started":"2021-08-05T02:49:32.456911Z","shell.execute_reply":"2021-08-05T02:49:39.385652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_image_data.head()","metadata":{"execution":{"iopub.status.busy":"2021-08-05T02:49:39.388052Z","iopub.execute_input":"2021-08-05T02:49:39.388341Z","iopub.status.idle":"2021-08-05T02:49:39.406783Z","shell.execute_reply.started":"2021-08-05T02:49:39.388315Z","shell.execute_reply":"2021-08-05T02:49:39.405762Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"imgs = []\nimage_paths = train_image_data['image_path'].values\n# ex ('../input/siim-covid19-detection/train/5776db0cec75/81456c9c5423/000a312787f2.dcm')\n\nthickness = 10\nscale = 5\n\n\nfor i in range(8):\n    image_path = random.choice(image_paths)\n    print(image_path)\n    img = dicom2array(path=image_path)\n    img = cv2.resize(img, None, fx=1/scale, fy=1/scale)\n    img = np.stack([img, img, img], axis=-1)\n    for i in train_image_data.loc[train_image_data['image_path'] == image_path].split_label.values[0]:\n        if i[0] == 'opacity':\n            img = cv2.rectangle(img, (int(float(i[2])/ scale), int(float(i[3])/ scale)),\n                                     (int(float(i[4])/ scale), int(float(i[5])/ scale)),\n                                     [0, 255, 0], thickness)\n    img = cv2.resize(img, (500, 500))\n    imgs.append(img)\n            \nplot_imgs(imgs, cmap=None)\n\n###\n# split_label.values ex) ['opacity', '1', '789.28836', '582.43035', '1815.94498', '2499.73327']","metadata":{"execution":{"iopub.status.busy":"2021-08-05T02:49:44.780231Z","iopub.execute_input":"2021-08-05T02:49:44.780635Z","iopub.status.idle":"2021-08-05T02:49:48.72767Z","shell.execute_reply.started":"2021-08-05T02:49:44.780599Z","shell.execute_reply":"2021-08-05T02:49:48.724227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Albumenatations","metadata":{}},{"cell_type":"markdown","source":"sample1 https://propen.dream-target.jp/blog/python_albumentations  \nsample2 https://qiita.com/Takayoshi_Makabe/items/79c8a5ba692aa94043f7","metadata":{}},{"cell_type":"markdown","source":"## Modeling","metadata":{}},{"cell_type":"markdown","source":"###  model YOLO5","metadata":{}},{"cell_type":"markdown","source":"Download Yolov% repository in temp directory","metadata":{}},{"cell_type":"code","source":"%cd ../kaggle\n#!mkdir tmp\n%cd tmp","metadata":{"execution":{"iopub.status.busy":"2021-08-06T01:52:31.310578Z","iopub.execute_input":"2021-08-06T01:52:31.311042Z","iopub.status.idle":"2021-08-06T01:52:31.322358Z","shell.execute_reply.started":"2021-08-06T01:52:31.311011Z","shell.execute_reply":"2021-08-06T01:52:31.321096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Download YOLOv5\n!git clone https://github.com/ultralytics/yolov5\n%cd yolov5\n# Install dependecies\n%pip install -qr requirements.txt\n%cd ../\nimport torch\n# 学習回すときにGPUをONにする\nprint(f\"Setup complete. Using torch {torch.__version__} ({torch.cuda.get_device_properties(0).name if torch.cuda.is_available() else 'CPU'})\")","metadata":{"execution":{"iopub.status.busy":"2021-08-06T01:52:56.313593Z","iopub.execute_input":"2021-08-06T01:52:56.313943Z","iopub.status.idle":"2021-08-06T01:53:06.314064Z","shell.execute_reply.started":"2021-08-06T01:52:56.313911Z","shell.execute_reply":"2021-08-06T01:53:06.313298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Chose model YOLOv5s、YOLOv5m、YOLOv5l、YOLOv5x (big order ⇄ GPU memory)  \nuse YOLOv5s first","metadata":{}},{"cell_type":"markdown","source":"Prepare Folder ","metadata":{}},{"cell_type":"markdown","source":"* required structure for dataset directory\n\n```\n/parent_folder\n    /dataset\n         /images\n             /train\n             /val\n         /labels\n             /train\n             /val\n    /yolov5\n```","metadata":{}},{"cell_type":"markdown","source":"## Hyperparameters Set","metadata":{}},{"cell_type":"code","source":"TRAIN_PATH = 'input/siim-covid19-resized-to-256px-jpg/train/'\nIMG_SIZE = 256\nBATCH_SIZE = 16\nEPOCHS = 10","metadata":{"execution":{"iopub.status.busy":"2021-08-06T01:30:10.193620Z","iopub.execute_input":"2021-08-06T01:30:10.193976Z","iopub.status.idle":"2021-08-06T01:30:10.198345Z","shell.execute_reply.started":"2021-08-06T01:30:10.193947Z","shell.execute_reply":"2021-08-06T01:30:10.197357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Prepare Dataset","metadata":{}},{"cell_type":"code","source":"%cd ..","metadata":{"execution":{"iopub.status.busy":"2021-08-06T01:31:40.013591Z","iopub.execute_input":"2021-08-06T01:31:40.014109Z","iopub.status.idle":"2021-08-06T01:31:40.020776Z","shell.execute_reply.started":"2021-08-06T01:31:40.014078Z","shell.execute_reply":"2021-08-06T01:31:40.019674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#%cd kaggle\ndf = pd.read_csv(\"input/siim-covid19-detection/train_image_level.csv\")","metadata":{"execution":{"iopub.status.busy":"2021-08-06T01:34:03.977665Z","iopub.execute_input":"2021-08-06T01:34:03.978041Z","iopub.status.idle":"2021-08-06T01:34:04.006715Z","shell.execute_reply.started":"2021-08-06T01:34:03.978006Z","shell.execute_reply":"2021-08-06T01:34:04.005795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df['id'] = df.apply(lambda row: row.id.split('_')[0], axis = 1)\ndf['path'] = df.apply(lambda row: TRAIN_PATH+row.id+'.jpg', axis=1)\ndf['image_level'] = df.apply(lambda row: row.label.split(' ')[0], axis=1)","metadata":{"execution":{"iopub.status.busy":"2021-08-06T01:34:04.314787Z","iopub.execute_input":"2021-08-06T01:34:04.315129Z","iopub.status.idle":"2021-08-06T01:34:04.646974Z","shell.execute_reply.started":"2021-08-06T01:34:04.315100Z","shell.execute_reply":"2021-08-06T01:34:04.645905Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Load meta.csv file","metadata":{}},{"cell_type":"code","source":"meta_df = pd.read_csv('input/siim-covid19-resized-to-256px-jpg/meta.csv')","metadata":{"execution":{"iopub.status.busy":"2021-08-06T01:34:06.963614Z","iopub.execute_input":"2021-08-06T01:34:06.963958Z","iopub.status.idle":"2021-08-06T01:34:06.992604Z","shell.execute_reply.started":"2021-08-06T01:34:06.963927Z","shell.execute_reply":"2021-08-06T01:34:06.991606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_meta_df = meta_df.loc[meta_df.split == 'train'] # select_train\ntrain_meta_df = train_meta_df.drop('split', axis=1) # delete_test\ntrain_meta_df.columns = ['id', 'dim0', 'dim1']","metadata":{"execution":{"iopub.status.busy":"2021-08-06T01:34:08.523674Z","iopub.execute_input":"2021-08-06T01:34:08.523998Z","iopub.status.idle":"2021-08-06T01:34:08.539048Z","shell.execute_reply.started":"2021-08-06T01:34:08.523970Z","shell.execute_reply":"2021-08-06T01:34:08.537979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Merge both the dataframes   why??\ndf = df.merge(train_meta_df, on='id', how=\"left\")\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2021-08-06T01:34:11.712339Z","iopub.execute_input":"2021-08-06T01:34:11.712692Z","iopub.status.idle":"2021-08-06T01:34:11.735983Z","shell.execute_reply.started":"2021-08-06T01:34:11.712660Z","shell.execute_reply":"2021-08-06T01:34:11.735067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Train-Validation split","metadata":{}},{"cell_type":"code","source":"train_df, valid_df = train_test_split(df, test_size=0.2, random_state=42, stratify=df.image_level.values)\n\n# ignore warning\ntrain_df.loc[:, 'split'] = 'train'\nvalid_df.loc[:, 'split'] = 'valid'\n\ndf = pd.concat([train_df, valid_df]).reset_index(drop=True)","metadata":{"execution":{"iopub.status.busy":"2021-08-06T01:35:29.060649Z","iopub.execute_input":"2021-08-06T01:35:29.061045Z","iopub.status.idle":"2021-08-06T01:35:29.090049Z","shell.execute_reply.started":"2021-08-06T01:35:29.061014Z","shell.execute_reply":"2021-08-06T01:35:29.088791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'Size of dataset: {len(df)}, training images: {len(train_df)}. validation images: {len(valid_df)}')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### prepare required folder structure","metadata":{}},{"cell_type":"code","source":"os.makedirs('tmp/covid/images/train', exist_ok=True)\nos.makedirs('tmp/covid/images/valid', exist_ok=True)\n! ls tmp/covid/images","metadata":{"execution":{"iopub.status.busy":"2021-08-06T01:39:51.494310Z","iopub.execute_input":"2021-08-06T01:39:51.494687Z","iopub.status.idle":"2021-08-06T01:39:52.227680Z","shell.execute_reply.started":"2021-08-06T01:39:51.494656Z","shell.execute_reply":"2021-08-06T01:39:52.226616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Move the images to relevant split folder.  (Need??)\nfrom tqdm import tqdm\nfrom shutil import copyfile\nfor i in tqdm(range(len(df))):\n    row = df.loc[i]\n    if row.split == 'train':\n        copyfile(row.path, f'tmp/covid/images/train/{row.id}.jpg')\n    else:\n        copyfile(row.path, f'tmp/covid/images/valid/{row.id}.jpg')","metadata":{"execution":{"iopub.status.busy":"2021-08-06T01:42:57.645591Z","iopub.execute_input":"2021-08-06T01:42:57.646234Z","iopub.status.idle":"2021-08-06T01:43:50.626484Z","shell.execute_reply.started":"2021-08-06T01:42:57.646184Z","shell.execute_reply":"2021-08-06T01:43:50.625414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Create .YANL file","metadata":{}},{"cell_type":"code","source":"%cd tmp","metadata":{"execution":{"iopub.status.busy":"2021-08-06T01:50:35.806109Z","iopub.execute_input":"2021-08-06T01:50:35.806452Z","iopub.status.idle":"2021-08-06T01:50:35.813146Z","shell.execute_reply.started":"2021-08-06T01:50:35.806423Z","shell.execute_reply":"2021-08-06T01:50:35.811956Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%pwd","metadata":{"execution":{"iopub.status.busy":"2021-08-06T01:55:56.764757Z","iopub.execute_input":"2021-08-06T01:55:56.765110Z","iopub.status.idle":"2021-08-06T01:55:56.771843Z","shell.execute_reply.started":"2021-08-06T01:55:56.765080Z","shell.execute_reply":"2021-08-06T01:55:56.770804Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import yaml\n\ndata_yaml = dict(\n    train = '../covid/images/train',\n    val = '../covid/images/valid',\n    nc = 2,\n    names = ['none', 'opacity']\n)\n\nwith open('data/data.yaml', 'w') as outfile:\n    yaml.dump(data_yaml, outfile, default_flow_style=True)\n\n%cat data/data.yaml","metadata":{"execution":{"iopub.status.busy":"2021-08-06T01:56:16.444329Z","iopub.execute_input":"2021-08-06T01:56:16.444704Z","iopub.status.idle":"2021-08-06T01:56:17.185210Z","shell.execute_reply.started":"2021-08-06T01:56:16.444667Z","shell.execute_reply":"2021-08-06T01:56:17.184039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Training","metadata":{}},{"cell_type":"markdown","source":"#### Use W&B","metadata":{}},{"cell_type":"markdown","source":"```\n--img {IMG_SIZE} \\ # Input image size.\n--batch {BATCH_SIZE} \\ # Batch size\n--epochs {EPOCHS} \\ # Number of epochs\n--data data.yaml \\ # Configuration file\n--weights yolov5s.pt \\ # Model name\n--save_period 1\\ # Save model after interval\n--project kaggle-siim-covid # W&B project name\n```","metadata":{}},{"cell_type":"code","source":"# 学習実行\n\"\"\"\n!python train.py --img {IMG_SIZE} \\\n                 --batch {BATCH_SIZE} \\\n                 --epochs {EPOCHS} \\\n                 --data data.yaml \\\n                 --weights yolov5s.pt \\\n                 --save_period 1\\\n                 --project kaggle-siim-covid\n\"\"\"","metadata":{"execution":{"iopub.status.busy":"2021-08-06T02:03:40.361044Z","iopub.execute_input":"2021-08-06T02:03:40.361380Z","iopub.status.idle":"2021-08-06T02:03:40.367486Z","shell.execute_reply.started":"2021-08-06T02:03:40.361349Z","shell.execute_reply":"2021-08-06T02:03:40.366457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Inference","metadata":{}},{"cell_type":"code","source":"TEST_PATH = '/kaggle/input/siim-covid19-resized-to-256px-jpg/test/' # absolute path\nMODEL_PATH = 'kaggle-siim-covid/exp/weights/best.pt'\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"```\n--weights {MODEL_PATH} \\ # path to the best model.\n--source {TEST_PATH} \\ # absolute path to the test images.\n--img {IMG_SIZE} \\ # Size of image\n--conf 0.281 \\ # Confidence threshold (default is 0.25)\n--iou-thres 0.5 \\ # IOU threshold (default is 0.45)\n--max-det 3 \\ # Number of detections per image (default is 1000) \n--save-txt \\ # Save predicted bounding box coordinates as txt files\n--save-conf # Save the confidence of prediction for each bounding box\n```","metadata":{}},{"cell_type":"code","source":"\"\"\"\n!python detect.py --weights {MODEL_PATH} \\\n                  --source {TEST_PATH} \\\n                  --img {IMG_SIZE} \\\n                  --conf 0.281 \\\n                  --iou-thres 0.5 \\\n                  --max-det 3 \\\n                  --save-txt \\\n                  --save-conf\n\"\"\"","metadata":{"execution":{"iopub.status.busy":"2021-08-06T02:05:19.842735Z","iopub.execute_input":"2021-08-06T02:05:19.843086Z","iopub.status.idle":"2021-08-06T02:05:19.850212Z","shell.execute_reply.started":"2021-08-06T02:05:19.843054Z","shell.execute_reply":"2021-08-06T02:05:19.849432Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## submit","metadata":{}},{"cell_type":"code","source":"submission_df = pd.read_csv(dataset_path/'sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2021-08-05T02:57:25.445906Z","iopub.execute_input":"2021-08-05T02:57:25.446585Z","iopub.status.idle":"2021-08-05T02:57:25.459822Z","shell.execute_reply.started":"2021-08-05T02:57:25.446549Z","shell.execute_reply":"2021-08-05T02:57:25.458996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df.head()","metadata":{"execution":{"iopub.status.busy":"2021-08-05T02:57:27.649805Z","iopub.execute_input":"2021-08-05T02:57:27.650201Z","iopub.status.idle":"2021-08-05T02:57:27.661932Z","shell.execute_reply.started":"2021-08-05T02:57:27.650168Z","shell.execute_reply":"2021-08-05T02:57:27.660898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df.iloc[2000:2010]","metadata":{"execution":{"iopub.status.busy":"2021-08-05T02:43:13.170243Z","iopub.execute_input":"2021-08-05T02:43:13.170603Z","iopub.status.idle":"2021-08-05T02:43:13.181874Z","shell.execute_reply.started":"2021-08-05T02:43:13.170571Z","shell.execute_reply":"2021-08-05T02:43:13.180829Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2021-08-05T02:43:46.403823Z","iopub.execute_input":"2021-08-05T02:43:46.404188Z","iopub.status.idle":"2021-08-05T02:43:46.418944Z","shell.execute_reply.started":"2021-08-05T02:43:46.404151Z","shell.execute_reply":"2021-08-05T02:43:46.417942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}