{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **SIIM COVID-19 Detectron2 Training**","metadata":{}},{"cell_type":"markdown","source":"## Inferance ,EDA and Dataset \n- [SIIM COVID-19 Detectron2 Inferance](https://www.kaggle.com/ammarnassanalhajali/siim-covid-19-detectron2-inferance)\n- [SIIM-FISABIO-RSNA COVID-19 Detection-EDA](https://www.kaggle.com/ammarnassanalhajali/siim-fisabio-rsna-covid-19-detection-eda)\n- [SIIM-COVID-19 Detection Training Labels (Dataset)](https://www.kaggle.com/ammarnassanalhajali/siimcovid19-detection-training-label)\n","metadata":{}},{"cell_type":"markdown","source":"### Hi kagglers, This is `training` notebook using `Detectron2`.\n\n> #### Thanks:\n> - https://www.kaggle.com/xhlulu/siim-covid19-resized-to-256px-jpg\n\n\n### Please if this kernel is useful, <font color='red'>please upvote !!</font>","metadata":{}},{"cell_type":"markdown","source":"# Detectron2\nDetectron2 is Facebook AI Research's next generation software system that implements state-of-the-art object detection algorithms. It is a ground-up rewrite of the previous version, Detectron, and it originates from maskrcnn-benchmark.","metadata":{}},{"cell_type":"markdown","source":"# Installation\n* detectron2 is not pre-installed in this kaggle docker, so let's install it.\n* we need to know CUDA and pytorch version to install correct detectron2.","metadata":{}},{"cell_type":"code","source":"!nvidia-smi","metadata":{"execution":{"iopub.status.busy":"2021-06-22T04:30:56.49372Z","iopub.execute_input":"2021-06-22T04:30:56.494165Z","iopub.status.idle":"2021-06-22T04:30:57.238144Z","shell.execute_reply.started":"2021-06-22T04:30:56.494031Z","shell.execute_reply":"2021-06-22T04:30:57.237126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!nvcc --version","metadata":{"execution":{"iopub.status.busy":"2021-06-22T04:30:57.242643Z","iopub.execute_input":"2021-06-22T04:30:57.243318Z","iopub.status.idle":"2021-06-22T04:30:57.902528Z","shell.execute_reply.started":"2021-06-22T04:30:57.243274Z","shell.execute_reply":"2021-06-22T04:30:57.901543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import torch, torchvision\nprint(torch.__version__, torch.cuda.is_available())","metadata":{"execution":{"iopub.status.busy":"2021-07-25T19:26:28.795927Z","iopub.execute_input":"2021-07-25T19:26:28.796284Z","iopub.status.idle":"2021-07-25T19:26:30.183923Z","shell.execute_reply.started":"2021-07-25T19:26:28.796189Z","shell.execute_reply":"2021-07-25T19:26:30.182994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* It seems CUDA=11.0 and torch==1.7.0 is used in this kaggle docker image.\n* See installation for details. https://detectron2.readthedocs.io/en/latest/tutorials/install.html","metadata":{}},{"cell_type":"markdown","source":"# Install Pre-Built Detectron2","metadata":{}},{"cell_type":"code","source":"!pip install detectron2 -f \\\n  https://dl.fbaipublicfiles.com/detectron2/wheels/cu110/torch1.7/index.html","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-07-25T19:26:36.158334Z","iopub.execute_input":"2021-07-25T19:26:36.158670Z","iopub.status.idle":"2021-07-25T19:27:15.788797Z","shell.execute_reply.started":"2021-07-25T19:26:36.158640Z","shell.execute_reply":"2021-07-25T19:27:15.787869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Import Libraries","metadata":{}},{"cell_type":"code","source":"\nimport numpy as np \nimport pandas as pd \nfrom datetime import datetime\nimport time\nfrom tqdm import tqdm_notebook as tqdm # progress bar\nimport matplotlib.pyplot as plt\n\nimport os, json, cv2, random\nimport skimage.io as io\nimport copy\nimport pickle\nfrom pathlib import Path\nfrom typing import Optional\nfrom tqdm import tqdm\n\n# torch\nimport torch\n\n\n\n# Albumenatations\nimport albumentations as A\nfrom albumentations.pytorch.transforms import ToTensorV2\n\n#from pycocotools.coco import COCO\nfrom sklearn.model_selection import StratifiedKFold\n\n# glob\nfrom glob import glob\n\n# numba\nimport numba\nfrom numba import jit\n\nimport warnings\nwarnings.filterwarnings('ignore') #Ignore \"future\" warnings and Data-Frame-Slicing warnings.\n\n\n# detectron2\nfrom detectron2.structures import BoxMode\nfrom detectron2 import model_zoo\nfrom detectron2.config import get_cfg\nfrom detectron2.data import DatasetCatalog, MetadataCatalog\nfrom detectron2.engine import DefaultPredictor, DefaultTrainer, launch\nfrom detectron2.evaluation import COCOEvaluator\nfrom detectron2.structures import BoxMode\nfrom detectron2.utils.visualizer import ColorMode\nfrom detectron2.utils.logger import setup_logger\nfrom detectron2.utils.visualizer import Visualizer\n\nfrom detectron2.data import DatasetCatalog, MetadataCatalog, build_detection_test_loader, build_detection_train_loader\nfrom detectron2.data import detection_utils as utils\n\n\nfrom detectron2.data import DatasetCatalog, MetadataCatalog, build_detection_test_loader, build_detection_train_loader\nfrom detectron2.data import detection_utils as utils\nimport detectron2.data.transforms as T\nfrom detectron2.evaluation import COCOEvaluator, inference_on_dataset\n\nsetup_logger()","metadata":{"execution":{"iopub.status.busy":"2021-07-25T19:27:21.059146Z","iopub.execute_input":"2021-07-25T19:27:21.059515Z","iopub.status.idle":"2021-07-25T19:27:23.910655Z","shell.execute_reply.started":"2021-07-25T19:27:21.059481Z","shell.execute_reply":"2021-07-25T19:27:23.909877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Loading","metadata":{}},{"cell_type":"code","source":"# --- Read data ---\nimgdir = \"../input/siim-covid19-resized-1024px\"\n# Read in the data CSV files\ntrain_df = pd.read_csv(\"../input/siimcovid19-detection-training-label/train_image_df.csv\")\nprint(train_df['integer_label'].value_counts())\nlen(train_df)\n\ntrain_df=train_df[train_df['integer_label']!=2]\nprint(train_df['integer_label'].value_counts())\n\ntrain_df.loc[train_df['integer_label'] ==2, 'integer_label'] = 0\ntrain_df.loc[train_df['integer_label'] ==1, 'integer_label'] = 0\ntrain_df.loc[train_df['integer_label'] ==3, 'integer_label'] = 1\nprint(train_df['integer_label'].value_counts())","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-07-25T19:27:33.827203Z","iopub.execute_input":"2021-07-25T19:27:33.827562Z","iopub.status.idle":"2021-07-25T19:27:33.957850Z","shell.execute_reply.started":"2021-07-25T19:27:33.827530Z","shell.execute_reply":"2021-07-25T19:27:33.957038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# configs","metadata":{}},{"cell_type":"code","source":"# --- configs ---\nthing_classes = [\n    \"atypical\",\n    \"typical\"\n]\n\ndebug=False\n#split_mode=\"all_train\" # Or  valid20\nsplit_mode=\"valid20\"\n\n\ncategory_name_to_id = {class_name: index for index, class_name in enumerate(thing_classes)}\ncategory_name_to_id","metadata":{"execution":{"iopub.status.busy":"2021-07-25T19:27:37.588520Z","iopub.execute_input":"2021-07-25T19:27:37.588842Z","iopub.status.idle":"2021-07-25T19:27:37.596451Z","shell.execute_reply.started":"2021-07-25T19:27:37.588812Z","shell.execute_reply":"2021-07-25T19:27:37.594691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data preparation\n* `detectron2` provides high-level API for training custom dataset.\n\nTo define custom dataset, we need to create **list of dict** (`dataset_dicts`) where each dict contains following:\n\n - file_name: file name of the image.\n - image_id: id of the image, index is used here.\n - height: height of the image.\n - width: width of the image.\n - annotation: This is the ground truth annotation data for object detection, which contains following\n     - bbox: bounding box pixel location with shape (n_boxes, 4)\n     - bbox_mode: `BoxMode.XYXY_ABS` is used here, meaning that absolute value of (x_min, y_min, x_max, y_max) annotation is used in the `bbox`.\n     - category_id: class label id for each bounding box, with shape (n_boxes,)\n\n`get_COVID19_data_dicts` is for train dataset preparation and `get_COVID19_data_dicts_test` is for test dataset preparation.","metadata":{}},{"cell_type":"code","source":"from glob import glob\n\ndef get_COVID19_data_dicts(\n    imgdir: Path,\n    train_df: pd.DataFrame,\n    use_cache: bool = True,\n    target_indices: Optional[np.ndarray] = None,\n    debug: bool = False,\n    data_type:str=\"train\"\n   \n):\n\n    cache_path = Path(\".\") / f\"dataset_dicts_cache_{data_type}.pkl\"\n    if not use_cache or not cache_path.exists():\n        print(\"Creating data...\")\n        df_meta = pd.read_csv(\"../input/siim-covid19-resized-1024px/meta.csv\")\n        train_meta=df_meta[df_meta.split==\"train\"]\n        if debug:\n            train_meta = train_meta.iloc[:100]  # For debug....\n\n        # Load 1 image to get image size.\n        image_id = train_meta.iloc[0,0]\n        #image_path = str(imgdir / \"train\" / f\"{image_id}.jpg\")\n        image_path = str(f'../input/siim-covid19-resized-1024px/train/{image_id}.jpg')\n        image = cv2.imread(image_path)\n        resized_height, resized_width, ch = image.shape\n        print(f\"image shape: {image.shape}\")\n\n        dataset_dicts = []\n        for index, train_meta_row in tqdm(train_meta.iterrows(), total=len(train_meta)):\n            record = {}\n            image_id, height, width,s = train_meta_row.values\n            #filename = str(imgdir / \"train\" / f\"{image_id}.jpg\")\n            filename = str(f'../input/siim-covid19-resized-1024px/train/{image_id}.jpg')\n            record[\"file_name\"] = filename\n            record[\"image_id\"] = image_id\n            record[\"height\"] = height\n            record[\"width\"] = width\n            #record[\"height\"] = resized_height\n            #record[\"width\"] = resized_width\n            objs = []\n            for index2, row in train_df.query(\"id == @image_id\").iterrows():\n                # print(row)\n                # print(row[\"class_name\"])\n                # class_name = row[\"class_name\"]\n                class_id = row[\"integer_label\"]\n                if class_id == 2: # NO class\n                    # It is \"No finding\"\n \n                    # Use this No finding class with the bbox covering all image area.\n                    #bbox_resized = [0, 0, resized_width, resized_height]\n                    bbox_resized = [50, 50, 200, 200]\n                    obj = {\n                        \"bbox\": bbox_resized,\n                        \"bbox_mode\": BoxMode.XYXY_ABS,\n                        \"category_id\": class_id,\n                    }\n                    #objs.append(obj)\n\n                else:\n                    # bbox_original = [int(row[\"x_min\"]), int(row[\"y_min\"]), int(row[\"x_max\"]), int(row[\"y_max\"])]\n                    h_ratio = resized_height / height\n                    w_ratio = resized_width / width\n                    bbox_resized = [\n                        float(row[\"x_min\"]) * w_ratio,\n                        float(row[\"y_min\"]) * h_ratio,\n                        float(row[\"x_max\"]) * w_ratio,\n                        float(row[\"y_max\"]) * h_ratio,\n                    ]\n                    obj = {\n                        \"bbox\": bbox_resized,\n                        \"bbox_mode\": BoxMode.XYXY_ABS,\n                        \"category_id\": class_id,\n                    }\n                    objs.append(obj)\n            record[\"annotations\"] = objs\n            dataset_dicts.append(record)\n        with open(cache_path, mode=\"wb\") as f:\n            pickle.dump(dataset_dicts, f)\n\n    print(f\"Load from cache {cache_path}\")\n    with open(cache_path, mode=\"rb\") as f:\n        dataset_dicts = pickle.load(f)\n    if target_indices is not None:\n        dataset_dicts = [dataset_dicts[i] for i in target_indices]\n    return dataset_dicts\n\n\ndef get_COVID19_data_dicts_test(\n    imgdir: Path, test_meta: pd.DataFrame, use_cache: bool = True, debug: bool = False,\n):\n    debug_str = f\"_debug{int(debug)}\"\n    cache_path = Path(\".\") / f\"dataset_dicts_cache_test.pkl\"\n    if not use_cache or not cache_path.exists():\n        print(\"Creating data...\")\n        # test_meta = pd.read_csv(imgdir / \"test_meta.csv\")\n        df_meta = pd.read_csv(\"../input/siim-covid19-resized-1024px/meta.csv\")\n        test_meta=df_meta[df_meta.split==\"test\"]\n        if debug:\n            test_meta = test_meta.iloc[:100]  # For debug....\n        # Load 1 image to get image size.\n        image_id = test_meta.iloc[0,0]\n        #image_path = str(imgdir / \"test\" / f\"{image_id}.jpg\")\n        image_path = str(f'../input/siim-covid19-resized-1024px/test/{image_id}.jpg')\n        image = cv2.imread(image_path)\n        resized_height, resized_width, ch = image.shape\n        #print(f\"image shape: {image.shape}\")\n\n        dataset_dicts = []\n        for index, test_meta_row in tqdm(test_meta.iterrows(), total=len(test_meta)):\n            record = {}\n\n            image_id, height, width,s = test_meta_row.values\n            #filename = str(imgdir / \"test\" / f\"{image_id}.jpg\")\n            filename = str(f'../input/siim-covid19-resized-1024px/test/{image_id}.jpg')\n            record[\"file_name\"] = filename\n            # record[\"image_id\"] = index\n            record[\"image_id\"] = image_id\n            record[\"height\"] = height\n            record[\"width\"] = width\n            #record[\"height\"] = resized_height\n            #record[\"width\"] = resized_width\n            # objs = []\n            # record[\"annotations\"] = objs\n            dataset_dicts.append(record)\n        with open(cache_path, mode=\"wb\") as f:\n            pickle.dump(dataset_dicts, f)\n\n    #print(f\"Load from cache {cache_path}\")\n    with open(cache_path, mode=\"rb\") as f:\n        dataset_dicts = pickle.load(f)\n    return dataset_dicts","metadata":{"execution":{"iopub.status.busy":"2021-07-25T19:27:43.960099Z","iopub.execute_input":"2021-07-25T19:27:43.960465Z","iopub.status.idle":"2021-07-25T19:27:43.982655Z","shell.execute_reply.started":"2021-07-25T19:27:43.960434Z","shell.execute_reply":"2021-07-25T19:27:43.981700Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"if split_mode == \"all_train\":\n    DatasetCatalog.register(\n        \"COVID19_data_train\",\n        lambda: get_COVID19_data_dicts(\n            imgdir,\n            train_df,\n            debug=debug,\n            data_type=\"train\"\n        ),\n    )\n    MetadataCatalog.get(\"COVID19_data_train\").set(thing_classes=thing_classes)\n    \n    \n    dataset_dicts_train = DatasetCatalog.get(\"COVID19_data_train\")\n    metadata_dicts_train = MetadataCatalog.get(\"COVID19_data_train\")\n    \n    \nelif split_mode == \"valid20\":\n\n    n_dataset = len(\n        get_COVID19_data_dicts(\n            imgdir, train_df, debug=debug,data_type=\"All\"\n        )\n    )\n    n_train = int(n_dataset * 0.80)\n    print(\"n_dataset\", n_dataset, \"n_train\", n_train)\n    rs = np.random.RandomState(42)\n    inds = rs.permutation(n_dataset)\n    train_inds, valid_inds = inds[:n_train], inds[n_train:]\n\n    DatasetCatalog.register(\n        \"COVID19_data_train\",\n        lambda: get_COVID19_data_dicts(\n            imgdir,\n            train_df,\n            target_indices=train_inds,\n            debug=debug,\n            data_type=\"train\"\n        ),\n    )\n    MetadataCatalog.get(\"COVID19_data_train\").set(thing_classes=thing_classes)\n    \n\n    DatasetCatalog.register(\n        \"COVID19_data_valid\",\n        lambda: get_COVID19_data_dicts(\n            imgdir,\n            train_df,\n            target_indices=valid_inds,\n            debug=debug,\n            data_type=\"val\"\n            ),\n        )\n    MetadataCatalog.get(\"COVID19_data_valid\").set(thing_classes=thing_classes)\n    \n    dataset_dicts_train = DatasetCatalog.get(\"COVID19_data_train\")\n    metadata_dicts_train = MetadataCatalog.get(\"COVID19_data_train\")\n\n    dataset_dicts_valid = DatasetCatalog.get(\"COVID19_data_valid\")\n    metadata_dicts_valid = MetadataCatalog.get(\"COVID19_data_valid\")\n    \nelse:\n    raise ValueError(f\"[ERROR] Unexpected value split_mode={split_mode}\")","metadata":{"execution":{"iopub.status.busy":"2021-07-25T19:27:51.395892Z","iopub.execute_input":"2021-07-25T19:27:51.396279Z","iopub.status.idle":"2021-07-25T19:28:46.030707Z","shell.execute_reply.started":"2021-07-25T19:27:51.396214Z","shell.execute_reply":"2021-07-25T19:28:46.029986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"data_vis\"></a>\n# Data Visualization\n\nIt's also very easy to visualize prepared training dataset with `detectron2`.<br/>\nIt provides `Visualizer` class, we can use it to draw an image with bounding box as following.","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(1, 1, figsize =(10,10))\nindices=[ax[0][0]]\ni=-1\nfor d in random.sample(dataset_dicts_train, 1):\n    i=i+1    \n    img = cv2.imread(d[\"file_name\"])\n    v = Visualizer(img[:, :, ::-1],\n                   metadata=metadata_dicts_train, \n                   scale=0.3, \n                   instance_mode=ColorMode.IMAGE_BW   # remove the colors of unsegmented pixels. This option is only available for segmentation models\n    )\n    out = v.draw_dataset_dict(d)\n    indices[i].grid(False)\n    indices[i].axis('off')\n    indices[i].imshow(out.get_image()[:, :, ::-1])\n    \nfig.savefig('sample1.png')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-07-25T13:28:46.735467Z","iopub.execute_input":"2021-07-25T13:28:46.735910Z","iopub.status.idle":"2021-07-25T13:28:46.943740Z","shell.execute_reply.started":"2021-07-25T13:28:46.735879Z","shell.execute_reply":"2021-07-25T13:28:46.941309Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nfor d in random.sample(dataset_dicts_train, 1):\n    img = cv2.imread(d[\"file_name\"])\n    visualizer = Visualizer(img[:, :, ::-1], metadata_dicts_train,scale=1.0)\n    vis = visualizer.draw_dataset_dict(d)\n    plt.imshow(vis.get_image()[:, :, ::-1])\n    ","metadata":{"execution":{"iopub.status.busy":"2021-07-25T19:29:01.273961Z","iopub.execute_input":"2021-07-25T19:29:01.274333Z","iopub.status.idle":"2021-07-25T19:29:02.945300Z","shell.execute_reply.started":"2021-07-25T19:29:01.274300Z","shell.execute_reply":"2021-07-25T19:29:02.944491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for d in random.sample(dataset_dicts_train, 1):\n    img = cv2.imread(d[\"file_name\"])\n    v = Visualizer(img[:, :, ::-1], metadata_dicts_train,scale=1.0)\n    out = v.draw_instance_predictions(outputs[\"instances\"].to(\"cpu\"))\n    boxes = v._convert_boxes(outputs[\"instances\"].pred_boxes.to('cpu')).squeeze()\n    for box in boxes:\n        out = v.draw_text(f\"{box}\", (box[0], box[1]),font_size=10)\nplt.imshow(out.get_image()[:, :, ::-1])","metadata":{"execution":{"iopub.status.busy":"2021-07-25T13:29:06.789618Z","iopub.execute_input":"2021-07-25T13:29:06.790087Z","iopub.status.idle":"2021-07-25T13:29:06.923470Z","shell.execute_reply.started":"2021-07-25T13:29:06.790055Z","shell.execute_reply":"2021-07-25T13:29:06.921824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Augmentation\nThe dataset is transformed by changing the brighness and flipping the image with 50% probability.","metadata":{}},{"cell_type":"code","source":"def custom_mapper(dataset_dict):\n    \n    dataset_dict = copy.deepcopy(dataset_dict)\n    image = utils.read_image(dataset_dict[\"file_name\"], format=\"BGR\")\n    transform_list = [#T.Resize((1024,1024)),\n                      T.RandomBrightness(0.9, 1.1),\n                      T.RandomFlip(prob=0.5, horizontal=False, vertical=True),\n                      T.RandomFlip(prob=0.5, horizontal=True, vertical=False)\n                      ]\n    image, transforms = T.apply_transform_gens(transform_list, image)\n    dataset_dict[\"image\"] = torch.as_tensor(image.transpose(2, 0, 1).astype(\"float32\"))\n\n    annos = [\n        utils.transform_instance_annotations(obj, transforms, image.shape[:2])\n        for obj in dataset_dict.pop(\"annotations\")\n        if obj.get(\"iscrowd\", 0) == 0\n    ]\n    instances = utils.annotations_to_instances(annos, image.shape[:2])\n    dataset_dict[\"instances\"] = utils.filter_empty_instances(instances)\n    return dataset_dict\nclass AugTrainer(DefaultTrainer):\n    @classmethod\n    def build_train_loader(cls, cfg):\n        return build_detection_train_loader(cfg, mapper=custom_mapper)","metadata":{"execution":{"iopub.status.busy":"2021-07-25T19:29:10.544111Z","iopub.execute_input":"2021-07-25T19:29:10.544523Z","iopub.status.idle":"2021-07-25T19:29:10.554732Z","shell.execute_reply.started":"2021-07-25T19:29:10.544486Z","shell.execute_reply":"2021-07-25T19:29:10.553426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training","metadata":{}},{"cell_type":"code","source":"cfg = get_cfg()\nconfig_name = \"COCO-Detection/faster_rcnn_R_50_FPN_3x.yaml\" \n#config_name = \"COCO-Detection/faster_rcnn_X_101_32x8d_FPN_3x.yaml\"\n#config_name = \"COCO-Detection/faster_rcnn_R_101_C4_3x.yaml\"\n\ncfg.merge_from_file(model_zoo.get_config_file(config_name))\n\ncfg.DATASETS.TRAIN = (\"COVID19_data_train\",)\n\nif split_mode == \"all_train\":\n    cfg.DATASETS.TEST = ()\nelse:\n    cfg.DATASETS.TEST = (\"COVID19_data_valid\",)\n    cfg.TEST.EVAL_PERIOD = 1000\n\ncfg.DATALOADER.NUM_WORKERS = 2\n#cfg.MODEL.WEIGHTS = model_zoo.get_checkpoint_url(config_name)\ncfg.MODEL.WEIGHTS=\"../input/1siim-covid19-detectron2-weights/output/model_final.pth\"\n\n\ncfg.SOLVER.IMS_PER_BATCH = 4\ncfg.SOLVER.BASE_LR = 0.00025\n\ncfg.SOLVER.WARMUP_ITERS = 1000\ncfg.SOLVER.MAX_ITER = 8000 #adjust up if val mAP is still rising, adjust down if overfit\n#cfg.SOLVER.STEPS = (100, 500) # must be less than  MAX_ITER \n#cfg.SOLVER.GAMMA = 0.05\n\n\ncfg.SOLVER.CHECKPOINT_PERIOD = 100000  # Small value=Frequent save need a lot of storage.\ncfg.MODEL.ROI_HEADS.BATCH_SIZE_PER_IMAGE = 128\ncfg.MODEL.ROI_HEADS.NUM_CLASSES = 2\n\n\nos.makedirs(cfg.OUTPUT_DIR, exist_ok=True)\n\n\n#Training using custom trainer defined above\n#trainer = AugTrainer(cfg) \ntrainer = DefaultTrainer(cfg) \ntrainer.resume_or_load(resume=False)\ntrainer.train()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2021-07-25T19:29:15.450483Z","iopub.execute_input":"2021-07-25T19:29:15.450812Z","iopub.status.idle":"2021-07-25T22:11:33.780269Z","shell.execute_reply.started":"2021-07-25T19:29:15.450780Z","shell.execute_reply":"2021-07-25T22:11:33.779316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Evaluator","metadata":{}},{"cell_type":"markdown","source":"* Famouns dataset's evaluator is already implemented in detectron2.\n* For example, many kinds of AP (Average Precision) is calculted in COCOEvaluator.\n* COCOEvaluator only calculates AP with IoU from 0.50 to 0.95","metadata":{}},{"cell_type":"code","source":"evaluator = COCOEvaluator(\"COVID19_data_valid\", cfg, False, output_dir=\"./output/\")\nval_loader = build_detection_test_loader(cfg, \"COVID19_data_valid\")\ninference_on_dataset(trainer.model, val_loader, evaluator)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-07-25T22:13:12.936465Z","iopub.execute_input":"2021-07-25T22:13:12.936829Z","iopub.status.idle":"2021-07-25T22:19:10.260202Z","shell.execute_reply.started":"2021-07-25T22:13:12.936795Z","shell.execute_reply":"2021-07-25T22:19:10.259130Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictor = DefaultPredictor(cfg)","metadata":{"execution":{"iopub.status.busy":"2021-07-25T22:26:19.390134Z","iopub.execute_input":"2021-07-25T22:26:19.390518Z","iopub.status.idle":"2021-07-25T22:26:20.448897Z","shell.execute_reply.started":"2021-07-25T22:26:19.390480Z","shell.execute_reply":"2021-07-25T22:26:20.447941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('hi')","metadata":{"execution":{"iopub.status.busy":"2021-07-25T16:30:23.466057Z","iopub.execute_input":"2021-07-25T16:30:23.466439Z","iopub.status.idle":"2021-07-25T16:30:23.476905Z","shell.execute_reply.started":"2021-07-25T16:30:23.466407Z","shell.execute_reply":"2021-07-25T16:30:23.475638Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(dataset_dicts_valid[0:1])","metadata":{"execution":{"iopub.status.busy":"2021-07-25T18:35:32.604149Z","iopub.execute_input":"2021-07-25T18:35:32.604526Z","iopub.status.idle":"2021-07-25T18:35:32.615196Z","shell.execute_reply.started":"2021-07-25T18:35:32.604495Z","shell.execute_reply":"2021-07-25T18:35:32.611198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install opencv-contrib-python","metadata":{"execution":{"iopub.status.busy":"2021-07-25T18:48:50.062036Z","iopub.execute_input":"2021-07-25T18:48:50.062433Z","iopub.status.idle":"2021-07-25T18:49:05.357976Z","shell.execute_reply.started":"2021-07-25T18:48:50.062400Z","shell.execute_reply":"2021-07-25T18:49:05.356736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# labels visualize\n\nimport matplotlib.pyplot as plt\nfor d in random.sample(dataset_dicts_valid, 1):\n    img = cv2.imread(d[\"file_name\"])\n    boxes=d[\"annotations\"]\n    category=-1\n    for box in boxes:\n        category=box['category_id']\n    font = cv2.FONT_HERSHEY_SIMPLEX    \n    print(\"Class :\", category)\n    for box in boxes:\n            print(\"uff\")\n            print(box)\n            cv2.rectangle(img, (10, 10), (2000, 2000), (0, 255, 0), 5)\n            cv2.putText(img,\"atypical\",(1000, 1000),font,20,(0,0,255),4,cv2.LINE_AA)\n           \n\n    plt.imshow(img)","metadata":{"execution":{"iopub.status.busy":"2021-07-25T19:17:01.577044Z","iopub.execute_input":"2021-07-25T19:17:01.577430Z","iopub.status.idle":"2021-07-25T19:17:02.446618Z","shell.execute_reply.started":"2021-07-25T19:17:01.577396Z","shell.execute_reply":"2021-07-25T19:17:02.445515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"## Visualize predictions\n\nfor d in random.sample(dataset_dicts_valid, 1):\n    im = cv2.imread(d[\"file_name\"])\n    #im=cv2.resize(im,(1024,1024))\n    #print('Image read')\n    \n    boxes=d[\"annotations\"]\n    category=-1\n    for box in boxes:\n        print('box')\n        print(box)\n        category=box['category_id']\n        \n    print(\"Ground Truth class :\",category)\n\n    outputs = predictor(im)\n    instances=outputs[\"instances\"].to(\"cpu\")\n    instances=instances[0:3]\n    #print(instances)\n    scores=outputs['instances'].scores.to(\"cpu\")\n    scores=scores[0:3]\n    detect=outputs['instances'].pred_classes.to(\"cpu\")\n    detect=detect[0:3]\n    v = Visualizer(im[:, :, ::-1],metadata=metadata_dicts_valid, scale=0.8)\n    out = v.draw_instance_predictions(instances)\n    boxes = v._convert_boxes(instances.pred_boxes.to('cpu')).squeeze()\n    i=0\n    for box in boxes:\n            out = v.draw_text(f\"{detect[i]}\", (box[0], box[1]),font_size=150)\n        \n    output=out.get_image()[:, :, ::-1]\n    plt.imshow(output)\n\n    ","metadata":{"execution":{"iopub.status.busy":"2021-07-25T22:36:24.034277Z","iopub.execute_input":"2021-07-25T22:36:24.034649Z","iopub.status.idle":"2021-07-25T22:36:26.410299Z","shell.execute_reply.started":"2021-07-25T22:36:24.034616Z","shell.execute_reply":"2021-07-25T22:36:26.409465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"torch.save(trainer.model, 'checkpoint_binary.pth')","metadata":{"execution":{"iopub.status.busy":"2021-07-25T22:39:18.326104Z","iopub.execute_input":"2021-07-25T22:39:18.326455Z","iopub.status.idle":"2021-07-25T22:39:18.757170Z","shell.execute_reply.started":"2021-07-25T22:39:18.326424Z","shell.execute_reply":"2021-07-25T22:39:18.756113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"out = v.draw_instance_predictions(outputs[\"instances\"].to(\"cpu\"))\nboxes = v._convert_boxes(outputs[\"instances\"].pred_boxes.to('cpu')).squeeze()\nfor box in boxes:\n    out = v.draw_text(f\"{box}\", (box[0], box[1]))\ncv2_imshow(out.get_image()[:, :, ::-1])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nmetrics_df = pd.read_json(\"./output/metrics.json\", orient=\"records\", lines=True)\nmdf = metrics_df.sort_values(\"iteration\")\nmdf.head(10).T","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-07-25T22:40:06.034519Z","iopub.execute_input":"2021-07-25T22:40:06.034858Z","iopub.status.idle":"2021-07-25T22:40:06.222877Z","shell.execute_reply.started":"2021-07-25T22:40:06.034827Z","shell.execute_reply":"2021-07-25T22:40:06.222137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 1. Loss curve\nfig, ax = plt.subplots()\n\nmdf1 = mdf[~mdf[\"total_loss\"].isna()]\nax.plot(mdf1[\"iteration\"], mdf1[\"total_loss\"], c=\"C0\", label=\"train\")\nif \"validation_loss\" in mdf.columns:\n    mdf2 = mdf[~mdf[\"validation_loss\"].isna()]\n    ax.plot(mdf2[\"iteration\"], mdf2[\"validation_loss\"], c=\"C1\", label=\"validation\")\n\n# ax.set_ylim([0, 0.5])\nax.legend()\nax.set_title(\"Loss curve\")\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-07-25T22:40:15.535101Z","iopub.execute_input":"2021-07-25T22:40:15.535549Z","iopub.status.idle":"2021-07-25T22:40:15.797382Z","shell.execute_reply.started":"2021-07-25T22:40:15.535509Z","shell.execute_reply":"2021-07-25T22:40:15.796267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 1. Loss curve\nfig, ax = plt.subplots()\n\nmdf1 = mdf[~mdf[\"fast_rcnn/cls_accuracy\"].isna()]\nax.plot(mdf1[\"iteration\"], mdf1[\"fast_rcnn/cls_accuracy\"], c=\"C0\", label=\"train\")\n# ax.set_ylim([0, 0.5])\nax.legend()\nax.set_title(\"Accuracy curve\")\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2021-07-25T22:40:54.114886Z","iopub.execute_input":"2021-07-25T22:40:54.115212Z","iopub.status.idle":"2021-07-25T22:40:54.261743Z","shell.execute_reply.started":"2021-07-25T22:40:54.115181Z","shell.execute_reply":"2021-07-25T22:40:54.260770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots()\nmdf1 = mdf[~mdf[\"fast_rcnn/fg_cls_accuracy\"].isna()]\nax.plot(mdf1[\"iteration\"], mdf1[\"fast_rcnn/cls_accuracy\"], c=\"C0\", label=\"valid\")\n# ax.set_ylim([0, 0.5])\nax.legend()\nax.set_title(\"fast_rcnn/fg_cls_accuracy\t\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-07-25T22:42:50.909643Z","iopub.execute_input":"2021-07-25T22:42:50.909980Z","iopub.status.idle":"2021-07-25T22:42:51.059705Z","shell.execute_reply.started":"2021-07-25T22:42:50.909934Z","shell.execute_reply":"2021-07-25T22:42:51.058740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots()\nmdf1 = mdf[~mdf[\"fast_rcnn/false_negative\"].isna()]\nax.plot(mdf1[\"iteration\"], mdf1[\"fast_rcnn/cls_accuracy\"], c=\"C0\", label=\"valid\")\n# ax.set_ylim([0, 0.5])\nax.legend()\nax.set_title(\"fast_rcnn/false_negative\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2021-07-25T22:43:52.719820Z","iopub.execute_input":"2021-07-25T22:43:52.720190Z","iopub.status.idle":"2021-07-25T22:43:52.910545Z","shell.execute_reply.started":"2021-07-25T22:43:52.720159Z","shell.execute_reply":"2021-07-25T22:43:52.909552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# References\n1. https://www.kaggle.com/ammarnassanalhajali/training-detectron2-for-blood-cells-detection\n1. https://www.kaggle.com/corochann/vinbigdata-detectron2-train\n","metadata":{}}]}