{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Thanks\nThanks to @Draconda, we found that the prior mask information is very effective for this task. On the one hand, we can use mask to directly regress values of yaw, pitch, roll, x, y, z, and mask as what @Draconda told. On the other hand, we can also concanate the mask information with the original image and send them to any network structure you currently design for prediction. Here, we share a simpler way to get accurate masks using detectron2"},{"metadata":{},"cell_type":"markdown","source":"## Requirements for Detectron2\n* Python ≥ 3.6\n* PyTorch ≥ 1.3\n* torchvision that matches the PyTorch installation. You can install them together at pytorch.org to make sure of this.\n* OpenCV, optional, needed by demo and visualization\n* pycocotools: pip install cython; pip install 'git+https://github.com/cocodataset/cocoapi.git#subdirectory=PythonAPI'\n* gcc & g++ ≥ 4.9"},{"metadata":{"trusted":false},"cell_type":"code","source":"pip install 'git+https://github.com/cocodataset/cocoapi.git#subdirectory=PythonAPI'","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":false},"cell_type":"code","source":"pip install 'git+https://github.com/facebookresearch/detectron2.git'","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Visualization of masks"},{"metadata":{"trusted":false},"cell_type":"code","source":"import torch\n\nfrom detectron2.config import get_cfg\nfrom detectron2.engine import DefaultPredictor\n\nif torch.cuda.is_available():\n    device = torch.device(\"cuda:{}\".format(0))\nelse:\n    device = torch.device(\"cpu\")\n\nprint(\"-> Loading model\")\ncfg = get_cfg()\ncfg.merge_from_file(\"../input/detectron2/configs/COCO-InstanceSegmentation/mask_rcnn_R_50_FPN_3x.yaml\")\n\ncfg.MODEL.DEVICE = str(device)\ncfg.MODEL.RPN.NMS_THRESH = 0.1\ncfg.MODEL.ROI_HEADS.SCORE_THRESH_TEST = 0.5\n\ncfg.MODEL.WEIGHTS = \"../input/parameters/model_final_f10217.pkl\"\n\nmodel = DefaultPredictor(cfg)","execution_count":null,"outputs":[]},{"metadata":{"trusted":false},"cell_type":"code","source":"import PIL.Image as Image\n\nfrom torchvision import transforms\n\ndefault_transform = transforms.Compose([transforms.ToTensor()])\n\ndef load_image(path, transform=default_transform):\n    image = Image.open(path)\n    return transform(image)","execution_count":null,"outputs":[]},{"metadata":{"trusted":false},"cell_type":"code","source":"image_path = '../input/pku-autonomous-driving/train_images/ID_7f6f07350.jpg'","execution_count":null,"outputs":[]},{"metadata":{"trusted":false},"cell_type":"code","source":"import cv2\n\nimage = cv2.imread(image_path)\noutputs = model(image)","execution_count":null,"outputs":[]},{"metadata":{"trusted":false},"cell_type":"code","source":"from detectron2.data import MetadataCatalog\nfrom matplotlib import pyplot as plt\n\nfrom detectron2.utils.visualizer import ColorMode\nfrom detectron2.utils.visualizer import Visualizer\n\nv = Visualizer(image[:, :, ::-1], metadata=MetadataCatalog.get(cfg.DATASETS.TRAIN[0]), scale=0.8, instance_mode=ColorMode.IMAGE_BW)\nv = v.draw_instance_predictions(outputs[\"instances\"].to(\"cpu\"))\nv = v.get_image()[:, :, ::-1]\n\nplt.imshow(v)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Only need one-channel mask\nSome kagglers find that the npy file is large. This is because the shape of the mask I store is the number of instances * image size. If you only care about the binary result of instance-background, you can do max pooling:"},{"metadata":{"trusted":false},"cell_type":"code","source":"mask = outputs[\"instances\"].pred_masks.sum(0) > 0","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Only need instances\nSome kaggler only want to use the detected instances as input to the model (without concatenating it with the original image), you can use opencv to crop:"},{"metadata":{"trusted":false},"cell_type":"code","source":"import numpy as np\n\nmask = torch.stack([mask, mask, mask], dim=2)\nmask = mask.cpu().numpy().astype(\"uint8\")\n\ninstances = cv2.multiply(image, mask)\nplt.imshow(instances)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Make mask predictions"},{"metadata":{"trusted":false},"cell_type":"code","source":"import os\nimport cv2\nimport pdb\nimport glob\nimport argparse\n\nimport numpy as np","execution_count":null,"outputs":[]},{"metadata":{"trusted":false},"cell_type":"code","source":"def make_multi_channel_masks(source_dir='../input/pku-autonomous-driving/train_images',\n                dist_dir='../input/pku-autonomous-driving/train_images_mask',\n                ext='jpg'):\n    \"\"\"Function to predict for a single image or folder of images\n    \"\"\"\n\n    # FINDING INPUT IMAGES\n    if os.path.isdir(source_dir):\n        # Searching folder for images\n        paths = glob.glob(os.path.join(source_dir, '*.{}'.format(ext)))\n        output_directory = dist_dir\n    else:\n        raise Exception(\"Can not find source_dir: {}\".format(source_dir))\n\n    if not os.path.exists(output_directory):\n        os.makedirs(output_directory)\n        \n    print(\"-> Predicting on {:d} test images\".format(len(paths)))\n\n    for idx, image_path in enumerate(paths):\n        image = cv2.imread(image_path)\n        outputs = model(image)\n\n        output_name = os.path.splitext(os.path.basename(image_path))[0]\n        name_dest_npy = os.path.join(output_directory, \"{}.npy\".format(output_name))\n        mask = outputs['instances'].pred_masks.cpu().numpy()\n        np.save(name_dest_npy, mask)\n\n        print(\"   Processed {:d} of {:d} images - saved prediction to {}\".format(\n                idx + 1, len(paths), name_dest_npy))\n\n    print('-> Done!')","execution_count":null,"outputs":[]},{"metadata":{"trusted":false},"cell_type":"code","source":"def make_single_channel_masks(source_dir='../input/pku-autonomous-driving/train_images',\n                dist_dir='../input/pku-autonomous-driving/train_images_mask',\n                ext='jpg'):\n    \"\"\"Function to predict for a single image or folder of images\n    \"\"\"\n\n    # FINDING INPUT IMAGES\n    if os.path.isdir(source_dir):\n        # Searching folder for images\n        paths = glob.glob(os.path.join(source_dir, '*.{}'.format(ext)))\n        output_directory = dist_dir\n    else:\n        raise Exception(\"Can not find source_dir: {}\".format(source_dir))\n\n    if not os.path.exists(output_directory):\n        os.makedirs(output_directory)\n        \n    print(\"-> Predicting on {:d} test images\".format(len(paths)))\n\n    for idx, image_path in enumerate(paths):\n        image = cv2.imread(image_path)\n        outputs = model(image)\n\n        output_name = os.path.splitext(os.path.basename(image_path))[0]\n        name_dest_npy = os.path.join(output_directory, \"{}.npy\".format(output_name))\n        mask = outputs[\"instances\"].pred_masks.sum(0) > 0\n        mask = mask.float().unsqueeze(0)\n        mask = mask.cpu().numpy()\n        np.save(name_dest_npy, mask)\n\n        print(\"   Processed {:d} of {:d} images - saved prediction to {}\".format(\n                idx + 1, len(paths), name_dest_npy))\n\n    print('-> Done!')","execution_count":null,"outputs":[]},{"metadata":{"trusted":false},"cell_type":"code","source":"def make_instances(source_dir='../input/pku-autonomous-driving/train_images',\n                dist_dir='../input/pku-autonomous-driving/train_images_mask',\n                ext='jpg'):\n    \"\"\"Function to predict for a single image or folder of images\n    \"\"\"\n\n    # FINDING INPUT IMAGES\n    if os.path.isdir(source_dir):\n        # Searching folder for images\n        paths = glob.glob(os.path.join(source_dir, '*.{}'.format(ext)))\n        output_directory = dist_dir\n    else:\n        raise Exception(\"Can not find source_dir: {}\".format(source_dir))\n\n    if not os.path.exists(output_directory):\n        os.makedirs(output_directory)\n        \n    print(\"-> Predicting on {:d} test images\".format(len(paths)))\n\n    for idx, image_path in enumerate(paths):\n        image = cv2.imread(image_path)\n        outputs = model(image)\n\n        output_name = os.path.splitext(os.path.basename(image_path))[0]\n        name_dest_jpg = os.path.join(output_directory, \"{}.jpg\".format(output_name))\n        mask = outputs[\"instances\"].pred_masks.sum(0) > 0\n        mask = torch.stack([mask, mask, mask], dim=2)\n        mask = mask.cpu().numpy().astype(\"uint8\")\n\n        instances = cv2.multiply(image, mask)\n        cv2.imwrite(name_dest_jpg, instances)\n\n        print(\"   Processed {:d} of {:d} images - saved prediction to {}\".format(\n                idx + 1, len(paths), name_dest_jpg))\n\n    print('-> Done!')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Choose a generator based on your needs:"},{"metadata":{"trusted":false},"cell_type":"code","source":"# make_multi_channel_masks()\n# make_single_channel_masks()\n# make_instances()","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.7.2"}},"nbformat":4,"nbformat_minor":1}