{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        os.path.join(dirname, filename)\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-01-04T10:34:20.484321Z","iopub.execute_input":"2023-01-04T10:34:20.484996Z","iopub.status.idle":"2023-01-04T10:36:33.345409Z","shell.execute_reply.started":"2023-01-04T10:34:20.484866Z","shell.execute_reply":"2023-01-04T10:36:33.344111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In this notebook, we want to use tensorflow object detection models.\n\nFirst things first, we must know what these pre-trained models can detect. These models are trained on COCO 2017 dataset which has 91 classes. I collected these classes in a text file and added a dummy term named __Background__ to the beginning of the file just to make the model work properly. The following function reads and returns the names of these classes from the file.","metadata":{}},{"cell_type":"code","source":"def read_class_names(path):\n    with open(path, 'r') as f:\n        classes_name = f.read().splitlines()\n    \n    return classes_name","metadata":{"execution":{"iopub.status.busy":"2023-01-04T10:36:33.347491Z","iopub.execute_input":"2023-01-04T10:36:33.347852Z","iopub.status.idle":"2023-01-04T10:36:33.352888Z","shell.execute_reply.started":"2023-01-04T10:36:33.347814Z","shell.execute_reply":"2023-01-04T10:36:33.351827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"classes_name = read_class_names('/kaggle/input/coco2017-classes/coconames.txt')\nclasses_name[:10]","metadata":{"execution":{"iopub.status.busy":"2023-01-04T10:36:33.354658Z","iopub.execute_input":"2023-01-04T10:36:33.355043Z","iopub.status.idle":"2023-01-04T10:36:33.376800Z","shell.execute_reply.started":"2023-01-04T10:36:33.354999Z","shell.execute_reply":"2023-01-04T10:36:33.375947Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The following function get a model url, extract it, save and return it.","metadata":{}},{"cell_type":"code","source":"import tensorflow as tf\nfrom tensorflow.python.keras.utils.data_utils import get_file\n\ndef get_model(model_URL):\n    model_name = os.path.basename(model_URL).split('.')[0]\n    get_path = get_file(fname=model_name, untar=True, origin=model_URL)\n    model = tf.saved_model.load(os.path.join(get_path, 'saved_model'))\n    return model","metadata":{"execution":{"iopub.status.busy":"2023-01-04T10:36:33.380442Z","iopub.execute_input":"2023-01-04T10:36:33.380692Z","iopub.status.idle":"2023-01-04T10:36:42.523149Z","shell.execute_reply.started":"2023-01-04T10:36:33.380668Z","shell.execute_reply":"2023-01-04T10:36:42.522135Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here, I am going to download and use EfficientDet D4 model. For more models you can check:\n\n[https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/tf2_detection_zoo.md](http://)","metadata":{}},{"cell_type":"code","source":"url = 'http://download.tensorflow.org/models/object_detection/tf2/20200711/efficientdet_d4_coco17_tpu-32.tar.gz'\n\nmodel = get_model(url)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T10:36:42.524759Z","iopub.execute_input":"2023-01-04T10:36:42.525633Z","iopub.status.idle":"2023-01-04T10:37:38.287972Z","shell.execute_reply.started":"2023-01-04T10:36:42.525586Z","shell.execute_reply":"2023-01-04T10:37:38.286859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And now we come to the most important part; We want to define a function that receives the image path, the model and some other arguments as input and detects the objects in the image and draws a bounding box around them.\n\nFirst, we read the image using the CV2 library which returns the image as a numpy array but the input of pre-trained TensorFlow2 detection models must be a tensor; So we convert numpy array to tensor and expand its dimention.\n\nThe output of the model is a dictionary whose three keys are important for us; 'detection_boxes', 'detection classes' and 'detection scores'. all of them are tensor so it's better to convert them to numpy array for easy use. Also, we consider a separate color for each class so that the name of the class and the bouning box can be identified with that color.","metadata":{}},{"cell_type":"markdown","source":"Finally, to optimize the bounding boxes, we use soft non-max suppresion algorithm and return the detected objects on the image along with the image itself.","metadata":{}},{"cell_type":"code","source":"import cv2\nfrom matplotlib import pyplot as plt\n\ndef object_detection(image_path, model, classes_name, max_output_size=50, iou_threshold=0.4, score_threshold=0.7, soft_nms_sigma=0.4):\n    image = cv2.imread(image_path)\n    image2tensor = tf.convert_to_tensor(image, dtype=tf.uint8)\n    image2tensor = image2tensor[tf.newaxis,...]\n    \n    detection = model(image2tensor)\n    bboxes = detection['detection_boxes'].numpy()[0]\n    class_indexes = detection['detection_classes'].numpy().astype(np.int32)[0]\n    class_scores = detection['detection_scores'].numpy()[0]\n    \n    selected_indices, _ = tf.image.non_max_suppression_with_scores(bboxes, scores=class_scores, max_output_size=max_output_size, iou_threshold=iou_threshold, score_threshold=score_threshold, soft_nms_sigma=soft_nms_sigma)\n    \n    img_h, img_w, img_c = image.shape\n    \n    classes_color = np.random.uniform(low=0, high=255, size=(len(classes_name),3))\n    \n    for i in selected_indices:\n        class_score = class_scores[i]\n        class_index = class_indexes[i]\n\n        class_label = classes_name[class_index]\n        class_color = classes_color[class_index]\n\n        bbox = bboxes[i].tolist()\n        ymin, xmin, ymax, xmax = bbox\n        ymin, xmin, ymax, xmax = int(ymin*img_h), int(xmin*img_w), int(ymax*img_h), int(xmax*img_w)   \n        cv2.rectangle(img=image, pt1=(xmin,ymin), pt2=(xmax,ymax), color=class_color, thickness=2)\n\n        display_txt = '{}: {}%'.format(class_label, class_score)\n        cv2.putText(img=image, text=display_txt, org=(xmin,ymin - 10), fontFace=cv2.FONT_HERSHEY_PLAIN, fontScale=2, color=class_color, thickness=2)\n    \n    \n    plt.figure(figsize=(15,15))\n    plt.imshow(cv2.cvtColor(image, cv2.COLOR_BGR2RGB))\n    plt.title(\"Object Detection Results\")\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-04T10:37:38.292036Z","iopub.execute_input":"2023-01-04T10:37:38.292643Z","iopub.status.idle":"2023-01-04T10:37:38.751456Z","shell.execute_reply.started":"2023-01-04T10:37:38.292604Z","shell.execute_reply":"2023-01-04T10:37:38.746890Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we can select one of the images and detects the objects in the image with the help of the function we have defined.\n\nObviously, some parameters of the soft-NMS algorithm are tricky and we have to try to find the best values by trial and error several times.","metadata":{}},{"cell_type":"code","source":"image_path = '/kaggle/input/open-images-2019-object-detection/test/00000b4dcff7f799.jpg'\n\nobject_detection(image_path, model, classes_name, max_output_size=20, iou_threshold=0.49, score_threshold=0.5, soft_nms_sigma=0.49)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T11:02:35.257904Z","iopub.execute_input":"2023-01-04T11:02:35.258292Z","iopub.status.idle":"2023-01-04T11:02:36.838296Z","shell.execute_reply.started":"2023-01-04T11:02:35.258260Z","shell.execute_reply":"2023-01-04T11:02:36.837157Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_path = '/kaggle/input/open-images-2019-object-detection/test/00001a21632de752.jpg'\n\nobject_detection(image_path, model, classes_name, max_output_size=50, iou_threshold=0.49, score_threshold=0.3, soft_nms_sigma=0.4)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T11:02:05.167584Z","iopub.execute_input":"2023-01-04T11:02:05.167993Z","iopub.status.idle":"2023-01-04T11:02:06.749314Z","shell.execute_reply.started":"2023-01-04T11:02:05.167959Z","shell.execute_reply":"2023-01-04T11:02:06.747516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_path = '/kaggle/input/open-images-2019-object-detection/test/0003d1c3be9ed3d6.jpg'\n\nobject_detection(image_path, model, classes_name, max_output_size=50, iou_threshold=0.49, score_threshold=0.6, soft_nms_sigma=0.49)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T10:43:04.400501Z","iopub.execute_input":"2023-01-04T10:43:04.400924Z","iopub.status.idle":"2023-01-04T10:43:05.836048Z","shell.execute_reply.started":"2023-01-04T10:43:04.400891Z","shell.execute_reply":"2023-01-04T10:43:05.834817Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_path = '/kaggle/input/open-images-2019-object-detection/test/000f579533500448.jpg'\n\nobject_detection(image_path, model, classes_name, max_output_size=50, iou_threshold=0.49, score_threshold=0.6, soft_nms_sigma=0.49)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T10:45:45.868524Z","iopub.execute_input":"2023-01-04T10:45:45.869564Z","iopub.status.idle":"2023-01-04T10:45:47.567295Z","shell.execute_reply.started":"2023-01-04T10:45:45.869527Z","shell.execute_reply":"2023-01-04T10:45:47.564857Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_path = '/kaggle/input/open-images-2019-object-detection/test/0026201b6d912efb.jpg'\n\nobject_detection(image_path, model, classes_name, max_output_size=50, iou_threshold=0.49, score_threshold=0.3, soft_nms_sigma=0.59)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T10:55:37.180083Z","iopub.execute_input":"2023-01-04T10:55:37.180652Z","iopub.status.idle":"2023-01-04T10:55:38.511120Z","shell.execute_reply.started":"2023-01-04T10:55:37.180597Z","shell.execute_reply":"2023-01-04T10:55:38.510247Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_path = '/kaggle/input/open-images-2019-object-detection/test/00246c2940caa984.jpg'\n\nobject_detection(image_path, model, classes_name, max_output_size=50, iou_threshold=0.49, score_threshold=0.4, soft_nms_sigma=0.49)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T10:56:21.910855Z","iopub.execute_input":"2023-01-04T10:56:21.911552Z","iopub.status.idle":"2023-01-04T10:56:23.288396Z","shell.execute_reply.started":"2023-01-04T10:56:21.911516Z","shell.execute_reply":"2023-01-04T10:56:23.287492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_path = '/kaggle/input/open-images-2019-object-detection/test/001b5a3eae515477.jpg'\n\nobject_detection(image_path, model, classes_name, max_output_size=100, iou_threshold=0.49, score_threshold=0.2, soft_nms_sigma=0.4)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T10:57:55.509210Z","iopub.execute_input":"2023-01-04T10:57:55.509569Z","iopub.status.idle":"2023-01-04T10:57:56.963668Z","shell.execute_reply.started":"2023-01-04T10:57:55.509537Z","shell.execute_reply":"2023-01-04T10:57:56.962777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_path = '/kaggle/input/open-images-2019-object-detection/test/003177411d8813e7.jpg'\n\nobject_detection(image_path, model, classes_name, max_output_size=50, iou_threshold=0.49, score_threshold=0.3, soft_nms_sigma=0.4)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T10:59:20.664897Z","iopub.execute_input":"2023-01-04T10:59:20.665257Z","iopub.status.idle":"2023-01-04T10:59:22.163656Z","shell.execute_reply.started":"2023-01-04T10:59:20.665227Z","shell.execute_reply":"2023-01-04T10:59:22.162818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_path = '/kaggle/input/open-images-2019-object-detection/test/00379950569d024c.jpg'\n\nobject_detection(image_path, model, classes_name, max_output_size=50, iou_threshold=0.49, score_threshold=0.3, soft_nms_sigma=0.4)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T11:00:07.009898Z","iopub.execute_input":"2023-01-04T11:00:07.010341Z","iopub.status.idle":"2023-01-04T11:00:08.422288Z","shell.execute_reply.started":"2023-01-04T11:00:07.010305Z","shell.execute_reply":"2023-01-04T11:00:08.420803Z"},"trusted":true},"execution_count":null,"outputs":[]}]}