{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"#### I tried to make the code of this Notebook as simple as possible. If you think you learned something from it. Do Upvote. ⬆️ 😊","metadata":{}},{"cell_type":"markdown","source":"![](https://wallpaperaccess.com/full/218450.jpg)","metadata":{}},{"cell_type":"markdown","source":"## Required Libraries","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\n\nimport matplotlib.pyplot as plt\nimport plotly.express as px\nimport plotly.graph_objects as go\nimport seaborn as sns\nsns.set_style('darkgrid')\n\nfrom PIL import Image, ImageDraw\nimport tensorflow as tf\n\nimport os\nimport ast  ## Change str -> list.\nimport sys\nimport time\n\nimport warnings\nwarnings.filterwarnings('ignore')\n\nimport greatbarrierreef","metadata":{"execution":{"iopub.status.busy":"2021-12-12T11:50:29.405146Z","iopub.execute_input":"2021-12-12T11:50:29.406066Z","iopub.status.idle":"2021-12-12T11:50:38.358223Z","shell.execute_reply.started":"2021-12-12T11:50:29.405955Z","shell.execute_reply":"2021-12-12T11:50:38.357089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Files\n#### **train** - Folder containing training set photos of the form video_{video_id}/{video_frame_number}.jpg.\n\n#### **[train/test].csv** - Metadata for the images. As with other test files, most of the test metadata data is only available to your notebook upon submission. Just the first few rows available for download.\n\n> #### **video_id** - ID number of the video the image was part of. The video ids are not meaningfully ordered.\n> #### **video_frame** - The frame number of the image within the video. Expect to see occasional gaps in the frame number from when the diver surfaced.\n> #### **sequence** - ID of a gap-free subset of a given video. The sequence ids are not meaningfully ordered.\n> #### **sequence_frame** - The frame number within a given sequence.\n> #### **image_id** - ID code for the image, in the format '{video_id}-{video_frame}'\n> #### **annotations** - The bounding boxes of any starfish detections in a string format that can be evaluated directly with Python. Does not use the same format as the predictions you will submit. Not available in test.csv. A bounding box is described by the pixel coordinate (x_min, y_min) of its upper left corner within the image together with its width and height in pixels.\n#### **example_sample_submission.csv** - A sample submission file in the correct format. The actual sample submission will be provided by the API; this is only provided to illustrate how to properly format predictions. The submission format is further described on the Evaluation page.\n\n#### **example_test.npy** - Sample data that will be served by the example API.\n\n#### **greatbarrierreef** - The image delivery API that will serve the test set pixel arrays. You may need Python 3.7 and a Linux environment to run the example offline without errors.","metadata":{}},{"cell_type":"markdown","source":"## Loading Data","metadata":{}},{"cell_type":"code","source":"df_train = pd.read_csv('../input/tensorflow-great-barrier-reef/train.csv')\ndf_train['img_path'] = os.path.join('../input/tensorflow-great-barrier-reef/train_images')+\"/video_\"+df_train.video_id.astype(str)+\"/\"+df_train.video_frame.astype(str)+\".jpg\"\ndf_train.head()","metadata":{"execution":{"iopub.status.busy":"2021-12-12T11:50:38.360611Z","iopub.execute_input":"2021-12-12T11:50:38.360967Z","iopub.status.idle":"2021-12-12T11:50:38.512925Z","shell.execute_reply.started":"2021-12-12T11:50:38.3609Z","shell.execute_reply":"2021-12-12T11:50:38.511628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Exploratory Data Analysis","metadata":{}},{"cell_type":"markdown","source":"#### Let's count how many images are there from each of the three videos.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(8,5))\nsns.countplot(df_train['video_id'], color='#2196F3')","metadata":{"execution":{"iopub.status.busy":"2021-12-12T11:50:38.515165Z","iopub.execute_input":"2021-12-12T11:50:38.515676Z","iopub.status.idle":"2021-12-12T11:50:38.796656Z","shell.execute_reply.started":"2021-12-12T11:50:38.515632Z","shell.execute_reply":"2021-12-12T11:50:38.79577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Now, let's have a look at how many images with bounding boxes are available.","metadata":{}},{"cell_type":"code","source":"with_annotation = len(df_train[df_train['annotations'] != '[]'])\nwithout_annotation = len(df_train[df_train['annotations'] == '[]'])\n\nlabels = ['Without Bounding Box', 'With Bounding Box']\n\nfig = go.Figure([go.Bar(x=labels, \n                        y=[without_annotation, with_annotation], width=0.6)])\n\nfig.update_layout(autosize=False, width=700, height=400, margin=dict(l=60, r=60, b=50, t=50, pad=4))\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2021-12-12T11:50:38.800251Z","iopub.execute_input":"2021-12-12T11:50:38.800534Z","iopub.status.idle":"2021-12-12T11:50:38.926839Z","shell.execute_reply.started":"2021-12-12T11:50:38.800491Z","shell.execute_reply":"2021-12-12T11:50:38.92586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Clearly, very few number of images have annotations as compared to without annotations images.","metadata":{}},{"cell_type":"markdown","source":"#### Now, we'll going to find out how many bouning boxes are there in each annotation columns.\n#### In order to do that we can simply count the number of opening curly brackets - '{' in annotation column.\n> **Example** :-  `df_train['annotations'][12843]` *where 12843 is a random number.*\n\n> Gives -> \n\n> \"[{'x': 338, 'y': 229, 'width': 45, 'height': 27},\n>\n> {'x': 357, 'y': 285, 'width': 37, 'height': 38}, \n> \n> {'x': 173, 'y': 588, 'width': 65, 'height': 56}, \n> \n> {'x': 234, 'y': 598, 'width': 31, 'height': 26}]\" \n\n> **We can see that it has 4 bounding box coordinates, the simple way to get this is to count number of opening curly brackets.**","metadata":{}},{"cell_type":"code","source":"# creating new column which contains the total number of bounding boxes\ndf_train['No_bbox'] = df_train['annotations'].apply(lambda x:x.count('{')) \n\n# Example\n\nn = df_train['No_bbox'][12843]\nprint(df_train['annotations'][12843])\nprint(f'Number of bounding boxes are : {n}.')","metadata":{"execution":{"iopub.status.busy":"2021-12-12T11:50:38.928774Z","iopub.execute_input":"2021-12-12T11:50:38.929129Z","iopub.status.idle":"2021-12-12T11:50:38.960285Z","shell.execute_reply.started":"2021-12-12T11:50:38.929083Z","shell.execute_reply":"2021-12-12T11:50:38.959316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = px.bar(df_train['No_bbox'].value_counts().drop(0), title='Count of Bounding Boxes')\nfig.update_layout(autosize=False, width=700, height=400, margin=dict(l=60, r=60, b=50, t=50, pad=4))\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2021-12-12T11:50:38.961794Z","iopub.execute_input":"2021-12-12T11:50:38.962117Z","iopub.status.idle":"2021-12-12T11:50:39.855878Z","shell.execute_reply.started":"2021-12-12T11:50:38.962074Z","shell.execute_reply":"2021-12-12T11:50:39.854873Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Most of the images has only single Bounding Box. Very few have more than 5 BBox.","metadata":{}},{"cell_type":"markdown","source":"#### Now, lets change 'annotations' from string to list data type using ***ast***.","metadata":{}},{"cell_type":"code","source":"df_train['annotations'] = df_train['annotations'].apply(ast.literal_eval)\ndf_train.head()","metadata":{"execution":{"iopub.status.busy":"2021-12-12T11:50:39.857553Z","iopub.execute_input":"2021-12-12T11:50:39.859132Z","iopub.status.idle":"2021-12-12T11:50:40.255948Z","shell.execute_reply.started":"2021-12-12T11:50:39.859084Z","shell.execute_reply":"2021-12-12T11:50:40.254845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Creating new DataFrame which carry images containing more than 1 Bounding Boxes (You could take any value you like) and then use that row's 'video_id' to see image with Bounding Boxes.","metadata":{}},{"cell_type":"code","source":"df2 = df_train[df_train['annotations'].astype(str) != \"[]\"]\ndf2 = df2[df2['No_bbox'] == 5]\ndf2.head()","metadata":{"execution":{"iopub.status.busy":"2021-12-12T11:50:40.257592Z","iopub.execute_input":"2021-12-12T11:50:40.257933Z","iopub.status.idle":"2021-12-12T11:50:40.321852Z","shell.execute_reply.started":"2021-12-12T11:50:40.25787Z","shell.execute_reply":"2021-12-12T11:50:40.320979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Visualization","metadata":{}},{"cell_type":"code","source":"def img_viz(df, id):\n    image = df_train['img_path'][id]\n    img = Image.open(image)\n    \n    for box in df_train['annotations'][id]:\n        shape = [box['x'], box['y'], box['x']+box['width'], box['y']+box['height']]\n        ImageDraw.Draw(img).rectangle(shape, outline =\"red\", width=3)\n    display(img)","metadata":{"execution":{"iopub.status.busy":"2021-12-12T11:50:40.323439Z","iopub.execute_input":"2021-12-12T11:50:40.323838Z","iopub.status.idle":"2021-12-12T11:50:40.332453Z","shell.execute_reply.started":"2021-12-12T11:50:40.323792Z","shell.execute_reply":"2021-12-12T11:50:40.330174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img_viz(df_train, id=5474)","metadata":{"execution":{"iopub.status.busy":"2021-12-12T11:50:40.336764Z","iopub.execute_input":"2021-12-12T11:50:40.338136Z","iopub.status.idle":"2021-12-12T11:50:40.756142Z","shell.execute_reply.started":"2021-12-12T11:50:40.338083Z","shell.execute_reply":"2021-12-12T11:50:40.755052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Awesome!!!","metadata":{}},{"cell_type":"markdown","source":"## Model\n#### Load the TensorFlow COTS detection model into memory and define some util functions for running inference.\n#### Read more about it [here](https://www.kaggle.com/khanhlvg/cots-detection-w-tensorflow-object-detection-api).","metadata":{}},{"cell_type":"code","source":"INPUT_DIR = '../input/tensorflow-great-barrier-reef/'\nsys.path.insert(0, INPUT_DIR)","metadata":{"execution":{"iopub.status.busy":"2021-12-12T11:50:40.757372Z","iopub.execute_input":"2021-12-12T11:50:40.757669Z","iopub.status.idle":"2021-12-12T11:50:40.767998Z","shell.execute_reply.started":"2021-12-12T11:50:40.757624Z","shell.execute_reply":"2021-12-12T11:50:40.76573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"MODEL_DIR = '../input/cots-detection-w-tensorflow-object-detection-api/cots_efficientdet_d0'\nstart_time = time.time()\ntf.keras.backend.clear_session()\ndetect_fn_tf_odt = tf.saved_model.load(os.path.join(os.path.join(MODEL_DIR, 'output'), 'saved_model'))\nend_time = time.time()\nelapsed_time = end_time - start_time\nprint('Elapsed time: ' + str(elapsed_time) + 's')","metadata":{"execution":{"iopub.status.busy":"2021-12-12T11:50:40.76935Z","iopub.execute_input":"2021-12-12T11:50:40.770227Z","iopub.status.idle":"2021-12-12T11:51:21.521144Z","shell.execute_reply.started":"2021-12-12T11:50:40.770137Z","shell.execute_reply":"2021-12-12T11:51:21.519982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Helper functions","metadata":{}},{"cell_type":"code","source":"def load_image_into_numpy_array(path):\n    \"\"\"Load an image from file into a numpy array.\n\n    Puts image into numpy array to feed into tensorflow graph.\n    Note that by convention we put it into a numpy array with shape\n    (height, width, channels), where channels=3 for RGB.\n\n    Args:\n    path: a file path (this can be local or on colossus)\n\n    Returns:\n    uint8 numpy array with shape (img_height, img_width, 3)\n    \"\"\"\n    img_data = tf.io.gfile.GFile(path, 'rb').read()\n    image = Image.open(io.BytesIO(img_data))\n    (im_width, im_height) = image.size\n    \n    return np.array(image.getdata()).reshape(\n      (im_height, im_width, 3)).astype(np.uint8)\n\ndef detect(image_np):\n    \"\"\"Detect COTS from a given numpy image.\"\"\"\n\n    input_tensor = np.expand_dims(image_np, 0)\n    start_time = time.time()\n    detections = detect_fn_tf_odt(input_tensor)\n    return detections","metadata":{"execution":{"iopub.status.busy":"2021-12-12T11:51:21.522691Z","iopub.execute_input":"2021-12-12T11:51:21.524098Z","iopub.status.idle":"2021-12-12T11:51:21.533524Z","shell.execute_reply.started":"2021-12-12T11:51:21.52405Z","shell.execute_reply":"2021-12-12T11:51:21.532201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Creating Environment","metadata":{}},{"cell_type":"code","source":"env = greatbarrierreef.make_env()   # initialize the environment\niter_test = env.iter_test()    # an iterator which loops over the test set and sample submission","metadata":{"execution":{"iopub.status.busy":"2021-12-12T11:51:21.53512Z","iopub.execute_input":"2021-12-12T11:51:21.535555Z","iopub.status.idle":"2021-12-12T11:51:21.546991Z","shell.execute_reply.started":"2021-12-12T11:51:21.535509Z","shell.execute_reply":"2021-12-12T11:51:21.54598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Prediction | Submission","metadata":{}},{"cell_type":"code","source":"DETECTION_THRESHOLD = 0.13\n\nsubmission_dict = {\n    'id': [],\n    'prediction_string': [],\n}\n\nfor (image_np, sample_prediction_df) in iter_test:\n    height, width, _ = image_np.shape\n    \n    # Run object detection using the TensorFlow model.\n    detections = detect(image_np)\n    \n    # Parse the detection result and generate a prediction string.\n    num_detections = detections['num_detections'][0].numpy().astype(np.int32)\n    predictions = []\n    for index in range(num_detections):\n        score = detections['detection_scores'][0][index].numpy()\n        if score < DETECTION_THRESHOLD:\n            continue\n\n        bbox = detections['detection_boxes'][0][index].numpy()\n        y_min = int(bbox[0] * height)\n        x_min = int(bbox[1] * width)\n        y_max = int(bbox[2] * height)\n        x_max = int(bbox[3] * width)\n        \n        bbox_width = x_max - x_min\n        bbox_height = y_max - y_min\n        \n        predictions.append('{:.2f} {} {} {} {}'.format(score, x_min, y_min, bbox_width, bbox_height))\n    \n    # Generate the submission data.\n    prediction_str = ' '.join(predictions)\n    sample_prediction_df['annotations'] = prediction_str\n    env.predict(sample_prediction_df)\n\n    print('Prediction:', prediction_str)","metadata":{"execution":{"iopub.status.busy":"2021-12-12T11:51:21.548906Z","iopub.execute_input":"2021-12-12T11:51:21.549348Z","iopub.status.idle":"2021-12-12T11:51:31.994074Z","shell.execute_reply.started":"2021-12-12T11:51:21.549267Z","shell.execute_reply":"2021-12-12T11:51:31.992847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## References\n* [COTS detection w/ TensorFlow Object Detection API](https://www.kaggle.com/khanhlvg/cots-detection-w-tensorflow-object-detection-api)\n* [EDA: Let's understand the data - protect the reef](https://www.kaggle.com/casfranco/eda-let-s-understand-the-data-protect-the-reef)\n* [Inference using EfficientDet-D0 model](https://www.kaggle.com/khanhlvg/inference-using-efficientdet-d0-model-tensorflow/notebook)","metadata":{}}]}