{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What's the Notebook for?\n\nThis competition has 11,898 bounding boxes and 4,919 annotated. It looks sufficient to train the model, but considering they are sourced from almost identical successive frames, it is doubtful that they are sufficient.\n\nCan we inprove the detection model by adding new synthesized images in GAN? This is the first motivation of this approach.\n\nHere, I used the model Pedestrian-Syntheis-GAN [1], which was used to generate pseudo-images of pedestrians. The model can be trained if we prepare an image in which the BBox and the target are replaced by noise, so I trained it using a partial dataset of the competition (about 1,200 sampled clips of size 256x256).\n\n[1] https://arxiv.org/abs/1804.02047","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2021-12-12T02:17:07.661246Z","iopub.execute_input":"2021-12-12T02:17:07.661854Z","iopub.status.idle":"2021-12-12T02:17:07.689439Z","shell.execute_reply.started":"2021-12-12T02:17:07.661758Z","shell.execute_reply":"2021-12-12T02:17:07.688800Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from pathlib import Path\nROOT_DIR = Path('/kaggle')\nINPUT_DIR = ROOT_DIR / 'input'\nWORK_DIR = ROOT_DIR / 'working'\n\ntrain = pd.read_csv(INPUT_DIR / 'tensorflow-great-barrier-reef/train.csv')","metadata":{"execution":{"iopub.status.busy":"2021-12-12T02:17:11.413471Z","iopub.execute_input":"2021-12-12T02:17:11.414150Z","iopub.status.idle":"2021-12-12T02:17:11.467272Z","shell.execute_reply.started":"2021-12-12T02:17:11.414112Z","shell.execute_reply":"2021-12-12T02:17:11.466543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train['n_bbox'] = train['annotations'].apply(lambda x: len(eval(x)))","metadata":{"execution":{"iopub.status.busy":"2021-12-12T02:17:12.593415Z","iopub.execute_input":"2021-12-12T02:17:12.593863Z","iopub.status.idle":"2021-12-12T02:17:12.815794Z","shell.execute_reply.started":"2021-12-12T02:17:12.593826Z","shell.execute_reply":"2021-12-12T02:17:12.815102Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# number of bounding boxes\ntrain['n_bbox'].sum()","metadata":{"execution":{"iopub.status.busy":"2021-12-12T02:17:12.977445Z","iopub.execute_input":"2021-12-12T02:17:12.978171Z","iopub.status.idle":"2021-12-12T02:17:12.987260Z","shell.execute_reply.started":"2021-12-12T02:17:12.978119Z","shell.execute_reply":"2021-12-12T02:17:12.986539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# number of annotated frames\nlen(train.query('n_bbox > 0'))","metadata":{"execution":{"iopub.status.busy":"2021-12-12T02:17:13.402586Z","iopub.execute_input":"2021-12-12T02:17:13.403127Z","iopub.status.idle":"2021-12-12T02:17:13.417689Z","shell.execute_reply.started":"2021-12-12T02:17:13.403088Z","shell.execute_reply":"2021-12-12T02:17:13.416988Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"image_path = INPUT_DIR / 'starfish-generate-psgan-dataset/psgan_datasets/images/train/10.png'","metadata":{"execution":{"iopub.status.busy":"2021-12-12T02:18:08.188780Z","iopub.execute_input":"2021-12-12T02:18:08.189072Z","iopub.status.idle":"2021-12-12T02:18:08.194480Z","shell.execute_reply.started":"2021-12-12T02:18:08.189027Z","shell.execute_reply":"2021-12-12T02:18:08.193459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Input Image for Pedestrian-Synthesis GAN (PS-GAN)","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots()\nimg = plt.imread(image_path)\nax.imshow(img)\nplt.axis('off');","metadata":{"execution":{"iopub.status.busy":"2021-12-12T02:18:09.667531Z","iopub.execute_input":"2021-12-12T02:18:09.668099Z","iopub.status.idle":"2021-12-12T02:18:09.982542Z","shell.execute_reply.started":"2021-12-12T02:18:09.668042Z","shell.execute_reply":"2021-12-12T02:18:09.981343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Synthesized Results\n\nHere, I can show part of the synthesized results.","metadata":{}},{"cell_type":"code","source":"import os\n\n\ndef plot_real_and_fake_image(image_id,\n                             root_path='/kaggle/input/starfish-psgan-results/results/img_128_v7/test_latest/images'):\n    \n    image_paths = [\n        os.path.join(root_path, f'{image_id}_real_B.png'),\n        os.path.join(root_path, f'{image_id}_real_A.png'),\n        os.path.join(root_path, f'{image_id}_fake_B.png'),\n    ]\n    \n    fig, axs = plt.subplots(1, 3, figsize=(9, 3))\n    \n    labels = ['real', 'input', 'synthesized']\n    for ax, image_path, label in zip(axs, image_paths, labels):\n        img = plt.imread(image_path)\n        ax.imshow(img)\n        ax.set_title(label, fontsize=14)\n        ax.axis('off')\n        \n    plt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2021-12-12T02:32:47.011267Z","iopub.execute_input":"2021-12-12T02:32:47.011525Z","iopub.status.idle":"2021-12-12T02:32:47.018180Z","shell.execute_reply.started":"2021-12-12T02:32:47.011498Z","shell.execute_reply":"2021-12-12T02:32:47.017351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Sample of Synthesized Images","metadata":{}},{"cell_type":"code","source":"plot_real_and_fake_image(91)","metadata":{"execution":{"iopub.status.busy":"2021-12-12T02:50:02.856406Z","iopub.execute_input":"2021-12-12T02:50:02.856666Z","iopub.status.idle":"2021-12-12T02:50:03.213710Z","shell.execute_reply.started":"2021-12-12T02:50:02.856639Z","shell.execute_reply":"2021-12-12T02:50:03.212946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_real_and_fake_image(67)","metadata":{"execution":{"iopub.status.busy":"2021-12-12T02:48:33.942025Z","iopub.execute_input":"2021-12-12T02:48:33.942733Z","iopub.status.idle":"2021-12-12T02:48:34.265500Z","shell.execute_reply.started":"2021-12-12T02:48:33.942694Z","shell.execute_reply":"2021-12-12T02:48:34.263937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_real_and_fake_image(22)","metadata":{"execution":{"iopub.status.busy":"2021-12-12T02:47:47.197993Z","iopub.execute_input":"2021-12-12T02:47:47.198587Z","iopub.status.idle":"2021-12-12T02:47:47.550439Z","shell.execute_reply.started":"2021-12-12T02:47:47.198548Z","shell.execute_reply":"2021-12-12T02:47:47.549104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_real_and_fake_image(63)","metadata":{"execution":{"iopub.status.busy":"2021-12-12T02:47:51.031791Z","iopub.execute_input":"2021-12-12T02:47:51.032334Z","iopub.status.idle":"2021-12-12T02:47:51.517444Z","shell.execute_reply.started":"2021-12-12T02:47:51.032295Z","shell.execute_reply":"2021-12-12T02:47:51.514168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I trained the model for about 150 epochs, but the latest model's discriminator seems overfitting because almost perfectly discriminates fake and real. So I used weight of 125 epoch.\n\nThe model sometimes generates COTS-like object, sometimes not. Since I fed only 1200 input data, I think the model are overfitted. It sometimes outputs completely fake-looking images when fed new image which doesn't apper in train dataset.\n\nI doubt this model can generate the data useful for object detection.","metadata":{}},{"cell_type":"markdown","source":"## Sample of Poorly Sinthesized Images","metadata":{}},{"cell_type":"code","source":"plot_real_and_fake_image(92)","metadata":{"execution":{"iopub.status.busy":"2021-12-12T02:38:33.116415Z","iopub.execute_input":"2021-12-12T02:38:33.116682Z","iopub.status.idle":"2021-12-12T02:38:33.618470Z","shell.execute_reply.started":"2021-12-12T02:38:33.116652Z","shell.execute_reply":"2021-12-12T02:38:33.616880Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_real_and_fake_image(15)","metadata":{"execution":{"iopub.status.busy":"2021-12-12T02:42:43.748327Z","iopub.execute_input":"2021-12-12T02:42:43.748855Z","iopub.status.idle":"2021-12-12T02:42:44.101389Z","shell.execute_reply.started":"2021-12-12T02:42:43.748815Z","shell.execute_reply":"2021-12-12T02:42:44.100744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Reference\n\nThe original repository [2] has many bugs, so I forked and modified the repository [3]. \n(The modified repository might still contain many bugs.)\n\n* [2] https://github.com/yueruchen/Pedestrian-Synthesis-GAN\n* [3] https://github.com/bilzard/Pedestrian-Synthesis-GAN","metadata":{}}]}