{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# iWildCam 2021 - Starter Notebook\n\n## Let's contribute study for wild animals from here with camera traps traps dataset!📷","metadata":{}},{"cell_type":"code","source":"import collections\nfrom datetime import datetime as dt\nimport gc\nimport glob \nimport json\nimport os\nimport warnings\nwarnings.filterwarnings('ignore')\n\nimport cv2\nfrom imblearn.under_sampling import RandomUnderSampler\nfrom IPython.display import YouTubeVideo\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nfrom PIL import Image, ImageDraw\nimport plotly.express as px\nimport seaborn as sns\n\n%matplotlib inline\n\nSEED = 2021","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Contents\n- [Competition goal and metric](#0)\n- [Motivation of competition](#1)\n- [Explanation for image file](#2)\n- [Explanation for submission file](#3)\n- [Explanation for metadata file](#4)\n- [Exploratory data analysis](#5)\n- [Sample solution](#6)","metadata":{}},{"cell_type":"markdown","source":"### <u>About this notebook</u>\nI want to give you an overview of the prior knowledge and data that might be needed in this challenge!\n\nThis contest aims to detect wildlife in trapped picture at new monitoring locations. Now, we can get great insights about wildlife by camera traps. Camera trap is popular method and there are so many data in the world. But due to the so large number of data, it seems that the data was not always effectively accessed and utilized.\n\nIn this competition, we aim to develop a model that effectively classifies animals taken at different observation points　in the world. This challenge will surely bring great insights.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true}},{"cell_type":"markdown","source":"<a id=\"0\"></a> <br>\n# <div class=\"alert alert-block alert-warning\">Competition goal and metric</div>\n\nThe goal of this competition is to categorize species and count the number of individuals across image bursts of camera trap. Image bursts of camera trap are assigned unique ID. The individual images that make up image bursts are also assigned ID. Using the image burst as input, count up how many animals of the 204 annotated species are present and output as CSV file.\n\n<img src=\"https://raw.githubusercontent.com/tasotasoso/kaggle_media/main/iwildcam2021/task_image.png\n\" width=\"***500***\">\n\nEvaluation will be done by Mean Columnwise Root Mean Squared Error(MCRMSE). \n\n$$\n\\frac{1}{m}\\sum_{j=1}^{m}\\sqrt{\\frac{1}{n}\\sum_{i=1}^{n}(x_{ij}-y_{ij})^2}\n$$\n\nj represents a species, i represents a sequence, x_ij is the predicted count for that species in that sequence, and y_ij is the ground truth count. This is an index that takes the RMSE for each species and then averages it across species.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"1\"></a> <br>\n# <div class=\"alert alert-block alert-info\">Motivation of competition</div>\n\n## How can we take wildlife pictures and use?\nWe can take wildlife pictures by camera trap. Camera traps have infrared sensor or motion sensor, so they can detect animals. When animals come near, camera take their pictures. Using camera traps, we can monitor wildlifes continuously ,at several point and at the same time. So we can understand how animals run their life in the area researchers interest.[1]\n\nFor example, in Kaen Krachan National Park in Thailand, indian sinatra are directlly observed. So it was pointed out that there is no longer any possibility. But by using camera trap, we could confirm thir existence. [2]\n\nTraditionally, camera traps have been considered an excellent method for investigating ecological information about wildlife in a certain area.　Data are used for population estimation, calculating population index, 24hours monitorling and so on.[3]\n\nResently there has been a movement to make effective use of photos from camera traps using machine learning. One of the example is [4]. Google successed to access so much wildlife photo knowledge. Following Wilflife Insights is the platform of taking the initiative.","metadata":{}},{"cell_type":"code","source":"YouTubeVideo('qKgRbkCkRFY')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thus, creating a model for effectively classifying wild animals is a very significant effort.\n\n### Referece\n\n[1]https://www.wwf.org.uk/project/conservationtechnology/camera-trap\n\n[2]https://www.wwf.or.jp/campaign/2015_camera/\n\n[3]https://en.wikipedia.org/wiki/Camera_trap\n\n[4]https://www.blog.google/products/earth/ai-finds-where-the-wild-things-are/","metadata":{}},{"cell_type":"markdown","source":"<a id=\"2\"></a> <br>\n# <div class=\"alert alert-block alert-info\">Explanation for image file</div>\n\nFirst, let's take a look at what kind of images are available.","metadata":{}},{"cell_type":"code","source":"TRAIN_DATA_PATH = '../input/iwildcam2021-fgvc8/train/'\nTEST_DATA_PATH = '../input/iwildcam2021-fgvc8/test/'\n\ntrain_jpeg = glob.glob(TRAIN_DATA_PATH + '*')\ntest_jpeg = glob.glob(TEST_DATA_PATH + '*')\n\nprint(\"number of train jpeg data:\", len(train_jpeg))\nprint(\"number of test jpeg data:\", len(test_jpeg))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Train data","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize=(25, 16))\nfor i,im_path in enumerate(train_jpeg[:16]):\n    ax = fig.add_subplot(4, 4, i+1, xticks=[], yticks=[])\n    im = Image.open(im_path)\n    im = im.resize((480,270))\n    plt.imshow(im)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(25, 16))\nfor i,im_path in enumerate(train_jpeg[16:32]):\n    ax = fig.add_subplot(4, 4, i+1, xticks=[], yticks=[])\n    im = Image.open(im_path)\n    im = im.resize((480,270))\n    plt.imshow(im)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Test data","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize=(25, 16))\nfor i,im_path in enumerate(test_jpeg[:16]):\n    ax = fig.add_subplot(4, 4, i+1, xticks=[], yticks=[])\n    im = Image.open(im_path)\n    im = im.resize((480,270))\n    plt.imshow(im)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"By camera traps, images are taken continuously because they are captured in bursts triggered by motion. Therefore, dataset also contains series of images, and in addition to IDs of image, IDs of the sequence are assigned.We will use IDs of image to load the image, on the other hands we will use IDs of sequence to submit.\n\nTake a look at [the submission file for details](#3).\n\nAdditionally, we will notice that some of the images are in color and some are in black and white. This is due to the difference in whether the images were taken during the day or at night.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"3\"></a> <br>\n# <div class=\"alert alert-block alert-success\">Explanation for submission file</div>\n\nExplanation for submission file\n\nIn iWildcam 2021 - FGVC8 conpetition, we have to detect the number of each animal species in the image sequences. On the other hand, [last year (in iWildcam 2020 - FGVC7)](https://www.kaggle.com/c/iwildcam-2020-fgvc7/overview/evaluation) we identified which category of animal was being shown. ","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(\"../input/iwildcam2021-fgvc8/sample_submission.csv\")\nsub.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Id columns represents id of sequence as we see in [Checking iwildcam2021_train_annotations.json](#4-1). PredictedX columns are number of individuals of species in sequence. There are 111057 sequences in test dataset, and 205 species in given dataset. ","metadata":{}},{"cell_type":"code","source":"sub.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"4\"></a> <br>\n# <div class=\"alert alert-block alert-success\">Explanation for metadata file</div>","metadata":{}},{"cell_type":"markdown","source":"<a id=\"4-1\"></a>\n## Checking iwildcam2021_train_annotations.json\n\n\nWe are provided annotation data for train data as \"iwildcam2021_train_annotations.json\". This json follows COCO-CameraTraps format with additional field.\n\nIf we load the json, we can find there are three key-values in it.","metadata":{}},{"cell_type":"code","source":"with open('../input/iwildcam2021-fgvc8/metadata/iwildcam2021_train_annotations.json', encoding='utf-8') as json_file:\n    train_annotations =json.load(json_file)\n    \ntrain_annotations.keys()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In images value, we can get data for each wildcan images. Wildcam will take several frames in a row. The value of 'seq_num_frames' key is the number of frames, 'id' is the id of the image, and 'seq_id' is the ID associated with the sequentially shot image. This 'seq_id' is the same as the 'Id' in the submission file.\n\nLet's extract the data corresponding to shot of seq_id:302ad820-7d42-11eb-8fb5-0242ac1c0002.","metadata":{}},{"cell_type":"code","source":"train_annotations_seq = train_annotations[\"images\"][94:104]\ntrain_annotations_seq","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If we see the images, we can find that they are a series of images.","metadata":{}},{"cell_type":"code","source":"train_images_seq = [(TRAIN_DATA_PATH+item[\"id\"]+'.jpg') for item in train_annotations_seq]\nimg_array = []\nsize = (480,270)\n\nfig = plt.figure(figsize=(25, 16))\nfor i,im_path in enumerate(train_images_seq):\n    ax = fig.add_subplot(4, 3, i+1, xticks=[], yticks=[])\n    im = Image.open(im_path)\n    im = im.resize(size)\n    plt.imshow(im)\n    \n    img_array.append(im)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The value of the categories key contains a list of annotated animal species. id 0 is empty. The id is up to 571, but there are 204 annotated species. So we have to classify for 204 species + empty at most (not 571 + empty)!","metadata":{}},{"cell_type":"code","source":"df_categories = pd.DataFrame.from_records(train_annotations[\"categories\"])\ndf_categories","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Each training image has at least one associated annotation. The annotatetd catefory_ids are in value of \"annotations\" key.","metadata":{}},{"cell_type":"code","source":"train_annotations[\"annotations\"][:10]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There seems to be data in train data for all annotated categories.","metadata":{}},{"cell_type":"code","source":"train_annotated_category = set([ annotation[\"category_id\"] for annotation in train_annotations[\"annotations\"]])\nlen(train_annotated_category)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"4-2\"></a> <br>\n## Checking iwildcam2021_test_information.json\n\nInformation for test dataset. The format is similar to iwildcam2021_train_annotations.json with only images key.","metadata":{}},{"cell_type":"code","source":"with open('../input/iwildcam2021-fgvc8/metadata/iwildcam2021_test_information.json', encoding='utf-8') as json_file:\n    test_information =json.load(json_file)\n    \ntest_information.keys()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_information['images'][:5]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"4-3\"></a> <br>\n## Checking iwildcam2021_megadetector_results.json\n\nWe can also use [Microsoft AI for Earth MegaDetector](https://github.com/microsoft/CameraTraps/blob/master/megadetector.md). This model is trained to detect animals, people, and vehicles in camera trap images using hundreds of thousands of bounding boxes from various ecosystems. The model does not identify animals, it only finds them.\n\nWe are provided some sample detection results as \"iwildcam2021_megadetector_results.json\".","metadata":{}},{"cell_type":"code","source":"with open('../input/iwildcam2021-fgvc8/metadata/iwildcam2021_megadetector_results.json', encoding='utf-8') as json_file:\n    megadetector_results =json.load(json_file)\n    \nmegadetector_results.keys()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are three key-value data in json.\n\nDetected result is in images value.","metadata":{}},{"cell_type":"code","source":"megadetector_results_df = pd.DataFrame(megadetector_results[\"images\"])\nmegadetector_results_df.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since there are 263504 detection data, it seems that all the data from train data and test data have been processed. So we do not need to run MegaDetector ourselves. If you want to finetune with the MegaDetector's weights or redo the estimation yourself, please refer to this [notebook](https://www.kaggle.com/nayuts/try-megadetector-crop-animals-on-kaggle-notebook).","metadata":{}},{"cell_type":"code","source":"print(f\"There are {len(megadetector_results_df)} detection data.\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Especially, detected bbox is in detections value.","metadata":{}},{"cell_type":"code","source":"megadetector_results_df.iloc[100][\"detections\"]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see the result like this.","metadata":{}},{"cell_type":"code","source":"#Refered: https://www.kaggle.com/qinhui1999/how-to-use-bbox-for-iwildcam-2020 \n\ndef draw_bboxs(detections_list, im):\n    \"\"\"\n    detections_list: list of set includes bbox.\n    im: image read by Pillow.\n    \"\"\"\n    \n    for detection in detections_list:\n        x1, y1,w_box, h_box = detection[\"bbox\"]\n        ymin,xmin,ymax, xmax=y1, x1, y1 + h_box, x1 + w_box\n        draw = ImageDraw.Draw(im)\n        \n        imageWidth=im.size[0]\n        imageHeight= im.size[1]\n        (left, right, top, bottom) = (xmin * imageWidth, xmax * imageWidth,\n                                      ymin * imageHeight, ymax * imageHeight)\n        \n        draw.line([(left, top), (left, bottom), (right, bottom),\n               (right, top), (left, top)], width=4, fill='Red')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's see 100th data of train dataset.\ndata_index = 100\n\n# Load 100th image data. \nim = Image.open(\"../input/iwildcam2021-fgvc8/train/\" + megadetector_results_df.loc[data_index]['id'] + \".jpg\")\nim = im.resize((480,270))\n\n# Overwrite bbox\ndraw_bboxs(megadetector_results_df.loc[data_index]['detections'], im)\n\n# Show\nplt.imshow(im)\nplt.title(f\"image {data_index} with bbox\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"megadetector_results_df.loc[data_index]['detections'][0][\"bbox\"]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It is also possible to crop the detected area like this. If you save the image, you can create  dataset.","metadata":{}},{"cell_type":"code","source":"def get_crop_area(bbox, image_size):\n    x1, y1,w_box, h_box = bbox\n    ymin,xmin,ymax, xmax = y1, x1, y1 + h_box, x1 + w_box\n    area = (xmin * image_size[0], ymin * image_size[1], \n            xmax * image_size[0], ymax * image_size[1])\n    return area\n\ncrop_area = get_crop_area(megadetector_results_df.loc[data_index]['detections'][0][\"bbox\"], im.size)\nim_croped = im.crop(crop_area)\nplt.imshow(im_croped)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"info and detection_categories values are incidental information.","metadata":{}},{"cell_type":"code","source":"megadetector_results[\"info\"]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"megadetector_results[\"detection_categories\"]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I created dataset that crop the detection area of MegaDetector. Because of the processing time involved, I wrote it in [separate notebook](https://www.kaggle.com/nayuts/256-x-256-cropped-images).","metadata":{}},{"cell_type":"markdown","source":"<a id=\"5\"></a> <br>\n# <div class=\"alert alert-block alert-success\">Exploratory data analysis</div>\n\nI'll check distribution of data in mainly three insight as following:\n- Category ID\n\n- Time point\n\n- Location\n\nAs far as the following EDA results are concerned, the given dataset seems be the same as last year. If you have interest, compare it to [my notes from last year](https://www.kaggle.com/nayuts/iwildcam-2020-overviewing-for-start). To a greater or lesser extent, we can use the findings of last year's competition.","metadata":{}},{"cell_type":"markdown","source":"## How many data are there per animal category Id?\n\nThere are a lot of categories in dataset. To confirme how many data are there in each categories, I plot barplot.","metadata":{}},{"cell_type":"code","source":"# Preperation for isualization\ndf_categories = pd.DataFrame(train_annotations[\"categories\"])\nlabels_id = [item[\"id\"] for item in train_annotations[\"categories\"]]\ncnt = collections.Counter([item[\"category_id\"] for item in train_annotations[\"annotations\"]])\ndf_categories_count = pd.DataFrame.from_dict(cnt, orient='index').reset_index()\ndf_categories_count = df_categories_count.rename(columns={'index':'id', 0:'count'})\n\ndf_categories_count = df_categories_count.merge(df_categories, on='id').sort_values(by=['count'], ascending=False)","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(30, 4))\nax = sns.barplot(x=\"id\", y=\"count\",data=df_categories_count, order=labels_id)\nax.set(ylabel='count')\nax.set(ylim=(0,80000))\nplt.title('distribution of count per id in train')","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since there are many categories,I will also provide a plotly interactive bar chart to make it easier to check the details.","metadata":{}},{"cell_type":"code","source":"fig = px.bar(df_categories_count, x=\"id\", y=\"count\", \n             title='distribution of count per id in train',\n             width=800, height=400, color='id')\nfig.show()","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The annotation data seems to be biased to some extent. To see the breakdown, let's look at the top 10 categories.\n\nmpty is the most, but annotations stating that animals are in the picture also seem to vary among the top 10.","metadata":{}},{"cell_type":"code","source":"df_categories_count.iloc[:10]","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"On the other hand, fewer categories have only about one sample. We need to be careful when splitting the dataset to train and validation data when training the model.","metadata":{}},{"cell_type":"code","source":"df_categories_count.iloc[-10:]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's look at the cumulative ratio. If we take the cumulative sum in order of increasing number, we can see that the 40th category reaches 95% and the 90th category reaches 99%.","metadata":{}},{"cell_type":"code","source":"# Refered https://www.kaggle.com/kushal1506/deciding-n-components-in-pca\n\nfig, ax = plt.subplots(figsize=(30, 10))\nxi = np.arange(1, len(df_categories_count)+1, step=1)\n\nplt.ylim(0.0,1.1)\nplt.plot(xi, df_categories_count[\"count\"].cumsum()/sum(df_categories_count[\"count\"]), marker='o', linestyle='--', color='b')\n\n\nplt.xlabel('Number of category', fontsize=30)\nplt.xticks(np.arange(0, len(df_categories_count), step=10)) #change from 0-based array index to 1-based human-readable label\nplt.ylabel('Accumulation Ratio (%)', fontsize=30)\nplt.title('Relationships when cumulative sums are taken in order of increasing categories.', fontsize=30)\n\nplt.axhline(y=0.99, color='g', linestyle='-')\nplt.text(0.5, 1.00, '99%', color = 'green', fontsize=30)\n\nplt.axhline(y=0.95, color='r', linestyle='-')\nplt.text(0.5, 0.92, '95%', color = 'red', fontsize=30)\n\nplt.axvline(x=40, color='g', linestyle='-')\nplt.text(40, 0.5, '40th', color = 'green', fontsize=30)\n         \nplt.axvline(x=90, color='g', linestyle='-')\nplt.text(90, 0.5, '90th', color = 'green', fontsize=30)\n\nax.grid(axis='x')\nplt.show()","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If we use such imbalanced data as it is, we may not be able to train our models well. For a quickly, using [RandomUnderSampler](https://imbalanced-learn.org/stable/under_sampling.html#controlled-under-sampling-techniques) of [imblearn](https://imbalanced-learn.org/stable/) may improve the situation.","metadata":{}},{"cell_type":"code","source":"# Convert annotation data to pandas DataFrame\ndf_train_annotations = pd.DataFrame(train_annotations[\"annotations\"])\n\n# Under sampling\nrus = RandomUnderSampler(random_state=SEED, replacement=True)\ndf_train_annotations_resampled, _ = rus.fit_resample(df_train_annotations, df_train_annotations[\"category_id\"])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_annotations_resampled.reset_index(drop=True)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## When did data taken?\n\nBecause animals can change their activity from time to time, we want to understand how data is distributed over time.","metadata":{}},{"cell_type":"code","source":"df_images = pd.DataFrame(train_annotations[\"images\"])\ndf_images_test = pd.DataFrame(test_information[\"images\"])","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"month_year = df_images['datetime'].map(lambda str: str[2:7])\nlabels_month_year = sorted(list(set(month_year)))\n\nmonth_year_test = df_images_test['datetime'].map(lambda str: str[2:7])","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Year and Month Perspective","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(1,2, figsize=(30,7))\nax = plt.subplot(1,2,1)\nax = plt.title('Count of train data per month & year')\nax = sns.countplot(month_year, order=labels_month_year)\nax.set(xlabel='YY-mm', ylabel='count')\nax.set(ylim=(0,50000))\n\nax = plt.subplot(1,2,2)\nax = plt.title('Count of test data per month & year')\n\nax = sns.countplot(month_year_test, order=labels_month_year)\nax.set(xlabel='YY-mm', ylabel='count')\nax.set(ylim=(0,50000))","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Data starts 2013-01 but there seems be some lacks. For example, train data between 2013-11 to 2014-02 are missing.\n\nAlso we can find that train data in between 2013-01 to 2013-07 are rich than other time point.\n\nTrain data covers test data in perspective of time point.","metadata":{}},{"cell_type":"markdown","source":"### Monthly perspective","metadata":{}},{"cell_type":"code","source":"labels_month = sorted(list(set(df_images['datetime'].map(lambda str: str[5:7]))))\n\nfig, ax = plt.subplots(1,2, figsize=(20,7))\nax = plt.subplot(1,2,1)\nplt.title('Count of train data per month')\nax = sns.countplot(df_images['datetime'].map(lambda str: str[5:7] ), order=labels_month)\nax.set(xlabel='mm', ylabel='count')\nax.set(ylim=(0,55000))\n\nax = plt.subplot(1,2,2)\nplt.title('Count of test data per month')\nax = sns.countplot(df_images_test['datetime'].map(lambda str: str[5:7] ), order=labels_month)\nax.set(xlabel='mm', ylabel='count')\nax.set(ylim=(0,55000))","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Train data are bias. In February, March, June and July, data are rich than other months.\n\nTrain data covers test data in perspective of month.\n\nData for November and December are missing. Do animals hibernate?","metadata":{}},{"cell_type":"markdown","source":"### Hourly perspectives","metadata":{}},{"cell_type":"code","source":"train_taken_hour = df_images['datetime'].map(lambda x: dt.strptime(x, '%Y-%m-%d %H:%M:%S.%f').hour)\ntest_taken_hour = df_images_test['datetime'].map(lambda x: dt.strptime(x, '%Y-%m-%d %H:%M:%S.%f').hour)\n\nfig, ax = plt.subplots(1,2, figsize=(20,7))\nax = plt.subplot(1,2,1)\nplt.title('Count of train data per hour')\nax = sns.countplot(train_taken_hour)\nax.set(xlabel='hour', ylabel='count')\nax.set(ylim=(0,20000))\n\nax = plt.subplot(1,2,2)\nplt.title('Count of test data per hour')\nax = sns.countplot(test_taken_hour)\nax.set(xlabel='hour', ylabel='count')\nax.set(ylim=(0,20000))","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If we decide arbitrarily during daytime and at night, we can also calculate diurnal and nocturnal data counts.\n\nFor example, we define \"during daytime\" is \"6-17 O'clock\" and \"at night\" is \"18-5 O'clock\",","metadata":{}},{"cell_type":"markdown","source":"### Day and night perspective","metadata":{}},{"cell_type":"code","source":"train_taken_phase = train_taken_hour.map(lambda x: \"daytime\" if x >= 6 and x < 18 else \"night\")\ntest_taken_phase = test_taken_hour.map(lambda x: \"daytime\" if x >= 6 and x < 18 else \"night\")","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(1,2, figsize=(20,7))\nax = plt.subplot(1,2,1)\nplt.title('Count of train data per phase')\nax = sns.countplot(train_taken_phase, order=[\"daytime\", \"night\"])\nax.set(xlabel='phase', ylabel='count')\nax.set(ylim=(0,200000))\n\nax = plt.subplot(1,2,2)\nplt.title('Count of test data per phase')\nax = sns.countplot(test_taken_phase, order=[\"daytime\", \"night\"])\nax.set(xlabel='phase', ylabel='count')\nax.set(ylim=(0,200000))","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"When we looked at given image files at the beginning of this notebook, we saw mixture of color and black and white images. Using the time of day feature, we may be able to successfully distinguish whether they are in color or not.","metadata":{}},{"cell_type":"markdown","source":"## Where did data taken?\n\nWe are required to detect photographs taken at different locations, but how distribute are data in perspect of location?","metadata":{}},{"cell_type":"code","source":"labels_location_train = sorted(list(set(df_images['location'])))\nlabels_location_test = sorted(list(set(df_images_test['location'])))\nlabels_location = labels_location_train + labels_location_test\n\nfig = plt.figure(figsize=(30, 4))\nax = sns.countplot(df_images['location'], order=labels_location)\nax.set(xlabel='location', ylabel='count')\nplt.title('Count of train data per location')","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(30, 4))\nax = sns.countplot(df_images_test['location'], order=labels_location)\nax.set(xlabel='location', ylabel='count')\nplt.title('Count of test data per location')","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Train data and test data seems be completelly taken in different locations.\n\nNumber of pictures are greatly differend by location.","metadata":{}},{"cell_type":"markdown","source":"<a id=\"6\"></a> <br>\n# <div class=\"alert alert-block alert-danger\">Sample solution</div>\n\nI have created sample solution, although it is not very accurate. Check out [this notebook](https://www.kaggle.com/nayuts/efficientnet-with-undersampling). It requires GPU to be turned on. I separated in order to people who forks this notebook and do some trial and error don't waste GPU time without realizing it.\n\n\nI haven't beaten kaggle_sample_all_zero_iwildcam_2021.csv yet, but I will publish the idea.\n\n1. First we crop the image based on the bbox detected by MegaDetector.\n2. In the training data, the correct answer labels are given as annotations, so we can use them to train the model.\n3. Classify the cropped images of the test data with the trained model.\n4. We choose the animal species and their counts of the image with the highest count among the images in the same image burst.","metadata":{}},{"cell_type":"markdown","source":"<img src=\"https://raw.githubusercontent.com/tasotasoso/kaggle_media/main/iwildcam2021/model_image.png\" width=\"***300***\">","metadata":{}},{"cell_type":"markdown","source":"## How can we improve accuracy?\n\nWe also check how empty the training data and test data are. We will also check how empty the training and test data is, because when we submit the results of our inference, we will fount that \"kaggle_sample_all_zero_iwildcam_2021.csv\" is very powerful. This should be because much of the testdata is mostly empty. From the detection results using MegaDetector, we can see roughly how much of the image is empty. Let's take a look.","metadata":{}},{"cell_type":"code","source":"def is_in_test(x):\n    if os.path.exists(TEST_DATA_PATH + x + \".jpg\"):\n        return True\n    else:\n        return False\n    \nare_images_in_test = [ is_in_test(x) for x in megadetector_results_df[\"id\"]]\nare_images_in_train = [not is_in_test for is_in_test in are_images_in_test]\n\ntrain_megadetector_results_df =  megadetector_results_df[are_images_in_train]\ntest_megadetector_results_df =  megadetector_results_df[are_images_in_test]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(15, 4))\nax = sns.countplot([len(detection) for detection in train_megadetector_results_df[\"detections\"]])\nax.set(xlabel='Number of detections', ylabel='count')\nplt.title('Distribution of number of detection by MegaDetector for train data.')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(15, 4))\nax = sns.countplot([len(detection) for detection in test_megadetector_results_df[\"detections\"]])\nax.set(xlabel='Number of detections', ylabel='count')\nplt.title('Distribution of number of detection by MegaDetector for test data.')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that the training data is nearly half empty, but the test data is almost half empty. In other words, in order to improve accuracy, it is better to submit all columns as zero as soon as possible if t can be determined to be empty with certainty.","metadata":{}}]}