{
  "id": 312904,
  "title": "ResNet50 Result from Notebook Provided for PyTorch (Public F1=0.612 at Epoch 12)",
  "url": "/competitions/herbarium-2022-fgvc9/discussion/312904",
  "author_name": "John Park",
  "post_date": "2022-03-14T17:30:57.004000",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>Thank you for participating in this competition! We are seeing considerable efforts and trials in the past few weeks, which is very exciting. The current top models look quite promising, considering that we have more than two months left to go.</p>\n<p>In this topic, I would like to briefly share the result of training a standard CNN with the <a href=\"https://www.kaggle.com/brendanrappazzo/herbarium-2022\" target=\"_blank\">Notebook</a> provided by our co-host, <a href=\"https://www.kaggle.com/brendanrappazzo\" target=\"_blank\">@brendanrappazzo</a>. Brendan is the 3rd place winner of the Herbarium 2021 competition, and had joined us to host this year's competition. The purpose of this post is to demonstrate how the notebook runs. Also, we would like to provide some reference points on how much time it requires to train the Herbarium 2022 dataset. Another purpose is to see the performance expectations of standard models on the dataset. </p>\n<p><a href=\"https://pytorch.org/hub/pytorch_vision_resnet/\" target=\"_blank\">ResNet50 from Pytorch</a> with the default setting from <a href=\"https://www.kaggle.com/brendanrappazzo/herbarium-2022\" target=\"_blank\">Brendan's notebook</a> was used for this experiment.</p>\n<p>Specifics are:</p>\n<ol>\n<li>ResNet50 pre-trained on ImageNet </li>\n<li>Cross entropy loss and SGD optimizer with momentum </li>\n<li>Augmentation<ul>\n<li>Normalization on ImageNet mean and std as suggested by <a href=\"https://pytorch.org/vision/stable/models.html\" target=\"_blank\">the official PyTorch documentation</a></li>\n<li>Random rotation 20 degrees, horizontal flip</li></ul></li>\n<li>Image resized to 380 by 380</li>\n<li>Batch size =32, Proxy batch size=256 (Gradient accumulation and weight update for every 8 batches)</li>\n</ol>\n<p>One training per epoch takes about 6-7 hours on the Kaggle notebook GPU. So, you can technically train it on a Kaggle notebook, saving the model every epoch and reloading the model to train. In that case, using full GPU hours you will be able to train ResNet50 for 5 epochs. After 5 epochs of training, ResNet50 gives you 0.49 Public F1, which is a decent score considering the short training time and basic settings. </p>\n<p>I posted the results as benchmarks on the leaderboard (epochs 5 and 10). Here are some more inference results below: </p>\n<table>\n<thead>\n<tr>\n<th>Epoch</th>\n<th>Public-F1</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>5</td>\n<td>0.490</td>\n</tr>\n<tr>\n<td>6</td>\n<td>0.531</td>\n</tr>\n<tr>\n<td>9</td>\n<td>0.587</td>\n</tr>\n<tr>\n<td>10</td>\n<td>0.597</td>\n</tr>\n<tr>\n<td>11</td>\n<td>0.606</td>\n</tr>\n<tr>\n<td>12</td>\n<td>0.612</td>\n</tr>\n</tbody>\n</table>\n<p>That's it! I hope this is helpful. I'll try to come back with other topics that could be more helpful to everyone. Please feel free to copy and edit <a href=\"https://www.kaggle.com/brendanrappazzo/herbarium-2022\" target=\"_blank\">Brendan's notebook</a> if you are using PyTorch and don't know where to start with. </p>",
  "messages": [
    {
      "id": 1722594,
      "postDate": "2022-03-14T17:30:57.003Z",
      "content": "<p>Hi everyone,</p>\n<p>Thank you for participating in this competition! We are seeing considerable efforts and trials in the past few weeks, which is very exciting. The current top models look quite promising, considering that we have more than two months left to go.</p>\n<p>In this topic, I would like to briefly share the result of training a standard CNN with the <a href=\"https://www.kaggle.com/brendanrappazzo/herbarium-2022\" target=\"_blank\">Notebook</a> provided by our co-host, <a href=\"https://www.kaggle.com/brendanrappazzo\" target=\"_blank\">@brendanrappazzo</a>. Brendan is the 3rd place winner of the Herbarium 2021 competition, and had joined us to host this year's competition. The purpose of this post is to demonstrate how the notebook runs. Also, we would like to provide some reference points on how much time it requires to train the Herbarium 2022 dataset. Another purpose is to see the performance expectations of standard models on the dataset. </p>\n<p><a href=\"https://pytorch.org/hub/pytorch_vision_resnet/\" target=\"_blank\">ResNet50 from Pytorch</a> with the default setting from <a href=\"https://www.kaggle.com/brendanrappazzo/herbarium-2022\" target=\"_blank\">Brendan's notebook</a> was used for this experiment.</p>\n<p>Specifics are:</p>\n<ol>\n<li>ResNet50 pre-trained on ImageNet </li>\n<li>Cross entropy loss and SGD optimizer with momentum </li>\n<li>Augmentation<ul>\n<li>Normalization on ImageNet mean and std as suggested by <a href=\"https://pytorch.org/vision/stable/models.html\" target=\"_blank\">the official PyTorch documentation</a></li>\n<li>Random rotation 20 degrees, horizontal flip</li></ul></li>\n<li>Image resized to 380 by 380</li>\n<li>Batch size =32, Proxy batch size=256 (Gradient accumulation and weight update for every 8 batches)</li>\n</ol>\n<p>One training per epoch takes about 6-7 hours on the Kaggle notebook GPU. So, you can technically train it on a Kaggle notebook, saving the model every epoch and reloading the model to train. In that case, using full GPU hours you will be able to train ResNet50 for 5 epochs. After 5 epochs of training, ResNet50 gives you 0.49 Public F1, which is a decent score considering the short training time and basic settings. </p>\n<p>I posted the results as benchmarks on the leaderboard (epochs 5 and 10). Here are some more inference results below: </p>\n<table>\n<thead>\n<tr>\n<th>Epoch</th>\n<th>Public-F1</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>5</td>\n<td>0.490</td>\n</tr>\n<tr>\n<td>6</td>\n<td>0.531</td>\n</tr>\n<tr>\n<td>9</td>\n<td>0.587</td>\n</tr>\n<tr>\n<td>10</td>\n<td>0.597</td>\n</tr>\n<tr>\n<td>11</td>\n<td>0.606</td>\n</tr>\n<tr>\n<td>12</td>\n<td>0.612</td>\n</tr>\n</tbody>\n</table>\n<p>That's it! I hope this is helpful. I'll try to come back with other topics that could be more helpful to everyone. Please feel free to copy and edit <a href=\"https://www.kaggle.com/brendanrappazzo/herbarium-2022\" target=\"_blank\">Brendan's notebook</a> if you are using PyTorch and don't know where to start with. </p>",
      "rawMarkdown": "Hi everyone,\n\nThank you for participating in this competition! We are seeing considerable efforts and trials in the past few weeks, which is very exciting. The current top models look quite promising, considering that we have more than two months left to go.\n\nIn this topic, I would like to briefly share the result of training a standard CNN with the [Notebook](https://www.kaggle.com/brendanrappazzo/herbarium-2022) provided by our co-host, @brendanrappazzo. Brendan is the 3rd place winner of the Herbarium 2021 competition, and had joined us to host this year's competition. The purpose of this post is to demonstrate how the notebook runs. Also, we would like to provide some reference points on how much time it requires to train the Herbarium 2022 dataset. Another purpose is to see the performance expectations of standard models on the dataset. \n\n[ResNet50 from Pytorch](https://pytorch.org/hub/pytorch_vision_resnet/) with the default setting from [Brendan's notebook](https://www.kaggle.com/brendanrappazzo/herbarium-2022) was used for this experiment.\n\nSpecifics are:\n1. ResNet50 pre-trained on ImageNet \n2. Cross entropy loss and SGD optimizer with momentum \n3. Augmentation\n  -  Normalization on ImageNet mean and std as suggested by [the official PyTorch documentation](https://pytorch.org/vision/stable/models.html)\n  -  Random rotation 20 degrees, horizontal flip\n4. Image resized to 380 by 380\n4. Batch size =32, Proxy batch size=256 (Gradient accumulation and weight update for every 8 batches)\n\nOne training per epoch takes about 6-7 hours on the Kaggle notebook GPU. So, you can technically train it on a Kaggle notebook, saving the model every epoch and reloading the model to train. In that case, using full GPU hours you will be able to train ResNet50 for 5 epochs. After 5 epochs of training, ResNet50 gives you 0.49 Public F1, which is a decent score considering the short training time and basic settings. \n\nI posted the results as benchmarks on the leaderboard (epochs 5 and 10). Here are some more inference results below: \n\n| Epoch | Public-F1  |\n| --- | --- |\n|  5|  0.490|\n|  6|  0.531|\n|  9|  0.587|\n|  10|  0.597|\n|  11|  0.606|\n|  12|  0.612|\n\nThat's it! I hope this is helpful. I'll try to come back with other topics that could be more helpful to everyone. Please feel free to copy and edit [Brendan's notebook](https://www.kaggle.com/brendanrappazzo/herbarium-2022) if you are using PyTorch and don't know where to start with. \n\n\n\n",
      "votes": 5
    },
    {
      "id": 1722750,
      "postDate": "2022-03-14T20:39:25.090Z",
      "content": "<p>On a slightly off topic, can I say it's so awesome to see an amazingly active host and a winner join the next run as a host too! </p>\n<p>Thanks for all the efforts, John! I'm super interested in contributing and participating here in the next few weeks! </p>",
      "rawMarkdown": "On a slightly off topic, can I say it's so awesome to see an amazingly active host and a winner join the next run as a host too! \n\nThanks for all the efforts, John! I'm super interested in contributing and participating here in the next few weeks! ",
      "votes": 2,
      "replies": [
        {
          "id": 1723561,
          "postDate": "2022-03-15T14:18:39.610Z",
          "content": "<p>Sanyam, you are very welcome! It is my pleasure. I would love to see more people using the herbarium image datasets, because it is such an intriguing dataset collected for centuries. The best part for the vision community is probably its extensive expert labeling, which many other image datasets lack and suffer from. The world's herbaria have 30M images for 400K species, on all public domain, so we are only limited by storage space and computing power. So much more to expect in the future as well! </p>",
          "rawMarkdown": "Sanyam, you are very welcome! It is my pleasure. I would love to see more people using the herbarium image datasets, because it is such an intriguing dataset collected for centuries. The best part for the vision community is probably its extensive expert labeling, which many other image datasets lack and suffer from. The world's herbaria have 30M images for 400K species, on all public domain, so we are only limited by storage space and computing power. So much more to expect in the future as well! ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1732660,
      "postDate": "2022-03-23T15:46:42.293Z",
      "content": "<p>Nice reference. Can't wait to start. The notebook is 404, can you update the notebook URL? </p>",
      "rawMarkdown": "Nice reference. Can't wait to start. The notebook is 404, can you update the notebook URL? ",
      "replies": [
        {
          "id": 1732684,
          "postDate": "2022-03-23T15:53:47.047Z",
          "content": "<p>Here is the link to the <a href=\"https://www.kaggle.com/code/brendanrappazzo/herbarium-2022-starter-sample-code\" target=\"_blank\">notebook</a>, you can also find it in the Code tab of the competition.</p>",
          "rawMarkdown": "Here is the link to the [notebook](https://www.kaggle.com/code/brendanrappazzo/herbarium-2022-starter-sample-code), you can also find it in the Code tab of the competition."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1722750,
      "author_name": "Sanyam Bhutani",
      "author_url": "",
      "post_date": "2022-03-14T20:39:25.090000",
      "content": "<p>On a slightly off topic, can I say it's so awesome to see an amazingly active host and a winner join the next run as a host too! </p>\n<p>Thanks for all the efforts, John! I'm super interested in contributing and participating here in the next few weeks! </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1723561,
          "author_name": "John Park",
          "author_url": "",
          "post_date": "2022-03-15T14:18:39.610000",
          "content": "<p>Sanyam, you are very welcome! It is my pleasure. I would love to see more people using the herbarium image datasets, because it is such an intriguing dataset collected for centuries. The best part for the vision community is probably its extensive expert labeling, which many other image datasets lack and suffer from. The world's herbaria have 30M images for 400K species, on all public domain, so we are only limited by storage space and computing power. So much more to expect in the future as well! </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1732660,
      "author_name": "Billion Zheng",
      "author_url": "",
      "post_date": "2022-03-23T15:46:42.293000",
      "content": "<p>Nice reference. Can't wait to start. The notebook is 404, can you update the notebook URL? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1732684,
          "author_name": "Riccardo de Lutio",
          "author_url": "",
          "post_date": "2022-03-23T15:53:47.047000",
          "content": "<p>Here is the link to the <a href=\"https://www.kaggle.com/code/brendanrappazzo/herbarium-2022-starter-sample-code\" target=\"_blank\">notebook</a>, you can also find it in the Code tab of the competition.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1722594": "Hi everyone,\n\nThank you for participating in this competition! We are seeing considerable efforts and trials in the past few weeks, which is very exciting. The current top models look quite promising, considering that we have more than two months left to go.\n\nIn this topic, I would like to briefly share the result of training a standard CNN with the [Notebook](https://www.kaggle.com/brendanrappazzo/herbarium-2022) provided by our co-host, @brendanrappazzo. Brendan is the 3rd place winner of the Herbarium 2021 competition, and had joined us to host this year's competition. The purpose of this post is to demonstrate how the notebook runs. Also, we would like to provide some reference points on how much time it requires to train the Herbarium 2022 dataset. Another purpose is to see the performance expectations of standard models on the dataset. \n\n[ResNet50 from Pytorch](https://pytorch.org/hub/pytorch_vision_resnet/) with the default setting from [Brendan's notebook](https://www.kaggle.com/brendanrappazzo/herbarium-2022) was used for this experiment.\n\nSpecifics are:\n1. ResNet50 pre-trained on ImageNet \n2. Cross entropy loss and SGD optimizer with momentum \n3. Augmentation\n  -  Normalization on ImageNet mean and std as suggested by [the official PyTorch documentation](https://pytorch.org/vision/stable/models.html)\n  -  Random rotation 20 degrees, horizontal flip\n4. Image resized to 380 by 380\n4. Batch size =32, Proxy batch size=256 (Gradient accumulation and weight update for every 8 batches)\n\nOne training per epoch takes about 6-7 hours on the Kaggle notebook GPU. So, you can technically train it on a Kaggle notebook, saving the model every epoch and reloading the model to train. In that case, using full GPU hours you will be able to train ResNet50 for 5 epochs. After 5 epochs of training, ResNet50 gives you 0.49 Public F1, which is a decent score considering the short training time and basic settings. \n\nI posted the results as benchmarks on the leaderboard (epochs 5 and 10). Here are some more inference results below: \n\n| Epoch | Public-F1  |\n| --- | --- |\n|  5|  0.490|\n|  6|  0.531|\n|  9|  0.587|\n|  10|  0.597|\n|  11|  0.606|\n|  12|  0.612|\n\nThat's it! I hope this is helpful. I'll try to come back with other topics that could be more helpful to everyone. Please feel free to copy and edit [Brendan's notebook](https://www.kaggle.com/brendanrappazzo/herbarium-2022) if you are using PyTorch and don't know where to start with. \n\n\n\n",
    "1722750": "On a slightly off topic, can I say it's so awesome to see an amazingly active host and a winner join the next run as a host too! \n\nThanks for all the efforts, John! I'm super interested in contributing and participating here in the next few weeks! ",
    "1732660": "Nice reference. Can't wait to start. The notebook is 404, can you update the notebook URL? "
  }
}