{
  "id": 154416,
  "title": "5th - baseline with 11 epochs CE+SE-ResNext50",
  "url": "/competitions/herbarium-2020-fgvc7/writeups/epoch11-vanilla-srx50-5th-baseline-with-11-epochs-",
  "author_name": "",
  "post_date": "2020-05-28T16:37:38.793Z",
  "votes": 10,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Congrats to all the winners. </p>\n\n<p>I enjoyed the last year Herbarium2019 where I managed to beat Facebook AI team while using 100x less compute (relaying on kernels only) to train final solution so i hoped i could do something similar this year. Oh boy, how wrong i was....</p>\n\n<p>The size of the dataset knocked me out immediately. Still i decided to run one experiment and that's to rerun the same code form last year and then to play with post-processing :). </p>\n\n<h2>Training the baseline</h2>\n\n<p>I borrowed 1080Ti for extended weekend and run 11 epochs :). I just used cross entropy, trained on whole dataset, had only some small augmentations to enable fast learning, and did not treat imbalance at all. Input size was 1/2 of the original, used accumulation to get to 256 batch size and scaled LR with cosine annealing.  You could think of my solution as a default or a baseline. </p>\n\n<h2>Squeezing out the baseline - submit all you got</h2>\n\n<h3>Comparing checkpoints</h3>\n\n<p>I submitted after every epoch and tried to get better understanding of what's going on.\nLast two checkpoints scored 0.64+ on public. I noticed that as my score increased with later checkpoints the number of predicted unique classes also linearly increased. It went from epoch 5, 0.55 and 18K to epoch 11, 0.64 and 24K</p>\n\n<h3>Postprocess 1</h3>\n\n<p>I multiplied softmax outputs  by the inverse of class frequency and got boost 0.05 and reaached 0.69+ with 27K unique classes.  </p>\n\n<h3>Postprocess 2</h3>\n\n<p>The classes in test set had the same distribution as in training set with max clipped at 10.  I calculated the difference between the expected counts and my predicted counts, then i adjusted  the probabilities accordingly. </p>\n\n<p>For example if I expected a count 10 but got 15 I needed a reduction of 0.66% (adjustment). To prevent to much fluctuation I sqrt normalized the adjustments and repeated it several times. I followed number of uniquely predicted classes and stopped when it taps the 32K. This gave me 0.04+ so i reached ~0.74. Horizontal flip gave me a bit extra 0.005.</p>",
  "messages": [
    {
      "id": "864878",
      "postDate": "05/28/2020 08:43:31",
      "content": "<p>Congrats to all the winners. </p>\n\n<p>I enjoyed the last year Herbarium2019 where I managed to beat Facebook AI team while using 100x less compute (relaying on kernels only) to train final solution so i hoped i could do something similar this year. Oh boy, how wrong i was....</p>\n\n<p>The size of the dataset knocked me out immediately. Still i decided to run one experiment and that's to rerun the same code form last year and then to play with post-processing :). </p>\n\n<h2>Training the baseline</h2>\n\n<p>I borrowed 1080Ti for extended weekend and run 11 epochs :). I just used cross entropy, trained on whole dataset, had only some small augmentations to enable fast learning, and did not treat imbalance at all. Input size was 1/2 of the original, used accumulation to get to 256 batch size and scaled LR with cosine annealing.  You could think of my solution as a default or a baseline. </p>\n\n<h2>Squeezing out the baseline - submit all you got</h2>\n\n<h3>Comparing checkpoints</h3>\n\n<p>I submitted after every epoch and tried to get better understanding of what's going on.\nLast two checkpoints scored 0.64+ on public. I noticed that as my score increased with later checkpoints the number of predicted unique classes also linearly increased. It went from epoch 5, 0.55 and 18K to epoch 11, 0.64 and 24K</p>\n\n<h3>Postprocess 1</h3>\n\n<p>I multiplied softmax outputs  by the inverse of class frequency and got boost 0.05 and reaached 0.69+ with 27K unique classes.  </p>\n\n<h3>Postprocess 2</h3>\n\n<p>The classes in test set had the same distribution as in training set with max clipped at 10.  I calculated the difference between the expected counts and my predicted counts, then i adjusted  the probabilities accordingly. </p>\n\n<p>For example if I expected a count 10 but got 15 I needed a reduction of 0.66% (adjustment). To prevent to much fluctuation I sqrt normalized the adjustments and repeated it several times. I followed number of uniquely predicted classes and stopped when it taps the 32K. This gave me 0.04+ so i reached ~0.74. Horizontal flip gave me a bit extra 0.005.</p>",
      "rawMarkdown": "Congrats to all the winners. \n\nI enjoyed the last year Herbarium2019 where I managed to beat Facebook AI team while using 100x less compute (relaying on kernels only) to train final solution so i hoped i could do something similar this year. Oh boy, how wrong i was....\n\nThe size of the dataset knocked me out immediately. Still i decided to run one experiment and that's to rerun the same code form last year and then to play with post-processing :). \n\n## Training the baseline\nI borrowed 1080Ti for extended weekend and run 11 epochs :). I just used cross entropy, trained on whole dataset, had only some small augmentations to enable fast learning, and did not treat imbalance at all. Input size was 1/2 of the original, used accumulation to get to 256 batch size and scaled LR with cosine annealing.  You could think of my solution as a default or a baseline. \n\n## Squeezing out the baseline - submit all you got\n### Comparing checkpoints\nI submitted after every epoch and tried to get better understanding of what's going on.\nLast two checkpoints scored 0.64+ on public. I noticed that as my score increased with later checkpoints the number of predicted unique classes also linearly increased. It went from epoch 5, 0.55 and 18K to epoch 11, 0.64 and 24K\n    \n### Postprocess 1\nI multiplied softmax outputs  by the inverse of class frequency and got boost 0.05 and reaached 0.69+ with 27K unique classes.  \n\n### Postprocess 2\nThe classes in test set had the same distribution as in training set with max clipped at 10.  I calculated the difference between the expected counts and my predicted counts, then i adjusted  the probabilities accordingly. \n\nFor example if I expected a count 10 but got 15 I needed a reduction of 0.66% (adjustment). To prevent to much fluctuation I sqrt normalized the adjustments and repeated it several times. I followed number of uniquely predicted classes and stopped when it taps the 32K. This gave me 0.04+ so i reached ~0.74. Horizontal flip gave me a bit extra 0.005.",
      "votes": null
    },
    {
      "id": "865222",
      "postDate": "05/28/2020 13:38:42",
      "content": "<p>good job, and good fighting spirit. 😄\nits not easy to compete on a large dataset against dudes with lots of GPU power  </p>\n\n<p>a simple trick to tackle large number of classes is doing bottleneck FC before the softmax\\arcface. this can almost double your batch size, and allow you to better utilize the GPU.</p>",
      "rawMarkdown": "good job, and good fighting spirit. 😄\nits not easy to compete on a large dataset against dudes with lots of GPU power  \n\na simple trick to tackle large number of classes is doing bottleneck FC before the softmax\\arcface. this can almost double your batch size, and allow you to better utilize the GPU.",
      "votes": null
    },
    {
      "id": "865252",
      "postDate": "05/28/2020 13:58:42",
      "content": "<p>thanks, \nI hope one day i get a chance to ride like grown up dudes :D</p>\n\n<p>For me adding FC bottleneck (512) reduces GPU consumption for 15% only and epoch time 10% (this could be due to poor CPU)</p>",
      "rawMarkdown": "thanks, \nI hope one day i get a chance to ride like grown up dudes :D\n\nFor me adding FC bottleneck (512) reduces GPU consumption for 15% only and epoch time 10% (this could be due to poor CPU)",
      "votes": null
    },
    {
      "id": "867981",
      "postDate": "05/30/2020 19:45:43",
      "content": "<p>Hey, <a href=\"/valanm\">@valanm</a> </p>\n\n<p>I am asking this everyone here :)\nHow do you did validation? How did you split data to train/val?</p>",
      "rawMarkdown": "Hey, @valanm \n\nI am asking this everyone here :)\nHow do you did validation? How did you split data to train/val?",
      "votes": null
    },
    {
      "id": "871328",
      "postDate": "06/02/2020 09:15:14",
      "content": "<p>i had no dedicated validation set- i didn't do it because i did not have enough resources for more than one run :) </p>",
      "rawMarkdown": "i had no dedicated validation set- i didn't do it because i did not have enough resources for more than one run :)",
      "votes": null
    },
    {
      "id": "875569",
      "postDate": "06/05/2020 23:27:23",
      "content": "<p>Hi,\nHow are you dealing with the classifier layer? I mean if I use a single layer it will have 2048*32093 = 65758557 parameters which is 2times huge than the feature extractors.</p>",
      "rawMarkdown": "Hi,\nHow are you dealing with the classifier layer? I mean if I use a single layer it will have 2048*32093 = 65758557 parameters which is 2times huge than the feature extractors.",
      "votes": null
    },
    {
      "id": "880430",
      "postDate": "06/10/2020 09:43:10",
      "content": "<p>FC params are much faster. you can also add FC 512 layer between 2048 and 32093 and reduce the number of params and hence increase the speed </p>",
      "rawMarkdown": "FC params are much faster. you can also add FC 512 layer between 2048 and 32093 and reduce the number of params and hence increase the speed",
      "votes": null
    },
    {
      "id": "880491",
      "postDate": "06/10/2020 10:46:25",
      "content": "<p>Thank you for your reply!! In fact at the very beginning I add a FC128 layer between features and results which after 7 epochs of training it only got 0.03 accuracy on public test cases😑 </p>",
      "rawMarkdown": "Thank you for your reply!! In fact at the very beginning I add a FC128 layer between features and results which after 7 epochs of training it only got 0.03 accuracy on public test cases😑",
      "votes": null
    },
    {
      "id": "880504",
      "postDate": "06/10/2020 10:55:38",
      "content": "<p>And finally I trained the model SEResNext50 with FC layer (2048* 32093) while after 8 epochs the public score was only around 0.15. Do you have any idea about the result?\nI used standard augmentations RandomFlip and resized crops...... The resolution was 332* 500. </p>",
      "rawMarkdown": "And finally I trained the model SEResNext50 with FC layer (2048* 32093) while after 8 epochs the public score was only around 0.15. Do you have any idea about the result?\nI used standard augmentations RandomFlip and resized crops...... The resolution was 332* 500.",
      "votes": null
    },
    {
      "id": "880566",
      "postDate": "06/10/2020 11:50:02",
      "content": "<p>find smaller dataset and experiment. there are many things that could go wrong. And you certainly did something wrong because i used the same arch and after one epoch i could score ~0.25</p>",
      "rawMarkdown": "find smaller dataset and experiment. there are many things that could go wrong. And you certainly did something wrong because i used the same arch and after one epoch i could score ~0.25",
      "votes": null
    },
    {
      "id": "880583",
      "postDate": "06/10/2020 11:59:54",
      "content": "<p>☹️ Thanks. I will have another look</p>",
      "rawMarkdown": "☹️ Thanks. I will have another look",
      "votes": null
    },
    {
      "id": "880585",
      "postDate": "06/10/2020 12:02:44",
      "content": "<p>I used mixed precision training since it saves half of the VRAM. Perhaps that's the reason, I will retrain the model with fp32 when my 2080Ti arrived.</p>",
      "rawMarkdown": "I used mixed precision training since it saves half of the VRAM. Perhaps that's the reason, I will retrain the model with fp32 when my 2080Ti arrived.",
      "votes": null
    },
    {
      "id": "880587",
      "postDate": "06/10/2020 12:04:47",
      "content": "<p>fp32 or fp16 it should not be so much difference - it gotta be something else</p>",
      "rawMarkdown": "fp32 or fp16 it should not be so much difference - it gotta be something else",
      "votes": null
    },
    {
      "id": "880589",
      "postDate": "06/10/2020 12:07:11",
      "content": "<p>😣 I will try with a pytorch pertained model instead of the net I implemented my self.</p>",
      "rawMarkdown": "😣 I will try with a pytorch pertained model instead of the net I implemented my self.",
      "votes": null
    },
    {
      "id": "1118237",
      "postDate": "12/18/2020 21:28:19",
      "content": "<p>did you find out what you got poor results?</p>",
      "rawMarkdown": "did you find out what you got poor results?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 865222,
      "author_name": "mrt23564",
      "author_url": "",
      "post_date": "05/28/2020 13:38:42",
      "content": "<p>good job, and good fighting spirit. 😄\nits not easy to compete on a large dataset against dudes with lots of GPU power  </p>\n\n<p>a simple trick to tackle large number of classes is doing bottleneck FC before the softmax\\arcface. this can almost double your batch size, and allow you to better utilize the GPU.</p>",
      "votes": null,
      "replies": [
        {
          "id": 865252,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "05/28/2020 13:58:42",
          "content": "<p>thanks, \nI hope one day i get a chance to ride like grown up dudes :D</p>\n\n<p>For me adding FC bottleneck (512) reduces GPU consumption for 15% only and epoch time 10% (this could be due to poor CPU)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 867981,
      "author_name": "discoholic",
      "author_url": "",
      "post_date": "05/30/2020 19:45:43",
      "content": "<p>Hey, <a href=\"/valanm\">@valanm</a> </p>\n\n<p>I am asking this everyone here :)\nHow do you did validation? How did you split data to train/val?</p>",
      "votes": null,
      "replies": [
        {
          "id": 871328,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "06/02/2020 09:15:14",
          "content": "<p>i had no dedicated validation set- i didn't do it because i did not have enough resources for more than one run :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 875569,
      "author_name": "chantlevdai",
      "author_url": "",
      "post_date": "06/05/2020 23:27:23",
      "content": "<p>Hi,\nHow are you dealing with the classifier layer? I mean if I use a single layer it will have 2048*32093 = 65758557 parameters which is 2times huge than the feature extractors.</p>",
      "votes": null,
      "replies": [
        {
          "id": 880430,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "06/10/2020 09:43:10",
          "content": "<p>FC params are much faster. you can also add FC 512 layer between 2048 and 32093 and reduce the number of params and hence increase the speed </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 880491,
          "author_name": "chantlevdai",
          "author_url": "",
          "post_date": "06/10/2020 10:46:25",
          "content": "<p>Thank you for your reply!! In fact at the very beginning I add a FC128 layer between features and results which after 7 epochs of training it only got 0.03 accuracy on public test cases😑 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 880504,
          "author_name": "chantlevdai",
          "author_url": "",
          "post_date": "06/10/2020 10:55:38",
          "content": "<p>And finally I trained the model SEResNext50 with FC layer (2048* 32093) while after 8 epochs the public score was only around 0.15. Do you have any idea about the result?\nI used standard augmentations RandomFlip and resized crops...... The resolution was 332* 500. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 880566,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "06/10/2020 11:50:02",
          "content": "<p>find smaller dataset and experiment. there are many things that could go wrong. And you certainly did something wrong because i used the same arch and after one epoch i could score ~0.25</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 880583,
          "author_name": "chantlevdai",
          "author_url": "",
          "post_date": "06/10/2020 11:59:54",
          "content": "<p>☹️ Thanks. I will have another look</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 880585,
          "author_name": "chantlevdai",
          "author_url": "",
          "post_date": "06/10/2020 12:02:44",
          "content": "<p>I used mixed precision training since it saves half of the VRAM. Perhaps that's the reason, I will retrain the model with fp32 when my 2080Ti arrived.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 880587,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "06/10/2020 12:04:47",
          "content": "<p>fp32 or fp16 it should not be so much difference - it gotta be something else</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 880589,
          "author_name": "chantlevdai",
          "author_url": "",
          "post_date": "06/10/2020 12:07:11",
          "content": "<p>😣 I will try with a pytorch pertained model instead of the net I implemented my self.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118237,
          "author_name": "thejravichandran",
          "author_url": "",
          "post_date": "12/18/2020 21:28:19",
          "content": "<p>did you find out what you got poor results?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "864878": "Congrats to all the winners. \n\nI enjoyed the last year Herbarium2019 where I managed to beat Facebook AI team while using 100x less compute (relaying on kernels only) to train final solution so i hoped i could do something similar this year. Oh boy, how wrong i was....\n\nThe size of the dataset knocked me out immediately. Still i decided to run one experiment and that's to rerun the same code form last year and then to play with post-processing :). \n\n## Training the baseline\nI borrowed 1080Ti for extended weekend and run 11 epochs :). I just used cross entropy, trained on whole dataset, had only some small augmentations to enable fast learning, and did not treat imbalance at all. Input size was 1/2 of the original, used accumulation to get to 256 batch size and scaled LR with cosine annealing.  You could think of my solution as a default or a baseline. \n\n## Squeezing out the baseline - submit all you got\n### Comparing checkpoints\nI submitted after every epoch and tried to get better understanding of what's going on.\nLast two checkpoints scored 0.64+ on public. I noticed that as my score increased with later checkpoints the number of predicted unique classes also linearly increased. It went from epoch 5, 0.55 and 18K to epoch 11, 0.64 and 24K\n    \n### Postprocess 1\nI multiplied softmax outputs  by the inverse of class frequency and got boost 0.05 and reaached 0.69+ with 27K unique classes.  \n\n### Postprocess 2\nThe classes in test set had the same distribution as in training set with max clipped at 10.  I calculated the difference between the expected counts and my predicted counts, then i adjusted  the probabilities accordingly. \n\nFor example if I expected a count 10 but got 15 I needed a reduction of 0.66% (adjustment). To prevent to much fluctuation I sqrt normalized the adjustments and repeated it several times. I followed number of uniquely predicted classes and stopped when it taps the 32K. This gave me 0.04+ so i reached ~0.74. Horizontal flip gave me a bit extra 0.005.",
    "865222": "good job, and good fighting spirit. 😄\nits not easy to compete on a large dataset against dudes with lots of GPU power  \n\na simple trick to tackle large number of classes is doing bottleneck FC before the softmax\\arcface. this can almost double your batch size, and allow you to better utilize the GPU.",
    "865252": "thanks, \nI hope one day i get a chance to ride like grown up dudes :D\n\nFor me adding FC bottleneck (512) reduces GPU consumption for 15% only and epoch time 10% (this could be due to poor CPU)",
    "867981": "Hey, @valanm \n\nI am asking this everyone here :)\nHow do you did validation? How did you split data to train/val?",
    "871328": "i had no dedicated validation set- i didn't do it because i did not have enough resources for more than one run :)",
    "875569": "Hi,\nHow are you dealing with the classifier layer? I mean if I use a single layer it will have 2048*32093 = 65758557 parameters which is 2times huge than the feature extractors.",
    "880430": "FC params are much faster. you can also add FC 512 layer between 2048 and 32093 and reduce the number of params and hence increase the speed",
    "880491": "Thank you for your reply!! In fact at the very beginning I add a FC128 layer between features and results which after 7 epochs of training it only got 0.03 accuracy on public test cases😑",
    "880504": "And finally I trained the model SEResNext50 with FC layer (2048* 32093) while after 8 epochs the public score was only around 0.15. Do you have any idea about the result?\nI used standard augmentations RandomFlip and resized crops...... The resolution was 332* 500.",
    "880566": "find smaller dataset and experiment. there are many things that could go wrong. And you certainly did something wrong because i used the same arch and after one epoch i could score ~0.25",
    "880583": "☹️ Thanks. I will have another look",
    "880585": "I used mixed precision training since it saves half of the VRAM. Perhaps that's the reason, I will retrain the model with fp32 when my 2080Ti arrived.",
    "880587": "fp32 or fp16 it should not be so much difference - it gotta be something else",
    "880589": "😣 I will try with a pytorch pertained model instead of the net I implemented my self.",
    "1118237": "did you find out what you got poor results?"
  },
  "source": "meta"
}