{
  "id": 142897,
  "title": "CE vanilla srx50 9 epochs = 0.698",
  "url": "/competitions/herbarium-2020-fgvc7/discussion/142897",
  "author_name": "",
  "post_date": "2020-04-12T18:06:16.487062100Z",
  "votes": 2,
  "comment_count": 16,
  "views": 0,
  "content": "<p>It seems to me that heavy lifting will decide this comp. </p>\n\n<p>i trained for two epochs and got 0.45 with srx50 and vanilla softmax. I got impression that 0.7 is doable with a standard CNN setting. Am I wrong? Anyone wanna confirm this? \nBesides, anyone willing to share hints or point to good reads for those who are here to learn :) </p>",
  "messages": [
    {
      "id": "805473",
      "postDate": "04/12/2020 18:06:16",
      "content": "<p>It seems to me that heavy lifting will decide this comp. </p>\n\n<p>i trained for two epochs and got 0.45 with srx50 and vanilla softmax. I got impression that 0.7 is doable with a standard CNN setting. Am I wrong? Anyone wanna confirm this? \nBesides, anyone willing to share hints or point to good reads for those who are here to learn :) </p>",
      "rawMarkdown": "It seems to me that heavy lifting will decide this comp. \n\ni trained for two epochs and got 0.45 with srx50 and vanilla softmax. I got impression that 0.7 is doable with a standard CNN setting. Am I wrong? Anyone wanna confirm this? \nBesides, anyone willing to share hints or point to good reads for those who are here to learn :)",
      "votes": null
    },
    {
      "id": "808550",
      "postDate": "04/15/2020 13:27:15",
      "content": "<p>ooh yeeaaah 0.7 is doable :) 5epochs give my current score 0.65.</p>",
      "rawMarkdown": "ooh yeeaaah 0.7 is doable :) 5epochs give my current score 0.65.",
      "votes": null
    },
    {
      "id": "809435",
      "postDate": "04/16/2020 07:28:05",
      "content": "<p>Nooice! What network is srx50 and what are loss function are you using? I'm struggling to get decent results with CE but I think I may have a bug somewhere.</p>",
      "rawMarkdown": "Nooice! What network is srx50 and what are loss function are you using? I'm struggling to get decent results with CE but I think I may have a bug somewhere.",
      "votes": null
    },
    {
      "id": "809465",
      "postDate": "04/16/2020 08:07:43",
      "content": "<p>it's CE. srx50 is good old SE-ResNext-50 - my favorite by far :)</p>\n\n<p>I wanna train a model that will score within 10% from the winning model while using 100x+ less resources. And all of this without experimenting but by using a simple baseline without tricks. That's my motivation</p>",
      "rawMarkdown": "it's CE. srx50 is good old SE-ResNext-50 - my favorite by far :)\n\nI wanna train a model that will score within 10% from the winning model while using 100x+ less resources. And all of this without experimenting but by using a simple baseline without tricks. That's my motivation",
      "votes": null
    },
    {
      "id": "818101",
      "postDate": "04/23/2020 16:18:20",
      "content": "<p>What image size are you using?</p>",
      "rawMarkdown": "What image size are you using?",
      "votes": null
    },
    {
      "id": "818183",
      "postDate": "04/23/2020 17:15:21",
      "content": "<p>Image size is 1/2 of the original. </p>",
      "rawMarkdown": "Image size is 1/2 of the original.",
      "votes": null
    },
    {
      "id": "842920",
      "postDate": "05/11/2020 17:46:01",
      "content": "<p>tricks are the heart of deep learning, you should use them :)</p>\n\n<p>not to discourage you, cause your goals are pretty neat, but our current top-1 model is faster then SE-ResNext-50</p>",
      "rawMarkdown": "tricks are the heart of deep learning, you should use them :)\n\nnot to discourage you, cause your goals are pretty neat, but our current top-1 model is faster then SE-ResNext-50",
      "votes": null
    },
    {
      "id": "842980",
      "postDate": "05/11/2020 18:28:44",
      "content": "<p>Thanks for the comment. Fully agree with your first sentence. Unfortunately, not everybody has resources for experimenting on 1M datasets with 1000px images and it's hard to know which tricks to use out of the box.</p>\n\n<p>Perhaps i was not clear enough, when i said resources  i meant resources for training. I would be happy to reach 73% with 2 days training on 1080Ti and no TTA, or postprocessing tricks. The \"100x less resources\" comment was a bait to get comments from guys at the top. :)</p>\n\n<p>Would i be wrong if i say \"you use your model\" :) ? Congrats on the score.</p>",
      "rawMarkdown": "Thanks for the comment. Fully agree with your first sentence. Unfortunately, not everybody has resources for experimenting on 1M datasets with 1000px images and it's hard to know which tricks to use out of the box.\n\nPerhaps i was not clear enough, when i said resources  i meant resources for training. I would be happy to reach 73% with 2 days training on 1080Ti and no TTA, or postprocessing tricks. The \"100x less resources\" comment was a bait to get comments from guys at the top. :)\n\nWould i be wrong if i say \"you use your model\" :) ? Congrats on the score.",
      "votes": null
    },
    {
      "id": "845324",
      "postDate": "05/13/2020 06:41:47",
      "content": "<p>indeed, part of our participation is trying to promote \"our model\".\nsimilar to the dudes who won ifood 2019 last year\n<a href=\"https://github.com/clovaai/assembled-cnn\">https://github.com/clovaai/assembled-cnn</a></p>\n\n<p>all the ResNext models (and efficientNet) were optimized to reduce FLOPS. however, GPUs are limited by memory access, not by FLOPS (number of calculation).\nhence, ResNext and efficientNet are slow on GPUs. real slow.</p>\n\n<p>congrats on you place in planet pathology. you are doing there currently better than me.\nI am quite frustrated there from the test set, which contains only (public) 900 pics. this makes it hard to do proper optimization, the way we can do in this competition.</p>",
      "rawMarkdown": "indeed, part of our participation is trying to promote \"our model\".\nsimilar to the dudes who won ifood 2019 last year\nhttps://github.com/clovaai/assembled-cnn\n\nall the ResNext models (and efficientNet) were optimized to reduce FLOPS. however, GPUs are limited by memory access, not by FLOPS (number of calculation).\nhence, ResNext and efficientNet are slow on GPUs. real slow.\n\ncongrats on you place in planet pathology. you are doing there currently better than me.\nI am quite frustrated there from the test set, which contains only (public) 900 pics. this makes it hard to do proper optimization, the way we can do in this competition.",
      "votes": null
    },
    {
      "id": "845442",
      "postDate": "05/13/2020 07:58:37",
      "content": "<p>Thanks for the comment. </p>\n\n<p>IMHO, Clova AI violated the rules in the ifood 2019 by using and not reporting external data before the deadline (it seems to me that they did not understand the rules).  Although I clearly indicated this, the organizers did not address the issue which is very disappointing. <a href=\"https://www.kaggle.com/c/ifood-2019-fgvc6/discussion/94425\">discussion</a>. </p>\n\n<p>Please note that my comment has nothing to do with the paper which i warmly recommend. </p>\n\n<p>Not willingly, my interests are in cutting the training time as i rely on free kaggle kernels (no local or cloud machines :) ). For instance, I used 9 epochs for Herbarium 2019 with srx101 (trained in a single kernel run) and managed to beat Facebook team who trained FixSenet154 for 150 epochs and larger input size :). </p>\n\n<p>As for plant pathology, we all overfit to the public LB :). That dataset is small enough to be used in kernels and that's why i am more competitive there. Simply, I was able to experiment more and try more tricks.</p>",
      "rawMarkdown": "Thanks for the comment. \n\nIMHO, Clova AI violated the rules in the ifood 2019 by using and not reporting external data before the deadline (it seems to me that they did not understand the rules).  Although I clearly indicated this, the organizers did not address the issue which is very disappointing. [discussion](https://www.kaggle.com/c/ifood-2019-fgvc6/discussion/94425 ). \n\nPlease note that my comment has nothing to do with the paper which i warmly recommend. \n\nNot willingly, my interests are in cutting the training time as i rely on free kaggle kernels (no local or cloud machines :) ). For instance, I used 9 epochs for Herbarium 2019 with srx101 (trained in a single kernel run) and managed to beat Facebook team who trained FixSenet154 for 150 epochs and larger input size :). \n\nAs for plant pathology, we all overfit to the public LB :). That dataset is small enough to be used in kernels and that's why i am more competitive there. Simply, I was able to experiment more and try more tricks.",
      "votes": null
    },
    {
      "id": "852388",
      "postDate": "05/18/2020 12:02:46",
      "content": "<p>Congrats on your position in Plant Pathology competition :)</p>\n\n<p>As for Herbarium 2020 competition, it clearly states that you can use pre-trained models from ImageNet, iNaturalist 2017-2018 or Herbarium 2019.\nWe are using pre-trains from ImageNet and iNaturalist 2018. we got almost similar scores, with iNat18 having slighly better scores on the public LB.</p>",
      "rawMarkdown": "Congrats on your position in Plant Pathology competition :)\n\nAs for Herbarium 2020 competition, it clearly states that you can use pre-trained models from ImageNet, iNaturalist 2017-2018 or Herbarium 2019.\nWe are using pre-trains from ImageNet and iNaturalist 2018. we got almost similar scores, with iNat18 having slighly better scores on the public LB.",
      "votes": null
    },
    {
      "id": "852424",
      "postDate": "05/18/2020 12:38:47",
      "content": "<p>thx</p>\n\n<p>was Herbarium 2019 helpful?</p>",
      "rawMarkdown": "thx\n\n was Herbarium 2019 helpful?",
      "votes": null
    },
    {
      "id": "852607",
      "postDate": "05/18/2020 14:54:47",
      "content": "<p>Haven't used Herbarium 2019 pretrain at all.</p>",
      "rawMarkdown": "Haven't used Herbarium 2019 pretrain at all.",
      "votes": null
    },
    {
      "id": "854318",
      "postDate": "05/20/2020 00:36:17",
      "content": "<p>Hi,</p>\n\n<p>I also joined just to get experience with bigger datasets. Could I ask some details on your training? For the 0.45 at two epochs, is this just all the data at 500x500 pixels, no data balancing? How long did that take on the 1080 TI? 10 hrs per epoch?</p>\n\n<p>Also, are you just training the head, or the whole model?</p>\n\n<p>Thanks,\nDaniel</p>",
      "rawMarkdown": "Hi,\n\nI also joined just to get experience with bigger datasets. Could I ask some details on your training? For the 0.45 at two epochs, is this just all the data at 500x500 pixels, no data balancing? How long did that take on the 1080 TI? 10 hrs per epoch?\n\nAlso, are you just training the head, or the whole model?\n\nThanks,\nDaniel",
      "votes": null
    },
    {
      "id": "862823",
      "postDate": "05/26/2020 21:17:27",
      "content": "<p>Hello. I am beginner in ML, please, could you say what means \"CE vanilla\", especially interesting what means \"CE\" abbreviation. I used Google, but seems like it not helps me much (Cross-Entropy, Context Encoder, Checkpoint Ensembles)?. <br>\nThanks.</p>",
      "rawMarkdown": "Hello. I am beginner in ML, please, could you say what means \"CE vanilla\", especially interesting what means \"CE\" abbreviation. I used Google, but seems like it not helps me much (Cross-Entropy, Context Encoder, Checkpoint Ensembles)?.  \nThanks.",
      "votes": null
    },
    {
      "id": "863273",
      "postDate": "05/27/2020 07:38:25",
      "content": "<p>Cross-Entropy</p>",
      "rawMarkdown": "Cross-Entropy",
      "votes": null
    },
    {
      "id": "863275",
      "postDate": "05/27/2020 07:42:02",
      "content": "<p><a href=\"/hussam789\">@hussam789</a> and <a href=\"/mrt23564\">@mrt23564</a> congrats. looking forward to your solution. :) </p>\n\n<p>I gave a single run to Tresnet-m448 but i have seen no accuracy improvements over srx50 (very similar scores). But the training speed and batch_size was a nice surprise :)</p>",
      "rawMarkdown": "hussam789 and @mrt23564 congrats. looking forward to your solution. :) \n\nI gave a single run to Tresnet-m448 but i have seen no accuracy improvements over srx50 (very similar scores). But the training speed and batch_size was a nice surprise :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 808550,
      "author_name": "valanm",
      "author_url": "",
      "post_date": "04/15/2020 13:27:15",
      "content": "<p>ooh yeeaaah 0.7 is doable :) 5epochs give my current score 0.65.</p>",
      "votes": null,
      "replies": [
        {
          "id": 809435,
          "author_name": "shaunspinelli",
          "author_url": "",
          "post_date": "04/16/2020 07:28:05",
          "content": "<p>Nooice! What network is srx50 and what are loss function are you using? I'm struggling to get decent results with CE but I think I may have a bug somewhere.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 809465,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "04/16/2020 08:07:43",
          "content": "<p>it's CE. srx50 is good old SE-ResNext-50 - my favorite by far :)</p>\n\n<p>I wanna train a model that will score within 10% from the winning model while using 100x+ less resources. And all of this without experimenting but by using a simple baseline without tricks. That's my motivation</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 842920,
          "author_name": "mrt23564",
          "author_url": "",
          "post_date": "05/11/2020 17:46:01",
          "content": "<p>tricks are the heart of deep learning, you should use them :)</p>\n\n<p>not to discourage you, cause your goals are pretty neat, but our current top-1 model is faster then SE-ResNext-50</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 842980,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "05/11/2020 18:28:44",
          "content": "<p>Thanks for the comment. Fully agree with your first sentence. Unfortunately, not everybody has resources for experimenting on 1M datasets with 1000px images and it's hard to know which tricks to use out of the box.</p>\n\n<p>Perhaps i was not clear enough, when i said resources  i meant resources for training. I would be happy to reach 73% with 2 days training on 1080Ti and no TTA, or postprocessing tricks. The \"100x less resources\" comment was a bait to get comments from guys at the top. :)</p>\n\n<p>Would i be wrong if i say \"you use your model\" :) ? Congrats on the score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 845324,
          "author_name": "mrt23564",
          "author_url": "",
          "post_date": "05/13/2020 06:41:47",
          "content": "<p>indeed, part of our participation is trying to promote \"our model\".\nsimilar to the dudes who won ifood 2019 last year\n<a href=\"https://github.com/clovaai/assembled-cnn\">https://github.com/clovaai/assembled-cnn</a></p>\n\n<p>all the ResNext models (and efficientNet) were optimized to reduce FLOPS. however, GPUs are limited by memory access, not by FLOPS (number of calculation).\nhence, ResNext and efficientNet are slow on GPUs. real slow.</p>\n\n<p>congrats on you place in planet pathology. you are doing there currently better than me.\nI am quite frustrated there from the test set, which contains only (public) 900 pics. this makes it hard to do proper optimization, the way we can do in this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 845442,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "05/13/2020 07:58:37",
          "content": "<p>Thanks for the comment. </p>\n\n<p>IMHO, Clova AI violated the rules in the ifood 2019 by using and not reporting external data before the deadline (it seems to me that they did not understand the rules).  Although I clearly indicated this, the organizers did not address the issue which is very disappointing. <a href=\"https://www.kaggle.com/c/ifood-2019-fgvc6/discussion/94425\">discussion</a>. </p>\n\n<p>Please note that my comment has nothing to do with the paper which i warmly recommend. </p>\n\n<p>Not willingly, my interests are in cutting the training time as i rely on free kaggle kernels (no local or cloud machines :) ). For instance, I used 9 epochs for Herbarium 2019 with srx101 (trained in a single kernel run) and managed to beat Facebook team who trained FixSenet154 for 150 epochs and larger input size :). </p>\n\n<p>As for plant pathology, we all overfit to the public LB :). That dataset is small enough to be used in kernels and that's why i am more competitive there. Simply, I was able to experiment more and try more tricks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 852388,
          "author_name": "hussam789",
          "author_url": "",
          "post_date": "05/18/2020 12:02:46",
          "content": "<p>Congrats on your position in Plant Pathology competition :)</p>\n\n<p>As for Herbarium 2020 competition, it clearly states that you can use pre-trained models from ImageNet, iNaturalist 2017-2018 or Herbarium 2019.\nWe are using pre-trains from ImageNet and iNaturalist 2018. we got almost similar scores, with iNat18 having slighly better scores on the public LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 852424,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "05/18/2020 12:38:47",
          "content": "<p>thx</p>\n\n<p>was Herbarium 2019 helpful?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 852607,
          "author_name": "hussam789",
          "author_url": "",
          "post_date": "05/18/2020 14:54:47",
          "content": "<p>Haven't used Herbarium 2019 pretrain at all.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 863275,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "05/27/2020 07:42:02",
          "content": "<p><a href=\"/hussam789\">@hussam789</a> and <a href=\"/mrt23564\">@mrt23564</a> congrats. looking forward to your solution. :) </p>\n\n<p>I gave a single run to Tresnet-m448 but i have seen no accuracy improvements over srx50 (very similar scores). But the training speed and batch_size was a nice surprise :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 818101,
      "author_name": "marcelosanchezortega",
      "author_url": "",
      "post_date": "04/23/2020 16:18:20",
      "content": "<p>What image size are you using?</p>",
      "votes": null,
      "replies": [
        {
          "id": 818183,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "04/23/2020 17:15:21",
          "content": "<p>Image size is 1/2 of the original. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 854318,
      "author_name": "datadote",
      "author_url": "",
      "post_date": "05/20/2020 00:36:17",
      "content": "<p>Hi,</p>\n\n<p>I also joined just to get experience with bigger datasets. Could I ask some details on your training? For the 0.45 at two epochs, is this just all the data at 500x500 pixels, no data balancing? How long did that take on the 1080 TI? 10 hrs per epoch?</p>\n\n<p>Also, are you just training the head, or the whole model?</p>\n\n<p>Thanks,\nDaniel</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 862823,
      "author_name": "aliaksandr960",
      "author_url": "",
      "post_date": "05/26/2020 21:17:27",
      "content": "<p>Hello. I am beginner in ML, please, could you say what means \"CE vanilla\", especially interesting what means \"CE\" abbreviation. I used Google, but seems like it not helps me much (Cross-Entropy, Context Encoder, Checkpoint Ensembles)?. <br>\nThanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 863273,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "05/27/2020 07:38:25",
          "content": "<p>Cross-Entropy</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "805473": "It seems to me that heavy lifting will decide this comp. \n\ni trained for two epochs and got 0.45 with srx50 and vanilla softmax. I got impression that 0.7 is doable with a standard CNN setting. Am I wrong? Anyone wanna confirm this? \nBesides, anyone willing to share hints or point to good reads for those who are here to learn :)",
    "808550": "ooh yeeaaah 0.7 is doable :) 5epochs give my current score 0.65.",
    "809435": "Nooice! What network is srx50 and what are loss function are you using? I'm struggling to get decent results with CE but I think I may have a bug somewhere.",
    "809465": "it's CE. srx50 is good old SE-ResNext-50 - my favorite by far :)\n\nI wanna train a model that will score within 10% from the winning model while using 100x+ less resources. And all of this without experimenting but by using a simple baseline without tricks. That's my motivation",
    "818101": "What image size are you using?",
    "818183": "Image size is 1/2 of the original.",
    "842920": "tricks are the heart of deep learning, you should use them :)\n\nnot to discourage you, cause your goals are pretty neat, but our current top-1 model is faster then SE-ResNext-50",
    "842980": "Thanks for the comment. Fully agree with your first sentence. Unfortunately, not everybody has resources for experimenting on 1M datasets with 1000px images and it's hard to know which tricks to use out of the box.\n\nPerhaps i was not clear enough, when i said resources  i meant resources for training. I would be happy to reach 73% with 2 days training on 1080Ti and no TTA, or postprocessing tricks. The \"100x less resources\" comment was a bait to get comments from guys at the top. :)\n\nWould i be wrong if i say \"you use your model\" :) ? Congrats on the score.",
    "845324": "indeed, part of our participation is trying to promote \"our model\".\nsimilar to the dudes who won ifood 2019 last year\nhttps://github.com/clovaai/assembled-cnn\n\nall the ResNext models (and efficientNet) were optimized to reduce FLOPS. however, GPUs are limited by memory access, not by FLOPS (number of calculation).\nhence, ResNext and efficientNet are slow on GPUs. real slow.\n\ncongrats on you place in planet pathology. you are doing there currently better than me.\nI am quite frustrated there from the test set, which contains only (public) 900 pics. this makes it hard to do proper optimization, the way we can do in this competition.",
    "845442": "Thanks for the comment. \n\nIMHO, Clova AI violated the rules in the ifood 2019 by using and not reporting external data before the deadline (it seems to me that they did not understand the rules).  Although I clearly indicated this, the organizers did not address the issue which is very disappointing. [discussion](https://www.kaggle.com/c/ifood-2019-fgvc6/discussion/94425 ). \n\nPlease note that my comment has nothing to do with the paper which i warmly recommend. \n\nNot willingly, my interests are in cutting the training time as i rely on free kaggle kernels (no local or cloud machines :) ). For instance, I used 9 epochs for Herbarium 2019 with srx101 (trained in a single kernel run) and managed to beat Facebook team who trained FixSenet154 for 150 epochs and larger input size :). \n\nAs for plant pathology, we all overfit to the public LB :). That dataset is small enough to be used in kernels and that's why i am more competitive there. Simply, I was able to experiment more and try more tricks.",
    "852388": "Congrats on your position in Plant Pathology competition :)\n\nAs for Herbarium 2020 competition, it clearly states that you can use pre-trained models from ImageNet, iNaturalist 2017-2018 or Herbarium 2019.\nWe are using pre-trains from ImageNet and iNaturalist 2018. we got almost similar scores, with iNat18 having slighly better scores on the public LB.",
    "852424": "thx\n\n was Herbarium 2019 helpful?",
    "852607": "Haven't used Herbarium 2019 pretrain at all.",
    "854318": "Hi,\n\nI also joined just to get experience with bigger datasets. Could I ask some details on your training? For the 0.45 at two epochs, is this just all the data at 500x500 pixels, no data balancing? How long did that take on the 1080 TI? 10 hrs per epoch?\n\nAlso, are you just training the head, or the whole model?\n\nThanks,\nDaniel",
    "862823": "Hello. I am beginner in ML, please, could you say what means \"CE vanilla\", especially interesting what means \"CE\" abbreviation. I used Google, but seems like it not helps me much (Cross-Entropy, Context Encoder, Checkpoint Ensembles)?.  \nThanks.",
    "863273": "Cross-Entropy",
    "863275": "hussam789 and @mrt23564 congrats. looking forward to your solution. :) \n\nI gave a single run to Tresnet-m448 but i have seen no accuracy improvements over srx50 (very similar scores). But the training speed and batch_size was a nice surprise :)"
  },
  "source": "meta"
}