{
  "id": 226645,
  "title": "the myth of resnet200d",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/226645",
  "author_name": "hengck23",
  "post_date": "2021-03-17T07:16:26.490000",
  "votes": 29,
  "comment_count": 6,
  "views": 0,
  "content": "<p><a href=\"https://arxiv.org/pdf/2103.07579.pdf\" target=\"_blank\">https://arxiv.org/pdf/2103.07579.pdf</a><br>\nRevisiting ResNets: Improved Training and Scaling Strategies</p>\n<p>it turns out that efficientnet is not necessarily better than resnet. it really depends on how you train them.</p>\n<p>automatic architecture search vs human design … back to square one????</p>\n<p><img src=\"https://i.ibb.co/tJHGFHf/Selection-035.png\" alt=\"\"> <br>\n<img src=\"https://i.ibb.co/mBGz2g8/Selection-034.png\" alt=\"\"><br>\n<img src=\"https://i.ibb.co/Gcv20ZJ/Selection-036.png\" alt=\"\"></p>\n<p>interesting \"Representations from supervised learning with improved training strategies rival or outperform representations from state-of-the-art self-supervised learning algorithms\"</p>",
  "messages": [
    {
      "id": 1241726,
      "postDate": "2021-03-17T07:16:26.490Z",
      "content": "<p><a href=\"https://arxiv.org/pdf/2103.07579.pdf\" target=\"_blank\">https://arxiv.org/pdf/2103.07579.pdf</a><br>\nRevisiting ResNets: Improved Training and Scaling Strategies</p>\n<p>it turns out that efficientnet is not necessarily better than resnet. it really depends on how you train them.</p>\n<p>automatic architecture search vs human design … back to square one????</p>\n<p><img src=\"https://i.ibb.co/tJHGFHf/Selection-035.png\" alt=\"\"> <br>\n<img src=\"https://i.ibb.co/mBGz2g8/Selection-034.png\" alt=\"\"><br>\n<img src=\"https://i.ibb.co/Gcv20ZJ/Selection-036.png\" alt=\"\"></p>\n<p>interesting \"Representations from supervised learning with improved training strategies rival or outperform representations from state-of-the-art self-supervised learning algorithms\"</p>",
      "rawMarkdown": "https://arxiv.org/pdf/2103.07579.pdf\nRevisiting ResNets: Improved Training and Scaling Strategies\n\nit turns out that efficientnet is not necessarily better than resnet. it really depends on how you train them.\n\nautomatic architecture search vs human design ... back to square one????\n\n![](https://i.ibb.co/tJHGFHf/Selection-035.png) \n![](https://i.ibb.co/mBGz2g8/Selection-034.png)\n![](https://i.ibb.co/Gcv20ZJ/Selection-036.png)\n\ninteresting \"Representations from supervised learning with improved training strategies rival or outperform representations from state-of-the-art self-supervised learning algorithms\"",
      "votes": 29
    },
    {
      "id": 1241934,
      "postDate": "2021-03-17T09:35:49.503Z",
      "content": "<p>To the Eff debate/defens: I trained a singel fold EffB3adv with valid score ~0.96 and public LB 0.963, B4ns valid score ~0.963, and LB public 0.965/Private 0.969 with B3+B1+B4 ensemble. That with seg prediction + img as input. <br>\nThe B1 reached val0.957 with only 5 epochs/~60min training.<br>\nWith some more time would ha tried more models/larger like ResNet-200D, this time I put my effort in small EfficientNet.<br>\nI also think with some more tuning I think I could have trained the Eff. higher. <br>\nUnfortunately, I didn’t ensemble with public ResNet models, maybe had given a better LB boost, finished the Eff training the last day.</p>",
      "rawMarkdown": "To the Eff debate/defens: I trained a singel fold EffB3adv with valid score ~0.96 and public LB 0.963, B4ns valid score ~0.963, and LB public 0.965/Private 0.969 with B3+B1+B4 ensemble. That with seg prediction + img as input. \nThe B1 reached val0.957 with only 5 epochs/~60min training.\nWith some more time would ha tried more models/larger like ResNet-200D, this time I put my effort in small EfficientNet.\nI also think with some more tuning I think I could have trained the Eff. higher. \nUnfortunately, I didn’t ensemble with public ResNet models, maybe had given a better LB boost, finished the Eff training the last day.",
      "votes": 1
    },
    {
      "id": 1241922,
      "postDate": "2021-03-17T09:25:33.980Z",
      "content": "<p>The models in this paper are resnet d. The tricks used also improve effnets, and I doubt they are new to any Kaggler with good experience in image classification competitions.  Anyway, it is good to see a paper confirming what we all experience in these competitions.</p>",
      "rawMarkdown": "The models in this paper are resnet d. The tricks used also improve effnets, and I doubt they are new to any Kaggler with good experience in image classification competitions.  Anyway, it is good to see a paper confirming what we all experience in these competitions.",
      "votes": 2
    },
    {
      "id": 1242060,
      "postDate": "2021-03-17T11:24:21.430Z",
      "content": "<p>I am no longer surprised that these models performed so well in this competition.</p>\n<p>The pre-trained weights on ImageNet used in this competition for ResNet-D comes from PyTorch Image Models (TIMM), according to the README of that repo, that weights were just added in Dec 18, 2020 and says:</p>\n<ul>\n<li>256x256 val, 0.94 crop (top-1) - 101D (82.33), 152D (83.08), 200D (83.25)</li>\n<li>288x288 val, 1.0 crop - 101D (82.64), 152D (83.48), 200D (83.76)</li>\n<li>320x320 val, 1.0 crop - 101D (83.00), 152D (83.66), 200D (84.01)</li>\n</ul>\n<p>There is no direct link to the paper, but whoever trained the models on ImageNet also used similar techniques for training, otherwise the performance would not be so high. Good to see, that the performance of ResNet-D is comparable to EfficientNet when trained using current state-of-the-art techniques. It is also interesting to me, that it is possible to achieve such high performance on ImageNet with a relatively small image size of 256x256, while EfficientNet requires almost twice the resolution for.</p>",
      "rawMarkdown": "I am no longer surprised that these models performed so well in this competition.\n\nThe pre-trained weights on ImageNet used in this competition for ResNet-D comes from PyTorch Image Models (TIMM), according to the README of that repo, that weights were just added in Dec 18, 2020 and says:\n- 256x256 val, 0.94 crop (top-1) - 101D (82.33), 152D (83.08), 200D (83.25)\n- 288x288 val, 1.0 crop - 101D (82.64), 152D (83.48), 200D (83.76)\n- 320x320 val, 1.0 crop - 101D (83.00), 152D (83.66), 200D (84.01)\n\nThere is no direct link to the paper, but whoever trained the models on ImageNet also used similar techniques for training, otherwise the performance would not be so high. Good to see, that the performance of ResNet-D is comparable to EfficientNet when trained using current state-of-the-art techniques. It is also interesting to me, that it is possible to achieve such high performance on ImageNet with a relatively small image size of 256x256, while EfficientNet requires almost twice the resolution for."
    },
    {
      "id": 1241839,
      "postDate": "2021-03-17T08:26:56.737Z",
      "content": "<p>Can you spell out what the link between ResNet-200D and the paper is? Is your point that ResNets do better than you'd think (incl. in this competition), given all the neural architecture search stuff? Or is ResNet-200D somehow even more closely linked to what is discussed in the paper?</p>",
      "rawMarkdown": "Can you spell out what the link between ResNet-200D and the paper is? Is your point that ResNets do better than you'd think (incl. in this competition), given all the neural architecture search stuff? Or is ResNet-200D somehow even more closely linked to what is discussed in the paper?"
    },
    {
      "id": 1241803,
      "postDate": "2021-03-17T07:56:33.567Z",
      "content": "<p>on top of that, I cannot seem to train EfficientNet well in this competition, maybe LB can hit 0.960 but the training process is extremely turbulent, with validation loss increasing for many epochs. I have tried various methods like freezing bn layers, using heavy dropout, but to no avail~~~ </p>",
      "rawMarkdown": "on top of that, I cannot seem to train EfficientNet well in this competition, maybe LB can hit 0.960 but the training process is extremely turbulent, with validation loss increasing for many epochs. I have tried various methods like freezing bn layers, using heavy dropout, but to no avail~~~ ",
      "replies": [
        {
          "id": 1241916,
          "postDate": "2021-03-17T09:20:37.593Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1241934,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2021-03-17T09:35:49.503000",
      "content": "<p>To the Eff debate/defens: I trained a singel fold EffB3adv with valid score ~0.96 and public LB 0.963, B4ns valid score ~0.963, and LB public 0.965/Private 0.969 with B3+B1+B4 ensemble. That with seg prediction + img as input. <br>\nThe B1 reached val0.957 with only 5 epochs/~60min training.<br>\nWith some more time would ha tried more models/larger like ResNet-200D, this time I put my effort in small EfficientNet.<br>\nI also think with some more tuning I think I could have trained the Eff. higher. <br>\nUnfortunately, I didn’t ensemble with public ResNet models, maybe had given a better LB boost, finished the Eff training the last day.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1241922,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2021-03-17T09:25:33.980000",
      "content": "<p>The models in this paper are resnet d. The tricks used also improve effnets, and I doubt they are new to any Kaggler with good experience in image classification competitions.  Anyway, it is good to see a paper confirming what we all experience in these competitions.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1242060,
      "author_name": "Konni",
      "author_url": "",
      "post_date": "2021-03-17T11:24:21.430000",
      "content": "<p>I am no longer surprised that these models performed so well in this competition.</p>\n<p>The pre-trained weights on ImageNet used in this competition for ResNet-D comes from PyTorch Image Models (TIMM), according to the README of that repo, that weights were just added in Dec 18, 2020 and says:</p>\n<ul>\n<li>256x256 val, 0.94 crop (top-1) - 101D (82.33), 152D (83.08), 200D (83.25)</li>\n<li>288x288 val, 1.0 crop - 101D (82.64), 152D (83.48), 200D (83.76)</li>\n<li>320x320 val, 1.0 crop - 101D (83.00), 152D (83.66), 200D (84.01)</li>\n</ul>\n<p>There is no direct link to the paper, but whoever trained the models on ImageNet also used similar techniques for training, otherwise the performance would not be so high. Good to see, that the performance of ResNet-D is comparable to EfficientNet when trained using current state-of-the-art techniques. It is also interesting to me, that it is possible to achieve such high performance on ImageNet with a relatively small image size of 256x256, while EfficientNet requires almost twice the resolution for.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1241839,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2021-03-17T08:26:56.737000",
      "content": "<p>Can you spell out what the link between ResNet-200D and the paper is? Is your point that ResNets do better than you'd think (incl. in this competition), given all the neural architecture search stuff? Or is ResNet-200D somehow even more closely linked to what is discussed in the paper?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1241803,
      "author_name": "gao-hongnan",
      "author_url": "",
      "post_date": "2021-03-17T07:56:33.567000",
      "content": "<p>on top of that, I cannot seem to train EfficientNet well in this competition, maybe LB can hit 0.960 but the training process is extremely turbulent, with validation loss increasing for many epochs. I have tried various methods like freezing bn layers, using heavy dropout, but to no avail~~~ </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1241916,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-17T09:20:37.593000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1241726": "https://arxiv.org/pdf/2103.07579.pdf\nRevisiting ResNets: Improved Training and Scaling Strategies\n\nit turns out that efficientnet is not necessarily better than resnet. it really depends on how you train them.\n\nautomatic architecture search vs human design ... back to square one????\n\n![](https://i.ibb.co/tJHGFHf/Selection-035.png) \n![](https://i.ibb.co/mBGz2g8/Selection-034.png)\n![](https://i.ibb.co/Gcv20ZJ/Selection-036.png)\n\ninteresting \"Representations from supervised learning with improved training strategies rival or outperform representations from state-of-the-art self-supervised learning algorithms\"",
    "1241934": "To the Eff debate/defens: I trained a singel fold EffB3adv with valid score ~0.96 and public LB 0.963, B4ns valid score ~0.963, and LB public 0.965/Private 0.969 with B3+B1+B4 ensemble. That with seg prediction + img as input. \nThe B1 reached val0.957 with only 5 epochs/~60min training.\nWith some more time would ha tried more models/larger like ResNet-200D, this time I put my effort in small EfficientNet.\nI also think with some more tuning I think I could have trained the Eff. higher. \nUnfortunately, I didn’t ensemble with public ResNet models, maybe had given a better LB boost, finished the Eff training the last day.",
    "1241922": "The models in this paper are resnet d. The tricks used also improve effnets, and I doubt they are new to any Kaggler with good experience in image classification competitions.  Anyway, it is good to see a paper confirming what we all experience in these competitions.",
    "1242060": "I am no longer surprised that these models performed so well in this competition.\n\nThe pre-trained weights on ImageNet used in this competition for ResNet-D comes from PyTorch Image Models (TIMM), according to the README of that repo, that weights were just added in Dec 18, 2020 and says:\n- 256x256 val, 0.94 crop (top-1) - 101D (82.33), 152D (83.08), 200D (83.25)\n- 288x288 val, 1.0 crop - 101D (82.64), 152D (83.48), 200D (83.76)\n- 320x320 val, 1.0 crop - 101D (83.00), 152D (83.66), 200D (84.01)\n\nThere is no direct link to the paper, but whoever trained the models on ImageNet also used similar techniques for training, otherwise the performance would not be so high. Good to see, that the performance of ResNet-D is comparable to EfficientNet when trained using current state-of-the-art techniques. It is also interesting to me, that it is possible to achieve such high performance on ImageNet with a relatively small image size of 256x256, while EfficientNet requires almost twice the resolution for.",
    "1241839": "Can you spell out what the link between ResNet-200D and the paper is? Is your point that ResNets do better than you'd think (incl. in this competition), given all the neural architecture search stuff? Or is ResNet-200D somehow even more closely linked to what is discussed in the paper?",
    "1241803": "on top of that, I cannot seem to train EfficientNet well in this competition, maybe LB can hit 0.960 but the training process is extremely turbulent, with validation loss increasing for many epochs. I have tried various methods like freezing bn layers, using heavy dropout, but to no avail~~~ "
  }
}