{
  "id": 212856,
  "title": "Do better ImageNet models really not do better on chest X-rays?",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/212856",
  "author_name": "Björn",
  "post_date": "2021-01-20T13:27:20.971000",
  "votes": 42,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Has anyone else looked into or thought about <a href=\"https://arxiv.org/abs/2101.06871\" target=\"_blank\">this recent claim</a> (or has experiments showing this is wrong besides e.g. <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/204950\" target=\"_blank\">this discussion post</a>) that higher performance on ImageNet does not translate to higher performance on medical chest X-ray imaging tasks (i.e. choosing your model based on ImageNet might not be best)? The authors claim that better performing architectures on ImageNet do not do better chest x-ray interpretation (both for pretrained and not) and that ImageNet pretraining helps more (!) for smaller models. That did seem potentially directly relevant to this competition.</p>\n<p>I've asked the authors <a href=\"https://twitter.com/BjornHolzhauer/status/1351792295042027520?s=19\" target=\"_blank\">on Twitter</a> whether that's just due to not training each architecture well. Maybe I misread, but it seems like they used 3 epochs at lr=1e-4 &amp; bs=16 for all architectures. I'd have thought that would hurt architectures like EfficientNet and larger networks within architectures (i.e. I'd assume this would hurt larger models more - especially without using a pre-trained model).</p>\n<p>Ross Wightman did point out <a href=\"https://t.co/Pf6Hm48FPH\" target=\"_blank\">another paper</a> that does look into whether architectures that perform well on ImageNet perform well for transfer learning (but not specifically for medical imaging, I believe). I had separately see <a href=\"https://arxiv.org/abs/2101.05913\" target=\"_blank\">this recent paper</a> that suggests that pre-trained models transfer better (i.e. being pre-trained is more valuable), if they have trained on really large datasets (i.e. even bigger than ImageNet) and are really large capacity models. I guess, there's also been a bunch of papers suggesting the value of pre-training for large architectures is more important, the smaller the new training data is, so results for a really huge chest x-ray dataset may not apply here (but again, bringing external data is of course an option).</p>",
  "messages": [
    {
      "id": 1161296,
      "postDate": "2021-01-20T13:27:20.973Z",
      "content": "<p>Has anyone else looked into or thought about <a href=\"https://arxiv.org/abs/2101.06871\" target=\"_blank\">this recent claim</a> (or has experiments showing this is wrong besides e.g. <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/204950\" target=\"_blank\">this discussion post</a>) that higher performance on ImageNet does not translate to higher performance on medical chest X-ray imaging tasks (i.e. choosing your model based on ImageNet might not be best)? The authors claim that better performing architectures on ImageNet do not do better chest x-ray interpretation (both for pretrained and not) and that ImageNet pretraining helps more (!) for smaller models. That did seem potentially directly relevant to this competition.</p>\n<p>I've asked the authors <a href=\"https://twitter.com/BjornHolzhauer/status/1351792295042027520?s=19\" target=\"_blank\">on Twitter</a> whether that's just due to not training each architecture well. Maybe I misread, but it seems like they used 3 epochs at lr=1e-4 &amp; bs=16 for all architectures. I'd have thought that would hurt architectures like EfficientNet and larger networks within architectures (i.e. I'd assume this would hurt larger models more - especially without using a pre-trained model).</p>\n<p>Ross Wightman did point out <a href=\"https://t.co/Pf6Hm48FPH\" target=\"_blank\">another paper</a> that does look into whether architectures that perform well on ImageNet perform well for transfer learning (but not specifically for medical imaging, I believe). I had separately see <a href=\"https://arxiv.org/abs/2101.05913\" target=\"_blank\">this recent paper</a> that suggests that pre-trained models transfer better (i.e. being pre-trained is more valuable), if they have trained on really large datasets (i.e. even bigger than ImageNet) and are really large capacity models. I guess, there's also been a bunch of papers suggesting the value of pre-training for large architectures is more important, the smaller the new training data is, so results for a really huge chest x-ray dataset may not apply here (but again, bringing external data is of course an option).</p>",
      "rawMarkdown": "Has anyone else looked into or thought about [this recent claim](https://arxiv.org/abs/2101.06871) (or has experiments showing this is wrong besides e.g. [this discussion post](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/204950)) that higher performance on ImageNet does not translate to higher performance on medical chest X-ray imaging tasks (i.e. choosing your model based on ImageNet might not be best)? The authors claim that better performing architectures on ImageNet do not do better chest x-ray interpretation (both for pretrained and not) and that ImageNet pretraining helps more (!) for smaller models. That did seem potentially directly relevant to this competition.\n\nI've asked the authors [on Twitter](https://twitter.com/BjornHolzhauer/status/1351792295042027520?s=19) whether that's just due to not training each architecture well. Maybe I misread, but it seems like they used 3 epochs at lr=1e-4 & bs=16 for all architectures. I'd have thought that would hurt architectures like EfficientNet and larger networks within architectures (i.e. I'd assume this would hurt larger models more - especially without using a pre-trained model).\n\nRoss Wightman did point out [another paper](https://t.co/Pf6Hm48FPH) that does look into whether architectures that perform well on ImageNet perform well for transfer learning (but not specifically for medical imaging, I believe). I had separately see [this recent paper](https://arxiv.org/abs/2101.05913) that suggests that pre-trained models transfer better (i.e. being pre-trained is more valuable), if they have trained on really large datasets (i.e. even bigger than ImageNet) and are really large capacity models. I guess, there's also been a bunch of papers suggesting the value of pre-training for large architectures is more important, the smaller the new training data is, so results for a really huge chest x-ray dataset may not apply here (but again, bringing external data is of course an option).",
      "votes": 42
    },
    {
      "id": 1161890,
      "postDate": "2021-01-20T19:55:54.907Z",
      "content": "<p>Ross has mentioned some criticisms of this study mostly pertaining to how poorly they evaluate hyperparameters. static learning rate of 1e-4 and batch size of 16 and no other published hyperparams makes it a significantly less robust indicator. </p>\n<p>He also pointed to a few other papers that have been done previously that cover similar topics. </p>\n<ul>\n<li>On Robustness and Transferability - <a href=\"https://arxiv.org/abs/2007.08558\" target=\"_blank\">https://arxiv.org/abs/2007.08558</a></li>\n<li>Done with ImageNet? - <a href=\"https://arxiv.org/abs/2006.07159\" target=\"_blank\">https://arxiv.org/abs/2006.07159</a></li>\n<li>Big Transfer - <a href=\"https://arxiv.org/abs/1912.11370\" target=\"_blank\">https://arxiv.org/abs/1912.11370</a></li>\n<li>Large scale study of representation learning - <a href=\"https://arxiv.org/abs/1910.04867\" target=\"_blank\">https://arxiv.org/abs/1910.04867</a></li>\n<li>Some of the work in Quoc Le group w/ Simon Kornblith</li>\n<li>Do better ImageNet models transfer better? - <a href=\"https://arxiv.org/abs/1805.08974\" target=\"_blank\">https://arxiv.org/abs/1805.08974</a></li>\n<li>Domain adaptive transfer - <a href=\"https://arxiv.org/abs/1811.07056\" target=\"_blank\">https://arxiv.org/abs/1811.07056</a></li>\n<li>Other papers w/ Simon and Hinton on representation learning, incl simclr constrastive learning have interseting insight wrt transfer as well</li>\n</ul>",
      "rawMarkdown": "Ross has mentioned some criticisms of this study mostly pertaining to how poorly they evaluate hyperparameters. static learning rate of 1e-4 and batch size of 16 and no other published hyperparams makes it a significantly less robust indicator. \n\nHe also pointed to a few other papers that have been done previously that cover similar topics. \n\n- On Robustness and Transferability - https://arxiv.org/abs/2007.08558\n- Done with ImageNet? - https://arxiv.org/abs/2006.07159\n- Big Transfer - https://arxiv.org/abs/1912.11370\n- Large scale study of representation learning - https://arxiv.org/abs/1910.04867\n- Some of the work in Quoc Le group w/ Simon Kornblith\n- Do better ImageNet models transfer better? - https://arxiv.org/abs/1805.08974\n- Domain adaptive transfer - https://arxiv.org/abs/1811.07056\n- Other papers w/ Simon and Hinton on representation learning, incl simclr constrastive learning have interseting insight wrt transfer as well",
      "votes": 11,
      "replies": [
        {
          "id": 1162369,
          "postDate": "2021-01-21T05:43:13.547Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1162368,
          "postDate": "2021-01-21T05:43:13.547Z",
          "content": "<blockquote>\n  <p>static learning rate of 1e-4 and batch size of 16 and no other published hyperparams makes it a significantly less robust indicator.</p>\n</blockquote>\n<p>I feel the same.</p>",
          "rawMarkdown": "> static learning rate of 1e-4 and batch size of 16 and no other published hyperparams makes it a significantly less robust indicator.\n\nI feel the same.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1161582,
      "postDate": "2021-01-20T16:20:30.710Z",
      "content": "<p>Wow. I feel very surprised by this paper. I have been working with X-rays for a few years now, and I was always confused why my InceptionV3 model beats all other ImageNet state-of-the art architectures when no pre-training is used. It is nice to see someone confirm that on other CXR dataset :)</p>",
      "rawMarkdown": "Wow. I feel very surprised by this paper. I have been working with X-rays for a few years now, and I was always confused why my InceptionV3 model beats all other ImageNet state-of-the art architectures when no pre-training is used. It is nice to see someone confirm that on other CXR dataset :)\n\n",
      "votes": 9,
      "replies": [
        {
          "id": 1161626,
          "postDate": "2021-01-20T16:54:31.693Z",
          "content": "<p>Just to confirm, there's invisible sarcasm-html tags here, right?</p>",
          "rawMarkdown": "Just to confirm, there's invisible sarcasm-html tags here, right?",
          "votes": 1
        },
        {
          "id": 1161787,
          "postDate": "2021-01-20T18:44:39.840Z",
          "content": "<p>No sarcasm intended. Most of the findings in the paper can be reproduced on other equivalent to CheXpert datasets.</p>",
          "rawMarkdown": "No sarcasm intended. Most of the findings in the paper can be reproduced on other equivalent to CheXpert datasets.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1162676,
      "postDate": "2021-01-21T09:09:16.413Z",
      "content": "<p>lesson learned: add inceptionV3 and densenet121 to your ensembles</p>",
      "rawMarkdown": "lesson learned: add inceptionV3 and densenet121 to your ensembles",
      "votes": 8,
      "replies": [
        {
          "id": 1215492,
          "postDate": "2021-02-23T16:54:54.250Z",
          "content": "<p>Most large models perform better, but why is densenet121 so much better than densenet201?</p>",
          "rawMarkdown": "Most large models perform better, but why is densenet121 so much better than densenet201?"
        }
      ]
    },
    {
      "id": 1205698,
      "postDate": "2021-02-16T22:18:42.560Z",
      "content": "<p>This was an amazing paper. I created a summary for it. Hope it helps.</p>\n<p>![<a href=\"https://drive.google.com/file/d/1RBUH30DJQtaNkNEuYz_3gYj-ZtlF0mLv/view?usp=sharing](paper\" target=\"_blank\">https://drive.google.com/file/d/1RBUH30DJQtaNkNEuYz_3gYj-ZtlF0mLv/view?usp=sharing](paper</a> summary)</p>",
      "rawMarkdown": "This was an amazing paper. I created a summary for it. Hope it helps.\n\n![https://drive.google.com/file/d/1RBUH30DJQtaNkNEuYz_3gYj-ZtlF0mLv/view?usp=sharing](paper summary)",
      "votes": 3
    },
    {
      "id": 1204254,
      "postDate": "2021-02-16T01:12:33.990Z",
      "content": "<p>Great paper indeed! <br>\nAlthough, I would have liked to see the results 'without' fine-tuning all the CNN parameters and just logistic regression setting results; to see the effectiveness of architectures.  </p>",
      "rawMarkdown": "Great paper indeed! \nAlthough, I would have liked to see the results 'without' fine-tuning all the CNN parameters and just logistic regression setting results; to see the effectiveness of architectures.  ",
      "votes": 1
    },
    {
      "id": 1189987,
      "postDate": "2021-02-07T11:46:12.263Z",
      "content": "<p>this is a fundamental and basic research question , similar example from my experience is efficientNetB0 and effcientNetB7 could not perform well on face recognition i.e getting embedding from them and comparing whether they are similar or not based on threshold. In Paper the authors claimed their models performed well on 5 out of 8 datasets including image Net , that is quite reasonable , but it did not worked out on my face recognition problem on which mobileNet performs greatly than efficienNet models . </p>",
      "rawMarkdown": "this is a fundamental and basic research question , similar example from my experience is efficientNetB0 and effcientNetB7 could not perform well on face recognition i.e getting embedding from them and comparing whether they are similar or not based on threshold. In Paper the authors claimed their models performed well on 5 out of 8 datasets including image Net , that is quite reasonable , but it did not worked out on my face recognition problem on which mobileNet performs greatly than efficienNet models . \n",
      "votes": 1
    },
    {
      "id": 1188416,
      "postDate": "2021-02-06T08:10:15.173Z",
      "content": "<p>One of the things I am reminded of with this paper is the general principle I learned early on that theoretically we can solve many of these problems with much smaller networks, less layers, less neurons, etc. but the big difficulty is finding the correct weights that generalize. In general it is easier to find good weights in larger networks, but it is on paper possible to get equivalent performance out of smaller networks. Initialization plays a very important role in this process, pretraining on imagenet is in some sense just a more well-informed initialization. </p>",
      "rawMarkdown": "One of the things I am reminded of with this paper is the general principle I learned early on that theoretically we can solve many of these problems with much smaller networks, less layers, less neurons, etc. but the big difficulty is finding the correct weights that generalize. In general it is easier to find good weights in larger networks, but it is on paper possible to get equivalent performance out of smaller networks. Initialization plays a very important role in this process, pretraining on imagenet is in some sense just a more well-informed initialization. ",
      "votes": 1
    },
    {
      "id": 1183289,
      "postDate": "2021-02-02T21:49:34.223Z",
      "content": "<p>Speaking as someone doing novel radiography + DL models with an insanely small dataset at work, I can confirm that for me, Chexnet worked a LOT better than imagenet (even with better normalization than what's done in most of the notebooks here). <br>\nYMMV ofc.</p>",
      "rawMarkdown": "Speaking as someone doing novel radiography + DL models with an insanely small dataset at work, I can confirm that for me, Chexnet worked a LOT better than imagenet (even with better normalization than what's done in most of the notebooks here). \nYMMV ofc.",
      "votes": 1
    },
    {
      "id": 1182689,
      "postDate": "2021-02-02T14:51:26.200Z",
      "content": "<p>Great findings,I will read these papers and think why these models which have high accuracy in Imagenet can't get high score in kaggle competitions.</p>",
      "rawMarkdown": "Great findings,I will read these papers and think why these models which have high accuracy in Imagenet can't get high score in kaggle competitions.",
      "votes": 1
    },
    {
      "id": 1161568,
      "postDate": "2021-01-20T16:11:45.330Z",
      "content": "<p>Nice work,so ,how do we choose the model?Using the small model?😂</p>",
      "rawMarkdown": "Nice work,so ,how do we choose the model?Using the small model?😂",
      "votes": 1,
      "replies": [
        {
          "id": 1162567,
          "postDate": "2021-01-21T08:19:56.427Z",
          "content": "<p>I guess repeating something like what this paper did, but better (i.e. getting rid of the dubious hyperparameters esp. learning rate schedule &amp; epoch number and setting those well for each architecture). I guess the whole crowd of participants trying things out is sort of a chaotic version of that. 🙄</p>",
          "rawMarkdown": "I guess repeating something like what this paper did, but better (i.e. getting rid of the dubious hyperparameters esp. learning rate schedule & epoch number and setting those well for each architecture). I guess the whole crowd of participants trying things out is sort of a chaotic version of that. 🙄"
        }
      ]
    },
    {
      "id": 1187998,
      "postDate": "2021-02-05T21:29:55.657Z",
      "content": "<p>I'd like to throw this in - <a href=\"https://openaccess.thecvf.com/content_CVPR_2020/html/Newell_How_Useful_Is_Self-Supervised_Pretraining_for_Visual_Tasks_CVPR_2020_paper.html\" target=\"_blank\">How useful is self-supervised pre-training</a></p>\n<p>Just something to be aware of before wasting previous GPU hours pretraining model.</p>",
      "rawMarkdown": "I'd like to throw this in - [How useful is self-supervised pre-training](https://openaccess.thecvf.com/content_CVPR_2020/html/Newell_How_Useful_Is_Self-Supervised_Pretraining_for_Visual_Tasks_CVPR_2020_paper.html)\n\nJust something to be aware of before wasting previous GPU hours pretraining model.",
      "votes": 2
    },
    {
      "id": 1161802,
      "postDate": "2021-01-20T18:54:39.477Z",
      "content": "<p>Interesting findings for sure. Surprised how underresearched this area is given that so few are actually doing just imagenet. Almost everyone is transferring away to different tasks.</p>\n<p>Even in this competition, people have claimed gains from one architecture vs the other but it seems fairly inconsistent based on size of model but possibly significant when going to different kinds of models, dense net, effnet, resnet, etc. </p>",
      "rawMarkdown": "Interesting findings for sure. Surprised how underresearched this area is given that so few are actually doing just imagenet. Almost everyone is transferring away to different tasks.\n\nEven in this competition, people have claimed gains from one architecture vs the other but it seems fairly inconsistent based on size of model but possibly significant when going to different kinds of models, dense net, effnet, resnet, etc. ",
      "votes": 2,
      "replies": [
        {
          "id": 1161834,
          "postDate": "2021-01-20T19:16:23.137Z",
          "content": "<p>I have been questioning this quite a lot lately. And I have a feeling that most of the improvement on ImageNet has come from better feature representations of multi-scale objects. Therefore, any other task which has multi-scale problem benefit from ImageNet advancements.</p>\n<p>On the other hand, most of the objects in chest X-rays are scale invariant and they do not benefit from that at all. This can also be seen in this competition as our best single models so far have been resnet-type architectures. </p>",
          "rawMarkdown": "I have been questioning this quite a lot lately. And I have a feeling that most of the improvement on ImageNet has come from better feature representations of multi-scale objects. Therefore, any other task which has multi-scale problem benefit from ImageNet advancements.\n\nOn the other hand, most of the objects in chest X-rays are scale invariant and they do not benefit from that at all. This can also be seen in this competition as our best single models so far have been resnet-type architectures. ",
          "votes": 7
        },
        {
          "id": 1161887,
          "postDate": "2021-01-20T19:50:18.610Z",
          "content": "<p>Yeah, plenty of people are trying to apply new methods to medical imaging, but it isnt operating in the same colorspace or constraints as normal photographs. it would be interesting if these results hold for other datasets, maybe it is that the benefits we see from new architectures just aren't pertinent to x-rays because of the characteristics you mentioned or other ones. </p>",
          "rawMarkdown": "Yeah, plenty of people are trying to apply new methods to medical imaging, but it isnt operating in the same colorspace or constraints as normal photographs. it would be interesting if these results hold for other datasets, maybe it is that the benefits we see from new architectures just aren't pertinent to x-rays because of the characteristics you mentioned or other ones. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1173438,
      "postDate": "2021-01-27T22:00:35.163Z",
      "content": "<p>Truly fascinating!</p>",
      "rawMarkdown": "Truly fascinating!"
    }
  ],
  "comments": [
    {
      "id": 1161890,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2021-01-20T19:55:54.907000",
      "content": "<p>Ross has mentioned some criticisms of this study mostly pertaining to how poorly they evaluate hyperparameters. static learning rate of 1e-4 and batch size of 16 and no other published hyperparams makes it a significantly less robust indicator. </p>\n<p>He also pointed to a few other papers that have been done previously that cover similar topics. </p>\n<ul>\n<li>On Robustness and Transferability - <a href=\"https://arxiv.org/abs/2007.08558\" target=\"_blank\">https://arxiv.org/abs/2007.08558</a></li>\n<li>Done with ImageNet? - <a href=\"https://arxiv.org/abs/2006.07159\" target=\"_blank\">https://arxiv.org/abs/2006.07159</a></li>\n<li>Big Transfer - <a href=\"https://arxiv.org/abs/1912.11370\" target=\"_blank\">https://arxiv.org/abs/1912.11370</a></li>\n<li>Large scale study of representation learning - <a href=\"https://arxiv.org/abs/1910.04867\" target=\"_blank\">https://arxiv.org/abs/1910.04867</a></li>\n<li>Some of the work in Quoc Le group w/ Simon Kornblith</li>\n<li>Do better ImageNet models transfer better? - <a href=\"https://arxiv.org/abs/1805.08974\" target=\"_blank\">https://arxiv.org/abs/1805.08974</a></li>\n<li>Domain adaptive transfer - <a href=\"https://arxiv.org/abs/1811.07056\" target=\"_blank\">https://arxiv.org/abs/1811.07056</a></li>\n<li>Other papers w/ Simon and Hinton on representation learning, incl simclr constrastive learning have interseting insight wrt transfer as well</li>\n</ul>",
      "votes": 11,
      "replies": [
        {
          "id": 1162369,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-01-21T05:43:13.547000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1162368,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2021-01-21T05:43:13.547000",
          "content": "<blockquote>\n  <p>static learning rate of 1e-4 and batch size of 16 and no other published hyperparams makes it a significantly less robust indicator.</p>\n</blockquote>\n<p>I feel the same.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1161582,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2021-01-20T16:20:30.710000",
      "content": "<p>Wow. I feel very surprised by this paper. I have been working with X-rays for a few years now, and I was always confused why my InceptionV3 model beats all other ImageNet state-of-the art architectures when no pre-training is used. It is nice to see someone confirm that on other CXR dataset :)</p>",
      "votes": 9,
      "replies": [
        {
          "id": 1161626,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2021-01-20T16:54:31.693000",
          "content": "<p>Just to confirm, there's invisible sarcasm-html tags here, right?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1161787,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2021-01-20T18:44:39.840000",
          "content": "<p>No sarcasm intended. Most of the findings in the paper can be reproduced on other equivalent to CheXpert datasets.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1162676,
      "author_name": "DatNT",
      "author_url": "",
      "post_date": "2021-01-21T09:09:16.413000",
      "content": "<p>lesson learned: add inceptionV3 and densenet121 to your ensembles</p>",
      "votes": 8,
      "replies": [
        {
          "id": 1215492,
          "author_name": "Alien",
          "author_url": "",
          "post_date": "2021-02-23T16:54:54.250000",
          "content": "<p>Most large models perform better, but why is densenet121 so much better than densenet201?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1205698,
      "author_name": "Dr. Amritpal Singh",
      "author_url": "",
      "post_date": "2021-02-16T22:18:42.560000",
      "content": "<p>This was an amazing paper. I created a summary for it. Hope it helps.</p>\n<p>![<a href=\"https://drive.google.com/file/d/1RBUH30DJQtaNkNEuYz_3gYj-ZtlF0mLv/view?usp=sharing](paper\" target=\"_blank\">https://drive.google.com/file/d/1RBUH30DJQtaNkNEuYz_3gYj-ZtlF0mLv/view?usp=sharing](paper</a> summary)</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1204254,
      "author_name": "Jaideep Murkute",
      "author_url": "",
      "post_date": "2021-02-16T01:12:33.990000",
      "content": "<p>Great paper indeed! <br>\nAlthough, I would have liked to see the results 'without' fine-tuning all the CNN parameters and just logistic regression setting results; to see the effectiveness of architectures.  </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1189987,
      "author_name": "u_ahmed khan",
      "author_url": "",
      "post_date": "2021-02-07T11:46:12.263000",
      "content": "<p>this is a fundamental and basic research question , similar example from my experience is efficientNetB0 and effcientNetB7 could not perform well on face recognition i.e getting embedding from them and comparing whether they are similar or not based on threshold. In Paper the authors claimed their models performed well on 5 out of 8 datasets including image Net , that is quite reasonable , but it did not worked out on my face recognition problem on which mobileNet performs greatly than efficienNet models . </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1188416,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2021-02-06T08:10:15.173000",
      "content": "<p>One of the things I am reminded of with this paper is the general principle I learned early on that theoretically we can solve many of these problems with much smaller networks, less layers, less neurons, etc. but the big difficulty is finding the correct weights that generalize. In general it is easier to find good weights in larger networks, but it is on paper possible to get equivalent performance out of smaller networks. Initialization plays a very important role in this process, pretraining on imagenet is in some sense just a more well-informed initialization. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1183289,
      "author_name": "Dan Ofer",
      "author_url": "",
      "post_date": "2021-02-02T21:49:34.223000",
      "content": "<p>Speaking as someone doing novel radiography + DL models with an insanely small dataset at work, I can confirm that for me, Chexnet worked a LOT better than imagenet (even with better normalization than what's done in most of the notebooks here). <br>\nYMMV ofc.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1182689,
      "author_name": "hujianxin",
      "author_url": "",
      "post_date": "2021-02-02T14:51:26.200000",
      "content": "<p>Great findings,I will read these papers and think why these models which have high accuracy in Imagenet can't get high score in kaggle competitions.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1161568,
      "author_name": "Bcw93",
      "author_url": "",
      "post_date": "2021-01-20T16:11:45.330000",
      "content": "<p>Nice work,so ,how do we choose the model?Using the small model?😂</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1162567,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2021-01-21T08:19:56.427000",
          "content": "<p>I guess repeating something like what this paper did, but better (i.e. getting rid of the dubious hyperparameters esp. learning rate schedule &amp; epoch number and setting those well for each architecture). I guess the whole crowd of participants trying things out is sort of a chaotic version of that. 🙄</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1187998,
      "author_name": "JunYong Tong",
      "author_url": "",
      "post_date": "2021-02-05T21:29:55.657000",
      "content": "<p>I'd like to throw this in - <a href=\"https://openaccess.thecvf.com/content_CVPR_2020/html/Newell_How_Useful_Is_Self-Supervised_Pretraining_for_Visual_Tasks_CVPR_2020_paper.html\" target=\"_blank\">How useful is self-supervised pre-training</a></p>\n<p>Just something to be aware of before wasting previous GPU hours pretraining model.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1161802,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2021-01-20T18:54:39.477000",
      "content": "<p>Interesting findings for sure. Surprised how underresearched this area is given that so few are actually doing just imagenet. Almost everyone is transferring away to different tasks.</p>\n<p>Even in this competition, people have claimed gains from one architecture vs the other but it seems fairly inconsistent based on size of model but possibly significant when going to different kinds of models, dense net, effnet, resnet, etc. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1161834,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "2021-01-20T19:16:23.137000",
          "content": "<p>I have been questioning this quite a lot lately. And I have a feeling that most of the improvement on ImageNet has come from better feature representations of multi-scale objects. Therefore, any other task which has multi-scale problem benefit from ImageNet advancements.</p>\n<p>On the other hand, most of the objects in chest X-rays are scale invariant and they do not benefit from that at all. This can also be seen in this competition as our best single models so far have been resnet-type architectures. </p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1161887,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2021-01-20T19:50:18.610000",
          "content": "<p>Yeah, plenty of people are trying to apply new methods to medical imaging, but it isnt operating in the same colorspace or constraints as normal photographs. it would be interesting if these results hold for other datasets, maybe it is that the benefits we see from new architectures just aren't pertinent to x-rays because of the characteristics you mentioned or other ones. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1173438,
      "author_name": "Shane Inglis",
      "author_url": "",
      "post_date": "2021-01-27T22:00:35.163000",
      "content": "<p>Truly fascinating!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1161296": "Has anyone else looked into or thought about [this recent claim](https://arxiv.org/abs/2101.06871) (or has experiments showing this is wrong besides e.g. [this discussion post](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/204950)) that higher performance on ImageNet does not translate to higher performance on medical chest X-ray imaging tasks (i.e. choosing your model based on ImageNet might not be best)? The authors claim that better performing architectures on ImageNet do not do better chest x-ray interpretation (both for pretrained and not) and that ImageNet pretraining helps more (!) for smaller models. That did seem potentially directly relevant to this competition.\n\nI've asked the authors [on Twitter](https://twitter.com/BjornHolzhauer/status/1351792295042027520?s=19) whether that's just due to not training each architecture well. Maybe I misread, but it seems like they used 3 epochs at lr=1e-4 & bs=16 for all architectures. I'd have thought that would hurt architectures like EfficientNet and larger networks within architectures (i.e. I'd assume this would hurt larger models more - especially without using a pre-trained model).\n\nRoss Wightman did point out [another paper](https://t.co/Pf6Hm48FPH) that does look into whether architectures that perform well on ImageNet perform well for transfer learning (but not specifically for medical imaging, I believe). I had separately see [this recent paper](https://arxiv.org/abs/2101.05913) that suggests that pre-trained models transfer better (i.e. being pre-trained is more valuable), if they have trained on really large datasets (i.e. even bigger than ImageNet) and are really large capacity models. I guess, there's also been a bunch of papers suggesting the value of pre-training for large architectures is more important, the smaller the new training data is, so results for a really huge chest x-ray dataset may not apply here (but again, bringing external data is of course an option).",
    "1161890": "Ross has mentioned some criticisms of this study mostly pertaining to how poorly they evaluate hyperparameters. static learning rate of 1e-4 and batch size of 16 and no other published hyperparams makes it a significantly less robust indicator. \n\nHe also pointed to a few other papers that have been done previously that cover similar topics. \n\n- On Robustness and Transferability - https://arxiv.org/abs/2007.08558\n- Done with ImageNet? - https://arxiv.org/abs/2006.07159\n- Big Transfer - https://arxiv.org/abs/1912.11370\n- Large scale study of representation learning - https://arxiv.org/abs/1910.04867\n- Some of the work in Quoc Le group w/ Simon Kornblith\n- Do better ImageNet models transfer better? - https://arxiv.org/abs/1805.08974\n- Domain adaptive transfer - https://arxiv.org/abs/1811.07056\n- Other papers w/ Simon and Hinton on representation learning, incl simclr constrastive learning have interseting insight wrt transfer as well",
    "1161582": "Wow. I feel very surprised by this paper. I have been working with X-rays for a few years now, and I was always confused why my InceptionV3 model beats all other ImageNet state-of-the art architectures when no pre-training is used. It is nice to see someone confirm that on other CXR dataset :)\n\n",
    "1162676": "lesson learned: add inceptionV3 and densenet121 to your ensembles",
    "1205698": "This was an amazing paper. I created a summary for it. Hope it helps.\n\n![https://drive.google.com/file/d/1RBUH30DJQtaNkNEuYz_3gYj-ZtlF0mLv/view?usp=sharing](paper summary)",
    "1204254": "Great paper indeed! \nAlthough, I would have liked to see the results 'without' fine-tuning all the CNN parameters and just logistic regression setting results; to see the effectiveness of architectures.  ",
    "1189987": "this is a fundamental and basic research question , similar example from my experience is efficientNetB0 and effcientNetB7 could not perform well on face recognition i.e getting embedding from them and comparing whether they are similar or not based on threshold. In Paper the authors claimed their models performed well on 5 out of 8 datasets including image Net , that is quite reasonable , but it did not worked out on my face recognition problem on which mobileNet performs greatly than efficienNet models . \n",
    "1188416": "One of the things I am reminded of with this paper is the general principle I learned early on that theoretically we can solve many of these problems with much smaller networks, less layers, less neurons, etc. but the big difficulty is finding the correct weights that generalize. In general it is easier to find good weights in larger networks, but it is on paper possible to get equivalent performance out of smaller networks. Initialization plays a very important role in this process, pretraining on imagenet is in some sense just a more well-informed initialization. ",
    "1183289": "Speaking as someone doing novel radiography + DL models with an insanely small dataset at work, I can confirm that for me, Chexnet worked a LOT better than imagenet (even with better normalization than what's done in most of the notebooks here). \nYMMV ofc.",
    "1182689": "Great findings,I will read these papers and think why these models which have high accuracy in Imagenet can't get high score in kaggle competitions.",
    "1161568": "Nice work,so ,how do we choose the model?Using the small model?😂",
    "1187998": "I'd like to throw this in - [How useful is self-supervised pre-training](https://openaccess.thecvf.com/content_CVPR_2020/html/Newell_How_Useful_Is_Self-Supervised_Pretraining_for_Visual_Tasks_CVPR_2020_paper.html)\n\nJust something to be aware of before wasting previous GPU hours pretraining model.",
    "1161802": "Interesting findings for sure. Surprised how underresearched this area is given that so few are actually doing just imagenet. Almost everyone is transferring away to different tasks.\n\nEven in this competition, people have claimed gains from one architecture vs the other but it seems fairly inconsistent based on size of model but possibly significant when going to different kinds of models, dense net, effnet, resnet, etc. ",
    "1173438": "Truly fascinating!"
  }
}