{
  "id": 217296,
  "title": "17 stage training",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/217296",
  "author_name": "",
  "post_date": "2021-02-06T08:00:37.851453Z",
  "votes": 27,
  "comment_count": 5,
  "views": 0,
  "content": "<p>One of the techniques that people seem to keep refining in this competition is various different stages of training. Student-teacher models, injecting information about the annotations, model distillation, training on subsets of data, using different augmentation intensities. These are potentially useful tools, but I want to warn against small gains that might be had here. </p>\n<p>From what I have seen there seems to be nearly immeasurable gains from these increasingly complex procedures and I think quite a bit of it can simply be explained by warm restarting models and continually cherry-picking checkpoints multiple stages through. </p>\n<p>To explain further what one can do to fool themselves into thinking they have found better performance is train a model for 35 epochs. Find the best validation performance was achieved on epoch 24. Take that epoch 24 checkpoint again and retry training with no other changes or trivial changes like a slightly lower learning rate or more intense augmentations and train another 5 epochs. This time you find some more progress has been made up to the checkpoint from epoch 28. </p>\n<p>You can repeat this process many times and even do multiple trials waiting to get a result that goes in the right direction for validation but it does not definitively prove that you have done something to improve your model, it just means you have found a way to further cherry pick in the direction of the validation set. </p>\n<p>Just a general warning, probably not worth permanently complicating your training pipeline because one time you got a .001 boost </p>",
  "messages": [
    {
      "id": "1188405",
      "postDate": "02/06/2021 08:00:37",
      "content": "<p>One of the techniques that people seem to keep refining in this competition is various different stages of training. Student-teacher models, injecting information about the annotations, model distillation, training on subsets of data, using different augmentation intensities. These are potentially useful tools, but I want to warn against small gains that might be had here. </p>\n<p>From what I have seen there seems to be nearly immeasurable gains from these increasingly complex procedures and I think quite a bit of it can simply be explained by warm restarting models and continually cherry-picking checkpoints multiple stages through. </p>\n<p>To explain further what one can do to fool themselves into thinking they have found better performance is train a model for 35 epochs. Find the best validation performance was achieved on epoch 24. Take that epoch 24 checkpoint again and retry training with no other changes or trivial changes like a slightly lower learning rate or more intense augmentations and train another 5 epochs. This time you find some more progress has been made up to the checkpoint from epoch 28. </p>\n<p>You can repeat this process many times and even do multiple trials waiting to get a result that goes in the right direction for validation but it does not definitively prove that you have done something to improve your model, it just means you have found a way to further cherry pick in the direction of the validation set. </p>\n<p>Just a general warning, probably not worth permanently complicating your training pipeline because one time you got a .001 boost </p>",
      "rawMarkdown": "One of the techniques that people seem to keep refining in this competition is various different stages of training. Student-teacher models, injecting information about the annotations, model distillation, training on subsets of data, using different augmentation intensities. These are potentially useful tools, but I want to warn against small gains that might be had here. \n\nFrom what I have seen there seems to be nearly immeasurable gains from these increasingly complex procedures and I think quite a bit of it can simply be explained by warm restarting models and continually cherry-picking checkpoints multiple stages through. \n\nTo explain further what one can do to fool themselves into thinking they have found better performance is train a model for 35 epochs. Find the best validation performance was achieved on epoch 24. Take that epoch 24 checkpoint again and retry training with no other changes or trivial changes like a slightly lower learning rate or more intense augmentations and train another 5 epochs. This time you find some more progress has been made up to the checkpoint from epoch 28. \n\nYou can repeat this process many times and even do multiple trials waiting to get a result that goes in the right direction for validation but it does not definitively prove that you have done something to improve your model, it just means you have found a way to further cherry pick in the direction of the validation set. \n\nJust a general warning, probably not worth permanently complicating your training pipeline because one time you got a .001 boost",
      "votes": null
    },
    {
      "id": "1188572",
      "postDate": "02/06/2021 10:58:22",
      "content": "<p>Hi,you mean is we use the trained model and adjust the lr to a min value,then we continue train the model from the best model checkpoint?So,how can we choost the min lr value?Thanks!</p>",
      "rawMarkdown": "Hi,you mean is we use the trained model and adjust the lr to a min value,then we continue train the model from the best model checkpoint?So,how can we choost the min lr value?Thanks!",
      "votes": null
    },
    {
      "id": "1189081",
      "postDate": "02/06/2021 17:43:46",
      "content": "<p>Tottally agree. CNNs training is like dark magic being a black box too. We do a lot of things and accidentanly may take better results by refining but actually it was just a luck event and we fool ourselfs that possible we found a new technique worth to be a research paper.</p>",
      "rawMarkdown": "Tottally agree. CNNs training is like dark magic being a black box too. We do a lot of things and accidentanly may take better results by refining but actually it was just a luck event and we fool ourselfs that possible we found a new technique worth to be a research paper.",
      "votes": null
    },
    {
      "id": "1189559",
      "postDate": "02/07/2021 05:41:05",
      "content": "<p>My overall point is you can continue to do more experiments but its sort of like p-hacking in statistics. If you do enough trials even with just random noise you expect a certain number of results to improve, but they are statistically the same thing, just noise in the right direction. </p>\n<p>I have no recommendation in terms of learning rate, I was just displaying that you can trick yourself by reading into results too much. People have been posting more and more stages of training and I have not seen any evidence that any of it really works. </p>",
      "rawMarkdown": "My overall point is you can continue to do more experiments but its sort of like p-hacking in statistics. If you do enough trials even with just random noise you expect a certain number of results to improve, but they are statistically the same thing, just noise in the right direction. \n\nI have no recommendation in terms of learning rate, I was just displaying that you can trick yourself by reading into results too much. People have been posting more and more stages of training and I have not seen any evidence that any of it really works.",
      "votes": null
    },
    {
      "id": "1190586",
      "postDate": "02/07/2021 19:57:21",
      "content": "<p>When I read the title of this discussion, my first thought was “oh no…” 😅<br>\nThank you for the excellent discussion points and please keep them coming</p>",
      "rawMarkdown": "When I read the title of this discussion, my first thought was “oh no...” 😅\nThank you for the excellent discussion points and please keep them coming",
      "votes": null
    },
    {
      "id": "1208061",
      "postDate": "02/18/2021 06:24:19",
      "content": "<p>Completely valid point. <br>\nIn a recent RPS competition I experience the same issue: growth of the model complexity from the some point adds nothing but noise.<br>\nIn RPS it was possible to detect it by visualization of agent distribution, and clearly see that it tends towards random behavior from some stage. But here it is easy to fool yourself by tiny CV boost as you mentioned.<br>\nBy the way, best strategy here is 73 stage training, but lets keep it secret :)</p>",
      "rawMarkdown": "Completely valid point. \nIn a recent RPS competition I experience the same issue: growth of the model complexity from the some point adds nothing but noise.\nIn RPS it was possible to detect it by visualization of agent distribution, and clearly see that it tends towards random behavior from some stage. But here it is easy to fool yourself by tiny CV boost as you mentioned.\nBy the way, best strategy here is 73 stage training, but lets keep it secret :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1188572,
      "author_name": "bcwang",
      "author_url": "",
      "post_date": "02/06/2021 10:58:22",
      "content": "<p>Hi,you mean is we use the trained model and adjust the lr to a min value,then we continue train the model from the best model checkpoint?So,how can we choost the min lr value?Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1189559,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "02/07/2021 05:41:05",
          "content": "<p>My overall point is you can continue to do more experiments but its sort of like p-hacking in statistics. If you do enough trials even with just random noise you expect a certain number of results to improve, but they are statistically the same thing, just noise in the right direction. </p>\n<p>I have no recommendation in terms of learning rate, I was just displaying that you can trick yourself by reading into results too much. People have been posting more and more stages of training and I have not seen any evidence that any of it really works. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1189081,
      "author_name": "",
      "author_url": "",
      "post_date": "02/06/2021 17:43:46",
      "content": "<p>Tottally agree. CNNs training is like dark magic being a black box too. We do a lot of things and accidentanly may take better results by refining but actually it was just a luck event and we fool ourselfs that possible we found a new technique worth to be a research paper.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1190586,
      "author_name": "reubenschmidt",
      "author_url": "",
      "post_date": "02/07/2021 19:57:21",
      "content": "<p>When I read the title of this discussion, my first thought was “oh no…” 😅<br>\nThank you for the excellent discussion points and please keep them coming</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1208061,
      "author_name": "glebkum",
      "author_url": "",
      "post_date": "02/18/2021 06:24:19",
      "content": "<p>Completely valid point. <br>\nIn a recent RPS competition I experience the same issue: growth of the model complexity from the some point adds nothing but noise.<br>\nIn RPS it was possible to detect it by visualization of agent distribution, and clearly see that it tends towards random behavior from some stage. But here it is easy to fool yourself by tiny CV boost as you mentioned.<br>\nBy the way, best strategy here is 73 stage training, but lets keep it secret :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1188405": "One of the techniques that people seem to keep refining in this competition is various different stages of training. Student-teacher models, injecting information about the annotations, model distillation, training on subsets of data, using different augmentation intensities. These are potentially useful tools, but I want to warn against small gains that might be had here. \n\nFrom what I have seen there seems to be nearly immeasurable gains from these increasingly complex procedures and I think quite a bit of it can simply be explained by warm restarting models and continually cherry-picking checkpoints multiple stages through. \n\nTo explain further what one can do to fool themselves into thinking they have found better performance is train a model for 35 epochs. Find the best validation performance was achieved on epoch 24. Take that epoch 24 checkpoint again and retry training with no other changes or trivial changes like a slightly lower learning rate or more intense augmentations and train another 5 epochs. This time you find some more progress has been made up to the checkpoint from epoch 28. \n\nYou can repeat this process many times and even do multiple trials waiting to get a result that goes in the right direction for validation but it does not definitively prove that you have done something to improve your model, it just means you have found a way to further cherry pick in the direction of the validation set. \n\nJust a general warning, probably not worth permanently complicating your training pipeline because one time you got a .001 boost",
    "1188572": "Hi,you mean is we use the trained model and adjust the lr to a min value,then we continue train the model from the best model checkpoint?So,how can we choost the min lr value?Thanks!",
    "1189081": "Tottally agree. CNNs training is like dark magic being a black box too. We do a lot of things and accidentanly may take better results by refining but actually it was just a luck event and we fool ourselfs that possible we found a new technique worth to be a research paper.",
    "1189559": "My overall point is you can continue to do more experiments but its sort of like p-hacking in statistics. If you do enough trials even with just random noise you expect a certain number of results to improve, but they are statistically the same thing, just noise in the right direction. \n\nI have no recommendation in terms of learning rate, I was just displaying that you can trick yourself by reading into results too much. People have been posting more and more stages of training and I have not seen any evidence that any of it really works.",
    "1190586": "When I read the title of this discussion, my first thought was “oh no...” 😅\nThank you for the excellent discussion points and please keep them coming",
    "1208061": "Completely valid point. \nIn a recent RPS competition I experience the same issue: growth of the model complexity from the some point adds nothing but noise.\nIn RPS it was possible to detect it by visualization of agent distribution, and clearly see that it tends towards random behavior from some stage. But here it is easy to fool yourself by tiny CV boost as you mentioned.\nBy the way, best strategy here is 73 stage training, but lets keep it secret :)"
  },
  "source": "meta"
}