{
  "id": 301452,
  "title": "About the Training Reproducibility of LB: .579",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/301452",
  "author_name": "Harshit Sheoran",
  "post_date": "2022-01-17T19:55:17.537000",
  "votes": 25,
  "comment_count": 7,
  "views": 0,
  "content": "<p>So, I took <a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a>'s training notebook to start with, trained the model with a bit bigger image size of 3840, got a score of .588 with no inference upscaling.<br>\nThen tried changing a few numbers to do further experiment, and  retrained model a lot of times, everytime it scored way below .588</p>\n<ol>\n<li>bs increase scored .495</li>\n<li>Rerun the same code again to score .493</li>\n<li>Trying out some hyperparams on inference side .492</li>\n<li>Trying Adam optimizer scoring .554</li>\n<li>ReTrying Adam optimizer with a different batch size .552 (This was a rerun of .588 rep fail with Adam)</li>\n<li>Changing Image sizes to 3600, 3000, 3200 all scoring less.</li>\n</ol>\n<p>Basically making it completely random when we get good scores and when we don't (except for Adam, we can put our hopes there).</p>\n<p>Then thought that YoloV5-S6 is not stable, ran M6 one time and scored .556 with image size 3200, ran L6 one time with image size 2560 and scored .427</p>\n<p>Why is there so much randomness that the same training can go from scoring .588 to .495 by just doubling the batch size and ofc, .588 retrain is also not happening.</p>\n<p>While this was happening, I submitted only when my CV was increasing.</p>\n<p>CV for .588:<br>\nP          R     mAP@.5 mAP@<br>\n0.864      0.644      0.725      0.339</p>\n<p>Same Fold Best CV but failed miserably on LB:<br>\nP           R             mAP@.5  mAP@<br>\n0.84      0.672      0.744      0.352</p>\n<p>Updates:</p>\n<ol>\n<li>Re-M6 .581</li>\n</ol>\n<p>Guidance for reducing this chaos of randomness would be appreciated.<br>\nThank You.</p>",
  "messages": [
    {
      "id": 1653703,
      "postDate": "2022-01-17T19:55:17.537Z",
      "content": "<p>So, I took <a href=\"https://www.kaggle.com/steamedsheep\" target=\"_blank\">@steamedsheep</a>'s training notebook to start with, trained the model with a bit bigger image size of 3840, got a score of .588 with no inference upscaling.<br>\nThen tried changing a few numbers to do further experiment, and  retrained model a lot of times, everytime it scored way below .588</p>\n<ol>\n<li>bs increase scored .495</li>\n<li>Rerun the same code again to score .493</li>\n<li>Trying out some hyperparams on inference side .492</li>\n<li>Trying Adam optimizer scoring .554</li>\n<li>ReTrying Adam optimizer with a different batch size .552 (This was a rerun of .588 rep fail with Adam)</li>\n<li>Changing Image sizes to 3600, 3000, 3200 all scoring less.</li>\n</ol>\n<p>Basically making it completely random when we get good scores and when we don't (except for Adam, we can put our hopes there).</p>\n<p>Then thought that YoloV5-S6 is not stable, ran M6 one time and scored .556 with image size 3200, ran L6 one time with image size 2560 and scored .427</p>\n<p>Why is there so much randomness that the same training can go from scoring .588 to .495 by just doubling the batch size and ofc, .588 retrain is also not happening.</p>\n<p>While this was happening, I submitted only when my CV was increasing.</p>\n<p>CV for .588:<br>\nP          R     mAP@.5 mAP@<br>\n0.864      0.644      0.725      0.339</p>\n<p>Same Fold Best CV but failed miserably on LB:<br>\nP           R             mAP@.5  mAP@<br>\n0.84      0.672      0.744      0.352</p>\n<p>Updates:</p>\n<ol>\n<li>Re-M6 .581</li>\n</ol>\n<p>Guidance for reducing this chaos of randomness would be appreciated.<br>\nThank You.</p>",
      "rawMarkdown": "So, I took @steamedsheep's training notebook to start with, trained the model with a bit bigger image size of 3840, got a score of .588 with no inference upscaling.\nThen tried changing a few numbers to do further experiment, and  retrained model a lot of times, everytime it scored way below .588\n1. bs increase scored .495\n2. Rerun the same code again to score .493\n3. Trying out some hyperparams on inference side .492\n3. Trying Adam optimizer scoring .554\n4. ReTrying Adam optimizer with a different batch size .552 (This was a rerun of .588 rep fail with Adam)\n5. Changing Image sizes to 3600, 3000, 3200 all scoring less.\n\nBasically making it completely random when we get good scores and when we don't (except for Adam, we can put our hopes there).\n\nThen thought that YoloV5-S6 is not stable, ran M6 one time and scored .556 with image size 3200, ran L6 one time with image size 2560 and scored .427\n\nWhy is there so much randomness that the same training can go from scoring .588 to .495 by just doubling the batch size and ofc, .588 retrain is also not happening.\n\nWhile this was happening, I submitted only when my CV was increasing.\n\nCV for .588:\nP          R     mAP@.5 mAP@\n0.864      0.644      0.725      0.339\n\nSame Fold Best CV but failed miserably on LB:\nP           R             mAP@.5  mAP@\n0.84      0.672      0.744      0.352\n\n\nUpdates:\n1. Re-M6 .581\n\nGuidance for reducing this chaos of randomness would be appreciated.\nThank You.",
      "votes": 25
    },
    {
      "id": 1653729,
      "postDate": "2022-01-17T20:19:14.863Z",
      "content": "<p>Welcome to the club 😄<br>\nThis is why I aked forTOP5 final model video on training data after competition :) I would like to see inference on training. It could be great research topic.</p>\n<p>I am really waiting for solution description and …. model validations procedure. This is part I have spent most of a time … unfortunately without success.  At one point, I even stated that there was something wrong with the implementation of the f2 metric on the Kaggle side, but I guess it's just my lack of knowledge in this area :) Today <a href=\"https://www.kaggle.com/lukaszborecki\" target=\"_blank\">@lukaszborecki</a> made interesting experiment - just submit crap model …. absolutely crap (with almost random bboxes) and it appeared that … is good (almost like first yolov5 models presented here) on LB :) </p>",
      "rawMarkdown": "Welcome to the club 😄\nThis is why I aked forTOP5 final model video on training data after competition :) I would like to see inference on training. It could be great research topic.\n\nI am really waiting for solution description and .... model validations procedure. This is part I have spent most of a time ... unfortunately without success.  At one point, I even stated that there was something wrong with the implementation of the f2 metric on the Kaggle side, but I guess it's just my lack of knowledge in this area :) Today @lukaszborecki made interesting experiment - just submit crap model .... absolutely crap (with almost random bboxes) and it appeared that ... is good (almost like first yolov5 models presented here) on LB :) ",
      "votes": 5,
      "replies": [
        {
          "id": 1653779,
          "postDate": "2022-01-17T21:17:45.997Z",
          "content": "<p>here is the image from inference Remek said and score is 0.45 but bboxes are crap and model misses bigger starfishes, medium one also <img src=\"https://i.ibb.co/bQxW0wX/results-12-3.png\" alt=\"\"></p>\n<p>and here how should look correct boxes</p>\n<p><img src=\"https://i.ibb.co/5G6b6Hk/results-12-3-1.png\" alt=\"\"></p>",
          "rawMarkdown": "here is the image from inference Remek said and score is 0.45 but bboxes are crap and model misses bigger starfishes, medium one also ![](https://i.ibb.co/bQxW0wX/results-12-3.png)\n\n\nand here how should look correct boxes\n\n![](https://i.ibb.co/5G6b6Hk/results-12-3-1.png)",
          "votes": 3
        },
        {
          "id": 1653790,
          "postDate": "2022-01-17T21:31:27.077Z",
          "content": "<p>Both photos are from valid …. and both models (first one … according to image resizing trends we just submit for fun) … got almost comparable score 😂</p>",
          "rawMarkdown": "Both photos are from valid .... and both models (first one ... according to image resizing trends we just submit for fun) ... got almost comparable score 😂"
        },
        {
          "id": 1654405,
          "postDate": "2022-01-18T13:26:14.880Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1656355,
          "postDate": "2022-01-19T09:40:12.630Z",
          "content": "<p>one sample doesn't prove enough, unless many more  enough.</p>",
          "rawMarkdown": "one sample doesn't prove enough, unless many more  enough."
        }
      ]
    },
    {
      "id": 1654417,
      "postDate": "2022-01-18T13:43:53.130Z",
      "content": "<p>What is CV?</p>",
      "rawMarkdown": "What is CV?"
    },
    {
      "id": 1653882,
      "postDate": "2022-01-18T00:24:31.700Z",
      "content": "<p>One possible reason for high CV but lower LB is that maybe you are overfitting Val data.</p>\n<p>BTW, maybe I am wrong. But it seems the seed of yolov5 model will change for each training. Therefore, there are randomness.</p>",
      "rawMarkdown": "One possible reason for high CV but lower LB is that maybe you are overfitting Val data.\n\nBTW, maybe I am wrong. But it seems the seed of yolov5 model will change for each training. Therefore, there are randomness."
    }
  ],
  "comments": [
    {
      "id": 1653729,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-01-17T20:19:14.863000",
      "content": "<p>Welcome to the club 😄<br>\nThis is why I aked forTOP5 final model video on training data after competition :) I would like to see inference on training. It could be great research topic.</p>\n<p>I am really waiting for solution description and …. model validations procedure. This is part I have spent most of a time … unfortunately without success.  At one point, I even stated that there was something wrong with the implementation of the f2 metric on the Kaggle side, but I guess it's just my lack of knowledge in this area :) Today <a href=\"https://www.kaggle.com/lukaszborecki\" target=\"_blank\">@lukaszborecki</a> made interesting experiment - just submit crap model …. absolutely crap (with almost random bboxes) and it appeared that … is good (almost like first yolov5 models presented here) on LB :) </p>",
      "votes": 5,
      "replies": [
        {
          "id": 1653779,
          "author_name": "Lukasz Borecki",
          "author_url": "",
          "post_date": "2022-01-17T21:17:45.997000",
          "content": "<p>here is the image from inference Remek said and score is 0.45 but bboxes are crap and model misses bigger starfishes, medium one also <img src=\"https://i.ibb.co/bQxW0wX/results-12-3.png\" alt=\"\"></p>\n<p>and here how should look correct boxes</p>\n<p><img src=\"https://i.ibb.co/5G6b6Hk/results-12-3-1.png\" alt=\"\"></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1653790,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-01-17T21:31:27.077000",
          "content": "<p>Both photos are from valid …. and both models (first one … according to image resizing trends we just submit for fun) … got almost comparable score 😂</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1654405,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-01-18T13:26:14.880000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1656355,
          "author_name": "dragon zhang",
          "author_url": "",
          "post_date": "2022-01-19T09:40:12.630000",
          "content": "<p>one sample doesn't prove enough, unless many more  enough.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1654417,
      "author_name": "Kira yang",
      "author_url": "",
      "post_date": "2022-01-18T13:43:53.130000",
      "content": "<p>What is CV?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1653882,
      "author_name": "DeepInvolution",
      "author_url": "",
      "post_date": "2022-01-18T00:24:31.700000",
      "content": "<p>One possible reason for high CV but lower LB is that maybe you are overfitting Val data.</p>\n<p>BTW, maybe I am wrong. But it seems the seed of yolov5 model will change for each training. Therefore, there are randomness.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1653703": "So, I took @steamedsheep's training notebook to start with, trained the model with a bit bigger image size of 3840, got a score of .588 with no inference upscaling.\nThen tried changing a few numbers to do further experiment, and  retrained model a lot of times, everytime it scored way below .588\n1. bs increase scored .495\n2. Rerun the same code again to score .493\n3. Trying out some hyperparams on inference side .492\n3. Trying Adam optimizer scoring .554\n4. ReTrying Adam optimizer with a different batch size .552 (This was a rerun of .588 rep fail with Adam)\n5. Changing Image sizes to 3600, 3000, 3200 all scoring less.\n\nBasically making it completely random when we get good scores and when we don't (except for Adam, we can put our hopes there).\n\nThen thought that YoloV5-S6 is not stable, ran M6 one time and scored .556 with image size 3200, ran L6 one time with image size 2560 and scored .427\n\nWhy is there so much randomness that the same training can go from scoring .588 to .495 by just doubling the batch size and ofc, .588 retrain is also not happening.\n\nWhile this was happening, I submitted only when my CV was increasing.\n\nCV for .588:\nP          R     mAP@.5 mAP@\n0.864      0.644      0.725      0.339\n\nSame Fold Best CV but failed miserably on LB:\nP           R             mAP@.5  mAP@\n0.84      0.672      0.744      0.352\n\n\nUpdates:\n1. Re-M6 .581\n\nGuidance for reducing this chaos of randomness would be appreciated.\nThank You.",
    "1653729": "Welcome to the club 😄\nThis is why I aked forTOP5 final model video on training data after competition :) I would like to see inference on training. It could be great research topic.\n\nI am really waiting for solution description and .... model validations procedure. This is part I have spent most of a time ... unfortunately without success.  At one point, I even stated that there was something wrong with the implementation of the f2 metric on the Kaggle side, but I guess it's just my lack of knowledge in this area :) Today @lukaszborecki made interesting experiment - just submit crap model .... absolutely crap (with almost random bboxes) and it appeared that ... is good (almost like first yolov5 models presented here) on LB :) ",
    "1654417": "What is CV?",
    "1653882": "One possible reason for high CV but lower LB is that maybe you are overfitting Val data.\n\nBTW, maybe I am wrong. But it seems the seed of yolov5 model will change for each training. Therefore, there are randomness."
  }
}