{
  "id": 304923,
  "title": "How to reproduce the best model? (YOLOv5 0.631)",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/304923",
  "author_name": "",
  "post_date": "2022-02-03T02:20:46.840511100Z",
  "votes": 12,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi guys.<br>\nI'm facing a weird situation I never encountered.<br>\nI trained a model accidentally got a very good LB 0.631, but I cannot reproduce it. I trained others and got 0.522, 0.589.</p>\n<p>I already checked a few common points about this.</p>\n<ul>\n<li>The random seed has already been set. Here is the code.</li>\n</ul>\n<pre><code>def init_seeds(seed=0):\n    import torch.backends.cudnn as cudnn\n    random.seed(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    cudnn.benchmark, cudnn.deterministic = (False, True) \n</code></pre>\n<ul>\n<li>Exactly the same hyperparameter setting in hyp.py.</li>\n<li>Almost the same performance in my evaluation dataset.<br>\n<img src=\"https://i.imgur.com/LwMhEm0.jpg\" alt=\"\"></li>\n<li>No multi-scale training and no TTA so far</li>\n</ul>\n<p>I know that we can never get the exact same result even with the same training parameter.<br>\nBut this is my first time encountering such a big difference in this.<br>\nHave anyone had a similar experience about it?</p>",
  "messages": [
    {
      "id": "1673818",
      "postDate": "02/03/2022 02:20:46",
      "content": "<p>Hi guys.<br>\nI'm facing a weird situation I never encountered.<br>\nI trained a model accidentally got a very good LB 0.631, but I cannot reproduce it. I trained others and got 0.522, 0.589.</p>\n<p>I already checked a few common points about this.</p>\n<ul>\n<li>The random seed has already been set. Here is the code.</li>\n</ul>\n<pre><code>def init_seeds(seed=0):\n    import torch.backends.cudnn as cudnn\n    random.seed(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    cudnn.benchmark, cudnn.deterministic = (False, True) \n</code></pre>\n<ul>\n<li>Exactly the same hyperparameter setting in hyp.py.</li>\n<li>Almost the same performance in my evaluation dataset.<br>\n<img src=\"https://i.imgur.com/LwMhEm0.jpg\" alt=\"\"></li>\n<li>No multi-scale training and no TTA so far</li>\n</ul>\n<p>I know that we can never get the exact same result even with the same training parameter.<br>\nBut this is my first time encountering such a big difference in this.<br>\nHave anyone had a similar experience about it?</p>",
      "rawMarkdown": "Hi guys.\nI'm facing a weird situation I never encountered.\nI trained a model accidentally got a very good LB 0.631, but I cannot reproduce it. I trained others and got 0.522, 0.589.\n\nI already checked a few common points about this.\n- The random seed has already been set. Here is the code.\n```\ndef init_seeds(seed=0):\n    import torch.backends.cudnn as cudnn\n    random.seed(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    cudnn.benchmark, cudnn.deterministic = (False, True) \n```\n- Exactly the same hyperparameter setting in hyp.py.\n- Almost the same performance in my evaluation dataset.\n![](https://i.imgur.com/LwMhEm0.jpg)\n- No multi-scale training and no TTA so far\n\nI know that we can never get the exact same result even with the same training parameter.\nBut this is my first time encountering such a big difference in this.\nHave anyone had a similar experience about it?",
      "votes": null
    },
    {
      "id": "1673828",
      "postDate": "02/03/2022 02:48:05",
      "content": "<p>Same here!!</p>",
      "rawMarkdown": "Same here!!",
      "votes": null
    },
    {
      "id": "1673853",
      "postDate": "02/03/2022 03:12:38",
      "content": "<p>\"weird situation I never encountered.\"<br>\nwhen data is little, weird things happens<br>\n(e.g. even learning rate matters!)<br>\ni don't think you can 100% control the randomness  <br>\ni thing the augmentation are not the same.<br>\n\"Almost the same performance in my evaluation dataset.\"<br>\nthis shows results are the same in your controlled environment. <br>\nif you are sure that your validation set is good, then this is the performance of your model.<br>\nnote: but if you examine the results, you don't see exact same results but same statistics (e.g. some fp,miss are random) somehow these differences are amplified by the public LB set becuase your validation set and public LB set are different</p>\n<hr>\n<p>\"Have anyone had a similar experience about it?\"<br>\ni think most kagglers are having the same issues </p>\n<hr>\n<p>this is an experience from work (numbers are fake and for illustrative purposes):<br>\ni build a face detectiion models:<br>\nmodel A: recall 99%, miss one black face<br>\nmodel B: recall 98%, miss 2 faces with spectacles<br>\nmodel C: recall 99%, miss one old faces (wrinkles)<br>\ncompetitor : 95%<br>\ni submit 3 models to the client. black box testing returns<br>\nmodel A: 87%<br>\nmodel B: 88%<br>\nmodel C: 90%<br>\ncompetitor : 95%<br>\nturns out that the in the train samples provided by the client, black faces, spectacles faces are rare classes.<br>\nin testing they have equal test images for each variations, e.g.<br>\n5% test = black faces<br>\n5% test = spectacles faces</p>",
      "rawMarkdown": "\"weird situation I never encountered.\"\n\nwhen data is little, weird things happens\n\n(e.g. even learning rate matters!)\n\ni don't think you can 100% control the randomness  \ni thing the augmentation are not the same.\n\n\"Almost the same performance in my evaluation dataset.\"\nthis shows results are the same in your controlled environment. \nif you are sure that your validation set is good, then this is the performance of your model.\n\nnote: but if you examine the results, you don't see exact same results but same statistics (e.g. some fp,miss are random) somehow these differences are amplified by the public LB set becuase your validation set and public LB set are different\n\n- - - -\n\n\n\"Have anyone had a similar experience about it?\"\n\ni think most kagglers are having the same issues \n\n\n\n- - - -\n\n\nthis is an experience from work (numbers are fake and for illustrative purposes):\ni build a face detectiion models:\nmodel A: recall 99%, miss one black face\nmodel B: recall 98%, miss 2 faces with spectacles\nmodel C: recall 99%, miss one old faces (wrinkles)\n\ncompetitor : 95%\n\ni submit 3 models to the client. black box testing returns\nmodel A: 87%\nmodel B: 88%\nmodel C: 90%\ncompetitor : 95%\n\nturns out that the in the train samples provided by the client, black faces, spectacles faces are rare classes.\nin testing they have equal test images for each variations, e.g.\n5% test = black faces\n5% test = spectacles faces",
      "votes": null
    },
    {
      "id": "1673875",
      "postDate": "02/03/2022 03:44:42",
      "content": "<p>Same here, my best model yolov5 with lb 0.62 and I can't reproduced it anymore. With the randomness, even the same data and hyperparameters, i wonder how the top lb can ensure their hyperparameters tuning works on lb. Only 5 submission a day, you can't try much as experiments as you want.</p>",
      "rawMarkdown": "Same here, my best model yolov5 with lb 0.62 and I can't reproduced it anymore. With the randomness, even the same data and hyperparameters, i wonder how the top lb can ensure their hyperparameters tuning works on lb. Only 5 submission a day, you can't try much as experiments as you want.",
      "votes": null
    },
    {
      "id": "1673884",
      "postDate": "02/03/2022 04:02:32",
      "content": "<p>The augmentation in yolov5 is random like mosaic or mixup. So each time you train it will be different. Just my guess :)</p>",
      "rawMarkdown": "The augmentation in yolov5 is random like mosaic or mixup. So each time you train it will be different. Just my guess :)",
      "votes": null
    },
    {
      "id": "1673954",
      "postDate": "02/03/2022 05:46:54",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23Thanks\" target=\"_blank\">@hengck23Thanks</a> to your advice<br>\nI got your point, it's a really great insight.<br>\nI also agree with the results are amplified by the public LB, but I got almost 20% difference in LB scores. So I wonder that maybe there is something I didn't notice it.</p>\n<p>And I guess the data in LB should be very large, because of the average running time. This one is also a reason I think shouldn't behave so much difference in models by same training hyperparameters.</p>\n<p><strong>\"I think most kagglers are having the same issues\"</strong><br>\nYeah! Exactly!</p>",
      "rawMarkdown": "hengck23Thanks to your advice\nI got your point, it's a really great insight.\nI also agree with the results are amplified by the public LB, but I got almost 20% difference in LB scores. So I wonder that maybe there is something I didn't notice it.\n\nAnd I guess the data in LB should be very large, because of the average running time. This one is also a reason I think shouldn't behave so much difference in models by same training hyperparameters.\n\n**\"I think most kagglers are having the same issues\"**\nYeah! Exactly!",
      "votes": null
    },
    {
      "id": "1673955",
      "postDate": "02/03/2022 05:48:33",
      "content": "<p>Yes, I see.<br>\nAugmentation must not be exactly same with each training.<br>\nI just feel so surprise why it can be so different!!!</p>",
      "rawMarkdown": "Yes, I see.\nAugmentation must not be exactly same with each training.\nI just feel so surprise why it can be so different!!!",
      "votes": null
    },
    {
      "id": "1673957",
      "postDate": "02/03/2022 05:49:43",
      "content": "<p>I hope you can find a way for it and share with it!!!<br>\nGood luck guys~~</p>",
      "rawMarkdown": "I hope you can find a way for it and share with it!!!\nGood luck guys~~",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1673828,
      "author_name": "chiehpower",
      "author_url": "",
      "post_date": "02/03/2022 02:48:05",
      "content": "<p>Same here!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1673853,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/03/2022 03:12:38",
      "content": "<p>\"weird situation I never encountered.\"<br>\nwhen data is little, weird things happens<br>\n(e.g. even learning rate matters!)<br>\ni don't think you can 100% control the randomness  <br>\ni thing the augmentation are not the same.<br>\n\"Almost the same performance in my evaluation dataset.\"<br>\nthis shows results are the same in your controlled environment. <br>\nif you are sure that your validation set is good, then this is the performance of your model.<br>\nnote: but if you examine the results, you don't see exact same results but same statistics (e.g. some fp,miss are random) somehow these differences are amplified by the public LB set becuase your validation set and public LB set are different</p>\n<hr>\n<p>\"Have anyone had a similar experience about it?\"<br>\ni think most kagglers are having the same issues </p>\n<hr>\n<p>this is an experience from work (numbers are fake and for illustrative purposes):<br>\ni build a face detectiion models:<br>\nmodel A: recall 99%, miss one black face<br>\nmodel B: recall 98%, miss 2 faces with spectacles<br>\nmodel C: recall 99%, miss one old faces (wrinkles)<br>\ncompetitor : 95%<br>\ni submit 3 models to the client. black box testing returns<br>\nmodel A: 87%<br>\nmodel B: 88%<br>\nmodel C: 90%<br>\ncompetitor : 95%<br>\nturns out that the in the train samples provided by the client, black faces, spectacles faces are rare classes.<br>\nin testing they have equal test images for each variations, e.g.<br>\n5% test = black faces<br>\n5% test = spectacles faces</p>",
      "votes": null,
      "replies": [
        {
          "id": 1673954,
          "author_name": "lilinchen",
          "author_url": "",
          "post_date": "02/03/2022 05:46:54",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23Thanks\" target=\"_blank\">@hengck23Thanks</a> to your advice<br>\nI got your point, it's a really great insight.<br>\nI also agree with the results are amplified by the public LB, but I got almost 20% difference in LB scores. So I wonder that maybe there is something I didn't notice it.</p>\n<p>And I guess the data in LB should be very large, because of the average running time. This one is also a reason I think shouldn't behave so much difference in models by same training hyperparameters.</p>\n<p><strong>\"I think most kagglers are having the same issues\"</strong><br>\nYeah! Exactly!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1673875,
      "author_name": "locbaop",
      "author_url": "",
      "post_date": "02/03/2022 03:44:42",
      "content": "<p>Same here, my best model yolov5 with lb 0.62 and I can't reproduced it anymore. With the randomness, even the same data and hyperparameters, i wonder how the top lb can ensure their hyperparameters tuning works on lb. Only 5 submission a day, you can't try much as experiments as you want.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1673957,
          "author_name": "lilinchen",
          "author_url": "",
          "post_date": "02/03/2022 05:49:43",
          "content": "<p>I hope you can find a way for it and share with it!!!<br>\nGood luck guys~~</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1673884,
      "author_name": "magiccard",
      "author_url": "",
      "post_date": "02/03/2022 04:02:32",
      "content": "<p>The augmentation in yolov5 is random like mosaic or mixup. So each time you train it will be different. Just my guess :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1673955,
          "author_name": "lilinchen",
          "author_url": "",
          "post_date": "02/03/2022 05:48:33",
          "content": "<p>Yes, I see.<br>\nAugmentation must not be exactly same with each training.<br>\nI just feel so surprise why it can be so different!!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1673818": "Hi guys.\nI'm facing a weird situation I never encountered.\nI trained a model accidentally got a very good LB 0.631, but I cannot reproduce it. I trained others and got 0.522, 0.589.\n\nI already checked a few common points about this.\n- The random seed has already been set. Here is the code.\n```\ndef init_seeds(seed=0):\n    import torch.backends.cudnn as cudnn\n    random.seed(seed)\n    np.random.seed(seed)\n    torch.manual_seed(seed)\n    cudnn.benchmark, cudnn.deterministic = (False, True) \n```\n- Exactly the same hyperparameter setting in hyp.py.\n- Almost the same performance in my evaluation dataset.\n![](https://i.imgur.com/LwMhEm0.jpg)\n- No multi-scale training and no TTA so far\n\nI know that we can never get the exact same result even with the same training parameter.\nBut this is my first time encountering such a big difference in this.\nHave anyone had a similar experience about it?",
    "1673828": "Same here!!",
    "1673853": "\"weird situation I never encountered.\"\n\nwhen data is little, weird things happens\n\n(e.g. even learning rate matters!)\n\ni don't think you can 100% control the randomness  \ni thing the augmentation are not the same.\n\n\"Almost the same performance in my evaluation dataset.\"\nthis shows results are the same in your controlled environment. \nif you are sure that your validation set is good, then this is the performance of your model.\n\nnote: but if you examine the results, you don't see exact same results but same statistics (e.g. some fp,miss are random) somehow these differences are amplified by the public LB set becuase your validation set and public LB set are different\n\n- - - -\n\n\n\"Have anyone had a similar experience about it?\"\n\ni think most kagglers are having the same issues \n\n\n\n- - - -\n\n\nthis is an experience from work (numbers are fake and for illustrative purposes):\ni build a face detectiion models:\nmodel A: recall 99%, miss one black face\nmodel B: recall 98%, miss 2 faces with spectacles\nmodel C: recall 99%, miss one old faces (wrinkles)\n\ncompetitor : 95%\n\ni submit 3 models to the client. black box testing returns\nmodel A: 87%\nmodel B: 88%\nmodel C: 90%\ncompetitor : 95%\n\nturns out that the in the train samples provided by the client, black faces, spectacles faces are rare classes.\nin testing they have equal test images for each variations, e.g.\n5% test = black faces\n5% test = spectacles faces",
    "1673875": "Same here, my best model yolov5 with lb 0.62 and I can't reproduced it anymore. With the randomness, even the same data and hyperparameters, i wonder how the top lb can ensure their hyperparameters tuning works on lb. Only 5 submission a day, you can't try much as experiments as you want.",
    "1673884": "The augmentation in yolov5 is random like mosaic or mixup. So each time you train it will be different. Just my guess :)",
    "1673954": "hengck23Thanks to your advice\nI got your point, it's a really great insight.\nI also agree with the results are amplified by the public LB, but I got almost 20% difference in LB scores. So I wonder that maybe there is something I didn't notice it.\n\nAnd I guess the data in LB should be very large, because of the average running time. This one is also a reason I think shouldn't behave so much difference in models by same training hyperparameters.\n\n**\"I think most kagglers are having the same issues\"**\nYeah! Exactly!",
    "1673955": "Yes, I see.\nAugmentation must not be exactly same with each training.\nI just feel so surprise why it can be so different!!!",
    "1673957": "I hope you can find a way for it and share with it!!!\nGood luck guys~~"
  },
  "source": "meta"
}