{
  "id": 549438,
  "title": "Identical Submissions Getting Such Different Scores",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/549438",
  "author_name": "",
  "post_date": "2024-12-02T12:18:07.706228100Z",
  "votes": null,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hey everyone,</p>\n<p>I’m a bit confused. I submitted the exact same notebook version twice, expecting similar scores, but the results were really different:</p>\n<ul>\n<li>First submission: 0.634</li>\n<li>Second submission: 0.477</li>\n</ul>\n<p>I thought that  only the inference part is run on 26% of the hidden dataset. So, is the model being retrained again during this process?<br>\nIf that’s the case, is it a common practice to set seeds to avoid this variability?<br>\nOr maybe the system evaluates different 26% subsets of the dataset each time?</p>\n<p>Has anyone experienced something similar? Would appreciate any insights on this.</p>\n<p>Thanks!</p>",
  "messages": [
    {
      "id": "3061149",
      "postDate": "12/02/2024 12:18:07",
      "content": "<p>Hey everyone,</p>\n<p>I’m a bit confused. I submitted the exact same notebook version twice, expecting similar scores, but the results were really different:</p>\n<ul>\n<li>First submission: 0.634</li>\n<li>Second submission: 0.477</li>\n</ul>\n<p>I thought that  only the inference part is run on 26% of the hidden dataset. So, is the model being retrained again during this process?<br>\nIf that’s the case, is it a common practice to set seeds to avoid this variability?<br>\nOr maybe the system evaluates different 26% subsets of the dataset each time?</p>\n<p>Has anyone experienced something similar? Would appreciate any insights on this.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hey everyone,\n\nI’m a bit confused. I submitted the exact same notebook version twice, expecting similar scores, but the results were really different:\n- First submission: 0.634\n- Second submission: 0.477\n\nI thought that  only the inference part is run on 26% of the hidden dataset. So, is the model being retrained again during this process?\nIf that’s the case, is it a common practice to set seeds to avoid this variability?\nOr maybe the system evaluates different 26% subsets of the dataset each time?\n\nHas anyone experienced something similar? Would appreciate any insights on this.\n\nThanks!",
      "votes": null
    },
    {
      "id": "3061193",
      "postDate": "12/02/2024 12:43:29",
      "content": "<p>No. And pretty sure is about your code.</p>",
      "rawMarkdown": "No. And pretty sure is about your code.",
      "votes": null
    },
    {
      "id": "3061194",
      "postDate": "12/02/2024 12:45:26",
      "content": "<p>I experienced getting the same score when submitting the same notebook. Do you have training included in your notebook? As far as I'm aware, the system evaluates the same 26% subset each time.</p>",
      "rawMarkdown": "I experienced getting the same score when submitting the same notebook. Do you have training included in your notebook? As far as I'm aware, the system evaluates the same 26% subset each time.",
      "votes": null
    },
    {
      "id": "3061215",
      "postDate": "12/02/2024 13:07:48",
      "content": "<p>Yeah, the training of the model is in the same notebook as the submission. </p>",
      "rawMarkdown": "Yeah, the training of the model is in the same notebook as the submission.",
      "votes": null
    },
    {
      "id": "3061234",
      "postDate": "12/02/2024 13:38:07",
      "content": "<p>Then that's why, as training using the same set-up can get stuck in different local minima and learn differently.</p>",
      "rawMarkdown": "Then that's why, as training using the same set-up can get stuck in different local minima and learn differently.",
      "votes": null
    },
    {
      "id": "3061454",
      "postDate": "12/02/2024 16:36:13",
      "content": "<p>\"If that’s the case, is it a common practice to set seeds to avoid this variability?\" <br>\nThe code is executed as you submit it upon hidden test set once. But if your code uses RNG sure, always set a seed.</p>",
      "rawMarkdown": "\"If that’s the case, is it a common practice to set seeds to avoid this variability?\" \nThe code is executed as you submit it upon hidden test set once. But if your code uses RNG sure, always set a seed.",
      "votes": null
    },
    {
      "id": "3061677",
      "postDate": "12/02/2024 23:27:58",
      "content": "<p>Agree with the assessments below…but, wow!  0.634 and 0.477?  That's a big difference.  Awful lot of variance in the training process.  If your models are that different each time, I would make sure you're saving the weights for each submission.  Although, I suppose setting seed achieves a similar result.  Are your local CV scores showing the same variance?</p>",
      "rawMarkdown": "Agree with the assessments below...but, wow!  0.634 and 0.477?  That's a big difference.  Awful lot of variance in the training process.  If your models are that different each time, I would make sure you're saving the weights for each submission.  Although, I suppose setting seed achieves a similar result.  Are your local CV scores showing the same variance?",
      "votes": null
    },
    {
      "id": "3061680",
      "postDate": "12/02/2024 23:30:58",
      "content": "<p>You might look at separating the training and inference into separate notebooks.  That would let you both train much longer and have a more complicated inference process.  (As well as fix the submission variance.)</p>",
      "rawMarkdown": "You might look at separating the training and inference into separate notebooks.  That would let you both train much longer and have a more complicated inference process.  (As well as fix the submission variance.)",
      "votes": null
    },
    {
      "id": "3061786",
      "postDate": "12/03/2024 01:43:41",
      "content": "<p>Thanks for clarifying! I get the part about local minima, I thought the training wasn’t repeated during the submission process. <br>\nI tested separating the training and loading my weights in the submission notebook, and now the submissions are scoring the same! Appreciate your suggestions 😀</p>",
      "rawMarkdown": "Thanks for clarifying! I get the part about local minima, I thought the training wasn’t repeated during the submission process. \nI tested separating the training and loading my weights in the submission notebook, and now the submissions are scoring the same! Appreciate your suggestions 😀",
      "votes": null
    },
    {
      "id": "3061804",
      "postDate": "12/03/2024 02:05:28",
      "content": "<p>The difference was definitely a lot!! For this submission, I was applying a lot of augmentations from monai:  </p>\n<pre><code>random_transforms = Compose([\n    RandCropByLabelClassesd(\n        keys=[, ],\n        label_key=,\n        spatial_size=[Config.PATCH_SIZE] * ,\n        num_classes=Config.NUM_CLASSES,\n        num_samples=\n    ),\n    RandRotate90d(keys=[, ], prob=, spatial_axes=[, ]),\n    RandFlipd(keys=[, ], prob=, spatial_axis=),\n    RandScaleIntensityd(keys=[], factors=, prob=),\n    RandAffined(keys=[, ], prob=, rotate_range=(, , ), scale_range=(, , )),\n    RandZoomd(keys=[, ], prob=, min_zoom=, max_zoom=),\n    RandGaussianNoised(keys=[], prob=, mean=, std=),\n    RandGaussianSmoothd(keys=[], prob=, sigma_x=(, ))\n])\n</code></pre>\n<p>My private score for this model was only 0.436,  I tested it on just one tomogram. But when the model was retrained during submission, it somehow ended up as a better performing model.  </p>\n<p>For my other submissions (without as many augmentations), the private scores have been more consistent, 0.6 to 0.7, with low 0.6s on the leaderboard.  </p>\n<p>I still need to experiment further about how these augmentations are affecting performance and see if I can consistently reproduce the 0.634 score.  </p>",
      "rawMarkdown": "The difference was definitely a lot!! For this submission, I was applying a lot of augmentations from monai:  \n\n```python\nrandom_transforms = Compose([\n    RandCropByLabelClassesd(\n        keys=[\"image\", \"label\"],\n        label_key=\"label\",\n        spatial_size=[Config.PATCH_SIZE] * 3,\n        num_classes=Config.NUM_CLASSES,\n        num_samples=16\n    ),\n    RandRotate90d(keys=[\"image\", \"label\"], prob=0.5, spatial_axes=[0, 2]),\n    RandFlipd(keys=[\"image\", \"label\"], prob=0.5, spatial_axis=0),\n    RandScaleIntensityd(keys=[\"image\"], factors=0.1, prob=0.5),\n    RandAffined(keys=[\"image\", \"label\"], prob=0.3, rotate_range=(0.1, 0.1, 0.1), scale_range=(0.1, 0.1, 0.1)),\n    RandZoomd(keys=[\"image\", \"label\"], prob=0.3, min_zoom=0.9, max_zoom=1.1),\n    RandGaussianNoised(keys=[\"image\"], prob=0.2, mean=0.0, std=0.1),\n    RandGaussianSmoothd(keys=[\"image\"], prob=0.2, sigma_x=(0.5, 1.5))\n])\n```\n\nMy private score for this model was only 0.436,  I tested it on just one tomogram. But when the model was retrained during submission, it somehow ended up as a better performing model.  \n\nFor my other submissions (without as many augmentations), the private scores have been more consistent, 0.6 to 0.7, with low 0.6s on the leaderboard.  \n\nI still need to experiment further about how these augmentations are affecting performance and see if I can consistently reproduce the 0.634 score.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3061193,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "12/02/2024 12:43:29",
      "content": "<p>No. And pretty sure is about your code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3061194,
      "author_name": "andreizamfir",
      "author_url": "",
      "post_date": "12/02/2024 12:45:26",
      "content": "<p>I experienced getting the same score when submitting the same notebook. Do you have training included in your notebook? As far as I'm aware, the system evaluates the same 26% subset each time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3061215,
          "author_name": "sersasj",
          "author_url": "",
          "post_date": "12/02/2024 13:07:48",
          "content": "<p>Yeah, the training of the model is in the same notebook as the submission. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3061234,
              "author_name": "andreizamfir",
              "author_url": "",
              "post_date": "12/02/2024 13:38:07",
              "content": "<p>Then that's why, as training using the same set-up can get stuck in different local minima and learn differently.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3061786,
                  "author_name": "sersasj",
                  "author_url": "",
                  "post_date": "12/03/2024 01:43:41",
                  "content": "<p>Thanks for clarifying! I get the part about local minima, I thought the training wasn’t repeated during the submission process. <br>\nI tested separating the training and loading my weights in the submission notebook, and now the submissions are scoring the same! Appreciate your suggestions 😀</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 3061680,
              "author_name": "davidlist",
              "author_url": "",
              "post_date": "12/02/2024 23:30:58",
              "content": "<p>You might look at separating the training and inference into separate notebooks.  That would let you both train much longer and have a more complicated inference process.  (As well as fix the submission variance.)</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3061454,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "12/02/2024 16:36:13",
      "content": "<p>\"If that’s the case, is it a common practice to set seeds to avoid this variability?\" <br>\nThe code is executed as you submit it upon hidden test set once. But if your code uses RNG sure, always set a seed.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3061677,
      "author_name": "davidlist",
      "author_url": "",
      "post_date": "12/02/2024 23:27:58",
      "content": "<p>Agree with the assessments below…but, wow!  0.634 and 0.477?  That's a big difference.  Awful lot of variance in the training process.  If your models are that different each time, I would make sure you're saving the weights for each submission.  Although, I suppose setting seed achieves a similar result.  Are your local CV scores showing the same variance?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3061804,
          "author_name": "sersasj",
          "author_url": "",
          "post_date": "12/03/2024 02:05:28",
          "content": "<p>The difference was definitely a lot!! For this submission, I was applying a lot of augmentations from monai:  </p>\n<pre><code>random_transforms = Compose([\n    RandCropByLabelClassesd(\n        keys=[, ],\n        label_key=,\n        spatial_size=[Config.PATCH_SIZE] * ,\n        num_classes=Config.NUM_CLASSES,\n        num_samples=\n    ),\n    RandRotate90d(keys=[, ], prob=, spatial_axes=[, ]),\n    RandFlipd(keys=[, ], prob=, spatial_axis=),\n    RandScaleIntensityd(keys=[], factors=, prob=),\n    RandAffined(keys=[, ], prob=, rotate_range=(, , ), scale_range=(, , )),\n    RandZoomd(keys=[, ], prob=, min_zoom=, max_zoom=),\n    RandGaussianNoised(keys=[], prob=, mean=, std=),\n    RandGaussianSmoothd(keys=[], prob=, sigma_x=(, ))\n])\n</code></pre>\n<p>My private score for this model was only 0.436,  I tested it on just one tomogram. But when the model was retrained during submission, it somehow ended up as a better performing model.  </p>\n<p>For my other submissions (without as many augmentations), the private scores have been more consistent, 0.6 to 0.7, with low 0.6s on the leaderboard.  </p>\n<p>I still need to experiment further about how these augmentations are affecting performance and see if I can consistently reproduce the 0.634 score.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3061149": "Hey everyone,\n\nI’m a bit confused. I submitted the exact same notebook version twice, expecting similar scores, but the results were really different:\n- First submission: 0.634\n- Second submission: 0.477\n\nI thought that  only the inference part is run on 26% of the hidden dataset. So, is the model being retrained again during this process?\nIf that’s the case, is it a common practice to set seeds to avoid this variability?\nOr maybe the system evaluates different 26% subsets of the dataset each time?\n\nHas anyone experienced something similar? Would appreciate any insights on this.\n\nThanks!",
    "3061193": "No. And pretty sure is about your code.",
    "3061194": "I experienced getting the same score when submitting the same notebook. Do you have training included in your notebook? As far as I'm aware, the system evaluates the same 26% subset each time.",
    "3061215": "Yeah, the training of the model is in the same notebook as the submission.",
    "3061234": "Then that's why, as training using the same set-up can get stuck in different local minima and learn differently.",
    "3061454": "\"If that’s the case, is it a common practice to set seeds to avoid this variability?\" \nThe code is executed as you submit it upon hidden test set once. But if your code uses RNG sure, always set a seed.",
    "3061677": "Agree with the assessments below...but, wow!  0.634 and 0.477?  That's a big difference.  Awful lot of variance in the training process.  If your models are that different each time, I would make sure you're saving the weights for each submission.  Although, I suppose setting seed achieves a similar result.  Are your local CV scores showing the same variance?",
    "3061680": "You might look at separating the training and inference into separate notebooks.  That would let you both train much longer and have a more complicated inference process.  (As well as fix the submission variance.)",
    "3061786": "Thanks for clarifying! I get the part about local minima, I thought the training wasn’t repeated during the submission process. \nI tested separating the training and loading my weights in the submission notebook, and now the submissions are scoring the same! Appreciate your suggestions 😀",
    "3061804": "The difference was definitely a lot!! For this submission, I was applying a lot of augmentations from monai:  \n\n```python\nrandom_transforms = Compose([\n    RandCropByLabelClassesd(\n        keys=[\"image\", \"label\"],\n        label_key=\"label\",\n        spatial_size=[Config.PATCH_SIZE] * 3,\n        num_classes=Config.NUM_CLASSES,\n        num_samples=16\n    ),\n    RandRotate90d(keys=[\"image\", \"label\"], prob=0.5, spatial_axes=[0, 2]),\n    RandFlipd(keys=[\"image\", \"label\"], prob=0.5, spatial_axis=0),\n    RandScaleIntensityd(keys=[\"image\"], factors=0.1, prob=0.5),\n    RandAffined(keys=[\"image\", \"label\"], prob=0.3, rotate_range=(0.1, 0.1, 0.1), scale_range=(0.1, 0.1, 0.1)),\n    RandZoomd(keys=[\"image\", \"label\"], prob=0.3, min_zoom=0.9, max_zoom=1.1),\n    RandGaussianNoised(keys=[\"image\"], prob=0.2, mean=0.0, std=0.1),\n    RandGaussianSmoothd(keys=[\"image\"], prob=0.2, sigma_x=(0.5, 1.5))\n])\n```\n\nMy private score for this model was only 0.436,  I tested it on just one tomogram. But when the model was retrained during submission, it somehow ended up as a better performing model.  \n\nFor my other submissions (without as many augmentations), the private scores have been more consistent, 0.6 to 0.7, with low 0.6s on the leaderboard.  \n\nI still need to experiment further about how these augmentations are affecting performance and see if I can consistently reproduce the 0.634 score."
  },
  "source": "meta"
}