{
  "id": 571196,
  "title": "Validation of models?",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/571196",
  "author_name": "",
  "post_date": "2025-04-01T20:13:51.249692600Z",
  "votes": 32,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Has anyone been able to find a reasonable way to validate their models/methods at this time. I dont often like to use the public LB as a method of validation, but for this comp I am finding it fairly difficult to validate with the huge gap between CV/LB scores. I also do wonder what the benefit of this is for competition hosts. I don't doubt their methods necessarily but more just curious. As I am sure data of this type at this scale can be a painful project! </p>",
  "messages": [
    {
      "id": "3167773",
      "postDate": "04/01/2025 20:13:51",
      "content": "<p>Has anyone been able to find a reasonable way to validate their models/methods at this time. I dont often like to use the public LB as a method of validation, but for this comp I am finding it fairly difficult to validate with the huge gap between CV/LB scores. I also do wonder what the benefit of this is for competition hosts. I don't doubt their methods necessarily but more just curious. As I am sure data of this type at this scale can be a painful project! </p>",
      "rawMarkdown": "Has anyone been able to find a reasonable way to validate their models/methods at this time. I dont often like to use the public LB as a method of validation, but for this comp I am finding it fairly difficult to validate with the huge gap between CV/LB scores. I also do wonder what the benefit of this is for competition hosts. I don't doubt their methods necessarily but more just curious. As I am sure data of this type at this scale can be a painful project!",
      "votes": null
    },
    {
      "id": "3167819",
      "postDate": "04/01/2025 21:19:13",
      "content": "<p>My understanding is that you're asking us why we structured the test set in a way that has caused disparity between CV/LB scores. This is a good question.</p>\n<p>To start, this is the information we have stated across the data page, the leaderboard, and the pinned forums: </p>\n<blockquote>\n  <p>the rerun test dataset contains approximately 900 tomograms. The test data only contain tomograms with one or zero motors.</p>\n</blockquote>\n<p>-</p>\n<blockquote>\n  <p>This leaderboard is calculated with approximately 30% of the test data. The final results will be based on the other 70%, so the final standings may be different.</p>\n</blockquote>\n<p>-</p>\n<blockquote>\n  <p>there is a population shift between the train and test set.</p>\n</blockquote>\n<p>-</p>\n<p>Firstly, the CV/LB disparity is mainly caused by the fact that there is little similarity between train and test tomograms. Our justification for this is that we want this competition to produce models that can generalize to the task of identifying motors on bacteria sampled through our microscope at BYU. Despite how restrictive this may sound, we have a diverse dataset because samplers do things differently. </p>\n<p>So to achieve this goal, we made sure that the train and test datasets would be different in the same way that our current data is different from the data we have yet to sample. If our test set were too similar to our train set, the competition would have little value since all the big brains the competition brings in to work on the problem would come up with a min-maxed solution for the metadata we provided them. That's about as useful as the U.S. Military's vehicle classification model in which both the train and test images had all tanks in the forest and all jeeps in deserts. </p>\n<p>Additionally, our public/private LB structure is included to punish overfitting on the test set. Scoring well on the public LB isn't enough to win because for all we know you're just finding the best random seed to optimize your public LB score. A private leaderboard causes those that base their submissions off metrics besides public LB to score better. </p>\n<p>I hope this helps clarify things! Let me know if you still have any questions. </p>",
      "rawMarkdown": "My understanding is that you're asking us why we structured the test set in a way that has caused disparity between CV/LB scores. This is a good question.\n\nTo start, this is the information we have stated across the data page, the leaderboard, and the pinned forums: \n\n>the rerun test dataset contains approximately 900 tomograms. The test data only contain tomograms with one or zero motors.\n\n-\n\n>This leaderboard is calculated with approximately 30% of the test data. The final results will be based on the other 70%, so the final standings may be different.\n\n-\n\n>there is a population shift between the train and test set.\n\n-\n\nFirstly, the CV/LB disparity is mainly caused by the fact that there is little similarity between train and test tomograms. Our justification for this is that we want this competition to produce models that can generalize to the task of identifying motors on bacteria sampled through our microscope at BYU. Despite how restrictive this may sound, we have a diverse dataset because samplers do things differently. \n\nSo to achieve this goal, we made sure that the train and test datasets would be different in the same way that our current data is different from the data we have yet to sample. If our test set were too similar to our train set, the competition would have little value since all the big brains the competition brings in to work on the problem would come up with a min-maxed solution for the metadata we provided them. That's about as useful as the U.S. Military's vehicle classification model in which both the train and test images had all tanks in the forest and all jeeps in deserts. \n\nAdditionally, our public/private LB structure is included to punish overfitting on the test set. Scoring well on the public LB isn't enough to win because for all we know you're just finding the best random seed to optimize your public LB score. A private leaderboard causes those that base their submissions off metrics besides public LB to score better. \n\nI hope this helps clarify things! Let me know if you still have any questions.",
      "votes": null
    },
    {
      "id": "3167884",
      "postDate": "04/01/2025 22:48:53",
      "content": "<p>Thank you for your reply!! The reasoning absolutely helps! I figured this was the logic we were employing, but it never hurts to doublecheck! We will have to get creative then! :)</p>",
      "rawMarkdown": "Thank you for your reply!! The reasoning absolutely helps! I figured this was the logic we were employing, but it never hurts to doublecheck! We will have to get creative then! :)",
      "votes": null
    },
    {
      "id": "3168084",
      "postDate": "04/02/2025 05:20:19",
      "content": "<blockquote>\n  <p>Additionally, our public/private LB structure is included to punish overfitting on the test set. Scoring well on the public LB isn't enough to win because for all we know you're just finding the best random seed to optimize your public LB score. A private leaderboard causes those that base their submissions off metrics besides public LB to score better.</p>\n</blockquote>\n<p>Usually a full test set is constructed, from which the public and private test set are selected at random. Are you indicating that this is not the case here?</p>\n<p>While I recognize that preventing overfitting on the public LB is important, if there is another population shift between public and private we can't really prepare for that in any sensible way by making our models more robust - it becomes more a matter of who guesses the nature of the population shift best.</p>",
      "rawMarkdown": ">Additionally, our public/private LB structure is included to punish overfitting on the test set. Scoring well on the public LB isn't enough to win because for all we know you're just finding the best random seed to optimize your public LB score. A private leaderboard causes those that base their submissions off metrics besides public LB to score better.\n\nUsually a full test set is constructed, from which the public and private test set are selected at random. Are you indicating that this is not the case here?\n\nWhile I recognize that preventing overfitting on the public LB is important, if there is another population shift between public and private we can't really prepare for that in any sensible way by making our models more robust - it becomes more a matter of who guesses the nature of the population shift best.",
      "votes": null
    },
    {
      "id": "3168088",
      "postDate": "04/02/2025 05:30:14",
      "content": "<p>I am not sure that is selected at random. I think he is suggesting there is distribution shift. There are ways to prepare for this sensibly but it would mean people would need to ignore the LB to some extent.</p>",
      "rawMarkdown": "I am not sure that is selected at random. I think he is suggesting there is distribution shift. There are ways to prepare for this sensibly but it would mean people would need to ignore the LB to some extent.",
      "votes": null
    },
    {
      "id": "3168091",
      "postDate": "04/02/2025 05:40:06",
      "content": "<p>I didn't mean to comment on the nature of the public/private LB structure. More on the reason for having a public/private split in the first place. However it seems like you already knew that so you can disregard what I said because it seems to have caused a bit of confusion haha</p>\n<p>We're trying to avoid sharing information about the exact nature of the test set to avoid any possible leakage. So for the moment I can't tell you the exact nature of the public/private split. However I can tell you that we did design our competition for the aim of rewarding generalizable models not guesswork. So the specific situation you are describing, where you need to guess the nature of the population shift, is something I don't think we should be concerned about. </p>",
      "rawMarkdown": "I didn't mean to comment on the nature of the public/private LB structure. More on the reason for having a public/private split in the first place. However it seems like you already knew that so you can disregard what I said because it seems to have caused a bit of confusion haha\n\nWe're trying to avoid sharing information about the exact nature of the test set to avoid any possible leakage. So for the moment I can't tell you the exact nature of the public/private split. However I can tell you that we did design our competition for the aim of rewarding generalizable models not guesswork. So the specific situation you are describing, where you need to guess the nature of the population shift, is something I don't think we should be concerned about.",
      "votes": null
    },
    {
      "id": "3168275",
      "postDate": "04/02/2025 09:43:46",
      "content": "<p>I've used some submissions to understand the test data's image and volume sizes. I observed that more than 50% of test images have dimensions larger than 1100 (either in X or Y).</p>\n<p>One challenge is that we don't have access to the actual voxel scale—though I'm wondering if it could be predicted somehow. In the training data, largest images have lower voxel spacing, meaning that after resizing, we don’t lose much detail. However, what if an image in test is both large and has high voxel spacing, resizing could cause significant detail loss.</p>\n<p>An alternative is to work with patches, but this approach might increase submission times, which can only be properly evaluated through actual test submissions :( </p>",
      "rawMarkdown": "I've used some submissions to understand the test data's image and volume sizes. I observed that more than 50% of test images have dimensions larger than 1100 (either in X or Y).\n\nOne challenge is that we don't have access to the actual voxel scale—though I'm wondering if it could be predicted somehow. In the training data, largest images have lower voxel spacing, meaning that after resizing, we don’t lose much detail. However, what if an image in test is both large and has high voxel spacing, resizing could cause significant detail loss.\n\nAn alternative is to work with patches, but this approach might increase submission times, which can only be properly evaluated through actual test submissions :(",
      "votes": null
    },
    {
      "id": "3168329",
      "postDate": "04/02/2025 10:32:05",
      "content": "<p>This is aligned with my experiments. You can probe even more, but beware of overfitting to public lb. We need to be creative!</p>",
      "rawMarkdown": "This is aligned with my experiments. You can probe even more, but beware of overfitting to public lb. We need to be creative!",
      "votes": null
    },
    {
      "id": "3168359",
      "postDate": "04/02/2025 10:59:50",
      "content": "<p>True! I'm not sure if the hosts appreciate such testing, but this data drift in the test set makes it more beneficial to adjust the solution/architecture based on the public LB rather than relying on CV. </p>",
      "rawMarkdown": "True! I'm not sure if the hosts appreciate such testing, but this data drift in the test set makes it more beneficial to adjust the solution/architecture based on the public LB rather than relying on CV.",
      "votes": null
    },
    {
      "id": "3168417",
      "postDate": "04/02/2025 12:10:02",
      "content": "<p>I thought of it as an exercise of reducing variables so a pertinent set of ‘Why?’ questions may arise, questions whose answers lead to what the host seeks.</p>",
      "rawMarkdown": "I thought of it as an exercise of reducing variables so a pertinent set of ‘Why?’ questions may arise, questions whose answers lead to what the host seeks.",
      "votes": null
    },
    {
      "id": "3168557",
      "postDate": "04/02/2025 14:59:33",
      "content": "<p>I talked with the team. We decided that we would clarify the public private issue. </p>\n<p>There is no intentional population shift between the public and private test set. It was generated using a random split. Any differences in distribution were due to unintentional variation. <br>\n<a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> </p>\n<p>Meaning that the experiments you are discussing are reasonably accurate as they are a random sample of the test set in its entirety<br>\n<a href=\"https://www.kaggle.com/andreizamfir\" target=\"_blank\">@andreizamfir</a>  <a href=\"https://www.kaggle.com/carlosperez97\" target=\"_blank\">@carlosperez97</a> </p>\n<p>Good luck! I hope this extra information helps!</p>",
      "rawMarkdown": "I talked with the team. We decided that we would clarify the public private issue. \n\nThere is no intentional population shift between the public and private test set. It was generated using a random split. Any differences in distribution were due to unintentional variation. \n@cody11null @jeroencottaar \n\nMeaning that the experiments you are discussing are reasonably accurate as they are a random sample of the test set in its entirety\n@andreizamfir  @carlosperez97 \n\nGood luck! I hope this extra information helps!",
      "votes": null
    },
    {
      "id": "3168593",
      "postDate": "04/02/2025 15:27:48",
      "content": "<p>Thanks for the clarification.</p>",
      "rawMarkdown": "Thanks for the clarification.",
      "votes": null
    },
    {
      "id": "3170144",
      "postDate": "04/04/2025 11:17:37",
      "content": "<p>I've found that 90% of tomograms in test have images with dimensions smaller than 1500 (either in X or Y, both dimensions are checked) and that 90% of the tomograms <strong>do not</strong> have images with dimensions smaller than 1300 (either in X or Y). We don't know which % will correspond to public and private but we have to keep this mind. </p>",
      "rawMarkdown": "I've found that 90% of tomograms in test have images with dimensions smaller than 1500 (either in X or Y, both dimensions are checked) and that 90% of the tomograms **do not** have images with dimensions smaller than 1300 (either in X or Y). We don't know which % will correspond to public and private but we have to keep this mind.",
      "votes": null
    },
    {
      "id": "3170749",
      "postDate": "04/05/2025 01:05:16",
      "content": "<p>Hi,<br>\nThanks for these informations.</p>\n<p>I have some questions: Do we know the voxel spacing in test set ? are shape and vxs constant across the test set ? what are theirs values ?</p>",
      "rawMarkdown": "Hi,\nThanks for these informations.\n\nI have some questions: Do we know the voxel spacing in test set ? are shape and vxs constant across the test set ? what are theirs values ?",
      "votes": null
    },
    {
      "id": "3171038",
      "postDate": "04/05/2025 08:12:15",
      "content": "<p>I have tried to work with patches, and even shared my work [<a href=\"https://www.kaggle.com/code/fautei/byu-yolo-sahi-submission-notebook]\" target=\"_blank\">https://www.kaggle.com/code/fautei/byu-yolo-sahi-submission-notebook]</a>, but results not improved compared to whole image. Can we make some suggestion about voxel spacing of test set compared to train based on this results?</p>",
      "rawMarkdown": "I have tried to work with patches, and even shared my work [https://www.kaggle.com/code/fautei/byu-yolo-sahi-submission-notebook], but results not improved compared to whole image. Can we make some suggestion about voxel spacing of test set compared to train based on this results?",
      "votes": null
    },
    {
      "id": "3171167",
      "postDate": "04/05/2025 11:09:43",
      "content": "<p>I have also attempted overlapping patches, and indeed, I barely found settings which maintained the score without patches. Afterall, it's a combination of both shape and voxel spacing which provides the scale of objects within the image (considering YOLO letterboxing to standard dimensions).</p>",
      "rawMarkdown": "I have also attempted overlapping patches, and indeed, I barely found settings which maintained the score without patches. Afterall, it's a combination of both shape and voxel spacing which provides the scale of objects within the image (considering YOLO letterboxing to standard dimensions).",
      "votes": null
    },
    {
      "id": "3177473",
      "postDate": "04/12/2025 20:00:08",
      "content": "<p>How are you able to figure that out? I only see the output of the submission notebook on the face test set with 3 tomograms. Is there a way to output log information from the submission running on the rerun test data (obviously without somehow cheating)? (I'm referring to your comment <a href=\"https://www.kaggle.com/carlosperez97\" target=\"_blank\">@carlosperez97</a> about the % of images with certain sizes)</p>",
      "rawMarkdown": "How are you able to figure that out? I only see the output of the submission notebook on the face test set with 3 tomograms. Is there a way to output log information from the submission running on the rerun test data (obviously without somehow cheating)? (I'm referring to your comment @carlosperez97 about the % of images with certain sizes)",
      "votes": null
    },
    {
      "id": "3177491",
      "postDate": "04/12/2025 21:02:59",
      "content": "<p>Returning -1, -1, -1 for all tomos gives you 0.000 lb. If you make conditions that throws an error, you get an error for your submission instead of a score. You can probe certain conditions happening or not like that.<br>\nSay you take the shape of the image, if they’re both over 1600, cause a break of the loop that causes an error, then you know there are images where that condition happens. If it all runs smoothly and you get 0.000, no images in the test set satisfy the condition.</p>",
      "rawMarkdown": "Returning -1, -1, -1 for all tomos gives you 0.000 lb. If you make conditions that throws an error, you get an error for your submission instead of a score. You can probe certain conditions happening or not like that.\nSay you take the shape of the image, if they’re both over 1600, cause a break of the loop that causes an error, then you know there are images where that condition happens. If it all runs smoothly and you get 0.000, no images in the test set satisfy the condition.",
      "votes": null
    },
    {
      "id": "3177499",
      "postDate": "04/12/2025 21:15:48",
      "content": "<p>That’s perfect, thanks a ton <a href=\"https://www.kaggle.com/andreizamfir\" target=\"_blank\">@andreizamfir</a>!</p>",
      "rawMarkdown": "That’s perfect, thanks a ton @andreizamfir!",
      "votes": null
    },
    {
      "id": "3177503",
      "postDate": "04/12/2025 21:28:27",
      "content": "<p>Yes, exactly as Andrei mentions. I added some assert and if one of them fails then I will see a Notebook failed execution. <a href=\"https://www.kaggle.com/bklynch\" target=\"_blank\">@bklynch</a> </p>",
      "rawMarkdown": "Yes, exactly as Andrei mentions. I added some assert and if one of them fails then I will see a Notebook failed execution. @bklynch",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3167819,
      "author_name": "andrewjdarley",
      "author_url": "",
      "post_date": "04/01/2025 21:19:13",
      "content": "<p>My understanding is that you're asking us why we structured the test set in a way that has caused disparity between CV/LB scores. This is a good question.</p>\n<p>To start, this is the information we have stated across the data page, the leaderboard, and the pinned forums: </p>\n<blockquote>\n  <p>the rerun test dataset contains approximately 900 tomograms. The test data only contain tomograms with one or zero motors.</p>\n</blockquote>\n<p>-</p>\n<blockquote>\n  <p>This leaderboard is calculated with approximately 30% of the test data. The final results will be based on the other 70%, so the final standings may be different.</p>\n</blockquote>\n<p>-</p>\n<blockquote>\n  <p>there is a population shift between the train and test set.</p>\n</blockquote>\n<p>-</p>\n<p>Firstly, the CV/LB disparity is mainly caused by the fact that there is little similarity between train and test tomograms. Our justification for this is that we want this competition to produce models that can generalize to the task of identifying motors on bacteria sampled through our microscope at BYU. Despite how restrictive this may sound, we have a diverse dataset because samplers do things differently. </p>\n<p>So to achieve this goal, we made sure that the train and test datasets would be different in the same way that our current data is different from the data we have yet to sample. If our test set were too similar to our train set, the competition would have little value since all the big brains the competition brings in to work on the problem would come up with a min-maxed solution for the metadata we provided them. That's about as useful as the U.S. Military's vehicle classification model in which both the train and test images had all tanks in the forest and all jeeps in deserts. </p>\n<p>Additionally, our public/private LB structure is included to punish overfitting on the test set. Scoring well on the public LB isn't enough to win because for all we know you're just finding the best random seed to optimize your public LB score. A private leaderboard causes those that base their submissions off metrics besides public LB to score better. </p>\n<p>I hope this helps clarify things! Let me know if you still have any questions. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3167884,
          "author_name": "cody11null",
          "author_url": "",
          "post_date": "04/01/2025 22:48:53",
          "content": "<p>Thank you for your reply!! The reasoning absolutely helps! I figured this was the logic we were employing, but it never hurts to doublecheck! We will have to get creative then! :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3168084,
          "author_name": "jeroencottaar",
          "author_url": "",
          "post_date": "04/02/2025 05:20:19",
          "content": "<blockquote>\n  <p>Additionally, our public/private LB structure is included to punish overfitting on the test set. Scoring well on the public LB isn't enough to win because for all we know you're just finding the best random seed to optimize your public LB score. A private leaderboard causes those that base their submissions off metrics besides public LB to score better.</p>\n</blockquote>\n<p>Usually a full test set is constructed, from which the public and private test set are selected at random. Are you indicating that this is not the case here?</p>\n<p>While I recognize that preventing overfitting on the public LB is important, if there is another population shift between public and private we can't really prepare for that in any sensible way by making our models more robust - it becomes more a matter of who guesses the nature of the population shift best.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3168088,
              "author_name": "cody11null",
              "author_url": "",
              "post_date": "04/02/2025 05:30:14",
              "content": "<p>I am not sure that is selected at random. I think he is suggesting there is distribution shift. There are ways to prepare for this sensibly but it would mean people would need to ignore the LB to some extent.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3168275,
                  "author_name": "carlosperez97",
                  "author_url": "",
                  "post_date": "04/02/2025 09:43:46",
                  "content": "<p>I've used some submissions to understand the test data's image and volume sizes. I observed that more than 50% of test images have dimensions larger than 1100 (either in X or Y).</p>\n<p>One challenge is that we don't have access to the actual voxel scale—though I'm wondering if it could be predicted somehow. In the training data, largest images have lower voxel spacing, meaning that after resizing, we don’t lose much detail. However, what if an image in test is both large and has high voxel spacing, resizing could cause significant detail loss.</p>\n<p>An alternative is to work with patches, but this approach might increase submission times, which can only be properly evaluated through actual test submissions :( </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3168329,
                      "author_name": "andreizamfir",
                      "author_url": "",
                      "post_date": "04/02/2025 10:32:05",
                      "content": "<p>This is aligned with my experiments. You can probe even more, but beware of overfitting to public lb. We need to be creative!</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3168359,
                          "author_name": "carlosperez97",
                          "author_url": "",
                          "post_date": "04/02/2025 10:59:50",
                          "content": "<p>True! I'm not sure if the hosts appreciate such testing, but this data drift in the test set makes it more beneficial to adjust the solution/architecture based on the public LB rather than relying on CV. </p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3168417,
                              "author_name": "andreizamfir",
                              "author_url": "",
                              "post_date": "04/02/2025 12:10:02",
                              "content": "<p>I thought of it as an exercise of reducing variables so a pertinent set of ‘Why?’ questions may arise, questions whose answers lead to what the host seeks.</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 3170144,
                                  "author_name": "carlosperez97",
                                  "author_url": "",
                                  "post_date": "04/04/2025 11:17:37",
                                  "content": "<p>I've found that 90% of tomograms in test have images with dimensions smaller than 1500 (either in X or Y, both dimensions are checked) and that 90% of the tomograms <strong>do not</strong> have images with dimensions smaller than 1300 (either in X or Y). We don't know which % will correspond to public and private but we have to keep this mind. </p>",
                                  "votes": null,
                                  "replies": [
                                    {
                                      "id": 3177473,
                                      "author_name": "bklynch",
                                      "author_url": "",
                                      "post_date": "04/12/2025 20:00:08",
                                      "content": "<p>How are you able to figure that out? I only see the output of the submission notebook on the face test set with 3 tomograms. Is there a way to output log information from the submission running on the rerun test data (obviously without somehow cheating)? (I'm referring to your comment <a href=\"https://www.kaggle.com/carlosperez97\" target=\"_blank\">@carlosperez97</a> about the % of images with certain sizes)</p>",
                                      "votes": null,
                                      "replies": [
                                        {
                                          "id": 3177491,
                                          "author_name": "andreizamfir",
                                          "author_url": "",
                                          "post_date": "04/12/2025 21:02:59",
                                          "content": "<p>Returning -1, -1, -1 for all tomos gives you 0.000 lb. If you make conditions that throws an error, you get an error for your submission instead of a score. You can probe certain conditions happening or not like that.<br>\nSay you take the shape of the image, if they’re both over 1600, cause a break of the loop that causes an error, then you know there are images where that condition happens. If it all runs smoothly and you get 0.000, no images in the test set satisfy the condition.</p>",
                                          "votes": null,
                                          "replies": [
                                            {
                                              "id": 3177499,
                                              "author_name": "bklynch",
                                              "author_url": "",
                                              "post_date": "04/12/2025 21:15:48",
                                              "content": "<p>That’s perfect, thanks a ton <a href=\"https://www.kaggle.com/andreizamfir\" target=\"_blank\">@andreizamfir</a>!</p>",
                                              "votes": null,
                                              "replies": [
                                                {
                                                  "id": 3177503,
                                                  "author_name": "carlosperez97",
                                                  "author_url": "",
                                                  "post_date": "04/12/2025 21:28:27",
                                                  "content": "<p>Yes, exactly as Andrei mentions. I added some assert and if one of them fails then I will see a Notebook failed execution. <a href=\"https://www.kaggle.com/bklynch\" target=\"_blank\">@bklynch</a> </p>",
                                                  "votes": null,
                                                  "replies": []
                                                }
                                              ]
                                            }
                                          ]
                                        }
                                      ]
                                    }
                                  ]
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    },
                    {
                      "id": 3171038,
                      "author_name": "fautei",
                      "author_url": "",
                      "post_date": "04/05/2025 08:12:15",
                      "content": "<p>I have tried to work with patches, and even shared my work [<a href=\"https://www.kaggle.com/code/fautei/byu-yolo-sahi-submission-notebook]\" target=\"_blank\">https://www.kaggle.com/code/fautei/byu-yolo-sahi-submission-notebook]</a>, but results not improved compared to whole image. Can we make some suggestion about voxel spacing of test set compared to train based on this results?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3171167,
                          "author_name": "andreizamfir",
                          "author_url": "",
                          "post_date": "04/05/2025 11:09:43",
                          "content": "<p>I have also attempted overlapping patches, and indeed, I barely found settings which maintained the score without patches. Afterall, it's a combination of both shape and voxel spacing which provides the scale of objects within the image (considering YOLO letterboxing to standard dimensions).</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            },
            {
              "id": 3168091,
              "author_name": "andrewjdarley",
              "author_url": "",
              "post_date": "04/02/2025 05:40:06",
              "content": "<p>I didn't mean to comment on the nature of the public/private LB structure. More on the reason for having a public/private split in the first place. However it seems like you already knew that so you can disregard what I said because it seems to have caused a bit of confusion haha</p>\n<p>We're trying to avoid sharing information about the exact nature of the test set to avoid any possible leakage. So for the moment I can't tell you the exact nature of the public/private split. However I can tell you that we did design our competition for the aim of rewarding generalizable models not guesswork. So the specific situation you are describing, where you need to guess the nature of the population shift, is something I don't think we should be concerned about. </p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3170749,
          "author_name": "drxc75",
          "author_url": "",
          "post_date": "04/05/2025 01:05:16",
          "content": "<p>Hi,<br>\nThanks for these informations.</p>\n<p>I have some questions: Do we know the voxel spacing in test set ? are shape and vxs constant across the test set ? what are theirs values ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3168557,
      "author_name": "andrewjdarley",
      "author_url": "",
      "post_date": "04/02/2025 14:59:33",
      "content": "<p>I talked with the team. We decided that we would clarify the public private issue. </p>\n<p>There is no intentional population shift between the public and private test set. It was generated using a random split. Any differences in distribution were due to unintentional variation. <br>\n<a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> </p>\n<p>Meaning that the experiments you are discussing are reasonably accurate as they are a random sample of the test set in its entirety<br>\n<a href=\"https://www.kaggle.com/andreizamfir\" target=\"_blank\">@andreizamfir</a>  <a href=\"https://www.kaggle.com/carlosperez97\" target=\"_blank\">@carlosperez97</a> </p>\n<p>Good luck! I hope this extra information helps!</p>",
      "votes": null,
      "replies": [
        {
          "id": 3168593,
          "author_name": "jeroencottaar",
          "author_url": "",
          "post_date": "04/02/2025 15:27:48",
          "content": "<p>Thanks for the clarification.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3167773": "Has anyone been able to find a reasonable way to validate their models/methods at this time. I dont often like to use the public LB as a method of validation, but for this comp I am finding it fairly difficult to validate with the huge gap between CV/LB scores. I also do wonder what the benefit of this is for competition hosts. I don't doubt their methods necessarily but more just curious. As I am sure data of this type at this scale can be a painful project!",
    "3167819": "My understanding is that you're asking us why we structured the test set in a way that has caused disparity between CV/LB scores. This is a good question.\n\nTo start, this is the information we have stated across the data page, the leaderboard, and the pinned forums: \n\n>the rerun test dataset contains approximately 900 tomograms. The test data only contain tomograms with one or zero motors.\n\n-\n\n>This leaderboard is calculated with approximately 30% of the test data. The final results will be based on the other 70%, so the final standings may be different.\n\n-\n\n>there is a population shift between the train and test set.\n\n-\n\nFirstly, the CV/LB disparity is mainly caused by the fact that there is little similarity between train and test tomograms. Our justification for this is that we want this competition to produce models that can generalize to the task of identifying motors on bacteria sampled through our microscope at BYU. Despite how restrictive this may sound, we have a diverse dataset because samplers do things differently. \n\nSo to achieve this goal, we made sure that the train and test datasets would be different in the same way that our current data is different from the data we have yet to sample. If our test set were too similar to our train set, the competition would have little value since all the big brains the competition brings in to work on the problem would come up with a min-maxed solution for the metadata we provided them. That's about as useful as the U.S. Military's vehicle classification model in which both the train and test images had all tanks in the forest and all jeeps in deserts. \n\nAdditionally, our public/private LB structure is included to punish overfitting on the test set. Scoring well on the public LB isn't enough to win because for all we know you're just finding the best random seed to optimize your public LB score. A private leaderboard causes those that base their submissions off metrics besides public LB to score better. \n\nI hope this helps clarify things! Let me know if you still have any questions.",
    "3167884": "Thank you for your reply!! The reasoning absolutely helps! I figured this was the logic we were employing, but it never hurts to doublecheck! We will have to get creative then! :)",
    "3168084": ">Additionally, our public/private LB structure is included to punish overfitting on the test set. Scoring well on the public LB isn't enough to win because for all we know you're just finding the best random seed to optimize your public LB score. A private leaderboard causes those that base their submissions off metrics besides public LB to score better.\n\nUsually a full test set is constructed, from which the public and private test set are selected at random. Are you indicating that this is not the case here?\n\nWhile I recognize that preventing overfitting on the public LB is important, if there is another population shift between public and private we can't really prepare for that in any sensible way by making our models more robust - it becomes more a matter of who guesses the nature of the population shift best.",
    "3168088": "I am not sure that is selected at random. I think he is suggesting there is distribution shift. There are ways to prepare for this sensibly but it would mean people would need to ignore the LB to some extent.",
    "3168091": "I didn't mean to comment on the nature of the public/private LB structure. More on the reason for having a public/private split in the first place. However it seems like you already knew that so you can disregard what I said because it seems to have caused a bit of confusion haha\n\nWe're trying to avoid sharing information about the exact nature of the test set to avoid any possible leakage. So for the moment I can't tell you the exact nature of the public/private split. However I can tell you that we did design our competition for the aim of rewarding generalizable models not guesswork. So the specific situation you are describing, where you need to guess the nature of the population shift, is something I don't think we should be concerned about.",
    "3168275": "I've used some submissions to understand the test data's image and volume sizes. I observed that more than 50% of test images have dimensions larger than 1100 (either in X or Y).\n\nOne challenge is that we don't have access to the actual voxel scale—though I'm wondering if it could be predicted somehow. In the training data, largest images have lower voxel spacing, meaning that after resizing, we don’t lose much detail. However, what if an image in test is both large and has high voxel spacing, resizing could cause significant detail loss.\n\nAn alternative is to work with patches, but this approach might increase submission times, which can only be properly evaluated through actual test submissions :(",
    "3168329": "This is aligned with my experiments. You can probe even more, but beware of overfitting to public lb. We need to be creative!",
    "3168359": "True! I'm not sure if the hosts appreciate such testing, but this data drift in the test set makes it more beneficial to adjust the solution/architecture based on the public LB rather than relying on CV.",
    "3168417": "I thought of it as an exercise of reducing variables so a pertinent set of ‘Why?’ questions may arise, questions whose answers lead to what the host seeks.",
    "3168557": "I talked with the team. We decided that we would clarify the public private issue. \n\nThere is no intentional population shift between the public and private test set. It was generated using a random split. Any differences in distribution were due to unintentional variation. \n@cody11null @jeroencottaar \n\nMeaning that the experiments you are discussing are reasonably accurate as they are a random sample of the test set in its entirety\n@andreizamfir  @carlosperez97 \n\nGood luck! I hope this extra information helps!",
    "3168593": "Thanks for the clarification.",
    "3170144": "I've found that 90% of tomograms in test have images with dimensions smaller than 1500 (either in X or Y, both dimensions are checked) and that 90% of the tomograms **do not** have images with dimensions smaller than 1300 (either in X or Y). We don't know which % will correspond to public and private but we have to keep this mind.",
    "3170749": "Hi,\nThanks for these informations.\n\nI have some questions: Do we know the voxel spacing in test set ? are shape and vxs constant across the test set ? what are theirs values ?",
    "3171038": "I have tried to work with patches, and even shared my work [https://www.kaggle.com/code/fautei/byu-yolo-sahi-submission-notebook], but results not improved compared to whole image. Can we make some suggestion about voxel spacing of test set compared to train based on this results?",
    "3171167": "I have also attempted overlapping patches, and indeed, I barely found settings which maintained the score without patches. Afterall, it's a combination of both shape and voxel spacing which provides the scale of objects within the image (considering YOLO letterboxing to standard dimensions).",
    "3177473": "How are you able to figure that out? I only see the output of the submission notebook on the face test set with 3 tomograms. Is there a way to output log information from the submission running on the rerun test data (obviously without somehow cheating)? (I'm referring to your comment @carlosperez97 about the % of images with certain sizes)",
    "3177491": "Returning -1, -1, -1 for all tomos gives you 0.000 lb. If you make conditions that throws an error, you get an error for your submission instead of a score. You can probe certain conditions happening or not like that.\nSay you take the shape of the image, if they’re both over 1600, cause a break of the loop that causes an error, then you know there are images where that condition happens. If it all runs smoothly and you get 0.000, no images in the test set satisfy the condition.",
    "3177499": "That’s perfect, thanks a ton @andreizamfir!",
    "3177503": "Yes, exactly as Andrei mentions. I added some assert and if one of them fails then I will see a Notebook failed execution. @bklynch"
  },
  "source": "meta"
}