{
  "id": 455716,
  "title": "Will the test set have a significant third dimension, gaps?",
  "url": "/competitions/blood-vessel-segmentation/discussion/455716",
  "author_name": "",
  "post_date": "2023-11-16T04:22:28.721704300Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Will the test set have a reliable third dimension?  In other words - the training set consists of a set of 2D images in sequence, and that sequence can be considered the third dimension. Also in the training set the sequence length is lengthy - 500 to 2200. And finally - in the training set there are no gaps in the sequence. So the third dimension is straightforward and significant.</p>\n<p>But the example test set is very different:<br>\na) The sequence length is 3 for both kidney_5 and kidney_6. Very different from the training set.<br>\nb) There are no gaps in the sequence, but is something we can count on?</p>\n<p>So my questions are:<br>\nc) Will the test set consist of sequences of images of a substantial length?<br>\nd) Can there be gaps in the sequence?</p>\n<p>I sort of was assuming the test set would have a significant third dimension without gaps. But now I'm wondering if that is the case.</p>\n<p>Thank you to the organizers!</p>",
  "messages": [
    {
      "id": "2526806",
      "postDate": "11/16/2023 04:22:28",
      "content": "<p>Will the test set have a reliable third dimension?  In other words - the training set consists of a set of 2D images in sequence, and that sequence can be considered the third dimension. Also in the training set the sequence length is lengthy - 500 to 2200. And finally - in the training set there are no gaps in the sequence. So the third dimension is straightforward and significant.</p>\n<p>But the example test set is very different:<br>\na) The sequence length is 3 for both kidney_5 and kidney_6. Very different from the training set.<br>\nb) There are no gaps in the sequence, but is something we can count on?</p>\n<p>So my questions are:<br>\nc) Will the test set consist of sequences of images of a substantial length?<br>\nd) Can there be gaps in the sequence?</p>\n<p>I sort of was assuming the test set would have a significant third dimension without gaps. But now I'm wondering if that is the case.</p>\n<p>Thank you to the organizers!</p>",
      "rawMarkdown": "Will the test set have a reliable third dimension?  In other words - the training set consists of a set of 2D images in sequence, and that sequence can be considered the third dimension. Also in the training set the sequence length is lengthy - 500 to 2200. And finally - in the training set there are no gaps in the sequence. So the third dimension is straightforward and significant.\n\nBut the example test set is very different:\na) The sequence length is 3 for both kidney_5 and kidney_6. Very different from the training set.\nb) There are no gaps in the sequence, but is something we can count on?\n\nSo my questions are:\nc) Will the test set consist of sequences of images of a substantial length?\nd) Can there be gaps in the sequence?\n\nI sort of was assuming the test set would have a significant third dimension without gaps. But now I'm wondering if that is the case.\n\nThank you to the organizers!",
      "votes": null
    },
    {
      "id": "2528351",
      "postDate": "11/17/2023 09:44:01",
      "content": "<p>You're right, this is very important. Having only a few slices precludes doing a 3D convolution approach, which might be natural for a meaningful 3D \"block\" of voxels, or any other methods that work in 3D. Working only on isolated 2D slices throws away the correlation between slices in the z direction. It goes to the heart of what kind of image sets this is supposed to be useful on: if it's supposed to work on individual 2D slices, that's fine. But if in real life it would work on a 3D set, then it will be perhaps severely limiting performance to make it work on individual 2D slices.</p>",
      "rawMarkdown": "You're right, this is very important. Having only a few slices precludes doing a 3D convolution approach, which might be natural for a meaningful 3D \"block\" of voxels, or any other methods that work in 3D. Working only on isolated 2D slices throws away the correlation between slices in the z direction. It goes to the heart of what kind of image sets this is supposed to be useful on: if it's supposed to work on individual 2D slices, that's fine. But if in real life it would work on a 3D set, then it will be perhaps severely limiting performance to make it work on individual 2D slices.",
      "votes": null
    },
    {
      "id": "2528740",
      "postDate": "11/17/2023 15:18:43",
      "content": "<p>I read this in the data tab of the competition:</p>\n<p>\"These example images are not meant to be representative of the test set with regard to scan resolution, beamline, or other such qualities. When your submission is scored, this example test data will be replaced with the full test set. The full test set contains about 1500 TIFF images\"</p>\n<p>And because there is two kidney sets (5 and 6), maybe some of the data will have significant 3rd dimension relationship, but another question would be is it half and half, Does it stop at 750 and then go to the next set? </p>\n<p>It could be beneficial to include 3rd dimensional relationships in the prediction process, but if we don't know where it starts and stops that is tricky.</p>",
      "rawMarkdown": "I read this in the data tab of the competition:\n\n\"These example images are not meant to be representative of the test set with regard to scan resolution, beamline, or other such qualities. When your submission is scored, this example test data will be replaced with the full test set. The full test set contains about 1500 TIFF images\"\n\nAnd because there is two kidney sets (5 and 6), maybe some of the data will have significant 3rd dimension relationship, but another question would be is it half and half, Does it stop at 750 and then go to the next set? \n\nIt could be beneficial to include 3rd dimensional relationships in the prediction process, but if we don't know where it starts and stops that is tricky.",
      "votes": null
    },
    {
      "id": "2529528",
      "postDate": "11/18/2023 10:44:11",
      "content": "<p>Yes, just a minimum depth would be a useful thing to know. As far as I can see, the comment about 1500 images doesn't give any hint of how many kidneys that is divided between… Providing only 3 test images seems to presuppose that we are not going to try and make good use of the 3rd dimension, but I feel that should surely give better results, so long as the data is actually available.</p>",
      "rawMarkdown": "Yes, just a minimum depth would be a useful thing to know. As far as I can see, the comment about 1500 images doesn't give any hint of how many kidneys that is divided between... Providing only 3 test images seems to presuppose that we are not going to try and make good use of the 3rd dimension, but I feel that should surely give better results, so long as the data is actually available.",
      "votes": null
    },
    {
      "id": "2532245",
      "postDate": "11/20/2023 21:27:34",
      "content": "<p>I've never properly submitted to a Kaggle competition before, so the answer to this may be obvious… but I tried submitting a notebook purely so that it would output the paths and sizes of the hidden test set, to understand if they will be properly 3D. But even run as a submission, it listed the same 6 files as I would get by running the notebook for myself, either interactively as a save/commit. In which case how could a \"real\" submission work anyway?</p>",
      "rawMarkdown": "I've never properly submitted to a Kaggle competition before, so the answer to this may be obvious... but I tried submitting a notebook purely so that it would output the paths and sizes of the hidden test set, to understand if they will be properly 3D. But even run as a submission, it listed the same 6 files as I would get by running the notebook for myself, either interactively as a save/commit. In which case how could a \"real\" submission work anyway?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2528351,
      "author_name": "charliewartnaby",
      "author_url": "",
      "post_date": "11/17/2023 09:44:01",
      "content": "<p>You're right, this is very important. Having only a few slices precludes doing a 3D convolution approach, which might be natural for a meaningful 3D \"block\" of voxels, or any other methods that work in 3D. Working only on isolated 2D slices throws away the correlation between slices in the z direction. It goes to the heart of what kind of image sets this is supposed to be useful on: if it's supposed to work on individual 2D slices, that's fine. But if in real life it would work on a 3D set, then it will be perhaps severely limiting performance to make it work on individual 2D slices.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2528740,
      "author_name": "peterwarren",
      "author_url": "",
      "post_date": "11/17/2023 15:18:43",
      "content": "<p>I read this in the data tab of the competition:</p>\n<p>\"These example images are not meant to be representative of the test set with regard to scan resolution, beamline, or other such qualities. When your submission is scored, this example test data will be replaced with the full test set. The full test set contains about 1500 TIFF images\"</p>\n<p>And because there is two kidney sets (5 and 6), maybe some of the data will have significant 3rd dimension relationship, but another question would be is it half and half, Does it stop at 750 and then go to the next set? </p>\n<p>It could be beneficial to include 3rd dimensional relationships in the prediction process, but if we don't know where it starts and stops that is tricky.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2529528,
          "author_name": "charliewartnaby",
          "author_url": "",
          "post_date": "11/18/2023 10:44:11",
          "content": "<p>Yes, just a minimum depth would be a useful thing to know. As far as I can see, the comment about 1500 images doesn't give any hint of how many kidneys that is divided between… Providing only 3 test images seems to presuppose that we are not going to try and make good use of the 3rd dimension, but I feel that should surely give better results, so long as the data is actually available.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2532245,
          "author_name": "charliewartnaby",
          "author_url": "",
          "post_date": "11/20/2023 21:27:34",
          "content": "<p>I've never properly submitted to a Kaggle competition before, so the answer to this may be obvious… but I tried submitting a notebook purely so that it would output the paths and sizes of the hidden test set, to understand if they will be properly 3D. But even run as a submission, it listed the same 6 files as I would get by running the notebook for myself, either interactively as a save/commit. In which case how could a \"real\" submission work anyway?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2526806": "Will the test set have a reliable third dimension?  In other words - the training set consists of a set of 2D images in sequence, and that sequence can be considered the third dimension. Also in the training set the sequence length is lengthy - 500 to 2200. And finally - in the training set there are no gaps in the sequence. So the third dimension is straightforward and significant.\n\nBut the example test set is very different:\na) The sequence length is 3 for both kidney_5 and kidney_6. Very different from the training set.\nb) There are no gaps in the sequence, but is something we can count on?\n\nSo my questions are:\nc) Will the test set consist of sequences of images of a substantial length?\nd) Can there be gaps in the sequence?\n\nI sort of was assuming the test set would have a significant third dimension without gaps. But now I'm wondering if that is the case.\n\nThank you to the organizers!",
    "2528351": "You're right, this is very important. Having only a few slices precludes doing a 3D convolution approach, which might be natural for a meaningful 3D \"block\" of voxels, or any other methods that work in 3D. Working only on isolated 2D slices throws away the correlation between slices in the z direction. It goes to the heart of what kind of image sets this is supposed to be useful on: if it's supposed to work on individual 2D slices, that's fine. But if in real life it would work on a 3D set, then it will be perhaps severely limiting performance to make it work on individual 2D slices.",
    "2528740": "I read this in the data tab of the competition:\n\n\"These example images are not meant to be representative of the test set with regard to scan resolution, beamline, or other such qualities. When your submission is scored, this example test data will be replaced with the full test set. The full test set contains about 1500 TIFF images\"\n\nAnd because there is two kidney sets (5 and 6), maybe some of the data will have significant 3rd dimension relationship, but another question would be is it half and half, Does it stop at 750 and then go to the next set? \n\nIt could be beneficial to include 3rd dimensional relationships in the prediction process, but if we don't know where it starts and stops that is tricky.",
    "2529528": "Yes, just a minimum depth would be a useful thing to know. As far as I can see, the comment about 1500 images doesn't give any hint of how many kidneys that is divided between... Providing only 3 test images seems to presuppose that we are not going to try and make good use of the 3rd dimension, but I feel that should surely give better results, so long as the data is actually available.",
    "2532245": "I've never properly submitted to a Kaggle competition before, so the answer to this may be obvious... but I tried submitting a notebook purely so that it would output the paths and sizes of the hidden test set, to understand if they will be properly 3D. But even run as a submission, it listed the same 6 files as I would get by running the notebook for myself, either interactively as a save/commit. In which case how could a \"real\" submission work anyway?"
  },
  "source": "meta"
}