{
  "id": 456596,
  "title": "Full testing set. 3D or 2D?",
  "url": "/competitions/blood-vessel-segmentation/discussion/456596",
  "author_name": "",
  "post_date": "2023-11-20T20:08:58.654410700Z",
  "votes": 4,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I'm sure I'm not the only one with this question.<br>\nThe given testing set is very small and obviously not representative of the “real” testing set that will be used at the end of the competition. However, it is unclear if the real/full dataset (it is mentioned it will be about 1500 images) will be a few random 2D slices of 3D volumes (essentially making three-dimensional data very difficult to exploit in the dataset) or if we can expect to have a sequence of 2D slices that can be stacked into reasonable 3D volumes.</p>\n<p>This is important because it allows us to plan to exploit three-dimensional structures of the inputs or not. They are obviously in the training set, but it's not clear at all if they will be in the testing set.</p>\n<p>Thanks in advance.</p>",
  "messages": [
    {
      "id": "2532196",
      "postDate": "11/20/2023 20:08:58",
      "content": "<p>I'm sure I'm not the only one with this question.<br>\nThe given testing set is very small and obviously not representative of the “real” testing set that will be used at the end of the competition. However, it is unclear if the real/full dataset (it is mentioned it will be about 1500 images) will be a few random 2D slices of 3D volumes (essentially making three-dimensional data very difficult to exploit in the dataset) or if we can expect to have a sequence of 2D slices that can be stacked into reasonable 3D volumes.</p>\n<p>This is important because it allows us to plan to exploit three-dimensional structures of the inputs or not. They are obviously in the training set, but it's not clear at all if they will be in the testing set.</p>\n<p>Thanks in advance.</p>",
      "rawMarkdown": "I'm sure I'm not the only one with this question.\nThe given testing set is very small and obviously not representative of the “real” testing set that will be used at the end of the competition. However, it is unclear if the real/full dataset (it is mentioned it will be about 1500 images) will be a few random 2D slices of 3D volumes (essentially making three-dimensional data very difficult to exploit in the dataset) or if we can expect to have a sequence of 2D slices that can be stacked into reasonable 3D volumes.\n\nThis is important because it allows us to plan to exploit three-dimensional structures of the inputs or not. They are obviously in the training set, but it's not clear at all if they will be in the testing set.\n\nThanks in advance.",
      "votes": null
    },
    {
      "id": "2532249",
      "postDate": "11/20/2023 21:35:56",
      "content": "<p>This is also asked in <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455716\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455716</a> . I agree it is crucial to know this. We have no idea how many kidneys the ~1500 test images cover or what the minimum number of slices per kidney is. If it's too \"thin\", it totally precludes some 3D approaches where we might need a decent number of voxels in the z-direction.</p>\n<p>I tried to just list the test files by making a dummy submission of a noteboook which listed what files were in 'test', but it just gave me the same 6 sample images as what we see in the public dataset. I'm a newbie, so I don't really understand how the real test set is substituted in place of the example test set -- maybe I'm seeing the output of a \"dry run\" of my notebook, and I'm not allowed to see the output of the true competition run?</p>",
      "rawMarkdown": "This is also asked in https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455716 . I agree it is crucial to know this. We have no idea how many kidneys the ~1500 test images cover or what the minimum number of slices per kidney is. If it's too \"thin\", it totally precludes some 3D approaches where we might need a decent number of voxels in the z-direction.\n\nI tried to just list the test files by making a dummy submission of a noteboook which listed what files were in 'test', but it just gave me the same 6 sample images as what we see in the public dataset. I'm a newbie, so I don't really understand how the real test set is substituted in place of the example test set -- maybe I'm seeing the output of a \"dry run\" of my notebook, and I'm not allowed to see the output of the true competition run?",
      "votes": null
    },
    {
      "id": "2532251",
      "postDate": "11/20/2023 21:39:09",
      "content": "<p>Same question posted here. Added to make sure not missing any information. <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456047\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456047</a></p>",
      "rawMarkdown": "Same question posted here. Added to make sure not missing any information. https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456047",
      "votes": null
    },
    {
      "id": "2532264",
      "postDate": "11/20/2023 22:11:34",
      "content": "<p>Hi all, thanks for the questions regarding test data. The quick response is that the training data represent the test data well in terms of the 3D sizes. i.e the 3rd dimension will be usable and quite possibly important for good segmentations. Indeed the 3D nature of this imaging technique is one of its key features. I am prepping a fuller post regarding the test and training datasets in the next day or two we just wanted to see what the common needs/qs were before replying to everyone too quickly. </p>",
      "rawMarkdown": "Hi all, thanks for the questions regarding test data. The quick response is that the training data represent the test data well in terms of the 3D sizes. i.e the 3rd dimension will be usable and quite possibly important for good segmentations. Indeed the 3D nature of this imaging technique is one of its key features. I am prepping a fuller post regarding the test and training datasets in the next day or two we just wanted to see what the common needs/qs were before replying to everyone too quickly.",
      "votes": null
    },
    {
      "id": "2532292",
      "postDate": "11/20/2023 23:14:21",
      "content": "<p>that will be incredibly helpful. Thank you tons :)</p>",
      "rawMarkdown": "that will be incredibly helpful. Thank you tons :)",
      "votes": null
    },
    {
      "id": "2532481",
      "postDate": "11/21/2023 05:36:47",
      "content": "<p>i thought the title of the competition is \"Segment vasculature in 3D scans of human kidney\"</p>",
      "rawMarkdown": "i thought the title of the competition is \"Segment vasculature in 3D scans of human kidney\"",
      "votes": null
    },
    {
      "id": "2532563",
      "postDate": "11/21/2023 06:43:09",
      "content": "<p>Ah that's good to know, thanks. However, as a newbie I'm worried about whether it will affect the submission process. I get the impression that Kaggle first puts a submitted notebook through a dry run on the publicly viewable test data, and only if it succeeds then runs it again with the hidden test data. If (say) I wanted to run something that needed batches of 16 images in the z-direction, it would never work on the example test data, so my notebook wouldn't proceed to the scored submission run. Or maybe I'd just have to output some dummy data for the example data so that it <em>appeared</em> to work on that, and then do my real processing on the hidden data once it got that far… However, I don't clearly understand how notebook execution works for competition submissions, it doesn't seem to be explained in detail (not specifically to this competition but generally on Kaggle) so I may well be getting the wrong end of the stick. (If anyone can point me to a detailed explanation of how notebooks are run on competition test data I'd be glad, e.g. am I even allowed to see notebook output corresponding to the hidden data, maybe not?)</p>",
      "rawMarkdown": "Ah that's good to know, thanks. However, as a newbie I'm worried about whether it will affect the submission process. I get the impression that Kaggle first puts a submitted notebook through a dry run on the publicly viewable test data, and only if it succeeds then runs it again with the hidden test data. If (say) I wanted to run something that needed batches of 16 images in the z-direction, it would never work on the example test data, so my notebook wouldn't proceed to the scored submission run. Or maybe I'd just have to output some dummy data for the example data so that it _appeared_ to work on that, and then do my real processing on the hidden data once it got that far... However, I don't clearly understand how notebook execution works for competition submissions, it doesn't seem to be explained in detail (not specifically to this competition but generally on Kaggle) so I may well be getting the wrong end of the stick. (If anyone can point me to a detailed explanation of how notebooks are run on competition test data I'd be glad, e.g. am I even allowed to see notebook output corresponding to the hidden data, maybe not?)",
      "votes": null
    },
    {
      "id": "2532593",
      "postDate": "11/21/2023 07:06:00",
      "content": "<p>Use the train data folder during your dummy run and the test folder when hidden test data is present. Just check the number of files in the test folder to determine which stage you are at.</p>",
      "rawMarkdown": "Use the train data folder during your dummy run and the test folder when hidden test data is present. Just check the number of files in the test folder to determine which stage you are at.",
      "votes": null
    },
    {
      "id": "2532641",
      "postDate": "11/21/2023 07:57:02",
      "content": "<p>\" it would never work on the example test data, \"</p>\n<p>it should something like</p>\n<pre><code>=\\\n            #local   development\n  #    #submit\n\n =='local':\n    valid_folder =    train folder\n    valid_image_range = (0,64)  #first 64 images\n\n =='submit':\n    is_not_server = .   glob test kidney_5 folder images  check  there are only 3 images\n     is_not_server:\n         valid_folder =    train folder\n         valid_image_range = (0,64)  #first 64 images\n   : \n         valid_folder =    test folder\n         valid_image_range = . glob   number of images .\n</code></pre>",
      "rawMarkdown": "\" it would never work on the example test data, \"\n\nit should something like\n\n\n```\nmode=\\\n  'local'          #local debug for development\n  #'submit'    #submit\n\nif mode=='local':\n    valid_folder = .... set to train folder\n    valid_image_range = (0,64)  #first 64 images\n\nif mode=='submit':\n    is_not_server = ... true if glob test kidney_5 folder images and check if there are only 3 images\n    if is_not_server:\n         valid_folder = .... set to train folder\n         valid_image_range = (0,64)  #first 64 images\n   else: \n         valid_folder = .... set to test folder\n         valid_image_range = ... glob to find number of images ...\n\n```",
      "votes": null
    },
    {
      "id": "2532652",
      "postDate": "11/21/2023 08:19:10",
      "content": "<p>Thanks guys, so I guess those approaches correspond to my <em>\"Or maybe I'd just have to output some dummy data for the example data so that it appeared to work on that, and then do my real processing on the hidden data once it got that far\"</em> kind of approach, i.e. deliberately doing something different to get past the first run.</p>\n<p>But can anyone confirm that is how Kaggle works with notebook competition submissions -- i.e. does it run my notebook once with the public data, verifying that succeeds (and maybe even checking I have a valid submission file for <em>that</em> data), before running it again on the hidden data? If so I need to at least output dummy data the first time that has the right IDs etc to get through to the 'hidden' scoring stage.</p>",
      "rawMarkdown": "Thanks guys, so I guess those approaches correspond to my _\"Or maybe I'd just have to output some dummy data for the example data so that it appeared to work on that, and then do my real processing on the hidden data once it got that far\"_ kind of approach, i.e. deliberately doing something different to get past the first run.\n\nBut can anyone confirm that is how Kaggle works with notebook competition submissions -- i.e. does it run my notebook once with the public data, verifying that succeeds (and maybe even checking I have a valid submission file for _that_ data), before running it again on the hidden data? If so I need to at least output dummy data the first time that has the right IDs etc to get through to the 'hidden' scoring stage.",
      "votes": null
    },
    {
      "id": "2538611",
      "postDate": "11/26/2023 10:21:26",
      "content": "<blockquote>\n  <p>But can anyone confirm that is how Kaggle works with notebook competition submissions</p>\n</blockquote>\n<p>Once you submit the notebook, it is only run on the hidden data (there is no checking with public data). Then the output file (<code>submission.csv</code>) is saved by Kaggle and used to compute the score via the metric for the competition.</p>",
      "rawMarkdown": ">But can anyone confirm that is how Kaggle works with notebook competition submissions\n\nOnce you submit the notebook, it is only run on the hidden data (there is no checking with public data). Then the output file (`submission.csv`) is saved by Kaggle and used to compute the score via the metric for the competition.",
      "votes": null
    },
    {
      "id": "2538657",
      "postDate": "11/26/2023 10:57:22",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/coderrkj\" target=\"_blank\">@coderrkj</a> . But do I get to <em>see</em> the competition run submission.csv, or other notebook output from running on the hidden test dataset, which would let me see the processed filenames for example? I did try submitting a dummy notebook but it only seemed to show me output from having processed the practice test dataset, not the real hidden one…</p>",
      "rawMarkdown": "Thanks @coderrkj . But do I get to *see* the competition run submission.csv, or other notebook output from running on the hidden test dataset, which would let me see the processed filenames for example? I did try submitting a dummy notebook but it only seemed to show me output from having processed the practice test dataset, not the real hidden one...",
      "votes": null
    },
    {
      "id": "2538665",
      "postDate": "11/26/2023 11:07:02",
      "content": "<p><a href=\"https://www.kaggle.com/charliewartnaby\" target=\"_blank\">@charliewartnaby</a> Yes, Kaggle will not let you see the <code>submission.csv</code> file generated by the hidden test set (only whether the run successfully completed and the public LB score). That will always be the case for all competitions.</p>",
      "rawMarkdown": "charliewartnaby Yes, Kaggle will not let you see the `submission.csv` file generated by the hidden test set (only whether the run successfully completed and the public LB score). That will always be the case for all competitions.",
      "votes": null
    },
    {
      "id": "2544941",
      "postDate": "12/01/2023 07:24:22",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/codeRKJ\" target=\"_blank\">@codeRKJ</a>, I can see that makes sense, otherwise information can be gleaned about the test set. (Sorry didn't get email notification about your reply!)</p>",
      "rawMarkdown": "Thanks @codeRKJ, I can see that makes sense, otherwise information can be gleaned about the test set. (Sorry didn't get email notification about your reply!)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2532249,
      "author_name": "charliewartnaby",
      "author_url": "",
      "post_date": "11/20/2023 21:35:56",
      "content": "<p>This is also asked in <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455716\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455716</a> . I agree it is crucial to know this. We have no idea how many kidneys the ~1500 test images cover or what the minimum number of slices per kidney is. If it's too \"thin\", it totally precludes some 3D approaches where we might need a decent number of voxels in the z-direction.</p>\n<p>I tried to just list the test files by making a dummy submission of a noteboook which listed what files were in 'test', but it just gave me the same 6 sample images as what we see in the public dataset. I'm a newbie, so I don't really understand how the real test set is substituted in place of the example test set -- maybe I'm seeing the output of a \"dry run\" of my notebook, and I'm not allowed to see the output of the true competition run?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2532251,
          "author_name": "dhinkris",
          "author_url": "",
          "post_date": "11/20/2023 21:39:09",
          "content": "<p>Same question posted here. Added to make sure not missing any information. <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456047\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456047</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2532264,
          "author_name": "clairewalsh",
          "author_url": "",
          "post_date": "11/20/2023 22:11:34",
          "content": "<p>Hi all, thanks for the questions regarding test data. The quick response is that the training data represent the test data well in terms of the 3D sizes. i.e the 3rd dimension will be usable and quite possibly important for good segmentations. Indeed the 3D nature of this imaging technique is one of its key features. I am prepping a fuller post regarding the test and training datasets in the next day or two we just wanted to see what the common needs/qs were before replying to everyone too quickly. </p>",
          "votes": null,
          "replies": [
            {
              "id": 2532292,
              "author_name": "dhinkris",
              "author_url": "",
              "post_date": "11/20/2023 23:14:21",
              "content": "<p>that will be incredibly helpful. Thank you tons :)</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2532563,
              "author_name": "charliewartnaby",
              "author_url": "",
              "post_date": "11/21/2023 06:43:09",
              "content": "<p>Ah that's good to know, thanks. However, as a newbie I'm worried about whether it will affect the submission process. I get the impression that Kaggle first puts a submitted notebook through a dry run on the publicly viewable test data, and only if it succeeds then runs it again with the hidden test data. If (say) I wanted to run something that needed batches of 16 images in the z-direction, it would never work on the example test data, so my notebook wouldn't proceed to the scored submission run. Or maybe I'd just have to output some dummy data for the example data so that it <em>appeared</em> to work on that, and then do my real processing on the hidden data once it got that far… However, I don't clearly understand how notebook execution works for competition submissions, it doesn't seem to be explained in detail (not specifically to this competition but generally on Kaggle) so I may well be getting the wrong end of the stick. (If anyone can point me to a detailed explanation of how notebooks are run on competition test data I'd be glad, e.g. am I even allowed to see notebook output corresponding to the hidden data, maybe not?)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2532593,
                  "author_name": "sakvaua",
                  "author_url": "",
                  "post_date": "11/21/2023 07:06:00",
                  "content": "<p>Use the train data folder during your dummy run and the test folder when hidden test data is present. Just check the number of files in the test folder to determine which stage you are at.</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2532641,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "11/21/2023 07:57:02",
                  "content": "<p>\" it would never work on the example test data, \"</p>\n<p>it should something like</p>\n<pre><code>=\\\n            #local   development\n  #    #submit\n\n =='local':\n    valid_folder =    train folder\n    valid_image_range = (0,64)  #first 64 images\n\n =='submit':\n    is_not_server = .   glob test kidney_5 folder images  check  there are only 3 images\n     is_not_server:\n         valid_folder =    train folder\n         valid_image_range = (0,64)  #first 64 images\n   : \n         valid_folder =    test folder\n         valid_image_range = . glob   number of images .\n</code></pre>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 2532652,
              "author_name": "charliewartnaby",
              "author_url": "",
              "post_date": "11/21/2023 08:19:10",
              "content": "<p>Thanks guys, so I guess those approaches correspond to my <em>\"Or maybe I'd just have to output some dummy data for the example data so that it appeared to work on that, and then do my real processing on the hidden data once it got that far\"</em> kind of approach, i.e. deliberately doing something different to get past the first run.</p>\n<p>But can anyone confirm that is how Kaggle works with notebook competition submissions -- i.e. does it run my notebook once with the public data, verifying that succeeds (and maybe even checking I have a valid submission file for <em>that</em> data), before running it again on the hidden data? If so I need to at least output dummy data the first time that has the right IDs etc to get through to the 'hidden' scoring stage.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2538611,
                  "author_name": "coderrkj",
                  "author_url": "",
                  "post_date": "11/26/2023 10:21:26",
                  "content": "<blockquote>\n  <p>But can anyone confirm that is how Kaggle works with notebook competition submissions</p>\n</blockquote>\n<p>Once you submit the notebook, it is only run on the hidden data (there is no checking with public data). Then the output file (<code>submission.csv</code>) is saved by Kaggle and used to compute the score via the metric for the competition.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2538657,
                      "author_name": "charliewartnaby",
                      "author_url": "",
                      "post_date": "11/26/2023 10:57:22",
                      "content": "<p>Thanks <a href=\"https://www.kaggle.com/coderrkj\" target=\"_blank\">@coderrkj</a> . But do I get to <em>see</em> the competition run submission.csv, or other notebook output from running on the hidden test dataset, which would let me see the processed filenames for example? I did try submitting a dummy notebook but it only seemed to show me output from having processed the practice test dataset, not the real hidden one…</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2538665,
                          "author_name": "coderrkj",
                          "author_url": "",
                          "post_date": "11/26/2023 11:07:02",
                          "content": "<p><a href=\"https://www.kaggle.com/charliewartnaby\" target=\"_blank\">@charliewartnaby</a> Yes, Kaggle will not let you see the <code>submission.csv</code> file generated by the hidden test set (only whether the run successfully completed and the public LB score). That will always be the case for all competitions.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2544941,
                              "author_name": "charliewartnaby",
                              "author_url": "",
                              "post_date": "12/01/2023 07:24:22",
                              "content": "<p>Thanks <a href=\"https://www.kaggle.com/codeRKJ\" target=\"_blank\">@codeRKJ</a>, I can see that makes sense, otherwise information can be gleaned about the test set. (Sorry didn't get email notification about your reply!)</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2532481,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/21/2023 05:36:47",
      "content": "<p>i thought the title of the competition is \"Segment vasculature in 3D scans of human kidney\"</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2532196": "I'm sure I'm not the only one with this question.\nThe given testing set is very small and obviously not representative of the “real” testing set that will be used at the end of the competition. However, it is unclear if the real/full dataset (it is mentioned it will be about 1500 images) will be a few random 2D slices of 3D volumes (essentially making three-dimensional data very difficult to exploit in the dataset) or if we can expect to have a sequence of 2D slices that can be stacked into reasonable 3D volumes.\n\nThis is important because it allows us to plan to exploit three-dimensional structures of the inputs or not. They are obviously in the training set, but it's not clear at all if they will be in the testing set.\n\nThanks in advance.",
    "2532249": "This is also asked in https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455716 . I agree it is crucial to know this. We have no idea how many kidneys the ~1500 test images cover or what the minimum number of slices per kidney is. If it's too \"thin\", it totally precludes some 3D approaches where we might need a decent number of voxels in the z-direction.\n\nI tried to just list the test files by making a dummy submission of a noteboook which listed what files were in 'test', but it just gave me the same 6 sample images as what we see in the public dataset. I'm a newbie, so I don't really understand how the real test set is substituted in place of the example test set -- maybe I'm seeing the output of a \"dry run\" of my notebook, and I'm not allowed to see the output of the true competition run?",
    "2532251": "Same question posted here. Added to make sure not missing any information. https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/456047",
    "2532264": "Hi all, thanks for the questions regarding test data. The quick response is that the training data represent the test data well in terms of the 3D sizes. i.e the 3rd dimension will be usable and quite possibly important for good segmentations. Indeed the 3D nature of this imaging technique is one of its key features. I am prepping a fuller post regarding the test and training datasets in the next day or two we just wanted to see what the common needs/qs were before replying to everyone too quickly.",
    "2532292": "that will be incredibly helpful. Thank you tons :)",
    "2532481": "i thought the title of the competition is \"Segment vasculature in 3D scans of human kidney\"",
    "2532563": "Ah that's good to know, thanks. However, as a newbie I'm worried about whether it will affect the submission process. I get the impression that Kaggle first puts a submitted notebook through a dry run on the publicly viewable test data, and only if it succeeds then runs it again with the hidden test data. If (say) I wanted to run something that needed batches of 16 images in the z-direction, it would never work on the example test data, so my notebook wouldn't proceed to the scored submission run. Or maybe I'd just have to output some dummy data for the example data so that it _appeared_ to work on that, and then do my real processing on the hidden data once it got that far... However, I don't clearly understand how notebook execution works for competition submissions, it doesn't seem to be explained in detail (not specifically to this competition but generally on Kaggle) so I may well be getting the wrong end of the stick. (If anyone can point me to a detailed explanation of how notebooks are run on competition test data I'd be glad, e.g. am I even allowed to see notebook output corresponding to the hidden data, maybe not?)",
    "2532593": "Use the train data folder during your dummy run and the test folder when hidden test data is present. Just check the number of files in the test folder to determine which stage you are at.",
    "2532641": "\" it would never work on the example test data, \"\n\nit should something like\n\n\n```\nmode=\\\n  'local'          #local debug for development\n  #'submit'    #submit\n\nif mode=='local':\n    valid_folder = .... set to train folder\n    valid_image_range = (0,64)  #first 64 images\n\nif mode=='submit':\n    is_not_server = ... true if glob test kidney_5 folder images and check if there are only 3 images\n    if is_not_server:\n         valid_folder = .... set to train folder\n         valid_image_range = (0,64)  #first 64 images\n   else: \n         valid_folder = .... set to test folder\n         valid_image_range = ... glob to find number of images ...\n\n```",
    "2532652": "Thanks guys, so I guess those approaches correspond to my _\"Or maybe I'd just have to output some dummy data for the example data so that it appeared to work on that, and then do my real processing on the hidden data once it got that far\"_ kind of approach, i.e. deliberately doing something different to get past the first run.\n\nBut can anyone confirm that is how Kaggle works with notebook competition submissions -- i.e. does it run my notebook once with the public data, verifying that succeeds (and maybe even checking I have a valid submission file for _that_ data), before running it again on the hidden data? If so I need to at least output dummy data the first time that has the right IDs etc to get through to the 'hidden' scoring stage.",
    "2538611": ">But can anyone confirm that is how Kaggle works with notebook competition submissions\n\nOnce you submit the notebook, it is only run on the hidden data (there is no checking with public data). Then the output file (`submission.csv`) is saved by Kaggle and used to compute the score via the metric for the competition.",
    "2538657": "Thanks @coderrkj . But do I get to *see* the competition run submission.csv, or other notebook output from running on the hidden test dataset, which would let me see the processed filenames for example? I did try submitting a dummy notebook but it only seemed to show me output from having processed the practice test dataset, not the real hidden one...",
    "2538665": "charliewartnaby Yes, Kaggle will not let you see the `submission.csv` file generated by the hidden test set (only whether the run successfully completed and the public LB score). That will always be the case for all competitions.",
    "2544941": "Thanks @codeRKJ, I can see that makes sense, otherwise information can be gleaned about the test set. (Sorry didn't get email notification about your reply!)"
  },
  "source": "meta"
}