{
  "id": 549936,
  "title": "Where is located the test data for submission ?",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/549936",
  "author_name": "",
  "post_date": "2024-12-04T16:09:55.592525400Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>For submission I use the following code snippet to retreive the test data:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3138065%2Fdf272cc70f9f6125958c74b9828b237e%2F2024-12-04%2017_06_02-notebookd1edafe1ae%20_%20Kaggle%20et%204%20pages%20de%20plus%20-%20Personnel%20%20Microsoft%20Edge.png?generation=1733328389262675&amp;alt=media\" alt=\"\"></p>\n<p>It works fine to do the prediction on the fake test set provided for dev purposes. But I was under the impression that this fake test set was supposed to be replaced by the real one when the notebook is submitted. I am missing something ? The submission.csv after the run of the notebook is populated with the fake test set results.</p>",
  "messages": [
    {
      "id": "3063572",
      "postDate": "12/04/2024 16:09:55",
      "content": "<p>Hi,</p>\n<p>For submission I use the following code snippet to retreive the test data:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3138065%2Fdf272cc70f9f6125958c74b9828b237e%2F2024-12-04%2017_06_02-notebookd1edafe1ae%20_%20Kaggle%20et%204%20pages%20de%20plus%20-%20Personnel%20%20Microsoft%20Edge.png?generation=1733328389262675&amp;alt=media\" alt=\"\"></p>\n<p>It works fine to do the prediction on the fake test set provided for dev purposes. But I was under the impression that this fake test set was supposed to be replaced by the real one when the notebook is submitted. I am missing something ? The submission.csv after the run of the notebook is populated with the fake test set results.</p>",
      "rawMarkdown": "Hi,\n\nFor submission I use the following code snippet to retreive the test data:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3138065%2Fdf272cc70f9f6125958c74b9828b237e%2F2024-12-04%2017_06_02-notebookd1edafe1ae%20_%20Kaggle%20et%204%20pages%20de%20plus%20-%20Personnel%20%20Microsoft%20Edge.png?generation=1733328389262675&alt=media)\n\nIt works fine to do the prediction on the fake test set provided for dev purposes. But I was under the impression that this fake test set was supposed to be replaced by the real one when the notebook is submitted. I am missing something ? The submission.csv after the run of the notebook is populated with the fake test set results.",
      "votes": null
    },
    {
      "id": "3063632",
      "postDate": "12/04/2024 17:24:11",
      "content": "<p>Are you sure crashes reading hidden test? As far I know is just like that, same folder, different samples.</p>",
      "rawMarkdown": "Are you sure crashes reading hidden test? As far I know is just like that, same folder, different samples.",
      "votes": null
    },
    {
      "id": "3063731",
      "postDate": "12/04/2024 19:57:39",
      "content": "<p>You won't be able to see the submission.csv file that's created when you submit your notebook to the competition.  What happens for me is that it saves a new version of my notebook and also submits it.  The new saved version will have a submission.csv file from the three \"test\" tomograms.  Is that what you're looking at?</p>",
      "rawMarkdown": "You won't be able to see the submission.csv file that's created when you submit your notebook to the competition.  What happens for me is that it saves a new version of my notebook and also submits it.  The new saved version will have a submission.csv file from the three \"test\" tomograms.  Is that what you're looking at?",
      "votes": null
    },
    {
      "id": "3063755",
      "postDate": "12/04/2024 20:27:16",
      "content": "<p>That makes more sense. I didn't read till the end.</p>",
      "rawMarkdown": "That makes more sense. I didn't read till the end.",
      "votes": null
    },
    {
      "id": "3063760",
      "postDate": "12/04/2024 20:30:10",
      "content": "<blockquote>\n  <p>You won't be able to see the submission.csv</p>\n</blockquote>\n<p>That is interesting, maybe I'm submitting the notebook in an incorect way ?</p>\n<p>What I did was the following:<br>\n1) Go to \"Submit to competition\" within the online notebook editor.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3138065%2Fd1b204bc41f0e54af43566e7e0f8e0fe%2F2024-12-04%2021_17_40-notebookd1edafe1ae%20_%20Kaggle%20et%204%20pages%20de%20plus%20-%20Personnel%20%20Microsoft%20Edge.png?generation=1733343731354901&amp;alt=media\" alt=\"\"></p>\n<p>2) Submit</p>\n<p>3) Review the notebook' submission logs.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3138065%2Febc4f7aefb0b0a35362685b19c25a596%2F2024-12-04%2021_19_38-notebookd1edafe1ae%20et%204%20pages%20de%20plus%20-%20Personnel%20%20Microsoft%20Edge.png?generation=1733343854595896&amp;alt=media\" alt=\"\"></p>\n<p>In the logs I just printed to the console the experiment names (which were the \"false\" test ones). Also the notebook runned in less than 10minutes, which I'm quite sure is not possible if it was running on the 500 test samples.</p>\n<p>Sorry if this seems like just a user-error (which is quite obviously the case) but I did not use Kaggle since a few years ago, so I'm not quite up to date with everything.</p>",
      "rawMarkdown": "> You won't be able to see the submission.csv\n\nThat is interesting, maybe I'm submitting the notebook in an incorect way ?\n\nWhat I did was the following:\n1) Go to \"Submit to competition\" within the online notebook editor.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3138065%2Fd1b204bc41f0e54af43566e7e0f8e0fe%2F2024-12-04%2021_17_40-notebookd1edafe1ae%20_%20Kaggle%20et%204%20pages%20de%20plus%20-%20Personnel%20%20Microsoft%20Edge.png?generation=1733343731354901&alt=media)\n\n2) Submit\n\n3) Review the notebook' submission logs.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3138065%2Febc4f7aefb0b0a35362685b19c25a596%2F2024-12-04%2021_19_38-notebookd1edafe1ae%20et%204%20pages%20de%20plus%20-%20Personnel%20%20Microsoft%20Edge.png?generation=1733343854595896&alt=media)\n\nIn the logs I just printed to the console the experiment names (which were the \"false\" test ones). Also the notebook runned in less than 10minutes, which I'm quite sure is not possible if it was running on the 500 test samples.\n\nSorry if this seems like just a user-error (which is quite obviously the case) but I did not use Kaggle since a few years ago, so I'm not quite up to date with everything.",
      "votes": null
    },
    {
      "id": "3063773",
      "postDate": "12/04/2024 20:44:18",
      "content": "<p>Yeah, those are the logs from the new version of your notebook that gets created.  They don't show you any of the output against the data used to grade your submission.   Check under the Submissions tab of the contest and see if you have anything listed there.  You should have entries for each submission with either a score or an error.</p>",
      "rawMarkdown": "Yeah, those are the logs from the new version of your notebook that gets created.  They don't show you any of the output against the data used to grade your submission.   Check under the Submissions tab of the contest and see if you have anything listed there.  You should have entries for each submission with either a score or an error.",
      "votes": null
    },
    {
      "id": "3063787",
      "postDate": "12/04/2024 21:23:21",
      "content": "<blockquote>\n  <p>Yeah, those are the logs from the new version of your notebook that gets created</p>\n</blockquote>\n<p>Okay thanks, now I get it. Basically the notebook is first runned on the public available dataset successfully and then again in another environement with the private dataset and something fail at that moment.</p>\n<p>I do have an error indeed, which is an unhandled exception. I think I know where it might be. Its a shame that we can't even get a stack trace. I guess its probably just to avoid people abusing the stack trace to retreive data from the test set. </p>\n<p>It was very helpful, thanks for your help.</p>",
      "rawMarkdown": "> Yeah, those are the logs from the new version of your notebook that gets created\n\nOkay thanks, now I get it. Basically the notebook is first runned on the public available dataset successfully and then again in another environement with the private dataset and something fail at that moment.\n\nI do have an error indeed, which is an unhandled exception. I think I know where it might be. Its a shame that we can't even get a stack trace. I guess its probably just to avoid people abusing the stack trace to retreive data from the test set. \n\nIt was very helpful, thanks for your help.",
      "votes": null
    },
    {
      "id": "3063799",
      "postDate": "12/04/2024 21:46:32",
      "content": "<p>The directory structure might have been setup with the expectation that we use copick to go through the files which may simplify some things.  I'm noticing that your code appears to make the assumption that every file it finds will be a zarr directory which may not be the case.  You might add some logic to check that the file you find is a directory and/or trap any exception and continue on to the next file.  (Note, I spent about 10 submissions early on with a very similar issue, although ultimately mine was that I was mistakenly trying to find files in the training directory.)</p>",
      "rawMarkdown": "The directory structure might have been setup with the expectation that we use copick to go through the files which may simplify some things.  I'm noticing that your code appears to make the assumption that every file it finds will be a zarr directory which may not be the case.  You might add some logic to check that the file you find is a directory and/or trap any exception and continue on to the next file.  (Note, I spent about 10 submissions early on with a very similar issue, although ultimately mine was that I was mistakenly trying to find files in the training directory.)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3063632,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "12/04/2024 17:24:11",
      "content": "<p>Are you sure crashes reading hidden test? As far I know is just like that, same folder, different samples.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3063731,
      "author_name": "davidlist",
      "author_url": "",
      "post_date": "12/04/2024 19:57:39",
      "content": "<p>You won't be able to see the submission.csv file that's created when you submit your notebook to the competition.  What happens for me is that it saves a new version of my notebook and also submits it.  The new saved version will have a submission.csv file from the three \"test\" tomograms.  Is that what you're looking at?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3063755,
          "author_name": "sacuscreed",
          "author_url": "",
          "post_date": "12/04/2024 20:27:16",
          "content": "<p>That makes more sense. I didn't read till the end.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 3063760,
          "author_name": "jgringet",
          "author_url": "",
          "post_date": "12/04/2024 20:30:10",
          "content": "<blockquote>\n  <p>You won't be able to see the submission.csv</p>\n</blockquote>\n<p>That is interesting, maybe I'm submitting the notebook in an incorect way ?</p>\n<p>What I did was the following:<br>\n1) Go to \"Submit to competition\" within the online notebook editor.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3138065%2Fd1b204bc41f0e54af43566e7e0f8e0fe%2F2024-12-04%2021_17_40-notebookd1edafe1ae%20_%20Kaggle%20et%204%20pages%20de%20plus%20-%20Personnel%20%20Microsoft%20Edge.png?generation=1733343731354901&amp;alt=media\" alt=\"\"></p>\n<p>2) Submit</p>\n<p>3) Review the notebook' submission logs.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3138065%2Febc4f7aefb0b0a35362685b19c25a596%2F2024-12-04%2021_19_38-notebookd1edafe1ae%20et%204%20pages%20de%20plus%20-%20Personnel%20%20Microsoft%20Edge.png?generation=1733343854595896&amp;alt=media\" alt=\"\"></p>\n<p>In the logs I just printed to the console the experiment names (which were the \"false\" test ones). Also the notebook runned in less than 10minutes, which I'm quite sure is not possible if it was running on the 500 test samples.</p>\n<p>Sorry if this seems like just a user-error (which is quite obviously the case) but I did not use Kaggle since a few years ago, so I'm not quite up to date with everything.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3063773,
              "author_name": "davidlist",
              "author_url": "",
              "post_date": "12/04/2024 20:44:18",
              "content": "<p>Yeah, those are the logs from the new version of your notebook that gets created.  They don't show you any of the output against the data used to grade your submission.   Check under the Submissions tab of the contest and see if you have anything listed there.  You should have entries for each submission with either a score or an error.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3063787,
                  "author_name": "jgringet",
                  "author_url": "",
                  "post_date": "12/04/2024 21:23:21",
                  "content": "<blockquote>\n  <p>Yeah, those are the logs from the new version of your notebook that gets created</p>\n</blockquote>\n<p>Okay thanks, now I get it. Basically the notebook is first runned on the public available dataset successfully and then again in another environement with the private dataset and something fail at that moment.</p>\n<p>I do have an error indeed, which is an unhandled exception. I think I know where it might be. Its a shame that we can't even get a stack trace. I guess its probably just to avoid people abusing the stack trace to retreive data from the test set. </p>\n<p>It was very helpful, thanks for your help.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3063799,
                      "author_name": "davidlist",
                      "author_url": "",
                      "post_date": "12/04/2024 21:46:32",
                      "content": "<p>The directory structure might have been setup with the expectation that we use copick to go through the files which may simplify some things.  I'm noticing that your code appears to make the assumption that every file it finds will be a zarr directory which may not be the case.  You might add some logic to check that the file you find is a directory and/or trap any exception and continue on to the next file.  (Note, I spent about 10 submissions early on with a very similar issue, although ultimately mine was that I was mistakenly trying to find files in the training directory.)</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3063572": "Hi,\n\nFor submission I use the following code snippet to retreive the test data:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3138065%2Fdf272cc70f9f6125958c74b9828b237e%2F2024-12-04%2017_06_02-notebookd1edafe1ae%20_%20Kaggle%20et%204%20pages%20de%20plus%20-%20Personnel%20%20Microsoft%20Edge.png?generation=1733328389262675&alt=media)\n\nIt works fine to do the prediction on the fake test set provided for dev purposes. But I was under the impression that this fake test set was supposed to be replaced by the real one when the notebook is submitted. I am missing something ? The submission.csv after the run of the notebook is populated with the fake test set results.",
    "3063632": "Are you sure crashes reading hidden test? As far I know is just like that, same folder, different samples.",
    "3063731": "You won't be able to see the submission.csv file that's created when you submit your notebook to the competition.  What happens for me is that it saves a new version of my notebook and also submits it.  The new saved version will have a submission.csv file from the three \"test\" tomograms.  Is that what you're looking at?",
    "3063755": "That makes more sense. I didn't read till the end.",
    "3063760": "> You won't be able to see the submission.csv\n\nThat is interesting, maybe I'm submitting the notebook in an incorect way ?\n\nWhat I did was the following:\n1) Go to \"Submit to competition\" within the online notebook editor.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3138065%2Fd1b204bc41f0e54af43566e7e0f8e0fe%2F2024-12-04%2021_17_40-notebookd1edafe1ae%20_%20Kaggle%20et%204%20pages%20de%20plus%20-%20Personnel%20%20Microsoft%20Edge.png?generation=1733343731354901&alt=media)\n\n2) Submit\n\n3) Review the notebook' submission logs.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3138065%2Febc4f7aefb0b0a35362685b19c25a596%2F2024-12-04%2021_19_38-notebookd1edafe1ae%20et%204%20pages%20de%20plus%20-%20Personnel%20%20Microsoft%20Edge.png?generation=1733343854595896&alt=media)\n\nIn the logs I just printed to the console the experiment names (which were the \"false\" test ones). Also the notebook runned in less than 10minutes, which I'm quite sure is not possible if it was running on the 500 test samples.\n\nSorry if this seems like just a user-error (which is quite obviously the case) but I did not use Kaggle since a few years ago, so I'm not quite up to date with everything.",
    "3063773": "Yeah, those are the logs from the new version of your notebook that gets created.  They don't show you any of the output against the data used to grade your submission.   Check under the Submissions tab of the contest and see if you have anything listed there.  You should have entries for each submission with either a score or an error.",
    "3063787": "> Yeah, those are the logs from the new version of your notebook that gets created\n\nOkay thanks, now I get it. Basically the notebook is first runned on the public available dataset successfully and then again in another environement with the private dataset and something fail at that moment.\n\nI do have an error indeed, which is an unhandled exception. I think I know where it might be. Its a shame that we can't even get a stack trace. I guess its probably just to avoid people abusing the stack trace to retreive data from the test set. \n\nIt was very helpful, thanks for your help.",
    "3063799": "The directory structure might have been setup with the expectation that we use copick to go through the files which may simplify some things.  I'm noticing that your code appears to make the assumption that every file it finds will be a zarr directory which may not be the case.  You might add some logic to check that the file you find is a directory and/or trap any exception and continue on to the next file.  (Note, I spent about 10 submissions early on with a very similar issue, although ultimately mine was that I was mistakenly trying to find files in the training directory.)"
  },
  "source": "meta"
}