{
  "id": 267375,
  "title": "Use your own preprocessed dataset",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/267375",
  "author_name": "",
  "post_date": "2021-08-23T00:03:26.244898600Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>It seems I cannot reference my own preprocessed dataset (DICOM -&gt; PNG) in my notebook when I submit my results. I get a \"Notebook Threw Exception\" error in the submissions panel and I think its due to the reference. Can someone confirm this is the case?</p>\n<p>I am just a bit confused because it says in the rules that you can use pretrained models, but because internet access is disabled we would need to load them as a Kaggle dataset, so I'm not sure how is anyone actually using pretrained models when we can't actually reference datasets other than the competition dataset when submitting?</p>",
  "messages": [
    {
      "id": "1486401",
      "postDate": "08/23/2021 00:03:26",
      "content": "<p>It seems I cannot reference my own preprocessed dataset (DICOM -&gt; PNG) in my notebook when I submit my results. I get a \"Notebook Threw Exception\" error in the submissions panel and I think its due to the reference. Can someone confirm this is the case?</p>\n<p>I am just a bit confused because it says in the rules that you can use pretrained models, but because internet access is disabled we would need to load them as a Kaggle dataset, so I'm not sure how is anyone actually using pretrained models when we can't actually reference datasets other than the competition dataset when submitting?</p>",
      "rawMarkdown": "It seems I cannot reference my own preprocessed dataset (DICOM -> PNG) in my notebook when I submit my results. I get a \"Notebook Threw Exception\" error in the submissions panel and I think its due to the reference. Can someone confirm this is the case?\n\nI am just a bit confused because it says in the rules that you can use pretrained models, but because internet access is disabled we would need to load them as a Kaggle dataset, so I'm not sure how is anyone actually using pretrained models when we can't actually reference datasets other than the competition dataset when submitting?",
      "votes": null
    },
    {
      "id": "1486417",
      "postDate": "08/23/2021 00:38:04",
      "content": "<p>same happened to me. Originally I  don't know the reason. Today, when I tried to submit result, It complains about none version dataset! which is pretrained weights.</p>",
      "rawMarkdown": "same happened to me. Originally I  don't know the reason. Today, when I tried to submit result, It complains about none version dataset! which is pretrained weights.",
      "votes": null
    },
    {
      "id": "1487253",
      "postDate": "08/23/2021 14:00:07",
      "content": "<p>Think I understand why, the private dataset they use when testing your submission is not the same. I see when I go to submit predictions it says \"In this competition, we will privately re-run your selected Notebook Version with a hidden test set substituted into the competition dataset. We then extract your chosen Output File from the re-run and use that to determine your score.\" </p>\n<p>So they are <em>adding</em> something or changing something about the dataset in the submission environment. For me, this means my code would throw an exception because it is looking for test cases by id, and if a sample is missing then there would be an exception. Conversely, if there was an added test case then my code would not produce the required output for that test case.</p>",
      "rawMarkdown": "Think I understand why, the private dataset they use when testing your submission is not the same. I see when I go to submit predictions it says \"In this competition, we will privately re-run your selected Notebook Version with a hidden test set substituted into the competition dataset. We then extract your chosen Output File from the re-run and use that to determine your score.\" \n\nSo they are _adding_ something or changing something about the dataset in the submission environment. For me, this means my code would throw an exception because it is looking for test cases by id, and if a sample is missing then there would be an exception. Conversely, if there was an added test case then my code would not produce the required output for that test case.",
      "votes": null
    },
    {
      "id": "1487254",
      "postDate": "08/23/2021 14:00:47",
      "content": "<p>On that note, I really wish there was better error reporting for things like this such as a saved stacktrace for the submission, would make debugging these things a lot easier than just shooting in the dark…</p>",
      "rawMarkdown": "On that note, I really wish there was better error reporting for things like this such as a saved stacktrace for the submission, would make debugging these things a lot easier than just shooting in the dark...",
      "votes": null
    },
    {
      "id": "1487293",
      "postDate": "08/23/2021 14:28:00",
      "content": "<p>I think the fundamental reason why they do this is </p>\n<ol>\n<li>prevent spamming public leaderboard and getting lucky </li>\n<li>Per the rules, you are allowed to access \"Freely &amp; publicly available external data is allowed, including pre-trained models\", which does <em>not</em> include preprocessed private datasets of your own. I think this is becaue allowing this would allow people to bypass the 9 hour computation limit for notebook submissions, because you could do the dataset processing in a separate notebook and then just reference the results in your submission as a dataset. </li>\n</ol>\n<p>There are probably ways you could drastically increase the total computation time spent on any given submission using the above reasoning, which is not the point of the competition.</p>",
      "rawMarkdown": "I think the fundamental reason why they do this is \n1. prevent spamming public leaderboard and getting lucky \n2. Per the rules, you are allowed to access \"Freely & publicly available external data is allowed, including pre-trained models\", which does _not_ include preprocessed private datasets of your own. I think this is becaue allowing this would allow people to bypass the 9 hour computation limit for notebook submissions, because you could do the dataset processing in a separate notebook and then just reference the results in your submission as a dataset. \n\nThere are probably ways you could drastically increase the total computation time spent on any given submission using the above reasoning, which is not the point of the competition.",
      "votes": null
    },
    {
      "id": "1490540",
      "postDate": "08/25/2021 17:16:04",
      "content": "<p>It seems that the notebook environment has recovered.  now I have tried other models with pretrained weights,  though the result is not good.</p>",
      "rawMarkdown": "It seems that the notebook environment has recovered.  now I have tried other models with pretrained weights,  though the result is not good.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1486417,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "08/23/2021 00:38:04",
      "content": "<p>same happened to me. Originally I  don't know the reason. Today, when I tried to submit result, It complains about none version dataset! which is pretrained weights.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1487253,
          "author_name": "d223chen",
          "author_url": "",
          "post_date": "08/23/2021 14:00:07",
          "content": "<p>Think I understand why, the private dataset they use when testing your submission is not the same. I see when I go to submit predictions it says \"In this competition, we will privately re-run your selected Notebook Version with a hidden test set substituted into the competition dataset. We then extract your chosen Output File from the re-run and use that to determine your score.\" </p>\n<p>So they are <em>adding</em> something or changing something about the dataset in the submission environment. For me, this means my code would throw an exception because it is looking for test cases by id, and if a sample is missing then there would be an exception. Conversely, if there was an added test case then my code would not produce the required output for that test case.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1487254,
          "author_name": "d223chen",
          "author_url": "",
          "post_date": "08/23/2021 14:00:47",
          "content": "<p>On that note, I really wish there was better error reporting for things like this such as a saved stacktrace for the submission, would make debugging these things a lot easier than just shooting in the dark…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1487293,
          "author_name": "d223chen",
          "author_url": "",
          "post_date": "08/23/2021 14:28:00",
          "content": "<p>I think the fundamental reason why they do this is </p>\n<ol>\n<li>prevent spamming public leaderboard and getting lucky </li>\n<li>Per the rules, you are allowed to access \"Freely &amp; publicly available external data is allowed, including pre-trained models\", which does <em>not</em> include preprocessed private datasets of your own. I think this is becaue allowing this would allow people to bypass the 9 hour computation limit for notebook submissions, because you could do the dataset processing in a separate notebook and then just reference the results in your submission as a dataset. </li>\n</ol>\n<p>There are probably ways you could drastically increase the total computation time spent on any given submission using the above reasoning, which is not the point of the competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1490540,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "08/25/2021 17:16:04",
      "content": "<p>It seems that the notebook environment has recovered.  now I have tried other models with pretrained weights,  though the result is not good.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1486401": "It seems I cannot reference my own preprocessed dataset (DICOM -> PNG) in my notebook when I submit my results. I get a \"Notebook Threw Exception\" error in the submissions panel and I think its due to the reference. Can someone confirm this is the case?\n\nI am just a bit confused because it says in the rules that you can use pretrained models, but because internet access is disabled we would need to load them as a Kaggle dataset, so I'm not sure how is anyone actually using pretrained models when we can't actually reference datasets other than the competition dataset when submitting?",
    "1486417": "same happened to me. Originally I  don't know the reason. Today, when I tried to submit result, It complains about none version dataset! which is pretrained weights.",
    "1487253": "Think I understand why, the private dataset they use when testing your submission is not the same. I see when I go to submit predictions it says \"In this competition, we will privately re-run your selected Notebook Version with a hidden test set substituted into the competition dataset. We then extract your chosen Output File from the re-run and use that to determine your score.\" \n\nSo they are _adding_ something or changing something about the dataset in the submission environment. For me, this means my code would throw an exception because it is looking for test cases by id, and if a sample is missing then there would be an exception. Conversely, if there was an added test case then my code would not produce the required output for that test case.",
    "1487254": "On that note, I really wish there was better error reporting for things like this such as a saved stacktrace for the submission, would make debugging these things a lot easier than just shooting in the dark...",
    "1487293": "I think the fundamental reason why they do this is \n1. prevent spamming public leaderboard and getting lucky \n2. Per the rules, you are allowed to access \"Freely & publicly available external data is allowed, including pre-trained models\", which does _not_ include preprocessed private datasets of your own. I think this is becaue allowing this would allow people to bypass the 9 hour computation limit for notebook submissions, because you could do the dataset processing in a separate notebook and then just reference the results in your submission as a dataset. \n\nThere are probably ways you could drastically increase the total computation time spent on any given submission using the above reasoning, which is not the point of the competition.",
    "1490540": "It seems that the notebook environment has recovered.  now I have tried other models with pretrained weights,  though the result is not good."
  },
  "source": "meta"
}