{
  "id": 277603,
  "title": "Help...🥕🥕 Submission Scoring Error..😭",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/277603",
  "author_name": "jypysk",
  "post_date": "2021-10-10T09:11:48.670000",
  "votes": 1,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Good morning to all kaggler! </p>\n<p>Im a very beginner to this kaggle competition, and facing a \"Submission Scoring Error\" now.</p>\n<p><a href=\"https://www.kaggle.com/jypysk/nokfold-test-submission-v0?scriptVersionId=76747394\" target=\"_blank\">https://www.kaggle.com/jypysk/nokfold-test-submission-v0?scriptVersionId=76747394</a><br>\n^ Here is my tiny and simple kaggle notebook for submission. </p>\n<p>So I've check, <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">https://www.kaggle.com/code-competition-debugging</a> &lt; here, </p>\n<ol>\n<li>all the dataset should be publicly opened.</li>\n<li>Should have 87 rows and 2 cols. (Checked the name of each column)</li>\n<li>Should name the csv file as \"submission.csv\" &amp; index=false.</li>\n</ol>\n<p>Can anyone help me with this problem..?<br>\nThank you for any help you can offer me. 🤒</p>",
  "messages": [
    {
      "id": 1540347,
      "postDate": "2021-10-10T12:04:33.763Z",
      "content": "<p>I am not sure, but do you consider hidden test set?</p>\n<blockquote>\n  <p>The competition data is defined by three cohorts: Training, Validation (Public), and Testing (Private). The “Training” and the “Validation” cohorts are provided to the participants, whereas the “Testing” cohort is kept hidden at all times, during and after the competition.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/data\" target=\"_blank\">https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/data</a></p>\n<p>You need to create inference notebook with whole processing pipeline like <a href=\"https://www.kaggle.com/ammarnassanalhajali/brain-tumor-3d-inference\" target=\"_blank\">this one</a>.</p>",
      "rawMarkdown": "I am not sure, but do you consider hidden test set?\n\n>  The competition data is defined by three cohorts: Training, Validation (Public), and Testing (Private). The “Training” and the “Validation” cohorts are provided to the participants, whereas the “Testing” cohort is kept hidden at all times, during and after the competition.\n\nhttps://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/data\n\nYou need to create inference notebook with whole processing pipeline like [this one](https://www.kaggle.com/ammarnassanalhajali/brain-tumor-3d-inference).\n",
      "votes": 1,
      "replies": [
        {
          "id": 1541919,
          "postDate": "2021-10-12T02:44:01.637Z",
          "content": "<p>Thank you for advice! Hm… I didn't know theres a hidden test set… Hmm… Actually I followed this notebook <a href=\"https://www.kaggle.com/luongduongminh/brain-tumor-3d\" target=\"_blank\">https://www.kaggle.com/luongduongminh/brain-tumor-3d</a>. But I could not find big differences between this notebook and mine.. But as your comment, I should add whole precessing pipeline in mine! Thank you a lot for kind reply!!☺️</p>",
          "rawMarkdown": "Thank you for advice! Hm... I didn't know theres a hidden test set... Hmm... Actually I followed this notebook https://www.kaggle.com/luongduongminh/brain-tumor-3d. But I could not find big differences between this notebook and mine.. But as your comment, I should add whole precessing pipeline in mine! Thank you a lot for kind reply!!☺️"
        },
        {
          "id": 1542134,
          "postDate": "2021-10-12T06:53:37.030Z",
          "content": "<p>You should not follow the notebook you linked. This notebook just blend the predictions of other notebooks for public test set and the predictions for hidden (private) test set are all 0.5 (see section [14]). Section [14] prevents scoring error, but its private score is chance level.</p>",
          "rawMarkdown": "You should not follow the notebook you linked. This notebook just blend the predictions of other notebooks for public test set and the predictions for hidden (private) test set are all 0.5 (see section [14]). Section [14] prevents scoring error, but its private score is chance level.",
          "votes": 1
        },
        {
          "id": 1542181,
          "postDate": "2021-10-12T07:56:08.827Z",
          "content": "<p><a href=\"https://www.kaggle.com/jypysk\" target=\"_blank\">@jypysk</a> imho, a notebook with more upvotes is usually a better starting point than a notebook with higher public score. The notebook you followed is actually not useful at all. It doesn't show you how to preprocess the data, how to create and train a model, how to generate predictions from the trained model, etc. It literally just takes public test set predictions from other notebooks, blends them together and (maybe luckily?) gets a high score on the public leaderboard (and as <a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a> mentioned, it doesn't predict the private/hidden test set at all and just set all predictions to default value 0.5). The notebook in <a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a>'s suggestion may have lower score, but I think it is a much better starting point than the other notebook.</p>",
          "rawMarkdown": "@jypysk imho, a notebook with more upvotes is usually a better starting point than a notebook with higher public score. The notebook you followed is actually not useful at all. It doesn't show you how to preprocess the data, how to create and train a model, how to generate predictions from the trained model, etc. It literally just takes public test set predictions from other notebooks, blends them together and (maybe luckily?) gets a high score on the public leaderboard (and as @tomooinubushi mentioned, it doesn't predict the private/hidden test set at all and just set all predictions to default value 0.5). The notebook in @tomooinubushi's suggestion may have lower score, but I think it is a much better starting point than the other notebook.",
          "votes": 1
        },
        {
          "id": 1543165,
          "postDate": "2021-10-13T08:05:09.073Z",
          "content": "<p>Thank you for your kind reply!! Thanks to your help, I finally succeeded to submit! What a wonderful world~!~! Hope you have a great day!!😆</p>",
          "rawMarkdown": "Thank you for your kind reply!! Thanks to your help, I finally succeeded to submit! What a wonderful world~!~! Hope you have a great day!!😆",
          "votes": 1
        }
      ]
    },
    {
      "id": 1540257,
      "postDate": "2021-10-10T09:57:01.873Z",
      "content": "<p>I think you don't need to do this </p>\n<pre><code>Clell 11: BraTSID = [str(x).zfill(5) for x in BraTSID]\n</code></pre>",
      "rawMarkdown": "I think you don't need to do this \n\n```\nClell 11: BraTSID = [str(x).zfill(5) for x in BraTSID]\n```",
      "votes": 1,
      "replies": [
        {
          "id": 1540264,
          "postDate": "2021-10-10T10:09:15.830Z",
          "content": "<p>Thank you for your comment! But still… not working…\"Submission Scoring Error\". Do I have to load '/kaggle/input/test/' at least one time?</p>",
          "rawMarkdown": "Thank you for your comment! But still... not working...\"Submission Scoring Error\". Do I have to load '/kaggle/input/test/' at least one time?"
        }
      ]
    },
    {
      "id": 1540226,
      "postDate": "2021-10-10T09:11:48.670Z",
      "content": "<p>Good morning to all kaggler! </p>\n<p>Im a very beginner to this kaggle competition, and facing a \"Submission Scoring Error\" now.</p>\n<p><a href=\"https://www.kaggle.com/jypysk/nokfold-test-submission-v0?scriptVersionId=76747394\" target=\"_blank\">https://www.kaggle.com/jypysk/nokfold-test-submission-v0?scriptVersionId=76747394</a><br>\n^ Here is my tiny and simple kaggle notebook for submission. </p>\n<p>So I've check, <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">https://www.kaggle.com/code-competition-debugging</a> &lt; here, </p>\n<ol>\n<li>all the dataset should be publicly opened.</li>\n<li>Should have 87 rows and 2 cols. (Checked the name of each column)</li>\n<li>Should name the csv file as \"submission.csv\" &amp; index=false.</li>\n</ol>\n<p>Can anyone help me with this problem..?<br>\nThank you for any help you can offer me. 🤒</p>",
      "rawMarkdown": "Good morning to all kaggler! \n\nIm a very beginner to this kaggle competition, and facing a \"Submission Scoring Error\" now.\n\nhttps://www.kaggle.com/jypysk/nokfold-test-submission-v0?scriptVersionId=76747394\n^ Here is my tiny and simple kaggle notebook for submission. \n\nSo I've check, https://www.kaggle.com/code-competition-debugging < here, \n1. all the dataset should be publicly opened.\n2. Should have 87 rows and 2 cols. (Checked the name of each column)\n3. Should name the csv file as \"submission.csv\" & index=false.\n\nCan anyone help me with this problem..?\nThank you for any help you can offer me. 🤒",
      "votes": 1
    },
    {
      "id": 1542173,
      "postDate": "2021-10-12T07:47:21.277Z",
      "content": "<p>Try to put , index_col=\"BraTS21ID\" while saving the submission file. It works in my case!</p>",
      "rawMarkdown": "Try to put , index_col=\"BraTS21ID\" while saving the submission file. It works in my case!",
      "replies": [
        {
          "id": 1543167,
          "postDate": "2021-10-13T08:05:40.840Z",
          "content": "<p>Thanks! I applied your comment, and succeed to submit~! 😆</p>",
          "rawMarkdown": "Thanks! I applied your comment, and succeed to submit~! 😆",
          "votes": 1
        },
        {
          "id": 1543223,
          "postDate": "2021-10-13T09:01:58.320Z",
          "content": "<p>Your welcome!!😃 and you can vote me here , if it's helpful in your case !!</p>",
          "rawMarkdown": "Your welcome!!😃 and you can vote me here , if it's helpful in your case !!\n"
        }
      ]
    },
    {
      "id": 1540601,
      "postDate": "2021-10-10T16:44:26.017Z",
      "content": "<p>Try to keep column BraTSID as int or just copy 'sample_submission.csv' and fill your MGMT_values. And yes, consider also adding your whole processing pipeline, as you have been advised already.</p>",
      "rawMarkdown": "Try to keep column BraTSID as int or just copy 'sample_submission.csv' and fill your MGMT_values. And yes, consider also adding your whole processing pipeline, as you have been advised already.",
      "replies": [
        {
          "id": 1541917,
          "postDate": "2021-10-12T02:40:07.370Z",
          "content": "<p>Oh…. Thank you for advice!! I will pay 100% attention to adding whole processing pipeline on my notebook! 🙏</p>",
          "rawMarkdown": "Oh.... Thank you for advice!! I will pay 100% attention to adding whole processing pipeline on my notebook! 🙏"
        }
      ]
    },
    {
      "id": 1540309,
      "postDate": "2021-10-10T11:19:52.083Z",
      "content": "<p>May be the train step is hard,so simple ensemble is need for someone who want get high score in LB!</p>",
      "rawMarkdown": "May be the train step is hard,so simple ensemble is need for someone who want get high score in LB!",
      "replies": [
        {
          "id": 1541920,
          "postDate": "2021-10-12T02:45:19.137Z",
          "content": "<p>Um.. May I ask what \"LB\" is rerferred to?</p>",
          "rawMarkdown": "Um.. May I ask what \"LB\" is rerferred to?"
        },
        {
          "id": 1541984,
          "postDate": "2021-10-12T03:56:50.757Z",
          "content": "<p>LB is the public score</p>",
          "rawMarkdown": "LB is the public score"
        },
        {
          "id": 1543163,
          "postDate": "2021-10-13T08:03:37.313Z",
          "content": "<p>Oh! Thank you!!! BTW I finally submitted!! Thank you for your advice!!<br>\n😆😆</p>",
          "rawMarkdown": "Oh! Thank you!!! BTW I finally submitted!! Thank you for your advice!!\n😆😆"
        },
        {
          "id": 1543200,
          "postDate": "2021-10-13T08:37:10.183Z",
          "content": "<p>good job!</p>",
          "rawMarkdown": "good job!\n"
        }
      ]
    },
    {
      "id": 1540260,
      "postDate": "2021-10-10T10:02:56.543Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1540347,
      "author_name": "tomoo inubushi",
      "author_url": "",
      "post_date": "2021-10-10T12:04:33.763000",
      "content": "<p>I am not sure, but do you consider hidden test set?</p>\n<blockquote>\n  <p>The competition data is defined by three cohorts: Training, Validation (Public), and Testing (Private). The “Training” and the “Validation” cohorts are provided to the participants, whereas the “Testing” cohort is kept hidden at all times, during and after the competition.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/data\" target=\"_blank\">https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/data</a></p>\n<p>You need to create inference notebook with whole processing pipeline like <a href=\"https://www.kaggle.com/ammarnassanalhajali/brain-tumor-3d-inference\" target=\"_blank\">this one</a>.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1541919,
          "author_name": "jypysk",
          "author_url": "",
          "post_date": "2021-10-12T02:44:01.637000",
          "content": "<p>Thank you for advice! Hm… I didn't know theres a hidden test set… Hmm… Actually I followed this notebook <a href=\"https://www.kaggle.com/luongduongminh/brain-tumor-3d\" target=\"_blank\">https://www.kaggle.com/luongduongminh/brain-tumor-3d</a>. But I could not find big differences between this notebook and mine.. But as your comment, I should add whole precessing pipeline in mine! Thank you a lot for kind reply!!☺️</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1542134,
          "author_name": "tomoo inubushi",
          "author_url": "",
          "post_date": "2021-10-12T06:53:37.030000",
          "content": "<p>You should not follow the notebook you linked. This notebook just blend the predictions of other notebooks for public test set and the predictions for hidden (private) test set are all 0.5 (see section [14]). Section [14] prevents scoring error, but its private score is chance level.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1542181,
          "author_name": "NaN",
          "author_url": "",
          "post_date": "2021-10-12T07:56:08.827000",
          "content": "<p><a href=\"https://www.kaggle.com/jypysk\" target=\"_blank\">@jypysk</a> imho, a notebook with more upvotes is usually a better starting point than a notebook with higher public score. The notebook you followed is actually not useful at all. It doesn't show you how to preprocess the data, how to create and train a model, how to generate predictions from the trained model, etc. It literally just takes public test set predictions from other notebooks, blends them together and (maybe luckily?) gets a high score on the public leaderboard (and as <a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a> mentioned, it doesn't predict the private/hidden test set at all and just set all predictions to default value 0.5). The notebook in <a href=\"https://www.kaggle.com/tomooinubushi\" target=\"_blank\">@tomooinubushi</a>'s suggestion may have lower score, but I think it is a much better starting point than the other notebook.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1543165,
          "author_name": "jypysk",
          "author_url": "",
          "post_date": "2021-10-13T08:05:09.073000",
          "content": "<p>Thank you for your kind reply!! Thanks to your help, I finally succeeded to submit! What a wonderful world~!~! Hope you have a great day!!😆</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1540257,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2021-10-10T09:57:01.873000",
      "content": "<p>I think you don't need to do this </p>\n<pre><code>Clell 11: BraTSID = [str(x).zfill(5) for x in BraTSID]\n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 1540264,
          "author_name": "jypysk",
          "author_url": "",
          "post_date": "2021-10-10T10:09:15.830000",
          "content": "<p>Thank you for your comment! But still… not working…\"Submission Scoring Error\". Do I have to load '/kaggle/input/test/' at least one time?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1542173,
      "author_name": "Deepali Mishra",
      "author_url": "",
      "post_date": "2021-10-12T07:47:21.277000",
      "content": "<p>Try to put , index_col=\"BraTS21ID\" while saving the submission file. It works in my case!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1543167,
          "author_name": "jypysk",
          "author_url": "",
          "post_date": "2021-10-13T08:05:40.840000",
          "content": "<p>Thanks! I applied your comment, and succeed to submit~! 😆</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1543223,
          "author_name": "Deepali Mishra",
          "author_url": "",
          "post_date": "2021-10-13T09:01:58.320000",
          "content": "<p>Your welcome!!😃 and you can vote me here , if it's helpful in your case !!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1540601,
      "author_name": "Lyubomir Klyambarski",
      "author_url": "",
      "post_date": "2021-10-10T16:44:26.017000",
      "content": "<p>Try to keep column BraTSID as int or just copy 'sample_submission.csv' and fill your MGMT_values. And yes, consider also adding your whole processing pipeline, as you have been advised already.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1541917,
          "author_name": "jypysk",
          "author_url": "",
          "post_date": "2021-10-12T02:40:07.370000",
          "content": "<p>Oh…. Thank you for advice!! I will pay 100% attention to adding whole processing pipeline on my notebook! 🙏</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1540309,
      "author_name": "Ctrl_CV",
      "author_url": "",
      "post_date": "2021-10-10T11:19:52.083000",
      "content": "<p>May be the train step is hard,so simple ensemble is need for someone who want get high score in LB!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1541920,
          "author_name": "jypysk",
          "author_url": "",
          "post_date": "2021-10-12T02:45:19.137000",
          "content": "<p>Um.. May I ask what \"LB\" is rerferred to?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1541984,
          "author_name": "Ctrl_CV",
          "author_url": "",
          "post_date": "2021-10-12T03:56:50.757000",
          "content": "<p>LB is the public score</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1543163,
          "author_name": "jypysk",
          "author_url": "",
          "post_date": "2021-10-13T08:03:37.313000",
          "content": "<p>Oh! Thank you!!! BTW I finally submitted!! Thank you for your advice!!<br>\n😆😆</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1543200,
          "author_name": "Ctrl_CV",
          "author_url": "",
          "post_date": "2021-10-13T08:37:10.183000",
          "content": "<p>good job!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1540260,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-10-10T10:02:56.543000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1540347": "I am not sure, but do you consider hidden test set?\n\n>  The competition data is defined by three cohorts: Training, Validation (Public), and Testing (Private). The “Training” and the “Validation” cohorts are provided to the participants, whereas the “Testing” cohort is kept hidden at all times, during and after the competition.\n\nhttps://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/data\n\nYou need to create inference notebook with whole processing pipeline like [this one](https://www.kaggle.com/ammarnassanalhajali/brain-tumor-3d-inference).\n",
    "1540257": "I think you don't need to do this \n\n```\nClell 11: BraTSID = [str(x).zfill(5) for x in BraTSID]\n```",
    "1540226": "Good morning to all kaggler! \n\nIm a very beginner to this kaggle competition, and facing a \"Submission Scoring Error\" now.\n\nhttps://www.kaggle.com/jypysk/nokfold-test-submission-v0?scriptVersionId=76747394\n^ Here is my tiny and simple kaggle notebook for submission. \n\nSo I've check, https://www.kaggle.com/code-competition-debugging < here, \n1. all the dataset should be publicly opened.\n2. Should have 87 rows and 2 cols. (Checked the name of each column)\n3. Should name the csv file as \"submission.csv\" & index=false.\n\nCan anyone help me with this problem..?\nThank you for any help you can offer me. 🤒",
    "1542173": "Try to put , index_col=\"BraTS21ID\" while saving the submission file. It works in my case!",
    "1540601": "Try to keep column BraTSID as int or just copy 'sample_submission.csv' and fill your MGMT_values. And yes, consider also adding your whole processing pipeline, as you have been advised already.",
    "1540309": "May be the train step is hard,so simple ensemble is need for someone who want get high score in LB!",
    "1540260": ""
  }
}