{
  "id": 473373,
  "title": "Submission Scoring Error",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/473373",
  "author_name": "Ez",
  "post_date": "2024-02-04T14:37:44.847000",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I was getting a submission scoring error and some of the other participants are getting the same error. I am writing this notebook to give some tips on how to solve this error.</p>\n<p>First, make sure that you are following the competition's <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/overview/code-requirements\" target=\"_blank\">code requirements</a> and <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/overview/evaluation\" target=\"_blank\">evaluation</a> rules:</p>\n<ul>\n<li>Prediction probabilities for each row must sum to 1</li>\n<li>Submission file must be named <code>submission.csv</code></li>\n<li>Internet access disabled</li>\n<li>Freely &amp; publicly available external data is allowed, including pre-trained models</li>\n<li>CPU Notebook &lt;= 9 hours run-time</li>\n<li>GPU Notebook &lt;= 9 hours run-time</li>\n</ul>\n<p>Still getting the error? The test data provided is only one row data. But the competion uses hidden test data to score the submissions. So, your code should be general enough to handle any other test data they may provide.</p>\n<p>For example, the following code snippet is only considering the given test data and it is not able to handle other test data.</p>\n<pre><code>...\n\nspectrogram_id, eeg_id, patient_id = ,,\ny_pred = predict(eeg_id, spectrogram_id, patient_id)\n\ny_pred = pd.DataFrame() \nsubmission = pd.DataFrame(y_pred, columns=[, , , , , ]) \nsubmission[] = eeg_id\nsubmission.to_csv(, index=)\n</code></pre>\n<p>In the other hand, the following code snippet is general enough to handle any other test data they may provide.</p>\n<pre><code>...\n\ntest = pd.read_csv()\ny_pred = model.predict(test) \nsubmission = pd.DataFrame(y_pred, columns=[, , , , , ])\nsubmission[] = test[]  \nsubmission.to_csv(, index=)\n</code></pre>",
  "messages": [
    {
      "id": 2635651,
      "postDate": "2024-02-04T14:37:44.847Z",
      "content": "<p>I was getting a submission scoring error and some of the other participants are getting the same error. I am writing this notebook to give some tips on how to solve this error.</p>\n<p>First, make sure that you are following the competition's <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/overview/code-requirements\" target=\"_blank\">code requirements</a> and <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/overview/evaluation\" target=\"_blank\">evaluation</a> rules:</p>\n<ul>\n<li>Prediction probabilities for each row must sum to 1</li>\n<li>Submission file must be named <code>submission.csv</code></li>\n<li>Internet access disabled</li>\n<li>Freely &amp; publicly available external data is allowed, including pre-trained models</li>\n<li>CPU Notebook &lt;= 9 hours run-time</li>\n<li>GPU Notebook &lt;= 9 hours run-time</li>\n</ul>\n<p>Still getting the error? The test data provided is only one row data. But the competion uses hidden test data to score the submissions. So, your code should be general enough to handle any other test data they may provide.</p>\n<p>For example, the following code snippet is only considering the given test data and it is not able to handle other test data.</p>\n<pre><code>...\n\nspectrogram_id, eeg_id, patient_id = ,,\ny_pred = predict(eeg_id, spectrogram_id, patient_id)\n\ny_pred = pd.DataFrame() \nsubmission = pd.DataFrame(y_pred, columns=[, , , , , ]) \nsubmission[] = eeg_id\nsubmission.to_csv(, index=)\n</code></pre>\n<p>In the other hand, the following code snippet is general enough to handle any other test data they may provide.</p>\n<pre><code>...\n\ntest = pd.read_csv()\ny_pred = model.predict(test) \nsubmission = pd.DataFrame(y_pred, columns=[, , , , , ])\nsubmission[] = test[]  \nsubmission.to_csv(, index=)\n</code></pre>",
      "rawMarkdown": "I was getting a submission scoring error and some of the other participants are getting the same error. I am writing this notebook to give some tips on how to solve this error.\n\nFirst, make sure that you are following the competition's [code requirements](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/overview/code-requirements) and [evaluation](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/overview/evaluation) rules:\n\n-   Prediction probabilities for each row must sum to 1\n-   Submission file must be named `submission.csv`\n-   Internet access disabled\n-   Freely & publicly available external data is allowed, including pre-trained models\n-   CPU Notebook <= 9 hours run-time\n-   GPU Notebook <= 9 hours run-time\n\nStill getting the error? The test data provided is only one row data. But the competion uses hidden test data to score the submissions. So, your code should be general enough to handle any other test data they may provide.\n\nFor example, the following code snippet is only considering the given test data and it is not able to handle other test data.\n\n```py\n...\n\nspectrogram_id, eeg_id, patient_id = 853520,3911565283,6885\ny_pred = predict(eeg_id, spectrogram_id, patient_id)\n\ny_pred = pd.DataFrame() # returns a (len(test), 6) numpy array \nsubmission = pd.DataFrame(y_pred, columns=['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']) \nsubmission['eeg_id'] = eeg_id\nsubmission.to_csv('submission.csv', index=False)\n```\n\nIn the other hand, the following code snippet is general enough to handle any other test data they may provide.\n\n```py\n...\n\ntest = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/test.csv')\ny_pred = model.predict(test) # returns a (len(test), 6) numpy array  \nsubmission = pd.DataFrame(y_pred, columns=['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote'])\nsubmission['eeg_id'] = test['eeg_id']  \nsubmission.to_csv('submission.csv', index=False)\n```",
      "votes": 3
    },
    {
      "id": 2641070,
      "postDate": "2024-02-07T09:02:34.680Z",
      "content": "<p>Have you solved this problem? I manually set the score to 0.166 0.166 0.166 0.166 0.166 0.170 and successfully got a public score. But when I switched to predictive output, it failed.</p>",
      "rawMarkdown": "Have you solved this problem? I manually set the score to 0.166 0.166 0.166 0.166 0.166 0.170 and successfully got a public score. But when I switched to predictive output, it failed."
    },
    {
      "id": 2724976,
      "postDate": "2024-03-31T08:31:49.680Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2641070,
      "author_name": "zznznb",
      "author_url": "",
      "post_date": "2024-02-07T09:02:34.680000",
      "content": "<p>Have you solved this problem? I manually set the score to 0.166 0.166 0.166 0.166 0.166 0.170 and successfully got a public score. But when I switched to predictive output, it failed.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2724976,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-03-31T08:31:49.680000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2635651": "I was getting a submission scoring error and some of the other participants are getting the same error. I am writing this notebook to give some tips on how to solve this error.\n\nFirst, make sure that you are following the competition's [code requirements](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/overview/code-requirements) and [evaluation](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/overview/evaluation) rules:\n\n-   Prediction probabilities for each row must sum to 1\n-   Submission file must be named `submission.csv`\n-   Internet access disabled\n-   Freely & publicly available external data is allowed, including pre-trained models\n-   CPU Notebook <= 9 hours run-time\n-   GPU Notebook <= 9 hours run-time\n\nStill getting the error? The test data provided is only one row data. But the competion uses hidden test data to score the submissions. So, your code should be general enough to handle any other test data they may provide.\n\nFor example, the following code snippet is only considering the given test data and it is not able to handle other test data.\n\n```py\n...\n\nspectrogram_id, eeg_id, patient_id = 853520,3911565283,6885\ny_pred = predict(eeg_id, spectrogram_id, patient_id)\n\ny_pred = pd.DataFrame() # returns a (len(test), 6) numpy array \nsubmission = pd.DataFrame(y_pred, columns=['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']) \nsubmission['eeg_id'] = eeg_id\nsubmission.to_csv('submission.csv', index=False)\n```\n\nIn the other hand, the following code snippet is general enough to handle any other test data they may provide.\n\n```py\n...\n\ntest = pd.read_csv('/kaggle/input/hms-harmful-brain-activity-classification/test.csv')\ny_pred = model.predict(test) # returns a (len(test), 6) numpy array  \nsubmission = pd.DataFrame(y_pred, columns=['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote'])\nsubmission['eeg_id'] = test['eeg_id']  \nsubmission.to_csv('submission.csv', index=False)\n```",
    "2641070": "Have you solved this problem? I manually set the score to 0.166 0.166 0.166 0.166 0.166 0.170 and successfully got a public score. But when I switched to predictive output, it failed.",
    "2724976": ""
  }
}