{
  "id": 470772,
  "title": "Help with ensemble",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/470772",
  "author_name": "",
  "post_date": "2024-01-25T14:39:24.468485800Z",
  "votes": 10,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I have two different models in two different notebooks. Running each notebook separately works fine, and the submissions are completed without errors.</p>\n<p>Now, I'm trying to ensemble both models in a new notebook. I'm using the simplest approach: running each model separately, obtaining the predictions, and then blending them.</p>\n<p>However, I keep encountering a <strong>'Notebook Threw Exception'</strong> error. I've already wasted four submissions today and can't determine the cause of this issue.</p>\n<p>My primary theory is that the error is due to the sum of the probabilities in the final submission file not adding up to one for each row (although the first row does), so I wrote this code in the hope of ensuring that the submission will follow this rule:</p>\n<pre><code> index, row  submission.iterrows():\n    row_sum = row[target_votes].()\n     row_sum != :\n        submission.loc[index, target_votes] = row[target_votes] / row_sum\n</code></pre>\n<p>Despite this update, I'm still facing the same error. Can anyone assist me with this? Are there other areas in the notebook that I might be overlooking?</p>",
  "messages": [
    {
      "id": "2619557",
      "postDate": "01/25/2024 14:39:24",
      "content": "<p>I have two different models in two different notebooks. Running each notebook separately works fine, and the submissions are completed without errors.</p>\n<p>Now, I'm trying to ensemble both models in a new notebook. I'm using the simplest approach: running each model separately, obtaining the predictions, and then blending them.</p>\n<p>However, I keep encountering a <strong>'Notebook Threw Exception'</strong> error. I've already wasted four submissions today and can't determine the cause of this issue.</p>\n<p>My primary theory is that the error is due to the sum of the probabilities in the final submission file not adding up to one for each row (although the first row does), so I wrote this code in the hope of ensuring that the submission will follow this rule:</p>\n<pre><code> index, row  submission.iterrows():\n    row_sum = row[target_votes].()\n     row_sum != :\n        submission.loc[index, target_votes] = row[target_votes] / row_sum\n</code></pre>\n<p>Despite this update, I'm still facing the same error. Can anyone assist me with this? Are there other areas in the notebook that I might be overlooking?</p>",
      "rawMarkdown": "I have two different models in two different notebooks. Running each notebook separately works fine, and the submissions are completed without errors.\n\nNow, I'm trying to ensemble both models in a new notebook. I'm using the simplest approach: running each model separately, obtaining the predictions, and then blending them.\n\nHowever, I keep encountering a **'Notebook Threw Exception'** error. I've already wasted four submissions today and can't determine the cause of this issue.\n\nMy primary theory is that the error is due to the sum of the probabilities in the final submission file not adding up to one for each row (although the first row does), so I wrote this code in the hope of ensuring that the submission will follow this rule:\n\n```python\nfor index, row in submission.iterrows():\n    row_sum = row[target_votes].sum()\n    if row_sum != 0:\n        submission.loc[index, target_votes] = row[target_votes] / row_sum\n```\n\nDespite this update, I'm still facing the same error. Can anyone assist me with this? Are there other areas in the notebook that I might be overlooking?",
      "votes": null
    },
    {
      "id": "2619608",
      "postDate": "01/25/2024 15:10:24",
      "content": "<p>When you add the two submission files, do you make sure that each has the same row order for column <code>eeg_id</code> ?</p>\n<p>Are you facing memory errors? Are the subs using too much memory together?</p>\n<p>As a debugging tool, try inferring 10_000 train eeg and create a submission from that. Then inspect the sub file.</p>",
      "rawMarkdown": "When you add the two submission files, do you make sure that each has the same row order for column `eeg_id` ?\n\nAre you facing memory errors? Are the subs using too much memory together?\n\nAs a debugging tool, try inferring 10_000 train eeg and create a submission from that. Then inspect the sub file.",
      "votes": null
    },
    {
      "id": "2619644",
      "postDate": "01/25/2024 15:36:26",
      "content": "<p>You could see an example of such here! <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/470634\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/470634</a> </p>\n<p>2 problems I had were that the row should always sum to 1 and that my predictions were not averaged in the loop. This code corrects both of those issues for your example </p>",
      "rawMarkdown": "You could see an example of such here! https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/470634 \n\n2 problems I had were that the row should always sum to 1 and that my predictions were not averaged in the loop. This code corrects both of those issues for your example",
      "votes": null
    },
    {
      "id": "2619705",
      "postDate": "01/25/2024 16:25:31",
      "content": "<p>hi <a href=\"https://www.kaggle.com/yantxx\" target=\"_blank\">@yantxx</a>,<br>\nIf it is possible can you please provide more context around the issue like the exact error screenshot and the code snippet with details of variable target_votes dtype and all.</p>\n<ol>\n<li>From the description you have provided, it is entirely possible that submission csv file is getting created or saved incorrectly. </li>\n<li>Submission file value assignment - Vector dimension can be the issue. Check if you need to flatten.<br>\n3.You can put try catch block which can provide more error details rather generic. Or at least print statements.</li>\n</ol>\n<p>Hope this provide some insight.</p>\n<p>thanks!</p>",
      "rawMarkdown": "hi @yantxx,\nIf it is possible can you please provide more context around the issue like the exact error screenshot and the code snippet with details of variable target_votes dtype and all.\n\n1. From the description you have provided, it is entirely possible that submission csv file is getting created or saved incorrectly. \n2. Submission file value assignment - Vector dimension can be the issue. Check if you need to flatten.\n3.You can put try catch block which can provide more error details rather generic. Or at least print statements.\n\nHope this provide some insight.\n\nthanks!",
      "votes": null
    },
    {
      "id": "2619722",
      "postDate": "01/25/2024 16:36:08",
      "content": "<p><a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> Thanks, I actually used your code in my first attempt, but it didn't work for me.</p>\n<p>Also, I'm curious about why it worked in your notebook. If I'm not mistaken, this code here -&gt;</p>\n<pre><code>submission=pd.read_csv()\nlabels=[,,,,,]\n i  ((labels)):\n    submission[]=(test_preds[:,i] + preds_combine[:, i])/\nsubmission.to_csv(,index=)\ndisplay(submission.head())\n</code></pre>\n<p>does not guarantee that each row will have a sum of probabilities equal to 1, right? When we average the probabilities from two different models, the resulting averages may not perfectly align to sum up to 1, because the individual models might have different probability distributions.</p>\n<p>Sorry if I'm missing something obvious here.</p>",
      "rawMarkdown": "cody11null Thanks, I actually used your code in my first attempt, but it didn't work for me.\n\nAlso, I'm curious about why it worked in your notebook. If I'm not mistaken, this code here ->\n\n```python\nsubmission=pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/sample_submission.csv\")\nlabels=['seizure','lpd','gpd','lrda','grda','other']\nfor i in range(len(labels)):\n    submission[f'{labels[i]}_vote']=(test_preds[:,i] + preds_combine[:, i])/2\nsubmission.to_csv(\"submission.csv\",index=None)\ndisplay(submission.head())\n```\n\ndoes not guarantee that each row will have a sum of probabilities equal to 1, right? When we average the probabilities from two different models, the resulting averages may not perfectly align to sum up to 1, because the individual models might have different probability distributions.\n\nSorry if I'm missing something obvious here.",
      "votes": null
    },
    {
      "id": "2619731",
      "postDate": "01/25/2024 16:48:49",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<p><strong><code>When you add the two submission files, do you make sure that each has the same row order for column eeg_id ?</code></strong><br>\nthis is the last line of the notebook</p>\n<pre><code>\npreds_model2 = preds_model2 .set_index().loc[submission[]].reset_index()\npreds_model1 = preds_model1 .set_index().loc[submission[]].reset_index()\n\n\ncolumns_to_average = [, , , , , ]\n\n column  columns_to_average:\n    submission[column] = (preds_model2 [column] + preds_model1 [column]) / \n\n\n index, row  submission.iterrows():\n    row_sum = row[columns_to_average].()\n     row_sum != :\n        submission.loc[index, columns_to_average] = row[columns_to_average] / row_sum\n\nsubmission.to_csv(,index=)\n</code></pre>\n<p>I was confident that the row order would stay the same.</p>\n<p><strong><code>Are you facing memory errors?</code></strong><br>\nNo! The notebook runs smoothly</p>\n<p><strong><code>Are the subs using too much memory together?</code></strong><br>\nSorry if this is a basic question, but how can I check this? The 'output' size of the notebook does not exceed the limit.</p>\n<p><strong><code>As a debugging tool, try inferring 10_000 train eeg and create a submission from that. Then inspect the sub file.</code></strong><br>\nGreat idea!</p>",
      "rawMarkdown": "cdeotte \n\n**`When you add the two submission files, do you make sure that each has the same row order for column eeg_id ?`**\nthis is the last line of the notebook\n\n```python\n# ensure the order of eeg_id\npreds_model2 = preds_model2 .set_index('eeg_id').loc[submission['eeg_id']].reset_index()\npreds_model1 = preds_model1 .set_index('eeg_id').loc[submission['eeg_id']].reset_index()\n\n# blend\ncolumns_to_average = ['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']\n\nfor column in columns_to_average:\n    submission[column] = (preds_model2 [column] + preds_model1 [column]) / 2\n\n# ensure sum to 1\nfor index, row in submission.iterrows():\n    row_sum = row[columns_to_average].sum()\n    if row_sum != 0:\n        submission.loc[index, columns_to_average] = row[columns_to_average] / row_sum\n\nsubmission.to_csv(\"submission.csv\",index=None)\n```\nI was confident that the row order would stay the same.\n\n\n**`Are you facing memory errors?`**\nNo! The notebook runs smoothly\n\n**`Are the subs using too much memory together?`**\nSorry if this is a basic question, but how can I check this? The 'output' size of the notebook does not exceed the limit.\n\n**`As a debugging tool, try inferring 10_000 train eeg and create a submission from that. Then inspect the sub file.`**\nGreat idea!",
      "votes": null
    },
    {
      "id": "2619740",
      "postDate": "01/25/2024 16:54:11",
      "content": "<blockquote>\n  <p>No! The notebook runs smoothly</p>\n</blockquote>\n<p>How do we know this? When we run the notebook the size of test data is only 1 row. When we submit, the size of test data is hundreds or thousands of rows. Therefore submit will use more memory than commit and we cannot know how much. The best way to test this is to infer 10_000 train rows locally.</p>",
      "rawMarkdown": ">No! The notebook runs smoothly\n\nHow do we know this? When we run the notebook the size of test data is only 1 row. When we submit, the size of test data is hundreds or thousands of rows. Therefore submit will use more memory than commit and we cannot know how much. The best way to test this is to infer 10_000 train rows locally.",
      "votes": null
    },
    {
      "id": "2619742",
      "postDate": "01/25/2024 16:55:39",
      "content": "<p>Great insight, Chris. Thank you! </p>",
      "rawMarkdown": "Great insight, Chris. Thank you!",
      "votes": null
    },
    {
      "id": "2619757",
      "postDate": "01/25/2024 17:18:22",
      "content": "<p>I do believe you are correct, I had just noticed it after the submission had worked so I didn’t worry about it. If you look through the problem was for me was in saving the predictions from previous models. The best example would be the preds combine feature from model 1</p>",
      "rawMarkdown": "I do believe you are correct, I had just noticed it after the submission had worked so I didn’t worry about it. If you look through the problem was for me was in saving the predictions from previous models. The best example would be the preds combine feature from model 1",
      "votes": null
    },
    {
      "id": "2683352",
      "postDate": "03/06/2024 00:08:46",
      "content": "<p>My be cuda out of memory.</p>",
      "rawMarkdown": "My be cuda out of memory.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2619608,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "01/25/2024 15:10:24",
      "content": "<p>When you add the two submission files, do you make sure that each has the same row order for column <code>eeg_id</code> ?</p>\n<p>Are you facing memory errors? Are the subs using too much memory together?</p>\n<p>As a debugging tool, try inferring 10_000 train eeg and create a submission from that. Then inspect the sub file.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2619731,
          "author_name": "yantxx",
          "author_url": "",
          "post_date": "01/25/2024 16:48:49",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>\n<p><strong><code>When you add the two submission files, do you make sure that each has the same row order for column eeg_id ?</code></strong><br>\nthis is the last line of the notebook</p>\n<pre><code>\npreds_model2 = preds_model2 .set_index().loc[submission[]].reset_index()\npreds_model1 = preds_model1 .set_index().loc[submission[]].reset_index()\n\n\ncolumns_to_average = [, , , , , ]\n\n column  columns_to_average:\n    submission[column] = (preds_model2 [column] + preds_model1 [column]) / \n\n\n index, row  submission.iterrows():\n    row_sum = row[columns_to_average].()\n     row_sum != :\n        submission.loc[index, columns_to_average] = row[columns_to_average] / row_sum\n\nsubmission.to_csv(,index=)\n</code></pre>\n<p>I was confident that the row order would stay the same.</p>\n<p><strong><code>Are you facing memory errors?</code></strong><br>\nNo! The notebook runs smoothly</p>\n<p><strong><code>Are the subs using too much memory together?</code></strong><br>\nSorry if this is a basic question, but how can I check this? The 'output' size of the notebook does not exceed the limit.</p>\n<p><strong><code>As a debugging tool, try inferring 10_000 train eeg and create a submission from that. Then inspect the sub file.</code></strong><br>\nGreat idea!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2619740,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "01/25/2024 16:54:11",
              "content": "<blockquote>\n  <p>No! The notebook runs smoothly</p>\n</blockquote>\n<p>How do we know this? When we run the notebook the size of test data is only 1 row. When we submit, the size of test data is hundreds or thousands of rows. Therefore submit will use more memory than commit and we cannot know how much. The best way to test this is to infer 10_000 train rows locally.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2619742,
                  "author_name": "yantxx",
                  "author_url": "",
                  "post_date": "01/25/2024 16:55:39",
                  "content": "<p>Great insight, Chris. Thank you! </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2619644,
      "author_name": "cody11null",
      "author_url": "",
      "post_date": "01/25/2024 15:36:26",
      "content": "<p>You could see an example of such here! <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/470634\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/470634</a> </p>\n<p>2 problems I had were that the row should always sum to 1 and that my predictions were not averaged in the loop. This code corrects both of those issues for your example </p>",
      "votes": null,
      "replies": [
        {
          "id": 2619722,
          "author_name": "yantxx",
          "author_url": "",
          "post_date": "01/25/2024 16:36:08",
          "content": "<p><a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> Thanks, I actually used your code in my first attempt, but it didn't work for me.</p>\n<p>Also, I'm curious about why it worked in your notebook. If I'm not mistaken, this code here -&gt;</p>\n<pre><code>submission=pd.read_csv()\nlabels=[,,,,,]\n i  ((labels)):\n    submission[]=(test_preds[:,i] + preds_combine[:, i])/\nsubmission.to_csv(,index=)\ndisplay(submission.head())\n</code></pre>\n<p>does not guarantee that each row will have a sum of probabilities equal to 1, right? When we average the probabilities from two different models, the resulting averages may not perfectly align to sum up to 1, because the individual models might have different probability distributions.</p>\n<p>Sorry if I'm missing something obvious here.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2619757,
              "author_name": "cody11null",
              "author_url": "",
              "post_date": "01/25/2024 17:18:22",
              "content": "<p>I do believe you are correct, I had just noticed it after the submission had worked so I didn’t worry about it. If you look through the problem was for me was in saving the predictions from previous models. The best example would be the preds combine feature from model 1</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2619705,
      "author_name": "supplejade",
      "author_url": "",
      "post_date": "01/25/2024 16:25:31",
      "content": "<p>hi <a href=\"https://www.kaggle.com/yantxx\" target=\"_blank\">@yantxx</a>,<br>\nIf it is possible can you please provide more context around the issue like the exact error screenshot and the code snippet with details of variable target_votes dtype and all.</p>\n<ol>\n<li>From the description you have provided, it is entirely possible that submission csv file is getting created or saved incorrectly. </li>\n<li>Submission file value assignment - Vector dimension can be the issue. Check if you need to flatten.<br>\n3.You can put try catch block which can provide more error details rather generic. Or at least print statements.</li>\n</ol>\n<p>Hope this provide some insight.</p>\n<p>thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2683352,
      "author_name": "pkyangno1",
      "author_url": "",
      "post_date": "03/06/2024 00:08:46",
      "content": "<p>My be cuda out of memory.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2619557": "I have two different models in two different notebooks. Running each notebook separately works fine, and the submissions are completed without errors.\n\nNow, I'm trying to ensemble both models in a new notebook. I'm using the simplest approach: running each model separately, obtaining the predictions, and then blending them.\n\nHowever, I keep encountering a **'Notebook Threw Exception'** error. I've already wasted four submissions today and can't determine the cause of this issue.\n\nMy primary theory is that the error is due to the sum of the probabilities in the final submission file not adding up to one for each row (although the first row does), so I wrote this code in the hope of ensuring that the submission will follow this rule:\n\n```python\nfor index, row in submission.iterrows():\n    row_sum = row[target_votes].sum()\n    if row_sum != 0:\n        submission.loc[index, target_votes] = row[target_votes] / row_sum\n```\n\nDespite this update, I'm still facing the same error. Can anyone assist me with this? Are there other areas in the notebook that I might be overlooking?",
    "2619608": "When you add the two submission files, do you make sure that each has the same row order for column `eeg_id` ?\n\nAre you facing memory errors? Are the subs using too much memory together?\n\nAs a debugging tool, try inferring 10_000 train eeg and create a submission from that. Then inspect the sub file.",
    "2619644": "You could see an example of such here! https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/470634 \n\n2 problems I had were that the row should always sum to 1 and that my predictions were not averaged in the loop. This code corrects both of those issues for your example",
    "2619705": "hi @yantxx,\nIf it is possible can you please provide more context around the issue like the exact error screenshot and the code snippet with details of variable target_votes dtype and all.\n\n1. From the description you have provided, it is entirely possible that submission csv file is getting created or saved incorrectly. \n2. Submission file value assignment - Vector dimension can be the issue. Check if you need to flatten.\n3.You can put try catch block which can provide more error details rather generic. Or at least print statements.\n\nHope this provide some insight.\n\nthanks!",
    "2619722": "cody11null Thanks, I actually used your code in my first attempt, but it didn't work for me.\n\nAlso, I'm curious about why it worked in your notebook. If I'm not mistaken, this code here ->\n\n```python\nsubmission=pd.read_csv(\"/kaggle/input/hms-harmful-brain-activity-classification/sample_submission.csv\")\nlabels=['seizure','lpd','gpd','lrda','grda','other']\nfor i in range(len(labels)):\n    submission[f'{labels[i]}_vote']=(test_preds[:,i] + preds_combine[:, i])/2\nsubmission.to_csv(\"submission.csv\",index=None)\ndisplay(submission.head())\n```\n\ndoes not guarantee that each row will have a sum of probabilities equal to 1, right? When we average the probabilities from two different models, the resulting averages may not perfectly align to sum up to 1, because the individual models might have different probability distributions.\n\nSorry if I'm missing something obvious here.",
    "2619731": "cdeotte \n\n**`When you add the two submission files, do you make sure that each has the same row order for column eeg_id ?`**\nthis is the last line of the notebook\n\n```python\n# ensure the order of eeg_id\npreds_model2 = preds_model2 .set_index('eeg_id').loc[submission['eeg_id']].reset_index()\npreds_model1 = preds_model1 .set_index('eeg_id').loc[submission['eeg_id']].reset_index()\n\n# blend\ncolumns_to_average = ['seizure_vote', 'lpd_vote', 'gpd_vote', 'lrda_vote', 'grda_vote', 'other_vote']\n\nfor column in columns_to_average:\n    submission[column] = (preds_model2 [column] + preds_model1 [column]) / 2\n\n# ensure sum to 1\nfor index, row in submission.iterrows():\n    row_sum = row[columns_to_average].sum()\n    if row_sum != 0:\n        submission.loc[index, columns_to_average] = row[columns_to_average] / row_sum\n\nsubmission.to_csv(\"submission.csv\",index=None)\n```\nI was confident that the row order would stay the same.\n\n\n**`Are you facing memory errors?`**\nNo! The notebook runs smoothly\n\n**`Are the subs using too much memory together?`**\nSorry if this is a basic question, but how can I check this? The 'output' size of the notebook does not exceed the limit.\n\n**`As a debugging tool, try inferring 10_000 train eeg and create a submission from that. Then inspect the sub file.`**\nGreat idea!",
    "2619740": ">No! The notebook runs smoothly\n\nHow do we know this? When we run the notebook the size of test data is only 1 row. When we submit, the size of test data is hundreds or thousands of rows. Therefore submit will use more memory than commit and we cannot know how much. The best way to test this is to infer 10_000 train rows locally.",
    "2619742": "Great insight, Chris. Thank you!",
    "2619757": "I do believe you are correct, I had just noticed it after the submission had worked so I didn’t worry about it. If you look through the problem was for me was in saving the predictions from previous models. The best example would be the preds combine feature from model 1",
    "2683352": "My be cuda out of memory."
  },
  "source": "meta"
}