{
  "id": 490665,
  "title": "The easiest way to make a brutal ensemble",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/490665",
  "author_name": "",
  "post_date": "2024-04-03T06:19:27.197791900Z",
  "votes": 10,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hello everybody!<br>\nIn this competition, as in most others, one of the challenges for me was to make the inference work on the test data.<br>\nI understand why kaggle has to restrict access to test data, but it's annoying when you get Notebook Threw Exception and Notebook Out of Memory errors when everything works on sample test data. In this competition, both errors were due to Out of Memory. <br>\nTo avoid this I tried to use some tricks and got a new bug:<br>\nI tried to use create_spectrogram_with_cusignal directly when creating the dataset for tensorflows, but the problem was that it caused some kind of conflict that broke the notebook without error if creating more than one dataset. <br>\nThen I was advised in the comments to run each code separately in separate .py files. It worked.<br>\nSo how it works:</p>\n<p>You have working code, you save it as a .py file:</p>\n<p>%% writefile model1.py<br>\nYour code<br>\n…..<br>\nsubmission.to_csv(\"submission_model1.csv\",index=None)</p>\n<p>then perform<br>\n!python /kaggle/working/model1.py</p>\n<p>I first converted the code into a script in the settings, copied the script and pasted it into one cell, quick and easy.</p>\n<p>And all the code is executed, after execution the memory is freed (By the way, I don't know if it is completely, but I had enough).</p>\n<p>And do this for each model that needs to be done separately.</p>\n<p>It is important that all commands like !pip install and so on are easiest to execute in the main notebook.</p>\n<p>Good luck to everyone and a peaceful sky above your head!</p>",
  "messages": [
    {
      "id": "2732423",
      "postDate": "04/03/2024 06:19:27",
      "content": "<p>Hello everybody!<br>\nIn this competition, as in most others, one of the challenges for me was to make the inference work on the test data.<br>\nI understand why kaggle has to restrict access to test data, but it's annoying when you get Notebook Threw Exception and Notebook Out of Memory errors when everything works on sample test data. In this competition, both errors were due to Out of Memory. <br>\nTo avoid this I tried to use some tricks and got a new bug:<br>\nI tried to use create_spectrogram_with_cusignal directly when creating the dataset for tensorflows, but the problem was that it caused some kind of conflict that broke the notebook without error if creating more than one dataset. <br>\nThen I was advised in the comments to run each code separately in separate .py files. It worked.<br>\nSo how it works:</p>\n<p>You have working code, you save it as a .py file:</p>\n<p>%% writefile model1.py<br>\nYour code<br>\n…..<br>\nsubmission.to_csv(\"submission_model1.csv\",index=None)</p>\n<p>then perform<br>\n!python /kaggle/working/model1.py</p>\n<p>I first converted the code into a script in the settings, copied the script and pasted it into one cell, quick and easy.</p>\n<p>And all the code is executed, after execution the memory is freed (By the way, I don't know if it is completely, but I had enough).</p>\n<p>And do this for each model that needs to be done separately.</p>\n<p>It is important that all commands like !pip install and so on are easiest to execute in the main notebook.</p>\n<p>Good luck to everyone and a peaceful sky above your head!</p>",
      "rawMarkdown": "Hello everybody!\nIn this competition, as in most others, one of the challenges for me was to make the inference work on the test data.\nI understand why kaggle has to restrict access to test data, but it's annoying when you get Notebook Threw Exception and Notebook Out of Memory errors when everything works on sample test data. In this competition, both errors were due to Out of Memory. \nTo avoid this I tried to use some tricks and got a new bug:\nI tried to use create_spectrogram_with_cusignal directly when creating the dataset for tensorflows, but the problem was that it caused some kind of conflict that broke the notebook without error if creating more than one dataset. \nThen I was advised in the comments to run each code separately in separate .py files. It worked.\nSo how it works:\n\nYou have working code, you save it as a .py file:\n\n%% writefile model1.py\nYour code\n.....\nsubmission.to_csv(\"submission_model1.csv\",index=None)\n\nthen perform\n!python /kaggle/working/model1.py\n\nI first converted the code into a script in the settings, copied the script and pasted it into one cell, quick and easy.\n\nAnd all the code is executed, after execution the memory is freed (By the way, I don't know if it is completely, but I had enough).\n\nAnd do this for each model that needs to be done separately.\n\nIt is important that all commands like !pip install and so on are easiest to execute in the main notebook.\n\nGood luck to everyone and a peaceful sky above your head!",
      "votes": null
    },
    {
      "id": "2732692",
      "postDate": "04/03/2024 09:22:52",
      "content": "<p>yes this is how it should be done always, it also makes the code much less cluttered</p>",
      "rawMarkdown": "yes this is how it should be done always, it also makes the code much less cluttered",
      "votes": null
    },
    {
      "id": "2732705",
      "postDate": "04/03/2024 09:29:15",
      "content": "<p>Yes, there is another advantage to this approach if you only want to publish part of your pipeline without the risk of publishing too much:-)</p>",
      "rawMarkdown": "Yes, there is another advantage to this approach if you only want to publish part of your pipeline without the risk of publishing too much:-)",
      "votes": null
    },
    {
      "id": "2732995",
      "postDate": "04/03/2024 13:00:58",
      "content": "<p>This way feels weird and error prone to me. I would still take the time and prepare a clean inference notebook.</p>",
      "rawMarkdown": "This way feels weird and error prone to me. I would still take the time and prepare a clean inference notebook.",
      "votes": null
    },
    {
      "id": "2733023",
      "postDate": "04/03/2024 13:16:56",
      "content": "<p>In fact, this approach helps to avoid random errors in the namespace, and importantly helps to solve the memory problem, when we need to use different approaches for preprocessing, after that it is difficult to completely clear the memory, and this approach just executes the script, after which it should theoretically release all memory</p>",
      "rawMarkdown": "In fact, this approach helps to avoid random errors in the namespace, and importantly helps to solve the memory problem, when we need to use different approaches for preprocessing, after that it is difficult to completely clear the memory, and this approach just executes the script, after which it should theoretically release all memory",
      "votes": null
    },
    {
      "id": "2737222",
      "postDate": "04/05/2024 16:41:46",
      "content": "<p>Your sharing is appreciated!</p>",
      "rawMarkdown": "Your sharing is appreciated!",
      "votes": null
    },
    {
      "id": "2737371",
      "postDate": "04/05/2024 18:25:53",
      "content": "<blockquote>\n  <p>This way feels weird and error prone to me. I would still take the time and prepare a clean inference notebook.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> 🌝</p>",
      "rawMarkdown": "> This way feels weird and error prone to me. I would still take the time and prepare a clean inference notebook.\n\n\n@pheadrus 🌝",
      "votes": null
    },
    {
      "id": "2737382",
      "postDate": "04/05/2024 18:30:03",
      "content": "<p>I'm actually taking my words back. It would be way easier for us to upload scripts, write csv files and blend them lol</p>",
      "rawMarkdown": "I'm actually taking my words back. It would be way easier for us to upload scripts, write csv files and blend them lol",
      "votes": null
    },
    {
      "id": "2737403",
      "postDate": "04/05/2024 18:47:23",
      "content": "<p>That's right, I do it like this:<br>\nsub1=pd.read_csv(\"/kaggle/working/submission_model1.csv\")<br>\nsub2=pd.read_csv(\"/kaggle/working/submission_model2.csv\")<br>\nsub3=pd.read_csv(\"/kaggle/working/submission_model3.csv\")<br>\nsub4=pd.read_csv(\"/kaggle/working/submission_model4.csv\")<br>\nsub1_GPU100=pd.read_csv(\"/kaggle/working/submission_GPU100 .csv\")  <br>\nIf you save the scripts separately, the inference notebook will contain only imports and installation of the necessary libraries, execution commands and the ensemble itself. Beauty)</p>",
      "rawMarkdown": "That's right, I do it like this:\nsub1=pd.read_csv(\"/kaggle/working/submission_model1.csv\")\nsub2=pd.read_csv(\"/kaggle/working/submission_model2.csv\")\nsub3=pd.read_csv(\"/kaggle/working/submission_model3.csv\")\nsub4=pd.read_csv(\"/kaggle/working/submission_model4.csv\")\nsub1_GPU100=pd.read_csv(\"/kaggle/working/submission_GPU100 .csv\")  \nIf you save the scripts separately, the inference notebook will contain only imports and installation of the necessary libraries, execution commands and the ensemble itself. Beauty)",
      "votes": null
    },
    {
      "id": "2737700",
      "postDate": "04/05/2024 22:37:27",
      "content": "<p>This is the best way for me, as it avoid strange OOM issues. Also for this competion, seems we could ensemble lots of models..🥲</p>",
      "rawMarkdown": "This is the best way for me, as it avoid strange OOM issues. Also for this competion, seems we could ensemble lots of models..🥲",
      "votes": null
    },
    {
      "id": "2737717",
      "postDate": "04/05/2024 22:53:10",
      "content": "<p>Instead of using %%writefile cells, I prefer to create a separate 'library' dataset, which contains all python code. So, the main script focuses on installing necessary modules and combining the submissions. This makes it much cleaner and easier to maintain, especially in the case different models need different environments.</p>",
      "rawMarkdown": "Instead of using %%writefile cells, I prefer to create a separate 'library' dataset, which contains all python code. So, the main script focuses on installing necessary modules and combining the submissions. This makes it much cleaner and easier to maintain, especially in the case different models need different environments.",
      "votes": null
    },
    {
      "id": "2739043",
      "postDate": "04/06/2024 19:26:21",
      "content": "<p>Guys, it's better to create scripts in separate notebooks, I combined 6 in one and now the notebook is \"slowing down\".</p>",
      "rawMarkdown": "Guys, it's better to create scripts in separate notebooks, I combined 6 in one and now the notebook is \"slowing down\".",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2732692,
      "author_name": "nikhilmishradev",
      "author_url": "",
      "post_date": "04/03/2024 09:22:52",
      "content": "<p>yes this is how it should be done always, it also makes the code much less cluttered</p>",
      "votes": null,
      "replies": [
        {
          "id": 2732705,
          "author_name": "aikhmelnytskyy",
          "author_url": "",
          "post_date": "04/03/2024 09:29:15",
          "content": "<p>Yes, there is another advantage to this approach if you only want to publish part of your pipeline without the risk of publishing too much:-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2732995,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "04/03/2024 13:00:58",
          "content": "<p>This way feels weird and error prone to me. I would still take the time and prepare a clean inference notebook.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2733023,
              "author_name": "aikhmelnytskyy",
              "author_url": "",
              "post_date": "04/03/2024 13:16:56",
              "content": "<p>In fact, this approach helps to avoid random errors in the namespace, and importantly helps to solve the memory problem, when we need to use different approaches for preprocessing, after that it is difficult to completely clear the memory, and this approach just executes the script, after which it should theoretically release all memory</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2737371,
              "author_name": "benihime91",
              "author_url": "",
              "post_date": "04/05/2024 18:25:53",
              "content": "<blockquote>\n  <p>This way feels weird and error prone to me. I would still take the time and prepare a clean inference notebook.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> 🌝</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2737382,
                  "author_name": "gunesevitan",
                  "author_url": "",
                  "post_date": "04/05/2024 18:30:03",
                  "content": "<p>I'm actually taking my words back. It would be way easier for us to upload scripts, write csv files and blend them lol</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2737403,
                      "author_name": "aikhmelnytskyy",
                      "author_url": "",
                      "post_date": "04/05/2024 18:47:23",
                      "content": "<p>That's right, I do it like this:<br>\nsub1=pd.read_csv(\"/kaggle/working/submission_model1.csv\")<br>\nsub2=pd.read_csv(\"/kaggle/working/submission_model2.csv\")<br>\nsub3=pd.read_csv(\"/kaggle/working/submission_model3.csv\")<br>\nsub4=pd.read_csv(\"/kaggle/working/submission_model4.csv\")<br>\nsub1_GPU100=pd.read_csv(\"/kaggle/working/submission_GPU100 .csv\")  <br>\nIf you save the scripts separately, the inference notebook will contain only imports and installation of the necessary libraries, execution commands and the ensemble itself. Beauty)</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2737222,
      "author_name": "nartaa",
      "author_url": "",
      "post_date": "04/05/2024 16:41:46",
      "content": "<p>Your sharing is appreciated!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2737700,
      "author_name": "goldenlock",
      "author_url": "",
      "post_date": "04/05/2024 22:37:27",
      "content": "<p>This is the best way for me, as it avoid strange OOM issues. Also for this competion, seems we could ensemble lots of models..🥲</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2737717,
      "author_name": "kdmitrie",
      "author_url": "",
      "post_date": "04/05/2024 22:53:10",
      "content": "<p>Instead of using %%writefile cells, I prefer to create a separate 'library' dataset, which contains all python code. So, the main script focuses on installing necessary modules and combining the submissions. This makes it much cleaner and easier to maintain, especially in the case different models need different environments.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2739043,
      "author_name": "aikhmelnytskyy",
      "author_url": "",
      "post_date": "04/06/2024 19:26:21",
      "content": "<p>Guys, it's better to create scripts in separate notebooks, I combined 6 in one and now the notebook is \"slowing down\".</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2732423": "Hello everybody!\nIn this competition, as in most others, one of the challenges for me was to make the inference work on the test data.\nI understand why kaggle has to restrict access to test data, but it's annoying when you get Notebook Threw Exception and Notebook Out of Memory errors when everything works on sample test data. In this competition, both errors were due to Out of Memory. \nTo avoid this I tried to use some tricks and got a new bug:\nI tried to use create_spectrogram_with_cusignal directly when creating the dataset for tensorflows, but the problem was that it caused some kind of conflict that broke the notebook without error if creating more than one dataset. \nThen I was advised in the comments to run each code separately in separate .py files. It worked.\nSo how it works:\n\nYou have working code, you save it as a .py file:\n\n%% writefile model1.py\nYour code\n.....\nsubmission.to_csv(\"submission_model1.csv\",index=None)\n\nthen perform\n!python /kaggle/working/model1.py\n\nI first converted the code into a script in the settings, copied the script and pasted it into one cell, quick and easy.\n\nAnd all the code is executed, after execution the memory is freed (By the way, I don't know if it is completely, but I had enough).\n\nAnd do this for each model that needs to be done separately.\n\nIt is important that all commands like !pip install and so on are easiest to execute in the main notebook.\n\nGood luck to everyone and a peaceful sky above your head!",
    "2732692": "yes this is how it should be done always, it also makes the code much less cluttered",
    "2732705": "Yes, there is another advantage to this approach if you only want to publish part of your pipeline without the risk of publishing too much:-)",
    "2732995": "This way feels weird and error prone to me. I would still take the time and prepare a clean inference notebook.",
    "2733023": "In fact, this approach helps to avoid random errors in the namespace, and importantly helps to solve the memory problem, when we need to use different approaches for preprocessing, after that it is difficult to completely clear the memory, and this approach just executes the script, after which it should theoretically release all memory",
    "2737222": "Your sharing is appreciated!",
    "2737371": "> This way feels weird and error prone to me. I would still take the time and prepare a clean inference notebook.\n\n\n@pheadrus 🌝",
    "2737382": "I'm actually taking my words back. It would be way easier for us to upload scripts, write csv files and blend them lol",
    "2737403": "That's right, I do it like this:\nsub1=pd.read_csv(\"/kaggle/working/submission_model1.csv\")\nsub2=pd.read_csv(\"/kaggle/working/submission_model2.csv\")\nsub3=pd.read_csv(\"/kaggle/working/submission_model3.csv\")\nsub4=pd.read_csv(\"/kaggle/working/submission_model4.csv\")\nsub1_GPU100=pd.read_csv(\"/kaggle/working/submission_GPU100 .csv\")  \nIf you save the scripts separately, the inference notebook will contain only imports and installation of the necessary libraries, execution commands and the ensemble itself. Beauty)",
    "2737700": "This is the best way for me, as it avoid strange OOM issues. Also for this competion, seems we could ensemble lots of models..🥲",
    "2737717": "Instead of using %%writefile cells, I prefer to create a separate 'library' dataset, which contains all python code. So, the main script focuses on installing necessary modules and combining the submissions. This makes it much cleaner and easier to maintain, especially in the case different models need different environments.",
    "2739043": "Guys, it's better to create scripts in separate notebooks, I combined 6 in one and now the notebook is \"slowing down\"."
  },
  "source": "meta"
}