{
  "id": 552479,
  "title": "Debugging submission error: Out of memory",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/552479",
  "author_name": "",
  "post_date": "2024-12-20T00:29:26.841202800Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>My submission code is running out of memory. I use batch_size of 1, and it runs on the sample test set no problems. Does kaggle provide any other logs on the submission? All I have is an \"Out of memory\" error and I don't know where in my code this occurs.</p>\n<p>I've done all the basics, reduced batch size, preds from the gpu to cpu after prediction etc.</p>\n<p>Debugging the submission run is very difficult because the logs seem to be hidden (which makes sense given the test data must remain private).</p>",
  "messages": [
    {
      "id": "3076390",
      "postDate": "12/20/2024 00:29:26",
      "content": "<p>My submission code is running out of memory. I use batch_size of 1, and it runs on the sample test set no problems. Does kaggle provide any other logs on the submission? All I have is an \"Out of memory\" error and I don't know where in my code this occurs.</p>\n<p>I've done all the basics, reduced batch size, preds from the gpu to cpu after prediction etc.</p>\n<p>Debugging the submission run is very difficult because the logs seem to be hidden (which makes sense given the test data must remain private).</p>",
      "rawMarkdown": "My submission code is running out of memory. I use batch_size of 1, and it runs on the sample test set no problems. Does kaggle provide any other logs on the submission? All I have is an \"Out of memory\" error and I don't know where in my code this occurs.\n\nI've done all the basics, reduced batch size, preds from the gpu to cpu after prediction etc.\n\nDebugging the submission run is very difficult because the logs seem to be hidden (which makes sense given the test data must remain private).",
      "votes": null
    },
    {
      "id": "3076477",
      "postDate": "12/20/2024 01:39:42",
      "content": "<p>The primary difference between the test set and the hidden set is the number of samples, namely 3 vs. ~500.  Is there any place in your code where you're not releasing memory between samples?</p>",
      "rawMarkdown": "The primary difference between the test set and the hidden set is the number of samples, namely 3 vs. ~500.  Is there any place in your code where you're not releasing memory between samples?",
      "votes": null
    },
    {
      "id": "3076810",
      "postDate": "12/20/2024 09:32:30",
      "content": "<p>Print the memory usage before and after the inference, including gpu memory, then you may find some clues. There should be a memory leakage?</p>",
      "rawMarkdown": "Print the memory usage before and after the inference, including gpu memory, then you may find some clues. There should be a memory leakage?",
      "votes": null
    },
    {
      "id": "3083716",
      "postDate": "12/29/2024 23:30:28",
      "content": "<p>Managed to fix the \"Out of Memory\" error with smarter reading in/out of data when inferencing.</p>\n<p>Where can you find that there are 500 samples in the test set?</p>\n<p>I also have another question, if the public leaderboard is calculated on 24% of data and your submission notebook runs for 10 hours - does this mean it will fail the competition rules on the private leaderboard with 76% of the data? Extrapolating run times it would take 31 hours for the private leaderboard to run, thus failing the &lt;= 12 hours runtime.</p>",
      "rawMarkdown": "Managed to fix the \"Out of Memory\" error with smarter reading in/out of data when inferencing.\n\nWhere can you find that there are 500 samples in the test set?\n\nI also have another question, if the public leaderboard is calculated on 24% of data and your submission notebook runs for 10 hours - does this mean it will fail the competition rules on the private leaderboard with 76% of the data? Extrapolating run times it would take 31 hours for the private leaderboard to run, thus failing the <= 12 hours runtime.",
      "votes": null
    },
    {
      "id": "3083735",
      "postDate": "12/30/2024 01:08:05",
      "content": "<blockquote>\n  <p>Where can you find that there are 500 samples in the test set?</p>\n</blockquote>\n<p>Various places, but <a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> specifically confirms it here:<br>\n<a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544702\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544702</a></p>\n<blockquote>\n  <p>I also have another question, if the public leaderboard is calculated on 24% of data and your submission notebook runs for 10 hours - does this mean it will fail the competition rules on the private leaderboard with 76% of the data? Extrapolating run times it would take 31 hours for the private leaderboard to run, thus failing the &lt;= 12 hours runtime.</p>\n</blockquote>\n<p>When you submit, they run your notebook against all of the data, but only report on the 24% for now.  There are a few discussion threads on this as well.  For instance:<br>\n<a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/551447\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/551447</a></p>",
      "rawMarkdown": ">Where can you find that there are 500 samples in the test set?\n\nVarious places, but @kharrington specifically confirms it here:\n[https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544702](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544702)\n\n>I also have another question, if the public leaderboard is calculated on 24% of data and your submission notebook runs for 10 hours - does this mean it will fail the competition rules on the private leaderboard with 76% of the data? Extrapolating run times it would take 31 hours for the private leaderboard to run, thus failing the <= 12 hours runtime.\n\nWhen you submit, they run your notebook against all of the data, but only report on the 24% for now.  There are a few discussion threads on this as well.  For instance:\n[https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/551447](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/551447)",
      "votes": null
    },
    {
      "id": "3084074",
      "postDate": "12/30/2024 11:33:30",
      "content": "<p>Thanks re 500 samples.</p>\n<p>Shortly after I asked that question, I figured the only way would be they score everything and reveal 24% of data - silly question really. But thanks for clarifying anyway :)</p>",
      "rawMarkdown": "Thanks re 500 samples.\n\nShortly after I asked that question, I figured the only way would be they score everything and reveal 24% of data - silly question really. But thanks for clarifying anyway :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3076477,
      "author_name": "davidlist",
      "author_url": "",
      "post_date": "12/20/2024 01:39:42",
      "content": "<p>The primary difference between the test set and the hidden set is the number of samples, namely 3 vs. ~500.  Is there any place in your code where you're not releasing memory between samples?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3083716,
          "author_name": "homiecal",
          "author_url": "",
          "post_date": "12/29/2024 23:30:28",
          "content": "<p>Managed to fix the \"Out of Memory\" error with smarter reading in/out of data when inferencing.</p>\n<p>Where can you find that there are 500 samples in the test set?</p>\n<p>I also have another question, if the public leaderboard is calculated on 24% of data and your submission notebook runs for 10 hours - does this mean it will fail the competition rules on the private leaderboard with 76% of the data? Extrapolating run times it would take 31 hours for the private leaderboard to run, thus failing the &lt;= 12 hours runtime.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3083735,
              "author_name": "davidlist",
              "author_url": "",
              "post_date": "12/30/2024 01:08:05",
              "content": "<blockquote>\n  <p>Where can you find that there are 500 samples in the test set?</p>\n</blockquote>\n<p>Various places, but <a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> specifically confirms it here:<br>\n<a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544702\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544702</a></p>\n<blockquote>\n  <p>I also have another question, if the public leaderboard is calculated on 24% of data and your submission notebook runs for 10 hours - does this mean it will fail the competition rules on the private leaderboard with 76% of the data? Extrapolating run times it would take 31 hours for the private leaderboard to run, thus failing the &lt;= 12 hours runtime.</p>\n</blockquote>\n<p>When you submit, they run your notebook against all of the data, but only report on the 24% for now.  There are a few discussion threads on this as well.  For instance:<br>\n<a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/551447\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/551447</a></p>",
              "votes": null,
              "replies": [
                {
                  "id": 3084074,
                  "author_name": "homiecal",
                  "author_url": "",
                  "post_date": "12/30/2024 11:33:30",
                  "content": "<p>Thanks re 500 samples.</p>\n<p>Shortly after I asked that question, I figured the only way would be they score everything and reveal 24% of data - silly question really. But thanks for clarifying anyway :)</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3076810,
      "author_name": "yksinyoung",
      "author_url": "",
      "post_date": "12/20/2024 09:32:30",
      "content": "<p>Print the memory usage before and after the inference, including gpu memory, then you may find some clues. There should be a memory leakage?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3076390": "My submission code is running out of memory. I use batch_size of 1, and it runs on the sample test set no problems. Does kaggle provide any other logs on the submission? All I have is an \"Out of memory\" error and I don't know where in my code this occurs.\n\nI've done all the basics, reduced batch size, preds from the gpu to cpu after prediction etc.\n\nDebugging the submission run is very difficult because the logs seem to be hidden (which makes sense given the test data must remain private).",
    "3076477": "The primary difference between the test set and the hidden set is the number of samples, namely 3 vs. ~500.  Is there any place in your code where you're not releasing memory between samples?",
    "3076810": "Print the memory usage before and after the inference, including gpu memory, then you may find some clues. There should be a memory leakage?",
    "3083716": "Managed to fix the \"Out of Memory\" error with smarter reading in/out of data when inferencing.\n\nWhere can you find that there are 500 samples in the test set?\n\nI also have another question, if the public leaderboard is calculated on 24% of data and your submission notebook runs for 10 hours - does this mean it will fail the competition rules on the private leaderboard with 76% of the data? Extrapolating run times it would take 31 hours for the private leaderboard to run, thus failing the <= 12 hours runtime.",
    "3083735": ">Where can you find that there are 500 samples in the test set?\n\nVarious places, but @kharrington specifically confirms it here:\n[https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544702](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544702)\n\n>I also have another question, if the public leaderboard is calculated on 24% of data and your submission notebook runs for 10 hours - does this mean it will fail the competition rules on the private leaderboard with 76% of the data? Extrapolating run times it would take 31 hours for the private leaderboard to run, thus failing the <= 12 hours runtime.\n\nWhen you submit, they run your notebook against all of the data, but only report on the 24% for now.  There are a few discussion threads on this as well.  For instance:\n[https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/551447](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/551447)",
    "3084074": "Thanks re 500 samples.\n\nShortly after I asked that question, I figured the only way would be they score everything and reveal 24% of data - silly question really. But thanks for clarifying anyway :)"
  },
  "source": "meta"
}