{
  "id": 568144,
  "title": "How to solve \"Notebook out of disk\" error while scoring?",
  "url": "/competitions/stanford-rna-3d-folding/discussion/568144",
  "author_name": "doheon114",
  "post_date": "2025-03-14T03:42:47.040000",
  "votes": 0,
  "comment_count": 8,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15620263%2Fcd011876418ab2ac9a23d20d95c1418c%2F2025-03-14%20%2012.40.46.png?generation=1741923660575625&amp;alt=media\" alt=\"\"><br>\nWhile implementing NuFold for this competition, the notebook runs normally during submission. However, the scoring process gets stuck loading indefinitely and eventually results in a \"Notebook out of disk\" error. How can I resolve this?</p>",
  "messages": [
    {
      "id": 3149377,
      "postDate": "2025-03-14T05:50:12.440Z",
      "content": "<p>i suspect wrong submission format. you do scoring in local.</p>\n<p>metric code is available. try local scoring. for various rna length and monitor disk and ram usage.<br>\nthere is no problem at my side. try degenerate case like inf and zero. i. not sure how nufold handle  cases when it cannot make prediction, please check it.</p>\n<p>make sure no temp files in kaggle working folder. only submission. csv is allowed.<br>\nelse kaggle server cannot find  the csv file(well known kaggle bug)</p>\n<p>you. can do online debug,. eg. submit only those less than 100 in length and set rest to zero.  other conditions are: only first 10, only. casp16 test( which you can check with local implementation), etc … this. can help pin point which test seq got error.</p>\n<p>you can check csv file size, etc or other property  of solution. set different submission on different  condition. eg if file size is large raise error, else submit all zero. these are online debug tricks.</p>",
      "rawMarkdown": "i suspect wrong submission format. you do scoring in local.\n\n\nmetric code is available. try local scoring. for various rna length and monitor disk and ram usage.\nthere is no problem at my side. try degenerate case like inf and zero. i. not sure how nufold handle  cases when it cannot make prediction, please check it.\n\nmake sure no temp files in kaggle working folder. only submission. csv is allowed.\nelse kaggle server cannot find  the csv file(well known kaggle bug)\n\nyou. can do online debug,. eg. submit only those less than 100 in length and set rest to zero.  other conditions are: only first 10, only. casp16 test( which you can check with local implementation), etc ... this. can help pin point which test seq got error.\n\nyou can check csv file size, etc or other property  of solution. set different submission on different  condition. eg if file size is large raise error, else submit all zero. these are online debug tricks.",
      "votes": 1,
      "replies": [
        {
          "id": 3150042,
          "postDate": "2025-03-15T02:16:27.690Z",
          "content": "<p>Thanks for your comment.<br>\nI'm currently debugging several issues. After running some experiments, I'll share the results</p>",
          "rawMarkdown": "Thanks for your comment.\nI'm currently debugging several issues. After running some experiments, I'll share the results"
        },
        {
          "id": 3150464,
          "postDate": "2025-03-15T14:24:04.137Z",
          "content": "<p>When I predict sequences shorter than 500 in length with NuFold, it works fine. Thanks! Additionally, its performance, when combined with Ribonanzanet (for sequences longer than 500), was quite good.<br>\nFollowing your advice in <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568298</a>, I will try to maximize the RNA sequence length that NuFold can handle in this environment. Furthermore, I plan to explore using other models or MSA as well.</p>",
          "rawMarkdown": "When I predict sequences shorter than 500 in length with NuFold, it works fine. Thanks! Additionally, its performance, when combined with Ribonanzanet (for sequences longer than 500), was quite good.\nFollowing your advice in [https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568298](url), I will try to maximize the RNA sequence length that NuFold can handle in this environment. Furthermore, I plan to explore using other models or MSA as well."
        },
        {
          "id": 3150517,
          "postDate": "2025-03-15T15:10:31Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for the reminder about the working directory.  Not only will the tm-score function modify the shape of the submission df as I mention in <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568383\" target=\"_blank\">this post</a>, it will also add the following to the working directory:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11924794%2F3f9e58eff40e4224e0bdf65c05fa0ec7%2Ftm_score_output.png?generation=1742051410520829&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Thank you @hengck23 for the reminder about the working directory.  Not only will the tm-score function modify the shape of the submission df as I mention in [this post](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568383), it will also add the following to the working directory:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11924794%2F3f9e58eff40e4224e0bdf65c05fa0ec7%2Ftm_score_output.png?generation=1742051410520829&alt=media)\n\n"
        }
      ]
    },
    {
      "id": 3159096,
      "postDate": "2025-03-25T09:19:17.310Z",
      "content": "<p>Hi. I wanted to ask about the lb score your NuFold model can achieve. I recently tested NuFold using Kaggle MSA on CASP15 data, and here are the results:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16242374%2Fca4316619fc34de20f7f9579e39b7a60%2FWechatIMG355.jpg?generation=1742894035567095&amp;alt=media\" alt=\"\"><br>\nFor my Kaggle submission, I used NuFold to predict sequences shorter than 500 nucleotides and Ribonanzanet for longer sequences. This approach gave me an LB score of ~0.23.</p>",
      "rawMarkdown": "Hi. I wanted to ask about the lb score your NuFold model can achieve. I recently tested NuFold using Kaggle MSA on CASP15 data, and here are the results:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16242374%2Fca4316619fc34de20f7f9579e39b7a60%2FWechatIMG355.jpg?generation=1742894035567095&alt=media)\nFor my Kaggle submission, I used NuFold to predict sequences shorter than 500 nucleotides and Ribonanzanet for longer sequences. This approach gave me an LB score of ~0.23.",
      "replies": [
        {
          "id": 3159098,
          "postDate": "2025-03-25T09:27:27.413Z",
          "content": "<p>any method that use MSA is dependent on the search results.</p>\n<p>you can try to write to the author for the MSA script they use for RNA.</p>\n<hr>\n<p>on the other hand methods that do not use MSA may NOT get better results if test is OUTLIER</p>",
          "rawMarkdown": "any method that use MSA is dependent on the search results.\n\nyou can try to write to the author for the MSA script they use for RNA.\n\n---\n\non the other hand methods that do not use MSA may NOT get better results if test is OUTLIER",
          "replies": [
            {
              "id": 3159108,
              "postDate": "2025-03-25T09:39:56.287Z",
              "content": "<p>Thank you for your advice. I will ask authors for the script or the MSA they used on CASP15 data.</p>",
              "rawMarkdown": "Thank you for your advice. I will ask authors for the script or the MSA they used on CASP15 data."
            }
          ]
        }
      ]
    },
    {
      "id": 3149303,
      "postDate": "2025-03-14T03:42:47.040Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15620263%2Fcd011876418ab2ac9a23d20d95c1418c%2F2025-03-14%20%2012.40.46.png?generation=1741923660575625&amp;alt=media\" alt=\"\"><br>\nWhile implementing NuFold for this competition, the notebook runs normally during submission. However, the scoring process gets stuck loading indefinitely and eventually results in a \"Notebook out of disk\" error. How can I resolve this?</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15620263%2Fcd011876418ab2ac9a23d20d95c1418c%2F2025-03-14%20%2012.40.46.png?generation=1741923660575625&alt=media)\nWhile implementing NuFold for this competition, the notebook runs normally during submission. However, the scoring process gets stuck loading indefinitely and eventually results in a \"Notebook out of disk\" error. How can I resolve this?"
    },
    {
      "id": 3150041,
      "postDate": "2025-03-15T02:16:15.477Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3149377,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-03-14T05:50:12.440000",
      "content": "<p>i suspect wrong submission format. you do scoring in local.</p>\n<p>metric code is available. try local scoring. for various rna length and monitor disk and ram usage.<br>\nthere is no problem at my side. try degenerate case like inf and zero. i. not sure how nufold handle  cases when it cannot make prediction, please check it.</p>\n<p>make sure no temp files in kaggle working folder. only submission. csv is allowed.<br>\nelse kaggle server cannot find  the csv file(well known kaggle bug)</p>\n<p>you. can do online debug,. eg. submit only those less than 100 in length and set rest to zero.  other conditions are: only first 10, only. casp16 test( which you can check with local implementation), etc … this. can help pin point which test seq got error.</p>\n<p>you can check csv file size, etc or other property  of solution. set different submission on different  condition. eg if file size is large raise error, else submit all zero. these are online debug tricks.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3150042,
          "author_name": "doheon114",
          "author_url": "",
          "post_date": "2025-03-15T02:16:27.690000",
          "content": "<p>Thanks for your comment.<br>\nI'm currently debugging several issues. After running some experiments, I'll share the results</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3150464,
          "author_name": "doheon114",
          "author_url": "",
          "post_date": "2025-03-15T14:24:04.137000",
          "content": "<p>When I predict sequences shorter than 500 in length with NuFold, it works fine. Thanks! Additionally, its performance, when combined with Ribonanzanet (for sequences longer than 500), was quite good.<br>\nFollowing your advice in <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568298</a>, I will try to maximize the RNA sequence length that NuFold can handle in this environment. Furthermore, I plan to explore using other models or MSA as well.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3150517,
          "author_name": "ChrisK321",
          "author_url": "",
          "post_date": "2025-03-15T15:10:31",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for the reminder about the working directory.  Not only will the tm-score function modify the shape of the submission df as I mention in <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/568383\" target=\"_blank\">this post</a>, it will also add the following to the working directory:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11924794%2F3f9e58eff40e4224e0bdf65c05fa0ec7%2Ftm_score_output.png?generation=1742051410520829&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3159096,
      "author_name": "Bianco Chiu",
      "author_url": "",
      "post_date": "2025-03-25T09:19:17.310000",
      "content": "<p>Hi. I wanted to ask about the lb score your NuFold model can achieve. I recently tested NuFold using Kaggle MSA on CASP15 data, and here are the results:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16242374%2Fca4316619fc34de20f7f9579e39b7a60%2FWechatIMG355.jpg?generation=1742894035567095&amp;alt=media\" alt=\"\"><br>\nFor my Kaggle submission, I used NuFold to predict sequences shorter than 500 nucleotides and Ribonanzanet for longer sequences. This approach gave me an LB score of ~0.23.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3159098,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2025-03-25T09:27:27.413000",
          "content": "<p>any method that use MSA is dependent on the search results.</p>\n<p>you can try to write to the author for the MSA script they use for RNA.</p>\n<hr>\n<p>on the other hand methods that do not use MSA may NOT get better results if test is OUTLIER</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3159108,
              "author_name": "Bianco Chiu",
              "author_url": "",
              "post_date": "2025-03-25T09:39:56.287000",
              "content": "<p>Thank you for your advice. I will ask authors for the script or the MSA they used on CASP15 data.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3150041,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-03-15T02:16:15.477000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3149377": "i suspect wrong submission format. you do scoring in local.\n\n\nmetric code is available. try local scoring. for various rna length and monitor disk and ram usage.\nthere is no problem at my side. try degenerate case like inf and zero. i. not sure how nufold handle  cases when it cannot make prediction, please check it.\n\nmake sure no temp files in kaggle working folder. only submission. csv is allowed.\nelse kaggle server cannot find  the csv file(well known kaggle bug)\n\nyou. can do online debug,. eg. submit only those less than 100 in length and set rest to zero.  other conditions are: only first 10, only. casp16 test( which you can check with local implementation), etc ... this. can help pin point which test seq got error.\n\nyou can check csv file size, etc or other property  of solution. set different submission on different  condition. eg if file size is large raise error, else submit all zero. these are online debug tricks.",
    "3159096": "Hi. I wanted to ask about the lb score your NuFold model can achieve. I recently tested NuFold using Kaggle MSA on CASP15 data, and here are the results:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16242374%2Fca4316619fc34de20f7f9579e39b7a60%2FWechatIMG355.jpg?generation=1742894035567095&alt=media)\nFor my Kaggle submission, I used NuFold to predict sequences shorter than 500 nucleotides and Ribonanzanet for longer sequences. This approach gave me an LB score of ~0.23.",
    "3149303": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15620263%2Fcd011876418ab2ac9a23d20d95c1418c%2F2025-03-14%20%2012.40.46.png?generation=1741923660575625&alt=media)\nWhile implementing NuFold for this competition, the notebook runs normally during submission. However, the scoring process gets stuck loading indefinitely and eventually results in a \"Notebook out of disk\" error. How can I resolve this?",
    "3150041": ""
  }
}