{
  "id": 567502,
  "title": "NA loss during finetuning with max length of 1024?",
  "url": "/competitions/stanford-rna-3d-folding/discussion/567502",
  "author_name": "",
  "post_date": "2025-03-10T18:04:20.223524800Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I'm having a problem with finetuning RibonanzaNet 3D (float16) - at some point I start getting NA loss, although I have no problems with finetuning it with lesser max length like 768, 512, etc. Did anyone else encountered the same problem? I was using dRMAE loss.</p>",
  "messages": [
    {
      "id": "3146295",
      "postDate": "03/10/2025 18:04:20",
      "content": "<p>I'm having a problem with finetuning RibonanzaNet 3D (float16) - at some point I start getting NA loss, although I have no problems with finetuning it with lesser max length like 768, 512, etc. Did anyone else encountered the same problem? I was using dRMAE loss.</p>",
      "rawMarkdown": "I'm having a problem with finetuning RibonanzaNet 3D (float16) - at some point I start getting NA loss, although I have no problems with finetuning it with lesser max length like 768, 512, etc. Did anyone else encountered the same problem? I was using dRMAE loss.",
      "votes": null
    },
    {
      "id": "3146321",
      "postDate": "03/10/2025 18:27:16",
      "content": "<p>I also encountered a similar problem. 🤧</p>",
      "rawMarkdown": "I also encountered a similar problem. 🤧",
      "votes": null
    },
    {
      "id": "3146333",
      "postDate": "03/10/2025 18:44:38",
      "content": "<p>Do you have any idea what may cause this?</p>",
      "rawMarkdown": "Do you have any idea what may cause this?",
      "votes": null
    },
    {
      "id": "3146356",
      "postDate": "03/10/2025 19:06:34",
      "content": "<p>After asking DeepSeek I think it can be in calculate_distance_matrix:</p>\n<ol>\n<li><p>If the values in X or Y are very large, squaring them could lead to overflow, resulting in NaN when taking the square root.</p></li>\n<li><p>Although epsilon is added to ensure the argument of sqrt is non-negative, extreme numerical errors could still result in negative values.</p></li>\n</ol>\n<p>For me everything else looks pretty robust.</p>\n<p>EDIT: So you are training with single dRMAE and larger sequences. Thanks for the advice.</p>\n<p>EDIT2: And since the problem occurs for large sequences I would say option one is the winner. Why not absolute distances too to mitigate it?</p>\n<p>EDIT: What reminds me <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/565746\" target=\"_blank\"><strong><code>validation_labels.csv</code>: should -1e+18 be interpreted as NaN?</strong></a></p>",
      "rawMarkdown": "After asking DeepSeek I think it can be in calculate_distance_matrix:\n\n1. If the values in X or Y are very large, squaring them could lead to overflow, resulting in NaN when taking the square root.\n\n2. Although epsilon is added to ensure the argument of sqrt is non-negative, extreme numerical errors could still result in negative values.\n\nFor me everything else looks pretty robust.\n\nEDIT: So you are training with single dRMAE and larger sequences. Thanks for the advice.\n\nEDIT2: And since the problem occurs for large sequences I would say option one is the winner. Why not absolute distances too to mitigate it?\n\nEDIT: What reminds me [**`validation_labels.csv`: should -1e+18 be interpreted as NaN?**](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/565746)",
      "votes": null
    },
    {
      "id": "3146447",
      "postDate": "03/10/2025 22:13:59",
      "content": "<blockquote>\n  <p>What reminds me validation_labels.csv: should -1e+18 be interpreted as NaN?</p>\n</blockquote>\n<p>Yes !</p>",
      "rawMarkdown": "> What reminds me validation_labels.csv: should -1e+18 be interpreted as NaN?\n\nYes !",
      "votes": null
    },
    {
      "id": "3147120",
      "postDate": "03/11/2025 17:00:25",
      "content": "<p>i also face this <br>\nmy solution:&nbsp;</p>\n<pre><code>NaNs found  target coordinates!\nNaN loss encountered!\n</code></pre>\n<p>This happened&nbsp;we your loss is Nan <br>\nTo prevent this i use this logic&nbsp;in my case i'm train custom&nbsp;network&nbsp;so it's woks </p>\n<pre><code>         torch.isnan(seqs).():\n            ()\n         torch.isnan(coords).():\n            ()  \n         torch.isnan(outputs).():\n            ()\n\n        \n        valid_elements = (~torch.isnan(coords)).()\n         valid_elements &lt; config[]:\n            ()\n              \n</code></pre>",
      "rawMarkdown": "i also face this \nmy solution: \n\n```python\nNaNs found in target coordinates!\nNaN loss encountered!\n```\n\nThis happened we your loss is Nan \nTo prevent this i use this logic in my case i'm train custom network so it's woks \n```python\n\n        if torch.isnan(seqs).any():\n            print(\"NaNs found in input sequences!\")\n        if torch.isnan(coords).any():\n            print(\"NaNs found in target coordinates!\")  # Expected\n        if torch.isnan(outputs).any():\n            print(\"NaNs found in model outputs!\")\n\n        # --- Batch Skipping ---\n        valid_elements = (~torch.isnan(coords)).sum()\n        if valid_elements < config['min_valid_elements']:\n            print(f\"Skipping batch with only {valid_elements} valid elements.\")\n            continue  # Skip to the next batch\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3146321,
      "author_name": "quan0095",
      "author_url": "",
      "post_date": "03/10/2025 18:27:16",
      "content": "<p>I also encountered a similar problem. 🤧</p>",
      "votes": null,
      "replies": [
        {
          "id": 3146333,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "03/10/2025 18:44:38",
          "content": "<p>Do you have any idea what may cause this?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3146356,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "03/10/2025 19:06:34",
      "content": "<p>After asking DeepSeek I think it can be in calculate_distance_matrix:</p>\n<ol>\n<li><p>If the values in X or Y are very large, squaring them could lead to overflow, resulting in NaN when taking the square root.</p></li>\n<li><p>Although epsilon is added to ensure the argument of sqrt is non-negative, extreme numerical errors could still result in negative values.</p></li>\n</ol>\n<p>For me everything else looks pretty robust.</p>\n<p>EDIT: So you are training with single dRMAE and larger sequences. Thanks for the advice.</p>\n<p>EDIT2: And since the problem occurs for large sequences I would say option one is the winner. Why not absolute distances too to mitigate it?</p>\n<p>EDIT: What reminds me <a href=\"https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/565746\" target=\"_blank\"><strong><code>validation_labels.csv</code>: should -1e+18 be interpreted as NaN?</strong></a></p>",
      "votes": null,
      "replies": [
        {
          "id": 3146447,
          "author_name": "louisstefanuto",
          "author_url": "",
          "post_date": "03/10/2025 22:13:59",
          "content": "<blockquote>\n  <p>What reminds me validation_labels.csv: should -1e+18 be interpreted as NaN?</p>\n</blockquote>\n<p>Yes !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3147120,
      "author_name": "sangrampatil5150",
      "author_url": "",
      "post_date": "03/11/2025 17:00:25",
      "content": "<p>i also face this <br>\nmy solution:&nbsp;</p>\n<pre><code>NaNs found  target coordinates!\nNaN loss encountered!\n</code></pre>\n<p>This happened&nbsp;we your loss is Nan <br>\nTo prevent this i use this logic&nbsp;in my case i'm train custom&nbsp;network&nbsp;so it's woks </p>\n<pre><code>         torch.isnan(seqs).():\n            ()\n         torch.isnan(coords).():\n            ()  \n         torch.isnan(outputs).():\n            ()\n\n        \n        valid_elements = (~torch.isnan(coords)).()\n         valid_elements &lt; config[]:\n            ()\n              \n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3146295": "I'm having a problem with finetuning RibonanzaNet 3D (float16) - at some point I start getting NA loss, although I have no problems with finetuning it with lesser max length like 768, 512, etc. Did anyone else encountered the same problem? I was using dRMAE loss.",
    "3146321": "I also encountered a similar problem. 🤧",
    "3146333": "Do you have any idea what may cause this?",
    "3146356": "After asking DeepSeek I think it can be in calculate_distance_matrix:\n\n1. If the values in X or Y are very large, squaring them could lead to overflow, resulting in NaN when taking the square root.\n\n2. Although epsilon is added to ensure the argument of sqrt is non-negative, extreme numerical errors could still result in negative values.\n\nFor me everything else looks pretty robust.\n\nEDIT: So you are training with single dRMAE and larger sequences. Thanks for the advice.\n\nEDIT2: And since the problem occurs for large sequences I would say option one is the winner. Why not absolute distances too to mitigate it?\n\nEDIT: What reminds me [**`validation_labels.csv`: should -1e+18 be interpreted as NaN?**](https://www.kaggle.com/competitions/stanford-rna-3d-folding/discussion/565746)",
    "3146447": "> What reminds me validation_labels.csv: should -1e+18 be interpreted as NaN?\n\nYes !",
    "3147120": "i also face this \nmy solution: \n\n```python\nNaNs found in target coordinates!\nNaN loss encountered!\n```\n\nThis happened we your loss is Nan \nTo prevent this i use this logic in my case i'm train custom network so it's woks \n```python\n\n        if torch.isnan(seqs).any():\n            print(\"NaNs found in input sequences!\")\n        if torch.isnan(coords).any():\n            print(\"NaNs found in target coordinates!\")  # Expected\n        if torch.isnan(outputs).any():\n            print(\"NaNs found in model outputs!\")\n\n        # --- Batch Skipping ---\n        valid_elements = (~torch.isnan(coords)).sum()\n        if valid_elements < config['min_valid_elements']:\n            print(f\"Skipping batch with only {valid_elements} valid elements.\")\n            continue  # Skip to the next batch\n```"
  },
  "source": "meta"
}