{
  "id": 585206,
  "title": "Newbie GNN model, I would love to have some feedback on my prediciton model.",
  "url": "/competitions/stanford-rna-3d-folding/discussion/585206",
  "author_name": "",
  "post_date": "2025-06-18T17:38:43.442263Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I'm VERY new to coding and machine learning (I have a biochemistry background), and this is my first stab at a Kaggle competition project.<br>\nI built a graph neural network (GNN) model with a lot of help from ChatGPT. The model uses a combination of sequence-based and structural features, including k-mer embeddings, minimum free energy (MFE), base-pairing probability matrices (BPPM), multiple sequence alignment (MSA)-based entropy, mutual information (MI) matrices, and position-specific scoring matrices (PSSMs).</p>\n<p>My TM-scores seem surprisingly high (~0.8), and I’m not sure if I’ve made a mistake or if the model is actually doing something meaningful. I’d really appreciate it if anyone more experienced could take a look and let me know whether this approach holds any real value or if I’ve overlooked something important.</p>\n<p>Here’s a link to my notebook: <a href=\"https://github.com/shendong124/Springboard/blob/main/Capstone03/Capstone%203.ipynb\" target=\"_blank\">https://github.com/shendong124/Springboard/blob/main/Capstone03/Capstone%203.ipynb</a></p>\n<p>Thanks so much in advance for your time and guidance!<br>\nShen</p>",
  "messages": [
    {
      "id": "3227245",
      "postDate": "06/18/2025 17:38:43",
      "content": "<p>Hi everyone,</p>\n<p>I'm VERY new to coding and machine learning (I have a biochemistry background), and this is my first stab at a Kaggle competition project.<br>\nI built a graph neural network (GNN) model with a lot of help from ChatGPT. The model uses a combination of sequence-based and structural features, including k-mer embeddings, minimum free energy (MFE), base-pairing probability matrices (BPPM), multiple sequence alignment (MSA)-based entropy, mutual information (MI) matrices, and position-specific scoring matrices (PSSMs).</p>\n<p>My TM-scores seem surprisingly high (~0.8), and I’m not sure if I’ve made a mistake or if the model is actually doing something meaningful. I’d really appreciate it if anyone more experienced could take a look and let me know whether this approach holds any real value or if I’ve overlooked something important.</p>\n<p>Here’s a link to my notebook: <a href=\"https://github.com/shendong124/Springboard/blob/main/Capstone03/Capstone%203.ipynb\" target=\"_blank\">https://github.com/shendong124/Springboard/blob/main/Capstone03/Capstone%203.ipynb</a></p>\n<p>Thanks so much in advance for your time and guidance!<br>\nShen</p>",
      "rawMarkdown": "Hi everyone,\n\nI'm VERY new to coding and machine learning (I have a biochemistry background), and this is my first stab at a Kaggle competition project.\nI built a graph neural network (GNN) model with a lot of help from ChatGPT. The model uses a combination of sequence-based and structural features, including k-mer embeddings, minimum free energy (MFE), base-pairing probability matrices (BPPM), multiple sequence alignment (MSA)-based entropy, mutual information (MI) matrices, and position-specific scoring matrices (PSSMs).\n\nMy TM-scores seem surprisingly high (~0.8), and I’m not sure if I’ve made a mistake or if the model is actually doing something meaningful. I’d really appreciate it if anyone more experienced could take a look and let me know whether this approach holds any real value or if I’ve overlooked something important.\n\nHere’s a link to my notebook: https://github.com/shendong124/Springboard/blob/main/Capstone03/Capstone%203.ipynb\n\nThanks so much in advance for your time and guidance!\nShen",
      "votes": null
    },
    {
      "id": "3227844",
      "postDate": "06/19/2025 10:58:48",
      "content": "<p>If I understand the code correctly…<br>\nIf you have normalized coordinates, you need to turn coordinates back to de-normalized coordinates before calculating TM-score.</p>",
      "rawMarkdown": "If I understand the code correctly...\nIf you have normalized coordinates, you need to turn coordinates back to de-normalized coordinates before calculating TM-score.",
      "votes": null
    },
    {
      "id": "3227949",
      "postDate": "06/19/2025 13:25:43",
      "content": "<p>I think there are two things:</p>\n<ol>\n<li>Validation set has 11 unique sequences and 6 of them are overlapping with training set.</li>\n<li>Your TM-score calculation is not corresponding to the one mentioned by the host. Your calculation:</li>\n</ol>\n<pre><code>\n\n\n tm_score(pred, target):\n     = pred.detach().cpu().numpy()\n     = target.detach().cpu().numpy()\n     = min(len(pred), len(target))  # Lalign = Lref\n\n    \n     L &gt;= :\n         = . * (L - .) ** . - .\n     L &lt; :\n         = .\n     L &lt; :\n         = .\n     L &lt; :\n         = .\n     L &lt; :\n         = .\n    :  #  &lt;= L &lt; \n         = .\n\n     = .\n     i in range(L):\n         = np.linalg.norm(pred[i] - target[i])\n         +=  / ( + (dist / d0) ** )\n     score / L\n</code></pre>\n<p>For TM-score calculation, we need to superimpose two structures first and then calculate the distance by the formulation. The superimposition step is done with USAlign. Here is how TM-score is calculated: <a href=\"https://www.kaggle.com/code/metric/ribonanza-tm-score\" target=\"_blank\">https://www.kaggle.com/code/metric/ribonanza-tm-score</a></p>",
      "rawMarkdown": "I think there are two things:\n\n1. Validation set has 11 unique sequences and 6 of them are overlapping with training set.\n2. Your TM-score calculation is not corresponding to the one mentioned by the host. Your calculation:\n```\n# -------------------------\n# TM-score calculation\n# -------------------------\ndef tm_score(pred, target):\n    pred = pred.detach().cpu().numpy()\n    target = target.detach().cpu().numpy()\n    L = min(len(pred), len(target))  # Lalign = Lref\n\n    # Apply piecewise formula for d0\n    if L >= 30:\n        d0 = 0.6 * (L - 0.5) ** 0.5 - 2.5\n    elif L < 12:\n        d0 = 0.3\n    elif L < 16:\n        d0 = 0.4\n    elif L < 20:\n        d0 = 0.5\n    elif L < 24:\n        d0 = 0.6\n    else:  # 24 <= L < 30\n        d0 = 0.7\n\n    score = 0.0\n    for i in range(L):\n        dist = np.linalg.norm(pred[i] - target[i])\n        score += 1 / (1 + (dist / d0) ** 2)\n    return score / L\n```\nFor TM-score calculation, we need to superimpose two structures first and then calculate the distance by the formulation. The superimposition step is done with USAlign. Here is how TM-score is calculated: https://www.kaggle.com/code/metric/ribonanza-tm-score",
      "votes": null
    },
    {
      "id": "3230605",
      "postDate": "06/23/2025 08:37:08",
      "content": "<p>Except some wrong implementation in your metric, </p>\n<p>try: <br>\ndelaunay connection<br>\nknn edge<br>\ngraphsage<br>\npositional embedding<br>\ntry to change graph to heterogeneous graph<br>\ndirected graph<br>\nmore <a href=\"https://pytorch-geometric.readthedocs.io/en/latest/modules/nn.html\" target=\"_blank\">pyg docs </a><br>\npoint sampling</p>\n<p>you can refer my <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/583128\" target=\"_blank\">solution in BYU comp</a>, I think some ingredients can improve the result in this comp </p>",
      "rawMarkdown": "Except some wrong implementation in your metric, \n\ntry: \ndelaunay connection\nknn edge\ngraphsage\npositional embedding\ntry to change graph to heterogeneous graph\ndirected graph\nmore [pyg docs ](https://pytorch-geometric.readthedocs.io/en/latest/modules/nn.html)\npoint sampling\n\nyou can refer my [solution in BYU comp](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/583128), I think some ingredients can improve the result in this comp",
      "votes": null
    },
    {
      "id": "3231180",
      "postDate": "06/24/2025 03:39:23",
      "content": "<p>From my perspective, GNN may not be the most suitable approach for capturing RNA spatial features. This is because RNA folding patterns are governed by some thermodynamic principles. If you examine the attention module implementation in the RibonanzaNet repository, you'll notice how it specifically addresses spatial information capture. GNN needs more feature enginnering.</p>",
      "rawMarkdown": "From my perspective, GNN may not be the most suitable approach for capturing RNA spatial features. This is because RNA folding patterns are governed by some thermodynamic principles. If you examine the attention module implementation in the RibonanzaNet repository, you'll notice how it specifically addresses spatial information capture. GNN needs more feature enginnering.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3227844,
      "author_name": "odat1248",
      "author_url": "",
      "post_date": "06/19/2025 10:58:48",
      "content": "<p>If I understand the code correctly…<br>\nIf you have normalized coordinates, you need to turn coordinates back to de-normalized coordinates before calculating TM-score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3227949,
      "author_name": "nguyenhoa",
      "author_url": "",
      "post_date": "06/19/2025 13:25:43",
      "content": "<p>I think there are two things:</p>\n<ol>\n<li>Validation set has 11 unique sequences and 6 of them are overlapping with training set.</li>\n<li>Your TM-score calculation is not corresponding to the one mentioned by the host. Your calculation:</li>\n</ol>\n<pre><code>\n\n\n tm_score(pred, target):\n     = pred.detach().cpu().numpy()\n     = target.detach().cpu().numpy()\n     = min(len(pred), len(target))  # Lalign = Lref\n\n    \n     L &gt;= :\n         = . * (L - .) ** . - .\n     L &lt; :\n         = .\n     L &lt; :\n         = .\n     L &lt; :\n         = .\n     L &lt; :\n         = .\n    :  #  &lt;= L &lt; \n         = .\n\n     = .\n     i in range(L):\n         = np.linalg.norm(pred[i] - target[i])\n         +=  / ( + (dist / d0) ** )\n     score / L\n</code></pre>\n<p>For TM-score calculation, we need to superimpose two structures first and then calculate the distance by the formulation. The superimposition step is done with USAlign. Here is how TM-score is calculated: <a href=\"https://www.kaggle.com/code/metric/ribonanza-tm-score\" target=\"_blank\">https://www.kaggle.com/code/metric/ribonanza-tm-score</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3230605,
      "author_name": "tom99763",
      "author_url": "",
      "post_date": "06/23/2025 08:37:08",
      "content": "<p>Except some wrong implementation in your metric, </p>\n<p>try: <br>\ndelaunay connection<br>\nknn edge<br>\ngraphsage<br>\npositional embedding<br>\ntry to change graph to heterogeneous graph<br>\ndirected graph<br>\nmore <a href=\"https://pytorch-geometric.readthedocs.io/en/latest/modules/nn.html\" target=\"_blank\">pyg docs </a><br>\npoint sampling</p>\n<p>you can refer my <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/583128\" target=\"_blank\">solution in BYU comp</a>, I think some ingredients can improve the result in this comp </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3231180,
      "author_name": "sweetyheehee",
      "author_url": "",
      "post_date": "06/24/2025 03:39:23",
      "content": "<p>From my perspective, GNN may not be the most suitable approach for capturing RNA spatial features. This is because RNA folding patterns are governed by some thermodynamic principles. If you examine the attention module implementation in the RibonanzaNet repository, you'll notice how it specifically addresses spatial information capture. GNN needs more feature enginnering.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3227245": "Hi everyone,\n\nI'm VERY new to coding and machine learning (I have a biochemistry background), and this is my first stab at a Kaggle competition project.\nI built a graph neural network (GNN) model with a lot of help from ChatGPT. The model uses a combination of sequence-based and structural features, including k-mer embeddings, minimum free energy (MFE), base-pairing probability matrices (BPPM), multiple sequence alignment (MSA)-based entropy, mutual information (MI) matrices, and position-specific scoring matrices (PSSMs).\n\nMy TM-scores seem surprisingly high (~0.8), and I’m not sure if I’ve made a mistake or if the model is actually doing something meaningful. I’d really appreciate it if anyone more experienced could take a look and let me know whether this approach holds any real value or if I’ve overlooked something important.\n\nHere’s a link to my notebook: https://github.com/shendong124/Springboard/blob/main/Capstone03/Capstone%203.ipynb\n\nThanks so much in advance for your time and guidance!\nShen",
    "3227844": "If I understand the code correctly...\nIf you have normalized coordinates, you need to turn coordinates back to de-normalized coordinates before calculating TM-score.",
    "3227949": "I think there are two things:\n\n1. Validation set has 11 unique sequences and 6 of them are overlapping with training set.\n2. Your TM-score calculation is not corresponding to the one mentioned by the host. Your calculation:\n```\n# -------------------------\n# TM-score calculation\n# -------------------------\ndef tm_score(pred, target):\n    pred = pred.detach().cpu().numpy()\n    target = target.detach().cpu().numpy()\n    L = min(len(pred), len(target))  # Lalign = Lref\n\n    # Apply piecewise formula for d0\n    if L >= 30:\n        d0 = 0.6 * (L - 0.5) ** 0.5 - 2.5\n    elif L < 12:\n        d0 = 0.3\n    elif L < 16:\n        d0 = 0.4\n    elif L < 20:\n        d0 = 0.5\n    elif L < 24:\n        d0 = 0.6\n    else:  # 24 <= L < 30\n        d0 = 0.7\n\n    score = 0.0\n    for i in range(L):\n        dist = np.linalg.norm(pred[i] - target[i])\n        score += 1 / (1 + (dist / d0) ** 2)\n    return score / L\n```\nFor TM-score calculation, we need to superimpose two structures first and then calculate the distance by the formulation. The superimposition step is done with USAlign. Here is how TM-score is calculated: https://www.kaggle.com/code/metric/ribonanza-tm-score",
    "3230605": "Except some wrong implementation in your metric, \n\ntry: \ndelaunay connection\nknn edge\ngraphsage\npositional embedding\ntry to change graph to heterogeneous graph\ndirected graph\nmore [pyg docs ](https://pytorch-geometric.readthedocs.io/en/latest/modules/nn.html)\npoint sampling\n\nyou can refer my [solution in BYU comp](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/583128), I think some ingredients can improve the result in this comp",
    "3231180": "From my perspective, GNN may not be the most suitable approach for capturing RNA spatial features. This is because RNA folding patterns are governed by some thermodynamic principles. If you examine the attention module implementation in the RibonanzaNet repository, you'll notice how it specifically addresses spatial information capture. GNN needs more feature enginnering."
  },
  "source": "meta"
}