{
  "id": 610629,
  "title": "Help us discover new RNA folds in Nature",
  "url": "/competitions/stanford-rna-3d-folding/discussion/610629",
  "author_name": "",
  "post_date": "2025-10-05T02:54:04.648109500Z",
  "votes": 8,
  "comment_count": 7,
  "views": 0,
  "content": "<p>One of the key goals of 3D RNA structure prediction is to identify all the RNA segments across natural RNA sequences that might have well-defined 3D structure.  We'd like your help to get us there!</p>\n<p>We'd like to get Kaggle codes run across all natural RNA sequences. For the highest confidence predictions that are novel, we could prioritize these segments for high resolution experimental 3D structure determination.</p>\n<p>We're one step away.  The thing is: <em>we don't want a lot of false positives</em>. So we'd love to have each notebook give a confidence score that predicts the TM-score that the target would get when compared to an experimentally solved structure. </p>\n<p>We'd especially like to identify cases where TM is predicted to be &gt;0.45, the typical cutoff for a reasonable fold prediction.</p>\n<p>If you'd like to help, take your favorite notebook and attach it to this data set</p>\n<p><a href=\"https://www.kaggle.com/datasets/rhijudas/stanford-rna-3d-folding-pdb-summer2025\" target=\"_blank\">https://www.kaggle.com/datasets/rhijudas/stanford-rna-3d-folding-pdb-summer2025</a></p>\n<p>It holds 13 targets that have come out in the PDB since the close of the training phase.</p>\n<p>And then see if you can come up with a confidence score that achieves a high precision in predicting which of the 13 predictions actually has Tm&gt;0.45.  </p>\n<p>Here's an example based on a notebook that runs template search with MMseqs2 and then estimates confidence based on the fraction of the target covered by the template: </p>\n<p><a href=\"https://www.kaggle.com/code/rhijudas/confidence-score-mmseqs2-3d-rna-template-id\" target=\"_blank\">https://www.kaggle.com/code/rhijudas/confidence-score-mmseqs2-3d-rna-template-id</a></p>\n<p>The notebook gets mean 3D accuracy of Tm = 0.368 and Precision, Recall, and F1 of confidence score accuracy of 1.0</p>\n<p>Your notebooks should do better in terms of 3D accuracy and, hopefully maintain precision of 1.0. Please post statistics and notebook links here if you make progress!</p>\n<p>(<em>Precision, recall, and F1 values updated after fix to notebook's comparison of ref and pred.</em>)</p>",
  "messages": [
    {
      "id": "3298264",
      "postDate": "10/05/2025 02:54:04",
      "content": "<p>One of the key goals of 3D RNA structure prediction is to identify all the RNA segments across natural RNA sequences that might have well-defined 3D structure.  We'd like your help to get us there!</p>\n<p>We'd like to get Kaggle codes run across all natural RNA sequences. For the highest confidence predictions that are novel, we could prioritize these segments for high resolution experimental 3D structure determination.</p>\n<p>We're one step away.  The thing is: <em>we don't want a lot of false positives</em>. So we'd love to have each notebook give a confidence score that predicts the TM-score that the target would get when compared to an experimentally solved structure. </p>\n<p>We'd especially like to identify cases where TM is predicted to be &gt;0.45, the typical cutoff for a reasonable fold prediction.</p>\n<p>If you'd like to help, take your favorite notebook and attach it to this data set</p>\n<p><a href=\"https://www.kaggle.com/datasets/rhijudas/stanford-rna-3d-folding-pdb-summer2025\" target=\"_blank\">https://www.kaggle.com/datasets/rhijudas/stanford-rna-3d-folding-pdb-summer2025</a></p>\n<p>It holds 13 targets that have come out in the PDB since the close of the training phase.</p>\n<p>And then see if you can come up with a confidence score that achieves a high precision in predicting which of the 13 predictions actually has Tm&gt;0.45.  </p>\n<p>Here's an example based on a notebook that runs template search with MMseqs2 and then estimates confidence based on the fraction of the target covered by the template: </p>\n<p><a href=\"https://www.kaggle.com/code/rhijudas/confidence-score-mmseqs2-3d-rna-template-id\" target=\"_blank\">https://www.kaggle.com/code/rhijudas/confidence-score-mmseqs2-3d-rna-template-id</a></p>\n<p>The notebook gets mean 3D accuracy of Tm = 0.368 and Precision, Recall, and F1 of confidence score accuracy of 1.0</p>\n<p>Your notebooks should do better in terms of 3D accuracy and, hopefully maintain precision of 1.0. Please post statistics and notebook links here if you make progress!</p>\n<p>(<em>Precision, recall, and F1 values updated after fix to notebook's comparison of ref and pred.</em>)</p>",
      "rawMarkdown": "One of the key goals of 3D RNA structure prediction is to identify all the RNA segments across natural RNA sequences that might have well-defined 3D structure.  We'd like your help to get us there!\n\nWe'd like to get Kaggle codes run across all natural RNA sequences. For the highest confidence predictions that are novel, we could prioritize these segments for high resolution experimental 3D structure determination.\n\nWe're one step away.  The thing is: *we don't want a lot of false positives*. So we'd love to have each notebook give a confidence score that predicts the TM-score that the target would get when compared to an experimentally solved structure. \n\nWe'd especially like to identify cases where TM is predicted to be >0.45, the typical cutoff for a reasonable fold prediction.\n\nIf you'd like to help, take your favorite notebook and attach it to this data set\n\nhttps://www.kaggle.com/datasets/rhijudas/stanford-rna-3d-folding-pdb-summer2025\n\nIt holds 13 targets that have come out in the PDB since the close of the training phase.\n\nAnd then see if you can come up with a confidence score that achieves a high precision in predicting which of the 13 predictions actually has Tm>0.45.  \n\nHere's an example based on a notebook that runs template search with MMseqs2 and then estimates confidence based on the fraction of the target covered by the template: \n\nhttps://www.kaggle.com/code/rhijudas/confidence-score-mmseqs2-3d-rna-template-id\n \nThe notebook gets mean 3D accuracy of Tm = 0.368 and Precision, Recall, and F1 of confidence score accuracy of 1.0\n\nYour notebooks should do better in terms of 3D accuracy and, hopefully maintain precision of 1.0. Please post statistics and notebook links here if you make progress!\n\n(*Precision, recall, and F1 values updated after fix to notebook's comparison of ref and pred.*)",
      "votes": null
    },
    {
      "id": "3298525",
      "postDate": "10/05/2025 18:35:44",
      "content": "<p>Hi, I'm not entirely sure what output format you would need. I ran a simulation on these 13 targets and got results like the following (see attachment)(the TM-scores were calculated only for informational purposes, using the tm_align library)</p>\n<table>\n<thead>\n<tr>\n<th>MAX_SCORE</th>\n<th>CONFIDENCE(SUM)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.488270057555273</td>\n<td>0.938955225050449</td>\n</tr>\n<tr>\n<td>0.331262390937843</td>\n<td>0.514765247702599</td>\n</tr>\n<tr>\n<td>0.370054298339892</td>\n<td>0.642883621156216</td>\n</tr>\n<tr>\n<td>0.283441785709035</td>\n<td>0.719378303736448</td>\n</tr>\n<tr>\n<td>0.249054987745784</td>\n<td>0.642032139003277</td>\n</tr>\n<tr>\n<td>0.798223644046592</td>\n<td>0.456463485956192</td>\n</tr>\n<tr>\n<td>0.636943080236158</td>\n<td>0.660994231700897</td>\n</tr>\n<tr>\n<td>0.548786849733515</td>\n<td>0.721843557432294</td>\n</tr>\n<tr>\n<td>0.344998093726893</td>\n<td>0.57373333349824</td>\n</tr>\n<tr>\n<td>0.741880618462633</td>\n<td>0.804388243705034</td>\n</tr>\n<tr>\n<td>0.625168415230373</td>\n<td>0.724774695932865</td>\n</tr>\n<tr>\n<td>0.317178276413226</td>\n<td>0.724590353667736</td>\n</tr>\n<tr>\n<td>0.763040445662887</td>\n<td>0.99974361560453</td>\n</tr>\n<tr>\n<td><strong>mean</strong></td>\n<td></td>\n</tr>\n<tr>\n<td>0.499869457215393</td>\n<td>0.701888158011291</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "Hi, I'm not entirely sure what output format you would need. I ran a simulation on these 13 targets and got results like the following (see attachment)(the TM-scores were calculated only for informational purposes, using the tm_align library)\n\n| MAX_SCORE       | CONFIDENCE(SUM)   |\n|-----------------|-----------------|\n| 0.488270057555273 | 0.938955225050449 |\n| 0.331262390937843 | 0.514765247702599 |\n| 0.370054298339892 | 0.642883621156216 |\n| 0.283441785709035 | 0.719378303736448 |\n| 0.249054987745784 | 0.642032139003277 |\n| 0.798223644046592 | 0.456463485956192 |\n| 0.636943080236158 | 0.660994231700897 |\n| 0.548786849733515 | 0.721843557432294 |\n| 0.344998093726893 | 0.57373333349824  |\n| 0.741880618462633 | 0.804388243705034 |\n| 0.625168415230373 | 0.724774695932865 |\n| 0.317178276413226 | 0.724590353667736 |\n| 0.763040445662887 | 0.99974361560453  |\n| **mean**          |                 |\n| 0.499869457215393 | 0.701888158011291 |",
      "votes": null
    },
    {
      "id": "3298545",
      "postDate": "10/05/2025 19:33:08",
      "content": "<p>Thanks! </p>\n<p>Do you see a way to define a confidence score such that when it is high, and only when it is high, it accurately predicts that the prediction has TM-score &gt; 0.45?</p>\n<p>For example in my <a href=\"https://www.kaggle.com/code/rhijudas/confidence-score-mmseqs2-3d-rna-template-id\" target=\"_blank\">notebook</a>, I calculated the fraction of the target covered by the template (<code>target_frac_template</code>), and was able to achieve this:</p>\n<pre><code>     TMscore  target_frac_template\n       G4P  .              .\n       G4R  .              .\n       G4Q  .              .\n       IWF  .              .\n       J09  .              .\n       MMG  .              .\n       MME  .              .\n       VQV  .              .\n       E9Q  .              .\n       J4N  .              .\n      J4O  .              .\n      JGM  .              .\n      LKU  .              .\n</code></pre>\n<p>In particular, precision is 100% – there were 6 targets with <code>target_frac_template</code>&gt;0.45, and all 6 indeed had <code>TM-score</code>&gt;0.45 when compared to the experimental structure.</p>\n<pre><code> TM-score: .\n: . (/)\n   : . (/)\n Score : .\n</code></pre>\n<p>That's nice except that my notebook only found a template for 6 of the 13 cases, and its mean TM-score is only 0.368. </p>\n<p>Most Kaggle notebooks (including yours) will have higher mean TM-score -- can they also define a confidence score that would achieve 100% precision in predicting which targets have TM-score &gt; 0.45?</p>\n<p>P.S. your question prompted me to find a bug in my notebook's comparison of predicted TM-score to actual TM-score. Thanks!</p>",
      "rawMarkdown": "Thanks! \n\nDo you see a way to define a confidence score such that when it is high, and only when it is high, it accurately predicts that the prediction has TM-score > 0.45?\n\nFor example in my [notebook](https://www.kaggle.com/code/rhijudas/confidence-score-mmseqs2-3d-rna-template-id), I calculated the fraction of the target covered by the template (`target_frac_template`), and was able to achieve this:\n\n```\n   target_id  TMscore  target_frac_template\n0       9G4P  0.02752              0.000000\n1       9G4R  0.02251              0.000000\n2       9G4Q  0.02234              0.000000\n3       9IWF  0.02473              0.000000\n4       9J09  0.02990              0.000000\n5       9MMG  0.91128              0.894828\n6       9MME  0.02618              0.000000\n7       8VQV  0.92762              1.000000\n8       9E9Q  0.57674              1.000000\n9       9J4N  0.70860              1.000000\n10      9J4O  0.60490              0.876404\n11      9JGM  0.02941              0.000000\n12      9LKU  0.87714              0.953846\n```\n\nIn particular, precision is 100% – there were 6 targets with `target_frac_template`>0.45, and all 6 indeed had `TM-score`>0.45 when compared to the experimental structure.\n\n```\nMean TM-score: 0.3684\nPrecision: 1.0000 (6/6)\nRecall   : 1.0000 (6/6)\nF1 Score : 1.0000\n```\n\nThat's nice except that my notebook only found a template for 6 of the 13 cases, and its mean TM-score is only 0.368. \n\nMost Kaggle notebooks (including yours) will have higher mean TM-score -- can they also define a confidence score that would achieve 100% precision in predicting which targets have TM-score > 0.45?\n\nP.S. your question prompted me to find a bug in my notebook's comparison of predicted TM-score to actual TM-score. Thanks!",
      "votes": null
    },
    {
      "id": "3298549",
      "postDate": "10/05/2025 19:48:02",
      "content": "<h2>🧬 Template-Based vs. Protenix Model Comparison</h2>\n<p>Below is a comparison of <strong>TM-scores</strong> between the <strong>Template-Based Model (TBM)</strong> and <strong>Protenix</strong> across several targets.  <br>\nThe <strong><code>best_model</code></strong> column highlights which model achieved the higher TM-score for each case.</p>\n<table>\n<thead>\n<tr>\n<th>target_id</th>\n<th>TM_TBM</th>\n<th>TM_Protenix</th>\n<th>best_model</th>\n<th>best_TM</th>\n<th>target_frac_template</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>8VQV</td>\n<td>0.91708</td>\n<td>0.74352</td>\n<td><strong>TBM</strong></td>\n<td>0.91708</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9E9Q</td>\n<td>0.56803</td>\n<td>0.60023</td>\n<td><strong>Protenix</strong></td>\n<td>0.60023</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9G4P</td>\n<td>0.46631</td>\n<td>0.31505</td>\n<td><strong>TBM</strong></td>\n<td>0.46631</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9G4Q</td>\n<td>0.33465</td>\n<td>0.43069</td>\n<td><strong>Protenix</strong></td>\n<td>0.43069</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9G4R</td>\n<td>0.23963</td>\n<td>0.43340</td>\n<td><strong>Protenix</strong></td>\n<td>0.43340</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9IWF</td>\n<td>0.66022</td>\n<td>0.59655</td>\n<td><strong>TBM</strong></td>\n<td>0.66022</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9J09</td>\n<td>0.26842</td>\n<td>0.38660</td>\n<td><strong>Protenix</strong></td>\n<td>0.38660</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9J4N</td>\n<td>0.73014</td>\n<td>0.74350</td>\n<td><strong>Protenix</strong></td>\n<td>0.74350</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9J4O</td>\n<td>0.69472</td>\n<td>0.64824</td>\n<td><strong>TBM</strong></td>\n<td>0.69472</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9JGM</td>\n<td>0.66172</td>\n<td>0.76514</td>\n<td><strong>Protenix</strong></td>\n<td>0.76514</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9LKU</td>\n<td>0.83085</td>\n<td>0.86004</td>\n<td><strong>Protenix</strong></td>\n<td>0.86004</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9MME</td>\n<td>0.74523</td>\n<td>0.20920</td>\n<td><strong>TBM</strong></td>\n<td>0.74523</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9MMG</td>\n<td>0.91277</td>\n<td>0.24754</td>\n<td><strong>TBM</strong></td>\n<td>0.91277</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p>Precision: 0.7692<br>\nRecall: 1.0000<br>\nF1 Score: 0.8696</p>\n<h3>✅ Observations</h3>\n<ul>\n<li><strong>Protenix</strong> outperforms TBM on <strong>6 out of 13</strong> targets, particularly on <strong>complex or non-template-driven</strong> folds.  </li>\n<li><strong>TBM</strong> remains stronger on <strong>template-rich and well-structured</strong> targets (e.g., 8VQV, 9MMG).  </li>\n</ul>\n<h3>💡 Takeaway</h3>\n<p>A <strong>hybrid ensemble</strong> approach—using <strong>TBM for high-template coverage</strong> targets and <strong>Protenix for difficult/no-template</strong> ones—achieves more balanced and robust overall performance.  </p>\n<h3>Hybrid approach  (TBM and Proteinx)</h3>\n<p><em>TM score</em> <strong>Private</strong>: 0.61793  and <strong>public</strong> 0.65366 </p>",
      "rawMarkdown": "## 🧬 Template-Based vs. Protenix Model Comparison\n\nBelow is a comparison of **TM-scores** between the **Template-Based Model (TBM)** and **Protenix** across several targets.  \nThe **`best_model`** column highlights which model achieved the higher TM-score for each case.\n\n| target_id | TM_TBM | TM_Protenix | best_model | best_TM | target_frac_template |\n|:-----------|--------:|-------------:|:------------|---------:|----------------------:|\n| 8VQV | 0.91708 | 0.74352 | **TBM** | 0.91708 | 1 |\n| 9E9Q | 0.56803 | 0.60023 | **Protenix** | 0.60023 | 1 |\n| 9G4P | 0.46631 | 0.31505 | **TBM** | 0.46631 | 1 |\n| 9G4Q | 0.33465 | 0.43069 | **Protenix** | 0.43069 | 1 |\n| 9G4R | 0.23963 | 0.43340 | **Protenix** | 0.43340 | 1 |\n| 9IWF | 0.66022 | 0.59655 | **TBM** | 0.66022 | 1 |\n| 9J09 | 0.26842 | 0.38660 | **Protenix** | 0.38660 | 1 |\n| 9J4N | 0.73014 | 0.74350 | **Protenix** | 0.74350 | 1 |\n| 9J4O | 0.69472 | 0.64824 | **TBM** | 0.69472 | 1 |\n| 9JGM | 0.66172 | 0.76514 | **Protenix** | 0.76514 | 1 |\n| 9LKU | 0.83085 | 0.86004 | **Protenix** | 0.86004 | 1 |\n| 9MME | 0.74523 | 0.20920 | **TBM** | 0.74523 | 1 |\n| 9MMG | 0.91277 | 0.24754 | **TBM** | 0.91277 | 1 |\n\n---\nPrecision: 0.7692\nRecall: 1.0000\nF1 Score: 0.8696\n\n### ✅ Observations\n- **Protenix** outperforms TBM on **6 out of 13** targets, particularly on **complex or non-template-driven** folds.  \n- **TBM** remains stronger on **template-rich and well-structured** targets (e.g., 8VQV, 9MMG).  \n\n### 💡 Takeaway\nA **hybrid ensemble** approach—using **TBM for high-template coverage** targets and **Protenix for difficult/no-template** ones—achieves more balanced and robust overall performance.  \n\n### Hybrid approach  (TBM and Proteinx)\n*TM score* **Private**: 0.61793  and **public** 0.65366",
      "votes": null
    },
    {
      "id": "3298946",
      "postDate": "10/06/2025 19:23:22",
      "content": "<p>Arun, thanks! Is there some score you could compute in a blind setting (i.e. without knowing the experimental structure) that would allow for selection of all 10/13 targets with actual TM&gt;0.45 with 100% precision? </p>",
      "rawMarkdown": "Arun, thanks! Is there some score you could compute in a blind setting (i.e. without knowing the experimental structure) that would allow for selection of all 10/13 targets with actual TM>0.45 with 100% precision?",
      "votes": null
    },
    {
      "id": "3299043",
      "postDate": "10/07/2025 04:07:02",
      "content": "<p>Tanks  guys for the </p>",
      "rawMarkdown": "Tanks  guys for the",
      "votes": null
    },
    {
      "id": "3299315",
      "postDate": "10/07/2025 18:49:18",
      "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> I'm working on it</p>",
      "rawMarkdown": "rhijudas I'm working on it",
      "votes": null
    },
    {
      "id": "3301495",
      "postDate": "10/13/2025 13:59:24",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a>! </p>\n<p>Here's my attempt at the confidence scoring challenge.</p>\n<p>I developed a hybrid system that combines template-based search (MMseqs2) with a reference-based fallback, using multi-factor confidence scoring designed with a precision-first philosophy.</p>\n<p>Results on the 13 PDB summer 2025 targets:</p>\n<p>Mean TM-score: 0.5377 (+46% improvement over baseline 0.3684)<br>\nPrecision: 1.0000 (zero false positives!)<br>\nRecall: 0.7143 (5/7 targets with TM &gt; 0.45 correctly identified)<br>\nF1 Score: 0.8333</p>\n<p>The system correctly identifies 5 high-confidence targets (9MMG, 8VQV, 9J4N, 9LKU, 9E9Q) and appropriately flags 2 borderline cases (9MME and 9J4O) as low confidence due to genuine prediction uncertainty.</p>\n<p><a href=\"https://www.kaggle.com/code/fernandosr85/hybrid-confidence-scoring-system\" target=\"_blank\">https://www.kaggle.com/code/fernandosr85/hybrid-confidence-scoring-system</a></p>",
      "rawMarkdown": "Hi @rhijudas! \n\nHere's my attempt at the confidence scoring challenge.\n\nI developed a hybrid system that combines template-based search (MMseqs2) with a reference-based fallback, using multi-factor confidence scoring designed with a precision-first philosophy.\n\nResults on the 13 PDB summer 2025 targets:\n\nMean TM-score: 0.5377 (+46% improvement over baseline 0.3684)\nPrecision: 1.0000 (zero false positives!)\nRecall: 0.7143 (5/7 targets with TM > 0.45 correctly identified)\nF1 Score: 0.8333\n\nThe system correctly identifies 5 high-confidence targets (9MMG, 8VQV, 9J4N, 9LKU, 9E9Q) and appropriately flags 2 borderline cases (9MME and 9J4O) as low confidence due to genuine prediction uncertainty.\n\nhttps://www.kaggle.com/code/fernandosr85/hybrid-confidence-scoring-system",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3298525,
      "author_name": "raulenrique",
      "author_url": "",
      "post_date": "10/05/2025 18:35:44",
      "content": "<p>Hi, I'm not entirely sure what output format you would need. I ran a simulation on these 13 targets and got results like the following (see attachment)(the TM-scores were calculated only for informational purposes, using the tm_align library)</p>\n<table>\n<thead>\n<tr>\n<th>MAX_SCORE</th>\n<th>CONFIDENCE(SUM)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.488270057555273</td>\n<td>0.938955225050449</td>\n</tr>\n<tr>\n<td>0.331262390937843</td>\n<td>0.514765247702599</td>\n</tr>\n<tr>\n<td>0.370054298339892</td>\n<td>0.642883621156216</td>\n</tr>\n<tr>\n<td>0.283441785709035</td>\n<td>0.719378303736448</td>\n</tr>\n<tr>\n<td>0.249054987745784</td>\n<td>0.642032139003277</td>\n</tr>\n<tr>\n<td>0.798223644046592</td>\n<td>0.456463485956192</td>\n</tr>\n<tr>\n<td>0.636943080236158</td>\n<td>0.660994231700897</td>\n</tr>\n<tr>\n<td>0.548786849733515</td>\n<td>0.721843557432294</td>\n</tr>\n<tr>\n<td>0.344998093726893</td>\n<td>0.57373333349824</td>\n</tr>\n<tr>\n<td>0.741880618462633</td>\n<td>0.804388243705034</td>\n</tr>\n<tr>\n<td>0.625168415230373</td>\n<td>0.724774695932865</td>\n</tr>\n<tr>\n<td>0.317178276413226</td>\n<td>0.724590353667736</td>\n</tr>\n<tr>\n<td>0.763040445662887</td>\n<td>0.99974361560453</td>\n</tr>\n<tr>\n<td><strong>mean</strong></td>\n<td></td>\n</tr>\n<tr>\n<td>0.499869457215393</td>\n<td>0.701888158011291</td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": [
        {
          "id": 3298545,
          "author_name": "rhijudas",
          "author_url": "",
          "post_date": "10/05/2025 19:33:08",
          "content": "<p>Thanks! </p>\n<p>Do you see a way to define a confidence score such that when it is high, and only when it is high, it accurately predicts that the prediction has TM-score &gt; 0.45?</p>\n<p>For example in my <a href=\"https://www.kaggle.com/code/rhijudas/confidence-score-mmseqs2-3d-rna-template-id\" target=\"_blank\">notebook</a>, I calculated the fraction of the target covered by the template (<code>target_frac_template</code>), and was able to achieve this:</p>\n<pre><code>     TMscore  target_frac_template\n       G4P  .              .\n       G4R  .              .\n       G4Q  .              .\n       IWF  .              .\n       J09  .              .\n       MMG  .              .\n       MME  .              .\n       VQV  .              .\n       E9Q  .              .\n       J4N  .              .\n      J4O  .              .\n      JGM  .              .\n      LKU  .              .\n</code></pre>\n<p>In particular, precision is 100% – there were 6 targets with <code>target_frac_template</code>&gt;0.45, and all 6 indeed had <code>TM-score</code>&gt;0.45 when compared to the experimental structure.</p>\n<pre><code> TM-score: .\n: . (/)\n   : . (/)\n Score : .\n</code></pre>\n<p>That's nice except that my notebook only found a template for 6 of the 13 cases, and its mean TM-score is only 0.368. </p>\n<p>Most Kaggle notebooks (including yours) will have higher mean TM-score -- can they also define a confidence score that would achieve 100% precision in predicting which targets have TM-score &gt; 0.45?</p>\n<p>P.S. your question prompted me to find a bug in my notebook's comparison of predicted TM-score to actual TM-score. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3298549,
      "author_name": "arunodhayan",
      "author_url": "",
      "post_date": "10/05/2025 19:48:02",
      "content": "<h2>🧬 Template-Based vs. Protenix Model Comparison</h2>\n<p>Below is a comparison of <strong>TM-scores</strong> between the <strong>Template-Based Model (TBM)</strong> and <strong>Protenix</strong> across several targets.  <br>\nThe <strong><code>best_model</code></strong> column highlights which model achieved the higher TM-score for each case.</p>\n<table>\n<thead>\n<tr>\n<th>target_id</th>\n<th>TM_TBM</th>\n<th>TM_Protenix</th>\n<th>best_model</th>\n<th>best_TM</th>\n<th>target_frac_template</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>8VQV</td>\n<td>0.91708</td>\n<td>0.74352</td>\n<td><strong>TBM</strong></td>\n<td>0.91708</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9E9Q</td>\n<td>0.56803</td>\n<td>0.60023</td>\n<td><strong>Protenix</strong></td>\n<td>0.60023</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9G4P</td>\n<td>0.46631</td>\n<td>0.31505</td>\n<td><strong>TBM</strong></td>\n<td>0.46631</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9G4Q</td>\n<td>0.33465</td>\n<td>0.43069</td>\n<td><strong>Protenix</strong></td>\n<td>0.43069</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9G4R</td>\n<td>0.23963</td>\n<td>0.43340</td>\n<td><strong>Protenix</strong></td>\n<td>0.43340</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9IWF</td>\n<td>0.66022</td>\n<td>0.59655</td>\n<td><strong>TBM</strong></td>\n<td>0.66022</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9J09</td>\n<td>0.26842</td>\n<td>0.38660</td>\n<td><strong>Protenix</strong></td>\n<td>0.38660</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9J4N</td>\n<td>0.73014</td>\n<td>0.74350</td>\n<td><strong>Protenix</strong></td>\n<td>0.74350</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9J4O</td>\n<td>0.69472</td>\n<td>0.64824</td>\n<td><strong>TBM</strong></td>\n<td>0.69472</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9JGM</td>\n<td>0.66172</td>\n<td>0.76514</td>\n<td><strong>Protenix</strong></td>\n<td>0.76514</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9LKU</td>\n<td>0.83085</td>\n<td>0.86004</td>\n<td><strong>Protenix</strong></td>\n<td>0.86004</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9MME</td>\n<td>0.74523</td>\n<td>0.20920</td>\n<td><strong>TBM</strong></td>\n<td>0.74523</td>\n<td>1</td>\n</tr>\n<tr>\n<td>9MMG</td>\n<td>0.91277</td>\n<td>0.24754</td>\n<td><strong>TBM</strong></td>\n<td>0.91277</td>\n<td>1</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<p>Precision: 0.7692<br>\nRecall: 1.0000<br>\nF1 Score: 0.8696</p>\n<h3>✅ Observations</h3>\n<ul>\n<li><strong>Protenix</strong> outperforms TBM on <strong>6 out of 13</strong> targets, particularly on <strong>complex or non-template-driven</strong> folds.  </li>\n<li><strong>TBM</strong> remains stronger on <strong>template-rich and well-structured</strong> targets (e.g., 8VQV, 9MMG).  </li>\n</ul>\n<h3>💡 Takeaway</h3>\n<p>A <strong>hybrid ensemble</strong> approach—using <strong>TBM for high-template coverage</strong> targets and <strong>Protenix for difficult/no-template</strong> ones—achieves more balanced and robust overall performance.  </p>\n<h3>Hybrid approach  (TBM and Proteinx)</h3>\n<p><em>TM score</em> <strong>Private</strong>: 0.61793  and <strong>public</strong> 0.65366 </p>",
      "votes": null,
      "replies": [
        {
          "id": 3298946,
          "author_name": "rhijudas",
          "author_url": "",
          "post_date": "10/06/2025 19:23:22",
          "content": "<p>Arun, thanks! Is there some score you could compute in a blind setting (i.e. without knowing the experimental structure) that would allow for selection of all 10/13 targets with actual TM&gt;0.45 with 100% precision? </p>",
          "votes": null,
          "replies": [
            {
              "id": 3299315,
              "author_name": "arunodhayan",
              "author_url": "",
              "post_date": "10/07/2025 18:49:18",
              "content": "<p><a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a> I'm working on it</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3299043,
      "author_name": "phumlani93tshabalala",
      "author_url": "",
      "post_date": "10/07/2025 04:07:02",
      "content": "<p>Tanks  guys for the </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3301495,
      "author_name": "fernandosr85",
      "author_url": "",
      "post_date": "10/13/2025 13:59:24",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/rhijudas\" target=\"_blank\">@rhijudas</a>! </p>\n<p>Here's my attempt at the confidence scoring challenge.</p>\n<p>I developed a hybrid system that combines template-based search (MMseqs2) with a reference-based fallback, using multi-factor confidence scoring designed with a precision-first philosophy.</p>\n<p>Results on the 13 PDB summer 2025 targets:</p>\n<p>Mean TM-score: 0.5377 (+46% improvement over baseline 0.3684)<br>\nPrecision: 1.0000 (zero false positives!)<br>\nRecall: 0.7143 (5/7 targets with TM &gt; 0.45 correctly identified)<br>\nF1 Score: 0.8333</p>\n<p>The system correctly identifies 5 high-confidence targets (9MMG, 8VQV, 9J4N, 9LKU, 9E9Q) and appropriately flags 2 borderline cases (9MME and 9J4O) as low confidence due to genuine prediction uncertainty.</p>\n<p><a href=\"https://www.kaggle.com/code/fernandosr85/hybrid-confidence-scoring-system\" target=\"_blank\">https://www.kaggle.com/code/fernandosr85/hybrid-confidence-scoring-system</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3298264": "One of the key goals of 3D RNA structure prediction is to identify all the RNA segments across natural RNA sequences that might have well-defined 3D structure.  We'd like your help to get us there!\n\nWe'd like to get Kaggle codes run across all natural RNA sequences. For the highest confidence predictions that are novel, we could prioritize these segments for high resolution experimental 3D structure determination.\n\nWe're one step away.  The thing is: *we don't want a lot of false positives*. So we'd love to have each notebook give a confidence score that predicts the TM-score that the target would get when compared to an experimentally solved structure. \n\nWe'd especially like to identify cases where TM is predicted to be >0.45, the typical cutoff for a reasonable fold prediction.\n\nIf you'd like to help, take your favorite notebook and attach it to this data set\n\nhttps://www.kaggle.com/datasets/rhijudas/stanford-rna-3d-folding-pdb-summer2025\n\nIt holds 13 targets that have come out in the PDB since the close of the training phase.\n\nAnd then see if you can come up with a confidence score that achieves a high precision in predicting which of the 13 predictions actually has Tm>0.45.  \n\nHere's an example based on a notebook that runs template search with MMseqs2 and then estimates confidence based on the fraction of the target covered by the template: \n\nhttps://www.kaggle.com/code/rhijudas/confidence-score-mmseqs2-3d-rna-template-id\n \nThe notebook gets mean 3D accuracy of Tm = 0.368 and Precision, Recall, and F1 of confidence score accuracy of 1.0\n\nYour notebooks should do better in terms of 3D accuracy and, hopefully maintain precision of 1.0. Please post statistics and notebook links here if you make progress!\n\n(*Precision, recall, and F1 values updated after fix to notebook's comparison of ref and pred.*)",
    "3298525": "Hi, I'm not entirely sure what output format you would need. I ran a simulation on these 13 targets and got results like the following (see attachment)(the TM-scores were calculated only for informational purposes, using the tm_align library)\n\n| MAX_SCORE       | CONFIDENCE(SUM)   |\n|-----------------|-----------------|\n| 0.488270057555273 | 0.938955225050449 |\n| 0.331262390937843 | 0.514765247702599 |\n| 0.370054298339892 | 0.642883621156216 |\n| 0.283441785709035 | 0.719378303736448 |\n| 0.249054987745784 | 0.642032139003277 |\n| 0.798223644046592 | 0.456463485956192 |\n| 0.636943080236158 | 0.660994231700897 |\n| 0.548786849733515 | 0.721843557432294 |\n| 0.344998093726893 | 0.57373333349824  |\n| 0.741880618462633 | 0.804388243705034 |\n| 0.625168415230373 | 0.724774695932865 |\n| 0.317178276413226 | 0.724590353667736 |\n| 0.763040445662887 | 0.99974361560453  |\n| **mean**          |                 |\n| 0.499869457215393 | 0.701888158011291 |",
    "3298545": "Thanks! \n\nDo you see a way to define a confidence score such that when it is high, and only when it is high, it accurately predicts that the prediction has TM-score > 0.45?\n\nFor example in my [notebook](https://www.kaggle.com/code/rhijudas/confidence-score-mmseqs2-3d-rna-template-id), I calculated the fraction of the target covered by the template (`target_frac_template`), and was able to achieve this:\n\n```\n   target_id  TMscore  target_frac_template\n0       9G4P  0.02752              0.000000\n1       9G4R  0.02251              0.000000\n2       9G4Q  0.02234              0.000000\n3       9IWF  0.02473              0.000000\n4       9J09  0.02990              0.000000\n5       9MMG  0.91128              0.894828\n6       9MME  0.02618              0.000000\n7       8VQV  0.92762              1.000000\n8       9E9Q  0.57674              1.000000\n9       9J4N  0.70860              1.000000\n10      9J4O  0.60490              0.876404\n11      9JGM  0.02941              0.000000\n12      9LKU  0.87714              0.953846\n```\n\nIn particular, precision is 100% – there were 6 targets with `target_frac_template`>0.45, and all 6 indeed had `TM-score`>0.45 when compared to the experimental structure.\n\n```\nMean TM-score: 0.3684\nPrecision: 1.0000 (6/6)\nRecall   : 1.0000 (6/6)\nF1 Score : 1.0000\n```\n\nThat's nice except that my notebook only found a template for 6 of the 13 cases, and its mean TM-score is only 0.368. \n\nMost Kaggle notebooks (including yours) will have higher mean TM-score -- can they also define a confidence score that would achieve 100% precision in predicting which targets have TM-score > 0.45?\n\nP.S. your question prompted me to find a bug in my notebook's comparison of predicted TM-score to actual TM-score. Thanks!",
    "3298549": "## 🧬 Template-Based vs. Protenix Model Comparison\n\nBelow is a comparison of **TM-scores** between the **Template-Based Model (TBM)** and **Protenix** across several targets.  \nThe **`best_model`** column highlights which model achieved the higher TM-score for each case.\n\n| target_id | TM_TBM | TM_Protenix | best_model | best_TM | target_frac_template |\n|:-----------|--------:|-------------:|:------------|---------:|----------------------:|\n| 8VQV | 0.91708 | 0.74352 | **TBM** | 0.91708 | 1 |\n| 9E9Q | 0.56803 | 0.60023 | **Protenix** | 0.60023 | 1 |\n| 9G4P | 0.46631 | 0.31505 | **TBM** | 0.46631 | 1 |\n| 9G4Q | 0.33465 | 0.43069 | **Protenix** | 0.43069 | 1 |\n| 9G4R | 0.23963 | 0.43340 | **Protenix** | 0.43340 | 1 |\n| 9IWF | 0.66022 | 0.59655 | **TBM** | 0.66022 | 1 |\n| 9J09 | 0.26842 | 0.38660 | **Protenix** | 0.38660 | 1 |\n| 9J4N | 0.73014 | 0.74350 | **Protenix** | 0.74350 | 1 |\n| 9J4O | 0.69472 | 0.64824 | **TBM** | 0.69472 | 1 |\n| 9JGM | 0.66172 | 0.76514 | **Protenix** | 0.76514 | 1 |\n| 9LKU | 0.83085 | 0.86004 | **Protenix** | 0.86004 | 1 |\n| 9MME | 0.74523 | 0.20920 | **TBM** | 0.74523 | 1 |\n| 9MMG | 0.91277 | 0.24754 | **TBM** | 0.91277 | 1 |\n\n---\nPrecision: 0.7692\nRecall: 1.0000\nF1 Score: 0.8696\n\n### ✅ Observations\n- **Protenix** outperforms TBM on **6 out of 13** targets, particularly on **complex or non-template-driven** folds.  \n- **TBM** remains stronger on **template-rich and well-structured** targets (e.g., 8VQV, 9MMG).  \n\n### 💡 Takeaway\nA **hybrid ensemble** approach—using **TBM for high-template coverage** targets and **Protenix for difficult/no-template** ones—achieves more balanced and robust overall performance.  \n\n### Hybrid approach  (TBM and Proteinx)\n*TM score* **Private**: 0.61793  and **public** 0.65366",
    "3298946": "Arun, thanks! Is there some score you could compute in a blind setting (i.e. without knowing the experimental structure) that would allow for selection of all 10/13 targets with actual TM>0.45 with 100% precision?",
    "3299043": "Tanks  guys for the",
    "3299315": "rhijudas I'm working on it",
    "3301495": "Hi @rhijudas! \n\nHere's my attempt at the confidence scoring challenge.\n\nI developed a hybrid system that combines template-based search (MMseqs2) with a reference-based fallback, using multi-factor confidence scoring designed with a precision-first philosophy.\n\nResults on the 13 PDB summer 2025 targets:\n\nMean TM-score: 0.5377 (+46% improvement over baseline 0.3684)\nPrecision: 1.0000 (zero false positives!)\nRecall: 0.7143 (5/7 targets with TM > 0.45 correctly identified)\nF1 Score: 0.8333\n\nThe system correctly identifies 5 high-confidence targets (9MMG, 8VQV, 9J4N, 9LKU, 9E9Q) and appropriately flags 2 borderline cases (9MME and 9J4O) as low confidence due to genuine prediction uncertainty.\n\nhttps://www.kaggle.com/code/fernandosr85/hybrid-confidence-scoring-system"
  },
  "source": "meta"
}