{
  "id": 568220,
  "title": "TM Score - Is the Usalign executable of the notebook modified ?",
  "url": "/competitions/stanford-rna-3d-folding/discussion/568220",
  "author_name": "",
  "post_date": "2025-03-14T16:17:02.145822700Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I am trying to replicate the evaluation pipeline outside of Kaggle.</p>\n<hr>\n<p>In the metrics section we can read :</p>\n<blockquote>\n  <p>d0 is a distance scaling factor in Angstroms, defined as:</p>\n  <p>𝑑0=0.6(𝐿ref−0.5)1/2−2.5<br>\n  for Lref ≥ 30; and d0 = 0.3, 0.4, 0.5, 0.6, or 0.7 for Lref &lt;12, 12-15, 16-19, 20-23, or 24-29, respectively.</p>\n  <p>The rotation and translation of predicted structures to align with experimental reference structures are carried out by US-align. To match default settings, as used in the CASP competitions, the alignment will be sequence-independent.</p>\n  <p>For each target RNA sequence, you will submit 5 predictions and your final score will be the average of best-of-5 TM-scores of all targets. For a few targets, multiple slightly different structures have been captured experimentally; your predictions' scores will be based on the best TM-score compared to each of these reference structures.</p>\n</blockquote>\n<p>As showed by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> the implementation of USalign uses other parameters :</p>\n<pre><code>{\n    D0_MIN=; \n\n    Lnorm=len;            \n    (Lnorm&lt;=) d0=;\n     (Lnorm&gt;&amp;&amp;Lnorm&lt;=) d0=;\n     (Lnorm&gt;&amp;&amp;Lnorm&lt;=) d0=;\n     (Lnorm&gt;&amp;&amp;Lnorm&lt;=) d0=;\n     (Lnorm&gt;&amp;&amp;Lnorm&lt;)  d0=;\n     d0=(*((Lnorm*), /));\n\n    d0_search=d0;\n     (d0_search&gt;)   d0_search=;\n     (d0_search&lt;) d0_search=;\n}\n</code></pre>\n<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> The metrics notebook provided by the hosts come with a dataset that contains a build of USalign. It is a modified version of USalign with the announced parameters, or is it a non-modified version ? (in that case which parameters should we use?)</p>\n<p>Thanks</p>",
  "messages": [
    {
      "id": "3149727",
      "postDate": "03/14/2025 16:17:02",
      "content": "<p>I am trying to replicate the evaluation pipeline outside of Kaggle.</p>\n<hr>\n<p>In the metrics section we can read :</p>\n<blockquote>\n  <p>d0 is a distance scaling factor in Angstroms, defined as:</p>\n  <p>𝑑0=0.6(𝐿ref−0.5)1/2−2.5<br>\n  for Lref ≥ 30; and d0 = 0.3, 0.4, 0.5, 0.6, or 0.7 for Lref &lt;12, 12-15, 16-19, 20-23, or 24-29, respectively.</p>\n  <p>The rotation and translation of predicted structures to align with experimental reference structures are carried out by US-align. To match default settings, as used in the CASP competitions, the alignment will be sequence-independent.</p>\n  <p>For each target RNA sequence, you will submit 5 predictions and your final score will be the average of best-of-5 TM-scores of all targets. For a few targets, multiple slightly different structures have been captured experimentally; your predictions' scores will be based on the best TM-score compared to each of these reference structures.</p>\n</blockquote>\n<p>As showed by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> the implementation of USalign uses other parameters :</p>\n<pre><code>{\n    D0_MIN=; \n\n    Lnorm=len;            \n    (Lnorm&lt;=) d0=;\n     (Lnorm&gt;&amp;&amp;Lnorm&lt;=) d0=;\n     (Lnorm&gt;&amp;&amp;Lnorm&lt;=) d0=;\n     (Lnorm&gt;&amp;&amp;Lnorm&lt;=) d0=;\n     (Lnorm&gt;&amp;&amp;Lnorm&lt;)  d0=;\n     d0=(*((Lnorm*), /));\n\n    d0_search=d0;\n     (d0_search&gt;)   d0_search=;\n     (d0_search&lt;) d0_search=;\n}\n</code></pre>\n<p><a href=\"https://www.kaggle.com/shujun717\" target=\"_blank\">@shujun717</a> The metrics notebook provided by the hosts come with a dataset that contains a build of USalign. It is a modified version of USalign with the announced parameters, or is it a non-modified version ? (in that case which parameters should we use?)</p>\n<p>Thanks</p>",
      "rawMarkdown": "I am trying to replicate the evaluation pipeline outside of Kaggle.\n\n---\n\nIn the metrics section we can read :\n\n>d0 is a distance scaling factor in Angstroms, defined as:\n>\n> 𝑑0=0.6(𝐿ref−0.5)1/2−2.5\nfor Lref ≥ 30; and d0 = 0.3, 0.4, 0.5, 0.6, or 0.7 for Lref <12, 12-15, 16-19, 20-23, or 24-29, respectively.\n\n>The rotation and translation of predicted structures to align with experimental reference structures are carried out by US-align. To match default settings, as used in the CASP competitions, the alignment will be sequence-independent.\n\n>For each target RNA sequence, you will submit 5 predictions and your final score will be the average of best-of-5 TM-scores of all targets. For a few targets, multiple slightly different structures have been captured experimentally; your predictions' scores will be based on the best TM-score compared to each of these reference structures.\n\nAs showed by @hengck23 the implementation of USalign uses other parameters :\n\n```cpp\nvoid parameter_set4final_C3prime(const double len, double &D0_MIN,\n    double &Lnorm, double &d0, double &d0_search)\n{\n    D0_MIN=0.3; \n \n    Lnorm=len;            //normalize TMscore by this in searching\n    if(Lnorm<=11) d0=0.3;\n    else if(Lnorm>11&&Lnorm<=15) d0=0.4;\n    else if(Lnorm>15&&Lnorm<=19) d0=0.5;\n    else if(Lnorm>19&&Lnorm<=23) d0=0.6;\n    else if(Lnorm>23&&Lnorm<30)  d0=0.7;\n    else d0=(0.6*pow((Lnorm*1.0-0.5), 1.0/2)-2.5);\n\n    d0_search=d0;\n    if (d0_search>8)   d0_search=8;\n    if (d0_search<4.5) d0_search=4.5;\n}\n```\n\n@shujun717 The metrics notebook provided by the hosts come with a dataset that contains a build of USalign. It is a modified version of USalign with the announced parameters, or is it a non-modified version ? (in that case which parameters should we use?)\n\nThanks",
      "votes": null
    },
    {
      "id": "3149869",
      "postDate": "03/14/2025 19:20:37",
      "content": "<p>i am sorry that i have quote the wrong function.<br>\ni have corrected the previous post:</p>\n<pre><code> standard_TMscore( **r1,  **r2,  **xtm,  **ytm,\n     **xt,  **x,  **y,  xlen,  ylen,  invmap[],\n    &amp; L_ali, &amp; RMSD,  D0_MIN,  Lnorm,  d0,\n     d0_search,  score_d8,  t[],  u[][],\n      mol_type)\n{\n    D0_MIN = ;\n    Lnorm = ylen;\n     (mol_type&gt;) \n    {\n             (Lnorm&lt;=) d0=; \n         (Lnorm&gt; &amp;&amp; Lnorm&lt;=) d0=;\n         (Lnorm&gt; &amp;&amp; Lnorm&lt;=) d0=;\n         (Lnorm&gt; &amp;&amp; Lnorm&lt;=) d0=;\n         (Lnorm&gt; &amp;&amp; Lnorm&lt;)  d0=;\n         d0=(*pow((Lnorm*), /));\n    }\n    \n    {\n         (Lnorm &gt; ) d0=(*pow((Lnorm*), /) );\n         d0 = D0_MIN;\n         (d0 &lt; D0_MIN) d0 = D0_MIN;\n    }\n     d0_input = d0;\n</code></pre>\n<p>i verified that usalign compiled by kaggle and by my local machine gives exact results.<br>\n(you can run download and run the given kaggle binary)</p>",
      "rawMarkdown": "i am sorry that i have quote the wrong function.\ni have corrected the previous post:\n\n```\n\n\ndouble standard_TMscore(double **r1, double **r2, double **xtm, double **ytm,\n    double **xt, double **x, double **y, int xlen, int ylen, int invmap[],\n    int& L_ali, double& RMSD, double D0_MIN, double Lnorm, double d0,\n    double d0_search, double score_d8, double t[3], double u[3][3],\n    const int mol_type)\n{\n    D0_MIN = 0.5;\n    Lnorm = ylen;\n    if (mol_type>0) // RNA\n    {\n        if     (Lnorm<=11) d0=0.3; \n        else if(Lnorm>11 && Lnorm<=15) d0=0.4;\n        else if(Lnorm>15 && Lnorm<=19) d0=0.5;\n        else if(Lnorm>19 && Lnorm<=23) d0=0.6;\n        else if(Lnorm>23 && Lnorm<30)  d0=0.7;\n        else d0=(0.6*pow((Lnorm*1.0-0.5), 1.0/2)-2.5);\n    }\n    else\n    {\n        if (Lnorm > 21) d0=(1.24*pow((Lnorm*1.0-15), 1.0/3) -1.8);\n        else d0 = D0_MIN;\n        if (d0 < D0_MIN) d0 = D0_MIN;\n    }\n    double d0_input = d0;// Scaled by seq_min\n\n```\n\ni verified that usalign compiled by kaggle and by my local machine gives exact results.\n(you can run download and run the given kaggle binary)",
      "votes": null
    },
    {
      "id": "3149927",
      "postDate": "03/14/2025 21:17:49",
      "content": "<p>I am on Mac so I need to rebuild the binary ;) thanks for your feedback</p>",
      "rawMarkdown": "I am on Mac so I need to rebuild the binary ;) thanks for your feedback",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3149869,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/14/2025 19:20:37",
      "content": "<p>i am sorry that i have quote the wrong function.<br>\ni have corrected the previous post:</p>\n<pre><code> standard_TMscore( **r1,  **r2,  **xtm,  **ytm,\n     **xt,  **x,  **y,  xlen,  ylen,  invmap[],\n    &amp; L_ali, &amp; RMSD,  D0_MIN,  Lnorm,  d0,\n     d0_search,  score_d8,  t[],  u[][],\n      mol_type)\n{\n    D0_MIN = ;\n    Lnorm = ylen;\n     (mol_type&gt;) \n    {\n             (Lnorm&lt;=) d0=; \n         (Lnorm&gt; &amp;&amp; Lnorm&lt;=) d0=;\n         (Lnorm&gt; &amp;&amp; Lnorm&lt;=) d0=;\n         (Lnorm&gt; &amp;&amp; Lnorm&lt;=) d0=;\n         (Lnorm&gt; &amp;&amp; Lnorm&lt;)  d0=;\n         d0=(*pow((Lnorm*), /));\n    }\n    \n    {\n         (Lnorm &gt; ) d0=(*pow((Lnorm*), /) );\n         d0 = D0_MIN;\n         (d0 &lt; D0_MIN) d0 = D0_MIN;\n    }\n     d0_input = d0;\n</code></pre>\n<p>i verified that usalign compiled by kaggle and by my local machine gives exact results.<br>\n(you can run download and run the given kaggle binary)</p>",
      "votes": null,
      "replies": [
        {
          "id": 3149927,
          "author_name": "louisstefanuto",
          "author_url": "",
          "post_date": "03/14/2025 21:17:49",
          "content": "<p>I am on Mac so I need to rebuild the binary ;) thanks for your feedback</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3149727": "I am trying to replicate the evaluation pipeline outside of Kaggle.\n\n---\n\nIn the metrics section we can read :\n\n>d0 is a distance scaling factor in Angstroms, defined as:\n>\n> 𝑑0=0.6(𝐿ref−0.5)1/2−2.5\nfor Lref ≥ 30; and d0 = 0.3, 0.4, 0.5, 0.6, or 0.7 for Lref <12, 12-15, 16-19, 20-23, or 24-29, respectively.\n\n>The rotation and translation of predicted structures to align with experimental reference structures are carried out by US-align. To match default settings, as used in the CASP competitions, the alignment will be sequence-independent.\n\n>For each target RNA sequence, you will submit 5 predictions and your final score will be the average of best-of-5 TM-scores of all targets. For a few targets, multiple slightly different structures have been captured experimentally; your predictions' scores will be based on the best TM-score compared to each of these reference structures.\n\nAs showed by @hengck23 the implementation of USalign uses other parameters :\n\n```cpp\nvoid parameter_set4final_C3prime(const double len, double &D0_MIN,\n    double &Lnorm, double &d0, double &d0_search)\n{\n    D0_MIN=0.3; \n \n    Lnorm=len;            //normalize TMscore by this in searching\n    if(Lnorm<=11) d0=0.3;\n    else if(Lnorm>11&&Lnorm<=15) d0=0.4;\n    else if(Lnorm>15&&Lnorm<=19) d0=0.5;\n    else if(Lnorm>19&&Lnorm<=23) d0=0.6;\n    else if(Lnorm>23&&Lnorm<30)  d0=0.7;\n    else d0=(0.6*pow((Lnorm*1.0-0.5), 1.0/2)-2.5);\n\n    d0_search=d0;\n    if (d0_search>8)   d0_search=8;\n    if (d0_search<4.5) d0_search=4.5;\n}\n```\n\n@shujun717 The metrics notebook provided by the hosts come with a dataset that contains a build of USalign. It is a modified version of USalign with the announced parameters, or is it a non-modified version ? (in that case which parameters should we use?)\n\nThanks",
    "3149869": "i am sorry that i have quote the wrong function.\ni have corrected the previous post:\n\n```\n\n\ndouble standard_TMscore(double **r1, double **r2, double **xtm, double **ytm,\n    double **xt, double **x, double **y, int xlen, int ylen, int invmap[],\n    int& L_ali, double& RMSD, double D0_MIN, double Lnorm, double d0,\n    double d0_search, double score_d8, double t[3], double u[3][3],\n    const int mol_type)\n{\n    D0_MIN = 0.5;\n    Lnorm = ylen;\n    if (mol_type>0) // RNA\n    {\n        if     (Lnorm<=11) d0=0.3; \n        else if(Lnorm>11 && Lnorm<=15) d0=0.4;\n        else if(Lnorm>15 && Lnorm<=19) d0=0.5;\n        else if(Lnorm>19 && Lnorm<=23) d0=0.6;\n        else if(Lnorm>23 && Lnorm<30)  d0=0.7;\n        else d0=(0.6*pow((Lnorm*1.0-0.5), 1.0/2)-2.5);\n    }\n    else\n    {\n        if (Lnorm > 21) d0=(1.24*pow((Lnorm*1.0-15), 1.0/3) -1.8);\n        else d0 = D0_MIN;\n        if (d0 < D0_MIN) d0 = D0_MIN;\n    }\n    double d0_input = d0;// Scaled by seq_min\n\n```\n\ni verified that usalign compiled by kaggle and by my local machine gives exact results.\n(you can run download and run the given kaggle binary)",
    "3149927": "I am on Mac so I need to rebuild the binary ;) thanks for your feedback"
  },
  "source": "meta"
}