{
  "id": 460172,
  "title": "Best train-data-only scores?",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/460172",
  "author_name": "",
  "post_date": "2023-12-08T05:08:58.929162800Z",
  "votes": 7,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I was very suspicious of BPP, so for the final submission, I had two ensembles, one with only train-data models and one that included BPP. I would like to know if anyone could score better using only the train data. I got 0.14683 for my ensemble (30 models). Again, only train data, no bpp, CapR, etc. Good thing I also submitted a bpp-included ensemble, lol</p>",
  "messages": [
    {
      "id": "2553220",
      "postDate": "12/08/2023 05:08:58",
      "content": "<p>I was very suspicious of BPP, so for the final submission, I had two ensembles, one with only train-data models and one that included BPP. I would like to know if anyone could score better using only the train data. I got 0.14683 for my ensemble (30 models). Again, only train data, no bpp, CapR, etc. Good thing I also submitted a bpp-included ensemble, lol</p>",
      "rawMarkdown": "I was very suspicious of BPP, so for the final submission, I had two ensembles, one with only train-data models and one that included BPP. I would like to know if anyone could score better using only the train data. I got 0.14683 for my ensemble (30 models). Again, only train data, no bpp, CapR, etc. Good thing I also submitted a bpp-included ensemble, lol",
      "votes": null
    },
    {
      "id": "2553493",
      "postDate": "12/08/2023 09:52:23",
      "content": "<p>Currently writing the solution, but a SINGLE model that is highly unfinished because of time constraints and high complexity/training requirements is around  0.13709 public/ 0.14366 private. This is score without any BPP's, any augmentation, no pre or post processing and, and only training on SN_filter&gt;1 data (so no pseudo-labeling or low SN_filter data correction), just simply feeding the sequence and getting reactivity estimates. This model was also underfitted, as training it in the last day only by extending epochs from last checkpoint provided better results with no signs of overfitting.</p>",
      "rawMarkdown": "Currently writing the solution, but a SINGLE model that is highly unfinished because of time constraints and high complexity/training requirements is around  0.13709 public/ 0.14366 private. This is score without any BPP's, any augmentation, no pre or post processing and, and only training on SN_filter>1 data (so no pseudo-labeling or low SN_filter data correction), just simply feeding the sequence and getting reactivity estimates. This model was also underfitted, as training it in the last day only by extending epochs from last checkpoint provided better results with no signs of overfitting.",
      "votes": null
    },
    {
      "id": "2553514",
      "postDate": "12/08/2023 10:14:46",
      "content": "<p>Damn, this is an amazing result. I hope you will publish it, I really want to see it.<br>\n(imho much more impressive than bpp-based top solutions tbh)</p>",
      "rawMarkdown": "Damn, this is an amazing result. I hope you will publish it, I really want to see it.\n(imho much more impressive than bpp-based top solutions tbh)",
      "votes": null
    },
    {
      "id": "2553889",
      "postDate": "12/08/2023 15:30:26",
      "content": "<p>Is it a deberta?</p>",
      "rawMarkdown": "Is it a deberta?",
      "votes": null
    },
    {
      "id": "2554076",
      "postDate": "12/08/2023 19:02:02",
      "content": "<p>It is a modified twin-tower architecture based on AlphaFold and its derivatives OpenComplex/RhoFold.</p>",
      "rawMarkdown": "It is a modified twin-tower architecture based on AlphaFold and its derivatives OpenComplex/RhoFold.",
      "votes": null
    },
    {
      "id": "2554193",
      "postDate": "12/08/2023 22:04:03",
      "content": "<p>I tried deberta but it didn't generalize well to longer sequences, have you been able to make it work? </p>",
      "rawMarkdown": "I tried deberta but it didn't generalize well to longer sequences, have you been able to make it work?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2553493,
      "author_name": "dankrstev",
      "author_url": "",
      "post_date": "12/08/2023 09:52:23",
      "content": "<p>Currently writing the solution, but a SINGLE model that is highly unfinished because of time constraints and high complexity/training requirements is around  0.13709 public/ 0.14366 private. This is score without any BPP's, any augmentation, no pre or post processing and, and only training on SN_filter&gt;1 data (so no pseudo-labeling or low SN_filter data correction), just simply feeding the sequence and getting reactivity estimates. This model was also underfitted, as training it in the last day only by extending epochs from last checkpoint provided better results with no signs of overfitting.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2553514,
          "author_name": "shlomoron",
          "author_url": "",
          "post_date": "12/08/2023 10:14:46",
          "content": "<p>Damn, this is an amazing result. I hope you will publish it, I really want to see it.<br>\n(imho much more impressive than bpp-based top solutions tbh)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2553889,
          "author_name": "shujun717",
          "author_url": "",
          "post_date": "12/08/2023 15:30:26",
          "content": "<p>Is it a deberta?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2554076,
              "author_name": "dankrstev",
              "author_url": "",
              "post_date": "12/08/2023 19:02:02",
              "content": "<p>It is a modified twin-tower architecture based on AlphaFold and its derivatives OpenComplex/RhoFold.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2554193,
              "author_name": "thedrcat",
              "author_url": "",
              "post_date": "12/08/2023 22:04:03",
              "content": "<p>I tried deberta but it didn't generalize well to longer sequences, have you been able to make it work? </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2553220": "I was very suspicious of BPP, so for the final submission, I had two ensembles, one with only train-data models and one that included BPP. I would like to know if anyone could score better using only the train data. I got 0.14683 for my ensemble (30 models). Again, only train data, no bpp, CapR, etc. Good thing I also submitted a bpp-included ensemble, lol",
    "2553493": "Currently writing the solution, but a SINGLE model that is highly unfinished because of time constraints and high complexity/training requirements is around  0.13709 public/ 0.14366 private. This is score without any BPP's, any augmentation, no pre or post processing and, and only training on SN_filter>1 data (so no pseudo-labeling or low SN_filter data correction), just simply feeding the sequence and getting reactivity estimates. This model was also underfitted, as training it in the last day only by extending epochs from last checkpoint provided better results with no signs of overfitting.",
    "2553514": "Damn, this is an amazing result. I hope you will publish it, I really want to see it.\n(imho much more impressive than bpp-based top solutions tbh)",
    "2553889": "Is it a deberta?",
    "2554076": "It is a modified twin-tower architecture based on AlphaFold and its derivatives OpenComplex/RhoFold.",
    "2554193": "I tried deberta but it didn't generalize well to longer sequences, have you been able to make it work?"
  },
  "source": "meta"
}