{
  "id": 456079,
  "title": "Anyone tried augmented bpp's from vienna, eternafold and others?",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/456079",
  "author_name": "",
  "post_date": "2023-11-18T00:18:42.806535300Z",
  "votes": 6,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Pretty much same as the title. <br>\nIn my experiments, only thing that boosts is BPP. <br>\nI've not used arnie to generate other bpp's till now.</p>",
  "messages": [
    {
      "id": "2529109",
      "postDate": "11/18/2023 00:18:42",
      "content": "<p>Pretty much same as the title. <br>\nIn my experiments, only thing that boosts is BPP. <br>\nI've not used arnie to generate other bpp's till now.</p>",
      "rawMarkdown": "Pretty much same as the title. \nIn my experiments, only thing that boosts is BPP. \nI've not used arnie to generate other bpp's till now.",
      "votes": null
    },
    {
      "id": "2532272",
      "postDate": "11/20/2023 22:39:09",
      "content": "<p>How are you embeddings the [seq_length, seq_length] tensor into your sequence model? Do you have any tips to change this 2D tensor into 1D for input to a sequence model?</p>",
      "rawMarkdown": "How are you embeddings the [seq_length, seq_length] tensor into your sequence model? Do you have any tips to change this 2D tensor into 1D for input to a sequence model?",
      "votes": null
    },
    {
      "id": "2532284",
      "postDate": "11/20/2023 22:50:29",
      "content": "<p>Well, when you think about it, the BPP values are really just attentions scores, no? ;-)</p>",
      "rawMarkdown": "Well, when you think about it, the BPP values are really just attentions scores, no? ;-)",
      "votes": null
    },
    {
      "id": "2532802",
      "postDate": "11/21/2023 10:40:54",
      "content": "<p>In addition to what <a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a> said, you can take sum/max/second max, etc., across the columns. This is the simplest way to get a [seq_length, #features] from [seq_length, seq_length] and will give you an immediate boost, even if not as good as more sophisticated ways.</p>",
      "rawMarkdown": "In addition to what @fnands said, you can take sum/max/second max, etc., across the columns. This is the simplest way to get a [seq_length, #features] from [seq_length, seq_length] and will give you an immediate boost, even if not as good as more sophisticated ways.",
      "votes": null
    },
    {
      "id": "2533382",
      "postDate": "11/21/2023 20:56:50",
      "content": "<p>I have the same experience.  This has been troubling me a lot.  I've also been optimizing secondary structures considering the DMS and 2A3 scores, which basically leaks information into cross validation.  Still it's not as good as using BPP in cross-validation.</p>\n<p>I think what I observed translates to this:  RNA folding has multiple conformations and considering only a single conformation is not enough.</p>\n<p>This competition is rather challenging.</p>",
      "rawMarkdown": "I have the same experience.  This has been troubling me a lot.  I've also been optimizing secondary structures considering the DMS and 2A3 scores, which basically leaks information into cross validation.  Still it's not as good as using BPP in cross-validation.\n\nI think what I observed translates to this:  RNA folding has multiple conformations and considering only a single conformation is not enough.\n\nThis competition is rather challenging.",
      "votes": null
    },
    {
      "id": "2535693",
      "postDate": "11/23/2023 14:35:19",
      "content": "<p>Thanks for the advice!</p>",
      "rawMarkdown": "Thanks for the advice!",
      "votes": null
    },
    {
      "id": "2538273",
      "postDate": "11/26/2023 00:51:29",
      "content": "<p>Im sorry, pretty sure i've misunderstood something, we don't have bpps for the hidden lb data right? How are you using bpps as additional features then?</p>",
      "rawMarkdown": "Im sorry, pretty sure i've misunderstood something, we don't have bpps for the hidden lb data right? How are you using bpps as additional features then?",
      "votes": null
    },
    {
      "id": "2538530",
      "postDate": "11/26/2023 08:16:08",
      "content": "<p>From the Data page -</p>\n<p>\"Ribonanza_bpp_files - TXT files listing position pairs predicted to have non-zero Watson-Crick base pair probabilities by the LinearPartition-EternaFold package. Files are given for train and test sequences, indexed by sequence_id\"</p>",
      "rawMarkdown": "From the Data page -\n\n\"Ribonanza_bpp_files - TXT files listing position pairs predicted to have non-zero Watson-Crick base pair probabilities by the LinearPartition-EternaFold package. Files are given for train and test sequences, indexed by sequence_id\"",
      "votes": null
    },
    {
      "id": "2538547",
      "postDate": "11/26/2023 08:52:56",
      "content": "<p>But</p>\n<blockquote>\n  <p>Within this set, the majority (1,008,000 RNAs) will be experimentally synthesized and profiled after the Kaggle competition begins</p>\n</blockquote>\n<p>So will the bpps be updated on the hidden data too? </p>",
      "rawMarkdown": "But\n> Within this set, the majority (1,008,000 RNAs) will be experimentally synthesized and profiled after the Kaggle competition begins\n\nSo will the bpps be updated on the hidden data too?",
      "votes": null
    },
    {
      "id": "2538568",
      "postDate": "11/26/2023 09:17:33",
      "content": "<p>This is not a code competition. The test data is not hidden and we know exactly what are the sequences that we are trying to predict, thus we can generate and use any feature we want.</p>",
      "rawMarkdown": "This is not a code competition. The test data is not hidden and we know exactly what are the sequences that we are trying to predict, thus we can generate and use any feature we want.",
      "votes": null
    },
    {
      "id": "2538570",
      "postDate": "11/26/2023 09:20:15",
      "content": "<p>see <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/437481#2441007\" target=\"_blank\">posts around here</a><br>\nfrom the host:<br>\n\"After we get these new data, we'll rescore the existing submission files. Note that although data will be new, they will be for private leaderboard* test sequences for which the submissions already include predictions,\"</p>\n<p>\"There are numerous sequences in test_sequences.csv for which we are asking for submissions. But we don't have the data to evaluate your submissions for those sequences at the start of the competition. \"</p>\n<p>so the sequence ids are there for test in submissions, in bpps, etc. but the data to do the evaluation, the reactivity will be ongoing in the competition. </p>",
      "rawMarkdown": "see [posts around here](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/437481#2441007)\nfrom the host:\n\"After we get these new data, we'll rescore the existing submission files. Note that although data will be new, they will be for private leaderboard* test sequences for which the submissions already include predictions,\"\n\n\"There are numerous sequences in test_sequences.csv for which we are asking for submissions. But we don't have the data to evaluate your submissions for those sequences at the start of the competition. \"\n\nso the sequence ids are there for test in submissions, in bpps, etc. but the data to do the evaluation, the reactivity will be ongoing in the competition.",
      "votes": null
    },
    {
      "id": "2538671",
      "postDate": "11/26/2023 11:16:10",
      "content": "<p>The host refers to the labels. The sequences are the same. Note that you can submit a csv file without any code. As I said, this is not a code competition.</p>",
      "rawMarkdown": "The host refers to the labels. The sequences are the same. Note that you can submit a csv file without any code. As I said, this is not a code competition.",
      "votes": null
    },
    {
      "id": "2538697",
      "postDate": "11/26/2023 11:52:02",
      "content": "<p>thank you!</p>",
      "rawMarkdown": "thank you!",
      "votes": null
    },
    {
      "id": "2538725",
      "postDate": "11/26/2023 12:38:15",
      "content": "<p>The BPPs are not extracted from experiments, they are predicted with EternaFold. </p>\n<p>From the data page: </p>\n<blockquote>\n  <p>on-zero Watson-Crick base pair probabilities by the LinearPartition-EternaFold package</p>\n</blockquote>",
      "rawMarkdown": "The BPPs are not extracted from experiments, they are predicted with EternaFold. \n\nFrom the data page: \n>on-zero Watson-Crick base pair probabilities by the LinearPartition-EternaFold package",
      "votes": null
    },
    {
      "id": "2539218",
      "postDate": "11/26/2023 20:50:41",
      "content": "<p>A followup question. Would you recommend embedding the index of the other pair or just the float probability? </p>",
      "rawMarkdown": "A followup question. Would you recommend embedding the index of the other pair or just the float probability?",
      "votes": null
    },
    {
      "id": "2539222",
      "postDate": "11/26/2023 20:54:05",
      "content": "<p>Try and find out! :)</p>",
      "rawMarkdown": "Try and find out! :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2532272,
      "author_name": "themecheng",
      "author_url": "",
      "post_date": "11/20/2023 22:39:09",
      "content": "<p>How are you embeddings the [seq_length, seq_length] tensor into your sequence model? Do you have any tips to change this 2D tensor into 1D for input to a sequence model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2532284,
          "author_name": "fnands",
          "author_url": "",
          "post_date": "11/20/2023 22:50:29",
          "content": "<p>Well, when you think about it, the BPP values are really just attentions scores, no? ;-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2532802,
          "author_name": "shlomoron",
          "author_url": "",
          "post_date": "11/21/2023 10:40:54",
          "content": "<p>In addition to what <a href=\"https://www.kaggle.com/fnands\" target=\"_blank\">@fnands</a> said, you can take sum/max/second max, etc., across the columns. This is the simplest way to get a [seq_length, #features] from [seq_length, seq_length] and will give you an immediate boost, even if not as good as more sophisticated ways.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2535693,
              "author_name": "themecheng",
              "author_url": "",
              "post_date": "11/23/2023 14:35:19",
              "content": "<p>Thanks for the advice!</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2539218,
              "author_name": "themecheng",
              "author_url": "",
              "post_date": "11/26/2023 20:50:41",
              "content": "<p>A followup question. Would you recommend embedding the index of the other pair or just the float probability? </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2539222,
                  "author_name": "shlomoron",
                  "author_url": "",
                  "post_date": "11/26/2023 20:54:05",
                  "content": "<p>Try and find out! :)</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2533382,
      "author_name": "aaalgo",
      "author_url": "",
      "post_date": "11/21/2023 20:56:50",
      "content": "<p>I have the same experience.  This has been troubling me a lot.  I've also been optimizing secondary structures considering the DMS and 2A3 scores, which basically leaks information into cross validation.  Still it's not as good as using BPP in cross-validation.</p>\n<p>I think what I observed translates to this:  RNA folding has multiple conformations and considering only a single conformation is not enough.</p>\n<p>This competition is rather challenging.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2538273,
      "author_name": "jainam213",
      "author_url": "",
      "post_date": "11/26/2023 00:51:29",
      "content": "<p>Im sorry, pretty sure i've misunderstood something, we don't have bpps for the hidden lb data right? How are you using bpps as additional features then?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2538530,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "11/26/2023 08:16:08",
          "content": "<p>From the Data page -</p>\n<p>\"Ribonanza_bpp_files - TXT files listing position pairs predicted to have non-zero Watson-Crick base pair probabilities by the LinearPartition-EternaFold package. Files are given for train and test sequences, indexed by sequence_id\"</p>",
          "votes": null,
          "replies": [
            {
              "id": 2538547,
              "author_name": "jainam213",
              "author_url": "",
              "post_date": "11/26/2023 08:52:56",
              "content": "<p>But</p>\n<blockquote>\n  <p>Within this set, the majority (1,008,000 RNAs) will be experimentally synthesized and profiled after the Kaggle competition begins</p>\n</blockquote>\n<p>So will the bpps be updated on the hidden data too? </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2538568,
                  "author_name": "shlomoron",
                  "author_url": "",
                  "post_date": "11/26/2023 09:17:33",
                  "content": "<p>This is not a code competition. The test data is not hidden and we know exactly what are the sequences that we are trying to predict, thus we can generate and use any feature we want.</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2538570,
                  "author_name": "something4kag",
                  "author_url": "",
                  "post_date": "11/26/2023 09:20:15",
                  "content": "<p>see <a href=\"https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/437481#2441007\" target=\"_blank\">posts around here</a><br>\nfrom the host:<br>\n\"After we get these new data, we'll rescore the existing submission files. Note that although data will be new, they will be for private leaderboard* test sequences for which the submissions already include predictions,\"</p>\n<p>\"There are numerous sequences in test_sequences.csv for which we are asking for submissions. But we don't have the data to evaluate your submissions for those sequences at the start of the competition. \"</p>\n<p>so the sequence ids are there for test in submissions, in bpps, etc. but the data to do the evaluation, the reactivity will be ongoing in the competition. </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2538671,
                      "author_name": "shlomoron",
                      "author_url": "",
                      "post_date": "11/26/2023 11:16:10",
                      "content": "<p>The host refers to the labels. The sequences are the same. Note that you can submit a csv file without any code. As I said, this is not a code competition.</p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 2538697,
                      "author_name": "jainam213",
                      "author_url": "",
                      "post_date": "11/26/2023 11:52:02",
                      "content": "<p>thank you!</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                },
                {
                  "id": 2538725,
                  "author_name": "fnands",
                  "author_url": "",
                  "post_date": "11/26/2023 12:38:15",
                  "content": "<p>The BPPs are not extracted from experiments, they are predicted with EternaFold. </p>\n<p>From the data page: </p>\n<blockquote>\n  <p>on-zero Watson-Crick base pair probabilities by the LinearPartition-EternaFold package</p>\n</blockquote>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2529109": "Pretty much same as the title. \nIn my experiments, only thing that boosts is BPP. \nI've not used arnie to generate other bpp's till now.",
    "2532272": "How are you embeddings the [seq_length, seq_length] tensor into your sequence model? Do you have any tips to change this 2D tensor into 1D for input to a sequence model?",
    "2532284": "Well, when you think about it, the BPP values are really just attentions scores, no? ;-)",
    "2532802": "In addition to what @fnands said, you can take sum/max/second max, etc., across the columns. This is the simplest way to get a [seq_length, #features] from [seq_length, seq_length] and will give you an immediate boost, even if not as good as more sophisticated ways.",
    "2533382": "I have the same experience.  This has been troubling me a lot.  I've also been optimizing secondary structures considering the DMS and 2A3 scores, which basically leaks information into cross validation.  Still it's not as good as using BPP in cross-validation.\n\nI think what I observed translates to this:  RNA folding has multiple conformations and considering only a single conformation is not enough.\n\nThis competition is rather challenging.",
    "2535693": "Thanks for the advice!",
    "2538273": "Im sorry, pretty sure i've misunderstood something, we don't have bpps for the hidden lb data right? How are you using bpps as additional features then?",
    "2538530": "From the Data page -\n\n\"Ribonanza_bpp_files - TXT files listing position pairs predicted to have non-zero Watson-Crick base pair probabilities by the LinearPartition-EternaFold package. Files are given for train and test sequences, indexed by sequence_id\"",
    "2538547": "But\n> Within this set, the majority (1,008,000 RNAs) will be experimentally synthesized and profiled after the Kaggle competition begins\n\nSo will the bpps be updated on the hidden data too?",
    "2538568": "This is not a code competition. The test data is not hidden and we know exactly what are the sequences that we are trying to predict, thus we can generate and use any feature we want.",
    "2538570": "see [posts around here](https://www.kaggle.com/competitions/stanford-ribonanza-rna-folding/discussion/437481#2441007)\nfrom the host:\n\"After we get these new data, we'll rescore the existing submission files. Note that although data will be new, they will be for private leaderboard* test sequences for which the submissions already include predictions,\"\n\n\"There are numerous sequences in test_sequences.csv for which we are asking for submissions. But we don't have the data to evaluate your submissions for those sequences at the start of the competition. \"\n\nso the sequence ids are there for test in submissions, in bpps, etc. but the data to do the evaluation, the reactivity will be ongoing in the competition.",
    "2538671": "The host refers to the labels. The sequences are the same. Note that you can submit a csv file without any code. As I said, this is not a code competition.",
    "2538697": "thank you!",
    "2538725": "The BPPs are not extracted from experiments, they are predicted with EternaFold. \n\nFrom the data page: \n>on-zero Watson-Crick base pair probabilities by the LinearPartition-EternaFold package",
    "2539218": "A followup question. Would you recommend embedding the index of the other pair or just the float probability?",
    "2539222": "Try and find out! :)"
  },
  "source": "meta"
}