{
  "id": 459077,
  "title": "Reactivity_xxx and Reactivity_dms_Map relation",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/459077",
  "author_name": "",
  "post_date": "2023-12-03T12:34:13.419115600Z",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello!<br>\nWhat is the relation between reactivities_001, 002 … and dms_map?<br>\nIf you look at the notebook i seem to be predicting a number of reactivates for each prediction<br>\nare these 001,002 etc? how do i calculate the dms_map from these ? average/mean??</p>\n<p>Also if we want to use pseudo lables we would need 001,002 values correct? but the predictions are not always 206 reactivates? how do we use the pseudo lables in that case</p>\n<p>I think im understanding something fundamentally incorrect would be glad if someone could help out </p>",
  "messages": [
    {
      "id": "2547328",
      "postDate": "12/03/2023 12:34:13",
      "content": "<p>Hello!<br>\nWhat is the relation between reactivities_001, 002 … and dms_map?<br>\nIf you look at the notebook i seem to be predicting a number of reactivates for each prediction<br>\nare these 001,002 etc? how do i calculate the dms_map from these ? average/mean??</p>\n<p>Also if we want to use pseudo lables we would need 001,002 values correct? but the predictions are not always 206 reactivates? how do we use the pseudo lables in that case</p>\n<p>I think im understanding something fundamentally incorrect would be glad if someone could help out </p>",
      "rawMarkdown": "Hello!\nWhat is the relation between reactivities_001, 002 ... and dms_map?\nIf you look at the notebook i seem to be predicting a number of reactivates for each prediction\nare these 001,002 etc? how do i calculate the dms_map from these ? average/mean??\n\nAlso if we want to use pseudo lables we would need 001,002 values correct? but the predictions are not always 206 reactivates? how do we use the pseudo lables in that case\n\nI think im understanding something fundamentally incorrect would be glad if someone could help out",
      "votes": null
    },
    {
      "id": "2547374",
      "postDate": "12/03/2023 12:58:44",
      "content": "<p>We are predicting the values from the columns that says \"reactivity_0…\". Those valors corresponds to the map for the corresponding \"Experiment type\" column at each nucleotid position.</p>",
      "rawMarkdown": "We are predicting the values from the columns that says \"reactivity_0...\". Those valors corresponds to the map for the corresponding \"Experiment type\" column at each nucleotid position.",
      "votes": null
    },
    {
      "id": "2547379",
      "postDate": "12/03/2023 13:01:04",
      "content": "<p>\"Also if we want to use pseudo lables we would need 001,002 values correct? but the predictions are not always 206 reactivates? how do we use the pseudo lables in that case\"<br>\nBecause the sequences are not always 206 nucleotides long. The firsts sequence length reactivities columns are the needed ones.</p>",
      "rawMarkdown": "\"Also if we want to use pseudo lables we would need 001,002 values correct? but the predictions are not always 206 reactivates? how do we use the pseudo lables in that case\"\nBecause the sequences are not always 206 nucleotides long. The firsts sequence length reactivities columns are the needed ones.",
      "votes": null
    },
    {
      "id": "2547420",
      "postDate": "12/03/2023 13:40:44",
      "content": "<p>Thanks a ton!<br>\nso for the final submissions<br>\n If i understand correctly,  I should just take the first values here correct ? ie. [ 4.8976e-02, -1.4476e-02] and not averaging all the reactivities.<br>\nEssentially after predicting 001,002 etc how would you get the final submission for dmsmapand da3 map  for that prediction</p>\n<blockquote>\n  <p>TokenClassifierOutput(loss=None, logits=tensor([[[ 4.8976e-02, -1.4476e-02],<br>\n           [ 1.6617e-04, -5.4074e-05],<br>\n           [ 1.5560e-04, -1.0296e-04],<br>\n           [ 1.7451e-04, -9.6448e-06],<br>\n           [ 2.7752e-04, -1.0885e-04],<br>\n           [ 2.1534e-04, -1.2070e-04],<br>\n           [ 2.0868e-04, -1.1965e-04],<br>\n           [ 1.7301e-04, -8.5471e-05],<br>\n  ……</p>\n</blockquote>",
      "rawMarkdown": "Thanks a ton!\nso for the final submissions\n If i understand correctly,  I should just take the first values here correct ? ie. [ 4.8976e-02, -1.4476e-02] and not averaging all the reactivities.\nEssentially after predicting 001,002 etc how would you get the final submission for dmsmapand da3 map  for that prediction\n>TokenClassifierOutput(loss=None, logits=tensor([[[ 4.8976e-02, -1.4476e-02],\n         [ 1.6617e-04, -5.4074e-05],\n         [ 1.5560e-04, -1.0296e-04],\n         [ 1.7451e-04, -9.6448e-06],\n         [ 2.7752e-04, -1.0885e-04],\n         [ 2.1534e-04, -1.2070e-04],\n         [ 2.0868e-04, -1.1965e-04],\n         [ 1.7301e-04, -8.5471e-05],\n......",
      "votes": null
    },
    {
      "id": "2547435",
      "postDate": "12/03/2023 13:54:19",
      "content": "<p>I think you've got it but I'm not sure about that last part. Essentially your labels are listed by sequences at train.csv but you will need to submit the test outputs concatenated. The columns 'id_min' and 'id_max' will show you how much submission need to grow with each sequence.</p>",
      "rawMarkdown": "I think you've got it but I'm not sure about that last part. Essentially your labels are listed by sequences at train.csv but you will need to submit the test outputs concatenated. The columns 'id_min' and 'id_max' will show you how much submission need to grow with each sequence.",
      "votes": null
    },
    {
      "id": "2547450",
      "postDate": "12/03/2023 14:05:32",
      "content": "<p>Thanks a lot! I've finally understood it :)</p>\n<blockquote>\n  <p>The columns 'id_min' and 'id_max' will show you how much submission need to grow with each sequence.</p>\n</blockquote>\n<p>Not totally sure what you mean by grow, but if in above example if [[ 4.8976e-02, -1.4476e-02],]] should be what would go into submission.csv <br>\nand the remaining values [ 1.6617e-04, -5.4074e-05], [ 1.5560e-04, -1.0296e-04], …. are the predicted values for 001,002 respectively then we're on the same page </p>",
      "rawMarkdown": "Thanks a lot! I've finally understood it :)\n>The columns 'id_min' and 'id_max' will show you how much submission need to grow with each sequence.\n\n\nNot totally sure what you mean by grow, but if in above example if [[ 4.8976e-02, -1.4476e-02],]] should be what would go into submission.csv \nand the remaining values [ 1.6617e-04, -5.4074e-05], [ 1.5560e-04, -1.0296e-04], .... are the predicted values for 001,002 respectively then we're on the same page",
      "votes": null
    },
    {
      "id": "2547461",
      "postDate": "12/03/2023 14:21:03",
      "content": "<p>test_sequences.csv starts with the sequence with sequence_id = 'eee73c1836bc'<br>\nThat sequence is 177 nucleotides long. So it will produce an output of 177 values (for each experiment) that will be need to be identified at submission file by the indexes from id_min = 0 to id_max = 176.<br>\nNext sequence is the one with sequence_id = 'd2a929af7a97' and it's 177 nucleotides long too. That one will also produce 177 pairs of values that will need to be listed on submission on the next indexes from id_min = 177 to id_max = 353… and so on<br>\nThe total test sequences will produce 269796671 pairs of values in some hours on most of the cases. The columns id_min an id_max helps to keep evrything under control.</p>",
      "rawMarkdown": "test_sequences.csv starts with the sequence with sequence_id = 'eee73c1836bc'\nThat sequence is 177 nucleotides long. So it will produce an output of 177 values (for each experiment) that will be need to be identified at submission file by the indexes from id_min = 0 to id_max = 176.\nNext sequence is the one with sequence_id = 'd2a929af7a97' and it's 177 nucleotides long too. That one will also produce 177 pairs of values that will need to be listed on submission on the next indexes from id_min = 177 to id_max = 353... and so on\nThe total test sequences will produce 269796671 pairs of values in some hours on most of the cases. The columns id_min an id_max helps to keep evrything under control.",
      "votes": null
    },
    {
      "id": "2547473",
      "postDate": "12/03/2023 14:35:49",
      "content": "<p>aaah thanks! </p>",
      "rawMarkdown": "aaah thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2547374,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "12/03/2023 12:58:44",
      "content": "<p>We are predicting the values from the columns that says \"reactivity_0…\". Those valors corresponds to the map for the corresponding \"Experiment type\" column at each nucleotid position.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2547379,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "12/03/2023 13:01:04",
      "content": "<p>\"Also if we want to use pseudo lables we would need 001,002 values correct? but the predictions are not always 206 reactivates? how do we use the pseudo lables in that case\"<br>\nBecause the sequences are not always 206 nucleotides long. The firsts sequence length reactivities columns are the needed ones.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2547420,
          "author_name": "dhruvdhilla",
          "author_url": "",
          "post_date": "12/03/2023 13:40:44",
          "content": "<p>Thanks a ton!<br>\nso for the final submissions<br>\n If i understand correctly,  I should just take the first values here correct ? ie. [ 4.8976e-02, -1.4476e-02] and not averaging all the reactivities.<br>\nEssentially after predicting 001,002 etc how would you get the final submission for dmsmapand da3 map  for that prediction</p>\n<blockquote>\n  <p>TokenClassifierOutput(loss=None, logits=tensor([[[ 4.8976e-02, -1.4476e-02],<br>\n           [ 1.6617e-04, -5.4074e-05],<br>\n           [ 1.5560e-04, -1.0296e-04],<br>\n           [ 1.7451e-04, -9.6448e-06],<br>\n           [ 2.7752e-04, -1.0885e-04],<br>\n           [ 2.1534e-04, -1.2070e-04],<br>\n           [ 2.0868e-04, -1.1965e-04],<br>\n           [ 1.7301e-04, -8.5471e-05],<br>\n  ……</p>\n</blockquote>",
          "votes": null,
          "replies": [
            {
              "id": 2547435,
              "author_name": "sacuscreed",
              "author_url": "",
              "post_date": "12/03/2023 13:54:19",
              "content": "<p>I think you've got it but I'm not sure about that last part. Essentially your labels are listed by sequences at train.csv but you will need to submit the test outputs concatenated. The columns 'id_min' and 'id_max' will show you how much submission need to grow with each sequence.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2547450,
                  "author_name": "dhruvdhilla",
                  "author_url": "",
                  "post_date": "12/03/2023 14:05:32",
                  "content": "<p>Thanks a lot! I've finally understood it :)</p>\n<blockquote>\n  <p>The columns 'id_min' and 'id_max' will show you how much submission need to grow with each sequence.</p>\n</blockquote>\n<p>Not totally sure what you mean by grow, but if in above example if [[ 4.8976e-02, -1.4476e-02],]] should be what would go into submission.csv <br>\nand the remaining values [ 1.6617e-04, -5.4074e-05], [ 1.5560e-04, -1.0296e-04], …. are the predicted values for 001,002 respectively then we're on the same page </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2547461,
                      "author_name": "sacuscreed",
                      "author_url": "",
                      "post_date": "12/03/2023 14:21:03",
                      "content": "<p>test_sequences.csv starts with the sequence with sequence_id = 'eee73c1836bc'<br>\nThat sequence is 177 nucleotides long. So it will produce an output of 177 values (for each experiment) that will be need to be identified at submission file by the indexes from id_min = 0 to id_max = 176.<br>\nNext sequence is the one with sequence_id = 'd2a929af7a97' and it's 177 nucleotides long too. That one will also produce 177 pairs of values that will need to be listed on submission on the next indexes from id_min = 177 to id_max = 353… and so on<br>\nThe total test sequences will produce 269796671 pairs of values in some hours on most of the cases. The columns id_min an id_max helps to keep evrything under control.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2547473,
                          "author_name": "dhruvdhilla",
                          "author_url": "",
                          "post_date": "12/03/2023 14:35:49",
                          "content": "<p>aaah thanks! </p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2547328": "Hello!\nWhat is the relation between reactivities_001, 002 ... and dms_map?\nIf you look at the notebook i seem to be predicting a number of reactivates for each prediction\nare these 001,002 etc? how do i calculate the dms_map from these ? average/mean??\n\nAlso if we want to use pseudo lables we would need 001,002 values correct? but the predictions are not always 206 reactivates? how do we use the pseudo lables in that case\n\nI think im understanding something fundamentally incorrect would be glad if someone could help out",
    "2547374": "We are predicting the values from the columns that says \"reactivity_0...\". Those valors corresponds to the map for the corresponding \"Experiment type\" column at each nucleotid position.",
    "2547379": "\"Also if we want to use pseudo lables we would need 001,002 values correct? but the predictions are not always 206 reactivates? how do we use the pseudo lables in that case\"\nBecause the sequences are not always 206 nucleotides long. The firsts sequence length reactivities columns are the needed ones.",
    "2547420": "Thanks a ton!\nso for the final submissions\n If i understand correctly,  I should just take the first values here correct ? ie. [ 4.8976e-02, -1.4476e-02] and not averaging all the reactivities.\nEssentially after predicting 001,002 etc how would you get the final submission for dmsmapand da3 map  for that prediction\n>TokenClassifierOutput(loss=None, logits=tensor([[[ 4.8976e-02, -1.4476e-02],\n         [ 1.6617e-04, -5.4074e-05],\n         [ 1.5560e-04, -1.0296e-04],\n         [ 1.7451e-04, -9.6448e-06],\n         [ 2.7752e-04, -1.0885e-04],\n         [ 2.1534e-04, -1.2070e-04],\n         [ 2.0868e-04, -1.1965e-04],\n         [ 1.7301e-04, -8.5471e-05],\n......",
    "2547435": "I think you've got it but I'm not sure about that last part. Essentially your labels are listed by sequences at train.csv but you will need to submit the test outputs concatenated. The columns 'id_min' and 'id_max' will show you how much submission need to grow with each sequence.",
    "2547450": "Thanks a lot! I've finally understood it :)\n>The columns 'id_min' and 'id_max' will show you how much submission need to grow with each sequence.\n\n\nNot totally sure what you mean by grow, but if in above example if [[ 4.8976e-02, -1.4476e-02],]] should be what would go into submission.csv \nand the remaining values [ 1.6617e-04, -5.4074e-05], [ 1.5560e-04, -1.0296e-04], .... are the predicted values for 001,002 respectively then we're on the same page",
    "2547461": "test_sequences.csv starts with the sequence with sequence_id = 'eee73c1836bc'\nThat sequence is 177 nucleotides long. So it will produce an output of 177 values (for each experiment) that will be need to be identified at submission file by the indexes from id_min = 0 to id_max = 176.\nNext sequence is the one with sequence_id = 'd2a929af7a97' and it's 177 nucleotides long too. That one will also produce 177 pairs of values that will need to be listed on submission on the next indexes from id_min = 177 to id_max = 353... and so on\nThe total test sequences will produce 269796671 pairs of values in some hours on most of the cases. The columns id_min an id_max helps to keep evrything under control.",
    "2547473": "aaah thanks!"
  },
  "source": "meta"
}