{
  "id": 437850,
  "title": "Beginner Question – What is the target?",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/437850",
  "author_name": "Danil Nizamov",
  "post_date": "2023-09-08T11:42:07.744000",
  "votes": 2,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hi everyone! I am not a biologist and I am not sure if I could show some result in this competition, but I would love to try. I can't understand how do I obtain the prediction target from the dataset. Do I need some specific biology knowledge for doing this?</p>",
  "messages": [
    {
      "id": 2429118,
      "postDate": "2023-09-08T11:42:07.743Z",
      "content": "<p>Hi everyone! I am not a biologist and I am not sure if I could show some result in this competition, but I would love to try. I can't understand how do I obtain the prediction target from the dataset. Do I need some specific biology knowledge for doing this?</p>",
      "rawMarkdown": "Hi everyone! I am not a biologist and I am not sure if I could show some result in this competition, but I would love to try. I can't understand how do I obtain the prediction target from the dataset. Do I need some specific biology knowledge for doing this?",
      "votes": 2
    },
    {
      "id": 2434422,
      "postDate": "2023-09-12T09:42:58.103Z",
      "content": "<p>Am i right?</p>\n<p>[GGGAA….]<br>\n[reactivity_0001, reactivity_0002,    reactivity_0003, reactivity_0004, reactivity_0005….]</p>\n<p>and thats what we should to predict?</p>",
      "rawMarkdown": "Am i right?\n\n[GGGAA....]\n[reactivity_0001, reactivity_0002,\treactivity_0003, reactivity_0004, reactivity_0005....]\n\nand thats what we should to predict?",
      "votes": 1,
      "replies": [
        {
          "id": 2434698,
          "postDate": "2023-09-12T13:14:15.140Z",
          "content": "<p>Correct. You are predicting the reactivity at each position.</p>",
          "rawMarkdown": "Correct. You are predicting the reactivity at each position.",
          "votes": 1
        },
        {
          "id": 2434873,
          "postDate": "2023-09-12T15:15:57.950Z",
          "content": "<p>And in the submission file you have to have a running id from 0 to the end, increment for each reactivity. Inclusive id_min and id_max (index the RNA with index - id_min). So, keep the same order as the test_sequences, else you'll garble everything.</p>",
          "rawMarkdown": "And in the submission file you have to have a running id from 0 to the end, increment for each reactivity. Inclusive id_min and id_max (index the RNA with index - id_min). So, keep the same order as the test_sequences, else you'll garble everything.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2429420,
      "postDate": "2023-09-08T15:34:52.213Z",
      "content": "<p>While you don't need to be a full-fledge molecular biologist to compete, it does help to understand some of the basic concepts of bioinformatics and where it differs from general Data Science.</p>\n<p>We're trying to predict the \"structural properties\" of RNA, essentially.  As such, it does help to have a modest understanding of thermodynamic relations in regards to RNA.</p>",
      "rawMarkdown": "While you don't need to be a full-fledge molecular biologist to compete, it does help to understand some of the basic concepts of bioinformatics and where it differs from general Data Science.\n\nWe're trying to predict the \"structural properties\" of RNA, essentially.  As such, it does help to have a modest understanding of thermodynamic relations in regards to RNA.\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 2429746,
          "postDate": "2023-09-08T19:12:14.227Z",
          "content": "<p>Thanks for the context! My education is not related to biology in any way, so I would probably skip this competition by now. Anyway it is not for an ML beginner for sure 😅</p>",
          "rawMarkdown": "Thanks for the context! My education is not related to biology in any way, so I would probably skip this competition by now. Anyway it is not for an ML beginner for sure 😅",
          "replies": [
            {
              "id": 2429749,
              "postDate": "2023-09-08T19:17:55.620Z",
              "content": "<p>If you're a beginner, there are several starter competitions available.  I'd also highly recommend the Playground Series competitions.  They are more challenging, but with the aim of learning new techniques to help you become a better data scientist!</p>",
              "rawMarkdown": "If you're a beginner, there are several starter competitions available.  I'd also highly recommend the Playground Series competitions.  They are more challenging, but with the aim of learning new techniques to help you become a better data scientist!",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2429647,
      "postDate": "2023-09-08T17:46:16.227Z",
      "content": "<p>According to the description:</p>\n<blockquote>\n  <p>From the Data page: At each position of each RNA sequence, there will be two ground truth values, corresponding to reactivity determined from two kinds of chemical mapping experiments, DMS_MaP and 2A3_MaP.</p>\n</blockquote>\n<p>The reactivity values can be seen in the training_data.csv</p>\n<p><strong>however</strong> what is not clear to me, is why they are clamping the predicted outputs to [0,1) when the reported reactivities in the csv can be <strong>negative</strong></p>",
      "rawMarkdown": "According to the description:\n> From the Data page: At each position of each RNA sequence, there will be two ground truth values, corresponding to reactivity determined from two kinds of chemical mapping experiments, DMS_MaP and 2A3_MaP.\n\nThe reactivity values can be seen in the training_data.csv\n\n**however** what is not clear to me, is why they are clamping the predicted outputs to [0,1) when the reported reactivities in the csv can be **negative**",
      "replies": [
        {
          "id": 2429772,
          "postDate": "2023-09-08T19:40:56.423Z",
          "content": "<p>Any negative reported reactivities are experimental artifacts, not real values. Generally speaking, the experimental procedure used to generate these data allow us to chemically modify RNA nucleotides with dimethyl sulfate (DMS) signal molecules (see <a href=\"https://www.nature.com/articles/nmeth.4057\" target=\"_blank\">this paper</a> if you're interested in more detail). Unpaired bases can accept the DMS modification, and paired bases cannot. The modifications can be quantified with genetic sequencing techniques, and we can normalize the data to a range from 1 (indicating a base is highly reactive) to 0 (indicating a base is not reactive). Values within this range can indicate that sometimes this base is paired, and sometimes it's not. RNA can form a wide range of structures for a given sequence, with different likelihoods of forming each structure in certain conditions.</p>",
          "rawMarkdown": "Any negative reported reactivities are experimental artifacts, not real values. Generally speaking, the experimental procedure used to generate these data allow us to chemically modify RNA nucleotides with dimethyl sulfate (DMS) signal molecules (see [this paper](https://www.nature.com/articles/nmeth.4057) if you're interested in more detail). Unpaired bases can accept the DMS modification, and paired bases cannot. The modifications can be quantified with genetic sequencing techniques, and we can normalize the data to a range from 1 (indicating a base is highly reactive) to 0 (indicating a base is not reactive). Values within this range can indicate that sometimes this base is paired, and sometimes it's not. RNA can form a wide range of structures for a given sequence, with different likelihoods of forming each structure in certain conditions.",
          "votes": 6,
          "replies": [
            {
              "id": 2429874,
              "postDate": "2023-09-08T21:32:29.603Z",
              "content": "<p>Thanks for the in-depth answer! Makes sense… 👍</p>",
              "rawMarkdown": "Thanks for the in-depth answer! Makes sense... 👍"
            },
            {
              "id": 2429894,
              "postDate": "2023-09-08T21:59:45.927Z",
              "content": "<p>Could we have more insight into the data preprocessing and normalization?</p>\n<p>My main question is: Are negative number or numbers larger than one meaningful? Or can they be clipped to be between 0 and 1 without any loss of information?</p>",
              "rawMarkdown": "Could we have more insight into the data preprocessing and normalization?\n\nMy main question is: Are negative number or numbers larger than one meaningful? Or can they be clipped to be between 0 and 1 without any loss of information?",
              "votes": 5
            },
            {
              "id": 2430288,
              "postDate": "2023-09-09T07:24:39.270Z",
              "content": "<p>So during evaluation, all values should be capped between [0,1], right?</p>",
              "rawMarkdown": "So during evaluation, all values should be capped between [0,1], right?"
            },
            {
              "id": 2469616,
              "postDate": "2023-10-06T13:30:46.707Z",
              "content": "<p>I have the same question.</p>",
              "rawMarkdown": "I have the same question.",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 2430604,
          "postDate": "2023-09-09T12:05:33.937Z",
          "content": "<p>\"<em>there will be two ground truth values, corresponding to reactivity determined from two kinds of chemical mapping experiments, DMS_MaP and 2A3_MaP</em>\"</p>\n<p>They cannot be seen in the downloadable file. Nor \"future\".<br>\n<strong>Edit</strong>: id, id_min and id_max are missing, too</p>",
          "rawMarkdown": "\"*there will be two ground truth values, corresponding to reactivity determined from two kinds of chemical mapping experiments, DMS_MaP and 2A3_MaP*\"\n\nThey cannot be seen in the downloadable file. Nor \"future\".\n**Edit**: id, id_min and id_max are missing, too",
          "replies": [
            {
              "id": 2431242,
              "postDate": "2023-09-09T23:52:00.383Z",
              "content": "<p>The measured values are found in the reactivity_0001 through reactivity_0170 columns in the training data (as well as reactivity error columns which provide experimental error). Each row of training data will either be providing results for either DMS or 2A3, as specified by the experiment_type column. \"future\" is not relevant to training - it indicates whether the data will only be available in the future (all the training data is already available!). id/id_min/id_max are basically just bookkeeping that allow you to connect rows in your submission to rows in the test data (eg, the second sequence is id_min 177 through id_max 353, which means id 177 in your submission corresponds to the predicted value at the first position in the second sequence, and 353 corresponds to the predicted value at the last position of the second sequence).</p>",
              "rawMarkdown": "The measured values are found in the reactivity_0001 through reactivity_0170 columns in the training data (as well as reactivity error columns which provide experimental error). Each row of training data will either be providing results for either DMS or 2A3, as specified by the experiment_type column. \"future\" is not relevant to training - it indicates whether the data will only be available in the future (all the training data is already available!). id/id_min/id_max are basically just bookkeeping that allow you to connect rows in your submission to rows in the test data (eg, the second sequence is id_min 177 through id_max 353, which means id 177 in your submission corresponds to the predicted value at the first position in the second sequence, and 353 corresponds to the predicted value at the last position of the second sequence).",
              "votes": 5
            },
            {
              "id": 2434992,
              "postDate": "2023-09-12T16:24:35.847Z",
              "content": "<p>CORRECTION: reactivity_0001 through reactivity_0206 - I misread the available columns</p>",
              "rawMarkdown": "CORRECTION: reactivity_0001 through reactivity_0206 - I misread the available columns",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2434422,
      "author_name": "Andrey Nemerov",
      "author_url": "",
      "post_date": "2023-09-12T09:42:58.103000",
      "content": "<p>Am i right?</p>\n<p>[GGGAA….]<br>\n[reactivity_0001, reactivity_0002,    reactivity_0003, reactivity_0004, reactivity_0005….]</p>\n<p>and thats what we should to predict?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2434698,
          "author_name": "DigitalEmbrace",
          "author_url": "",
          "post_date": "2023-09-12T13:14:15.140000",
          "content": "<p>Correct. You are predicting the reactivity at each position.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2434873,
          "author_name": "Tord Malmgren",
          "author_url": "",
          "post_date": "2023-09-12T15:15:57.950000",
          "content": "<p>And in the submission file you have to have a running id from 0 to the end, increment for each reactivity. Inclusive id_min and id_max (index the RNA with index - id_min). So, keep the same order as the test_sequences, else you'll garble everything.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2429420,
      "author_name": "Kent RB",
      "author_url": "",
      "post_date": "2023-09-08T15:34:52.213000",
      "content": "<p>While you don't need to be a full-fledge molecular biologist to compete, it does help to understand some of the basic concepts of bioinformatics and where it differs from general Data Science.</p>\n<p>We're trying to predict the \"structural properties\" of RNA, essentially.  As such, it does help to have a modest understanding of thermodynamic relations in regards to RNA.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2429746,
          "author_name": "Danil Nizamov",
          "author_url": "",
          "post_date": "2023-09-08T19:12:14.227000",
          "content": "<p>Thanks for the context! My education is not related to biology in any way, so I would probably skip this competition by now. Anyway it is not for an ML beginner for sure 😅</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2429749,
              "author_name": "Kent RB",
              "author_url": "",
              "post_date": "2023-09-08T19:17:55.620000",
              "content": "<p>If you're a beginner, there are several starter competitions available.  I'd also highly recommend the Playground Series competitions.  They are more challenging, but with the aim of learning new techniques to help you become a better data scientist!</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2429647,
      "author_name": "E/S Pronk",
      "author_url": "",
      "post_date": "2023-09-08T17:46:16.227000",
      "content": "<p>According to the description:</p>\n<blockquote>\n  <p>From the Data page: At each position of each RNA sequence, there will be two ground truth values, corresponding to reactivity determined from two kinds of chemical mapping experiments, DMS_MaP and 2A3_MaP.</p>\n</blockquote>\n<p>The reactivity values can be seen in the training_data.csv</p>\n<p><strong>however</strong> what is not clear to me, is why they are clamping the predicted outputs to [0,1) when the reported reactivities in the csv can be <strong>negative</strong></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2429772,
          "author_name": "Thomas",
          "author_url": "",
          "post_date": "2023-09-08T19:40:56.423000",
          "content": "<p>Any negative reported reactivities are experimental artifacts, not real values. Generally speaking, the experimental procedure used to generate these data allow us to chemically modify RNA nucleotides with dimethyl sulfate (DMS) signal molecules (see <a href=\"https://www.nature.com/articles/nmeth.4057\" target=\"_blank\">this paper</a> if you're interested in more detail). Unpaired bases can accept the DMS modification, and paired bases cannot. The modifications can be quantified with genetic sequencing techniques, and we can normalize the data to a range from 1 (indicating a base is highly reactive) to 0 (indicating a base is not reactive). Values within this range can indicate that sometimes this base is paired, and sometimes it's not. RNA can form a wide range of structures for a given sequence, with different likelihoods of forming each structure in certain conditions.</p>",
          "votes": 6,
          "replies": [
            {
              "id": 2429874,
              "author_name": "E/S Pronk",
              "author_url": "",
              "post_date": "2023-09-08T21:32:29.603000",
              "content": "<p>Thanks for the in-depth answer! Makes sense… 👍</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2429894,
              "author_name": "Matt",
              "author_url": "",
              "post_date": "2023-09-08T21:59:45.927000",
              "content": "<p>Could we have more insight into the data preprocessing and normalization?</p>\n<p>My main question is: Are negative number or numbers larger than one meaningful? Or can they be clipped to be between 0 and 1 without any loss of information?</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2430288,
              "author_name": "Ai Vu Hong",
              "author_url": "",
              "post_date": "2023-09-09T07:24:39.270000",
              "content": "<p>So during evaluation, all values should be capped between [0,1], right?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2469616,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-10-06T13:30:46.707000",
              "content": "<p>I have the same question.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2430604,
          "author_name": "Tord Malmgren",
          "author_url": "",
          "post_date": "2023-09-09T12:05:33.937000",
          "content": "<p>\"<em>there will be two ground truth values, corresponding to reactivity determined from two kinds of chemical mapping experiments, DMS_MaP and 2A3_MaP</em>\"</p>\n<p>They cannot be seen in the downloadable file. Nor \"future\".<br>\n<strong>Edit</strong>: id, id_min and id_max are missing, too</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2431242,
              "author_name": "Jonathan Romano",
              "author_url": "",
              "post_date": "2023-09-09T23:52:00.383000",
              "content": "<p>The measured values are found in the reactivity_0001 through reactivity_0170 columns in the training data (as well as reactivity error columns which provide experimental error). Each row of training data will either be providing results for either DMS or 2A3, as specified by the experiment_type column. \"future\" is not relevant to training - it indicates whether the data will only be available in the future (all the training data is already available!). id/id_min/id_max are basically just bookkeeping that allow you to connect rows in your submission to rows in the test data (eg, the second sequence is id_min 177 through id_max 353, which means id 177 in your submission corresponds to the predicted value at the first position in the second sequence, and 353 corresponds to the predicted value at the last position of the second sequence).</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2434992,
              "author_name": "Jonathan Romano",
              "author_url": "",
              "post_date": "2023-09-12T16:24:35.847000",
              "content": "<p>CORRECTION: reactivity_0001 through reactivity_0206 - I misread the available columns</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2429118": "Hi everyone! I am not a biologist and I am not sure if I could show some result in this competition, but I would love to try. I can't understand how do I obtain the prediction target from the dataset. Do I need some specific biology knowledge for doing this?",
    "2434422": "Am i right?\n\n[GGGAA....]\n[reactivity_0001, reactivity_0002,\treactivity_0003, reactivity_0004, reactivity_0005....]\n\nand thats what we should to predict?",
    "2429420": "While you don't need to be a full-fledge molecular biologist to compete, it does help to understand some of the basic concepts of bioinformatics and where it differs from general Data Science.\n\nWe're trying to predict the \"structural properties\" of RNA, essentially.  As such, it does help to have a modest understanding of thermodynamic relations in regards to RNA.\n\n",
    "2429647": "According to the description:\n> From the Data page: At each position of each RNA sequence, there will be two ground truth values, corresponding to reactivity determined from two kinds of chemical mapping experiments, DMS_MaP and 2A3_MaP.\n\nThe reactivity values can be seen in the training_data.csv\n\n**however** what is not clear to me, is why they are clamping the predicted outputs to [0,1) when the reported reactivities in the csv can be **negative**"
  }
}