{
  "id": 454773,
  "title": "What are the input, label ?",
  "url": "/competitions/stanford-ribonanza-rna-folding/discussion/454773",
  "author_name": "",
  "post_date": "2023-11-11T17:35:01.363911600Z",
  "votes": 2,
  "comment_count": 12,
  "views": 0,
  "content": "<p>What are the target to compare with the output. We test if it's a one of the 2 types indicated by the experiment field ?</p>",
  "messages": [
    {
      "id": "2521432",
      "postDate": "11/11/2023 17:35:01",
      "content": "<p>What are the target to compare with the output. We test if it's a one of the 2 types indicated by the experiment field ?</p>",
      "rawMarkdown": "What are the target to compare with the output. We test if it's a one of the 2 types indicated by the experiment field ?",
      "votes": null
    },
    {
      "id": "2522067",
      "postDate": "11/12/2023 11:01:33",
      "content": "<p>No, the target are the values for both experiments 2A3 and DMS. The values at columns like 'reactivity_0…' (not the ones with reactivity_error_0…). So for each RNA sequence you must predict the set of reactivities corresponding to each nucleotide and experiment. It's a regression model.</p>",
      "rawMarkdown": "No, the target are the values for both experiments 2A3 and DMS. The values at columns like 'reactivity_0...' (not the ones with reactivity_error_0...). So for each RNA sequence you must predict the set of reactivities corresponding to each nucleotide and experiment. It's a regression model.",
      "votes": null
    },
    {
      "id": "2522070",
      "postDate": "11/12/2023 11:04:33",
      "content": "<p>The test submission is a bit more complex since you need to submit the concatenated predictions of each test sequence.</p>",
      "rawMarkdown": "The test submission is a bit more complex since you need to submit the concatenated predictions of each test sequence.",
      "votes": null
    },
    {
      "id": "2522073",
      "postDate": "11/12/2023 11:09:01",
      "content": "<p>You have to predict 2 output for  2A3 and DMS. But the value is only one for each nucleotide. How to compare the output with target for calculating the cost in a Neural Network.</p>",
      "rawMarkdown": "You have to predict 2 output for  2A3 and DMS. But the value is only one for each nucleotide. How to compare the output with target for calculating the cost in a Neural Network.",
      "votes": null
    },
    {
      "id": "2522078",
      "postDate": "11/12/2023 11:11:45",
      "content": "<p>Ah now I understand your question. The train are divided in half for each experiment. The first half corresponds to 2A3 while second half to DMS but both are the same sequences. So join them in order to have both labels for each sequence.</p>",
      "rawMarkdown": "Ah now I understand your question. The train are divided in half for each experiment. The first half corresponds to 2A3 while second half to DMS but both are the same sequences. So join them in order to have both labels for each sequence.",
      "votes": null
    },
    {
      "id": "2522080",
      "postDate": "11/12/2023 11:15:24",
      "content": "<p>Am I able to succeed with a 4.6ghz Ryzen 5800X this task ? There is a lot of data.</p>",
      "rawMarkdown": "Am I able to succeed with a 4.6ghz Ryzen 5800X this task ? There is a lot of data.",
      "votes": null
    },
    {
      "id": "2522086",
      "postDate": "11/12/2023 11:21:14",
      "content": "<p>I'm completely casual in hardware sorry. You have weeckly quota in the kaggle GPU and just for learn I'm training in my personal computer with a RTX 3060.<br>\nIf you are asking about memmory if you can read the whole csv you can store half csv with double columns. Other option could be train one model per experiment.</p>",
      "rawMarkdown": "I'm completely casual in hardware sorry. You have weeckly quota in the kaggle GPU and just for learn I'm training in my personal computer with a RTX 3060.\nIf you are asking about memmory if you can read the whole csv you can store half csv with double columns. Other option could be train one model per experiment.",
      "votes": null
    },
    {
      "id": "2522095",
      "postDate": "11/12/2023 11:34:44",
      "content": "<p>With kaggle gpu you have to open a notebook and put your code. It's that ?</p>",
      "rawMarkdown": "With kaggle gpu you have to open a notebook and put your code. It's that ?",
      "votes": null
    },
    {
      "id": "2522118",
      "postDate": "11/12/2023 11:57:30",
      "content": "<p>And select the accelerator, yes.</p>",
      "rawMarkdown": "And select the accelerator, yes.",
      "votes": null
    },
    {
      "id": "2523140",
      "postDate": "11/13/2023 09:48:35",
      "content": "<p>What is your time of prediction ? I calculated and it takes 2 days or more with notebook .</p>",
      "rawMarkdown": "What is your time of prediction ? I calculated and it takes 2 days or more with notebook .",
      "votes": null
    },
    {
      "id": "2523162",
      "postDate": "11/13/2023 10:19:23",
      "content": "<p>Should be around a few hours even with kaggle GPU with a basic model. Be sure you are using GPU. There are some inference exemples on code section.</p>",
      "rawMarkdown": "Should be around a few hours even with kaggle GPU with a basic model. Be sure you are using GPU. There are some inference exemples on code section.",
      "votes": null
    },
    {
      "id": "2523261",
      "postDate": "11/13/2023 11:31:12",
      "content": "<p>what do you mean by basic model ?</p>",
      "rawMarkdown": "what do you mean by basic model ?",
      "votes": null
    },
    {
      "id": "2523821",
      "postDate": "11/13/2023 19:26:27",
      "content": "<p>Depending on your approach. If you are doing an ensemble of very heavey models or just a single moderate large  model.</p>",
      "rawMarkdown": "Depending on your approach. If you are doing an ensemble of very heavey models or just a single moderate large  model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2522067,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "11/12/2023 11:01:33",
      "content": "<p>No, the target are the values for both experiments 2A3 and DMS. The values at columns like 'reactivity_0…' (not the ones with reactivity_error_0…). So for each RNA sequence you must predict the set of reactivities corresponding to each nucleotide and experiment. It's a regression model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2522073,
          "author_name": "rgismeyssonnier",
          "author_url": "",
          "post_date": "11/12/2023 11:09:01",
          "content": "<p>You have to predict 2 output for  2A3 and DMS. But the value is only one for each nucleotide. How to compare the output with target for calculating the cost in a Neural Network.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2522078,
              "author_name": "sacuscreed",
              "author_url": "",
              "post_date": "11/12/2023 11:11:45",
              "content": "<p>Ah now I understand your question. The train are divided in half for each experiment. The first half corresponds to 2A3 while second half to DMS but both are the same sequences. So join them in order to have both labels for each sequence.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2522080,
                  "author_name": "rgismeyssonnier",
                  "author_url": "",
                  "post_date": "11/12/2023 11:15:24",
                  "content": "<p>Am I able to succeed with a 4.6ghz Ryzen 5800X this task ? There is a lot of data.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2522086,
                      "author_name": "sacuscreed",
                      "author_url": "",
                      "post_date": "11/12/2023 11:21:14",
                      "content": "<p>I'm completely casual in hardware sorry. You have weeckly quota in the kaggle GPU and just for learn I'm training in my personal computer with a RTX 3060.<br>\nIf you are asking about memmory if you can read the whole csv you can store half csv with double columns. Other option could be train one model per experiment.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2522095,
                          "author_name": "rgismeyssonnier",
                          "author_url": "",
                          "post_date": "11/12/2023 11:34:44",
                          "content": "<p>With kaggle gpu you have to open a notebook and put your code. It's that ?</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2522118,
                              "author_name": "sacuscreed",
                              "author_url": "",
                              "post_date": "11/12/2023 11:57:30",
                              "content": "<p>And select the accelerator, yes.</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2522070,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "11/12/2023 11:04:33",
      "content": "<p>The test submission is a bit more complex since you need to submit the concatenated predictions of each test sequence.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2523140,
          "author_name": "rgismeyssonnier",
          "author_url": "",
          "post_date": "11/13/2023 09:48:35",
          "content": "<p>What is your time of prediction ? I calculated and it takes 2 days or more with notebook .</p>",
          "votes": null,
          "replies": [
            {
              "id": 2523162,
              "author_name": "sacuscreed",
              "author_url": "",
              "post_date": "11/13/2023 10:19:23",
              "content": "<p>Should be around a few hours even with kaggle GPU with a basic model. Be sure you are using GPU. There are some inference exemples on code section.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2523261,
                  "author_name": "rgismeyssonnier",
                  "author_url": "",
                  "post_date": "11/13/2023 11:31:12",
                  "content": "<p>what do you mean by basic model ?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2523821,
                      "author_name": "sacuscreed",
                      "author_url": "",
                      "post_date": "11/13/2023 19:26:27",
                      "content": "<p>Depending on your approach. If you are doing an ensemble of very heavey models or just a single moderate large  model.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2521432": "What are the target to compare with the output. We test if it's a one of the 2 types indicated by the experiment field ?",
    "2522067": "No, the target are the values for both experiments 2A3 and DMS. The values at columns like 'reactivity_0...' (not the ones with reactivity_error_0...). So for each RNA sequence you must predict the set of reactivities corresponding to each nucleotide and experiment. It's a regression model.",
    "2522070": "The test submission is a bit more complex since you need to submit the concatenated predictions of each test sequence.",
    "2522073": "You have to predict 2 output for  2A3 and DMS. But the value is only one for each nucleotide. How to compare the output with target for calculating the cost in a Neural Network.",
    "2522078": "Ah now I understand your question. The train are divided in half for each experiment. The first half corresponds to 2A3 while second half to DMS but both are the same sequences. So join them in order to have both labels for each sequence.",
    "2522080": "Am I able to succeed with a 4.6ghz Ryzen 5800X this task ? There is a lot of data.",
    "2522086": "I'm completely casual in hardware sorry. You have weeckly quota in the kaggle GPU and just for learn I'm training in my personal computer with a RTX 3060.\nIf you are asking about memmory if you can read the whole csv you can store half csv with double columns. Other option could be train one model per experiment.",
    "2522095": "With kaggle gpu you have to open a notebook and put your code. It's that ?",
    "2522118": "And select the accelerator, yes.",
    "2523140": "What is your time of prediction ? I calculated and it takes 2 days or more with notebook .",
    "2523162": "Should be around a few hours even with kaggle GPU with a basic model. Be sure you are using GPU. There are some inference exemples on code section.",
    "2523261": "what do you mean by basic model ?",
    "2523821": "Depending on your approach. If you are doing an ensemble of very heavey models or just a single moderate large  model."
  },
  "source": "meta"
}