{
  "id": 565885,
  "title": "Limited data",
  "url": "/competitions/stanford-rna-3d-folding/discussion/565885",
  "author_name": "",
  "post_date": "2025-03-02T15:11:04.696069500Z",
  "votes": null,
  "comment_count": 8,
  "views": 0,
  "content": "<p>hi guys </p>\n<p>when I look at our training set, it's only around 800 data points. do you think we should request more data ? with 800 data points I don't think we can touch on those deep learning model or I'm wrong please correct me</p>",
  "messages": [
    {
      "id": "3138443",
      "postDate": "03/02/2025 15:11:04",
      "content": "<p>hi guys </p>\n<p>when I look at our training set, it's only around 800 data points. do you think we should request more data ? with 800 data points I don't think we can touch on those deep learning model or I'm wrong please correct me</p>",
      "rawMarkdown": "hi guys \n\nwhen I look at our training set, it's only around 800 data points. do you think we should request more data ? with 800 data points I don't think we can touch on those deep learning model or I'm wrong please correct me",
      "votes": null
    },
    {
      "id": "3138525",
      "postDate": "03/02/2025 16:35:26",
      "content": "<p>In the data section you can see the following:</p>\n<blockquote>\n  <p>The developers of RFdiffusion have made available a synthetic data set of over 400,000 RNA structures here.</p>\n</blockquote>\n<p>link to the data - <a href=\"https://www.kaggle.com/datasets/andrewfavor/uw-synthetic-rna-structures\" target=\"_blank\">https://www.kaggle.com/datasets/andrewfavor/uw-synthetic-rna-structures</a></p>",
      "rawMarkdown": "In the data section you can see the following:\n>The developers of RFdiffusion have made available a synthetic data set of over 400,000 RNA structures here.\n\nlink to the data - https://www.kaggle.com/datasets/andrewfavor/uw-synthetic-rna-structures",
      "votes": null
    },
    {
      "id": "3138555",
      "postDate": "03/02/2025 17:24:50",
      "content": "<p>thank you I want to make sure we only predict x_1 y_1 and z_1 right ?</p>",
      "rawMarkdown": "thank you I want to make sure we only predict x_1 y_1 and z_1 right ?",
      "votes": null
    },
    {
      "id": "3138600",
      "postDate": "03/02/2025 18:20:17",
      "content": "<p>yes, and you have to make 5 predictions per sequence (x1,y1,z1 … x5,y5,z5)</p>",
      "rawMarkdown": "yes, and you have to make 5 predictions per sequence (x1,y1,z1 ... x5,y5,z5)",
      "votes": null
    },
    {
      "id": "3138612",
      "postDate": "03/02/2025 18:24:15",
      "content": "<p>But our train data stop at _1 I see nothing after that soooooo how can I predict _2 and above ?</p>",
      "rawMarkdown": "But our train data stop at _1 I see nothing after that soooooo how can I predict _2 and above ?",
      "votes": null
    },
    {
      "id": "3138667",
      "postDate": "03/02/2025 19:30:13",
      "content": "<p><a href=\"https://www.kaggle.com/louisstefanuto\" target=\"_blank\">@louisstefanuto</a> <a href=\"https://www.kaggle.com/amitaharoni\" target=\"_blank\">@amitaharoni</a> </p>",
      "rawMarkdown": "louisstefanuto @amitaharoni",
      "votes": null
    },
    {
      "id": "3138799",
      "postDate": "03/02/2025 22:57:57",
      "content": "<p>yes but one sequence can take multiple shapes, usually there is only one, but it can take many<br>\nfor instance if you check the validation_labels.csv file, you will see that there is actually 40 x 3 columns. Meaning that some sequences have been experimentally detected in 40 different shapes</p>",
      "rawMarkdown": "yes but one sequence can take multiple shapes, usually there is only one, but it can take many\nfor instance if you check the validation_labels.csv file, you will see that there is actually 40 x 3 columns. Meaning that some sequences have been experimentally detected in 40 different shapes",
      "votes": null
    },
    {
      "id": "3139044",
      "postDate": "03/03/2025 06:20:10",
      "content": "<p>There is no such a thing as requesting more data in any Kaggle competition, and especially so in this one. The number of structurally characterized RNAs is small, and that's the way it is.</p>",
      "rawMarkdown": "There is no such a thing as requesting more data in any Kaggle competition, and especially so in this one. The number of structurally characterized RNAs is small, and that's the way it is.",
      "votes": null
    },
    {
      "id": "3139062",
      "postDate": "03/03/2025 06:42:49",
      "content": "<p>thanks for letting me know just that when I look at monthly competitions by Kaggle they do provide extra data sometimes </p>",
      "rawMarkdown": "thanks for letting me know just that when I look at monthly competitions by Kaggle they do provide extra data sometimes",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3138525,
      "author_name": "amitaharoni",
      "author_url": "",
      "post_date": "03/02/2025 16:35:26",
      "content": "<p>In the data section you can see the following:</p>\n<blockquote>\n  <p>The developers of RFdiffusion have made available a synthetic data set of over 400,000 RNA structures here.</p>\n</blockquote>\n<p>link to the data - <a href=\"https://www.kaggle.com/datasets/andrewfavor/uw-synthetic-rna-structures\" target=\"_blank\">https://www.kaggle.com/datasets/andrewfavor/uw-synthetic-rna-structures</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 3138555,
          "author_name": "theonetheonly",
          "author_url": "",
          "post_date": "03/02/2025 17:24:50",
          "content": "<p>thank you I want to make sure we only predict x_1 y_1 and z_1 right ?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3138600,
              "author_name": "louisstefanuto",
              "author_url": "",
              "post_date": "03/02/2025 18:20:17",
              "content": "<p>yes, and you have to make 5 predictions per sequence (x1,y1,z1 … x5,y5,z5)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3138612,
                  "author_name": "theonetheonly",
                  "author_url": "",
                  "post_date": "03/02/2025 18:24:15",
                  "content": "<p>But our train data stop at _1 I see nothing after that soooooo how can I predict _2 and above ?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3138667,
                      "author_name": "theonetheonly",
                      "author_url": "",
                      "post_date": "03/02/2025 19:30:13",
                      "content": "<p><a href=\"https://www.kaggle.com/louisstefanuto\" target=\"_blank\">@louisstefanuto</a> <a href=\"https://www.kaggle.com/amitaharoni\" target=\"_blank\">@amitaharoni</a> </p>",
                      "votes": null,
                      "replies": []
                    },
                    {
                      "id": 3138799,
                      "author_name": "louisstefanuto",
                      "author_url": "",
                      "post_date": "03/02/2025 22:57:57",
                      "content": "<p>yes but one sequence can take multiple shapes, usually there is only one, but it can take many<br>\nfor instance if you check the validation_labels.csv file, you will see that there is actually 40 x 3 columns. Meaning that some sequences have been experimentally detected in 40 different shapes</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3139044,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "03/03/2025 06:20:10",
      "content": "<p>There is no such a thing as requesting more data in any Kaggle competition, and especially so in this one. The number of structurally characterized RNAs is small, and that's the way it is.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3139062,
          "author_name": "theonetheonly",
          "author_url": "",
          "post_date": "03/03/2025 06:42:49",
          "content": "<p>thanks for letting me know just that when I look at monthly competitions by Kaggle they do provide extra data sometimes </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3138443": "hi guys \n\nwhen I look at our training set, it's only around 800 data points. do you think we should request more data ? with 800 data points I don't think we can touch on those deep learning model or I'm wrong please correct me",
    "3138525": "In the data section you can see the following:\n>The developers of RFdiffusion have made available a synthetic data set of over 400,000 RNA structures here.\n\nlink to the data - https://www.kaggle.com/datasets/andrewfavor/uw-synthetic-rna-structures",
    "3138555": "thank you I want to make sure we only predict x_1 y_1 and z_1 right ?",
    "3138600": "yes, and you have to make 5 predictions per sequence (x1,y1,z1 ... x5,y5,z5)",
    "3138612": "But our train data stop at _1 I see nothing after that soooooo how can I predict _2 and above ?",
    "3138667": "louisstefanuto @amitaharoni",
    "3138799": "yes but one sequence can take multiple shapes, usually there is only one, but it can take many\nfor instance if you check the validation_labels.csv file, you will see that there is actually 40 x 3 columns. Meaning that some sequences have been experimentally detected in 40 different shapes",
    "3139044": "There is no such a thing as requesting more data in any Kaggle competition, and especially so in this one. The number of structurally characterized RNAs is small, and that's the way it is.",
    "3139062": "thanks for letting me know just that when I look at monthly competitions by Kaggle they do provide extra data sometimes"
  },
  "source": "meta"
}