{
  "id": 566529,
  "title": "Day 1: General understanding of the data",
  "url": "/competitions/stanford-rna-3d-folding/discussion/566529",
  "author_name": "Pastor Soto",
  "post_date": "2025-03-06T00:18:03.683000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>My first day at the competition involves understand the data and making a sample submission.</p>\n<p>I am replicating the code of this <a href=\"https://www.kaggle.com/code/olaflundstrom/stanford-rna-3d-folding-kaggle-competition\" target=\"_blank\">notebook</a>, the approach is to code along without copying and pasting ask LLM LLMs once I don't understand something  and make improvements when needed</p>\n<h2>Discussion points:</h2>\n<ul>\n<li>The data contains the sequence and the position of the RNA in a 3-dimensional space (x, y, z).</li>\n<li>Preprocessing the data to map the labels with the training sequence</li>\n<li>Normalize the values to make easier comparisons between the data</li>\n<li>Feature engineering can be started with one-hot encoding, GC proportion, normalized  position of each nucleotide in the sequence</li>\n</ul>\n<h2>Next steps:</h2>\n<ul>\n<li>Make the submission with the code of the notebook</li>\n<li>Make small changes to evaluate the results</li>\n<li>Set an evaluation system on the validation dataset</li>\n<li>EDA to understand the data</li>\n<li>Ribonanza similarities of the dataset</li>\n</ul>",
  "messages": [
    {
      "id": 3141835,
      "postDate": "2025-03-06T00:18:03.683Z",
      "content": "<p>My first day at the competition involves understand the data and making a sample submission.</p>\n<p>I am replicating the code of this <a href=\"https://www.kaggle.com/code/olaflundstrom/stanford-rna-3d-folding-kaggle-competition\" target=\"_blank\">notebook</a>, the approach is to code along without copying and pasting ask LLM LLMs once I don't understand something  and make improvements when needed</p>\n<h2>Discussion points:</h2>\n<ul>\n<li>The data contains the sequence and the position of the RNA in a 3-dimensional space (x, y, z).</li>\n<li>Preprocessing the data to map the labels with the training sequence</li>\n<li>Normalize the values to make easier comparisons between the data</li>\n<li>Feature engineering can be started with one-hot encoding, GC proportion, normalized  position of each nucleotide in the sequence</li>\n</ul>\n<h2>Next steps:</h2>\n<ul>\n<li>Make the submission with the code of the notebook</li>\n<li>Make small changes to evaluate the results</li>\n<li>Set an evaluation system on the validation dataset</li>\n<li>EDA to understand the data</li>\n<li>Ribonanza similarities of the dataset</li>\n</ul>",
      "rawMarkdown": "My first day at the competition involves understand the data and making a sample submission.\n\nI am replicating the code of this [notebook](https://www.kaggle.com/code/olaflundstrom/stanford-rna-3d-folding-kaggle-competition), the approach is to code along without copying and pasting ask LLM LLMs once I don't understand something  and make improvements when needed\n\n## Discussion points:\n\n- The data contains the sequence and the position of the RNA in a 3-dimensional space (x, y, z).\n- Preprocessing the data to map the labels with the training sequence\n- Normalize the values to make easier comparisons between the data\n- Feature engineering can be started with one-hot encoding, GC proportion, normalized  position of each nucleotide in the sequence\n\n## Next steps:\n- Make the submission with the code of the notebook\n- Make small changes to evaluate the results\n- Set an evaluation system on the validation dataset\n- EDA to understand the data\n- Ribonanza similarities of the dataset",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3141835": "My first day at the competition involves understand the data and making a sample submission.\n\nI am replicating the code of this [notebook](https://www.kaggle.com/code/olaflundstrom/stanford-rna-3d-folding-kaggle-competition), the approach is to code along without copying and pasting ask LLM LLMs once I don't understand something  and make improvements when needed\n\n## Discussion points:\n\n- The data contains the sequence and the position of the RNA in a 3-dimensional space (x, y, z).\n- Preprocessing the data to map the labels with the training sequence\n- Normalize the values to make easier comparisons between the data\n- Feature engineering can be started with one-hot encoding, GC proportion, normalized  position of each nucleotide in the sequence\n\n## Next steps:\n- Make the submission with the code of the notebook\n- Make small changes to evaluate the results\n- Set an evaluation system on the validation dataset\n- EDA to understand the data\n- Ribonanza similarities of the dataset"
  }
}