{
  "id": 569070,
  "title": "Confused about validation strategy",
  "url": "/competitions/stanford-rna-3d-folding/discussion/569070",
  "author_name": "moth",
  "post_date": "2025-03-19T19:42:14.936000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I am performing validation as proposed by the competition's host in his training code of <code>RibonanzaNet</code> <a href=\"https://www.kaggle.com/code/shujun717/ribonanzanet-3d-finetune\" target=\"_blank\">here</a>. In this code, he splits the training data in <code>train_sequences.csv</code> and <code>train_labels.csv</code> by dates (<code>2020-01-01</code>, <code>2022-05-01</code>) to obtain a new train and validation sets. After the data is split these are the following shapes:</p>\n<pre><code>df_train_sequences : (, )\n—————————————————————————————————————————————————————\ndf_train_labels : (, )\n—————————————————————————————————————————————————————\ndf_valid_sequences : (, )\n—————————————————————————————————————————————————————\ndf_valid_labels : (, )\n—————————————————————————————————————————————————————\n</code></pre>\n<p>However, we also have the <code>validation_labels.csv</code> which is a different file with sequences not present in the <code>train_labels.csv</code> file.</p>\n<p>My questions are:</p>\n<ol>\n<li>I haven't find a proper way to correctly apply the competition's metric to the validation set that comes from splitting the training data (here I do not refer to <code>validation_labels.csv</code>). For instance, I predict all coordinates for my validation set (roughly 26k samples) so the evaluation metric takes forever. Has anyone applied the metric to the validation set at training time? How are you evaluating your model performance on unseen predictions? The host evaluates the best model as the one with the best validation loss (not best validation metric).</li>\n<li>In addition, I quite not understand the purpose of the 40 different sets of points per row in the <code>validation_labels.csv</code>. Will we have 40 different estimates for every nucleotide base (A,C,G,U) as labels?</li>\n<li>What is the purpose of splitting the training data into <code>train</code> and <code>validation</code> sets if we already have a <code>validation_labels.csv</code> file that can be used as a validation set?</li>\n</ol>\n<p>Any clarifications on these matters is greatly appreciated!</p>",
  "messages": [
    {
      "id": 3154309,
      "postDate": "2025-03-19T19:42:14.937Z",
      "content": "<p>I am performing validation as proposed by the competition's host in his training code of <code>RibonanzaNet</code> <a href=\"https://www.kaggle.com/code/shujun717/ribonanzanet-3d-finetune\" target=\"_blank\">here</a>. In this code, he splits the training data in <code>train_sequences.csv</code> and <code>train_labels.csv</code> by dates (<code>2020-01-01</code>, <code>2022-05-01</code>) to obtain a new train and validation sets. After the data is split these are the following shapes:</p>\n<pre><code>df_train_sequences : (, )\n—————————————————————————————————————————————————————\ndf_train_labels : (, )\n—————————————————————————————————————————————————————\ndf_valid_sequences : (, )\n—————————————————————————————————————————————————————\ndf_valid_labels : (, )\n—————————————————————————————————————————————————————\n</code></pre>\n<p>However, we also have the <code>validation_labels.csv</code> which is a different file with sequences not present in the <code>train_labels.csv</code> file.</p>\n<p>My questions are:</p>\n<ol>\n<li>I haven't find a proper way to correctly apply the competition's metric to the validation set that comes from splitting the training data (here I do not refer to <code>validation_labels.csv</code>). For instance, I predict all coordinates for my validation set (roughly 26k samples) so the evaluation metric takes forever. Has anyone applied the metric to the validation set at training time? How are you evaluating your model performance on unseen predictions? The host evaluates the best model as the one with the best validation loss (not best validation metric).</li>\n<li>In addition, I quite not understand the purpose of the 40 different sets of points per row in the <code>validation_labels.csv</code>. Will we have 40 different estimates for every nucleotide base (A,C,G,U) as labels?</li>\n<li>What is the purpose of splitting the training data into <code>train</code> and <code>validation</code> sets if we already have a <code>validation_labels.csv</code> file that can be used as a validation set?</li>\n</ol>\n<p>Any clarifications on these matters is greatly appreciated!</p>",
      "rawMarkdown": "I am performing validation as proposed by the competition's host in his training code of `RibonanzaNet` [here](https://www.kaggle.com/code/shujun717/ribonanzanet-3d-finetune). In this code, he splits the training data in `train_sequences.csv` and `train_labels.csv` by dates (`2020-01-01`, `2022-05-01`) to obtain a new train and validation sets. After the data is split these are the following shapes:\n```\ndf_train_sequences shape: (542, 5)\n—————————————————————————————————————————————————————\ndf_train_labels shape: (84566, 11)\n—————————————————————————————————————————————————————\ndf_valid_sequences shape: (80, 5)\n—————————————————————————————————————————————————————\ndf_valid_labels shape: (26071, 12)\n—————————————————————————————————————————————————————\n```\n\nHowever, we also have the `validation_labels.csv` which is a different file with sequences not present in the `train_labels.csv` file.\n\nMy questions are:\n1. I haven't find a proper way to correctly apply the competition's metric to the validation set that comes from splitting the training data (here I do not refer to `validation_labels.csv`). For instance, I predict all coordinates for my validation set (roughly 26k samples) so the evaluation metric takes forever. Has anyone applied the metric to the validation set at training time? How are you evaluating your model performance on unseen predictions? The host evaluates the best model as the one with the best validation loss (not best validation metric).\n2. In addition, I quite not understand the purpose of the 40 different sets of points per row in the `validation_labels.csv`. Will we have 40 different estimates for every nucleotide base (A,C,G,U) as labels?\n3. What is the purpose of splitting the training data into `train` and `validation` sets if we already have a `validation_labels.csv` file that can be used as a validation set?\n\nAny clarifications on these matters is greatly appreciated!",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3154309": "I am performing validation as proposed by the competition's host in his training code of `RibonanzaNet` [here](https://www.kaggle.com/code/shujun717/ribonanzanet-3d-finetune). In this code, he splits the training data in `train_sequences.csv` and `train_labels.csv` by dates (`2020-01-01`, `2022-05-01`) to obtain a new train and validation sets. After the data is split these are the following shapes:\n```\ndf_train_sequences shape: (542, 5)\n—————————————————————————————————————————————————————\ndf_train_labels shape: (84566, 11)\n—————————————————————————————————————————————————————\ndf_valid_sequences shape: (80, 5)\n—————————————————————————————————————————————————————\ndf_valid_labels shape: (26071, 12)\n—————————————————————————————————————————————————————\n```\n\nHowever, we also have the `validation_labels.csv` which is a different file with sequences not present in the `train_labels.csv` file.\n\nMy questions are:\n1. I haven't find a proper way to correctly apply the competition's metric to the validation set that comes from splitting the training data (here I do not refer to `validation_labels.csv`). For instance, I predict all coordinates for my validation set (roughly 26k samples) so the evaluation metric takes forever. Has anyone applied the metric to the validation set at training time? How are you evaluating your model performance on unseen predictions? The host evaluates the best model as the one with the best validation loss (not best validation metric).\n2. In addition, I quite not understand the purpose of the 40 different sets of points per row in the `validation_labels.csv`. Will we have 40 different estimates for every nucleotide base (A,C,G,U) as labels?\n3. What is the purpose of splitting the training data into `train` and `validation` sets if we already have a `validation_labels.csv` file that can be used as a validation set?\n\nAny clarifications on these matters is greatly appreciated!"
  }
}