{
  "id": 578148,
  "title": "fine-tuning model mystery",
  "url": "/competitions/stanford-rna-3d-folding/discussion/578148",
  "author_name": "doheon114",
  "post_date": "2025-05-09T04:30:41.037000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I'm currently working on enhancing performance by fine-tuning open-source models that were not trained on recent RNA PDB entries (many of them have not been fine-tuned with 2023 data).</p>\n<p>However, when I fine-tune these models using recent PDB data (mainly from 2023), they almost always overfit to the CASP15 validation set and even to the 2024 RNA data extracted from the training set (v1)—but not to the hidden LB dataset. In fact, the LB score consistently decreases as I apply more fine-tuned open-source models.</p>\n<p>The figure below shows the results of my fine-tuning experiments:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15620263%2Fb431cb6dced16485f69f4ae7c592caed%2F2025-05-09%20%201.27.20.png?generation=1746764887308012&amp;alt=media\" alt=\"\"></p>\n<p>From this, I suspect that my fine-tuned models are overfitting to the CASP15 dataset, which has a cutoff around 2022—this means my ~2023-based fine-tuning data could be indirectly \"cheating.\" If this hypothesis is correct, I would expect the fine-tuned models to perform worse on RNA data from 2024 too.</p>\n<p>But here's the mystery: for the <strong>2024 RNA samples</strong> extracted from the training set, the fine-tuned model actually performs better than the original pre-trained model! This contradicts the overfitting hypothesis and leaves me puzzled.</p>",
  "messages": [
    {
      "id": 3198104,
      "postDate": "2025-05-09T04:30:41.037Z",
      "content": "<p>I'm currently working on enhancing performance by fine-tuning open-source models that were not trained on recent RNA PDB entries (many of them have not been fine-tuned with 2023 data).</p>\n<p>However, when I fine-tune these models using recent PDB data (mainly from 2023), they almost always overfit to the CASP15 validation set and even to the 2024 RNA data extracted from the training set (v1)—but not to the hidden LB dataset. In fact, the LB score consistently decreases as I apply more fine-tuned open-source models.</p>\n<p>The figure below shows the results of my fine-tuning experiments:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15620263%2Fb431cb6dced16485f69f4ae7c592caed%2F2025-05-09%20%201.27.20.png?generation=1746764887308012&amp;alt=media\" alt=\"\"></p>\n<p>From this, I suspect that my fine-tuned models are overfitting to the CASP15 dataset, which has a cutoff around 2022—this means my ~2023-based fine-tuning data could be indirectly \"cheating.\" If this hypothesis is correct, I would expect the fine-tuned models to perform worse on RNA data from 2024 too.</p>\n<p>But here's the mystery: for the <strong>2024 RNA samples</strong> extracted from the training set, the fine-tuned model actually performs better than the original pre-trained model! This contradicts the overfitting hypothesis and leaves me puzzled.</p>",
      "rawMarkdown": "I'm currently working on enhancing performance by fine-tuning open-source models that were not trained on recent RNA PDB entries (many of them have not been fine-tuned with 2023 data).\n\nHowever, when I fine-tune these models using recent PDB data (mainly from 2023), they almost always overfit to the CASP15 validation set and even to the 2024 RNA data extracted from the training set (v1)—but not to the hidden LB dataset. In fact, the LB score consistently decreases as I apply more fine-tuned open-source models.\n\nThe figure below shows the results of my fine-tuning experiments:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15620263%2Fb431cb6dced16485f69f4ae7c592caed%2F2025-05-09%20%201.27.20.png?generation=1746764887308012&alt=media)\n\nFrom this, I suspect that my fine-tuned models are overfitting to the CASP15 dataset, which has a cutoff around 2022—this means my ~2023-based fine-tuning data could be indirectly \"cheating.\" If this hypothesis is correct, I would expect the fine-tuned models to perform worse on RNA data from 2024 too.\n\nBut here's the mystery: for the **2024 RNA samples** extracted from the training set, the fine-tuned model actually performs better than the original pre-trained model! This contradicts the overfitting hypothesis and leaves me puzzled.\n\n",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3198104": "I'm currently working on enhancing performance by fine-tuning open-source models that were not trained on recent RNA PDB entries (many of them have not been fine-tuned with 2023 data).\n\nHowever, when I fine-tune these models using recent PDB data (mainly from 2023), they almost always overfit to the CASP15 validation set and even to the 2024 RNA data extracted from the training set (v1)—but not to the hidden LB dataset. In fact, the LB score consistently decreases as I apply more fine-tuned open-source models.\n\nThe figure below shows the results of my fine-tuning experiments:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15620263%2Fb431cb6dced16485f69f4ae7c592caed%2F2025-05-09%20%201.27.20.png?generation=1746764887308012&alt=media)\n\nFrom this, I suspect that my fine-tuned models are overfitting to the CASP15 dataset, which has a cutoff around 2022—this means my ~2023-based fine-tuning data could be indirectly \"cheating.\" If this hypothesis is correct, I would expect the fine-tuned models to perform worse on RNA data from 2024 too.\n\nBut here's the mystery: for the **2024 RNA samples** extracted from the training set, the fine-tuned model actually performs better than the original pre-trained model! This contradicts the overfitting hypothesis and leaves me puzzled.\n\n"
  }
}