{
  "id": 584728,
  "title": "Is well correlating model validation in the current setting possible?",
  "url": "/competitions/drw-crypto-market-prediction/discussion/584728",
  "author_name": "",
  "post_date": "2025-06-15T15:21:04.041712400Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Since no participant openly described or mentioned any method with a solid <a href=\"https://www.kaggle.com/competitions/drw-crypto-market-prediction/\" target=\"_blank\">CV-LB Correlation</a> I went back to square one and began to rethink the validation process in this competition. </p>\n<p>What is odd about the competitions LB and training data split is that they are almost equally long (Training data has 525887, LB Data has 538150 rows). If we visuallize the split it would approximately look as following (note that the host declared they would use the newer data for the final scores while using 49% for Public LB):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15117559%2F8715f3d96c9e6357b803483c92953086%2FUnbenanntes%20Diagramm.drawio.png?generation=1749999421106376&amp;alt=media\" alt=\"Competition Data Split\"></p>\n<p>So I believe if we want to simulate our final score in a somewhat realistic matter we would have to do something like </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15117559%2F0bef0a71d33538c5c9ec3e1d77735141%2FUnbenanntes%20Diagramm.drawio(1).png?generation=1750000551674968&amp;alt=media\" alt=\"Realistic LB Score Data Split\"></p>\n<p>which cuts half of our data and will likely force us to use either the Public LB as a sort of split which is not feasible for proper HPO or we will have to use way smaller cv splits, which would decrease our actual training data percentage even lower, to a point where we would use less training datapoints then we would finally predict with the model for private LB. </p>\n<p>What do you think about this?</p>",
  "messages": [
    {
      "id": "3224858",
      "postDate": "06/15/2025 15:21:04",
      "content": "<p>Since no participant openly described or mentioned any method with a solid <a href=\"https://www.kaggle.com/competitions/drw-crypto-market-prediction/\" target=\"_blank\">CV-LB Correlation</a> I went back to square one and began to rethink the validation process in this competition. </p>\n<p>What is odd about the competitions LB and training data split is that they are almost equally long (Training data has 525887, LB Data has 538150 rows). If we visuallize the split it would approximately look as following (note that the host declared they would use the newer data for the final scores while using 49% for Public LB):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15117559%2F8715f3d96c9e6357b803483c92953086%2FUnbenanntes%20Diagramm.drawio.png?generation=1749999421106376&amp;alt=media\" alt=\"Competition Data Split\"></p>\n<p>So I believe if we want to simulate our final score in a somewhat realistic matter we would have to do something like </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15117559%2F0bef0a71d33538c5c9ec3e1d77735141%2FUnbenanntes%20Diagramm.drawio(1).png?generation=1750000551674968&amp;alt=media\" alt=\"Realistic LB Score Data Split\"></p>\n<p>which cuts half of our data and will likely force us to use either the Public LB as a sort of split which is not feasible for proper HPO or we will have to use way smaller cv splits, which would decrease our actual training data percentage even lower, to a point where we would use less training datapoints then we would finally predict with the model for private LB. </p>\n<p>What do you think about this?</p>",
      "rawMarkdown": "Since no participant openly described or mentioned any method with a solid [CV-LB Correlation](https://www.kaggle.com/competitions/drw-crypto-market-prediction/) I went back to square one and began to rethink the validation process in this competition. \n\nWhat is odd about the competitions LB and training data split is that they are almost equally long (Training data has 525887, LB Data has 538150 rows). If we visuallize the split it would approximately look as following (note that the host declared they would use the newer data for the final scores while using 49% for Public LB):\n![Competition Data Split](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15117559%2F8715f3d96c9e6357b803483c92953086%2FUnbenanntes%20Diagramm.drawio.png?generation=1749999421106376&alt=media)\n\nSo I believe if we want to simulate our final score in a somewhat realistic matter we would have to do something like \n\n![Realistic LB Score Data Split](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15117559%2F0bef0a71d33538c5c9ec3e1d77735141%2FUnbenanntes%20Diagramm.drawio(1).png?generation=1750000551674968&alt=media)\n\nwhich cuts half of our data and will likely force us to use either the Public LB as a sort of split which is not feasible for proper HPO or we will have to use way smaller cv splits, which would decrease our actual training data percentage even lower, to a point where we would use less training datapoints then we would finally predict with the model for private LB. \n\nWhat do you think about this?",
      "votes": null
    },
    {
      "id": "3224918",
      "postDate": "06/15/2025 16:27:50",
      "content": "<p>I tried it on my last submission today.<br>\nHere a result:<br>\ntrain: 0.7<br>\nval: 0.125<br>\ntest: 0.08848</p>\n<p>I used only 1 lgbm model.</p>",
      "rawMarkdown": "I tried it on my last submission today.\nHere a result:\ntrain: 0.7\nval: 0.125\ntest: 0.08848\n\nI used only 1 lgbm model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3224918,
      "author_name": "nguyennguyen599",
      "author_url": "",
      "post_date": "06/15/2025 16:27:50",
      "content": "<p>I tried it on my last submission today.<br>\nHere a result:<br>\ntrain: 0.7<br>\nval: 0.125<br>\ntest: 0.08848</p>\n<p>I used only 1 lgbm model.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3224858": "Since no participant openly described or mentioned any method with a solid [CV-LB Correlation](https://www.kaggle.com/competitions/drw-crypto-market-prediction/) I went back to square one and began to rethink the validation process in this competition. \n\nWhat is odd about the competitions LB and training data split is that they are almost equally long (Training data has 525887, LB Data has 538150 rows). If we visuallize the split it would approximately look as following (note that the host declared they would use the newer data for the final scores while using 49% for Public LB):\n![Competition Data Split](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15117559%2F8715f3d96c9e6357b803483c92953086%2FUnbenanntes%20Diagramm.drawio.png?generation=1749999421106376&alt=media)\n\nSo I believe if we want to simulate our final score in a somewhat realistic matter we would have to do something like \n\n![Realistic LB Score Data Split](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15117559%2F0bef0a71d33538c5c9ec3e1d77735141%2FUnbenanntes%20Diagramm.drawio(1).png?generation=1750000551674968&alt=media)\n\nwhich cuts half of our data and will likely force us to use either the Public LB as a sort of split which is not feasible for proper HPO or we will have to use way smaller cv splits, which would decrease our actual training data percentage even lower, to a point where we would use less training datapoints then we would finally predict with the model for private LB. \n\nWhat do you think about this?",
    "3224918": "I tried it on my last submission today.\nHere a result:\ntrain: 0.7\nval: 0.125\ntest: 0.08848\n\nI used only 1 lgbm model."
  },
  "source": "meta"
}