{
  "id": 593920,
  "title": "Wow OK: ideas about what has happened here?",
  "url": "/competitions/drw-crypto-market-prediction/discussion/593920",
  "author_name": "",
  "post_date": "2025-07-31T11:09:00.248573500Z",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I know crypto data is notoriously noisy and regime-dependent, but the majority of submissions on the LB have dropped/ascended 100+ places! In comparison to the few other Kaggle competitions whose leaderboards I've seen, this is a lot.</p>\n<p>For those who maintained their positions in the top 50/top 30/top 10, what did you do to stabilize your public/private correlation? I would love to know. I wonder if I may have hurt myself by trying to squeeze signal out of too many features in my ensemble, which lost predictive power between public and private test sets.</p>",
  "messages": [
    {
      "id": "3258864",
      "postDate": "07/31/2025 11:09:00",
      "content": "<p>I know crypto data is notoriously noisy and regime-dependent, but the majority of submissions on the LB have dropped/ascended 100+ places! In comparison to the few other Kaggle competitions whose leaderboards I've seen, this is a lot.</p>\n<p>For those who maintained their positions in the top 50/top 30/top 10, what did you do to stabilize your public/private correlation? I would love to know. I wonder if I may have hurt myself by trying to squeeze signal out of too many features in my ensemble, which lost predictive power between public and private test sets.</p>",
      "rawMarkdown": "I know crypto data is notoriously noisy and regime-dependent, but the majority of submissions on the LB have dropped/ascended 100+ places! In comparison to the few other Kaggle competitions whose leaderboards I've seen, this is a lot.\n\nFor those who maintained their positions in the top 50/top 30/top 10, what did you do to stabilize your public/private correlation? I would love to know. I wonder if I may have hurt myself by trying to squeeze signal out of too many features in my ensemble, which lost predictive power between public and private test sets.",
      "votes": null
    },
    {
      "id": "3258900",
      "postDate": "07/31/2025 12:42:34",
      "content": "<p>It sounds like your score drop might be related to the curse of dimensionality especially if too many features were involved in the ensemble. My own approach was fairly simple, and I suspect adding more models might have improved my result. Out of curiosity, how many models did you include in your ensemble?</p>",
      "rawMarkdown": "It sounds like your score drop might be related to the curse of dimensionality especially if too many features were involved in the ensemble. My own approach was fairly simple, and I suspect adding more models might have improved my result. Out of curiosity, how many models did you include in your ensemble?",
      "votes": null
    },
    {
      "id": "3258989",
      "postDate": "07/31/2025 16:40:27",
      "content": "<p>There was no way to select the best score correctly. Public and private scores are negatively correlated. I have scores 0.05 - 0.11 and 0.12 - 0.06 on public and private lb. </p>",
      "rawMarkdown": "There was no way to select the best score correctly. Public and private scores are negatively correlated. I have scores 0.05 - 0.11 and 0.12 - 0.06 on public and private lb.",
      "votes": null
    },
    {
      "id": "3258990",
      "postDate": "07/31/2025 16:44:05",
      "content": "<p>Similar story for me, I have 0.12 - 0.06 public vs. 0.06 - 0.10 for private, with rough anti-correlation. The worst part is that my CV and the public LB were actually pretty well-correlated too.</p>",
      "rawMarkdown": "Similar story for me, I have 0.12 - 0.06 public vs. 0.06 - 0.10 for private, with rough anti-correlation. The worst part is that my CV and the public LB were actually pretty well-correlated too.",
      "votes": null
    },
    {
      "id": "3258997",
      "postDate": "07/31/2025 16:56:45",
      "content": "<p>You are right. I noticed the same thing in my submission. It's interesting that even with ensembling multiple models, there was still anti-correlation.</p>",
      "rawMarkdown": "You are right. I noticed the same thing in my submission. It's interesting that even with ensembling multiple models, there was still anti-correlation.",
      "votes": null
    },
    {
      "id": "3259041",
      "postDate": "07/31/2025 18:26:40",
      "content": "<p>It does seem like public vs private score is not super correlated, maybe because this is an anti-overfit problem than an underfit problem.</p>",
      "rawMarkdown": "It does seem like public vs private score is not super correlated, maybe because this is an anti-overfit problem than an underfit problem.",
      "votes": null
    },
    {
      "id": "3259043",
      "postDate": "07/31/2025 18:34:31",
      "content": "<p>Maybe for highly noisy or anti-overfit competition, it would be \"better\" to incorporate IS score to balance luck in OOS and research skill in IS</p>",
      "rawMarkdown": "Maybe for highly noisy or anti-overfit competition, it would be \"better\" to incorporate IS score to balance luck in OOS and research skill in IS",
      "votes": null
    },
    {
      "id": "3259044",
      "postDate": "07/31/2025 18:38:09",
      "content": "<p>This would be nice</p>",
      "rawMarkdown": "This would be nice",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3258900,
      "author_name": "byunjins",
      "author_url": "",
      "post_date": "07/31/2025 12:42:34",
      "content": "<p>It sounds like your score drop might be related to the curse of dimensionality especially if too many features were involved in the ensemble. My own approach was fairly simple, and I suspect adding more models might have improved my result. Out of curiosity, how many models did you include in your ensemble?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3258989,
          "author_name": "jankowalski2000",
          "author_url": "",
          "post_date": "07/31/2025 16:40:27",
          "content": "<p>There was no way to select the best score correctly. Public and private scores are negatively correlated. I have scores 0.05 - 0.11 and 0.12 - 0.06 on public and private lb. </p>",
          "votes": null,
          "replies": [
            {
              "id": 3258990,
              "author_name": "sarahjeffreson",
              "author_url": "",
              "post_date": "07/31/2025 16:44:05",
              "content": "<p>Similar story for me, I have 0.12 - 0.06 public vs. 0.06 - 0.10 for private, with rough anti-correlation. The worst part is that my CV and the public LB were actually pretty well-correlated too.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3258997,
              "author_name": "byunjins",
              "author_url": "",
              "post_date": "07/31/2025 16:56:45",
              "content": "<p>You are right. I noticed the same thing in my submission. It's interesting that even with ensembling multiple models, there was still anti-correlation.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3259041,
      "author_name": "alexzhongs",
      "author_url": "",
      "post_date": "07/31/2025 18:26:40",
      "content": "<p>It does seem like public vs private score is not super correlated, maybe because this is an anti-overfit problem than an underfit problem.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3259043,
          "author_name": "alexzhongs",
          "author_url": "",
          "post_date": "07/31/2025 18:34:31",
          "content": "<p>Maybe for highly noisy or anti-overfit competition, it would be \"better\" to incorporate IS score to balance luck in OOS and research skill in IS</p>",
          "votes": null,
          "replies": [
            {
              "id": 3259044,
              "author_name": "sarahjeffreson",
              "author_url": "",
              "post_date": "07/31/2025 18:38:09",
              "content": "<p>This would be nice</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3258864": "I know crypto data is notoriously noisy and regime-dependent, but the majority of submissions on the LB have dropped/ascended 100+ places! In comparison to the few other Kaggle competitions whose leaderboards I've seen, this is a lot.\n\nFor those who maintained their positions in the top 50/top 30/top 10, what did you do to stabilize your public/private correlation? I would love to know. I wonder if I may have hurt myself by trying to squeeze signal out of too many features in my ensemble, which lost predictive power between public and private test sets.",
    "3258900": "It sounds like your score drop might be related to the curse of dimensionality especially if too many features were involved in the ensemble. My own approach was fairly simple, and I suspect adding more models might have improved my result. Out of curiosity, how many models did you include in your ensemble?",
    "3258989": "There was no way to select the best score correctly. Public and private scores are negatively correlated. I have scores 0.05 - 0.11 and 0.12 - 0.06 on public and private lb.",
    "3258990": "Similar story for me, I have 0.12 - 0.06 public vs. 0.06 - 0.10 for private, with rough anti-correlation. The worst part is that my CV and the public LB were actually pretty well-correlated too.",
    "3258997": "You are right. I noticed the same thing in my submission. It's interesting that even with ensembling multiple models, there was still anti-correlation.",
    "3259041": "It does seem like public vs private score is not super correlated, maybe because this is an anti-overfit problem than an underfit problem.",
    "3259043": "Maybe for highly noisy or anti-overfit competition, it would be \"better\" to incorporate IS score to balance luck in OOS and research skill in IS",
    "3259044": "This would be nice"
  },
  "source": "meta"
}