{
  "id": 591217,
  "title": "No reliable CV?",
  "url": "/competitions/aeroclub-recsys-2025/discussion/591217",
  "author_name": "",
  "post_date": "2025-07-26T08:24:32.801780Z",
  "votes": 4,
  "comment_count": 22,
  "views": 0,
  "content": "<p>I tried various CV methods (groupkfold, time-wise) but none seems quite as robust in my opinion. only one that somewhat works is groupkfold CV, but I am not 100% confident in that either. </p>\n<p>Has anyone figured out a better way to construct robust validation? </p>",
  "messages": [
    {
      "id": "3254307",
      "postDate": "07/26/2025 08:24:32",
      "content": "<p>I tried various CV methods (groupkfold, time-wise) but none seems quite as robust in my opinion. only one that somewhat works is groupkfold CV, but I am not 100% confident in that either. </p>\n<p>Has anyone figured out a better way to construct robust validation? </p>",
      "rawMarkdown": "I tried various CV methods (groupkfold, time-wise) but none seems quite as robust in my opinion. only one that somewhat works is groupkfold CV, but I am not 100% confident in that either. \n\nHas anyone figured out a better way to construct robust validation?",
      "votes": null
    },
    {
      "id": "3254375",
      "postDate": "07/26/2025 11:45:33",
      "content": "<p>Same here. May local CV could achieve 0.6 without data leakage, but only get 0.51 in test.</p>",
      "rawMarkdown": "Same here. May local CV could achieve 0.6 without data leakage, but only get 0.51 in test.",
      "votes": null
    },
    {
      "id": "3254376",
      "postDate": "07/26/2025 11:46:37",
      "content": "<p>But I haven't tried k-fold or other model ensemble things yet, not sure would it affect.</p>",
      "rawMarkdown": "But I haven't tried k-fold or other model ensemble things yet, not sure would it affect.",
      "votes": null
    },
    {
      "id": "3254489",
      "postDate": "07/26/2025 16:00:07",
      "content": "<p>Thanks. Whats your CV scheme BTW?</p>",
      "rawMarkdown": "Thanks. Whats your CV scheme BTW?",
      "votes": null
    },
    {
      "id": "3254491",
      "postDate": "07/26/2025 16:08:43",
      "content": "<p>not sure. but maybe normal 5-fold. Other than that, It will take too long to train the proper model.</p>",
      "rawMarkdown": "not sure. but maybe normal 5-fold. Other than that, It will take too long to train the proper model.",
      "votes": null
    },
    {
      "id": "3254492",
      "postDate": "07/26/2025 16:09:48",
      "content": "<p>how about you? BTW how did your CV compare to test?</p>",
      "rawMarkdown": "how about you? BTW how did your CV compare to test?",
      "votes": null
    },
    {
      "id": "3254521",
      "postDate": "07/26/2025 17:01:55",
      "content": "<p>Group kfold, cv 0.6 lb 0.52 </p>",
      "rawMarkdown": "Group kfold, cv 0.6 lb 0.52",
      "votes": null
    },
    {
      "id": "3254560",
      "postDate": "07/26/2025 18:36:39",
      "content": "<p>Same here. 0.599 and 0.5 in test. Thank you for sharing, because I spend 2 days searching for leakage. And sure no leak.  I think it is because ticket offer changed after summer has gone. (Just few train test split, not k fold)</p>",
      "rawMarkdown": "Same here. 0.599 and 0.5 in test. Thank you for sharing, because I spend 2 days searching for leakage. And sure no leak.  I think it is because ticket offer changed after summer has gone. (Just few train test split, not k fold)",
      "votes": null
    },
    {
      "id": "3254583",
      "postDate": "07/26/2025 19:54:03",
      "content": "<p>Thanks for answering!</p>",
      "rawMarkdown": "Thanks for answering!",
      "votes": null
    },
    {
      "id": "3254651",
      "postDate": "07/27/2025 00:20:50",
      "content": "<p>do you exclude group with size &lt;= 10 when calculating hitrate@3?</p>",
      "rawMarkdown": "do you exclude group with size <= 10 when calculating hitrate@3?",
      "votes": null
    },
    {
      "id": "3254721",
      "postDate": "07/27/2025 05:21:23",
      "content": "<p>Yes, it's possible that the proportion of larger groups has increased, for example. However, my local validation was showing 0.511–0.514 with lb 0.512 until recently. Then the validation rose to 0.520 while the lb was 0.513.</p>",
      "rawMarkdown": "Yes, it's possible that the proportion of larger groups has increased, for example. However, my local validation was showing 0.511–0.514 with lb 0.512 until recently. Then the validation rose to 0.520 while the lb was 0.513.",
      "votes": null
    },
    {
      "id": "3254774",
      "postDate": "07/27/2025 07:58:22",
      "content": "<p>Yes</p>\n<pre><code> ():\n    \n    \n    val_df = pl.DataFrame({\n        : groups_np,\n        : labels_np,\n        : preds_np\n    })\n\n    \n    val_df = val_df.sort([, ], descending=[, ])\n\n    \n    val_df = val_df.with_columns([\n        pl.col().rank(method=, descending=).over().alias()\n    ])\n\n    \n    group_size_df = (\n        val_df.group_by()\n              .agg(pl.().alias())\n    )\n\n    \n    hit_df = (\n        val_df.(pl.col() == )\n              .join(group_size_df, on=)\n              .with_columns([\n                  (pl.col() &lt;=top).cast(pl.Int8).alias()\n              ])\n    )\n\n    \n    cond = pl.col() &gt; min_group_size\n     max_group_size   :\n        cond = cond &amp; (pl.col() &lt;= max_group_size)\n    hit_filtered = hit_df.(cond)\n\n    \n    num_groups = hit_filtered.height\n    num_hits = hit_filtered[].()\n     num_groups &gt; :\n        hitrate = num_hits / num_groups\n    :\n        hitrate = \n\n    \n    ()\n\n     hitrate\n</code></pre>",
      "rawMarkdown": "Yes\n\n\n\n```python\ndef compute_hitrate_at_3(groups_np, labels_np, preds_np,top =3, min_group_size=10, max_group_size=None):\n    \"\"\"\n    計算 HitRate@3，可自訂 group size 範圍\n    參數:\n        groups_np: np.array of ranker_id\n        labels_np: np.array of true labels (0/1)\n        preds_np: np.array of predicted scores\n        min_group_size: int, 最小 group size (預設=10)\n        max_group_size: int or None, 最大 group size (預設=None，不限)\n    回傳:\n        hitrate: float\n    \"\"\"\n    # 建立 DataFrame\n    val_df = pl.DataFrame({\n        \"ranker_id\": groups_np,\n        \"label\": labels_np,\n        \"score\": preds_np\n    })\n\n    # 排序\n    val_df = val_df.sort([\"ranker_id\", \"score\"], descending=[False, True])\n\n    # rank_in_group: 分數最高 rank=1\n    val_df = val_df.with_columns([\n        pl.col(\"score\").rank(method=\"ordinal\", descending=True).over(\"ranker_id\").alias(\"rank_in_group\")\n    ])\n\n    # 計算每組大小\n    group_size_df = (\n        val_df.group_by(\"ranker_id\")\n              .agg(pl.len().alias(\"group_size\"))\n    )\n\n    # 對每組找 label=1 的 row\n    hit_df = (\n        val_df.filter(pl.col(\"label\") == 1)\n              .join(group_size_df, on=\"ranker_id\")\n              .with_columns([\n                  (pl.col(\"rank_in_group\") <=top).cast(pl.Int8).alias(\"is_hit\")\n              ])\n    )\n\n    # 篩選 group size\n    cond = pl.col(\"group_size\") > min_group_size\n    if max_group_size is not None:\n        cond = cond & (pl.col(\"group_size\") <= max_group_size)\n    hit_filtered = hit_df.filter(cond)\n\n    # 計算\n    num_groups = hit_filtered.height\n    num_hits = hit_filtered[\"is_hit\"].sum()\n    if num_groups > 0:\n        hitrate = num_hits / num_groups\n    else:\n        hitrate = 0.0\n\n    # 輸出\n    print(f\"✅ HitRate@3 (groups size in [{min_group_size}, {max_group_size or 'inf'}]): {hitrate:.4f}\")\n\n    return hitrate\n```",
      "votes": null
    },
    {
      "id": "3254804",
      "postDate": "07/27/2025 09:04:46",
      "content": "<p>Interesting! Anyway, thanks for sharing!</p>",
      "rawMarkdown": "Interesting! Anyway, thanks for sharing!",
      "votes": null
    },
    {
      "id": "3254817",
      "postDate": "07/27/2025 09:39:50",
      "content": "<p>The function looks fine. Did you by any chance use information from the validation set when constructing your features?</p>",
      "rawMarkdown": "The function looks fine. Did you by any chance use information from the validation set when constructing your features?",
      "votes": null
    },
    {
      "id": "3259448",
      "postDate": "08/01/2025 13:29:21",
      "content": "<p>CV | LB<br>\n42.1 | 42.0<br>\n46.2 | 46.1<br>\n56.6 | 48.0<br>\n60.4 | 49.2</p>",
      "rawMarkdown": "CV | LB\n42.1 | 42.0\n46.2 | 46.1\n56.6 | 48.0\n60.4 | 49.2",
      "votes": null
    },
    {
      "id": "3262524",
      "postDate": "08/03/2025 19:23:49",
      "content": "<p>your correlation seems quite nice, what's your CV scheme? </p>",
      "rawMarkdown": "your correlation seems quite nice, what's your CV scheme?",
      "votes": null
    },
    {
      "id": "3262734",
      "postDate": "08/04/2025 07:44:31",
      "content": "<p>Group 10 folds</p>",
      "rawMarkdown": "Group 10 folds",
      "votes": null
    },
    {
      "id": "3262773",
      "postDate": "08/04/2025 08:52:08",
      "content": "<p>I will ensemble 10 models and I am done. 0.51+ is out of my reach. </p>",
      "rawMarkdown": "I will ensemble 10 models and I am done. 0.51+ is out of my reach.",
      "votes": null
    },
    {
      "id": "3262861",
      "postDate": "08/04/2025 11:34:13",
      "content": "<p>woah! 10 gkf on entire data? How long does it take to train? </p>",
      "rawMarkdown": "woah! 10 gkf on entire data? How long does it take to train?",
      "votes": null
    },
    {
      "id": "3262867",
      "postDate": "08/04/2025 11:48:07",
      "content": "<p>10 days if not parallel. Now, I use a sample to train, but validate on full. </p>",
      "rawMarkdown": "10 days if not parallel. Now, I use a sample to train, but validate on full.",
      "votes": null
    },
    {
      "id": "3262875",
      "postDate": "08/04/2025 12:03:50",
      "content": "<p>Do you use JSONs Raw Additional Data? I don't. </p>",
      "rawMarkdown": "Do you use JSONs Raw Additional Data? I don't.",
      "votes": null
    },
    {
      "id": "3262919",
      "postDate": "08/04/2025 13:25:32",
      "content": "<p>Very little json data, i am not sure if too many things can be extracted and successfully merged with the training data. </p>",
      "rawMarkdown": "Very little json data, i am not sure if too many things can be extracted and successfully merged with the training data.",
      "votes": null
    },
    {
      "id": "3266076",
      "postDate": "08/08/2025 11:11:59",
      "content": "<p>You can easily achieve very strong CV scores using hyperparameter optimization (HPO) with higher tree depths and group-based cross-validation (without leakage). However, the fact that such high scores are possible strongly suggests there's inherent leakage, regardless of the CV strategy used.<br>\nThis seems to be the kind of competition where it's safer to stick with a highly regularized, small-sized model to avoid overfitting to the private test set. I expect we'll see a leaderboard reshuffle.<br>\nMy approach is to focus purely on achieving the best CV performance using group-fold CV and a compact model, while completely ignoring the private leaderboard. There’s really no other safe option here.</p>",
      "rawMarkdown": "You can easily achieve very strong CV scores using hyperparameter optimization (HPO) with higher tree depths and group-based cross-validation (without leakage). However, the fact that such high scores are possible strongly suggests there's inherent leakage, regardless of the CV strategy used.\nThis seems to be the kind of competition where it's safer to stick with a highly regularized, small-sized model to avoid overfitting to the private test set. I expect we'll see a leaderboard reshuffle.\nMy approach is to focus purely on achieving the best CV performance using group-fold CV and a compact model, while completely ignoring the private leaderboard. There’s really no other safe option here.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3254375,
      "author_name": "dingyangwang",
      "author_url": "",
      "post_date": "07/26/2025 11:45:33",
      "content": "<p>Same here. May local CV could achieve 0.6 without data leakage, but only get 0.51 in test.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3254376,
          "author_name": "dingyangwang",
          "author_url": "",
          "post_date": "07/26/2025 11:46:37",
          "content": "<p>But I haven't tried k-fold or other model ensemble things yet, not sure would it affect.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3254489,
              "author_name": "pheadrus",
              "author_url": "",
              "post_date": "07/26/2025 16:00:07",
              "content": "<p>Thanks. Whats your CV scheme BTW?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3254491,
                  "author_name": "dingyangwang",
                  "author_url": "",
                  "post_date": "07/26/2025 16:08:43",
                  "content": "<p>not sure. but maybe normal 5-fold. Other than that, It will take too long to train the proper model.</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 3254492,
                  "author_name": "dingyangwang",
                  "author_url": "",
                  "post_date": "07/26/2025 16:09:48",
                  "content": "<p>how about you? BTW how did your CV compare to test?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3254521,
                      "author_name": "pheadrus",
                      "author_url": "",
                      "post_date": "07/26/2025 17:01:55",
                      "content": "<p>Group kfold, cv 0.6 lb 0.52 </p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3254583,
                          "author_name": "dingyangwang",
                          "author_url": "",
                          "post_date": "07/26/2025 19:54:03",
                          "content": "<p>Thanks for answering!</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 3254560,
          "author_name": "sergeyqt2024",
          "author_url": "",
          "post_date": "07/26/2025 18:36:39",
          "content": "<p>Same here. 0.599 and 0.5 in test. Thank you for sharing, because I spend 2 days searching for leakage. And sure no leak.  I think it is because ticket offer changed after summer has gone. (Just few train test split, not k fold)</p>",
          "votes": null,
          "replies": [
            {
              "id": 3254721,
              "author_name": "mikhailgolubchik",
              "author_url": "",
              "post_date": "07/27/2025 05:21:23",
              "content": "<p>Yes, it's possible that the proportion of larger groups has increased, for example. However, my local validation was showing 0.511–0.514 with lb 0.512 until recently. Then the validation rose to 0.520 while the lb was 0.513.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3254804,
                  "author_name": "dingyangwang",
                  "author_url": "",
                  "post_date": "07/27/2025 09:04:46",
                  "content": "<p>Interesting! Anyway, thanks for sharing!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 3254651,
          "author_name": "mango789",
          "author_url": "",
          "post_date": "07/27/2025 00:20:50",
          "content": "<p>do you exclude group with size &lt;= 10 when calculating hitrate@3?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3254774,
              "author_name": "dingyangwang",
              "author_url": "",
              "post_date": "07/27/2025 07:58:22",
              "content": "<p>Yes</p>\n<pre><code> ():\n    \n    \n    val_df = pl.DataFrame({\n        : groups_np,\n        : labels_np,\n        : preds_np\n    })\n\n    \n    val_df = val_df.sort([, ], descending=[, ])\n\n    \n    val_df = val_df.with_columns([\n        pl.col().rank(method=, descending=).over().alias()\n    ])\n\n    \n    group_size_df = (\n        val_df.group_by()\n              .agg(pl.().alias())\n    )\n\n    \n    hit_df = (\n        val_df.(pl.col() == )\n              .join(group_size_df, on=)\n              .with_columns([\n                  (pl.col() &lt;=top).cast(pl.Int8).alias()\n              ])\n    )\n\n    \n    cond = pl.col() &gt; min_group_size\n     max_group_size   :\n        cond = cond &amp; (pl.col() &lt;= max_group_size)\n    hit_filtered = hit_df.(cond)\n\n    \n    num_groups = hit_filtered.height\n    num_hits = hit_filtered[].()\n     num_groups &gt; :\n        hitrate = num_hits / num_groups\n    :\n        hitrate = \n\n    \n    ()\n\n     hitrate\n</code></pre>",
              "votes": null,
              "replies": [
                {
                  "id": 3254817,
                  "author_name": "mango789",
                  "author_url": "",
                  "post_date": "07/27/2025 09:39:50",
                  "content": "<p>The function looks fine. Did you by any chance use information from the validation set when constructing your features?</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3259448,
      "author_name": "jankowalski2000",
      "author_url": "",
      "post_date": "08/01/2025 13:29:21",
      "content": "<p>CV | LB<br>\n42.1 | 42.0<br>\n46.2 | 46.1<br>\n56.6 | 48.0<br>\n60.4 | 49.2</p>",
      "votes": null,
      "replies": [
        {
          "id": 3262524,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "08/03/2025 19:23:49",
          "content": "<p>your correlation seems quite nice, what's your CV scheme? </p>",
          "votes": null,
          "replies": [
            {
              "id": 3262734,
              "author_name": "jankowalski2000",
              "author_url": "",
              "post_date": "08/04/2025 07:44:31",
              "content": "<p>Group 10 folds</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3262773,
                  "author_name": "jankowalski2000",
                  "author_url": "",
                  "post_date": "08/04/2025 08:52:08",
                  "content": "<p>I will ensemble 10 models and I am done. 0.51+ is out of my reach. </p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 3262861,
                  "author_name": "pheadrus",
                  "author_url": "",
                  "post_date": "08/04/2025 11:34:13",
                  "content": "<p>woah! 10 gkf on entire data? How long does it take to train? </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3262867,
                      "author_name": "jankowalski2000",
                      "author_url": "",
                      "post_date": "08/04/2025 11:48:07",
                      "content": "<p>10 days if not parallel. Now, I use a sample to train, but validate on full. </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            },
            {
              "id": 3262875,
              "author_name": "jankowalski2000",
              "author_url": "",
              "post_date": "08/04/2025 12:03:50",
              "content": "<p>Do you use JSONs Raw Additional Data? I don't. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3262919,
                  "author_name": "pheadrus",
                  "author_url": "",
                  "post_date": "08/04/2025 13:25:32",
                  "content": "<p>Very little json data, i am not sure if too many things can be extracted and successfully merged with the training data. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3266076,
      "author_name": "jamalsaeedi",
      "author_url": "",
      "post_date": "08/08/2025 11:11:59",
      "content": "<p>You can easily achieve very strong CV scores using hyperparameter optimization (HPO) with higher tree depths and group-based cross-validation (without leakage). However, the fact that such high scores are possible strongly suggests there's inherent leakage, regardless of the CV strategy used.<br>\nThis seems to be the kind of competition where it's safer to stick with a highly regularized, small-sized model to avoid overfitting to the private test set. I expect we'll see a leaderboard reshuffle.<br>\nMy approach is to focus purely on achieving the best CV performance using group-fold CV and a compact model, while completely ignoring the private leaderboard. There’s really no other safe option here.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3254307": "I tried various CV methods (groupkfold, time-wise) but none seems quite as robust in my opinion. only one that somewhat works is groupkfold CV, but I am not 100% confident in that either. \n\nHas anyone figured out a better way to construct robust validation?",
    "3254375": "Same here. May local CV could achieve 0.6 without data leakage, but only get 0.51 in test.",
    "3254376": "But I haven't tried k-fold or other model ensemble things yet, not sure would it affect.",
    "3254489": "Thanks. Whats your CV scheme BTW?",
    "3254491": "not sure. but maybe normal 5-fold. Other than that, It will take too long to train the proper model.",
    "3254492": "how about you? BTW how did your CV compare to test?",
    "3254521": "Group kfold, cv 0.6 lb 0.52",
    "3254560": "Same here. 0.599 and 0.5 in test. Thank you for sharing, because I spend 2 days searching for leakage. And sure no leak.  I think it is because ticket offer changed after summer has gone. (Just few train test split, not k fold)",
    "3254583": "Thanks for answering!",
    "3254651": "do you exclude group with size <= 10 when calculating hitrate@3?",
    "3254721": "Yes, it's possible that the proportion of larger groups has increased, for example. However, my local validation was showing 0.511–0.514 with lb 0.512 until recently. Then the validation rose to 0.520 while the lb was 0.513.",
    "3254774": "Yes\n\n\n\n```python\ndef compute_hitrate_at_3(groups_np, labels_np, preds_np,top =3, min_group_size=10, max_group_size=None):\n    \"\"\"\n    計算 HitRate@3，可自訂 group size 範圍\n    參數:\n        groups_np: np.array of ranker_id\n        labels_np: np.array of true labels (0/1)\n        preds_np: np.array of predicted scores\n        min_group_size: int, 最小 group size (預設=10)\n        max_group_size: int or None, 最大 group size (預設=None，不限)\n    回傳:\n        hitrate: float\n    \"\"\"\n    # 建立 DataFrame\n    val_df = pl.DataFrame({\n        \"ranker_id\": groups_np,\n        \"label\": labels_np,\n        \"score\": preds_np\n    })\n\n    # 排序\n    val_df = val_df.sort([\"ranker_id\", \"score\"], descending=[False, True])\n\n    # rank_in_group: 分數最高 rank=1\n    val_df = val_df.with_columns([\n        pl.col(\"score\").rank(method=\"ordinal\", descending=True).over(\"ranker_id\").alias(\"rank_in_group\")\n    ])\n\n    # 計算每組大小\n    group_size_df = (\n        val_df.group_by(\"ranker_id\")\n              .agg(pl.len().alias(\"group_size\"))\n    )\n\n    # 對每組找 label=1 的 row\n    hit_df = (\n        val_df.filter(pl.col(\"label\") == 1)\n              .join(group_size_df, on=\"ranker_id\")\n              .with_columns([\n                  (pl.col(\"rank_in_group\") <=top).cast(pl.Int8).alias(\"is_hit\")\n              ])\n    )\n\n    # 篩選 group size\n    cond = pl.col(\"group_size\") > min_group_size\n    if max_group_size is not None:\n        cond = cond & (pl.col(\"group_size\") <= max_group_size)\n    hit_filtered = hit_df.filter(cond)\n\n    # 計算\n    num_groups = hit_filtered.height\n    num_hits = hit_filtered[\"is_hit\"].sum()\n    if num_groups > 0:\n        hitrate = num_hits / num_groups\n    else:\n        hitrate = 0.0\n\n    # 輸出\n    print(f\"✅ HitRate@3 (groups size in [{min_group_size}, {max_group_size or 'inf'}]): {hitrate:.4f}\")\n\n    return hitrate\n```",
    "3254804": "Interesting! Anyway, thanks for sharing!",
    "3254817": "The function looks fine. Did you by any chance use information from the validation set when constructing your features?",
    "3259448": "CV | LB\n42.1 | 42.0\n46.2 | 46.1\n56.6 | 48.0\n60.4 | 49.2",
    "3262524": "your correlation seems quite nice, what's your CV scheme?",
    "3262734": "Group 10 folds",
    "3262773": "I will ensemble 10 models and I am done. 0.51+ is out of my reach.",
    "3262861": "woah! 10 gkf on entire data? How long does it take to train?",
    "3262867": "10 days if not parallel. Now, I use a sample to train, but validate on full.",
    "3262875": "Do you use JSONs Raw Additional Data? I don't.",
    "3262919": "Very little json data, i am not sure if too many things can be extracted and successfully merged with the training data.",
    "3266076": "You can easily achieve very strong CV scores using hyperparameter optimization (HPO) with higher tree depths and group-based cross-validation (without leakage). However, the fact that such high scores are possible strongly suggests there's inherent leakage, regardless of the CV strategy used.\nThis seems to be the kind of competition where it's safer to stick with a highly regularized, small-sized model to avoid overfitting to the private test set. I expect we'll see a leaderboard reshuffle.\nMy approach is to focus purely on achieving the best CV performance using group-fold CV and a compact model, while completely ignoring the private leaderboard. There’s really no other safe option here."
  },
  "source": "meta"
}