{
  "id": 599559,
  "title": "What ensemble methods worked best for you?",
  "url": "/competitions/aeroclub-recsys-2025/discussion/599559",
  "author_name": "",
  "post_date": "2025-08-17T11:22:04.391633800Z",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>Now that the competition has finished, I’m curious about how much ensemble helped for you.</p>\n<p>In our case, we tried several methods (Spearman correlation, RRF, and a few others), but the LB improvement was very small (around +0.001 ~ +0.002).</p>\n<p>From looking at the final leaderboard, I noticed that some teams must have gained much more from ensembling, boosts of +0.004 ~ +0.005, or even higher.</p>\n<p>What ensemble strategies worked best for you here? Weighted averaging, rank averaging, stacking/blending, or something else?</p>\n<p>And when you built your ensemble, how did you decide which models to include and which to leave out?</p>",
  "messages": [
    {
      "id": "3270801",
      "postDate": "08/17/2025 11:22:04",
      "content": "<p>Hi everyone,</p>\n<p>Now that the competition has finished, I’m curious about how much ensemble helped for you.</p>\n<p>In our case, we tried several methods (Spearman correlation, RRF, and a few others), but the LB improvement was very small (around +0.001 ~ +0.002).</p>\n<p>From looking at the final leaderboard, I noticed that some teams must have gained much more from ensembling, boosts of +0.004 ~ +0.005, or even higher.</p>\n<p>What ensemble strategies worked best for you here? Weighted averaging, rank averaging, stacking/blending, or something else?</p>\n<p>And when you built your ensemble, how did you decide which models to include and which to leave out?</p>",
      "rawMarkdown": "Hi everyone,\n\nNow that the competition has finished, I’m curious about how much ensemble helped for you.\n\nIn our case, we tried several methods (Spearman correlation, RRF, and a few others), but the LB improvement was very small (around +0.001 ~ +0.002).\n\nFrom looking at the final leaderboard, I noticed that some teams must have gained much more from ensembling, boosts of +0.004 ~ +0.005, or even higher.\n\nWhat ensemble strategies worked best for you here? Weighted averaging, rank averaging, stacking/blending, or something else?\n\nAnd when you built your ensemble, how did you decide which models to include and which to leave out?",
      "votes": null
    },
    {
      "id": "3270829",
      "postDate": "08/17/2025 12:38:51",
      "content": "<p>I used optuna to optimize 30+ 40+ models it can improve valid score but overfit local valid, after game finish by late submission I found only 5 top score model ensemble with equal weights could achive best.  PB from 5446 to 54725.  Actually before last submission gpt5 suggested me to choose 5 top model ensemble version. below is what gpt5 said:<br>\nSpecial Considerations for Hit@3<br>\nFor highly correlated models, equal-weight ensembling is generally fine in Hit@3, as long as their error distributions aren’t 100% identical — even if the performance gain is small, it can still reduce the risk of a correct sample being unexpectedly pushed out of the top 3.</p>\n<p>Hit@3 is more sensitive to rank fluctuations than to fine-grained score differences. Optuna’s weight tuning essentially chases maximum scores rather than preserving a stable top‑3 structure — that’s why your leaderboard score drops online.</p>\n<p>🛠 Decision Recommendations</p>\n<table>\n<thead>\n<tr>\n<th>Plan</th>\n<th>Offline Mean Score</th>\n<th>Public LB Performance Expectation</th>\n<th>Private Stability</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>5-model high-score equal weight</td>\n<td>Slightly lower</td>\n<td>Likely to match or slightly exceed best single model</td>\n<td>High</td>\n</tr>\n<tr>\n<td>11-model Optuna</td>\n<td>Slightly higher</td>\n<td>Public LB signal lower than single model</td>\n<td>Medium–low</td>\n</tr>\n<tr>\n<td>37-model Optuna</td>\n<td>Highest</td>\n<td>Lowest on Public LB</td>\n<td>Low</td>\n</tr>\n<tr>\n<td>Given your current evidence, that 0.002 offline gap is not worth sacrificing stability and generalization, especially when the Private leaderboard holds 70% of the weight.</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<p>💡 Final Submission Approach<br>\nFirst choice: 5-model high-score equal-weight ensemble (In Hit@3, losing a few thousandths in score is fine — the real risk is rank displacement).</p>\n<p>Extra insurance: Blend the “best single model” and the 5‑model equal-weight ensemble at 0.5 / 0.5 to hedge against distribution shift.</p>\n<p>Avoid using Optuna for weight tuning at this stage — your online history shows it introduces noise rather than value.</p>",
      "rawMarkdown": "I used optuna to optimize 30+ 40+ models it can improve valid score but overfit local valid, after game finish by late submission I found only 5 top score model ensemble with equal weights could achive best.  PB from 5446 to 54725.  Actually before last submission gpt5 suggested me to choose 5 top model ensemble version. below is what gpt5 said:\nSpecial Considerations for Hit@3\nFor highly correlated models, equal-weight ensembling is generally fine in Hit@3, as long as their error distributions aren’t 100% identical — even if the performance gain is small, it can still reduce the risk of a correct sample being unexpectedly pushed out of the top 3.\n\nHit@3 is more sensitive to rank fluctuations than to fine-grained score differences. Optuna’s weight tuning essentially chases maximum scores rather than preserving a stable top‑3 structure — that’s why your leaderboard score drops online.\n\n🛠 Decision Recommendations\nPlan\t| Offline Mean Score\t| Public LB Performance Expectation |\tPrivate Stability\n-- | -- | --| -- |\n5-model high-score equal weight |\tSlightly lower |\tLikely to match or slightly exceed best single model\t| High\n11-model Optuna\t| Slightly higher\t| Public LB signal lower than single model\t| Medium–low\n37-model Optuna |\tHighest |\tLowest on Public LB |\tLow\nGiven your current evidence, that 0.002 offline gap is not worth sacrificing stability and generalization, especially when the Private leaderboard holds 70% of the weight.\n\n💡 Final Submission Approach\nFirst choice: 5-model high-score equal-weight ensemble (In Hit@3, losing a few thousandths in score is fine — the real risk is rank displacement).\n\nExtra insurance: Blend the “best single model” and the 5‑model equal-weight ensemble at 0.5 / 0.5 to hedge against distribution shift.\n\nAvoid using Optuna for weight tuning at this stage — your online history shows it introduces noise rather than value.",
      "votes": null
    },
    {
      "id": "3270838",
      "postDate": "08/17/2025 12:57:19",
      "content": "<p>Thanks for sharing! I also just tried averaging the top single models, but the boost was still only around 0.001–0.002. This is probably because the models use very similar features…<br>\nbtw, were the models you used for the ensemble also very similar, or did they use different features?</p>",
      "rawMarkdown": "Thanks for sharing! I also just tried averaging the top single models, but the boost was still only around 0.001–0.002. This is probably because the models use very similar features...\nbtw, were the models you used for the ensemble also very similar, or did they use different features?",
      "votes": null
    },
    {
      "id": "3270848",
      "postDate": "08/17/2025 13:30:45",
      "content": "<p>For top 5 models. <br>\nYes they are using same features but different loss_funcation like rank and classification. <br>\n  \"v27.obj-ndcg.tm-xgb\", #0.504745<br>\n  \"v27.obj-pairwise.tm-xgb\", #0.503980<br>\n  \"v27.obj-map.tm-xgb\", #0.503621<br>\n  \"v27.tm-lgb\", #0.500270<br>\n  \"v27.task-cls.tm-xgb\", #0.498966<br>\nLB is similar for </p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>PB</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>best single</td>\n<td>54156</td>\n<td>54062</td>\n</tr>\n<tr>\n<td>best 5 ensemble</td>\n<td>54725</td>\n<td>54053</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "For top 5 models. \nYes they are using same features but different loss_funcation like rank and classification. \n  \"v27.obj-ndcg.tm-xgb\", #0.504745\n  \"v27.obj-pairwise.tm-xgb\", #0.503980\n  \"v27.obj-map.tm-xgb\", #0.503621\n  \"v27.tm-lgb\", #0.500270\n  \"v27.task-cls.tm-xgb\", #0.498966\nLB is similar for \n\n|model | PB | LB | \n| --- | --- | --- |\n| best single |  54156 | 54062 |\n| best 5 ensemble | 54725 | 54053 |",
      "votes": null
    },
    {
      "id": "3270849",
      "postDate": "08/17/2025 13:39:41",
      "content": "<p>I optimized weights across buckets of group sizes. My hypothesis was that different models (10 in total) had different strengths based on group sizes since they were trained with different values of num_pairs parameter in xgb and lgbm. I saw a reasonable boost i think with ensembling 5 models or so, but the perf plateaued after that. Though my CV continued to increase a little. My final sub with 10 models is what had the best CV and the best pvt LB. </p>",
      "rawMarkdown": "I optimized weights across buckets of group sizes. My hypothesis was that different models (10 in total) had different strengths based on group sizes since they were trained with different values of num_pairs parameter in xgb and lgbm. I saw a reasonable boost i think with ensembling 5 models or so, but the perf plateaued after that. Though my CV continued to increase a little. My final sub with 10 models is what had the best CV and the best pvt LB.",
      "votes": null
    },
    {
      "id": "3270857",
      "postDate": "08/17/2025 13:48:34",
      "content": "<p>Thanks a lot! I learned something really useful — how ensembling models trained with the same features but different loss functions can help. Really appreciate you sharing your results!</p>",
      "rawMarkdown": "Thanks a lot! I learned something really useful — how ensembling models trained with the same features but different loss functions can help. Really appreciate you sharing your results!",
      "votes": null
    },
    {
      "id": "3270861",
      "postDate": "08/17/2025 13:52:35",
      "content": "<p>Thanks for sharing! I had thought about training models focused on different group sizes, but I hadn’t considered ensembling them this way.😲</p>",
      "rawMarkdown": "Thanks for sharing! I had thought about training models focused on different group sizes, but I hadn’t considered ensembling them this way.😲",
      "votes": null
    },
    {
      "id": "3270864",
      "postDate": "08/17/2025 13:56:19",
      "content": "<p>and btw, this is actually the first time I’ve learned about the num_pairs parameter🥺</p>",
      "rawMarkdown": "and btw, this is actually the first time I’ve learned about the num_pairs parameter🥺",
      "votes": null
    },
    {
      "id": "3271170",
      "postDate": "08/18/2025 08:15:04",
      "content": "<p>I took 5-7 models (all Light GBM),  use min-max scale for their prediction and took mean of this over models as final.    Checked median, rank and some more and this way was the best</p>",
      "rawMarkdown": "I took 5-7 models (all Light GBM),  use min-max scale for their prediction and took mean of this over models as final.    Checked median, rank and some more and this way was the best",
      "votes": null
    },
    {
      "id": "3271188",
      "postDate": "08/18/2025 09:10:56",
      "content": "<p>Thanks. btw, my lightgbm model doesn't perform well, around 0.510~0.514 in public LB…</p>",
      "rawMarkdown": "Thanks. btw, my lightgbm model doesn't perform well, around 0.510~0.514 in public LB...",
      "votes": null
    },
    {
      "id": "3271192",
      "postDate": "08/18/2025 09:32:39",
      "content": "<p>I tried Catboost, LGBM and some simple handmade kind of Bert2rec.   Bert2rec is very bad (it is no surprise, but I thoght  may be it can produce some features) , CB is not good here and I drop it  and LGBM bring me 10th place.  Will consider   Xboost for future :).   </p>",
      "rawMarkdown": "I tried Catboost, LGBM and some simple handmade kind of Bert2rec.   Bert2rec is very bad (it is no surprise, but I thoght  may be it can produce some features) , CB is not good here and I drop it  and LGBM bring me 10th place.  Will consider   Xboost for future :).",
      "votes": null
    },
    {
      "id": "3271205",
      "postDate": "08/18/2025 10:09:46",
      "content": "<p>yes, you can try it. I think xgb is more stable and won't encounter bugs when training in gpu😂</p>",
      "rawMarkdown": "yes, you can try it. I think xgb is more stable and won't encounter bugs when training in gpu😂",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3270829,
      "author_name": "goldenlock",
      "author_url": "",
      "post_date": "08/17/2025 12:38:51",
      "content": "<p>I used optuna to optimize 30+ 40+ models it can improve valid score but overfit local valid, after game finish by late submission I found only 5 top score model ensemble with equal weights could achive best.  PB from 5446 to 54725.  Actually before last submission gpt5 suggested me to choose 5 top model ensemble version. below is what gpt5 said:<br>\nSpecial Considerations for Hit@3<br>\nFor highly correlated models, equal-weight ensembling is generally fine in Hit@3, as long as their error distributions aren’t 100% identical — even if the performance gain is small, it can still reduce the risk of a correct sample being unexpectedly pushed out of the top 3.</p>\n<p>Hit@3 is more sensitive to rank fluctuations than to fine-grained score differences. Optuna’s weight tuning essentially chases maximum scores rather than preserving a stable top‑3 structure — that’s why your leaderboard score drops online.</p>\n<p>🛠 Decision Recommendations</p>\n<table>\n<thead>\n<tr>\n<th>Plan</th>\n<th>Offline Mean Score</th>\n<th>Public LB Performance Expectation</th>\n<th>Private Stability</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>5-model high-score equal weight</td>\n<td>Slightly lower</td>\n<td>Likely to match or slightly exceed best single model</td>\n<td>High</td>\n</tr>\n<tr>\n<td>11-model Optuna</td>\n<td>Slightly higher</td>\n<td>Public LB signal lower than single model</td>\n<td>Medium–low</td>\n</tr>\n<tr>\n<td>37-model Optuna</td>\n<td>Highest</td>\n<td>Lowest on Public LB</td>\n<td>Low</td>\n</tr>\n<tr>\n<td>Given your current evidence, that 0.002 offline gap is not worth sacrificing stability and generalization, especially when the Private leaderboard holds 70% of the weight.</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n</tbody>\n</table>\n<p>💡 Final Submission Approach<br>\nFirst choice: 5-model high-score equal-weight ensemble (In Hit@3, losing a few thousandths in score is fine — the real risk is rank displacement).</p>\n<p>Extra insurance: Blend the “best single model” and the 5‑model equal-weight ensemble at 0.5 / 0.5 to hedge against distribution shift.</p>\n<p>Avoid using Optuna for weight tuning at this stage — your online history shows it introduces noise rather than value.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3270838,
          "author_name": "mango789",
          "author_url": "",
          "post_date": "08/17/2025 12:57:19",
          "content": "<p>Thanks for sharing! I also just tried averaging the top single models, but the boost was still only around 0.001–0.002. This is probably because the models use very similar features…<br>\nbtw, were the models you used for the ensemble also very similar, or did they use different features?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3270848,
              "author_name": "goldenlock",
              "author_url": "",
              "post_date": "08/17/2025 13:30:45",
              "content": "<p>For top 5 models. <br>\nYes they are using same features but different loss_funcation like rank and classification. <br>\n  \"v27.obj-ndcg.tm-xgb\", #0.504745<br>\n  \"v27.obj-pairwise.tm-xgb\", #0.503980<br>\n  \"v27.obj-map.tm-xgb\", #0.503621<br>\n  \"v27.tm-lgb\", #0.500270<br>\n  \"v27.task-cls.tm-xgb\", #0.498966<br>\nLB is similar for </p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>PB</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>best single</td>\n<td>54156</td>\n<td>54062</td>\n</tr>\n<tr>\n<td>best 5 ensemble</td>\n<td>54725</td>\n<td>54053</td>\n</tr>\n</tbody>\n</table>",
              "votes": null,
              "replies": [
                {
                  "id": 3270857,
                  "author_name": "mango789",
                  "author_url": "",
                  "post_date": "08/17/2025 13:48:34",
                  "content": "<p>Thanks a lot! I learned something really useful — how ensembling models trained with the same features but different loss functions can help. Really appreciate you sharing your results!</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3270849,
      "author_name": "pheadrus",
      "author_url": "",
      "post_date": "08/17/2025 13:39:41",
      "content": "<p>I optimized weights across buckets of group sizes. My hypothesis was that different models (10 in total) had different strengths based on group sizes since they were trained with different values of num_pairs parameter in xgb and lgbm. I saw a reasonable boost i think with ensembling 5 models or so, but the perf plateaued after that. Though my CV continued to increase a little. My final sub with 10 models is what had the best CV and the best pvt LB. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3270861,
          "author_name": "mango789",
          "author_url": "",
          "post_date": "08/17/2025 13:52:35",
          "content": "<p>Thanks for sharing! I had thought about training models focused on different group sizes, but I hadn’t considered ensembling them this way.😲</p>",
          "votes": null,
          "replies": [
            {
              "id": 3270864,
              "author_name": "mango789",
              "author_url": "",
              "post_date": "08/17/2025 13:56:19",
              "content": "<p>and btw, this is actually the first time I’ve learned about the num_pairs parameter🥺</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3271170,
      "author_name": "sergeyqt2024",
      "author_url": "",
      "post_date": "08/18/2025 08:15:04",
      "content": "<p>I took 5-7 models (all Light GBM),  use min-max scale for their prediction and took mean of this over models as final.    Checked median, rank and some more and this way was the best</p>",
      "votes": null,
      "replies": [
        {
          "id": 3271188,
          "author_name": "mango789",
          "author_url": "",
          "post_date": "08/18/2025 09:10:56",
          "content": "<p>Thanks. btw, my lightgbm model doesn't perform well, around 0.510~0.514 in public LB…</p>",
          "votes": null,
          "replies": [
            {
              "id": 3271192,
              "author_name": "sergeyqt2024",
              "author_url": "",
              "post_date": "08/18/2025 09:32:39",
              "content": "<p>I tried Catboost, LGBM and some simple handmade kind of Bert2rec.   Bert2rec is very bad (it is no surprise, but I thoght  may be it can produce some features) , CB is not good here and I drop it  and LGBM bring me 10th place.  Will consider   Xboost for future :).   </p>",
              "votes": null,
              "replies": [
                {
                  "id": 3271205,
                  "author_name": "mango789",
                  "author_url": "",
                  "post_date": "08/18/2025 10:09:46",
                  "content": "<p>yes, you can try it. I think xgb is more stable and won't encounter bugs when training in gpu😂</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3270801": "Hi everyone,\n\nNow that the competition has finished, I’m curious about how much ensemble helped for you.\n\nIn our case, we tried several methods (Spearman correlation, RRF, and a few others), but the LB improvement was very small (around +0.001 ~ +0.002).\n\nFrom looking at the final leaderboard, I noticed that some teams must have gained much more from ensembling, boosts of +0.004 ~ +0.005, or even higher.\n\nWhat ensemble strategies worked best for you here? Weighted averaging, rank averaging, stacking/blending, or something else?\n\nAnd when you built your ensemble, how did you decide which models to include and which to leave out?",
    "3270829": "I used optuna to optimize 30+ 40+ models it can improve valid score but overfit local valid, after game finish by late submission I found only 5 top score model ensemble with equal weights could achive best.  PB from 5446 to 54725.  Actually before last submission gpt5 suggested me to choose 5 top model ensemble version. below is what gpt5 said:\nSpecial Considerations for Hit@3\nFor highly correlated models, equal-weight ensembling is generally fine in Hit@3, as long as their error distributions aren’t 100% identical — even if the performance gain is small, it can still reduce the risk of a correct sample being unexpectedly pushed out of the top 3.\n\nHit@3 is more sensitive to rank fluctuations than to fine-grained score differences. Optuna’s weight tuning essentially chases maximum scores rather than preserving a stable top‑3 structure — that’s why your leaderboard score drops online.\n\n🛠 Decision Recommendations\nPlan\t| Offline Mean Score\t| Public LB Performance Expectation |\tPrivate Stability\n-- | -- | --| -- |\n5-model high-score equal weight |\tSlightly lower |\tLikely to match or slightly exceed best single model\t| High\n11-model Optuna\t| Slightly higher\t| Public LB signal lower than single model\t| Medium–low\n37-model Optuna |\tHighest |\tLowest on Public LB |\tLow\nGiven your current evidence, that 0.002 offline gap is not worth sacrificing stability and generalization, especially when the Private leaderboard holds 70% of the weight.\n\n💡 Final Submission Approach\nFirst choice: 5-model high-score equal-weight ensemble (In Hit@3, losing a few thousandths in score is fine — the real risk is rank displacement).\n\nExtra insurance: Blend the “best single model” and the 5‑model equal-weight ensemble at 0.5 / 0.5 to hedge against distribution shift.\n\nAvoid using Optuna for weight tuning at this stage — your online history shows it introduces noise rather than value.",
    "3270838": "Thanks for sharing! I also just tried averaging the top single models, but the boost was still only around 0.001–0.002. This is probably because the models use very similar features...\nbtw, were the models you used for the ensemble also very similar, or did they use different features?",
    "3270848": "For top 5 models. \nYes they are using same features but different loss_funcation like rank and classification. \n  \"v27.obj-ndcg.tm-xgb\", #0.504745\n  \"v27.obj-pairwise.tm-xgb\", #0.503980\n  \"v27.obj-map.tm-xgb\", #0.503621\n  \"v27.tm-lgb\", #0.500270\n  \"v27.task-cls.tm-xgb\", #0.498966\nLB is similar for \n\n|model | PB | LB | \n| --- | --- | --- |\n| best single |  54156 | 54062 |\n| best 5 ensemble | 54725 | 54053 |",
    "3270849": "I optimized weights across buckets of group sizes. My hypothesis was that different models (10 in total) had different strengths based on group sizes since they were trained with different values of num_pairs parameter in xgb and lgbm. I saw a reasonable boost i think with ensembling 5 models or so, but the perf plateaued after that. Though my CV continued to increase a little. My final sub with 10 models is what had the best CV and the best pvt LB.",
    "3270857": "Thanks a lot! I learned something really useful — how ensembling models trained with the same features but different loss functions can help. Really appreciate you sharing your results!",
    "3270861": "Thanks for sharing! I had thought about training models focused on different group sizes, but I hadn’t considered ensembling them this way.😲",
    "3270864": "and btw, this is actually the first time I’ve learned about the num_pairs parameter🥺",
    "3271170": "I took 5-7 models (all Light GBM),  use min-max scale for their prediction and took mean of this over models as final.    Checked median, rank and some more and this way was the best",
    "3271188": "Thanks. btw, my lightgbm model doesn't perform well, around 0.510~0.514 in public LB...",
    "3271192": "I tried Catboost, LGBM and some simple handmade kind of Bert2rec.   Bert2rec is very bad (it is no surprise, but I thoght  may be it can produce some features) , CB is not good here and I drop it  and LGBM bring me 10th place.  Will consider   Xboost for future :).",
    "3271205": "yes, you can try it. I think xgb is more stable and won't encounter bugs when training in gpu😂"
  },
  "source": "meta"
}