{
  "id": 599666,
  "title": "1st solution using count-based features (Single model PB 0.54942)",
  "url": "/competitions/aeroclub-recsys-2025/writeups/1st-solution-using-count-based-features",
  "author_name": "",
  "post_date": "2025-08-18T03:45:59.820Z",
  "votes": 25,
  "comment_count": 10,
  "views": 0,
  "content": "<hr>\n<p>First, I want to express my sincere gratitude to the organizers for providing such a large-scale, real-world dataset. It offered a rare opportunity to dive deep into practical challenges and learn from them.</p>\n<p>I also want to thank the authors of several outstanding notebooks that guided my journey:</p>\n<ul>\n<li>I began with <a href=\"https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars\" target=\"_blank\">XGBoost Ranker with Polars</a> a clean and efficient starting point that helped me grasp the essentials of ranking models. Thank you <a href=\"https://www.kaggle.com/ka1242\" target=\"_blank\">@ka1242</a>!</li>\n<li>The most creative idea, in my opinion, came from <a href=\"https://www.kaggle.com/code/mango789/xgboost-ranker-rule-based-rerank\" target=\"_blank\">reranking approach</a> . The rule-based reranking worked like a charm, and without that notebook, I wouldn’t have realized how crucial <code>flight_hash</code> is to the final performance. Huge thanks to <a href=\"https://www.kaggle.com/mango789\" target=\"_blank\">@mango789</a> !</li>\n</ul>\n<hr>\n<h2>Approach overview and note on test data usage</h2>\n<p>My main single-model solution initially leveraged a leak-like feature: I included the test data from the very beginning of the competition. Count-based categorical encoding built on the full dataset noticeably improved both local validation and the public leaderboard.</p>\n<p>I recall a discussion where the host mentioned that using the test set was acceptable because it contains no labels. While some recommendation competitions forbid using the test data, I didn’t find such a restriction here. Still, in real scenarios we cannot rely on future data, and we usually predict day by day using only past labels. A streaming setup would better reflect production.</p>\n<p>Below I share results for the same single model trained without the test data (i.e., I removed the test set from the count-based encodings and training input).</p>\n<hr>\n<h2>Results without using test data</h2>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>valid</th>\n<th>LB</th>\n<th>PB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>best single model</td>\n<td>0.504745</td>\n<td>0.54062</td>\n<td>0.54156</td>\n</tr>\n<tr>\n<td>best single model without testdata</td>\n<td>0.478704</td>\n<td>0.51978</td>\n<td>0.52692</td>\n</tr>\n</tbody>\n</table>\n<p>This is an ongoing write-up; I may refine it or add more data stats later.</p>\n<hr>\n<h2>Key components</h2>\n<h3>Validation strategy</h3>\n<ul>\n<li><strong>Temporal split:</strong> This is a recommendation task, so datetime order matters. The dataset spans <code>requestDate</code> from 2024-05-17 03:03:08 to 2024-12-31 18:54:00, which I index as days 1–229.</li>\n<li><strong>Fold design:</strong> I used a single temporal fold: train on days 1–103, validate on days 104–166, and test on days 167–229.</li>\n<li><strong>Rationale:</strong> Validation and test windows have similar lengths, and local validation generally matched the LB, especially for larger improvements. When there was a mismatch, I trusted local validation. One exception was the final Optuna weight tuning for ensembles, which can overfit to the validation set.</li>\n</ul>\n<h3>Count-based categorical encoding</h3>\n<ul>\n<li><strong>XGBoost / LightGBM:</strong> I applied count-based encoding to all categorical features, then treated all columns as numerical — no categorical indices were passed into training.</li>\n<li><strong>CatBoost:</strong> I also experimented with exactly the same count-based encoding here, converting everything to numerical features. However, in my tests, all variations yielded very similar results, with no meaningful gain over the baseline.</li>\n</ul>\n<h3>Time features</h3>\n<ul>\n<li><strong>Importance:</strong> Time-derived features gave the largest lift compared to using only the provided CSV fields.</li>\n<li><strong>Cyclical encoding:</strong> I added sine/cosine transforms so the model can capture wrap-around (e.g., hours 23 and 0 are adjacent).</li>\n</ul>\n<h3>Duration and distance features</h3>\n<ul>\n<li><strong>Accurate duration:</strong> I converted local times to UTC to compute true elapsed durations.</li>\n<li><strong>Distance-based signals:</strong> I computed airport distances from lat/lon and derived related features such as <code>direct_price_per_km</code>.</li>\n</ul>\n<h3>Simple rank-based features on important base features</h3>\n<ul>\n<li><strong>Motivation:</strong> Ranking proxies (per group) for price, duration, and itinerary complexity helped the ranker focus on relative ordering within a user/query.</li>\n</ul>\n<hr>\n<h2>Ranking features code (Polars)</h2>\n<pre><code> polars  pl\n\n\nrank_order = {\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n}\n\n () :\n    exprs = []\n     col, order  rank_order.items():\n         col  df.columns:\n            exprs.append(\n                pl.col(col)\n                  .rank(method=, descending=(order == ))\n                  .over(group_col)\n                  .alias()\n            )\n     df.with_columns(exprs)\n</code></pre>\n<p>There are more experiments to be done, for example adding  'miniRules0_statusInfos': 'desc',  'miniRules1_statusInfos': 'desc' will further improve local valid and online PB, unfortunately LB drop:)    </p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>valid</th>\n<th>LB</th>\n<th>PB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>best single model</td>\n<td>0.504745</td>\n<td>0.54062</td>\n<td>0.54156</td>\n</tr>\n<tr>\n<td>adding more rank feats like  miniRules*_statusInfos</td>\n<td>0.506949</td>\n<td>0.53557</td>\n<td>0.54284</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h3>Additonal ranking features based on (ranker_id, flight_hash)</h3>\n<p>As we fould rule based on flight_hash could improve local and online, we could add them to our feats so tree models could learn more.    <br>\nSo in additon to add_ranking_feats(df, 'ranker_id')  I also added add_ranking_feats(df, (ranker_id, flight_hash))</p>\n<pre><code>COLS_TO_COMPARE = [\n    ,\n    ,\n    ,\n    ,\n    ,\n    \n]\n\ndf = df.with_columns(\n    pl.concat_str(\n        [pl.col(c).cast().fill_null()  c  COLS_TO_COMPARE]\n    ).alias()\n)\n\ndf = add_ranking_feats(df, )\ndf = add_ranking_feats(df, [, ])\n</code></pre>\n<hr>\n<h3>Couting of important base features groupby profileId and compayID</h3>\n<pre><code>\n ():\n  exprs = []\n   col  rank_order.keys():\n     col  df.columns:\n      exprs.extend([\n        pl.col(col).mean().over(group_col).alias(),\n        pl.col(col).().over(group_col).alias(),\n        pl.col(col).().over(group_col).alias(),\n        pl.col(col).std().over(group_col).alias(),\n        pl.col(col).median().over(group_col).alias(),\n      ])\n  df = df.with_columns(exprs)\n  exprs = []\n   col  rank_order.keys():\n     col  df.columns:\n      exprs.extend([\n        ((pl.col(col) - pl.col()) / (pl.col() + )).alias(),\n        ((pl.col(col) - pl.col()) / (pl.col() - pl.col() + )).alias(),\n        (pl.col(col) / (pl.col() + )).alias(),\n      ])\n\n  df = df.with_columns(exprs)\n   df\n\ndf = add_stats_feats(df, )\ndf = add_stats_feats(df, )\n</code></pre>\n<h3>Ensembel</h3>\n<p>During the competition, I trained 41 models and tuned their weights with Optuna, but none beat a simple ensemble of my top five single models. For simplicity, I report that 5‑model ensemble result (late submission) here: all models use the same features but differ by learner and objective. XGBoost consistently performed best, with LightGBM adding useful diversity; CatBoost underperformed. The ensemble includes five models—XGBoost with rank:ndcg, rank:map, and rank:pairwise objectives, an XGBoost binary classifier, and LightGBM with LambdaRank (LightGBM’s xendcg performed much worse). The best setup uses equal weights across the five models. </p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>valid</th>\n<th>LB</th>\n<th>PB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>5 top single models</td>\n<td>0.5101</td>\n<td>0.54053</td>\n<td>0.54725</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h3>Further improments</h3>\n<p>I intentionally avoided using any stat features derived from labels or selected outcomes, as including them could lead the model to \"see\" future labels during training, resulting in overfitting and poor generalization on unseen data. </p>\n<p>However, one promising workaround is to construct <strong>windowed features</strong>—that is, using only historical statistics available prior to each instance's timestamp. This technique is effectively demonstrated in <a href=\"https://www.kaggle.com/code/mikhailgolubchik/sm-xgboost-single\" target=\"_blank\">Mikhail Golubchik’s XGBoost notebook</a>. Thanks to <a href=\"https://www.kaggle.com/mikhailgolubchik\" target=\"_blank\">@mikhailgolubchik</a> for the inspiration!</p>\n<p>I haven’t implemented this yet, mainly because it’s time-consuming and I was concerned about a potential mismatch between training and test data: for training, you can use stats up to the instance time, but for test data, you're limited to stats from earlier days only. Still, this approach is definitely worth exploring further.</p>\n<p>I merged the code from <a href=\"https://www.kaggle.com/mikhailgolubchik\" target=\"_blank\">@mikhailgolubchik</a>'s notebook and changed the feats to be used(source_cols) a bit , notice you could investigate more feats.  </p>\n<pre><code> FLAGS.history_avg:\n  test = util.get_test(df)\n\n  df = util.get_nontest(df)\n  train = util.get_train(df)\n\n    FLAGS.online:\n    valid = util.get_valid(df)\n    test = pl.concat([valid, test], =)\n\n  source_cols = [\n      ,\n      ,\n      ,\n      ,\n      ,\n      ,\n      ,\n      ,\n      ,\n      ,\n      ,\n      ,\n  ]\n\n  train, df_stats_pr = make_history_avg(train,\n                                      =source_cols,\n                                      =,\n                                      =)\n  train, df_stats_co = make_history_avg(train,\n                                      =source_cols,\n                                      =,\n                                      =)\n\n  test = test.join(df_stats_pr, =, =)\n  test = test.join(df_stats_co, =, =)\n\n  df = gz.align_and_concat([train, test])\n</code></pre>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>valid</th>\n<th>LB</th>\n<th>PB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>add selected/pos history avg feats</td>\n<td>0.5120</td>\n<td>0.54889</td>\n<td>0.54942</td>\n</tr>\n<tr>\n<td>4 history avg feats xgb ensemble</td>\n<td></td>\n<td>0.55201</td>\n<td>0.55291</td>\n</tr>\n</tbody>\n</table>\n<p>Well such great improvment, PB up about 0.008, so personally I think we might have over 0.6 score on PB but &gt; 0.7 might be a bit too hard but not mission impossible:) For this dataset about 50% profileId in test exists in the train dataset, I did not investiage for each profileId how many ranker_ids exits, if there are many then may be we can also use more user history book info.  Another finding is with model perform better we do not need rerank anymore.  </p>\n<h3>Opensource the solution</h3>\n<p><a href=\"https://www.kaggle.com/code/goldenlock/aeroclub-recsys-2025-1st-solution\" target=\"_blank\">https://www.kaggle.com/code/goldenlock/aeroclub-recsys-2025-1st-solution</a>  <br>\nNotice aeroclub-recsys-2025-model2 has the 4 xgb models which ensemble with PB 55291 and the submission.parquet.   <br>\nPlease refer to the notebook and its README.  </p>",
  "messages": [
    {
      "id": "3271087",
      "postDate": "08/18/2025 03:45:09",
      "content": "<hr>\n<p>First, I want to express my sincere gratitude to the organizers for providing such a large-scale, real-world dataset. It offered a rare opportunity to dive deep into practical challenges and learn from them.</p>\n<p>I also want to thank the authors of several outstanding notebooks that guided my journey:</p>\n<ul>\n<li>I began with <a href=\"https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars\" target=\"_blank\">XGBoost Ranker with Polars</a> a clean and efficient starting point that helped me grasp the essentials of ranking models. Thank you <a href=\"https://www.kaggle.com/ka1242\" target=\"_blank\">@ka1242</a>!</li>\n<li>The most creative idea, in my opinion, came from <a href=\"https://www.kaggle.com/code/mango789/xgboost-ranker-rule-based-rerank\" target=\"_blank\">reranking approach</a> . The rule-based reranking worked like a charm, and without that notebook, I wouldn’t have realized how crucial <code>flight_hash</code> is to the final performance. Huge thanks to <a href=\"https://www.kaggle.com/mango789\" target=\"_blank\">@mango789</a> !</li>\n</ul>\n<hr>\n<h2>Approach overview and note on test data usage</h2>\n<p>My main single-model solution initially leveraged a leak-like feature: I included the test data from the very beginning of the competition. Count-based categorical encoding built on the full dataset noticeably improved both local validation and the public leaderboard.</p>\n<p>I recall a discussion where the host mentioned that using the test set was acceptable because it contains no labels. While some recommendation competitions forbid using the test data, I didn’t find such a restriction here. Still, in real scenarios we cannot rely on future data, and we usually predict day by day using only past labels. A streaming setup would better reflect production.</p>\n<p>Below I share results for the same single model trained without the test data (i.e., I removed the test set from the count-based encodings and training input).</p>\n<hr>\n<h2>Results without using test data</h2>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>valid</th>\n<th>LB</th>\n<th>PB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>best single model</td>\n<td>0.504745</td>\n<td>0.54062</td>\n<td>0.54156</td>\n</tr>\n<tr>\n<td>best single model without testdata</td>\n<td>0.478704</td>\n<td>0.51978</td>\n<td>0.52692</td>\n</tr>\n</tbody>\n</table>\n<p>This is an ongoing write-up; I may refine it or add more data stats later.</p>\n<hr>\n<h2>Key components</h2>\n<h3>Validation strategy</h3>\n<ul>\n<li><strong>Temporal split:</strong> This is a recommendation task, so datetime order matters. The dataset spans <code>requestDate</code> from 2024-05-17 03:03:08 to 2024-12-31 18:54:00, which I index as days 1–229.</li>\n<li><strong>Fold design:</strong> I used a single temporal fold: train on days 1–103, validate on days 104–166, and test on days 167–229.</li>\n<li><strong>Rationale:</strong> Validation and test windows have similar lengths, and local validation generally matched the LB, especially for larger improvements. When there was a mismatch, I trusted local validation. One exception was the final Optuna weight tuning for ensembles, which can overfit to the validation set.</li>\n</ul>\n<h3>Count-based categorical encoding</h3>\n<ul>\n<li><strong>XGBoost / LightGBM:</strong> I applied count-based encoding to all categorical features, then treated all columns as numerical — no categorical indices were passed into training.</li>\n<li><strong>CatBoost:</strong> I also experimented with exactly the same count-based encoding here, converting everything to numerical features. However, in my tests, all variations yielded very similar results, with no meaningful gain over the baseline.</li>\n</ul>\n<h3>Time features</h3>\n<ul>\n<li><strong>Importance:</strong> Time-derived features gave the largest lift compared to using only the provided CSV fields.</li>\n<li><strong>Cyclical encoding:</strong> I added sine/cosine transforms so the model can capture wrap-around (e.g., hours 23 and 0 are adjacent).</li>\n</ul>\n<h3>Duration and distance features</h3>\n<ul>\n<li><strong>Accurate duration:</strong> I converted local times to UTC to compute true elapsed durations.</li>\n<li><strong>Distance-based signals:</strong> I computed airport distances from lat/lon and derived related features such as <code>direct_price_per_km</code>.</li>\n</ul>\n<h3>Simple rank-based features on important base features</h3>\n<ul>\n<li><strong>Motivation:</strong> Ranking proxies (per group) for price, duration, and itinerary complexity helped the ranker focus on relative ordering within a user/query.</li>\n</ul>\n<hr>\n<h2>Ranking features code (Polars)</h2>\n<pre><code> polars  pl\n\n\nrank_order = {\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n}\n\n () :\n    exprs = []\n     col, order  rank_order.items():\n         col  df.columns:\n            exprs.append(\n                pl.col(col)\n                  .rank(method=, descending=(order == ))\n                  .over(group_col)\n                  .alias()\n            )\n     df.with_columns(exprs)\n</code></pre>\n<p>There are more experiments to be done, for example adding  'miniRules0_statusInfos': 'desc',  'miniRules1_statusInfos': 'desc' will further improve local valid and online PB, unfortunately LB drop:)    </p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>valid</th>\n<th>LB</th>\n<th>PB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>best single model</td>\n<td>0.504745</td>\n<td>0.54062</td>\n<td>0.54156</td>\n</tr>\n<tr>\n<td>adding more rank feats like  miniRules*_statusInfos</td>\n<td>0.506949</td>\n<td>0.53557</td>\n<td>0.54284</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h3>Additonal ranking features based on (ranker_id, flight_hash)</h3>\n<p>As we fould rule based on flight_hash could improve local and online, we could add them to our feats so tree models could learn more.    <br>\nSo in additon to add_ranking_feats(df, 'ranker_id')  I also added add_ranking_feats(df, (ranker_id, flight_hash))</p>\n<pre><code>COLS_TO_COMPARE = [\n    ,\n    ,\n    ,\n    ,\n    ,\n    \n]\n\ndf = df.with_columns(\n    pl.concat_str(\n        [pl.col(c).cast().fill_null()  c  COLS_TO_COMPARE]\n    ).alias()\n)\n\ndf = add_ranking_feats(df, )\ndf = add_ranking_feats(df, [, ])\n</code></pre>\n<hr>\n<h3>Couting of important base features groupby profileId and compayID</h3>\n<pre><code>\n ():\n  exprs = []\n   col  rank_order.keys():\n     col  df.columns:\n      exprs.extend([\n        pl.col(col).mean().over(group_col).alias(),\n        pl.col(col).().over(group_col).alias(),\n        pl.col(col).().over(group_col).alias(),\n        pl.col(col).std().over(group_col).alias(),\n        pl.col(col).median().over(group_col).alias(),\n      ])\n  df = df.with_columns(exprs)\n  exprs = []\n   col  rank_order.keys():\n     col  df.columns:\n      exprs.extend([\n        ((pl.col(col) - pl.col()) / (pl.col() + )).alias(),\n        ((pl.col(col) - pl.col()) / (pl.col() - pl.col() + )).alias(),\n        (pl.col(col) / (pl.col() + )).alias(),\n      ])\n\n  df = df.with_columns(exprs)\n   df\n\ndf = add_stats_feats(df, )\ndf = add_stats_feats(df, )\n</code></pre>\n<h3>Ensembel</h3>\n<p>During the competition, I trained 41 models and tuned their weights with Optuna, but none beat a simple ensemble of my top five single models. For simplicity, I report that 5‑model ensemble result (late submission) here: all models use the same features but differ by learner and objective. XGBoost consistently performed best, with LightGBM adding useful diversity; CatBoost underperformed. The ensemble includes five models—XGBoost with rank:ndcg, rank:map, and rank:pairwise objectives, an XGBoost binary classifier, and LightGBM with LambdaRank (LightGBM’s xendcg performed much worse). The best setup uses equal weights across the five models. </p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>valid</th>\n<th>LB</th>\n<th>PB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>5 top single models</td>\n<td>0.5101</td>\n<td>0.54053</td>\n<td>0.54725</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h3>Further improments</h3>\n<p>I intentionally avoided using any stat features derived from labels or selected outcomes, as including them could lead the model to \"see\" future labels during training, resulting in overfitting and poor generalization on unseen data. </p>\n<p>However, one promising workaround is to construct <strong>windowed features</strong>—that is, using only historical statistics available prior to each instance's timestamp. This technique is effectively demonstrated in <a href=\"https://www.kaggle.com/code/mikhailgolubchik/sm-xgboost-single\" target=\"_blank\">Mikhail Golubchik’s XGBoost notebook</a>. Thanks to <a href=\"https://www.kaggle.com/mikhailgolubchik\" target=\"_blank\">@mikhailgolubchik</a> for the inspiration!</p>\n<p>I haven’t implemented this yet, mainly because it’s time-consuming and I was concerned about a potential mismatch between training and test data: for training, you can use stats up to the instance time, but for test data, you're limited to stats from earlier days only. Still, this approach is definitely worth exploring further.</p>\n<p>I merged the code from <a href=\"https://www.kaggle.com/mikhailgolubchik\" target=\"_blank\">@mikhailgolubchik</a>'s notebook and changed the feats to be used(source_cols) a bit , notice you could investigate more feats.  </p>\n<pre><code> FLAGS.history_avg:\n  test = util.get_test(df)\n\n  df = util.get_nontest(df)\n  train = util.get_train(df)\n\n    FLAGS.online:\n    valid = util.get_valid(df)\n    test = pl.concat([valid, test], =)\n\n  source_cols = [\n      ,\n      ,\n      ,\n      ,\n      ,\n      ,\n      ,\n      ,\n      ,\n      ,\n      ,\n      ,\n  ]\n\n  train, df_stats_pr = make_history_avg(train,\n                                      =source_cols,\n                                      =,\n                                      =)\n  train, df_stats_co = make_history_avg(train,\n                                      =source_cols,\n                                      =,\n                                      =)\n\n  test = test.join(df_stats_pr, =, =)\n  test = test.join(df_stats_co, =, =)\n\n  df = gz.align_and_concat([train, test])\n</code></pre>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>valid</th>\n<th>LB</th>\n<th>PB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>add selected/pos history avg feats</td>\n<td>0.5120</td>\n<td>0.54889</td>\n<td>0.54942</td>\n</tr>\n<tr>\n<td>4 history avg feats xgb ensemble</td>\n<td></td>\n<td>0.55201</td>\n<td>0.55291</td>\n</tr>\n</tbody>\n</table>\n<p>Well such great improvment, PB up about 0.008, so personally I think we might have over 0.6 score on PB but &gt; 0.7 might be a bit too hard but not mission impossible:) For this dataset about 50% profileId in test exists in the train dataset, I did not investiage for each profileId how many ranker_ids exits, if there are many then may be we can also use more user history book info.  Another finding is with model perform better we do not need rerank anymore.  </p>\n<h3>Opensource the solution</h3>\n<p><a href=\"https://www.kaggle.com/code/goldenlock/aeroclub-recsys-2025-1st-solution\" target=\"_blank\">https://www.kaggle.com/code/goldenlock/aeroclub-recsys-2025-1st-solution</a>  <br>\nNotice aeroclub-recsys-2025-model2 has the 4 xgb models which ensemble with PB 55291 and the submission.parquet.   <br>\nPlease refer to the notebook and its README.  </p>",
      "rawMarkdown": "First, I want to express my sincere gratitude to the organizers for providing such a large-scale, real-world dataset. It offered a rare opportunity to dive deep into practical challenges and learn from them.\n\nI also want to thank the authors of several outstanding notebooks that guided my journey:\n\n- I began with [XGBoost Ranker with Polars](https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars) a clean and efficient starting point that helped me grasp the essentials of ranking models. Thank you @ka1242!\n- The most creative idea, in my opinion, came from [reranking approach](https://www.kaggle.com/code/mango789/xgboost-ranker-rule-based-rerank) . The rule-based reranking worked like a charm, and without that notebook, I wouldn’t have realized how crucial `flight_hash` is to the final performance. Huge thanks to @mango789 !\n\n---\n\n## Approach overview and note on test data usage\n\nMy main single-model solution initially leveraged a leak-like feature: I included the test data from the very beginning of the competition. Count-based categorical encoding built on the full dataset noticeably improved both local validation and the public leaderboard.\n\nI recall a discussion where the host mentioned that using the test set was acceptable because it contains no labels. While some recommendation competitions forbid using the test data, I didn’t find such a restriction here. Still, in real scenarios we cannot rely on future data, and we usually predict day by day using only past labels. A streaming setup would better reflect production.\n\nBelow I share results for the same single model trained without the test data (i.e., I removed the test set from the count-based encodings and training input).\n\n---\n\n## Results without using test data\n\n| model | valid | LB  | PB |\n| --- | --- | --- | -- |\n| best single model | 0.504745 | 0.54062 | 0.54156 | \n| best single model without testdata | 0.478704 | 0.51978 | 0.52692 |\n\nThis is an ongoing write-up; I may refine it or add more data stats later.\n\n---\n\n## Key components\n\n### Validation strategy\n\n- **Temporal split:** This is a recommendation task, so datetime order matters. The dataset spans `requestDate` from 2024-05-17 03:03:08 to 2024-12-31 18:54:00, which I index as days 1–229.\n- **Fold design:** I used a single temporal fold: train on days 1–103, validate on days 104–166, and test on days 167–229.\n- **Rationale:** Validation and test windows have similar lengths, and local validation generally matched the LB, especially for larger improvements. When there was a mismatch, I trusted local validation. One exception was the final Optuna weight tuning for ensembles, which can overfit to the validation set.\n\n### Count-based categorical encoding\n\n- **XGBoost / LightGBM:** I applied count-based encoding to all categorical features, then treated all columns as numerical — no categorical indices were passed into training.\n- **CatBoost:** I also experimented with exactly the same count-based encoding here, converting everything to numerical features. However, in my tests, all variations yielded very similar results, with no meaningful gain over the baseline.\n\n### Time features\n\n- **Importance:** Time-derived features gave the largest lift compared to using only the provided CSV fields.\n- **Cyclical encoding:** I added sine/cosine transforms so the model can capture wrap-around (e.g., hours 23 and 0 are adjacent).\n\n### Duration and distance features\n\n- **Accurate duration:** I converted local times to UTC to compute true elapsed durations.\n- **Distance-based signals:** I computed airport distances from lat/lon and derived related features such as `direct_price_per_km`.\n\n### Simple rank-based features on important base features\n\n- **Motivation:** Ranking proxies (per group) for price, duration, and itinerary complexity helped the ranker focus on relative ordering within a user/query.\n\n---\n\n## Ranking features code (Polars)\n\n```python\nimport polars as pl\n\n# Columns to rank and their desired order within each group (e.g., ranker_id/session)\nrank_order = {\n    'totalPrice': 'asc',\n    'flight_duration_total': 'asc',\n    'book_lead_time_hours': 'desc',\n    'flight_duration_travel_ratio': 'asc',\n    'seg_legs_all_count': 'asc',\n    'avg_cabin_legs_all': 'desc',\n    'avg_baggage_count_legs_all': 'desc',\n    'avg_baggage_weight_legs_all': 'desc',\n    'direct_price_per_km': 'asc',\n}\n\ndef add_ranking_feats(df, group_col='ranker_id') :\n    exprs = []\n    for col, order in rank_order.items():\n        if col in df.columns:\n            exprs.append(\n                pl.col(col)\n                  .rank(method='average', descending=(order == 'desc'))\n                  .over(group_col)\n                  .alias(f'rank_{col}')\n            )\n    return df.with_columns(exprs)\n```\nThere are more experiments to be done, for example adding  'miniRules0_statusInfos': 'desc',  'miniRules1_statusInfos': 'desc' will further improve local valid and online PB, unfortunately LB drop:)    \n| model | valid | LB  | PB |\n| --- | --- | --- | -- |\n| best single model | 0.504745 | 0.54062 | 0.54156 | \n| adding more rank feats like  miniRules*_statusInfos| 0.506949 | 0.53557 | 0.54284 |\n\n---\n### Additonal ranking features based on (ranker_id, flight_hash)  \nAs we fould rule based on flight_hash could improve local and online, we could add them to our feats so tree models could learn more.    \nSo in additon to add_ranking_feats(df, 'ranker_id')  I also added add_ranking_feats(df, (ranker_id, flight_hash))\n\n```python\nCOLS_TO_COMPARE = [\n    \"legs0_departureAt\",\n    \"legs0_arrivalAt\",\n    \"legs1_departureAt\",\n    \"legs1_arrivalAt\",\n    \"legs0_segments0_flightNumber\",\n    \"legs1_segments0_flightNumber\"\n]\n\ndf = df.with_columns(\n    pl.concat_str(\n        [pl.col(c).cast(str).fill_null(\"NULL\") for c in COLS_TO_COMPARE]\n    ).alias(\"flight_hash\")\n)\n\ndf = add_ranking_feats(df, 'ranker_id')\ndf = add_ranking_feats(df, ['ranker_id', 'flight_hash'])\n```\n--- \n### Couting of important base features groupby profileId and compayID  \n```python\n#Notice not consider label/selected and df is (train and test) combined\ndef add_stats_feats(df, group_col='profileId'):\n  exprs = []\n  for col in rank_order.keys():\n    if col in df.columns:\n      exprs.extend([\n        pl.col(col).mean().over(group_col).alias(f\"avg_{col}_{group_col}_stats\"),\n        pl.col(col).min().over(group_col).alias(f\"min_{col}_{group_col}_stats\"),\n        pl.col(col).max().over(group_col).alias(f\"max_{col}_{group_col}_stats\"),\n        pl.col(col).std().over(group_col).alias(f\"std_{col}_{group_col}_stats\"),\n        pl.col(col).median().over(group_col).alias(f\"median_{col}_{group_col}_stats\"),\n      ])\n  df = df.with_columns(exprs)\n  exprs = []\n  for col in rank_order.keys():\n    if col in df.columns:\n      exprs.extend([\n        ((pl.col(col) - pl.col(f\"avg_{col}_{group_col}_stats\")) / (pl.col(f\"std_{col}_{group_col}_stats\") + 1e-5)).alias(f\"{col}_zscore_{group_col}_stats\"),\n        ((pl.col(col) - pl.col(f\"min_{col}_{group_col}_stats\")) / (pl.col(f\"max_{col}_{group_col}_stats\") - pl.col(f\"min_{col}_{group_col}_stats\") + 1e-5)).alias(f\"{col}_minmax_{group_col}_stats\"),\n        (pl.col(col) / (pl.col(f\"avg_{col}_{group_col}_stats\") + 1e-5)).alias(f\"{col}_{group_col}_stats_ratio\"),\n      ])\n     \n  df = df.with_columns(exprs)\n  return df\n\ndf = add_stats_feats(df, 'pfofileId')\ndf = add_stats_feats(df, 'compnayID')\n```\n\n### Ensembel\nDuring the competition, I trained 41 models and tuned their weights with Optuna, but none beat a simple ensemble of my top five single models. For simplicity, I report that 5‑model ensemble result (late submission) here: all models use the same features but differ by learner and objective. XGBoost consistently performed best, with LightGBM adding useful diversity; CatBoost underperformed. The ensemble includes five models—XGBoost with rank:ndcg, rank:map, and rank:pairwise objectives, an XGBoost binary classifier, and LightGBM with LambdaRank (LightGBM’s xendcg performed much worse). The best setup uses equal weights across the five models. \n| model | valid | LB | PB |\n| --- | --- | --- |  --- |\n| 5 top single models |  0.5101  | 0.54053 | 0.54725 |\n\n---\n### Further improments \nI intentionally avoided using any stat features derived from labels or selected outcomes, as including them could lead the model to \"see\" future labels during training, resulting in overfitting and poor generalization on unseen data. \n\nHowever, one promising workaround is to construct **windowed features**—that is, using only historical statistics available prior to each instance's timestamp. This technique is effectively demonstrated in [Mikhail Golubchik’s XGBoost notebook](https://www.kaggle.com/code/mikhailgolubchik/sm-xgboost-single). Thanks to @mikhailgolubchik for the inspiration!\n\nI haven’t implemented this yet, mainly because it’s time-consuming and I was concerned about a potential mismatch between training and test data: for training, you can use stats up to the instance time, but for test data, you're limited to stats from earlier days only. Still, this approach is definitely worth exploring further.\n\nI merged the code from @mikhailgolubchik's notebook and changed the feats to be used(source_cols) a bit , notice you could investigate more feats.  \n```\nif FLAGS.history_avg:\n  test = util.get_test(df)\n  \n  df = util.get_nontest(df)\n  train = util.get_train(df)\n  \n  if not FLAGS.online:\n    valid = util.get_valid(df)\n    test = pl.concat([valid, test], how='vertical')\n\n  source_cols = [\n      'time_legs0_departureAt_hour',\n      'time_legs1_departureAt_hour',\n      'time_legs0_arrivalAt_hour',\n      'time_legs1_arrivalAt_hour',\n      'rank_totalPrice',\n      'rank_flight_duration_total',\n      'avg_cabin_legs_all',\n      'avg_baggage_count_legs_all',\n      'avg_baggage_weight_legs_all',\n      'direct_price_per_km',\n      'miniRules1_statusInfos',\n      'miniRules0_statusInfos',\n  ]\n\n  train, df_stats_pr = make_history_avg(train,\n                                      source_cols=source_cols,\n                                      group_col=\"uid\",\n                                      suffix='_uid')\n  train, df_stats_co = make_history_avg(train,\n                                      source_cols=source_cols,\n                                      group_col=\"companyID\",\n                                      suffix='_company')\n  \n  test = test.join(df_stats_pr, on='uid', how='left')\n  test = test.join(df_stats_co, on='companyID', how='left')\n\n  df = gz.align_and_concat([train, test])\n```\n| model | valid | LB | PB |\n| --- | --- | --- |  --- |\n| add selected/pos history avg feats |  0.5120  | 0.54889 | 0.54942 |  \n| 4 history avg feats xgb ensemble|   | 0.55201 | 0.55291 |  \n\nWell such great improvment, PB up about 0.008, so personally I think we might have over 0.6 score on PB but > 0.7 might be a bit too hard but not mission impossible:) For this dataset about 50% profileId in test exists in the train dataset, I did not investiage for each profileId how many ranker_ids exits, if there are many then may be we can also use more user history book info.  Another finding is with model perform better we do not need rerank anymore.  \n\n### Opensource the solution  \nhttps://www.kaggle.com/code/goldenlock/aeroclub-recsys-2025-1st-solution  \nNotice aeroclub-recsys-2025-model2 has the 4 xgb models which ensemble with PB 55291 and the submission.parquet.   \nPlease refer to the notebook and its README.",
      "votes": null
    },
    {
      "id": "3271169",
      "postDate": "08/18/2025 08:09:03",
      "content": "<p>您好。 Congratulation！ </p>",
      "rawMarkdown": "您好。 Congratulation！",
      "votes": null
    },
    {
      "id": "3271431",
      "postDate": "08/18/2025 21:36:06",
      "content": "<blockquote>\n  <p>Still, in real scenarios we cannot rely on future data, and we usually predict day by day using only past labels. A streaming setup would better reflect production.</p>\n</blockquote>\n<p>Why your approach could work in real scenario in this specific case? In flight search, count-based encoding is realistic because flight schedules, route frequencies, and large volumes of daily requests are all available and some features can be pre-computed.</p>",
      "rawMarkdown": ">Still, in real scenarios we cannot rely on future data, and we usually predict day by day using only past labels. A streaming setup would better reflect production.\n\nWhy your approach could work in real scenario in this specific case? In flight search, count-based encoding is realistic because flight schedules, route frequencies, and large volumes of daily requests are all available and some features can be pre-computed.",
      "votes": null
    },
    {
      "id": "3271434",
      "postDate": "08/18/2025 21:47:02",
      "content": "<p>Thanks for sharing this!</p>",
      "rawMarkdown": "Thanks for sharing this!",
      "votes": null
    },
    {
      "id": "3271452",
      "postDate": "08/19/2025 00:05:57",
      "content": "<p>Yeah you are correct, we could have real time counting, but for my approach here the problem is I used counting of like day 100 to predict on day 99, however the total pipline could still work if I use test dataset but still if I count based only on instance history (not look ahead), so using just  previous window/history feature could be fine either counting on all expousure data or selected data.</p>",
      "rawMarkdown": "Yeah you are correct, we could have real time counting, but for my approach here the problem is I used counting of like day 100 to predict on day 99, however the total pipline could still work if I use test dataset but still if I count based only on instance history (not look ahead), so using just  previous window/history feature could be fine either counting on all expousure data or selected data.",
      "votes": null
    },
    {
      "id": "3271644",
      "postDate": "08/19/2025 10:24:16",
      "content": "<p>good work this is my first comment</p>",
      "rawMarkdown": "good work this is my first comment",
      "votes": null
    },
    {
      "id": "3272919",
      "postDate": "08/21/2025 18:48:29",
      "content": "<p>Nice, really interesting to see you approached this and to learn from it. Thanks for sharing! Congrats on winning!!</p>",
      "rawMarkdown": "Nice, really interesting to see you approached this and to learn from it. Thanks for sharing! Congrats on winning!!",
      "votes": null
    },
    {
      "id": "3273518",
      "postDate": "08/22/2025 20:11:23",
      "content": "<p>Thank you for sharing this it will help while learning!</p>",
      "rawMarkdown": "Thank you for sharing this it will help while learning!",
      "votes": null
    },
    {
      "id": "3273531",
      "postDate": "08/22/2025 20:33:35",
      "content": "<p>Super cool work !! thank you !!</p>",
      "rawMarkdown": "Super cool work !! thank you !!",
      "votes": null
    },
    {
      "id": "3274842",
      "postDate": "08/25/2025 14:02:30",
      "content": "<p>Thanks for sharing…it helps alot!</p>",
      "rawMarkdown": "Thanks for sharing...it helps alot!",
      "votes": null
    },
    {
      "id": "3276565",
      "postDate": "08/26/2025 18:00:30",
      "content": "<p>Nice and interesting </p>",
      "rawMarkdown": "Nice and interesting",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3271169,
      "author_name": "sergeyqt2024",
      "author_url": "",
      "post_date": "08/18/2025 08:09:03",
      "content": "<p>您好。 Congratulation！ </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3271431,
      "author_name": "samvelkoch",
      "author_url": "",
      "post_date": "08/18/2025 21:36:06",
      "content": "<blockquote>\n  <p>Still, in real scenarios we cannot rely on future data, and we usually predict day by day using only past labels. A streaming setup would better reflect production.</p>\n</blockquote>\n<p>Why your approach could work in real scenario in this specific case? In flight search, count-based encoding is realistic because flight schedules, route frequencies, and large volumes of daily requests are all available and some features can be pre-computed.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3271452,
          "author_name": "goldenlock",
          "author_url": "",
          "post_date": "08/19/2025 00:05:57",
          "content": "<p>Yeah you are correct, we could have real time counting, but for my approach here the problem is I used counting of like day 100 to predict on day 99, however the total pipline could still work if I use test dataset but still if I count based only on instance history (not look ahead), so using just  previous window/history feature could be fine either counting on all expousure data or selected data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3271434,
      "author_name": "irakozekelly",
      "author_url": "",
      "post_date": "08/18/2025 21:47:02",
      "content": "<p>Thanks for sharing this!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3271644,
      "author_name": "khyati267",
      "author_url": "",
      "post_date": "08/19/2025 10:24:16",
      "content": "<p>good work this is my first comment</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3272919,
      "author_name": "hypoxiic",
      "author_url": "",
      "post_date": "08/21/2025 18:48:29",
      "content": "<p>Nice, really interesting to see you approached this and to learn from it. Thanks for sharing! Congrats on winning!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3273518,
      "author_name": "irakozekelly",
      "author_url": "",
      "post_date": "08/22/2025 20:11:23",
      "content": "<p>Thank you for sharing this it will help while learning!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3273531,
      "author_name": "abdelhakouanzougui2",
      "author_url": "",
      "post_date": "08/22/2025 20:33:35",
      "content": "<p>Super cool work !! thank you !!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3274842,
      "author_name": "saraharshadbcs",
      "author_url": "",
      "post_date": "08/25/2025 14:02:30",
      "content": "<p>Thanks for sharing…it helps alot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3276565,
      "author_name": "sanjanakumari651",
      "author_url": "",
      "post_date": "08/26/2025 18:00:30",
      "content": "<p>Nice and interesting </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3271087": "First, I want to express my sincere gratitude to the organizers for providing such a large-scale, real-world dataset. It offered a rare opportunity to dive deep into practical challenges and learn from them.\n\nI also want to thank the authors of several outstanding notebooks that guided my journey:\n\n- I began with [XGBoost Ranker with Polars](https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars) a clean and efficient starting point that helped me grasp the essentials of ranking models. Thank you @ka1242!\n- The most creative idea, in my opinion, came from [reranking approach](https://www.kaggle.com/code/mango789/xgboost-ranker-rule-based-rerank) . The rule-based reranking worked like a charm, and without that notebook, I wouldn’t have realized how crucial `flight_hash` is to the final performance. Huge thanks to @mango789 !\n\n---\n\n## Approach overview and note on test data usage\n\nMy main single-model solution initially leveraged a leak-like feature: I included the test data from the very beginning of the competition. Count-based categorical encoding built on the full dataset noticeably improved both local validation and the public leaderboard.\n\nI recall a discussion where the host mentioned that using the test set was acceptable because it contains no labels. While some recommendation competitions forbid using the test data, I didn’t find such a restriction here. Still, in real scenarios we cannot rely on future data, and we usually predict day by day using only past labels. A streaming setup would better reflect production.\n\nBelow I share results for the same single model trained without the test data (i.e., I removed the test set from the count-based encodings and training input).\n\n---\n\n## Results without using test data\n\n| model | valid | LB  | PB |\n| --- | --- | --- | -- |\n| best single model | 0.504745 | 0.54062 | 0.54156 | \n| best single model without testdata | 0.478704 | 0.51978 | 0.52692 |\n\nThis is an ongoing write-up; I may refine it or add more data stats later.\n\n---\n\n## Key components\n\n### Validation strategy\n\n- **Temporal split:** This is a recommendation task, so datetime order matters. The dataset spans `requestDate` from 2024-05-17 03:03:08 to 2024-12-31 18:54:00, which I index as days 1–229.\n- **Fold design:** I used a single temporal fold: train on days 1–103, validate on days 104–166, and test on days 167–229.\n- **Rationale:** Validation and test windows have similar lengths, and local validation generally matched the LB, especially for larger improvements. When there was a mismatch, I trusted local validation. One exception was the final Optuna weight tuning for ensembles, which can overfit to the validation set.\n\n### Count-based categorical encoding\n\n- **XGBoost / LightGBM:** I applied count-based encoding to all categorical features, then treated all columns as numerical — no categorical indices were passed into training.\n- **CatBoost:** I also experimented with exactly the same count-based encoding here, converting everything to numerical features. However, in my tests, all variations yielded very similar results, with no meaningful gain over the baseline.\n\n### Time features\n\n- **Importance:** Time-derived features gave the largest lift compared to using only the provided CSV fields.\n- **Cyclical encoding:** I added sine/cosine transforms so the model can capture wrap-around (e.g., hours 23 and 0 are adjacent).\n\n### Duration and distance features\n\n- **Accurate duration:** I converted local times to UTC to compute true elapsed durations.\n- **Distance-based signals:** I computed airport distances from lat/lon and derived related features such as `direct_price_per_km`.\n\n### Simple rank-based features on important base features\n\n- **Motivation:** Ranking proxies (per group) for price, duration, and itinerary complexity helped the ranker focus on relative ordering within a user/query.\n\n---\n\n## Ranking features code (Polars)\n\n```python\nimport polars as pl\n\n# Columns to rank and their desired order within each group (e.g., ranker_id/session)\nrank_order = {\n    'totalPrice': 'asc',\n    'flight_duration_total': 'asc',\n    'book_lead_time_hours': 'desc',\n    'flight_duration_travel_ratio': 'asc',\n    'seg_legs_all_count': 'asc',\n    'avg_cabin_legs_all': 'desc',\n    'avg_baggage_count_legs_all': 'desc',\n    'avg_baggage_weight_legs_all': 'desc',\n    'direct_price_per_km': 'asc',\n}\n\ndef add_ranking_feats(df, group_col='ranker_id') :\n    exprs = []\n    for col, order in rank_order.items():\n        if col in df.columns:\n            exprs.append(\n                pl.col(col)\n                  .rank(method='average', descending=(order == 'desc'))\n                  .over(group_col)\n                  .alias(f'rank_{col}')\n            )\n    return df.with_columns(exprs)\n```\nThere are more experiments to be done, for example adding  'miniRules0_statusInfos': 'desc',  'miniRules1_statusInfos': 'desc' will further improve local valid and online PB, unfortunately LB drop:)    \n| model | valid | LB  | PB |\n| --- | --- | --- | -- |\n| best single model | 0.504745 | 0.54062 | 0.54156 | \n| adding more rank feats like  miniRules*_statusInfos| 0.506949 | 0.53557 | 0.54284 |\n\n---\n### Additonal ranking features based on (ranker_id, flight_hash)  \nAs we fould rule based on flight_hash could improve local and online, we could add them to our feats so tree models could learn more.    \nSo in additon to add_ranking_feats(df, 'ranker_id')  I also added add_ranking_feats(df, (ranker_id, flight_hash))\n\n```python\nCOLS_TO_COMPARE = [\n    \"legs0_departureAt\",\n    \"legs0_arrivalAt\",\n    \"legs1_departureAt\",\n    \"legs1_arrivalAt\",\n    \"legs0_segments0_flightNumber\",\n    \"legs1_segments0_flightNumber\"\n]\n\ndf = df.with_columns(\n    pl.concat_str(\n        [pl.col(c).cast(str).fill_null(\"NULL\") for c in COLS_TO_COMPARE]\n    ).alias(\"flight_hash\")\n)\n\ndf = add_ranking_feats(df, 'ranker_id')\ndf = add_ranking_feats(df, ['ranker_id', 'flight_hash'])\n```\n--- \n### Couting of important base features groupby profileId and compayID  \n```python\n#Notice not consider label/selected and df is (train and test) combined\ndef add_stats_feats(df, group_col='profileId'):\n  exprs = []\n  for col in rank_order.keys():\n    if col in df.columns:\n      exprs.extend([\n        pl.col(col).mean().over(group_col).alias(f\"avg_{col}_{group_col}_stats\"),\n        pl.col(col).min().over(group_col).alias(f\"min_{col}_{group_col}_stats\"),\n        pl.col(col).max().over(group_col).alias(f\"max_{col}_{group_col}_stats\"),\n        pl.col(col).std().over(group_col).alias(f\"std_{col}_{group_col}_stats\"),\n        pl.col(col).median().over(group_col).alias(f\"median_{col}_{group_col}_stats\"),\n      ])\n  df = df.with_columns(exprs)\n  exprs = []\n  for col in rank_order.keys():\n    if col in df.columns:\n      exprs.extend([\n        ((pl.col(col) - pl.col(f\"avg_{col}_{group_col}_stats\")) / (pl.col(f\"std_{col}_{group_col}_stats\") + 1e-5)).alias(f\"{col}_zscore_{group_col}_stats\"),\n        ((pl.col(col) - pl.col(f\"min_{col}_{group_col}_stats\")) / (pl.col(f\"max_{col}_{group_col}_stats\") - pl.col(f\"min_{col}_{group_col}_stats\") + 1e-5)).alias(f\"{col}_minmax_{group_col}_stats\"),\n        (pl.col(col) / (pl.col(f\"avg_{col}_{group_col}_stats\") + 1e-5)).alias(f\"{col}_{group_col}_stats_ratio\"),\n      ])\n     \n  df = df.with_columns(exprs)\n  return df\n\ndf = add_stats_feats(df, 'pfofileId')\ndf = add_stats_feats(df, 'compnayID')\n```\n\n### Ensembel\nDuring the competition, I trained 41 models and tuned their weights with Optuna, but none beat a simple ensemble of my top five single models. For simplicity, I report that 5‑model ensemble result (late submission) here: all models use the same features but differ by learner and objective. XGBoost consistently performed best, with LightGBM adding useful diversity; CatBoost underperformed. The ensemble includes five models—XGBoost with rank:ndcg, rank:map, and rank:pairwise objectives, an XGBoost binary classifier, and LightGBM with LambdaRank (LightGBM’s xendcg performed much worse). The best setup uses equal weights across the five models. \n| model | valid | LB | PB |\n| --- | --- | --- |  --- |\n| 5 top single models |  0.5101  | 0.54053 | 0.54725 |\n\n---\n### Further improments \nI intentionally avoided using any stat features derived from labels or selected outcomes, as including them could lead the model to \"see\" future labels during training, resulting in overfitting and poor generalization on unseen data. \n\nHowever, one promising workaround is to construct **windowed features**—that is, using only historical statistics available prior to each instance's timestamp. This technique is effectively demonstrated in [Mikhail Golubchik’s XGBoost notebook](https://www.kaggle.com/code/mikhailgolubchik/sm-xgboost-single). Thanks to @mikhailgolubchik for the inspiration!\n\nI haven’t implemented this yet, mainly because it’s time-consuming and I was concerned about a potential mismatch between training and test data: for training, you can use stats up to the instance time, but for test data, you're limited to stats from earlier days only. Still, this approach is definitely worth exploring further.\n\nI merged the code from @mikhailgolubchik's notebook and changed the feats to be used(source_cols) a bit , notice you could investigate more feats.  \n```\nif FLAGS.history_avg:\n  test = util.get_test(df)\n  \n  df = util.get_nontest(df)\n  train = util.get_train(df)\n  \n  if not FLAGS.online:\n    valid = util.get_valid(df)\n    test = pl.concat([valid, test], how='vertical')\n\n  source_cols = [\n      'time_legs0_departureAt_hour',\n      'time_legs1_departureAt_hour',\n      'time_legs0_arrivalAt_hour',\n      'time_legs1_arrivalAt_hour',\n      'rank_totalPrice',\n      'rank_flight_duration_total',\n      'avg_cabin_legs_all',\n      'avg_baggage_count_legs_all',\n      'avg_baggage_weight_legs_all',\n      'direct_price_per_km',\n      'miniRules1_statusInfos',\n      'miniRules0_statusInfos',\n  ]\n\n  train, df_stats_pr = make_history_avg(train,\n                                      source_cols=source_cols,\n                                      group_col=\"uid\",\n                                      suffix='_uid')\n  train, df_stats_co = make_history_avg(train,\n                                      source_cols=source_cols,\n                                      group_col=\"companyID\",\n                                      suffix='_company')\n  \n  test = test.join(df_stats_pr, on='uid', how='left')\n  test = test.join(df_stats_co, on='companyID', how='left')\n\n  df = gz.align_and_concat([train, test])\n```\n| model | valid | LB | PB |\n| --- | --- | --- |  --- |\n| add selected/pos history avg feats |  0.5120  | 0.54889 | 0.54942 |  \n| 4 history avg feats xgb ensemble|   | 0.55201 | 0.55291 |  \n\nWell such great improvment, PB up about 0.008, so personally I think we might have over 0.6 score on PB but > 0.7 might be a bit too hard but not mission impossible:) For this dataset about 50% profileId in test exists in the train dataset, I did not investiage for each profileId how many ranker_ids exits, if there are many then may be we can also use more user history book info.  Another finding is with model perform better we do not need rerank anymore.  \n\n### Opensource the solution  \nhttps://www.kaggle.com/code/goldenlock/aeroclub-recsys-2025-1st-solution  \nNotice aeroclub-recsys-2025-model2 has the 4 xgb models which ensemble with PB 55291 and the submission.parquet.   \nPlease refer to the notebook and its README.",
    "3271169": "您好。 Congratulation！",
    "3271431": ">Still, in real scenarios we cannot rely on future data, and we usually predict day by day using only past labels. A streaming setup would better reflect production.\n\nWhy your approach could work in real scenario in this specific case? In flight search, count-based encoding is realistic because flight schedules, route frequencies, and large volumes of daily requests are all available and some features can be pre-computed.",
    "3271434": "Thanks for sharing this!",
    "3271452": "Yeah you are correct, we could have real time counting, but for my approach here the problem is I used counting of like day 100 to predict on day 99, however the total pipline could still work if I use test dataset but still if I count based only on instance history (not look ahead), so using just  previous window/history feature could be fine either counting on all expousure data or selected data.",
    "3271644": "good work this is my first comment",
    "3272919": "Nice, really interesting to see you approached this and to learn from it. Thanks for sharing! Congrats on winning!!",
    "3273518": "Thank you for sharing this it will help while learning!",
    "3273531": "Super cool work !! thank you !!",
    "3274842": "Thanks for sharing...it helps alot!",
    "3276565": "Nice and interesting"
  },
  "source": "meta"
}