{
  "id": 600461,
  "title": "FlightRank 2025: 2nd Place Solution",
  "url": "/competitions/aeroclub-recsys-2025/writeups/flightrank-2025-2nd-place-solution",
  "author_name": "",
  "post_date": "2025-08-22T21:53:10.123Z",
  "votes": 8,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I want to thank the organizers for creating such an engaging real-world ranking challenge. The flight recommendation domain provided rich opportunities to explore user behavior modeling and ensemble strategies.</p>\n<p>EDIT: Code is available at repo <a href=\"url\" target=\"_blank\">(https://github.com/Rajneesh-Tiwari/FlightRank-2025-Aeroclub-RecSys-Cup/</a></p>\n<p>I also want to acknowledge several excellent notebooks that inspired parts of this solution:</p>\n<ul>\n<li>The initial XGBoost ranker implementations (<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars</a>)  that provided a solid foundation</li>\n<li>Creative approaches to behavioral feature engineering that highlighted the importance of user history</li>\n<li>Ensemble strategies that demonstrated the value of model diversity</li>\n</ul>\n<h2>Final Results</h2>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Valid (CV)</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Best Single Model</td>\n<td>0.6412</td>\n<td>0.53807</td>\n</tr>\n<tr>\n<td>13-Model Ensemble</td>\n<td>0.6615</td>\n<td><strong>0.54175</strong></td>\n</tr>\n</tbody>\n</table>\n<h2>Approach Overview</h2>\n<p>My solution centers around a multi-algorithm ensemble with bucket-wise optimization, recognizing that user ranking behavior changes dramatically based on the number of flight options available in each search session. The core insight was that someone choosing from 15 options behaves very differently than someone choosing from 150 options.</p>\n<h2>Key Components</h2>\n<h3>Validation Strategy</h3>\n<p><strong>GroupKFold Cross-Validation</strong>: Used 10-fold GroupKFold (5-fold for CatBoost) ensuring all flights from the same search session stay together. This prevents leakage and enables proper behavioral feature creation.</p>\n<p><strong>Rationale</strong>: Flight ranking is inherently grouped by search session, and we needed to prevent any information leakage between folds while creating behavioral features.</p>\n<h3>Feature Engineering: Cross-Validation Aware Aggregations</h3>\n<p>The biggest breakthrough was creating behavioral features without group info leakage. Traditional aggregations would use future information, so I built a system that calculates features differently for each fold. There could be some temporal leak here, but i was okay with it.</p>\n<pre><code> ():\n    \n    \n     i  ():\n        df_train_fold = original_train_df.(pl.col() != i)\n        df_val_fold = original_train_df.(pl.col() == i)\n\n        \n         feature_name, config  agg_configs.items():\n            agg_df = df_train_fold.group_by(config[]).agg(\n                (pl.col(config[]), config[])()\n            )\n            df_val_fold = df_val_fold.join(agg_df, how=)\n</code></pre>\n<p><strong>Key Behavioral Features</strong>:</p>\n<ul>\n<li><code>avg_price_percentile_by_user</code>: User's historical price sensitivity</li>\n<li><code>company_price_discipline_rate</code>: Company policy strictness  </li>\n<li><code>user_min_segments_selection_rate</code>: Preference for direct flights</li>\n<li><code>user_route_frequency</code>: Route familiarity patterns</li>\n</ul>\n<h3>Ranking and Percentile Features</h3>\n<p>Heavy focus on competitive positioning within each search session:</p>\n<pre><code> ():\n    expressions = []\n     col  cols_to_percentile:\n        expressions.append(\n            ((pl.col(col).rank().over(group_col) - ) / \n             (pl.col(col).count().over(group_col) - ) * )\n            .alias()\n        )\n     df.with_columns(expressions)\n</code></pre>\n<p><strong>High-Impact Ranking Features</strong>:</p>\n<ul>\n<li><code>totalPrice_percentile_in_group</code>: Price position within search</li>\n<li><code>position_within_segment_tier</code>: Ranking within same complexity level</li>\n<li><code>segment_tier</code>: Flight complexity relative to minimum required</li>\n</ul>\n<h3>Segment Intelligence Features</h3>\n<p>Approach to modeling flight complexity:</p>\n<pre><code> ():\n    expressions = [\n        \n        (pl.col() - pl.col().().over() + )\n        .alias(),\n\n        \n        pl.col().rank().over([, ])\n        .alias(),\n\n        \n        (pl.col() - pl.col().().over())\n        .alias()\n    ]\n     df.with_columns(expressions)\n</code></pre>\n<h3>Multi-Algorithm Ensemble (13 Models)</h3>\n<table>\n<thead>\n<tr>\n<th>Algorithm</th>\n<th>Models</th>\n<th>Cross-Validation</th>\n<th>Key Parameters</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>XGBoost</td>\n<td>7</td>\n<td>10-fold</td>\n<td>max_pairs: 8,16,32,64,96,128</td>\n</tr>\n<tr>\n<td>LightGBM</td>\n<td>4</td>\n<td>10-fold</td>\n<td>max_pairs: 8,32,64,96</td>\n</tr>\n<tr>\n<td>CatBoost</td>\n<td>2</td>\n<td>5-fold</td>\n<td>max_pairs: 32,128</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Best Single Model</strong>: XGBoost with 8 pairs and 3500 iterations achieved 0.53807 private LB.</p>\n<h3>Bucket-Wise Ensemble Optimization</h3>\n<p><strong>The Core Innovation</strong>: Instead of global ensemble weights, optimize different weights for different group sizes.</p>\n<pre><code>\nbrackets = [\n    (, ), (,), (, ), (, ), (,), \n    (, ), (,), (, ), (, ), (, ), \n    (, ), (, ), (, ), (, ), \n    (, ), (, ), (,), (, ),\n    (, ), (, ), (, ), (,),\n    (, ), (, ), (, ())\n]\n\n\n ():\n     ():\n        weights = [trial.suggest_float(, , ) \n                   model  model_cols]\n\n        blended_preds = np.zeros((bracket_df))\n         i, w  (weights):\n            blended_preds += w * bracket_df[model_cols[i]]\n\n         hitrate_at_3(bracket_df[], blended_preds, \n                           bracket_df[])\n</code></pre>\n<p><strong>Softmax Normalization</strong>: Before optimization, all model predictions are normalized using softmax within each search session, ensuring proper scaling across different algorithms.</p>\n<p><strong>Why This Works</strong>: Users behave differently when choosing from 15 vs 150 options. Small choice sets could favor convenience features, while large choice sets could increase price sensitivity.</p>\n<h3>Ensemble Results</h3>\n<table>\n<thead>\n<tr>\n<th>Component</th>\n<th>Contribution</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Best Single Model</td>\n<td>0.53807</td>\n</tr>\n<tr>\n<td>Bucket-Wise Optimization</td>\n<td><strong>0.54175</strong></td>\n</tr>\n</tbody>\n</table>\n<p>The bucket-wise approach provided approximately +0.008 improvement over simple averaging, which basically means not too much lift and perhaps I should have ensembled with another time-split model.</p>\n<h2>Implementation Details</h2>\n<p><strong>Memory Optimization</strong>: Used Polars throughout for efficient processing of 300+ features across millions of rows.</p>\n<p><strong>Training Time</strong>: </p>\n<ul>\n<li>Feature engineering: ~25 mins per model run</li>\n<li>Cross-validation: ~45 minutes per fold (XGBoost/LightGBM), ~90 minutes (CatBoost)  </li>\n<li>Ensemble optimization: ~3 hours</li>\n<li>Total: ~12-15 hours</li>\n</ul>\n<h2>Feature Importance Insights</h2>\n<p><strong>Top 10 Features</strong>:</p>\n<ol>\n<li><code>segment_distance_from_minimum</code> - Flight complexity deviation</li>\n<li><code>is_min_segments</code> - Direct flight indicator  </li>\n<li><code>segment_tier</code> - Complexity tier within search</li>\n<li><code>policy_flexibility_interaction</code> - Company policy patterns</li>\n<li><code>segment_price_category</code> - Combined complexity-price positioning</li>\n<li><code>chose_minimum_segments</code> - Historical direct flight preference</li>\n<li><code>is_only_direct_option</code> - Exclusive direct flight availability</li>\n<li><code>segment_flexibility_score</code> - User convenience-price trade-off</li>\n<li><code>company_avg_segment_tier_for_route</code> - Company route policies</li>\n<li><code>price_premium_per_extra_segment</code> - Convenience cost analysis</li>\n</ol>\n<p>Behavioral aggregation features were crucial for capturing user and company preferences, while competitive positioning features handled within-search dynamics.</p>\n<h2>Key Learnings</h2>\n<ol>\n<li><strong>Bucket-wise optimization</strong> provided significant improvement over global ensemble weights</li>\n<li><strong>Cross-validation aware behavioral features</strong> prevented overfitting but most likely had time leak</li>\n<li><strong>Softmax normalization</strong> was crucial for combining different algorithms effectively</li>\n<li><strong>Segment intelligence</strong> was more important than traditional flight features</li>\n<li><strong>Extended training iterations</strong> sometimes beat complex hyperparameter tuning</li>\n<li><strong>Multi-algorithm diversity</strong> provided better results than single-algorithm optimization</li>\n</ol>\n<h2>Things that I missed or did not work</h2>\n<ol>\n<li><strong>Time-splits</strong> - I thought of doing time splits and even trained 1 model with most recent 3 months data, it scored Pvt LB: 0.529, in hindsight I am sure the ensemble would be better if i had blended with this result. But i did not have a way to CV validate this.</li>\n<li><strong>Psuedo Labels</strong> - I tired doing pseudo labelling as well, but it was hard, as most of the pseudo selected samples seemed to be ones with smaller group sizes, so i wasn't convinced if that would help at all.</li>\n</ol>\n<p>PS: Written with help of Claude so please pardon the hyperbole if there are any. :) </p>",
  "messages": [
    {
      "id": "3273553",
      "postDate": "08/22/2025 21:45:01",
      "content": "<p>I want to thank the organizers for creating such an engaging real-world ranking challenge. The flight recommendation domain provided rich opportunities to explore user behavior modeling and ensemble strategies.</p>\n<p>EDIT: Code is available at repo <a href=\"url\" target=\"_blank\">(https://github.com/Rajneesh-Tiwari/FlightRank-2025-Aeroclub-RecSys-Cup/</a></p>\n<p>I also want to acknowledge several excellent notebooks that inspired parts of this solution:</p>\n<ul>\n<li>The initial XGBoost ranker implementations (<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars</a>)  that provided a solid foundation</li>\n<li>Creative approaches to behavioral feature engineering that highlighted the importance of user history</li>\n<li>Ensemble strategies that demonstrated the value of model diversity</li>\n</ul>\n<h2>Final Results</h2>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Valid (CV)</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Best Single Model</td>\n<td>0.6412</td>\n<td>0.53807</td>\n</tr>\n<tr>\n<td>13-Model Ensemble</td>\n<td>0.6615</td>\n<td><strong>0.54175</strong></td>\n</tr>\n</tbody>\n</table>\n<h2>Approach Overview</h2>\n<p>My solution centers around a multi-algorithm ensemble with bucket-wise optimization, recognizing that user ranking behavior changes dramatically based on the number of flight options available in each search session. The core insight was that someone choosing from 15 options behaves very differently than someone choosing from 150 options.</p>\n<h2>Key Components</h2>\n<h3>Validation Strategy</h3>\n<p><strong>GroupKFold Cross-Validation</strong>: Used 10-fold GroupKFold (5-fold for CatBoost) ensuring all flights from the same search session stay together. This prevents leakage and enables proper behavioral feature creation.</p>\n<p><strong>Rationale</strong>: Flight ranking is inherently grouped by search session, and we needed to prevent any information leakage between folds while creating behavioral features.</p>\n<h3>Feature Engineering: Cross-Validation Aware Aggregations</h3>\n<p>The biggest breakthrough was creating behavioral features without group info leakage. Traditional aggregations would use future information, so I built a system that calculates features differently for each fold. There could be some temporal leak here, but i was okay with it.</p>\n<pre><code> ():\n    \n    \n     i  ():\n        df_train_fold = original_train_df.(pl.col() != i)\n        df_val_fold = original_train_df.(pl.col() == i)\n\n        \n         feature_name, config  agg_configs.items():\n            agg_df = df_train_fold.group_by(config[]).agg(\n                (pl.col(config[]), config[])()\n            )\n            df_val_fold = df_val_fold.join(agg_df, how=)\n</code></pre>\n<p><strong>Key Behavioral Features</strong>:</p>\n<ul>\n<li><code>avg_price_percentile_by_user</code>: User's historical price sensitivity</li>\n<li><code>company_price_discipline_rate</code>: Company policy strictness  </li>\n<li><code>user_min_segments_selection_rate</code>: Preference for direct flights</li>\n<li><code>user_route_frequency</code>: Route familiarity patterns</li>\n</ul>\n<h3>Ranking and Percentile Features</h3>\n<p>Heavy focus on competitive positioning within each search session:</p>\n<pre><code> ():\n    expressions = []\n     col  cols_to_percentile:\n        expressions.append(\n            ((pl.col(col).rank().over(group_col) - ) / \n             (pl.col(col).count().over(group_col) - ) * )\n            .alias()\n        )\n     df.with_columns(expressions)\n</code></pre>\n<p><strong>High-Impact Ranking Features</strong>:</p>\n<ul>\n<li><code>totalPrice_percentile_in_group</code>: Price position within search</li>\n<li><code>position_within_segment_tier</code>: Ranking within same complexity level</li>\n<li><code>segment_tier</code>: Flight complexity relative to minimum required</li>\n</ul>\n<h3>Segment Intelligence Features</h3>\n<p>Approach to modeling flight complexity:</p>\n<pre><code> ():\n    expressions = [\n        \n        (pl.col() - pl.col().().over() + )\n        .alias(),\n\n        \n        pl.col().rank().over([, ])\n        .alias(),\n\n        \n        (pl.col() - pl.col().().over())\n        .alias()\n    ]\n     df.with_columns(expressions)\n</code></pre>\n<h3>Multi-Algorithm Ensemble (13 Models)</h3>\n<table>\n<thead>\n<tr>\n<th>Algorithm</th>\n<th>Models</th>\n<th>Cross-Validation</th>\n<th>Key Parameters</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>XGBoost</td>\n<td>7</td>\n<td>10-fold</td>\n<td>max_pairs: 8,16,32,64,96,128</td>\n</tr>\n<tr>\n<td>LightGBM</td>\n<td>4</td>\n<td>10-fold</td>\n<td>max_pairs: 8,32,64,96</td>\n</tr>\n<tr>\n<td>CatBoost</td>\n<td>2</td>\n<td>5-fold</td>\n<td>max_pairs: 32,128</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Best Single Model</strong>: XGBoost with 8 pairs and 3500 iterations achieved 0.53807 private LB.</p>\n<h3>Bucket-Wise Ensemble Optimization</h3>\n<p><strong>The Core Innovation</strong>: Instead of global ensemble weights, optimize different weights for different group sizes.</p>\n<pre><code>\nbrackets = [\n    (, ), (,), (, ), (, ), (,), \n    (, ), (,), (, ), (, ), (, ), \n    (, ), (, ), (, ), (, ), \n    (, ), (, ), (,), (, ),\n    (, ), (, ), (, ), (,),\n    (, ), (, ), (, ())\n]\n\n\n ():\n     ():\n        weights = [trial.suggest_float(, , ) \n                   model  model_cols]\n\n        blended_preds = np.zeros((bracket_df))\n         i, w  (weights):\n            blended_preds += w * bracket_df[model_cols[i]]\n\n         hitrate_at_3(bracket_df[], blended_preds, \n                           bracket_df[])\n</code></pre>\n<p><strong>Softmax Normalization</strong>: Before optimization, all model predictions are normalized using softmax within each search session, ensuring proper scaling across different algorithms.</p>\n<p><strong>Why This Works</strong>: Users behave differently when choosing from 15 vs 150 options. Small choice sets could favor convenience features, while large choice sets could increase price sensitivity.</p>\n<h3>Ensemble Results</h3>\n<table>\n<thead>\n<tr>\n<th>Component</th>\n<th>Contribution</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Best Single Model</td>\n<td>0.53807</td>\n</tr>\n<tr>\n<td>Bucket-Wise Optimization</td>\n<td><strong>0.54175</strong></td>\n</tr>\n</tbody>\n</table>\n<p>The bucket-wise approach provided approximately +0.008 improvement over simple averaging, which basically means not too much lift and perhaps I should have ensembled with another time-split model.</p>\n<h2>Implementation Details</h2>\n<p><strong>Memory Optimization</strong>: Used Polars throughout for efficient processing of 300+ features across millions of rows.</p>\n<p><strong>Training Time</strong>: </p>\n<ul>\n<li>Feature engineering: ~25 mins per model run</li>\n<li>Cross-validation: ~45 minutes per fold (XGBoost/LightGBM), ~90 minutes (CatBoost)  </li>\n<li>Ensemble optimization: ~3 hours</li>\n<li>Total: ~12-15 hours</li>\n</ul>\n<h2>Feature Importance Insights</h2>\n<p><strong>Top 10 Features</strong>:</p>\n<ol>\n<li><code>segment_distance_from_minimum</code> - Flight complexity deviation</li>\n<li><code>is_min_segments</code> - Direct flight indicator  </li>\n<li><code>segment_tier</code> - Complexity tier within search</li>\n<li><code>policy_flexibility_interaction</code> - Company policy patterns</li>\n<li><code>segment_price_category</code> - Combined complexity-price positioning</li>\n<li><code>chose_minimum_segments</code> - Historical direct flight preference</li>\n<li><code>is_only_direct_option</code> - Exclusive direct flight availability</li>\n<li><code>segment_flexibility_score</code> - User convenience-price trade-off</li>\n<li><code>company_avg_segment_tier_for_route</code> - Company route policies</li>\n<li><code>price_premium_per_extra_segment</code> - Convenience cost analysis</li>\n</ol>\n<p>Behavioral aggregation features were crucial for capturing user and company preferences, while competitive positioning features handled within-search dynamics.</p>\n<h2>Key Learnings</h2>\n<ol>\n<li><strong>Bucket-wise optimization</strong> provided significant improvement over global ensemble weights</li>\n<li><strong>Cross-validation aware behavioral features</strong> prevented overfitting but most likely had time leak</li>\n<li><strong>Softmax normalization</strong> was crucial for combining different algorithms effectively</li>\n<li><strong>Segment intelligence</strong> was more important than traditional flight features</li>\n<li><strong>Extended training iterations</strong> sometimes beat complex hyperparameter tuning</li>\n<li><strong>Multi-algorithm diversity</strong> provided better results than single-algorithm optimization</li>\n</ol>\n<h2>Things that I missed or did not work</h2>\n<ol>\n<li><strong>Time-splits</strong> - I thought of doing time splits and even trained 1 model with most recent 3 months data, it scored Pvt LB: 0.529, in hindsight I am sure the ensemble would be better if i had blended with this result. But i did not have a way to CV validate this.</li>\n<li><strong>Psuedo Labels</strong> - I tired doing pseudo labelling as well, but it was hard, as most of the pseudo selected samples seemed to be ones with smaller group sizes, so i wasn't convinced if that would help at all.</li>\n</ol>\n<p>PS: Written with help of Claude so please pardon the hyperbole if there are any. :) </p>",
      "rawMarkdown": "I want to thank the organizers for creating such an engaging real-world ranking challenge. The flight recommendation domain provided rich opportunities to explore user behavior modeling and ensemble strategies.\n\nEDIT: Code is available at repo [(https://github.com/Rajneesh-Tiwari/FlightRank-2025-Aeroclub-RecSys-Cup/](url)\n\nI also want to acknowledge several excellent notebooks that inspired parts of this solution:\n- The initial XGBoost ranker implementations ([https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars](url))  that provided a solid foundation\n- Creative approaches to behavioral feature engineering that highlighted the importance of user history\n- Ensemble strategies that demonstrated the value of model diversity\n\n## Final Results\n\n| Model | Valid (CV) | Private LB |\n|-------|------------|------------|\n| Best Single Model | 0.6412 | 0.53807 |\n| 13-Model Ensemble | 0.6615 | **0.54175** |\n\n## Approach Overview\n\nMy solution centers around a multi-algorithm ensemble with bucket-wise optimization, recognizing that user ranking behavior changes dramatically based on the number of flight options available in each search session. The core insight was that someone choosing from 15 options behaves very differently than someone choosing from 150 options.\n\n## Key Components\n\n### Validation Strategy\n\n**GroupKFold Cross-Validation**: Used 10-fold GroupKFold (5-fold for CatBoost) ensuring all flights from the same search session stay together. This prevents leakage and enables proper behavioral feature creation.\n\n**Rationale**: Flight ranking is inherently grouped by search session, and we needed to prevent any information leakage between folds while creating behavioral features.\n\n### Feature Engineering: Cross-Validation Aware Aggregations\n\nThe biggest breakthrough was creating behavioral features without group info leakage. Traditional aggregations would use future information, so I built a system that calculates features differently for each fold. There could be some temporal leak here, but i was okay with it.\n\n```python\ndef create_cv_aware_aggregate_features_no_leakage(\n    train_df, test_df, agg_configs\n):\n    # For each fold, calculate behavioral features using only \n    # training data from other folds\n    for i in range(10):\n        df_train_fold = original_train_df.filter(pl.col('fold') != i)\n        df_val_fold = original_train_df.filter(pl.col('fold') == i)\n        \n        # Calculate user/company behavior on training data only\n        for feature_name, config in agg_configs.items():\n            agg_df = df_train_fold.group_by(config['group_by']).agg(\n                getattr(pl.col(config['agg_col']), config['agg_func'])()\n            )\n            df_val_fold = df_val_fold.join(agg_df, how='left')\n```\n\n**Key Behavioral Features**:\n- `avg_price_percentile_by_user`: User's historical price sensitivity\n- `company_price_discipline_rate`: Company policy strictness  \n- `user_min_segments_selection_rate`: Preference for direct flights\n- `user_route_frequency`: Route familiarity patterns\n\n### Ranking and Percentile Features\n\nHeavy focus on competitive positioning within each search session:\n\n```python\ndef get_percentile_features(df, cols_to_percentile, group_col='ranker_id'):\n    expressions = []\n    for col in cols_to_percentile:\n        expressions.append(\n            ((pl.col(col).rank(\"ordinal\").over(group_col) - 1.0) / \n             (pl.col(col).count().over(group_col) - 1.0) * 100.0)\n            .alias(f'{col}_percentile_in_group')\n        )\n    return df.with_columns(expressions)\n```\n\n**High-Impact Ranking Features**:\n- `totalPrice_percentile_in_group`: Price position within search\n- `position_within_segment_tier`: Ranking within same complexity level\n- `segment_tier`: Flight complexity relative to minimum required\n\n### Segment Intelligence Features\n\nApproach to modeling flight complexity:\n\n```python\ndef create_segment_tier_position_features(df):\n    expressions = [\n        # Segment tier (1=minimum, 2=min+1, etc.)\n        (pl.col('total_segments') - pl.col('total_segments').min().over('ranker_id') + 1)\n        .alias('segment_tier'),\n        \n        # Position within same complexity level\n        pl.col('totalPrice').rank('ordinal').over(['ranker_id', 'total_segments'])\n        .alias('position_within_segment_tier'),\n        \n        # Price premium for convenience\n        (pl.col('totalPrice') - pl.col('totalPrice').min().over('ranker_id'))\n        .alias('price_premium_vs_min_segments')\n    ]\n    return df.with_columns(expressions)\n```\n\n### Multi-Algorithm Ensemble (13 Models)\n\n| Algorithm | Models | Cross-Validation | Key Parameters |\n|-----------|--------|------------------|----------------|\n| XGBoost | 7 | 10-fold | max_pairs: 8,16,32,64,96,128 |\n| LightGBM | 4 | 10-fold |max_pairs: 8,32,64,96 |\n| CatBoost | 2 | 5-fold | max_pairs: 32,128 |\n\n**Best Single Model**: XGBoost with 8 pairs and 3500 iterations achieved 0.53807 private LB.\n\n### Bucket-Wise Ensemble Optimization\n\n**The Core Innovation**: Instead of global ensemble weights, optimize different weights for different group sizes.\n\n```python\n# 25 buckets based on search session size\nbrackets = [\n    (10, 14), (14,18), (18, 22), (22, 27), (27,32), \n    (32, 37), (37,42), (42, 48), (48, 56), (56, 66), \n    (66, 78), (78, 98), (98, 120), (120, 155), \n    (155, 205), (205, 280), (280,361), (361, 501),\n    (501, 625), (625, 750), (750, 1001), (1001,1500),\n    (1500, 2501), (2501, 3500), (3500, float('inf'))\n]\n\n# Optuna optimization with 1500 trials per bucket\ndef find_optimal_weights_optuna(bracket_df, model_cols, n_trials=1500):\n    def objective(trial):\n        weights = [trial.suggest_float(f'weight_{model}', 0.0, 1.0) \n                  for model in model_cols]\n        \n        blended_preds = np.zeros(len(bracket_df))\n        for i, w in enumerate(weights):\n            blended_preds += w * bracket_df[model_cols[i]]\n            \n        return hitrate_at_3(bracket_df['selected'], blended_preds, \n                           bracket_df['group_key'])\n```\n\n**Softmax Normalization**: Before optimization, all model predictions are normalized using softmax within each search session, ensuring proper scaling across different algorithms.\n\n**Why This Works**: Users behave differently when choosing from 15 vs 150 options. Small choice sets could favor convenience features, while large choice sets could increase price sensitivity.\n\n### Ensemble Results\n\n| Component | Contribution |\n|-----------|-------------|\n| Best Single Model | 0.53807 |\n| Bucket-Wise Optimization | **0.54175** |\n\nThe bucket-wise approach provided approximately +0.008 improvement over simple averaging, which basically means not too much lift and perhaps I should have ensembled with another time-split model.\n\n## Implementation Details\n\n**Memory Optimization**: Used Polars throughout for efficient processing of 300+ features across millions of rows.\n\n**Training Time**: \n- Feature engineering: ~25 mins per model run\n- Cross-validation: ~45 minutes per fold (XGBoost/LightGBM), ~90 minutes (CatBoost)  \n- Ensemble optimization: ~3 hours\n- Total: ~12-15 hours\n\n## Feature Importance Insights\n\n**Top 10 Features**:\n1. `segment_distance_from_minimum` - Flight complexity deviation\n2. `is_min_segments` - Direct flight indicator  \n3. `segment_tier` - Complexity tier within search\n4. `policy_flexibility_interaction` - Company policy patterns\n5. `segment_price_category` - Combined complexity-price positioning\n6. `chose_minimum_segments` - Historical direct flight preference\n7. `is_only_direct_option` - Exclusive direct flight availability\n8. `segment_flexibility_score` - User convenience-price trade-off\n9. `company_avg_segment_tier_for_route` - Company route policies\n10. `price_premium_per_extra_segment` - Convenience cost analysis\n\nBehavioral aggregation features were crucial for capturing user and company preferences, while competitive positioning features handled within-search dynamics.\n\n## Key Learnings\n\n1. **Bucket-wise optimization** provided significant improvement over global ensemble weights\n2. **Cross-validation aware behavioral features** prevented overfitting but most likely had time leak\n3. **Softmax normalization** was crucial for combining different algorithms effectively\n4. **Segment intelligence** was more important than traditional flight features\n5. **Extended training iterations** sometimes beat complex hyperparameter tuning\n6. **Multi-algorithm diversity** provided better results than single-algorithm optimization\n\n## Things that I missed or did not work\n1. **Time-splits** - I thought of doing time splits and even trained 1 model with most recent 3 months data, it scored Pvt LB: 0.529, in hindsight I am sure the ensemble would be better if i had blended with this result. But i did not have a way to CV validate this.\n2. **Psuedo Labels** - I tired doing pseudo labelling as well, but it was hard, as most of the pseudo selected samples seemed to be ones with smaller group sizes, so i wasn't convinced if that would help at all.\n\nPS: Written with help of Claude so please pardon the hyperbole if there are any. :)",
      "votes": null
    },
    {
      "id": "3273646",
      "postDate": "08/23/2025 05:09:26",
      "content": "<p>thank you for sharing your knowledge!</p>",
      "rawMarkdown": "thank you for sharing your knowledge!",
      "votes": null
    },
    {
      "id": "3273647",
      "postDate": "08/23/2025 05:09:59",
      "content": "<p>is the code avaliable somewhere?</p>",
      "rawMarkdown": "is the code avaliable somewhere?",
      "votes": null
    },
    {
      "id": "3273679",
      "postDate": "08/23/2025 06:35:45",
      "content": "<p>Yes. Refining a few things. Will update writeup with github in a bit.</p>",
      "rawMarkdown": "Yes. Refining a few things. Will update writeup with github in a bit.",
      "votes": null
    },
    {
      "id": "3274047",
      "postDate": "08/23/2025 21:07:39",
      "content": "<p>I think a lot of people will learn from this, thank you!</p>",
      "rawMarkdown": "I think a lot of people will learn from this, thank you!",
      "votes": null
    },
    {
      "id": "3274049",
      "postDate": "08/23/2025 21:15:00",
      "content": "<p>Thanks a lot !!</p>",
      "rawMarkdown": "Thanks a lot !!",
      "votes": null
    },
    {
      "id": "3274126",
      "postDate": "08/24/2025 05:04:32",
      "content": "<p>Code repo added. </p>",
      "rawMarkdown": "Code repo added.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3273646,
      "author_name": "junhaochan",
      "author_url": "",
      "post_date": "08/23/2025 05:09:26",
      "content": "<p>thank you for sharing your knowledge!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3273647,
      "author_name": "junhaochan",
      "author_url": "",
      "post_date": "08/23/2025 05:09:59",
      "content": "<p>is the code avaliable somewhere?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3273679,
          "author_name": "pheadrus",
          "author_url": "",
          "post_date": "08/23/2025 06:35:45",
          "content": "<p>Yes. Refining a few things. Will update writeup with github in a bit.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3274126,
              "author_name": "pheadrus",
              "author_url": "",
              "post_date": "08/24/2025 05:04:32",
              "content": "<p>Code repo added. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3274047,
      "author_name": "irakozekelly",
      "author_url": "",
      "post_date": "08/23/2025 21:07:39",
      "content": "<p>I think a lot of people will learn from this, thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3274049,
      "author_name": "abdelhakouanzougui2",
      "author_url": "",
      "post_date": "08/23/2025 21:15:00",
      "content": "<p>Thanks a lot !!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3273553": "I want to thank the organizers for creating such an engaging real-world ranking challenge. The flight recommendation domain provided rich opportunities to explore user behavior modeling and ensemble strategies.\n\nEDIT: Code is available at repo [(https://github.com/Rajneesh-Tiwari/FlightRank-2025-Aeroclub-RecSys-Cup/](url)\n\nI also want to acknowledge several excellent notebooks that inspired parts of this solution:\n- The initial XGBoost ranker implementations ([https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars](url))  that provided a solid foundation\n- Creative approaches to behavioral feature engineering that highlighted the importance of user history\n- Ensemble strategies that demonstrated the value of model diversity\n\n## Final Results\n\n| Model | Valid (CV) | Private LB |\n|-------|------------|------------|\n| Best Single Model | 0.6412 | 0.53807 |\n| 13-Model Ensemble | 0.6615 | **0.54175** |\n\n## Approach Overview\n\nMy solution centers around a multi-algorithm ensemble with bucket-wise optimization, recognizing that user ranking behavior changes dramatically based on the number of flight options available in each search session. The core insight was that someone choosing from 15 options behaves very differently than someone choosing from 150 options.\n\n## Key Components\n\n### Validation Strategy\n\n**GroupKFold Cross-Validation**: Used 10-fold GroupKFold (5-fold for CatBoost) ensuring all flights from the same search session stay together. This prevents leakage and enables proper behavioral feature creation.\n\n**Rationale**: Flight ranking is inherently grouped by search session, and we needed to prevent any information leakage between folds while creating behavioral features.\n\n### Feature Engineering: Cross-Validation Aware Aggregations\n\nThe biggest breakthrough was creating behavioral features without group info leakage. Traditional aggregations would use future information, so I built a system that calculates features differently for each fold. There could be some temporal leak here, but i was okay with it.\n\n```python\ndef create_cv_aware_aggregate_features_no_leakage(\n    train_df, test_df, agg_configs\n):\n    # For each fold, calculate behavioral features using only \n    # training data from other folds\n    for i in range(10):\n        df_train_fold = original_train_df.filter(pl.col('fold') != i)\n        df_val_fold = original_train_df.filter(pl.col('fold') == i)\n        \n        # Calculate user/company behavior on training data only\n        for feature_name, config in agg_configs.items():\n            agg_df = df_train_fold.group_by(config['group_by']).agg(\n                getattr(pl.col(config['agg_col']), config['agg_func'])()\n            )\n            df_val_fold = df_val_fold.join(agg_df, how='left')\n```\n\n**Key Behavioral Features**:\n- `avg_price_percentile_by_user`: User's historical price sensitivity\n- `company_price_discipline_rate`: Company policy strictness  \n- `user_min_segments_selection_rate`: Preference for direct flights\n- `user_route_frequency`: Route familiarity patterns\n\n### Ranking and Percentile Features\n\nHeavy focus on competitive positioning within each search session:\n\n```python\ndef get_percentile_features(df, cols_to_percentile, group_col='ranker_id'):\n    expressions = []\n    for col in cols_to_percentile:\n        expressions.append(\n            ((pl.col(col).rank(\"ordinal\").over(group_col) - 1.0) / \n             (pl.col(col).count().over(group_col) - 1.0) * 100.0)\n            .alias(f'{col}_percentile_in_group')\n        )\n    return df.with_columns(expressions)\n```\n\n**High-Impact Ranking Features**:\n- `totalPrice_percentile_in_group`: Price position within search\n- `position_within_segment_tier`: Ranking within same complexity level\n- `segment_tier`: Flight complexity relative to minimum required\n\n### Segment Intelligence Features\n\nApproach to modeling flight complexity:\n\n```python\ndef create_segment_tier_position_features(df):\n    expressions = [\n        # Segment tier (1=minimum, 2=min+1, etc.)\n        (pl.col('total_segments') - pl.col('total_segments').min().over('ranker_id') + 1)\n        .alias('segment_tier'),\n        \n        # Position within same complexity level\n        pl.col('totalPrice').rank('ordinal').over(['ranker_id', 'total_segments'])\n        .alias('position_within_segment_tier'),\n        \n        # Price premium for convenience\n        (pl.col('totalPrice') - pl.col('totalPrice').min().over('ranker_id'))\n        .alias('price_premium_vs_min_segments')\n    ]\n    return df.with_columns(expressions)\n```\n\n### Multi-Algorithm Ensemble (13 Models)\n\n| Algorithm | Models | Cross-Validation | Key Parameters |\n|-----------|--------|------------------|----------------|\n| XGBoost | 7 | 10-fold | max_pairs: 8,16,32,64,96,128 |\n| LightGBM | 4 | 10-fold |max_pairs: 8,32,64,96 |\n| CatBoost | 2 | 5-fold | max_pairs: 32,128 |\n\n**Best Single Model**: XGBoost with 8 pairs and 3500 iterations achieved 0.53807 private LB.\n\n### Bucket-Wise Ensemble Optimization\n\n**The Core Innovation**: Instead of global ensemble weights, optimize different weights for different group sizes.\n\n```python\n# 25 buckets based on search session size\nbrackets = [\n    (10, 14), (14,18), (18, 22), (22, 27), (27,32), \n    (32, 37), (37,42), (42, 48), (48, 56), (56, 66), \n    (66, 78), (78, 98), (98, 120), (120, 155), \n    (155, 205), (205, 280), (280,361), (361, 501),\n    (501, 625), (625, 750), (750, 1001), (1001,1500),\n    (1500, 2501), (2501, 3500), (3500, float('inf'))\n]\n\n# Optuna optimization with 1500 trials per bucket\ndef find_optimal_weights_optuna(bracket_df, model_cols, n_trials=1500):\n    def objective(trial):\n        weights = [trial.suggest_float(f'weight_{model}', 0.0, 1.0) \n                  for model in model_cols]\n        \n        blended_preds = np.zeros(len(bracket_df))\n        for i, w in enumerate(weights):\n            blended_preds += w * bracket_df[model_cols[i]]\n            \n        return hitrate_at_3(bracket_df['selected'], blended_preds, \n                           bracket_df['group_key'])\n```\n\n**Softmax Normalization**: Before optimization, all model predictions are normalized using softmax within each search session, ensuring proper scaling across different algorithms.\n\n**Why This Works**: Users behave differently when choosing from 15 vs 150 options. Small choice sets could favor convenience features, while large choice sets could increase price sensitivity.\n\n### Ensemble Results\n\n| Component | Contribution |\n|-----------|-------------|\n| Best Single Model | 0.53807 |\n| Bucket-Wise Optimization | **0.54175** |\n\nThe bucket-wise approach provided approximately +0.008 improvement over simple averaging, which basically means not too much lift and perhaps I should have ensembled with another time-split model.\n\n## Implementation Details\n\n**Memory Optimization**: Used Polars throughout for efficient processing of 300+ features across millions of rows.\n\n**Training Time**: \n- Feature engineering: ~25 mins per model run\n- Cross-validation: ~45 minutes per fold (XGBoost/LightGBM), ~90 minutes (CatBoost)  \n- Ensemble optimization: ~3 hours\n- Total: ~12-15 hours\n\n## Feature Importance Insights\n\n**Top 10 Features**:\n1. `segment_distance_from_minimum` - Flight complexity deviation\n2. `is_min_segments` - Direct flight indicator  \n3. `segment_tier` - Complexity tier within search\n4. `policy_flexibility_interaction` - Company policy patterns\n5. `segment_price_category` - Combined complexity-price positioning\n6. `chose_minimum_segments` - Historical direct flight preference\n7. `is_only_direct_option` - Exclusive direct flight availability\n8. `segment_flexibility_score` - User convenience-price trade-off\n9. `company_avg_segment_tier_for_route` - Company route policies\n10. `price_premium_per_extra_segment` - Convenience cost analysis\n\nBehavioral aggregation features were crucial for capturing user and company preferences, while competitive positioning features handled within-search dynamics.\n\n## Key Learnings\n\n1. **Bucket-wise optimization** provided significant improvement over global ensemble weights\n2. **Cross-validation aware behavioral features** prevented overfitting but most likely had time leak\n3. **Softmax normalization** was crucial for combining different algorithms effectively\n4. **Segment intelligence** was more important than traditional flight features\n5. **Extended training iterations** sometimes beat complex hyperparameter tuning\n6. **Multi-algorithm diversity** provided better results than single-algorithm optimization\n\n## Things that I missed or did not work\n1. **Time-splits** - I thought of doing time splits and even trained 1 model with most recent 3 months data, it scored Pvt LB: 0.529, in hindsight I am sure the ensemble would be better if i had blended with this result. But i did not have a way to CV validate this.\n2. **Psuedo Labels** - I tired doing pseudo labelling as well, but it was hard, as most of the pseudo selected samples seemed to be ones with smaller group sizes, so i wasn't convinced if that would help at all.\n\nPS: Written with help of Claude so please pardon the hyperbole if there are any. :)",
    "3273646": "thank you for sharing your knowledge!",
    "3273647": "is the code avaliable somewhere?",
    "3273679": "Yes. Refining a few things. Will update writeup with github in a bit.",
    "3274047": "I think a lot of people will learn from this, thank you!",
    "3274049": "Thanks a lot !!",
    "3274126": "Code repo added."
  },
  "source": "meta"
}