{
  "id": 599750,
  "title": "3rd place solution",
  "url": "/competitions/aeroclub-recsys-2025/writeups/3rd-place-solution",
  "author_name": "",
  "post_date": "2025-08-24T05:29:00.500Z",
  "votes": 9,
  "comment_count": 2,
  "views": 0,
  "content": "<h1>Machine Learning Competition Solution</h1>\n<h2>Overview</h2>\n<p>We would like to thank the organizers for creating such an engaging and challenging competition. This document outlines our comprehensive approach to solving the flight ranking problem.</p>\n<p>Our solution began with the foundation of a public notebook (<a href=\"https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars\" target=\"_blank\">XGBoost Ranker</a>), though by the competition's end, our implementation had evolved significantly beyond the original approach.</p>\n<h2>Feature Engineering Strategy</h2>\n<p>Our feature engineering process was methodical and multi-faceted, focusing on capturing the complexity of flight booking behavior and preferences.</p>\n<h3>1. Core Flight Features</h3>\n<p><strong>Route Analysis:</strong></p>\n<ul>\n<li>Extracted segment counts and connection details for each flight leg</li>\n<li>Constructed complete flight paths (origin → connections → destination)</li>\n<li>Analyzed route structure and categorized connection types</li>\n</ul>\n<p><strong>Temporal Features:</strong></p>\n<ul>\n<li>Converted departure and arrival times to minutes for consistent processing</li>\n<li>Extracted time-based patterns including time of day, day of week, and seasonality</li>\n<li>Identified holidays and weekends to capture booking pattern variations</li>\n<li>Analyzed business-friendly timing preferences</li>\n</ul>\n<h3>2. Geographic &amp; Airport Features</h3>\n<p><strong>Distance Calculations:</strong></p>\n<ul>\n<li>Implemented Haversine formula for precise inter-airport distance measurements</li>\n<li>Calculated elevation differences between airports</li>\n<li>Identified intercontinental flights for specialized handling</li>\n</ul>\n<p><strong>Airport Characteristics:</strong></p>\n<ul>\n<li>Classified airports as hub versus regional facilities</li>\n<li>Analyzed route popularity and traffic patterns</li>\n<li>Implemented continental grouping for geographic insights</li>\n</ul>\n<h3>3. Advanced User Behavior Features</h3>\n<p><strong>Historical Patterns:</strong></p>\n<ul>\n<li>Analyzed user preferences across airlines, booking timing, and service classes</li>\n<li>Incorporated loyalty program analysis</li>\n<li>Evaluated corporate policy compliance patterns</li>\n</ul>\n<p><strong>Contextual Analysis:</strong></p>\n<ul>\n<li>Calculated relative rankings within search sessions (price, duration, convenience)</li>\n<li>Identified Pareto-optimal options in the choice set</li>\n<li>Analyzed price-quality trade-off preferences</li>\n</ul>\n<h3>4. Flight Quality &amp; Convenience Metrics</h3>\n<p><strong>Quality Scoring:</strong></p>\n<ul>\n<li>Mapped aircraft quality characteristics (wide-body vs narrow-body)</li>\n<li>Assessed schedule convenience (business hours vs red-eye flights)</li>\n<li>Analyzed fare flexibility options</li>\n<li>Evaluated connection time comfort levels</li>\n</ul>\n<p><strong>Pricing Analytics:</strong></p>\n<ul>\n<li>Calculated price-per-kilometer and price-per-minute efficiency ratios</li>\n<li>Measured deviation from group medians</li>\n<li>Analyzed competitive positioning within search results</li>\n</ul>\n<h3>5. Temporal Split &amp; Historical Counters</h3>\n<p>Following successful strategies employed by other competition winners, we implemented historical counters with multiple categorical feature combinations. Our approach involved:</p>\n<ul>\n<li>Using 50% of the time-sorted training dataset as historical reference</li>\n<li>Training on the remaining 50% to maintain temporal validity</li>\n<li>Creating interaction features across multiple categorical dimensions</li>\n</ul>\n<h3>6. Data Sampling Strategy</h3>\n<p>To ensure training stability and computational efficiency:</p>\n<ul>\n<li>Limited group sizes to 64 samples maximum</li>\n<li>Preserved all positive examples (selected=1) to maintain signal strength</li>\n<li>Applied random sampling to negative examples within each group</li>\n</ul>\n<h2>Model Architecture</h2>\n<h3>7. Model Training</h3>\n<p>Our ensemble approach incorporated multiple ranking algorithms:</p>\n<p><strong>Primary Models:</strong></p>\n<ul>\n<li><strong>CatBoostRanker</strong> with three different loss functions:<ul>\n<li>QueryCrossEntropy (alpha=1.0) - primary loss function</li>\n<li>PairLogitPairwise - secondary approach</li>\n<li>YetiRankPairwise - tertiary validation</li></ul></li>\n<li><strong>LightGBM Ranker</strong> implemented by Farukcan</li>\n</ul>\n<h3>8. Ensemble Strategy</h3>\n<p>Our final prediction system employed a normalized score ensemble:</p>\n<ul>\n<li>Normalized all individual ranker scores to ensure comparable scales</li>\n<li>Combined normalized scores using optimized ratios</li>\n<li>Selected ensemble weights primarily through leaderboard validation</li>\n<li>Maintained simplicity in the ensemble approach for robust generalization</li>\n</ul>\n<h2>Key Insights</h2>\n<p>The success of our solution stemmed from several critical discoveries:</p>\n<ol>\n<li><p><strong>Price and Duration Features Were Paramount</strong>: As expected, flight price and duration emerged as the most important predictive features. We experimented with multiple variations of these features including:</p>\n<ul>\n<li>Absolute values (raw price and duration)</li>\n<li>Relative measures (compared to group averages)</li>\n<li>Ranked positions within search results</li></ul>\n<p>This comprehensive treatment of the core decision factors proved essential for model performance.</p></li>\n<li><p><strong>Historical Counters Provided Substantial Gains</strong>: Historical counters delivered significant performance improvements, contributing approximately 200 features to our final dataset. However, this represents only a fraction of the possible categorical combinations available in the data. There remains substantial opportunity for further score improvements by exploring additional feature combinations, though computational constraints limited our exploration.</p></li>\n<li><p><strong>Multi-Loss Function Ensemble Benefits</strong>: A key discovery was that combining models trained with different loss functions, even when using identical training data, yielded meaningful performance gains. This suggests that different loss functions capture complementary aspects of the ranking problem, making ensemble diversity valuable beyond just data variation.</p></li>\n</ol>\n<p>Thanks for attention</p>",
  "messages": [
    {
      "id": "3271368",
      "postDate": "08/18/2025 17:25:27",
      "content": "<h1>Machine Learning Competition Solution</h1>\n<h2>Overview</h2>\n<p>We would like to thank the organizers for creating such an engaging and challenging competition. This document outlines our comprehensive approach to solving the flight ranking problem.</p>\n<p>Our solution began with the foundation of a public notebook (<a href=\"https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars\" target=\"_blank\">XGBoost Ranker</a>), though by the competition's end, our implementation had evolved significantly beyond the original approach.</p>\n<h2>Feature Engineering Strategy</h2>\n<p>Our feature engineering process was methodical and multi-faceted, focusing on capturing the complexity of flight booking behavior and preferences.</p>\n<h3>1. Core Flight Features</h3>\n<p><strong>Route Analysis:</strong></p>\n<ul>\n<li>Extracted segment counts and connection details for each flight leg</li>\n<li>Constructed complete flight paths (origin → connections → destination)</li>\n<li>Analyzed route structure and categorized connection types</li>\n</ul>\n<p><strong>Temporal Features:</strong></p>\n<ul>\n<li>Converted departure and arrival times to minutes for consistent processing</li>\n<li>Extracted time-based patterns including time of day, day of week, and seasonality</li>\n<li>Identified holidays and weekends to capture booking pattern variations</li>\n<li>Analyzed business-friendly timing preferences</li>\n</ul>\n<h3>2. Geographic &amp; Airport Features</h3>\n<p><strong>Distance Calculations:</strong></p>\n<ul>\n<li>Implemented Haversine formula for precise inter-airport distance measurements</li>\n<li>Calculated elevation differences between airports</li>\n<li>Identified intercontinental flights for specialized handling</li>\n</ul>\n<p><strong>Airport Characteristics:</strong></p>\n<ul>\n<li>Classified airports as hub versus regional facilities</li>\n<li>Analyzed route popularity and traffic patterns</li>\n<li>Implemented continental grouping for geographic insights</li>\n</ul>\n<h3>3. Advanced User Behavior Features</h3>\n<p><strong>Historical Patterns:</strong></p>\n<ul>\n<li>Analyzed user preferences across airlines, booking timing, and service classes</li>\n<li>Incorporated loyalty program analysis</li>\n<li>Evaluated corporate policy compliance patterns</li>\n</ul>\n<p><strong>Contextual Analysis:</strong></p>\n<ul>\n<li>Calculated relative rankings within search sessions (price, duration, convenience)</li>\n<li>Identified Pareto-optimal options in the choice set</li>\n<li>Analyzed price-quality trade-off preferences</li>\n</ul>\n<h3>4. Flight Quality &amp; Convenience Metrics</h3>\n<p><strong>Quality Scoring:</strong></p>\n<ul>\n<li>Mapped aircraft quality characteristics (wide-body vs narrow-body)</li>\n<li>Assessed schedule convenience (business hours vs red-eye flights)</li>\n<li>Analyzed fare flexibility options</li>\n<li>Evaluated connection time comfort levels</li>\n</ul>\n<p><strong>Pricing Analytics:</strong></p>\n<ul>\n<li>Calculated price-per-kilometer and price-per-minute efficiency ratios</li>\n<li>Measured deviation from group medians</li>\n<li>Analyzed competitive positioning within search results</li>\n</ul>\n<h3>5. Temporal Split &amp; Historical Counters</h3>\n<p>Following successful strategies employed by other competition winners, we implemented historical counters with multiple categorical feature combinations. Our approach involved:</p>\n<ul>\n<li>Using 50% of the time-sorted training dataset as historical reference</li>\n<li>Training on the remaining 50% to maintain temporal validity</li>\n<li>Creating interaction features across multiple categorical dimensions</li>\n</ul>\n<h3>6. Data Sampling Strategy</h3>\n<p>To ensure training stability and computational efficiency:</p>\n<ul>\n<li>Limited group sizes to 64 samples maximum</li>\n<li>Preserved all positive examples (selected=1) to maintain signal strength</li>\n<li>Applied random sampling to negative examples within each group</li>\n</ul>\n<h2>Model Architecture</h2>\n<h3>7. Model Training</h3>\n<p>Our ensemble approach incorporated multiple ranking algorithms:</p>\n<p><strong>Primary Models:</strong></p>\n<ul>\n<li><strong>CatBoostRanker</strong> with three different loss functions:<ul>\n<li>QueryCrossEntropy (alpha=1.0) - primary loss function</li>\n<li>PairLogitPairwise - secondary approach</li>\n<li>YetiRankPairwise - tertiary validation</li></ul></li>\n<li><strong>LightGBM Ranker</strong> implemented by Farukcan</li>\n</ul>\n<h3>8. Ensemble Strategy</h3>\n<p>Our final prediction system employed a normalized score ensemble:</p>\n<ul>\n<li>Normalized all individual ranker scores to ensure comparable scales</li>\n<li>Combined normalized scores using optimized ratios</li>\n<li>Selected ensemble weights primarily through leaderboard validation</li>\n<li>Maintained simplicity in the ensemble approach for robust generalization</li>\n</ul>\n<h2>Key Insights</h2>\n<p>The success of our solution stemmed from several critical discoveries:</p>\n<ol>\n<li><p><strong>Price and Duration Features Were Paramount</strong>: As expected, flight price and duration emerged as the most important predictive features. We experimented with multiple variations of these features including:</p>\n<ul>\n<li>Absolute values (raw price and duration)</li>\n<li>Relative measures (compared to group averages)</li>\n<li>Ranked positions within search results</li></ul>\n<p>This comprehensive treatment of the core decision factors proved essential for model performance.</p></li>\n<li><p><strong>Historical Counters Provided Substantial Gains</strong>: Historical counters delivered significant performance improvements, contributing approximately 200 features to our final dataset. However, this represents only a fraction of the possible categorical combinations available in the data. There remains substantial opportunity for further score improvements by exploring additional feature combinations, though computational constraints limited our exploration.</p></li>\n<li><p><strong>Multi-Loss Function Ensemble Benefits</strong>: A key discovery was that combining models trained with different loss functions, even when using identical training data, yielded meaningful performance gains. This suggests that different loss functions capture complementary aspects of the ranking problem, making ensemble diversity valuable beyond just data variation.</p></li>\n</ol>\n<p>Thanks for attention</p>",
      "rawMarkdown": "# Machine Learning Competition Solution\n\n## Overview\n\nWe would like to thank the organizers for creating such an engaging and challenging competition. This document outlines our comprehensive approach to solving the flight ranking problem.\n\nOur solution began with the foundation of a public notebook ([XGBoost Ranker](https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars)), though by the competition's end, our implementation had evolved significantly beyond the original approach.\n\n## Feature Engineering Strategy\n\nOur feature engineering process was methodical and multi-faceted, focusing on capturing the complexity of flight booking behavior and preferences.\n\n### 1. Core Flight Features\n\n**Route Analysis:**\n- Extracted segment counts and connection details for each flight leg\n- Constructed complete flight paths (origin → connections → destination)\n- Analyzed route structure and categorized connection types\n\n**Temporal Features:**\n- Converted departure and arrival times to minutes for consistent processing\n- Extracted time-based patterns including time of day, day of week, and seasonality\n- Identified holidays and weekends to capture booking pattern variations\n- Analyzed business-friendly timing preferences\n\n### 2. Geographic & Airport Features\n\n**Distance Calculations:**\n- Implemented Haversine formula for precise inter-airport distance measurements\n- Calculated elevation differences between airports\n- Identified intercontinental flights for specialized handling\n\n**Airport Characteristics:**\n- Classified airports as hub versus regional facilities\n- Analyzed route popularity and traffic patterns\n- Implemented continental grouping for geographic insights\n\n### 3. Advanced User Behavior Features\n\n**Historical Patterns:**\n- Analyzed user preferences across airlines, booking timing, and service classes\n- Incorporated loyalty program analysis\n- Evaluated corporate policy compliance patterns\n\n**Contextual Analysis:**\n- Calculated relative rankings within search sessions (price, duration, convenience)\n- Identified Pareto-optimal options in the choice set\n- Analyzed price-quality trade-off preferences\n\n### 4. Flight Quality & Convenience Metrics\n\n**Quality Scoring:**\n- Mapped aircraft quality characteristics (wide-body vs narrow-body)\n- Assessed schedule convenience (business hours vs red-eye flights)\n- Analyzed fare flexibility options\n- Evaluated connection time comfort levels\n\n**Pricing Analytics:**\n- Calculated price-per-kilometer and price-per-minute efficiency ratios\n- Measured deviation from group medians\n- Analyzed competitive positioning within search results\n\n### 5. Temporal Split & Historical Counters\n\nFollowing successful strategies employed by other competition winners, we implemented historical counters with multiple categorical feature combinations. Our approach involved:\n\n- Using 50% of the time-sorted training dataset as historical reference\n- Training on the remaining 50% to maintain temporal validity\n- Creating interaction features across multiple categorical dimensions\n\n### 6. Data Sampling Strategy\n\nTo ensure training stability and computational efficiency:\n\n- Limited group sizes to 64 samples maximum\n- Preserved all positive examples (selected=1) to maintain signal strength\n- Applied random sampling to negative examples within each group\n\n## Model Architecture\n\n### 7. Model Training\n\nOur ensemble approach incorporated multiple ranking algorithms:\n\n**Primary Models:**\n- **CatBoostRanker** with three different loss functions:\n  - QueryCrossEntropy (alpha=1.0) - primary loss function\n  - PairLogitPairwise - secondary approach\n  - YetiRankPairwise - tertiary validation\n- **LightGBM Ranker** implemented by Farukcan\n\n### 8. Ensemble Strategy\n\nOur final prediction system employed a normalized score ensemble:\n\n- Normalized all individual ranker scores to ensure comparable scales\n- Combined normalized scores using optimized ratios\n- Selected ensemble weights primarily through leaderboard validation\n- Maintained simplicity in the ensemble approach for robust generalization\n\n## Key Insights\n\nThe success of our solution stemmed from several critical discoveries:\n\n1. **Price and Duration Features Were Paramount**: As expected, flight price and duration emerged as the most important predictive features. We experimented with multiple variations of these features including:\n   - Absolute values (raw price and duration)\n   - Relative measures (compared to group averages)\n   - Ranked positions within search results\n   \n   This comprehensive treatment of the core decision factors proved essential for model performance.\n\n2. **Historical Counters Provided Substantial Gains**: Historical counters delivered significant performance improvements, contributing approximately 200 features to our final dataset. However, this represents only a fraction of the possible categorical combinations available in the data. There remains substantial opportunity for further score improvements by exploring additional feature combinations, though computational constraints limited our exploration.\n\n3. **Multi-Loss Function Ensemble Benefits**: A key discovery was that combining models trained with different loss functions, even when using identical training data, yielded meaningful performance gains. This suggests that different loss functions capture complementary aspects of the ranking problem, making ensemble diversity valuable beyond just data variation.\n\nThanks for attention",
      "votes": null
    },
    {
      "id": "3271373",
      "postDate": "08/18/2025 17:44:06",
      "content": "<p>Excellent work, so many valuable insights!<br>\nVery elaborate solution.</p>",
      "rawMarkdown": "Excellent work, so many valuable insights!\nVery elaborate solution.",
      "votes": null
    },
    {
      "id": "3273648",
      "postDate": "08/23/2025 05:11:00",
      "content": "<p>thank you for sharing your insights, is the code avaliable somewhere?</p>",
      "rawMarkdown": "thank you for sharing your insights, is the code avaliable somewhere?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3271373,
      "author_name": "mikhailgolubchik",
      "author_url": "",
      "post_date": "08/18/2025 17:44:06",
      "content": "<p>Excellent work, so many valuable insights!<br>\nVery elaborate solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3273648,
      "author_name": "junhaochan",
      "author_url": "",
      "post_date": "08/23/2025 05:11:00",
      "content": "<p>thank you for sharing your insights, is the code avaliable somewhere?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3271368": "# Machine Learning Competition Solution\n\n## Overview\n\nWe would like to thank the organizers for creating such an engaging and challenging competition. This document outlines our comprehensive approach to solving the flight ranking problem.\n\nOur solution began with the foundation of a public notebook ([XGBoost Ranker](https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars)), though by the competition's end, our implementation had evolved significantly beyond the original approach.\n\n## Feature Engineering Strategy\n\nOur feature engineering process was methodical and multi-faceted, focusing on capturing the complexity of flight booking behavior and preferences.\n\n### 1. Core Flight Features\n\n**Route Analysis:**\n- Extracted segment counts and connection details for each flight leg\n- Constructed complete flight paths (origin → connections → destination)\n- Analyzed route structure and categorized connection types\n\n**Temporal Features:**\n- Converted departure and arrival times to minutes for consistent processing\n- Extracted time-based patterns including time of day, day of week, and seasonality\n- Identified holidays and weekends to capture booking pattern variations\n- Analyzed business-friendly timing preferences\n\n### 2. Geographic & Airport Features\n\n**Distance Calculations:**\n- Implemented Haversine formula for precise inter-airport distance measurements\n- Calculated elevation differences between airports\n- Identified intercontinental flights for specialized handling\n\n**Airport Characteristics:**\n- Classified airports as hub versus regional facilities\n- Analyzed route popularity and traffic patterns\n- Implemented continental grouping for geographic insights\n\n### 3. Advanced User Behavior Features\n\n**Historical Patterns:**\n- Analyzed user preferences across airlines, booking timing, and service classes\n- Incorporated loyalty program analysis\n- Evaluated corporate policy compliance patterns\n\n**Contextual Analysis:**\n- Calculated relative rankings within search sessions (price, duration, convenience)\n- Identified Pareto-optimal options in the choice set\n- Analyzed price-quality trade-off preferences\n\n### 4. Flight Quality & Convenience Metrics\n\n**Quality Scoring:**\n- Mapped aircraft quality characteristics (wide-body vs narrow-body)\n- Assessed schedule convenience (business hours vs red-eye flights)\n- Analyzed fare flexibility options\n- Evaluated connection time comfort levels\n\n**Pricing Analytics:**\n- Calculated price-per-kilometer and price-per-minute efficiency ratios\n- Measured deviation from group medians\n- Analyzed competitive positioning within search results\n\n### 5. Temporal Split & Historical Counters\n\nFollowing successful strategies employed by other competition winners, we implemented historical counters with multiple categorical feature combinations. Our approach involved:\n\n- Using 50% of the time-sorted training dataset as historical reference\n- Training on the remaining 50% to maintain temporal validity\n- Creating interaction features across multiple categorical dimensions\n\n### 6. Data Sampling Strategy\n\nTo ensure training stability and computational efficiency:\n\n- Limited group sizes to 64 samples maximum\n- Preserved all positive examples (selected=1) to maintain signal strength\n- Applied random sampling to negative examples within each group\n\n## Model Architecture\n\n### 7. Model Training\n\nOur ensemble approach incorporated multiple ranking algorithms:\n\n**Primary Models:**\n- **CatBoostRanker** with three different loss functions:\n  - QueryCrossEntropy (alpha=1.0) - primary loss function\n  - PairLogitPairwise - secondary approach\n  - YetiRankPairwise - tertiary validation\n- **LightGBM Ranker** implemented by Farukcan\n\n### 8. Ensemble Strategy\n\nOur final prediction system employed a normalized score ensemble:\n\n- Normalized all individual ranker scores to ensure comparable scales\n- Combined normalized scores using optimized ratios\n- Selected ensemble weights primarily through leaderboard validation\n- Maintained simplicity in the ensemble approach for robust generalization\n\n## Key Insights\n\nThe success of our solution stemmed from several critical discoveries:\n\n1. **Price and Duration Features Were Paramount**: As expected, flight price and duration emerged as the most important predictive features. We experimented with multiple variations of these features including:\n   - Absolute values (raw price and duration)\n   - Relative measures (compared to group averages)\n   - Ranked positions within search results\n   \n   This comprehensive treatment of the core decision factors proved essential for model performance.\n\n2. **Historical Counters Provided Substantial Gains**: Historical counters delivered significant performance improvements, contributing approximately 200 features to our final dataset. However, this represents only a fraction of the possible categorical combinations available in the data. There remains substantial opportunity for further score improvements by exploring additional feature combinations, though computational constraints limited our exploration.\n\n3. **Multi-Loss Function Ensemble Benefits**: A key discovery was that combining models trained with different loss functions, even when using identical training data, yielded meaningful performance gains. This suggests that different loss functions capture complementary aspects of the ranking problem, making ensemble diversity valuable beyond just data variation.\n\nThanks for attention",
    "3271373": "Excellent work, so many valuable insights!\nVery elaborate solution.",
    "3273648": "thank you for sharing your insights, is the code avaliable somewhere?"
  },
  "source": "meta"
}