{
  "id": 599687,
  "title": "4th solution: XGBoost at the Core of Flight Recommendation",
  "url": "/competitions/aeroclub-recsys-2025/writeups/4th-solution-xgboost-model",
  "author_name": "",
  "post_date": "2025-08-18T11:56:08.713Z",
  "votes": 11,
  "comment_count": 1,
  "views": 0,
  "content": "<p>First of all, I’d like to thank the organizers for hosting this recommendation competition. As a newcomer to both Kaggle and recommendation systems, I feel honored to work with a real-world industrial dataset and to develop and refine algorithms on it.</p>\n<p>I used <a href=\"https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars\" target=\"_blank\">@ka1242’s notebook</a> as my starting point, and my entire pipeline is built upon it. Many thanks to <a href=\"https://www.kaggle.com/ka1242\" target=\"_blank\">@ka1242</a> for sharing this solid foundation, which allowed me to focus on further improvements and model refinement.</p>\n<p><strong>You can find the source code for training our best xgboost model in this repo</strong> <a href=\"https://github.com/mango7789/FlightRank-2025-Aeroclub-RecSys-Cup-4th-Solution\" target=\"_blank\">https://github.com/mango7789/FlightRank-2025-Aeroclub-RecSys-Cup-4th-Solution</a></p>\n<h2>Data Split and Validation</h2>\n<p>Since the order of occurrence can influence recommendation performance, I followed the train–validation split strategy used in the starter notebook. Specifically, the last ~10% of the training data is held out for validation.</p>\n<h2>Feature Engineering</h2>\n<p>I constructed a wide range of features for this task. Broadly, they can be grouped into the following categories: <strong>segments, time, rank, price, cabin class, company, carrier, frequent flyer, profile, airport, corporateTariffCode, group, label, option, and route.</strong> Most of these are relatively straightforward and intuitive to design, so here I will highlight a few representative ones:</p>\n<ul>\n<li><strong>Rank</strong>: derived from the within-query ranks of price and duration, each normalized separately, and further combined through interaction features.  </li>\n<li><strong>Cabin class</strong>: added a feature indicating whether all segments belong to the same (economic) class.  </li>\n<li><strong>Company</strong>: occurrence counts, average price, and proportion of direct flights — this category provided notable improvement.  </li>\n<li><strong>Carrier</strong>: selection frequency and ratio.  </li>\n<li><strong>Profile</strong>: only counted occurrences, as additional statistics were not effective due to sparsity.  </li>\n<li><strong>Airport</strong>: enriched with external data from <a href=\"https://www.kaggle.com/competitions/aeroclub-recsys-2025/discussion/585877\" target=\"_blank\">this discussion</a>, including indicators such as whether the flight is international and whether it crosses time zones. I also experimented with haversine distance and “same departure–return airport” as a proxy for convenience.  </li>\n<li><strong>CorporateTariffCode</strong>: usage frequency.  </li>\n<li><strong>Route</strong>: aggregated statistics such as average price, average duration, direct-flight ratio, and corporate-code ratio.  </li>\n<li><strong>Group</strong>: aggregated stats for each <code>ranker_id</code>.  </li>\n</ul>\n<p>The <strong>label</strong> and <strong>option</strong> features are derived from insights in the raw json files. Each json file contains additional information, such as a label for each flight combination. These labels likely correspond to the information displayed on the flight page during ticket purchase and may influence user choice. I extracted these labels from the json files; however, as noted in <a href=\"https://www.kaggle.com/competitions/aeroclub-recsys-2025/discussion/590264#3253916\" target=\"_blank\">this discussion</a>, a perfect match to the parquet data is not possible. Instead, I applied a coarse-grained row-order alignment, which introduces roughly 10% mismatches. Despite this noise, including these features still led to some model improvement.  </p>\n<p>Upon inspecting the raw json, I observed that a single <strong>flight option</strong> can have multiple <strong>pricings</strong>, likely due to different promotional policies. In the parquet data, multiple records may correspond to the same option with different pricings, which are <strong>mutually exclusive</strong> since a user can select only one flight and one pricing at a time. Therefore, it is important to construct features grouped by option, such as maximum, minimum, average, or rank within the option. As with labels, perfect alignment between the parquet and json files is infeasible, so I used the <code>flight_hash</code> as a surrogate to approximate the option.</p>\n<p>Originally, the <code>flight_hash</code> was used in reranking as a penalty for flights within the same option (<a href=\"https://www.kaggle.com/code/mango789/xgboost-ranker-rule-based-rerank\" target=\"_blank\">example here</a>). After confirming its usefulness in reranking, I extended its use to create additional groupby features for the model.</p>\n<p>I also extracted <code>searchType</code> and <code>age</code> from the json files, but for reasons I do not fully understand, these features did not improve model performance.</p>\n<h2>Other Notes</h2>\n<p>Most of my work in this competition focused on feature engineering, training XGBoost and LGBM models, tuning parameters, and experimenting with objectives such as <code>rank:pairwise</code>, <code>rank:ndcg</code>, and <code>lambdarank</code>. For data transformations, I noticed that price and duration follow a long-tail distribution, so I applied log transformations.  </p>\n<p>I also explored other ideas, such as pseudo-labeling and adding weights to each group, but these approaches did not improve performance on my side.  </p>\n<p>Additionally, I considered training a rerank deep learning model to fine-tune scores after the coarse-grained GBDT predictions, as I observed a hitrate@20 above 0.8 in local CV. However, due to time constraints, this was not implemented.</p>\n<p>Regarding the <code>bySelf</code> column, all tickets in the training data are booked <code>bySelf</code>, while in the test data there are about 200,000 rows that are not <code>bySelf</code>. I believe results could improve if we had non-<code>bySelf</code> tickets in the training data. It might also be beneficial to create some human-crafted rules for non-<code>bySelf</code> ticket selection based on empirical company choices.</p>\n<p>Other external data, such as more informative airport data, flight number (for delay statistics) and aircraft code (to distinguish low-cost, standard, and premium flights), may also help improve performance, I think.</p>\n<p>To be honest, my best model was trained three weeks ago, and further attempts at feature engineering or adjustments did not yield improvement. This may be due to the sparsity in users' historical behavior, suggesting that the practical limit for GBDT methods on this dataset is around 0.53–0.54.</p>\n<h2>Deep Learning Ranker</h2>\n<p>To explore ways of enhancing the module's capabilities, I tested several architectures with different configurations. The most effective one turned out to be a hybrid neural ranking architecture that combines:</p>\n<p>Linear terms:</p>\n<ul>\n<li>One-hot style embeddings for categorical features (dimension = 1).</li>\n<li>A linear projection of numeric features.</li>\n</ul>\n<p>Embeddings for categorical features:</p>\n<ul>\n<li>Each categorical feature is mapped to a dense embedding (shared size = 64).</li>\n<li>These embeddings are concatenated with normalized numeric features.</li>\n</ul>\n<p>Deep MLP tower:</p>\n<ul>\n<li>Several fully connected layers (512 → 256 → 128 → 64) with BatchNorm, ReLU, and Dropout.</li>\n<li>Produces nonlinear interactions between categorical and numeric features.</li>\n</ul>\n<p>Output layer:</p>\n<ul>\n<li>Final linear unit outputs a single relevance score per item.</li>\n<li>Scores are trained with a pairwise ranking loss (xRankNet variant).</li>\n</ul>\n<p>The deep learning model showed relatively weak performance (0.48–0.495), emphasizing that deeper architectures are not always more effective than traditional methods. In many cases, it is more reliable to rely on well-established models such as XGBoost, LightGBM, and similar approaches.</p>\n<p>I also experimented with other architectures incorporating ideas such as SENet with normalization and dropout, bilinear interactions, self-attention layers, and transformer-based encoders for categorical and numerical features. However, in most cases, the results were not optimal or showed only minor variations in individual performance.</p>\n<h2>Ensemble</h2>\n<p>To further enhance performance and boost the overall score, I explored several ensemble strategies that combine the three previously trained models: XGBoost, LightGBM, and DL, with the most effective one being:</p>\n<p>First, the confidence score is computed by combining two signals: a base score that decreases linearly from 1 for the top item to 0 for the last item within each group, and an RRF score that gives higher weight to top-ranked items; the final confidence is a weighted mix of these two, balancing group position and reciprocal rank.</p>\n<p>Next, load multiple confidence files, aligning them, and then computing weights for each model based on two factors — uniqueness (how different it is from others, using PCA + Spearman residuals) and informativeness (entropy of its confidence distribution); these weights are combined to produce a final weighted confidence score, which is then ranked per group to generate the ensemble submission.</p>\n<p>This method achieved solid performance (0.536–0.537), but it required tremendous effort for only a marginal improvement over the standalone XGBoost model (0.53552).</p>",
  "messages": [
    {
      "id": "3271158",
      "postDate": "08/18/2025 07:35:54",
      "content": "<p>First of all, I’d like to thank the organizers for hosting this recommendation competition. As a newcomer to both Kaggle and recommendation systems, I feel honored to work with a real-world industrial dataset and to develop and refine algorithms on it.</p>\n<p>I used <a href=\"https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars\" target=\"_blank\">@ka1242’s notebook</a> as my starting point, and my entire pipeline is built upon it. Many thanks to <a href=\"https://www.kaggle.com/ka1242\" target=\"_blank\">@ka1242</a> for sharing this solid foundation, which allowed me to focus on further improvements and model refinement.</p>\n<p><strong>You can find the source code for training our best xgboost model in this repo</strong> <a href=\"https://github.com/mango7789/FlightRank-2025-Aeroclub-RecSys-Cup-4th-Solution\" target=\"_blank\">https://github.com/mango7789/FlightRank-2025-Aeroclub-RecSys-Cup-4th-Solution</a></p>\n<h2>Data Split and Validation</h2>\n<p>Since the order of occurrence can influence recommendation performance, I followed the train–validation split strategy used in the starter notebook. Specifically, the last ~10% of the training data is held out for validation.</p>\n<h2>Feature Engineering</h2>\n<p>I constructed a wide range of features for this task. Broadly, they can be grouped into the following categories: <strong>segments, time, rank, price, cabin class, company, carrier, frequent flyer, profile, airport, corporateTariffCode, group, label, option, and route.</strong> Most of these are relatively straightforward and intuitive to design, so here I will highlight a few representative ones:</p>\n<ul>\n<li><strong>Rank</strong>: derived from the within-query ranks of price and duration, each normalized separately, and further combined through interaction features.  </li>\n<li><strong>Cabin class</strong>: added a feature indicating whether all segments belong to the same (economic) class.  </li>\n<li><strong>Company</strong>: occurrence counts, average price, and proportion of direct flights — this category provided notable improvement.  </li>\n<li><strong>Carrier</strong>: selection frequency and ratio.  </li>\n<li><strong>Profile</strong>: only counted occurrences, as additional statistics were not effective due to sparsity.  </li>\n<li><strong>Airport</strong>: enriched with external data from <a href=\"https://www.kaggle.com/competitions/aeroclub-recsys-2025/discussion/585877\" target=\"_blank\">this discussion</a>, including indicators such as whether the flight is international and whether it crosses time zones. I also experimented with haversine distance and “same departure–return airport” as a proxy for convenience.  </li>\n<li><strong>CorporateTariffCode</strong>: usage frequency.  </li>\n<li><strong>Route</strong>: aggregated statistics such as average price, average duration, direct-flight ratio, and corporate-code ratio.  </li>\n<li><strong>Group</strong>: aggregated stats for each <code>ranker_id</code>.  </li>\n</ul>\n<p>The <strong>label</strong> and <strong>option</strong> features are derived from insights in the raw json files. Each json file contains additional information, such as a label for each flight combination. These labels likely correspond to the information displayed on the flight page during ticket purchase and may influence user choice. I extracted these labels from the json files; however, as noted in <a href=\"https://www.kaggle.com/competitions/aeroclub-recsys-2025/discussion/590264#3253916\" target=\"_blank\">this discussion</a>, a perfect match to the parquet data is not possible. Instead, I applied a coarse-grained row-order alignment, which introduces roughly 10% mismatches. Despite this noise, including these features still led to some model improvement.  </p>\n<p>Upon inspecting the raw json, I observed that a single <strong>flight option</strong> can have multiple <strong>pricings</strong>, likely due to different promotional policies. In the parquet data, multiple records may correspond to the same option with different pricings, which are <strong>mutually exclusive</strong> since a user can select only one flight and one pricing at a time. Therefore, it is important to construct features grouped by option, such as maximum, minimum, average, or rank within the option. As with labels, perfect alignment between the parquet and json files is infeasible, so I used the <code>flight_hash</code> as a surrogate to approximate the option.</p>\n<p>Originally, the <code>flight_hash</code> was used in reranking as a penalty for flights within the same option (<a href=\"https://www.kaggle.com/code/mango789/xgboost-ranker-rule-based-rerank\" target=\"_blank\">example here</a>). After confirming its usefulness in reranking, I extended its use to create additional groupby features for the model.</p>\n<p>I also extracted <code>searchType</code> and <code>age</code> from the json files, but for reasons I do not fully understand, these features did not improve model performance.</p>\n<h2>Other Notes</h2>\n<p>Most of my work in this competition focused on feature engineering, training XGBoost and LGBM models, tuning parameters, and experimenting with objectives such as <code>rank:pairwise</code>, <code>rank:ndcg</code>, and <code>lambdarank</code>. For data transformations, I noticed that price and duration follow a long-tail distribution, so I applied log transformations.  </p>\n<p>I also explored other ideas, such as pseudo-labeling and adding weights to each group, but these approaches did not improve performance on my side.  </p>\n<p>Additionally, I considered training a rerank deep learning model to fine-tune scores after the coarse-grained GBDT predictions, as I observed a hitrate@20 above 0.8 in local CV. However, due to time constraints, this was not implemented.</p>\n<p>Regarding the <code>bySelf</code> column, all tickets in the training data are booked <code>bySelf</code>, while in the test data there are about 200,000 rows that are not <code>bySelf</code>. I believe results could improve if we had non-<code>bySelf</code> tickets in the training data. It might also be beneficial to create some human-crafted rules for non-<code>bySelf</code> ticket selection based on empirical company choices.</p>\n<p>Other external data, such as more informative airport data, flight number (for delay statistics) and aircraft code (to distinguish low-cost, standard, and premium flights), may also help improve performance, I think.</p>\n<p>To be honest, my best model was trained three weeks ago, and further attempts at feature engineering or adjustments did not yield improvement. This may be due to the sparsity in users' historical behavior, suggesting that the practical limit for GBDT methods on this dataset is around 0.53–0.54.</p>\n<h2>Deep Learning Ranker</h2>\n<p>To explore ways of enhancing the module's capabilities, I tested several architectures with different configurations. The most effective one turned out to be a hybrid neural ranking architecture that combines:</p>\n<p>Linear terms:</p>\n<ul>\n<li>One-hot style embeddings for categorical features (dimension = 1).</li>\n<li>A linear projection of numeric features.</li>\n</ul>\n<p>Embeddings for categorical features:</p>\n<ul>\n<li>Each categorical feature is mapped to a dense embedding (shared size = 64).</li>\n<li>These embeddings are concatenated with normalized numeric features.</li>\n</ul>\n<p>Deep MLP tower:</p>\n<ul>\n<li>Several fully connected layers (512 → 256 → 128 → 64) with BatchNorm, ReLU, and Dropout.</li>\n<li>Produces nonlinear interactions between categorical and numeric features.</li>\n</ul>\n<p>Output layer:</p>\n<ul>\n<li>Final linear unit outputs a single relevance score per item.</li>\n<li>Scores are trained with a pairwise ranking loss (xRankNet variant).</li>\n</ul>\n<p>The deep learning model showed relatively weak performance (0.48–0.495), emphasizing that deeper architectures are not always more effective than traditional methods. In many cases, it is more reliable to rely on well-established models such as XGBoost, LightGBM, and similar approaches.</p>\n<p>I also experimented with other architectures incorporating ideas such as SENet with normalization and dropout, bilinear interactions, self-attention layers, and transformer-based encoders for categorical and numerical features. However, in most cases, the results were not optimal or showed only minor variations in individual performance.</p>\n<h2>Ensemble</h2>\n<p>To further enhance performance and boost the overall score, I explored several ensemble strategies that combine the three previously trained models: XGBoost, LightGBM, and DL, with the most effective one being:</p>\n<p>First, the confidence score is computed by combining two signals: a base score that decreases linearly from 1 for the top item to 0 for the last item within each group, and an RRF score that gives higher weight to top-ranked items; the final confidence is a weighted mix of these two, balancing group position and reciprocal rank.</p>\n<p>Next, load multiple confidence files, aligning them, and then computing weights for each model based on two factors — uniqueness (how different it is from others, using PCA + Spearman residuals) and informativeness (entropy of its confidence distribution); these weights are combined to produce a final weighted confidence score, which is then ranked per group to generate the ensemble submission.</p>\n<p>This method achieved solid performance (0.536–0.537), but it required tremendous effort for only a marginal improvement over the standalone XGBoost model (0.53552).</p>",
      "rawMarkdown": "First of all, I’d like to thank the organizers for hosting this recommendation competition. As a newcomer to both Kaggle and recommendation systems, I feel honored to work with a real-world industrial dataset and to develop and refine algorithms on it.\n\nI used [@ka1242’s notebook](https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars) as my starting point, and my entire pipeline is built upon it. Many thanks to @ka1242 for sharing this solid foundation, which allowed me to focus on further improvements and model refinement.\n\n**You can find the source code for training our best xgboost model in this repo** https://github.com/mango7789/FlightRank-2025-Aeroclub-RecSys-Cup-4th-Solution\n\n## Data Split and Validation\n\nSince the order of occurrence can influence recommendation performance, I followed the train–validation split strategy used in the starter notebook. Specifically, the last ~10% of the training data is held out for validation.\n\n## Feature Engineering\n\nI constructed a wide range of features for this task. Broadly, they can be grouped into the following categories: **segments, time, rank, price, cabin class, company, carrier, frequent flyer, profile, airport, corporateTariffCode, group, label, option, and route.** Most of these are relatively straightforward and intuitive to design, so here I will highlight a few representative ones:\n\n- **Rank**: derived from the within-query ranks of price and duration, each normalized separately, and further combined through interaction features.  \n- **Cabin class**: added a feature indicating whether all segments belong to the same (economic) class.  \n- **Company**: occurrence counts, average price, and proportion of direct flights — this category provided notable improvement.  \n- **Carrier**: selection frequency and ratio.  \n- **Profile**: only counted occurrences, as additional statistics were not effective due to sparsity.  \n- **Airport**: enriched with external data from [this discussion](https://www.kaggle.com/competitions/aeroclub-recsys-2025/discussion/585877), including indicators such as whether the flight is international and whether it crosses time zones. I also experimented with haversine distance and “same departure–return airport” as a proxy for convenience.  \n- **CorporateTariffCode**: usage frequency.  \n- **Route**: aggregated statistics such as average price, average duration, direct-flight ratio, and corporate-code ratio.  \n- **Group**: aggregated stats for each `ranker_id`.  \n\nThe **label** and **option** features are derived from insights in the raw json files. Each json file contains additional information, such as a label for each flight combination. These labels likely correspond to the information displayed on the flight page during ticket purchase and may influence user choice. I extracted these labels from the json files; however, as noted in [this discussion](https://www.kaggle.com/competitions/aeroclub-recsys-2025/discussion/590264#3253916), a perfect match to the parquet data is not possible. Instead, I applied a coarse-grained row-order alignment, which introduces roughly 10% mismatches. Despite this noise, including these features still led to some model improvement.  \n\nUpon inspecting the raw json, I observed that a single **flight option** can have multiple **pricings**, likely due to different promotional policies. In the parquet data, multiple records may correspond to the same option with different pricings, which are **mutually exclusive** since a user can select only one flight and one pricing at a time. Therefore, it is important to construct features grouped by option, such as maximum, minimum, average, or rank within the option. As with labels, perfect alignment between the parquet and json files is infeasible, so I used the `flight_hash` as a surrogate to approximate the option.\n\nOriginally, the `flight_hash` was used in reranking as a penalty for flights within the same option ([example here](https://www.kaggle.com/code/mango789/xgboost-ranker-rule-based-rerank)). After confirming its usefulness in reranking, I extended its use to create additional groupby features for the model.\n\nI also extracted `searchType` and `age` from the json files, but for reasons I do not fully understand, these features did not improve model performance.\n\n## Other Notes\n\nMost of my work in this competition focused on feature engineering, training XGBoost and LGBM models, tuning parameters, and experimenting with objectives such as `rank:pairwise`, `rank:ndcg`, and `lambdarank`. For data transformations, I noticed that price and duration follow a long-tail distribution, so I applied log transformations.  \n\nI also explored other ideas, such as pseudo-labeling and adding weights to each group, but these approaches did not improve performance on my side.  \n\nAdditionally, I considered training a rerank deep learning model to fine-tune scores after the coarse-grained GBDT predictions, as I observed a hitrate@20 above 0.8 in local CV. However, due to time constraints, this was not implemented.\n\nRegarding the `bySelf` column, all tickets in the training data are booked `bySelf`, while in the test data there are about 200,000 rows that are not `bySelf`. I believe results could improve if we had non-`bySelf` tickets in the training data. It might also be beneficial to create some human-crafted rules for non-`bySelf` ticket selection based on empirical company choices.\n\nOther external data, such as more informative airport data, flight number (for delay statistics) and aircraft code (to distinguish low-cost, standard, and premium flights), may also help improve performance, I think.\n\nTo be honest, my best model was trained three weeks ago, and further attempts at feature engineering or adjustments did not yield improvement. This may be due to the sparsity in users' historical behavior, suggesting that the practical limit for GBDT methods on this dataset is around 0.53–0.54.\n\n## Deep Learning Ranker\n\nTo explore ways of enhancing the module's capabilities, I tested several architectures with different configurations. The most effective one turned out to be a hybrid neural ranking architecture that combines:\n\nLinear terms:\n - One-hot style embeddings for categorical features (dimension = 1).\n - A linear projection of numeric features.\n\nEmbeddings for categorical features:\n - Each categorical feature is mapped to a dense embedding (shared size = 64).\n - These embeddings are concatenated with normalized numeric features.\n\nDeep MLP tower:\n - Several fully connected layers (512 → 256 → 128 → 64) with BatchNorm, ReLU, and Dropout.\n - Produces nonlinear interactions between categorical and numeric features.\n\nOutput layer:\n - Final linear unit outputs a single relevance score per item.\n - Scores are trained with a pairwise ranking loss (xRankNet variant).\n\nThe deep learning model showed relatively weak performance (0.48–0.495), emphasizing that deeper architectures are not always more effective than traditional methods. In many cases, it is more reliable to rely on well-established models such as XGBoost, LightGBM, and similar approaches.\n\nI also experimented with other architectures incorporating ideas such as SENet with normalization and dropout, bilinear interactions, self-attention layers, and transformer-based encoders for categorical and numerical features. However, in most cases, the results were not optimal or showed only minor variations in individual performance.\n\n## Ensemble\n\nTo further enhance performance and boost the overall score, I explored several ensemble strategies that combine the three previously trained models: XGBoost, LightGBM, and DL, with the most effective one being:\n\nFirst, the confidence score is computed by combining two signals: a base score that decreases linearly from 1 for the top item to 0 for the last item within each group, and an RRF score that gives higher weight to top-ranked items; the final confidence is a weighted mix of these two, balancing group position and reciprocal rank.\n\nNext, load multiple confidence files, aligning them, and then computing weights for each model based on two factors — uniqueness (how different it is from others, using PCA + Spearman residuals) and informativeness (entropy of its confidence distribution); these weights are combined to produce a final weighted confidence score, which is then ranked per group to generate the ensemble submission.\n\nThis method achieved solid performance (0.536–0.537), but it required tremendous effort for only a marginal improvement over the standalone XGBoost model (0.53552).",
      "votes": null
    },
    {
      "id": "3272114",
      "postDate": "08/20/2025 09:20:31",
      "content": "<p>Feel free to check out the source code repository: <a href=\"https://github.com/mango7789/FlightRank-2025-Aeroclub-RecSys-Cup-4th-Solution\" target=\"_blank\">FlightRank 2025 – 4th Solution</a>😘</p>",
      "rawMarkdown": "Feel free to check out the source code repository: [FlightRank 2025 – 4th Solution](https://github.com/mango7789/FlightRank-2025-Aeroclub-RecSys-Cup-4th-Solution)😘",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3272114,
      "author_name": "mango789",
      "author_url": "",
      "post_date": "08/20/2025 09:20:31",
      "content": "<p>Feel free to check out the source code repository: <a href=\"https://github.com/mango7789/FlightRank-2025-Aeroclub-RecSys-Cup-4th-Solution\" target=\"_blank\">FlightRank 2025 – 4th Solution</a>😘</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3271158": "First of all, I’d like to thank the organizers for hosting this recommendation competition. As a newcomer to both Kaggle and recommendation systems, I feel honored to work with a real-world industrial dataset and to develop and refine algorithms on it.\n\nI used [@ka1242’s notebook](https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars) as my starting point, and my entire pipeline is built upon it. Many thanks to @ka1242 for sharing this solid foundation, which allowed me to focus on further improvements and model refinement.\n\n**You can find the source code for training our best xgboost model in this repo** https://github.com/mango7789/FlightRank-2025-Aeroclub-RecSys-Cup-4th-Solution\n\n## Data Split and Validation\n\nSince the order of occurrence can influence recommendation performance, I followed the train–validation split strategy used in the starter notebook. Specifically, the last ~10% of the training data is held out for validation.\n\n## Feature Engineering\n\nI constructed a wide range of features for this task. Broadly, they can be grouped into the following categories: **segments, time, rank, price, cabin class, company, carrier, frequent flyer, profile, airport, corporateTariffCode, group, label, option, and route.** Most of these are relatively straightforward and intuitive to design, so here I will highlight a few representative ones:\n\n- **Rank**: derived from the within-query ranks of price and duration, each normalized separately, and further combined through interaction features.  \n- **Cabin class**: added a feature indicating whether all segments belong to the same (economic) class.  \n- **Company**: occurrence counts, average price, and proportion of direct flights — this category provided notable improvement.  \n- **Carrier**: selection frequency and ratio.  \n- **Profile**: only counted occurrences, as additional statistics were not effective due to sparsity.  \n- **Airport**: enriched with external data from [this discussion](https://www.kaggle.com/competitions/aeroclub-recsys-2025/discussion/585877), including indicators such as whether the flight is international and whether it crosses time zones. I also experimented with haversine distance and “same departure–return airport” as a proxy for convenience.  \n- **CorporateTariffCode**: usage frequency.  \n- **Route**: aggregated statistics such as average price, average duration, direct-flight ratio, and corporate-code ratio.  \n- **Group**: aggregated stats for each `ranker_id`.  \n\nThe **label** and **option** features are derived from insights in the raw json files. Each json file contains additional information, such as a label for each flight combination. These labels likely correspond to the information displayed on the flight page during ticket purchase and may influence user choice. I extracted these labels from the json files; however, as noted in [this discussion](https://www.kaggle.com/competitions/aeroclub-recsys-2025/discussion/590264#3253916), a perfect match to the parquet data is not possible. Instead, I applied a coarse-grained row-order alignment, which introduces roughly 10% mismatches. Despite this noise, including these features still led to some model improvement.  \n\nUpon inspecting the raw json, I observed that a single **flight option** can have multiple **pricings**, likely due to different promotional policies. In the parquet data, multiple records may correspond to the same option with different pricings, which are **mutually exclusive** since a user can select only one flight and one pricing at a time. Therefore, it is important to construct features grouped by option, such as maximum, minimum, average, or rank within the option. As with labels, perfect alignment between the parquet and json files is infeasible, so I used the `flight_hash` as a surrogate to approximate the option.\n\nOriginally, the `flight_hash` was used in reranking as a penalty for flights within the same option ([example here](https://www.kaggle.com/code/mango789/xgboost-ranker-rule-based-rerank)). After confirming its usefulness in reranking, I extended its use to create additional groupby features for the model.\n\nI also extracted `searchType` and `age` from the json files, but for reasons I do not fully understand, these features did not improve model performance.\n\n## Other Notes\n\nMost of my work in this competition focused on feature engineering, training XGBoost and LGBM models, tuning parameters, and experimenting with objectives such as `rank:pairwise`, `rank:ndcg`, and `lambdarank`. For data transformations, I noticed that price and duration follow a long-tail distribution, so I applied log transformations.  \n\nI also explored other ideas, such as pseudo-labeling and adding weights to each group, but these approaches did not improve performance on my side.  \n\nAdditionally, I considered training a rerank deep learning model to fine-tune scores after the coarse-grained GBDT predictions, as I observed a hitrate@20 above 0.8 in local CV. However, due to time constraints, this was not implemented.\n\nRegarding the `bySelf` column, all tickets in the training data are booked `bySelf`, while in the test data there are about 200,000 rows that are not `bySelf`. I believe results could improve if we had non-`bySelf` tickets in the training data. It might also be beneficial to create some human-crafted rules for non-`bySelf` ticket selection based on empirical company choices.\n\nOther external data, such as more informative airport data, flight number (for delay statistics) and aircraft code (to distinguish low-cost, standard, and premium flights), may also help improve performance, I think.\n\nTo be honest, my best model was trained three weeks ago, and further attempts at feature engineering or adjustments did not yield improvement. This may be due to the sparsity in users' historical behavior, suggesting that the practical limit for GBDT methods on this dataset is around 0.53–0.54.\n\n## Deep Learning Ranker\n\nTo explore ways of enhancing the module's capabilities, I tested several architectures with different configurations. The most effective one turned out to be a hybrid neural ranking architecture that combines:\n\nLinear terms:\n - One-hot style embeddings for categorical features (dimension = 1).\n - A linear projection of numeric features.\n\nEmbeddings for categorical features:\n - Each categorical feature is mapped to a dense embedding (shared size = 64).\n - These embeddings are concatenated with normalized numeric features.\n\nDeep MLP tower:\n - Several fully connected layers (512 → 256 → 128 → 64) with BatchNorm, ReLU, and Dropout.\n - Produces nonlinear interactions between categorical and numeric features.\n\nOutput layer:\n - Final linear unit outputs a single relevance score per item.\n - Scores are trained with a pairwise ranking loss (xRankNet variant).\n\nThe deep learning model showed relatively weak performance (0.48–0.495), emphasizing that deeper architectures are not always more effective than traditional methods. In many cases, it is more reliable to rely on well-established models such as XGBoost, LightGBM, and similar approaches.\n\nI also experimented with other architectures incorporating ideas such as SENet with normalization and dropout, bilinear interactions, self-attention layers, and transformer-based encoders for categorical and numerical features. However, in most cases, the results were not optimal or showed only minor variations in individual performance.\n\n## Ensemble\n\nTo further enhance performance and boost the overall score, I explored several ensemble strategies that combine the three previously trained models: XGBoost, LightGBM, and DL, with the most effective one being:\n\nFirst, the confidence score is computed by combining two signals: a base score that decreases linearly from 1 for the top item to 0 for the last item within each group, and an RRF score that gives higher weight to top-ranked items; the final confidence is a weighted mix of these two, balancing group position and reciprocal rank.\n\nNext, load multiple confidence files, aligning them, and then computing weights for each model based on two factors — uniqueness (how different it is from others, using PCA + Spearman residuals) and informativeness (entropy of its confidence distribution); these weights are combined to produce a final weighted confidence score, which is then ranked per group to generate the ensemble submission.\n\nThis method achieved solid performance (0.536–0.537), but it required tremendous effort for only a marginal improvement over the standalone XGBoost model (0.53552).",
    "3272114": "Feel free to check out the source code repository: [FlightRank 2025 – 4th Solution](https://github.com/mango7789/FlightRank-2025-Aeroclub-RecSys-Cup-4th-Solution)😘"
  },
  "source": "meta"
}