{
  "id": 600069,
  "title": "8-th place solution",
  "url": "/competitions/aeroclub-recsys-2025/discussion/600069",
  "author_name": "Mikhail Golubchik",
  "post_date": "2025-08-20T18:12:55.624000",
  "votes": 4,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Before diving into the solution description itself, I would first like to thank the competition organizers for an interesting task and for providing real, authentic data.</p>\n<h2>Solution Description</h2>\n<p><strong>Baseline</strong>. I began with the solution by Kirill Khoruzhii (@ka1242): 'AeroClub RecSys 2025 - XGBoost Ranking Baseline' (<a href=\"https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars)\" target=\"_blank\">https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars)</a>. I would like to thank him for an excellent set of features.</p>\n<p><strong>Models Used.</strong> I also experimented with using LightGBM both as a replacement for and in addition to XGBoost. However, it yielded roughly the same results, so I eventually discontinued its use by the end of the competition.</p>\n<p><strong>Important Features.</strong> As expected, the most important features turned out to be ticket prices and their derivatives. This includes metrics like the ratio to the median price within a group, among others. Features related to flight duration and the number of connections were also significant. Additionally, features indicating the possibility of ticket cancellation or exchange, as well as the ticket class, proved to be substantial contributors.</p>\n<p><strong>Feature Selection:</strong> Features were selected primarily based on intuition, driven by what we expected to be important for a person when choosing a flight ticket. This initial selection was then validated, and if the assumption about a feature's usefulness was confirmed, it was kept for the final model.</p>\n<p>In several cases, however, adding certain features—particularly categorical ones related to the customers' past choice history—led to severe overfitting. Although using such features seemed intuitively beneficial, they proved to be impractical in practice.</p>\n<p>For example, a factor like the price at which a customer/company previously booked a ticket turned out to be useful. Conversely, trying to incorporate the type of aircraft or the airline carrier a customer/company preferred in the past led to overfitting.</p>\n<p><strong>Overfitting:</strong> Features based on customer choice history led to overfitting. In fact, the first-place competition winner, <a href=\"https://www.kaggle.com/goldenlock\" target=\"_blank\">@goldenlock</a>, stated in his solution <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/aeroclub-recsys-2025/writeups/1st-solution-using-count-based-features</a> that he avoided using them entirely until the competition ended. However, after the competition concluded, he tested the approach I proposed for utilizing a subset of these features and managed to significantly improve the score of his final solution.</p>\n<p><strong>Mitigating Overfitting:</strong> The approach was based on a windowed method for generating customer history features. Specifically, for each group, features were constructed using data only from groups where a ticket was purchased by the same customers prior to the current ticket selection group</p>\n<p><strong>Historical Window Features:</strong> For the features used to calculate the history of option selection within a group, a specific company, or a user, the following metrics were generated: mean, median, and the number of groups in which a choice was previously made and this feature was present. The 25% and 75% quantiles were also generated. The historical features were based on the following features:</p>\n<ul>\n<li><p>'legs0_departureAt_hour', 'legs1_departureAt_hour', 'legs0_arrivalAt_hour', 'legs1_arrivalAt_hour' </p></li>\n<li><p>'price_rank', 'price_from_median', 'duration_rank', 'duration_from_median', </p></li>\n<li><p>'avg_cabin_class', 'baggage_total', 'l0_seg', </p></li>\n<li><p>'legs0_segments0_seatsAvailable', 'legs0_segments0_baggageAllowance_quantity', 'legs0_segments0_cabinClass', </p></li>\n<li><p>'miniRules1_statusInfos', 'miniRules0_statusInfos' </p></li>\n</ul>\n<h2>Training of Models</h2>\n<p><strong>Validation:</strong> When selecting features, the data was split into a training and test part, which was then combined with validation. As I now understand, this was probably a mistake, and validation and test should have been separated.<br>\nAt the same time, the score obtained on validation was almost the same as the one on the Public and Private Leaderboards. Moreover, improvements in local validation usually correlated with improvements on the LB.<br>\nLocal validation yielded about 0.535. For local validation, the last month of training data was used.</p>\n<p><strong>Training:</strong> Training was performed on the full training dataset without any use of test data. This included not using the test set for computing statistics or encoding categorical features.</p>\n<p><strong>Handling New Data in Test:</strong> To prevent the model from being significantly distorted by the appearance of new categories in categorical features, about 1% of categorical values in the training data were randomly replaced with missing values. If the proportion of such missing values was less than 1%, additional ones were added. Then, in the test set, unknown categories in categorical features were also marked as missing values (-1) during categorical encoding.</p>\n<p><strong>Reducing Group Size</strong>: While considering ways to preselect the most promising candidates for the final choice of the three most likely options a client would choose, I tried simply randomly reducing all groups to a maximum of 50 options. This involved downsampling negative examples (those not chosen by the client). Unexpectedly, this improved the score. I hypothesized that ranking became easier for the models, and the class imbalance was reduced. This is because the probability of correctly guessing the user's choice in a large group was low anyway. However, without reducing group size, models might have struggled more to learn anything useful from such a large group.</p>\n<p>A beneficial side effect was the reduction in memory usage during training and an increase in training speed. This, in turn, allowed for the use of more features. It's important to note that the group size reduction was performed after all features were calculated and statistics were collected. That is, the features were generated using the full groups. Reducing the group size before feature engineering led to worse model performance.</p>\n<h2>Possible Improvements</h2>\n<p>It would be beneficial to find a way to train the model without overfitting, not only on numerical data but also on historical categorical data. This was not achieved within the scope of this solution.</p>\n<p>Approaches that generate additional features using language models, or other transformers seem very promising. The idea is to use an attention mechanism to extract meaningful information from the full description of the offered options. <a href=\"https://www.kaggle.com/sergeyqt20244\" target=\"_blank\">@sergeyqt20244</a> mentioned attempts to apply Bert2rec, though so far without significant results: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/aeroclub-recsys-2025/discussion/599559#3271192</a></p>\n<p>Below is a notebook with my single model solution:<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/mikhailgolubchik/sm-xgboost-single</a></p>",
  "messages": [
    {
      "id": 3272336,
      "postDate": "2025-08-20T18:12:55.623Z",
      "content": "<p>Before diving into the solution description itself, I would first like to thank the competition organizers for an interesting task and for providing real, authentic data.</p>\n<h2>Solution Description</h2>\n<p><strong>Baseline</strong>. I began with the solution by Kirill Khoruzhii (@ka1242): 'AeroClub RecSys 2025 - XGBoost Ranking Baseline' (<a href=\"https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars)\" target=\"_blank\">https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars)</a>. I would like to thank him for an excellent set of features.</p>\n<p><strong>Models Used.</strong> I also experimented with using LightGBM both as a replacement for and in addition to XGBoost. However, it yielded roughly the same results, so I eventually discontinued its use by the end of the competition.</p>\n<p><strong>Important Features.</strong> As expected, the most important features turned out to be ticket prices and their derivatives. This includes metrics like the ratio to the median price within a group, among others. Features related to flight duration and the number of connections were also significant. Additionally, features indicating the possibility of ticket cancellation or exchange, as well as the ticket class, proved to be substantial contributors.</p>\n<p><strong>Feature Selection:</strong> Features were selected primarily based on intuition, driven by what we expected to be important for a person when choosing a flight ticket. This initial selection was then validated, and if the assumption about a feature's usefulness was confirmed, it was kept for the final model.</p>\n<p>In several cases, however, adding certain features—particularly categorical ones related to the customers' past choice history—led to severe overfitting. Although using such features seemed intuitively beneficial, they proved to be impractical in practice.</p>\n<p>For example, a factor like the price at which a customer/company previously booked a ticket turned out to be useful. Conversely, trying to incorporate the type of aircraft or the airline carrier a customer/company preferred in the past led to overfitting.</p>\n<p><strong>Overfitting:</strong> Features based on customer choice history led to overfitting. In fact, the first-place competition winner, <a href=\"https://www.kaggle.com/goldenlock\" target=\"_blank\">@goldenlock</a>, stated in his solution <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/aeroclub-recsys-2025/writeups/1st-solution-using-count-based-features</a> that he avoided using them entirely until the competition ended. However, after the competition concluded, he tested the approach I proposed for utilizing a subset of these features and managed to significantly improve the score of his final solution.</p>\n<p><strong>Mitigating Overfitting:</strong> The approach was based on a windowed method for generating customer history features. Specifically, for each group, features were constructed using data only from groups where a ticket was purchased by the same customers prior to the current ticket selection group</p>\n<p><strong>Historical Window Features:</strong> For the features used to calculate the history of option selection within a group, a specific company, or a user, the following metrics were generated: mean, median, and the number of groups in which a choice was previously made and this feature was present. The 25% and 75% quantiles were also generated. The historical features were based on the following features:</p>\n<ul>\n<li><p>'legs0_departureAt_hour', 'legs1_departureAt_hour', 'legs0_arrivalAt_hour', 'legs1_arrivalAt_hour' </p></li>\n<li><p>'price_rank', 'price_from_median', 'duration_rank', 'duration_from_median', </p></li>\n<li><p>'avg_cabin_class', 'baggage_total', 'l0_seg', </p></li>\n<li><p>'legs0_segments0_seatsAvailable', 'legs0_segments0_baggageAllowance_quantity', 'legs0_segments0_cabinClass', </p></li>\n<li><p>'miniRules1_statusInfos', 'miniRules0_statusInfos' </p></li>\n</ul>\n<h2>Training of Models</h2>\n<p><strong>Validation:</strong> When selecting features, the data was split into a training and test part, which was then combined with validation. As I now understand, this was probably a mistake, and validation and test should have been separated.<br>\nAt the same time, the score obtained on validation was almost the same as the one on the Public and Private Leaderboards. Moreover, improvements in local validation usually correlated with improvements on the LB.<br>\nLocal validation yielded about 0.535. For local validation, the last month of training data was used.</p>\n<p><strong>Training:</strong> Training was performed on the full training dataset without any use of test data. This included not using the test set for computing statistics or encoding categorical features.</p>\n<p><strong>Handling New Data in Test:</strong> To prevent the model from being significantly distorted by the appearance of new categories in categorical features, about 1% of categorical values in the training data were randomly replaced with missing values. If the proportion of such missing values was less than 1%, additional ones were added. Then, in the test set, unknown categories in categorical features were also marked as missing values (-1) during categorical encoding.</p>\n<p><strong>Reducing Group Size</strong>: While considering ways to preselect the most promising candidates for the final choice of the three most likely options a client would choose, I tried simply randomly reducing all groups to a maximum of 50 options. This involved downsampling negative examples (those not chosen by the client). Unexpectedly, this improved the score. I hypothesized that ranking became easier for the models, and the class imbalance was reduced. This is because the probability of correctly guessing the user's choice in a large group was low anyway. However, without reducing group size, models might have struggled more to learn anything useful from such a large group.</p>\n<p>A beneficial side effect was the reduction in memory usage during training and an increase in training speed. This, in turn, allowed for the use of more features. It's important to note that the group size reduction was performed after all features were calculated and statistics were collected. That is, the features were generated using the full groups. Reducing the group size before feature engineering led to worse model performance.</p>\n<h2>Possible Improvements</h2>\n<p>It would be beneficial to find a way to train the model without overfitting, not only on numerical data but also on historical categorical data. This was not achieved within the scope of this solution.</p>\n<p>Approaches that generate additional features using language models, or other transformers seem very promising. The idea is to use an attention mechanism to extract meaningful information from the full description of the offered options. <a href=\"https://www.kaggle.com/sergeyqt20244\" target=\"_blank\">@sergeyqt20244</a> mentioned attempts to apply Bert2rec, though so far without significant results: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/competitions/aeroclub-recsys-2025/discussion/599559#3271192</a></p>\n<p>Below is a notebook with my single model solution:<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/mikhailgolubchik/sm-xgboost-single</a></p>",
      "rawMarkdown": "Before diving into the solution description itself, I would first like to thank the competition organizers for an interesting task and for providing real, authentic data.\n\n<h2>Solution Description</h2>\n\n**Baseline**. I began with the solution by Kirill Khoruzhii (@ka1242): 'AeroClub RecSys 2025 - XGBoost Ranking Baseline' (https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars). I would like to thank him for an excellent set of features.\n\n**Models Used.** I also experimented with using LightGBM both as a replacement for and in addition to XGBoost. However, it yielded roughly the same results, so I eventually discontinued its use by the end of the competition.\n\n**Important Features.** As expected, the most important features turned out to be ticket prices and their derivatives. This includes metrics like the ratio to the median price within a group, among others. Features related to flight duration and the number of connections were also significant. Additionally, features indicating the possibility of ticket cancellation or exchange, as well as the ticket class, proved to be substantial contributors.\n\n**Feature Selection:** Features were selected primarily based on intuition, driven by what we expected to be important for a person when choosing a flight ticket. This initial selection was then validated, and if the assumption about a feature's usefulness was confirmed, it was kept for the final model.\n\nIn several cases, however, adding certain features—particularly categorical ones related to the customers' past choice history—led to severe overfitting. Although using such features seemed intuitively beneficial, they proved to be impractical in practice.\n\nFor example, a factor like the price at which a customer/company previously booked a ticket turned out to be useful. Conversely, trying to incorporate the type of aircraft or the airline carrier a customer/company preferred in the past led to overfitting.\n\n**Overfitting:** Features based on customer choice history led to overfitting. In fact, the first-place competition winner, @goldenlock, stated in his solution [https://www.kaggle.com/competitions/aeroclub-recsys-2025/writeups/1st-solution-using-count-based-features](url) that he avoided using them entirely until the competition ended. However, after the competition concluded, he tested the approach I proposed for utilizing a subset of these features and managed to significantly improve the score of his final solution.\n\n**Mitigating Overfitting:** The approach was based on a windowed method for generating customer history features. Specifically, for each group, features were constructed using data only from groups where a ticket was purchased by the same customers prior to the current ticket selection group\n\n**Historical Window Features:** For the features used to calculate the history of option selection within a group, a specific company, or a user, the following metrics were generated: mean, median, and the number of groups in which a choice was previously made and this feature was present. The 25% and 75% quantiles were also generated. The historical features were based on the following features:\n\n- 'legs0_departureAt_hour', 'legs1_departureAt_hour', 'legs0_arrivalAt_hour', 'legs1_arrivalAt_hour' \n\n- 'price_rank', 'price_from_median', 'duration_rank', 'duration_from_median', \n\n- 'avg_cabin_class', 'baggage_total', 'l0_seg', \n\n- 'legs0_segments0_seatsAvailable', 'legs0_segments0_baggageAllowance_quantity', 'legs0_segments0_cabinClass', \n\n- 'miniRules1_statusInfos', 'miniRules0_statusInfos' \n\n<h2>Training of Models</h2>\n\n**Validation:** When selecting features, the data was split into a training and test part, which was then combined with validation. As I now understand, this was probably a mistake, and validation and test should have been separated.\nAt the same time, the score obtained on validation was almost the same as the one on the Public and Private Leaderboards. Moreover, improvements in local validation usually correlated with improvements on the LB.\nLocal validation yielded about 0.535. For local validation, the last month of training data was used.\n\n**Training:** Training was performed on the full training dataset without any use of test data. This included not using the test set for computing statistics or encoding categorical features.\n\n**Handling New Data in Test:** To prevent the model from being significantly distorted by the appearance of new categories in categorical features, about 1% of categorical values in the training data were randomly replaced with missing values. If the proportion of such missing values was less than 1%, additional ones were added. Then, in the test set, unknown categories in categorical features were also marked as missing values (-1) during categorical encoding.\n\n**Reducing Group Size**: While considering ways to preselect the most promising candidates for the final choice of the three most likely options a client would choose, I tried simply randomly reducing all groups to a maximum of 50 options. This involved downsampling negative examples (those not chosen by the client). Unexpectedly, this improved the score. I hypothesized that ranking became easier for the models, and the class imbalance was reduced. This is because the probability of correctly guessing the user's choice in a large group was low anyway. However, without reducing group size, models might have struggled more to learn anything useful from such a large group.\n\nA beneficial side effect was the reduction in memory usage during training and an increase in training speed. This, in turn, allowed for the use of more features. It's important to note that the group size reduction was performed after all features were calculated and statistics were collected. That is, the features were generated using the full groups. Reducing the group size before feature engineering led to worse model performance.\n\n<h2>Possible Improvements</h2>\n\nIt would be beneficial to find a way to train the model without overfitting, not only on numerical data but also on historical categorical data. This was not achieved within the scope of this solution.\n\nApproaches that generate additional features using language models, or other transformers seem very promising. The idea is to use an attention mechanism to extract meaningful information from the full description of the offered options. @sergeyqt20244 mentioned attempts to apply Bert2rec, though so far without significant results: [https://www.kaggle.com/competitions/aeroclub-recsys-2025/discussion/599559#3271192](url)\n\nBelow is a notebook with my single model solution:\n[https://www.kaggle.com/code/mikhailgolubchik/sm-xgboost-single](url)",
      "votes": 4
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3272336": "Before diving into the solution description itself, I would first like to thank the competition organizers for an interesting task and for providing real, authentic data.\n\n<h2>Solution Description</h2>\n\n**Baseline**. I began with the solution by Kirill Khoruzhii (@ka1242): 'AeroClub RecSys 2025 - XGBoost Ranking Baseline' (https://www.kaggle.com/code/ka1242/xgboost-ranker-with-polars). I would like to thank him for an excellent set of features.\n\n**Models Used.** I also experimented with using LightGBM both as a replacement for and in addition to XGBoost. However, it yielded roughly the same results, so I eventually discontinued its use by the end of the competition.\n\n**Important Features.** As expected, the most important features turned out to be ticket prices and their derivatives. This includes metrics like the ratio to the median price within a group, among others. Features related to flight duration and the number of connections were also significant. Additionally, features indicating the possibility of ticket cancellation or exchange, as well as the ticket class, proved to be substantial contributors.\n\n**Feature Selection:** Features were selected primarily based on intuition, driven by what we expected to be important for a person when choosing a flight ticket. This initial selection was then validated, and if the assumption about a feature's usefulness was confirmed, it was kept for the final model.\n\nIn several cases, however, adding certain features—particularly categorical ones related to the customers' past choice history—led to severe overfitting. Although using such features seemed intuitively beneficial, they proved to be impractical in practice.\n\nFor example, a factor like the price at which a customer/company previously booked a ticket turned out to be useful. Conversely, trying to incorporate the type of aircraft or the airline carrier a customer/company preferred in the past led to overfitting.\n\n**Overfitting:** Features based on customer choice history led to overfitting. In fact, the first-place competition winner, @goldenlock, stated in his solution [https://www.kaggle.com/competitions/aeroclub-recsys-2025/writeups/1st-solution-using-count-based-features](url) that he avoided using them entirely until the competition ended. However, after the competition concluded, he tested the approach I proposed for utilizing a subset of these features and managed to significantly improve the score of his final solution.\n\n**Mitigating Overfitting:** The approach was based on a windowed method for generating customer history features. Specifically, for each group, features were constructed using data only from groups where a ticket was purchased by the same customers prior to the current ticket selection group\n\n**Historical Window Features:** For the features used to calculate the history of option selection within a group, a specific company, or a user, the following metrics were generated: mean, median, and the number of groups in which a choice was previously made and this feature was present. The 25% and 75% quantiles were also generated. The historical features were based on the following features:\n\n- 'legs0_departureAt_hour', 'legs1_departureAt_hour', 'legs0_arrivalAt_hour', 'legs1_arrivalAt_hour' \n\n- 'price_rank', 'price_from_median', 'duration_rank', 'duration_from_median', \n\n- 'avg_cabin_class', 'baggage_total', 'l0_seg', \n\n- 'legs0_segments0_seatsAvailable', 'legs0_segments0_baggageAllowance_quantity', 'legs0_segments0_cabinClass', \n\n- 'miniRules1_statusInfos', 'miniRules0_statusInfos' \n\n<h2>Training of Models</h2>\n\n**Validation:** When selecting features, the data was split into a training and test part, which was then combined with validation. As I now understand, this was probably a mistake, and validation and test should have been separated.\nAt the same time, the score obtained on validation was almost the same as the one on the Public and Private Leaderboards. Moreover, improvements in local validation usually correlated with improvements on the LB.\nLocal validation yielded about 0.535. For local validation, the last month of training data was used.\n\n**Training:** Training was performed on the full training dataset without any use of test data. This included not using the test set for computing statistics or encoding categorical features.\n\n**Handling New Data in Test:** To prevent the model from being significantly distorted by the appearance of new categories in categorical features, about 1% of categorical values in the training data were randomly replaced with missing values. If the proportion of such missing values was less than 1%, additional ones were added. Then, in the test set, unknown categories in categorical features were also marked as missing values (-1) during categorical encoding.\n\n**Reducing Group Size**: While considering ways to preselect the most promising candidates for the final choice of the three most likely options a client would choose, I tried simply randomly reducing all groups to a maximum of 50 options. This involved downsampling negative examples (those not chosen by the client). Unexpectedly, this improved the score. I hypothesized that ranking became easier for the models, and the class imbalance was reduced. This is because the probability of correctly guessing the user's choice in a large group was low anyway. However, without reducing group size, models might have struggled more to learn anything useful from such a large group.\n\nA beneficial side effect was the reduction in memory usage during training and an increase in training speed. This, in turn, allowed for the use of more features. It's important to note that the group size reduction was performed after all features were calculated and statistics were collected. That is, the features were generated using the full groups. Reducing the group size before feature engineering led to worse model performance.\n\n<h2>Possible Improvements</h2>\n\nIt would be beneficial to find a way to train the model without overfitting, not only on numerical data but also on historical categorical data. This was not achieved within the scope of this solution.\n\nApproaches that generate additional features using language models, or other transformers seem very promising. The idea is to use an attention mechanism to extract meaningful information from the full description of the offered options. @sergeyqt20244 mentioned attempts to apply Bert2rec, though so far without significant results: [https://www.kaggle.com/competitions/aeroclub-recsys-2025/discussion/599559#3271192](url)\n\nBelow is a notebook with my single model solution:\n[https://www.kaggle.com/code/mikhailgolubchik/sm-xgboost-single](url)"
  }
}