{
  "id": 586149,
  "title": "Ideas Kitchen",
  "url": "/competitions/aeroclub-recsys-2025/discussion/586149",
  "author_name": "",
  "post_date": "2025-06-25T11:34:18.330354300Z",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>In this topic, I'd like to discuss ideas, RecSys model architectures and approaches worth trying in the current competition. These will help us move beyond the \"create features and feed them to boosting\" approach that clearly dominated public notebooks in the early days of the competition.</p>\n<p>One of the most crucial stages in recommendations for this task is candidate retrieval of items, among which our ranking models will select the best of the best. Think of this as a funnel aimed at reducing the number of possible options for ranking from thousands to hundreds / tens. The fewer variants we need to rank, the higher the probability of achieving better results.</p>\n<p>Consider to try a two-stage approach: fast selection of top-N candidates in the retrieval stage, then precise ranking of final candidates in the ranking stage. This allows for scalability and using more complex models in the second stage.</p>\n<p>The competition's specificity lies in the fact that our items are very similar to each other while simultaneously existing in several interconnected dimensions: route, date/time, carrier, conditions, etc. Think about how and by which characteristics you can select candidates, forming this funnel.</p>\n<p>We also have sparse and implicit signals. This might be helped by negative sampling, collapsing items into higher-level categories, popularity-based fallbacks, temporal filters.</p>\n<p>Users next. It's important to understand that we can think about users from several directions. For example, the famous 2024 paper \"Actions Speak Louder than Words\" suggests an idea to think about users as sequences of their actions. This is an extremely interesting and working idea, however, this competition doesn't provide much action data. It probably won't fit here. *<em>We considered to add clickstream data during the competition, but decided to postpone this idea until the next competition in the series</em>.</p>\n<p>Nevertheless, in the current data space you can try classic collaborative filtering methods, user2user and item2item matrices, experiment with user similarities based on their features and actions. One idea, for example, is to move from user level to company level.</p>\n<p>I encourage you to share your thoughts, links and ideas. In RecSys, there's still no silver bullet or universal solution for all situations. We have impressive RecSys research papers and lack of open implementations.  We'd really like to move beyond \"popularity bias\" as a most descent solution :)</p>",
  "messages": [
    {
      "id": "3232110",
      "postDate": "06/25/2025 11:34:18",
      "content": "<p>In this topic, I'd like to discuss ideas, RecSys model architectures and approaches worth trying in the current competition. These will help us move beyond the \"create features and feed them to boosting\" approach that clearly dominated public notebooks in the early days of the competition.</p>\n<p>One of the most crucial stages in recommendations for this task is candidate retrieval of items, among which our ranking models will select the best of the best. Think of this as a funnel aimed at reducing the number of possible options for ranking from thousands to hundreds / tens. The fewer variants we need to rank, the higher the probability of achieving better results.</p>\n<p>Consider to try a two-stage approach: fast selection of top-N candidates in the retrieval stage, then precise ranking of final candidates in the ranking stage. This allows for scalability and using more complex models in the second stage.</p>\n<p>The competition's specificity lies in the fact that our items are very similar to each other while simultaneously existing in several interconnected dimensions: route, date/time, carrier, conditions, etc. Think about how and by which characteristics you can select candidates, forming this funnel.</p>\n<p>We also have sparse and implicit signals. This might be helped by negative sampling, collapsing items into higher-level categories, popularity-based fallbacks, temporal filters.</p>\n<p>Users next. It's important to understand that we can think about users from several directions. For example, the famous 2024 paper \"Actions Speak Louder than Words\" suggests an idea to think about users as sequences of their actions. This is an extremely interesting and working idea, however, this competition doesn't provide much action data. It probably won't fit here. *<em>We considered to add clickstream data during the competition, but decided to postpone this idea until the next competition in the series</em>.</p>\n<p>Nevertheless, in the current data space you can try classic collaborative filtering methods, user2user and item2item matrices, experiment with user similarities based on their features and actions. One idea, for example, is to move from user level to company level.</p>\n<p>I encourage you to share your thoughts, links and ideas. In RecSys, there's still no silver bullet or universal solution for all situations. We have impressive RecSys research papers and lack of open implementations.  We'd really like to move beyond \"popularity bias\" as a most descent solution :)</p>",
      "rawMarkdown": "In this topic, I'd like to discuss ideas, RecSys model architectures and approaches worth trying in the current competition. These will help us move beyond the \"create features and feed them to boosting\" approach that clearly dominated public notebooks in the early days of the competition.\n\nOne of the most crucial stages in recommendations for this task is candidate retrieval of items, among which our ranking models will select the best of the best. Think of this as a funnel aimed at reducing the number of possible options for ranking from thousands to hundreds / tens. The fewer variants we need to rank, the higher the probability of achieving better results.\n\nConsider to try a two-stage approach: fast selection of top-N candidates in the retrieval stage, then precise ranking of final candidates in the ranking stage. This allows for scalability and using more complex models in the second stage.\n\nThe competition's specificity lies in the fact that our items are very similar to each other while simultaneously existing in several interconnected dimensions: route, date/time, carrier, conditions, etc. Think about how and by which characteristics you can select candidates, forming this funnel.\n\nWe also have sparse and implicit signals. This might be helped by negative sampling, collapsing items into higher-level categories, popularity-based fallbacks, temporal filters.\n\nUsers next. It's important to understand that we can think about users from several directions. For example, the famous 2024 paper \"Actions Speak Louder than Words\" suggests an idea to think about users as sequences of their actions. This is an extremely interesting and working idea, however, this competition doesn't provide much action data. It probably won't fit here. **We considered to add clickstream data during the competition, but decided to postpone this idea until the next competition in the series*.\n\nNevertheless, in the current data space you can try classic collaborative filtering methods, user2user and item2item matrices, experiment with user similarities based on their features and actions. One idea, for example, is to move from user level to company level.\n\nI encourage you to share your thoughts, links and ideas. In RecSys, there's still no silver bullet or universal solution for all situations. We have impressive RecSys research papers and lack of open implementations.  We'd really like to move beyond \"popularity bias\" as a most descent solution :)",
      "votes": null
    },
    {
      "id": "3233760",
      "postDate": "06/27/2025 06:09:28",
      "content": "<p>Just to strat discussion, congratulation to <a href=\"https://www.kaggle.com/insuperabile\" target=\"_blank\">@insuperabile</a> for crossing 0.5!!!!!!!    </p>\n<p>As far as I can see from open notebooks till now everyone is using \"classical\" gbm ranking and ( I am far from being expert in ml and recsys) it seems to me that +/- 0.5 is limit for this.  It is of cause 2 times better than dummy model and this is good, but it is far away from 0.7.   So need to use some other ideas for sure. </p>",
      "rawMarkdown": "Just to strat discussion, congratulation to @insuperabile for crossing 0.5!!!!!!!    \n\nAs far as I can see from open notebooks till now everyone is using \"classical\" gbm ranking and ( I am far from being expert in ml and recsys) it seems to me that +/- 0.5 is limit for this.  It is of cause 2 times better than dummy model and this is good, but it is far away from 0.7.   So need to use some other ideas for sure.",
      "votes": null
    },
    {
      "id": "3234161",
      "postDate": "06/27/2025 15:07:07",
      "content": "<p>thanks but i think there is a lot of feature engineering ideas more, and I didn't even look at json, maybe there's something there too.<br>\nfor 0.5 i did actually almost nothing)</p>",
      "rawMarkdown": "thanks but i think there is a lot of feature engineering ideas more, and I didn't even look at json, maybe there's something there too.\nfor 0.5 i did actually almost nothing)",
      "votes": null
    },
    {
      "id": "3242386",
      "postDate": "07/05/2025 20:13:46",
      "content": "<p>my score reduced after reducing candidates😂</p>",
      "rawMarkdown": "my score reduced after reducing candidates😂",
      "votes": null
    },
    {
      "id": "3242654",
      "postDate": "07/06/2025 07:28:10",
      "content": "<p>Actually I think it is useless here in this task. You have only max 300 candidates. Every model can handle them. Of casue idea is goid and there is a desire to apply it. But… The reduce stage is necessary when you have 1000000 candidates like fir example if you propose youtube videos to user. </p>",
      "rawMarkdown": "Actually I think it is useless here in this task. You have only max 300 candidates. Every model can handle them. Of casue idea is goid and there is a desire to apply it. But... The reduce stage is necessary when you have 1000000 candidates like fir example if you propose youtube videos to user.",
      "votes": null
    },
    {
      "id": "3277916",
      "postDate": "08/28/2025 22:50:23",
      "content": "<p>A solid RecSys approach here is a two-stage system: fast candidate retrieval then precise ranking, since items are so similar. Just like <a href=\"https://kitchenremodelingbaltimore.net/\" target=\"_blank\">kitchen remodeling Baltimore</a><br>\n projects, breaking work into stages gives better results.</p>",
      "rawMarkdown": "A solid RecSys approach here is a two-stage system: fast candidate retrieval then precise ranking, since items are so similar. Just like [kitchen remodeling Baltimore](https://kitchenremodelingbaltimore.net/)\n projects, breaking work into stages gives better results.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3233760,
      "author_name": "sergeyqt2024",
      "author_url": "",
      "post_date": "06/27/2025 06:09:28",
      "content": "<p>Just to strat discussion, congratulation to <a href=\"https://www.kaggle.com/insuperabile\" target=\"_blank\">@insuperabile</a> for crossing 0.5!!!!!!!    </p>\n<p>As far as I can see from open notebooks till now everyone is using \"classical\" gbm ranking and ( I am far from being expert in ml and recsys) it seems to me that +/- 0.5 is limit for this.  It is of cause 2 times better than dummy model and this is good, but it is far away from 0.7.   So need to use some other ideas for sure. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3234161,
          "author_name": "insuperabile",
          "author_url": "",
          "post_date": "06/27/2025 15:07:07",
          "content": "<p>thanks but i think there is a lot of feature engineering ideas more, and I didn't even look at json, maybe there's something there too.<br>\nfor 0.5 i did actually almost nothing)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3242386,
      "author_name": "meetbabariya",
      "author_url": "",
      "post_date": "07/05/2025 20:13:46",
      "content": "<p>my score reduced after reducing candidates😂</p>",
      "votes": null,
      "replies": [
        {
          "id": 3242654,
          "author_name": "sergeyqt2024",
          "author_url": "",
          "post_date": "07/06/2025 07:28:10",
          "content": "<p>Actually I think it is useless here in this task. You have only max 300 candidates. Every model can handle them. Of casue idea is goid and there is a desire to apply it. But… The reduce stage is necessary when you have 1000000 candidates like fir example if you propose youtube videos to user. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3277916,
      "author_name": "abdulhadi898",
      "author_url": "",
      "post_date": "08/28/2025 22:50:23",
      "content": "<p>A solid RecSys approach here is a two-stage system: fast candidate retrieval then precise ranking, since items are so similar. Just like <a href=\"https://kitchenremodelingbaltimore.net/\" target=\"_blank\">kitchen remodeling Baltimore</a><br>\n projects, breaking work into stages gives better results.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3232110": "In this topic, I'd like to discuss ideas, RecSys model architectures and approaches worth trying in the current competition. These will help us move beyond the \"create features and feed them to boosting\" approach that clearly dominated public notebooks in the early days of the competition.\n\nOne of the most crucial stages in recommendations for this task is candidate retrieval of items, among which our ranking models will select the best of the best. Think of this as a funnel aimed at reducing the number of possible options for ranking from thousands to hundreds / tens. The fewer variants we need to rank, the higher the probability of achieving better results.\n\nConsider to try a two-stage approach: fast selection of top-N candidates in the retrieval stage, then precise ranking of final candidates in the ranking stage. This allows for scalability and using more complex models in the second stage.\n\nThe competition's specificity lies in the fact that our items are very similar to each other while simultaneously existing in several interconnected dimensions: route, date/time, carrier, conditions, etc. Think about how and by which characteristics you can select candidates, forming this funnel.\n\nWe also have sparse and implicit signals. This might be helped by negative sampling, collapsing items into higher-level categories, popularity-based fallbacks, temporal filters.\n\nUsers next. It's important to understand that we can think about users from several directions. For example, the famous 2024 paper \"Actions Speak Louder than Words\" suggests an idea to think about users as sequences of their actions. This is an extremely interesting and working idea, however, this competition doesn't provide much action data. It probably won't fit here. **We considered to add clickstream data during the competition, but decided to postpone this idea until the next competition in the series*.\n\nNevertheless, in the current data space you can try classic collaborative filtering methods, user2user and item2item matrices, experiment with user similarities based on their features and actions. One idea, for example, is to move from user level to company level.\n\nI encourage you to share your thoughts, links and ideas. In RecSys, there's still no silver bullet or universal solution for all situations. We have impressive RecSys research papers and lack of open implementations.  We'd really like to move beyond \"popularity bias\" as a most descent solution :)",
    "3233760": "Just to strat discussion, congratulation to @insuperabile for crossing 0.5!!!!!!!    \n\nAs far as I can see from open notebooks till now everyone is using \"classical\" gbm ranking and ( I am far from being expert in ml and recsys) it seems to me that +/- 0.5 is limit for this.  It is of cause 2 times better than dummy model and this is good, but it is far away from 0.7.   So need to use some other ideas for sure.",
    "3234161": "thanks but i think there is a lot of feature engineering ideas more, and I didn't even look at json, maybe there's something there too.\nfor 0.5 i did actually almost nothing)",
    "3242386": "my score reduced after reducing candidates😂",
    "3242654": "Actually I think it is useless here in this task. You have only max 300 candidates. Every model can handle them. Of casue idea is goid and there is a desire to apply it. But... The reduce stage is necessary when you have 1000000 candidates like fir example if you propose youtube videos to user.",
    "3277916": "A solid RecSys approach here is a two-stage system: fast candidate retrieval then precise ranking, since items are so similar. Just like [kitchen remodeling Baltimore](https://kitchenremodelingbaltimore.net/)\n projects, breaking work into stages gives better results."
  },
  "source": "meta"
}