{
  "id": 364375,
  "title": "💡 Do not disregard longer sessions -- they contribute disproportionately to the competition metric!",
  "url": "/competitions/otto-recommender-system/discussion/364375",
  "author_name": "",
  "post_date": "2022-11-06T06:12:38.644017200Z",
  "votes": 23,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hey!</p>\n<p>Very often during modeling we might want to throw out data that is not convenient to work with.</p>\n<p>An excellent example of this is shortening sentences for NLP. We might have a lot of sentences with some word count, and then just a handful of sentences that are extremely long. If we are training a language model, we might be tempted to discard longer sequences.</p>\n<p>In this competition, we have a similar scenario. I emulated how a test set is likely to have been constructed. You can find the methodology in the <a href=\"https://www.kaggle.com/code/radek1/a-robust-local-validation-framework\" target=\"_blank\">💡A robust local validation framework 🚀🚀🚀 notebook</a>.</p>\n<p>Here is how input lengths might break down in test (with the column on the right being percentage of all input sequences).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F45b4dd7c2919349e7ea08aef42d74014%2Finput_length_in_test.png?generation=1667714619896082&amp;alt=media\" alt=\"\"></p>\n<p>We can see that the vast majority of sessions are short. We might want to discard the longer sequences.</p>\n<p>But that would be a grave mistake!</p>\n<p>This is again based on the input lengths we are likely to see in test. See how disproportionately longer sessions result in orders, the highest-valued predictions.</p>\n<p>From the competition description we know that this is how results are weighted: <code>{'clicks': 0.10, 'carts': 0.30, 'orders': 0.60}</code></p>\n<p>And now let's look at how <code>carts</code> are distributed by input length in the emulated test set</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F053d4bbab9f6c5487a6eeb58bb16eebf%2Finput_length_in_test_carts.png?generation=1667714909209643&amp;alt=media\" alt=\"\"></p>\n<p>And here is how the situation looks like for <code>orders</code> (the disproportion is even greater!)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fbe073f33d0144199d0eb4011270fdca1%2Finput_length_in_test_orders.png?generation=1667714972605857&amp;alt=media\" alt=\"\"></p>\n<p>This is not surprising in some sense. We might expect people who buy to spend more time on the website. Still, if we follow standard modeling practices and disregard those longer sequences, we would be severely limiting our potential in this competition!</p>\n<p>Those longer sequences contribute to a disproportionate degree to the overall score!</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": "2018894",
      "postDate": "11/06/2022 06:12:38",
      "content": "<p>Hey!</p>\n<p>Very often during modeling we might want to throw out data that is not convenient to work with.</p>\n<p>An excellent example of this is shortening sentences for NLP. We might have a lot of sentences with some word count, and then just a handful of sentences that are extremely long. If we are training a language model, we might be tempted to discard longer sequences.</p>\n<p>In this competition, we have a similar scenario. I emulated how a test set is likely to have been constructed. You can find the methodology in the <a href=\"https://www.kaggle.com/code/radek1/a-robust-local-validation-framework\" target=\"_blank\">💡A robust local validation framework 🚀🚀🚀 notebook</a>.</p>\n<p>Here is how input lengths might break down in test (with the column on the right being percentage of all input sequences).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F45b4dd7c2919349e7ea08aef42d74014%2Finput_length_in_test.png?generation=1667714619896082&amp;alt=media\" alt=\"\"></p>\n<p>We can see that the vast majority of sessions are short. We might want to discard the longer sequences.</p>\n<p>But that would be a grave mistake!</p>\n<p>This is again based on the input lengths we are likely to see in test. See how disproportionately longer sessions result in orders, the highest-valued predictions.</p>\n<p>From the competition description we know that this is how results are weighted: <code>{'clicks': 0.10, 'carts': 0.30, 'orders': 0.60}</code></p>\n<p>And now let's look at how <code>carts</code> are distributed by input length in the emulated test set</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F053d4bbab9f6c5487a6eeb58bb16eebf%2Finput_length_in_test_carts.png?generation=1667714909209643&amp;alt=media\" alt=\"\"></p>\n<p>And here is how the situation looks like for <code>orders</code> (the disproportion is even greater!)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fbe073f33d0144199d0eb4011270fdca1%2Finput_length_in_test_orders.png?generation=1667714972605857&amp;alt=media\" alt=\"\"></p>\n<p>This is not surprising in some sense. We might expect people who buy to spend more time on the website. Still, if we follow standard modeling practices and disregard those longer sequences, we would be severely limiting our potential in this competition!</p>\n<p>Those longer sequences contribute to a disproportionate degree to the overall score!</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "Hey!\n\nVery often during modeling we might want to throw out data that is not convenient to work with.\n\nAn excellent example of this is shortening sentences for NLP. We might have a lot of sentences with some word count, and then just a handful of sentences that are extremely long. If we are training a language model, we might be tempted to discard longer sequences.\n\nIn this competition, we have a similar scenario. I emulated how a test set is likely to have been constructed. You can find the methodology in the [💡A robust local validation framework 🚀🚀🚀 notebook](https://www.kaggle.com/code/radek1/a-robust-local-validation-framework).\n\nHere is how input lengths might break down in test (with the column on the right being percentage of all input sequences).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F45b4dd7c2919349e7ea08aef42d74014%2Finput_length_in_test.png?generation=1667714619896082&alt=media)\n\nWe can see that the vast majority of sessions are short. We might want to discard the longer sequences.\n\nBut that would be a grave mistake!\n\nThis is again based on the input lengths we are likely to see in test. See how disproportionately longer sessions result in orders, the highest-valued predictions.\n\nFrom the competition description we know that this is how results are weighted: `{'clicks': 0.10, 'carts': 0.30, 'orders': 0.60}`\n\nAnd now let's look at how `carts` are distributed by input length in the emulated test set\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F053d4bbab9f6c5487a6eeb58bb16eebf%2Finput_length_in_test_carts.png?generation=1667714909209643&alt=media)\n\nAnd here is how the situation looks like for `orders` (the disproportion is even greater!)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fbe073f33d0144199d0eb4011270fdca1%2Finput_length_in_test_orders.png?generation=1667714972605857&alt=media)\n\nThis is not surprising in some sense. We might expect people who buy to spend more time on the website. Still, if we follow standard modeling practices and disregard those longer sequences, we would be severely limiting our potential in this competition!\n\nThose longer sequences contribute to a disproportionate degree to the overall score!\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": null
    },
    {
      "id": "2033760",
      "postDate": "11/17/2022 14:02:52",
      "content": "<p>Thanks so much <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> for this wonderful post! This is very helpful!</p>\n<p>Here is how I understand your post, could you help me verify? </p>\n<ol>\n<li>\"longer sequences\" refer to a lot more events in a session like \"25-10000\" events (clicks, carts, or orders)</li>\n</ol>\n<p>a question here, what do you consider to be a long session?  \"25-10000\" is surely very very long, how about 10-25 and 5-10? do you consider them to be long?</p>\n<ol>\n<li><p>A session with more events not only provide more info to generate possible predictions but also indicate the user is more likely to put products into cart and even order.</p></li>\n<li><p>because of the point above and metric using weights which take carts and orders more seriously {'clicks': 0.10, 'carts': 0.30, 'orders': 0.60}, these longer sessions will help us score higher. Therefore we shouldn't discard but treasure longer sessions. </p></li>\n<li><p>especially when longer sessions on carts and orders have proportions which is not tiny as you showed in your charts above</p></li>\n</ol>",
      "rawMarkdown": "Thanks so much @radek1 for this wonderful post! This is very helpful!\n\nHere is how I understand your post, could you help me verify? \n\n1. \"longer sequences\" refer to a lot more events in a session like \"25-10000\" events (clicks, carts, or orders)\n\na question here, what do you consider to be a long session?  \"25-10000\" is surely very very long, how about 10-25 and 5-10? do you consider them to be long?\n\n2. A session with more events not only provide more info to generate possible predictions but also indicate the user is more likely to put products into cart and even order.\n\n3. because of the point above and metric using weights which take carts and orders more seriously {'clicks': 0.10, 'carts': 0.30, 'orders': 0.60}, these longer sessions will help us score higher. Therefore we shouldn't discard but treasure longer sessions. \n\n4. especially when longer sessions on carts and orders have proportions which is not tiny as you showed in your charts above",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2033760,
      "author_name": "danielliao",
      "author_url": "",
      "post_date": "11/17/2022 14:02:52",
      "content": "<p>Thanks so much <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> for this wonderful post! This is very helpful!</p>\n<p>Here is how I understand your post, could you help me verify? </p>\n<ol>\n<li>\"longer sequences\" refer to a lot more events in a session like \"25-10000\" events (clicks, carts, or orders)</li>\n</ol>\n<p>a question here, what do you consider to be a long session?  \"25-10000\" is surely very very long, how about 10-25 and 5-10? do you consider them to be long?</p>\n<ol>\n<li><p>A session with more events not only provide more info to generate possible predictions but also indicate the user is more likely to put products into cart and even order.</p></li>\n<li><p>because of the point above and metric using weights which take carts and orders more seriously {'clicks': 0.10, 'carts': 0.30, 'orders': 0.60}, these longer sessions will help us score higher. Therefore we shouldn't discard but treasure longer sessions. </p></li>\n<li><p>especially when longer sessions on carts and orders have proportions which is not tiny as you showed in your charts above</p></li>\n</ol>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2018894": "Hey!\n\nVery often during modeling we might want to throw out data that is not convenient to work with.\n\nAn excellent example of this is shortening sentences for NLP. We might have a lot of sentences with some word count, and then just a handful of sentences that are extremely long. If we are training a language model, we might be tempted to discard longer sequences.\n\nIn this competition, we have a similar scenario. I emulated how a test set is likely to have been constructed. You can find the methodology in the [💡A robust local validation framework 🚀🚀🚀 notebook](https://www.kaggle.com/code/radek1/a-robust-local-validation-framework).\n\nHere is how input lengths might break down in test (with the column on the right being percentage of all input sequences).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F45b4dd7c2919349e7ea08aef42d74014%2Finput_length_in_test.png?generation=1667714619896082&alt=media)\n\nWe can see that the vast majority of sessions are short. We might want to discard the longer sequences.\n\nBut that would be a grave mistake!\n\nThis is again based on the input lengths we are likely to see in test. See how disproportionately longer sessions result in orders, the highest-valued predictions.\n\nFrom the competition description we know that this is how results are weighted: `{'clicks': 0.10, 'carts': 0.30, 'orders': 0.60}`\n\nAnd now let's look at how `carts` are distributed by input length in the emulated test set\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F053d4bbab9f6c5487a6eeb58bb16eebf%2Finput_length_in_test_carts.png?generation=1667714909209643&alt=media)\n\nAnd here is how the situation looks like for `orders` (the disproportion is even greater!)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2Fbe073f33d0144199d0eb4011270fdca1%2Finput_length_in_test_orders.png?generation=1667714972605857&alt=media)\n\nThis is not surprising in some sense. We might expect people who buy to spend more time on the website. Still, if we follow standard modeling practices and disregard those longer sequences, we would be severely limiting our potential in this competition!\n\nThose longer sequences contribute to a disproportionate degree to the overall score!\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
    "2033760": "Thanks so much @radek1 for this wonderful post! This is very helpful!\n\nHere is how I understand your post, could you help me verify? \n\n1. \"longer sequences\" refer to a lot more events in a session like \"25-10000\" events (clicks, carts, or orders)\n\na question here, what do you consider to be a long session?  \"25-10000\" is surely very very long, how about 10-25 and 5-10? do you consider them to be long?\n\n2. A session with more events not only provide more info to generate possible predictions but also indicate the user is more likely to put products into cart and even order.\n\n3. because of the point above and metric using weights which take carts and orders more seriously {'clicks': 0.10, 'carts': 0.30, 'orders': 0.60}, these longer sessions will help us score higher. Therefore we shouldn't discard but treasure longer sessions. \n\n4. especially when longer sessions on carts and orders have proportions which is not tiny as you showed in your charts above"
  },
  "source": "meta"
}