{
  "id": 363965,
  "title": "Important information regarding test data from competition repository",
  "url": "/competitions/otto-recommender-system/discussion/363965",
  "author_name": "",
  "post_date": "2022-11-03T21:18:26.283065100Z",
  "votes": 21,
  "comment_count": 12,
  "views": 0,
  "content": "<p>There is an associated repository with a good README and some related code. Here is <a href=\"https://github.com/otto-de/recsys-dataset\" target=\"_blank\">the repo</a>.</p>\n<p>There is important information regarding the test set posted there that I have not come across here on Kaggle! Please take a look at the figure below<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F05ce69f1c8516ac1cbd1bc4f2730fd02%2Fsessions.png?generation=1667510035134467&amp;alt=media\" alt=\"\"></p>\n<p>The important piece of information is this -- the train session data that overlap with the test period was discarded! That makes sense, that is what we would expect.</p>\n<p>But the key piece of information is this -- all the sessions in the test data don't have truncated beginnings! Essentially, they are complete sessions, they are not residuals from sessions that started in the train session period.</p>\n<p>This is super important IMHO as it should impact how we think about the problem and might impact how we work with the data! Essentially, the train sessions might be truncated from the right, but may not be. The test sessions are guaranteed to not be truncated from the left (which is not obvious from the information on Kaggle) and are guaranteed to be truncated from the right (that is where the \"labels\" or should I rather say ground truth for the test data comes from!)</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": "2016239",
      "postDate": "11/03/2022 21:18:26",
      "content": "<p>There is an associated repository with a good README and some related code. Here is <a href=\"https://github.com/otto-de/recsys-dataset\" target=\"_blank\">the repo</a>.</p>\n<p>There is important information regarding the test set posted there that I have not come across here on Kaggle! Please take a look at the figure below<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F05ce69f1c8516ac1cbd1bc4f2730fd02%2Fsessions.png?generation=1667510035134467&amp;alt=media\" alt=\"\"></p>\n<p>The important piece of information is this -- the train session data that overlap with the test period was discarded! That makes sense, that is what we would expect.</p>\n<p>But the key piece of information is this -- all the sessions in the test data don't have truncated beginnings! Essentially, they are complete sessions, they are not residuals from sessions that started in the train session period.</p>\n<p>This is super important IMHO as it should impact how we think about the problem and might impact how we work with the data! Essentially, the train sessions might be truncated from the right, but may not be. The test sessions are guaranteed to not be truncated from the left (which is not obvious from the information on Kaggle) and are guaranteed to be truncated from the right (that is where the \"labels\" or should I rather say ground truth for the test data comes from!)</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "There is an associated repository with a good README and some related code. Here is [the repo](https://github.com/otto-de/recsys-dataset).\n\nThere is important information regarding the test set posted there that I have not come across here on Kaggle! Please take a look at the figure below\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F05ce69f1c8516ac1cbd1bc4f2730fd02%2Fsessions.png?generation=1667510035134467&alt=media)\n\nThe important piece of information is this -- the train session data that overlap with the test period was discarded! That makes sense, that is what we would expect.\n\nBut the key piece of information is this -- all the sessions in the test data don't have truncated beginnings! Essentially, they are complete sessions, they are not residuals from sessions that started in the train session period.\n\nThis is super important IMHO as it should impact how we think about the problem and might impact how we work with the data! Essentially, the train sessions might be truncated from the right, but may not be. The test sessions are guaranteed to not be truncated from the left (which is not obvious from the information on Kaggle) and are guaranteed to be truncated from the right (that is where the \"labels\" or should I rather say ground truth for the test data comes from!)\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": null
    },
    {
      "id": "2016269",
      "postDate": "11/03/2022 22:07:01",
      "content": "<p>To clarify: as I mentioned in my other <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363874#2016257\" target=\"_blank\">comment</a>, the ground truth also comes from the test week period and events that users created after the end of the week are not included in the evaluation. Please refer to the <a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py\" target=\"_blank\">testset.py</a> script in our public GitHub repository for more details.</p>",
      "rawMarkdown": "To clarify: as I mentioned in my other [comment](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363874#2016257), the ground truth also comes from the test week period and events that users created after the end of the week are not included in the evaluation. Please refer to the [testset.py](https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py) script in our public GitHub repository for more details.",
      "votes": null
    },
    {
      "id": "2016326",
      "postDate": "11/03/2022 23:21:21",
      "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> for this important distinction! Appreciate the clarification! 🙏</p>",
      "rawMarkdown": "Thank you very much @pnormann for this important distinction! Appreciate the clarification! 🙏",
      "votes": null
    },
    {
      "id": "2016398",
      "postDate": "11/04/2022 01:17:59",
      "content": "<p><a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> apologies, if I may please ask you. If there two the same AIDs in ground truth as 'cart', how are the hits calculated?</p>\n<p>If I predict two the same AIDs that happen two be in ground truth, do they count as two hits? Or would I get the same result if I predict just a single AID no matter how many times they appear in ground truth?</p>",
      "rawMarkdown": "pnormann apologies, if I may please ask you. If there two the same AIDs in ground truth as 'cart', how are the hits calculated?\n\nIf I predict two the same AIDs that happen two be in ground truth, do they count as two hits? Or would I get the same result if I predict just a single AID no matter how many times they appear in ground truth?",
      "votes": null
    },
    {
      "id": "2033364",
      "postDate": "11/17/2022 07:47:38",
      "content": "<p>i think duplicate aid in carts will not affect the result, as the evaluation code use \"set\"<br>\n<code>if 'carts' in labels and labels['carts']:\n        cart_hits = len(set(prediction['carts'][:k]).intersection(labels['carts']))\n    else:\n        cart_hits = None</code></p>",
      "rawMarkdown": "i think duplicate aid in carts will not affect the result, as the evaluation code use \"set\"\n`    if 'carts' in labels and labels['carts']:\n        cart_hits = len(set(prediction['carts'][:k]).intersection(labels['carts']))\n    else:\n        cart_hits = None`",
      "votes": null
    },
    {
      "id": "2033691",
      "postDate": "11/17/2022 13:02:12",
      "content": "<p>Thanks for sharing the script and here is the <a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/evaluate.py#L44\" target=\"_blank\">link to code</a>. </p>\n<p>as for <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>'s question, according to the src line above, when the ground truth has two duplicated aids, and Radek's prediction also has the aid appeared twice, that aid only count as one hit. </p>\n<pre><code>([, , ]).intersection([, , ]) \n\n</code></pre>",
      "rawMarkdown": "Thanks for sharing the script and here is the [link to code](https://github.com/otto-de/recsys-dataset/blob/main/src/evaluate.py#L44). \n\nas for @radek1's question, according to the src line above, when the ground truth has two duplicated aids, and Radek's prediction also has the aid appeared twice, that aid only count as one hit. \n\n```python\nset(['a', 'b', 'b']).intersection(['a', 'b', 'b']) \n# return {'a', 'b'}\n```",
      "votes": null
    },
    {
      "id": "2059540",
      "postDate": "12/09/2022 02:03:18",
      "content": "<p>Thanks for the info.<br>\nI’m wondering what the y-axis means. The users are different in train and test, right?</p>",
      "rawMarkdown": "Thanks for the info.\nI’m wondering what the y-axis means. The users are different in train and test, right?",
      "votes": null
    },
    {
      "id": "2059543",
      "postDate": "12/09/2022 02:09:31",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/aesoptacit\" target=\"_blank\">@aesoptacit</a>! I don't think it has much meaning, maybe one can imagine it is ordering the sessions by their idxs or something like that 🙂 </p>\n<p>Yes, you are right. The sessions (for the purposes of this competition, users) are disjoint between train and test.</p>\n<p>Glad that you found this useful! 🙌</p>",
      "rawMarkdown": "Hey @aesoptacit! I don't think it has much meaning, maybe one can imagine it is ordering the sessions by their idxs or something like that 🙂 \n\nYes, you are right. The sessions (for the purposes of this competition, users) are disjoint between train and test.\n\nGlad that you found this useful! 🙌",
      "votes": null
    },
    {
      "id": "2059548",
      "postDate": "12/09/2022 02:13:29",
      "content": "<p>I understood.<br>\nThank you for your quick response!!</p>",
      "rawMarkdown": "I understood.\nThank you for your quick response!!",
      "votes": null
    },
    {
      "id": "2059595",
      "postDate": "12/09/2022 04:16:32",
      "content": "<p>np at all, glad I could be of help! 🙂</p>",
      "rawMarkdown": "np at all, glad I could be of help! 🙂",
      "votes": null
    },
    {
      "id": "2061480",
      "postDate": "12/11/2022 07:02:16",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a>, can we assume users (or sessions) that we are supposed to predict are brand new users, i.e they have never had any events before? Or it can be a new session of an \"old\" user? Thank you.</p>",
      "rawMarkdown": "Hi @pnormann, can we assume users (or sessions) that we are supposed to predict are brand new users, i.e they have never had any events before? Or it can be a new session of an \"old\" user? Thank you.",
      "votes": null
    },
    {
      "id": "2062794",
      "postDate": "12/12/2022 12:01:10",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/quangphm\" target=\"_blank\">@quangphm</a>, yes, all test users are new in the sense that they have not interacted with the shop during the training period. They can, however, have had events before the data extraction period started and be returning customers with items already in their carts and wishlists.</p>",
      "rawMarkdown": "Hi @quangphm, yes, all test users are new in the sense that they have not interacted with the shop during the training period. They can, however, have had events before the data extraction period started and be returning customers with items already in their carts and wishlists.",
      "votes": null
    },
    {
      "id": "2062924",
      "postDate": "12/12/2022 14:20:34",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> for clarifying.</p>",
      "rawMarkdown": "Thanks @pnormann for clarifying.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2016269,
      "author_name": "pnormann",
      "author_url": "",
      "post_date": "11/03/2022 22:07:01",
      "content": "<p>To clarify: as I mentioned in my other <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363874#2016257\" target=\"_blank\">comment</a>, the ground truth also comes from the test week period and events that users created after the end of the week are not included in the evaluation. Please refer to the <a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py\" target=\"_blank\">testset.py</a> script in our public GitHub repository for more details.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2016326,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/03/2022 23:21:21",
          "content": "<p>Thank you very much <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> for this important distinction! Appreciate the clarification! 🙏</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2016398,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/04/2022 01:17:59",
          "content": "<p><a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> apologies, if I may please ask you. If there two the same AIDs in ground truth as 'cart', how are the hits calculated?</p>\n<p>If I predict two the same AIDs that happen two be in ground truth, do they count as two hits? Or would I get the same result if I predict just a single AID no matter how many times they appear in ground truth?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2033364,
          "author_name": "huangzchao",
          "author_url": "",
          "post_date": "11/17/2022 07:47:38",
          "content": "<p>i think duplicate aid in carts will not affect the result, as the evaluation code use \"set\"<br>\n<code>if 'carts' in labels and labels['carts']:\n        cart_hits = len(set(prediction['carts'][:k]).intersection(labels['carts']))\n    else:\n        cart_hits = None</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2033691,
          "author_name": "danielliao",
          "author_url": "",
          "post_date": "11/17/2022 13:02:12",
          "content": "<p>Thanks for sharing the script and here is the <a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/evaluate.py#L44\" target=\"_blank\">link to code</a>. </p>\n<p>as for <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>'s question, according to the src line above, when the ground truth has two duplicated aids, and Radek's prediction also has the aid appeared twice, that aid only count as one hit. </p>\n<pre><code>([, , ]).intersection([, , ]) \n\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2061480,
          "author_name": "quangphm",
          "author_url": "",
          "post_date": "12/11/2022 07:02:16",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a>, can we assume users (or sessions) that we are supposed to predict are brand new users, i.e they have never had any events before? Or it can be a new session of an \"old\" user? Thank you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2062794,
          "author_name": "pnormann",
          "author_url": "",
          "post_date": "12/12/2022 12:01:10",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/quangphm\" target=\"_blank\">@quangphm</a>, yes, all test users are new in the sense that they have not interacted with the shop during the training period. They can, however, have had events before the data extraction period started and be returning customers with items already in their carts and wishlists.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2062924,
          "author_name": "quangphm",
          "author_url": "",
          "post_date": "12/12/2022 14:20:34",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> for clarifying.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2059540,
      "author_name": "aesoptacit",
      "author_url": "",
      "post_date": "12/09/2022 02:03:18",
      "content": "<p>Thanks for the info.<br>\nI’m wondering what the y-axis means. The users are different in train and test, right?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2059543,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "12/09/2022 02:09:31",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/aesoptacit\" target=\"_blank\">@aesoptacit</a>! I don't think it has much meaning, maybe one can imagine it is ordering the sessions by their idxs or something like that 🙂 </p>\n<p>Yes, you are right. The sessions (for the purposes of this competition, users) are disjoint between train and test.</p>\n<p>Glad that you found this useful! 🙌</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2059548,
          "author_name": "aesoptacit",
          "author_url": "",
          "post_date": "12/09/2022 02:13:29",
          "content": "<p>I understood.<br>\nThank you for your quick response!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2059595,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "12/09/2022 04:16:32",
          "content": "<p>np at all, glad I could be of help! 🙂</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2016239": "There is an associated repository with a good README and some related code. Here is [the repo](https://github.com/otto-de/recsys-dataset).\n\nThere is important information regarding the test set posted there that I have not come across here on Kaggle! Please take a look at the figure below\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F05ce69f1c8516ac1cbd1bc4f2730fd02%2Fsessions.png?generation=1667510035134467&alt=media)\n\nThe important piece of information is this -- the train session data that overlap with the test period was discarded! That makes sense, that is what we would expect.\n\nBut the key piece of information is this -- all the sessions in the test data don't have truncated beginnings! Essentially, they are complete sessions, they are not residuals from sessions that started in the train session period.\n\nThis is super important IMHO as it should impact how we think about the problem and might impact how we work with the data! Essentially, the train sessions might be truncated from the right, but may not be. The test sessions are guaranteed to not be truncated from the left (which is not obvious from the information on Kaggle) and are guaranteed to be truncated from the right (that is where the \"labels\" or should I rather say ground truth for the test data comes from!)\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
    "2016269": "To clarify: as I mentioned in my other [comment](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363874#2016257), the ground truth also comes from the test week period and events that users created after the end of the week are not included in the evaluation. Please refer to the [testset.py](https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py) script in our public GitHub repository for more details.",
    "2016326": "Thank you very much @pnormann for this important distinction! Appreciate the clarification! 🙏",
    "2016398": "pnormann apologies, if I may please ask you. If there two the same AIDs in ground truth as 'cart', how are the hits calculated?\n\nIf I predict two the same AIDs that happen two be in ground truth, do they count as two hits? Or would I get the same result if I predict just a single AID no matter how many times they appear in ground truth?",
    "2033364": "i think duplicate aid in carts will not affect the result, as the evaluation code use \"set\"\n`    if 'carts' in labels and labels['carts']:\n        cart_hits = len(set(prediction['carts'][:k]).intersection(labels['carts']))\n    else:\n        cart_hits = None`",
    "2033691": "Thanks for sharing the script and here is the [link to code](https://github.com/otto-de/recsys-dataset/blob/main/src/evaluate.py#L44). \n\nas for @radek1's question, according to the src line above, when the ground truth has two duplicated aids, and Radek's prediction also has the aid appeared twice, that aid only count as one hit. \n\n```python\nset(['a', 'b', 'b']).intersection(['a', 'b', 'b']) \n# return {'a', 'b'}\n```",
    "2059540": "Thanks for the info.\nI’m wondering what the y-axis means. The users are different in train and test, right?",
    "2059543": "Hey @aesoptacit! I don't think it has much meaning, maybe one can imagine it is ordering the sessions by their idxs or something like that 🙂 \n\nYes, you are right. The sessions (for the purposes of this competition, users) are disjoint between train and test.\n\nGlad that you found this useful! 🙌",
    "2059548": "I understood.\nThank you for your quick response!!",
    "2059595": "np at all, glad I could be of help! 🙂",
    "2061480": "Hi @pnormann, can we assume users (or sessions) that we are supposed to predict are brand new users, i.e they have never had any events before? Or it can be a new session of an \"old\" user? Thank you.",
    "2062794": "Hi @quangphm, yes, all test users are new in the sense that they have not interacted with the shop during the training period. They can, however, have had events before the data extraction period started and be returning customers with items already in their carts and wishlists.",
    "2062924": "Thanks @pnormann for clarifying."
  },
  "source": "meta"
}