{
  "id": 363874,
  "title": "A top-down perspective on the current metric values",
  "url": "/competitions/otto-recommender-system/discussion/363874",
  "author_name": "",
  "post_date": "2022-11-03T13:44:14.957709500Z",
  "votes": 46,
  "comment_count": 24,
  "views": 0,
  "content": "<p>Let me lay out some assumptions:</p>\n<ul>\n<li>Sessions in the test set are truncated in a random moment of the session life. </li>\n<li>For sessions truncated at the end, the problem is easy, because all the products that will be eventually added to cart &amp; purchased, were already clicked. Such sessions get a recall score of 1.0 quite easily with a model based only on the current session data.</li>\n<li>For sessions truncated at the start, the problem VERY HARD, because we need to pick producs from thousands of available products, without prior user history (there is no user id!). For such sessions, a recall score of 0.0 is quite hard to beat (<a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/leaderboard\" target=\"_blank\">see the winning scores for H&amp;M competition - they are close to 0</a>). This is different metric, but shows how difficult such a problem is.</li>\n</ul>\n<p>My intuitions:</p>\n<ul>\n<li>If sessions are truncated randomly, then the distribution of the test set can be thought of as a balanced mixture of the above extreme cases. Hence, the simple baselines based only on session data should get around 0.5 score, which we can see on the leaderboard</li>\n<li>It is easy to build a model based on current session data, and everyone will do that. However, the winners will be decided by those who can actually predict the hard problem: sessions truncated at the start.</li>\n</ul>\n<p>[EDITS]</p>\n<ul>\n<li>as confirmed below, a session covers all events of a particular user. Hence, introducing the custom sub-session logic might be helpful in the modeling. The hypothesis here  is that users are much more likely to purchase products withing the current \"logical\" session, compared to historical sessions.</li>\n</ul>",
  "messages": [
    {
      "id": "2015712",
      "postDate": "11/03/2022 13:44:14",
      "content": "<p>Let me lay out some assumptions:</p>\n<ul>\n<li>Sessions in the test set are truncated in a random moment of the session life. </li>\n<li>For sessions truncated at the end, the problem is easy, because all the products that will be eventually added to cart &amp; purchased, were already clicked. Such sessions get a recall score of 1.0 quite easily with a model based only on the current session data.</li>\n<li>For sessions truncated at the start, the problem VERY HARD, because we need to pick producs from thousands of available products, without prior user history (there is no user id!). For such sessions, a recall score of 0.0 is quite hard to beat (<a href=\"https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/leaderboard\" target=\"_blank\">see the winning scores for H&amp;M competition - they are close to 0</a>). This is different metric, but shows how difficult such a problem is.</li>\n</ul>\n<p>My intuitions:</p>\n<ul>\n<li>If sessions are truncated randomly, then the distribution of the test set can be thought of as a balanced mixture of the above extreme cases. Hence, the simple baselines based only on session data should get around 0.5 score, which we can see on the leaderboard</li>\n<li>It is easy to build a model based on current session data, and everyone will do that. However, the winners will be decided by those who can actually predict the hard problem: sessions truncated at the start.</li>\n</ul>\n<p>[EDITS]</p>\n<ul>\n<li>as confirmed below, a session covers all events of a particular user. Hence, introducing the custom sub-session logic might be helpful in the modeling. The hypothesis here  is that users are much more likely to purchase products withing the current \"logical\" session, compared to historical sessions.</li>\n</ul>",
      "rawMarkdown": "Let me lay out some assumptions:\n- Sessions in the test set are truncated in a random moment of the session life. \n- For sessions truncated at the end, the problem is easy, because all the products that will be eventually added to cart & purchased, were already clicked. Such sessions get a recall score of 1.0 quite easily with a model based only on the current session data.\n- For sessions truncated at the start, the problem VERY HARD, because we need to pick producs from thousands of available products, without prior user history (there is no user id!). For such sessions, a recall score of 0.0 is quite hard to beat ([see the winning scores for H&M competition - they are close to 0](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/leaderboard)). This is different metric, but shows how difficult such a problem is.\n\nMy intuitions:\n- If sessions are truncated randomly, then the distribution of the test set can be thought of as a balanced mixture of the above extreme cases. Hence, the simple baselines based only on session data should get around 0.5 score, which we can see on the leaderboard\n- It is easy to build a model based on current session data, and everyone will do that. However, the winners will be decided by those who can actually predict the hard problem: sessions truncated at the start.\n\n\n[EDITS]\n- as confirmed below, a session covers all events of a particular user. Hence, introducing the custom sub-session logic might be helpful in the modeling. The hypothesis here  is that users are much more likely to purchase products withing the current \"logical\" session, compared to historical sessions.",
      "votes": null
    },
    {
      "id": "2015741",
      "postDate": "11/03/2022 14:08:32",
      "content": "<p>Are you sure that test sessions are truncated randomly? I didn't see that information. If that's the case, you are right. It's really hard to predict those sessions, besides predicting the most clicked/carted/ordered product might be the best choice for them.</p>\n<p>About H&amp;M competition, I don't know much about the business side of this application but recall@x makes more sense than map@x in my opinion.</p>",
      "rawMarkdown": "Are you sure that test sessions are truncated randomly? I didn't see that information. If that's the case, you are right. It's really hard to predict those sessions, besides predicting the most clicked/carted/ordered product might be the best choice for them.\n\nAbout H&M competition, I don't know much about the business side of this application but recall@x makes more sense than map@x in my opinion.",
      "votes": null
    },
    {
      "id": "2015758",
      "postDate": "11/03/2022 14:23:21",
      "content": "<p>I am not sure. It is my guess because:</p>\n<ul>\n<li>it makes sense</li>\n<li>scores align with the assumption :)</li>\n</ul>",
      "rawMarkdown": "I am not sure. It is my guess because:\n- it makes sense\n- scores align with the assumption :)",
      "votes": null
    },
    {
      "id": "2015773",
      "postDate": "11/03/2022 14:36:15",
      "content": "<p>Yep, it makes sense. Is it even possible to predict those sessions accurately? This reminds me what Spotify does when I create a new playlist. It doesn't recommend anything when the playlist is empty. When I add a song, it recommends similar songs from the same artist. As the playlist becomes larger, it recommends different songs from different artists. Spotify probably uses a representation to query but we don't have much here. I think we can assume that people would click/cart/order similar things and learn a representation from session sequences somehow…</p>",
      "rawMarkdown": "Yep, it makes sense. Is it even possible to predict those sessions accurately? This reminds me what Spotify does when I create a new playlist. It doesn't recommend anything when the playlist is empty. When I add a song, it recommends similar songs from the same artist. As the playlist becomes larger, it recommends different songs from different artists. Spotify probably uses a representation to query but we don't have much here. I think we can assume that people would click/cart/order similar things and learn a representation from session sequences somehow...",
      "votes": null
    },
    {
      "id": "2015813",
      "postDate": "11/03/2022 15:02:57",
      "content": "<p>Yep - we could use standard worv2vec approach, where we treat session sequence like a sentence. AirBnB published a nice article about it <a href=\"https://medium.com/airbnb-engineering/listing-embeddings-for-similar-listing-recommendations-and-real-time-personalization-in-search-601172f7603e\" target=\"_blank\">https://medium.com/airbnb-engineering/listing-embeddings-for-similar-listing-recommendations-and-real-time-personalization-in-search-601172f7603e</a></p>",
      "rawMarkdown": "Yep - we could use standard worv2vec approach, where we treat session sequence like a sentence. AirBnB published a nice article about it https://medium.com/airbnb-engineering/listing-embeddings-for-similar-listing-recommendations-and-real-time-personalization-in-search-601172f7603e",
      "votes": null
    },
    {
      "id": "2015821",
      "postDate": "11/03/2022 15:05:44",
      "content": "<p>The sessions are truncated randomly in that they start at an arbitrary point in time and end at an arbitrary point in time 🙂</p>\n<p>Since the sessions are a recording of \"user activity\" which can begin and end at any day, the truncation is random. It is a window into a view of user activity. I got this from doing my <a href=\"https://www.kaggle.com/code/radek1/eda-an-overview-of-the-full-dataset\" target=\"_blank\">EDA</a>.</p>\n<p>Now the interesting bit here also is that <code>test</code> sessions are much shorter than <code>train</code> session! (because the duration over which the <code>test</code> sessions were generated is much shorter!). I guess that is an interesting aspect of all of this to ponder as well 🙂</p>",
      "rawMarkdown": "The sessions are truncated randomly in that they start at an arbitrary point in time and end at an arbitrary point in time 🙂\n\nSince the sessions are a recording of \"user activity\" which can begin and end at any day, the truncation is random. It is a window into a view of user activity. I got this from doing my [EDA](https://www.kaggle.com/code/radek1/eda-an-overview-of-the-full-dataset).\n\nNow the interesting bit here also is that `test` sessions are much shorter than `train` session! (because the duration over which the `test` sessions were generated is much shorter!). I guess that is an interesting aspect of all of this to ponder as well 🙂",
      "votes": null
    },
    {
      "id": "2015825",
      "postDate": "11/03/2022 15:07:11",
      "content": "<p>This is a great overview of the situation, <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a>! Thx for posting! Definitely helped me organize my thinking on this problem thx to your comment! 🙏</p>",
      "rawMarkdown": "This is a great overview of the situation, @narsil! Thx for posting! Definitely helped me organize my thinking on this problem thx to your comment! 🙏",
      "votes": null
    },
    {
      "id": "2015897",
      "postDate": "11/03/2022 15:59:53",
      "content": "<p>You are right. We dont actually need the user id  - we get all events for a user within a tracking period. It means we can create sub-sessions and use such logic in the modelling. </p>",
      "rawMarkdown": "You are right. We dont actually need the user id  - we get all events for a user within a tracking period. It means we can create sub-sessions and use such logic in the modelling.",
      "votes": null
    },
    {
      "id": "2015898",
      "postDate": "11/03/2022 16:00:42",
      "content": "<p>Thank you. It seems to be a very interesting competition </p>",
      "rawMarkdown": "Thank you. It seems to be a very interesting competition",
      "votes": null
    },
    {
      "id": "2016226",
      "postDate": "11/03/2022 21:00:35",
      "content": "<p>Not exactly sure if I am looking into the correct place but if so, the sessions are truncated randomly. <br>\nthe code is here: <a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py#L25\" target=\"_blank\">https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py#L25</a></p>",
      "rawMarkdown": "Not exactly sure if I am looking into the correct place but if so, the sessions are truncated randomly. \nthe code is here: https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py#L25",
      "votes": null
    },
    {
      "id": "2016242",
      "postDate": "11/03/2022 21:24:54",
      "content": "<p>Nice find <a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a>! 🙌 Especially assuming <a href=\"https://github.com/otto-de/recsys-dataset/blob/5d9c5f6b0a82cb09f4b98542c77a774facaf51c7/src/testset.py#L34\" target=\"_blank\">the code below</a> was used to create the test set for Kaggle.</p>\n<p>This would indeed suggest that the sessions in test were randomly split.</p>\n<p>But the timestamps for test never exceed the 1 week period. 🤔 I wonder if the \"ground truth data\" also only comes from the 1 week period? I guess that would make sense.</p>\n<p>The test set would then be constructed as such:</p>\n<ul>\n<li>take session data from a one-week period (as opposed to session data that starts in the one week period but might run beyond the end date of the test week)</li>\n<li>randomly truncate it</li>\n</ul>",
      "rawMarkdown": "Nice find @snnclsr! 🙌 Especially assuming [the code below](https://github.com/otto-de/recsys-dataset/blob/5d9c5f6b0a82cb09f4b98542c77a774facaf51c7/src/testset.py#L34) was used to create the test set for Kaggle.\n\nThis would indeed suggest that the sessions in test were randomly split.\n\nBut the timestamps for test never exceed the 1 week period. 🤔 I wonder if the \"ground truth data\" also only comes from the 1 week period? I guess that would make sense.\n\nThe test set would then be constructed as such:\n- take session data from a one-week period (as opposed to session data that starts in the one week period but might run beyond the end date of the test week)\n- randomly truncate it",
      "votes": null
    },
    {
      "id": "2016243",
      "postDate": "11/03/2022 21:27:02",
      "content": "<p><a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> just as an FYI today I came across <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363965\" target=\"_blank\">this piece of information</a>, I think it might be useful to the reasoning you outline in your OP 🙂</p>\n<p>Essentially, train session data that would overlap with the test period has been thrown away, so all the sessions in test are never truncated from the left! They are always brand new sessions, though might be very short (which is the problem that you mention -- a session truncated at the start).</p>",
      "rawMarkdown": "narsil just as an FYI today I came across [this piece of information](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363965), I think it might be useful to the reasoning you outline in your OP 🙂\n\nEssentially, train session data that would overlap with the test period has been thrown away, so all the sessions in test are never truncated from the left! They are always brand new sessions, though might be very short (which is the problem that you mention -- a session truncated at the start).",
      "votes": null
    },
    {
      "id": "2016253",
      "postDate": "11/03/2022 21:41:14",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a>, thanks for sharing your thoughts! We can confirm your assumptions regarding the nature of the data, and as <a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a> mentioned, you can find the procedure of how we created the test set for the challenge in our GitHub repo <a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "Hey @narsil, thanks for sharing your thoughts! We can confirm your assumptions regarding the nature of the data, and as @snnclsr mentioned, you can find the procedure of how we created the test set for the challenge in our GitHub repo [here](https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py).",
      "votes": null
    },
    {
      "id": "2016257",
      "postDate": "11/03/2022 21:51:14",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, the truncated test sessions, including the ground truth, which you are supposed to predict, all come from the same week after the train set and never exceed the week's end date.</p>",
      "rawMarkdown": "Hey @radek1, the truncated test sessions, including the ground truth, which you are supposed to predict, all come from the same week after the train set and never exceed the week's end date.",
      "votes": null
    },
    {
      "id": "2016327",
      "postDate": "11/03/2022 23:23:08",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> for your answer!</p>",
      "rawMarkdown": "Thank you @pnormann for your answer!",
      "votes": null
    },
    {
      "id": "2016499",
      "postDate": "11/04/2022 04:39:51",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> for confirmation</p>",
      "rawMarkdown": "Thanks @pnormann for confirmation",
      "votes": null
    },
    {
      "id": "2016504",
      "postDate": "11/04/2022 04:42:27",
      "content": "<p>Thanks for confirmation. No truncation from the left is a logical choice and relflecting real life.  During the inference on production, we also have all user history up until that point.</p>",
      "rawMarkdown": "Thanks for confirmation. No truncation from the left is a logical choice and relflecting real life.  During the inference on production, we also have all user history up until that point.",
      "votes": null
    },
    {
      "id": "2017479",
      "postDate": "11/04/2022 20:45:10",
      "content": "<p>Hi narsil, <br>\nthanks for sharing your insights, they are great! I do have a question regarding the hypothesis at the end. Could you elaborate what you mean by \"users are much more likely to purchase products withing the current \"logical\" session, compared to historical sessions.\"?</p>",
      "rawMarkdown": "Hi narsil, \nthanks for sharing your insights, they are great! I do have a question regarding the hypothesis at the end. Could you elaborate what you mean by \"users are much more likely to purchase products withing the current \"logical\" session, compared to historical sessions.\"?",
      "votes": null
    },
    {
      "id": "2018387",
      "postDate": "11/05/2022 16:54:33",
      "content": "<p>When the uses closes his browser and goes on to some other activity, her attention is reset, and she \"forgets\" about products she was considering. Abandoning the session without adding to cart or purchasing is a negative signal</p>",
      "rawMarkdown": "When the uses closes his browser and goes on to some other activity, her attention is reset, and she \"forgets\" about products she was considering. Abandoning the session without adding to cart or purchasing is a negative signal",
      "votes": null
    },
    {
      "id": "2018395",
      "postDate": "11/05/2022 17:01:02",
      "content": "<p>Oh okay, I get it now thanks! But isnt that very speculative? This doesnt need to be the case for every person. I know someone (me lol) who clicks on a pair of shoes multiple times before he really buys them or considers putting them in the cart. Not everyone is \"impulsive\" (might be a to negative word) to put in the cart at first glance. :)</p>",
      "rawMarkdown": "Oh okay, I get it now thanks! But isnt that very speculative? This doesnt need to be the case for every person. I know someone (me lol) who clicks on a pair of shoes multiple times before he really buys them or considers putting them in the cart. Not everyone is \"impulsive\" (might be a to negative word) to put in the cart at first glance. :)",
      "votes": null
    },
    {
      "id": "2020005",
      "postDate": "11/07/2022 06:04:50",
      "content": "<p>It might be untrue for many users and still provide a strong signal to the model</p>",
      "rawMarkdown": "It might be untrue for many users and still provide a strong signal to the model",
      "votes": null
    },
    {
      "id": "2033197",
      "postDate": "11/17/2022 04:51:23",
      "content": "<p>Addition problem, for training set, when you create label, though random split, you will most probaly get sequence feature longer than 1 week. While in test set, the longest sequence feature, you just can get 1 week, which maybe worse for test prediction…</p>",
      "rawMarkdown": "Addition problem, for training set, when you create label, though random split, you will most probaly get sequence feature longer than 1 week. While in test set, the longest sequence feature, you just can get 1 week, which maybe worse for test prediction...",
      "votes": null
    },
    {
      "id": "2033243",
      "postDate": "11/17/2022 05:47:40",
      "content": "<p>Good analysis. Train data is 4 weeks and test data is 1 week. The <code>test.csv</code> are random uniform from length of user events in 1 week. So most test users have very little data compared with train data. Most test predictions will be \"cold start\" problem.</p>\n<p>I agree that the winning solutions will solve the \"cold start\" problem (i.e. predict test users with very little history data). </p>",
      "rawMarkdown": "Good analysis. Train data is 4 weeks and test data is 1 week. The `test.csv` are random uniform from length of user events in 1 week. So most test users have very little data compared with train data. Most test predictions will be \"cold start\" problem.\n\nI agree that the winning solutions will solve the \"cold start\" problem (i.e. predict test users with very little history data).",
      "votes": null
    },
    {
      "id": "2033300",
      "postDate": "11/17/2022 06:28:04",
      "content": "<p>Indeed. Feature definition mismatch between training and inference is one of the worst things to happen. You can mitigate it by directly controlling the timespan of data for feature calculation.</p>",
      "rawMarkdown": "Indeed. Feature definition mismatch between training and inference is one of the worst things to happen. You can mitigate it by directly controlling the timespan of data for feature calculation.",
      "votes": null
    },
    {
      "id": "2051043",
      "postDate": "12/01/2022 07:03:27",
      "content": "<p>Thanks Chris. Another interesting data property is that for every user we have at least the first product xe clicked. I ran some tests and it seems this is already a powerful context. Conceptally it resembles a bit the acquistion channel of the user (email, organic, etc), which normally is a strong feature in these types of problems.</p>",
      "rawMarkdown": "Thanks Chris. Another interesting data property is that for every user we have at least the first product xe clicked. I ran some tests and it seems this is already a powerful context. Conceptally it resembles a bit the acquistion channel of the user (email, organic, etc), which normally is a strong feature in these types of problems.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2015741,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "11/03/2022 14:08:32",
      "content": "<p>Are you sure that test sessions are truncated randomly? I didn't see that information. If that's the case, you are right. It's really hard to predict those sessions, besides predicting the most clicked/carted/ordered product might be the best choice for them.</p>\n<p>About H&amp;M competition, I don't know much about the business side of this application but recall@x makes more sense than map@x in my opinion.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2015758,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "11/03/2022 14:23:21",
          "content": "<p>I am not sure. It is my guess because:</p>\n<ul>\n<li>it makes sense</li>\n<li>scores align with the assumption :)</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2015773,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "11/03/2022 14:36:15",
          "content": "<p>Yep, it makes sense. Is it even possible to predict those sessions accurately? This reminds me what Spotify does when I create a new playlist. It doesn't recommend anything when the playlist is empty. When I add a song, it recommends similar songs from the same artist. As the playlist becomes larger, it recommends different songs from different artists. Spotify probably uses a representation to query but we don't have much here. I think we can assume that people would click/cart/order similar things and learn a representation from session sequences somehow…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2015813,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "11/03/2022 15:02:57",
          "content": "<p>Yep - we could use standard worv2vec approach, where we treat session sequence like a sentence. AirBnB published a nice article about it <a href=\"https://medium.com/airbnb-engineering/listing-embeddings-for-similar-listing-recommendations-and-real-time-personalization-in-search-601172f7603e\" target=\"_blank\">https://medium.com/airbnb-engineering/listing-embeddings-for-similar-listing-recommendations-and-real-time-personalization-in-search-601172f7603e</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2015821,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/03/2022 15:05:44",
          "content": "<p>The sessions are truncated randomly in that they start at an arbitrary point in time and end at an arbitrary point in time 🙂</p>\n<p>Since the sessions are a recording of \"user activity\" which can begin and end at any day, the truncation is random. It is a window into a view of user activity. I got this from doing my <a href=\"https://www.kaggle.com/code/radek1/eda-an-overview-of-the-full-dataset\" target=\"_blank\">EDA</a>.</p>\n<p>Now the interesting bit here also is that <code>test</code> sessions are much shorter than <code>train</code> session! (because the duration over which the <code>test</code> sessions were generated is much shorter!). I guess that is an interesting aspect of all of this to ponder as well 🙂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2015897,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "11/03/2022 15:59:53",
          "content": "<p>You are right. We dont actually need the user id  - we get all events for a user within a tracking period. It means we can create sub-sessions and use such logic in the modelling. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2016226,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "11/03/2022 21:00:35",
          "content": "<p>Not exactly sure if I am looking into the correct place but if so, the sessions are truncated randomly. <br>\nthe code is here: <a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py#L25\" target=\"_blank\">https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py#L25</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2016242,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/03/2022 21:24:54",
          "content": "<p>Nice find <a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a>! 🙌 Especially assuming <a href=\"https://github.com/otto-de/recsys-dataset/blob/5d9c5f6b0a82cb09f4b98542c77a774facaf51c7/src/testset.py#L34\" target=\"_blank\">the code below</a> was used to create the test set for Kaggle.</p>\n<p>This would indeed suggest that the sessions in test were randomly split.</p>\n<p>But the timestamps for test never exceed the 1 week period. 🤔 I wonder if the \"ground truth data\" also only comes from the 1 week period? I guess that would make sense.</p>\n<p>The test set would then be constructed as such:</p>\n<ul>\n<li>take session data from a one-week period (as opposed to session data that starts in the one week period but might run beyond the end date of the test week)</li>\n<li>randomly truncate it</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2016257,
          "author_name": "pnormann",
          "author_url": "",
          "post_date": "11/03/2022 21:51:14",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, the truncated test sessions, including the ground truth, which you are supposed to predict, all come from the same week after the train set and never exceed the week's end date.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2016327,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/03/2022 23:23:08",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> for your answer!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2015825,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "11/03/2022 15:07:11",
      "content": "<p>This is a great overview of the situation, <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a>! Thx for posting! Definitely helped me organize my thinking on this problem thx to your comment! 🙏</p>",
      "votes": null,
      "replies": [
        {
          "id": 2015898,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "11/03/2022 16:00:42",
          "content": "<p>Thank you. It seems to be a very interesting competition </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2016243,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/03/2022 21:27:02",
          "content": "<p><a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> just as an FYI today I came across <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363965\" target=\"_blank\">this piece of information</a>, I think it might be useful to the reasoning you outline in your OP 🙂</p>\n<p>Essentially, train session data that would overlap with the test period has been thrown away, so all the sessions in test are never truncated from the left! They are always brand new sessions, though might be very short (which is the problem that you mention -- a session truncated at the start).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2016504,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "11/04/2022 04:42:27",
          "content": "<p>Thanks for confirmation. No truncation from the left is a logical choice and relflecting real life.  During the inference on production, we also have all user history up until that point.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2016253,
      "author_name": "pnormann",
      "author_url": "",
      "post_date": "11/03/2022 21:41:14",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a>, thanks for sharing your thoughts! We can confirm your assumptions regarding the nature of the data, and as <a href=\"https://www.kaggle.com/snnclsr\" target=\"_blank\">@snnclsr</a> mentioned, you can find the procedure of how we created the test set for the challenge in our GitHub repo <a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py\" target=\"_blank\">here</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2016499,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "11/04/2022 04:39:51",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> for confirmation</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2017479,
      "author_name": "ianfit",
      "author_url": "",
      "post_date": "11/04/2022 20:45:10",
      "content": "<p>Hi narsil, <br>\nthanks for sharing your insights, they are great! I do have a question regarding the hypothesis at the end. Could you elaborate what you mean by \"users are much more likely to purchase products withing the current \"logical\" session, compared to historical sessions.\"?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2018387,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "11/05/2022 16:54:33",
          "content": "<p>When the uses closes his browser and goes on to some other activity, her attention is reset, and she \"forgets\" about products she was considering. Abandoning the session without adding to cart or purchasing is a negative signal</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2018395,
          "author_name": "ianfit",
          "author_url": "",
          "post_date": "11/05/2022 17:01:02",
          "content": "<p>Oh okay, I get it now thanks! But isnt that very speculative? This doesnt need to be the case for every person. I know someone (me lol) who clicks on a pair of shoes multiple times before he really buys them or considers putting them in the cart. Not everyone is \"impulsive\" (might be a to negative word) to put in the cart at first glance. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2020005,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "11/07/2022 06:04:50",
          "content": "<p>It might be untrue for many users and still provide a strong signal to the model</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2033197,
      "author_name": "huangzchao",
      "author_url": "",
      "post_date": "11/17/2022 04:51:23",
      "content": "<p>Addition problem, for training set, when you create label, though random split, you will most probaly get sequence feature longer than 1 week. While in test set, the longest sequence feature, you just can get 1 week, which maybe worse for test prediction…</p>",
      "votes": null,
      "replies": [
        {
          "id": 2033300,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "11/17/2022 06:28:04",
          "content": "<p>Indeed. Feature definition mismatch between training and inference is one of the worst things to happen. You can mitigate it by directly controlling the timespan of data for feature calculation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2033243,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/17/2022 05:47:40",
      "content": "<p>Good analysis. Train data is 4 weeks and test data is 1 week. The <code>test.csv</code> are random uniform from length of user events in 1 week. So most test users have very little data compared with train data. Most test predictions will be \"cold start\" problem.</p>\n<p>I agree that the winning solutions will solve the \"cold start\" problem (i.e. predict test users with very little history data). </p>",
      "votes": null,
      "replies": [
        {
          "id": 2051043,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "12/01/2022 07:03:27",
          "content": "<p>Thanks Chris. Another interesting data property is that for every user we have at least the first product xe clicked. I ran some tests and it seems this is already a powerful context. Conceptally it resembles a bit the acquistion channel of the user (email, organic, etc), which normally is a strong feature in these types of problems.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2015712": "Let me lay out some assumptions:\n- Sessions in the test set are truncated in a random moment of the session life. \n- For sessions truncated at the end, the problem is easy, because all the products that will be eventually added to cart & purchased, were already clicked. Such sessions get a recall score of 1.0 quite easily with a model based only on the current session data.\n- For sessions truncated at the start, the problem VERY HARD, because we need to pick producs from thousands of available products, without prior user history (there is no user id!). For such sessions, a recall score of 0.0 is quite hard to beat ([see the winning scores for H&M competition - they are close to 0](https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/leaderboard)). This is different metric, but shows how difficult such a problem is.\n\nMy intuitions:\n- If sessions are truncated randomly, then the distribution of the test set can be thought of as a balanced mixture of the above extreme cases. Hence, the simple baselines based only on session data should get around 0.5 score, which we can see on the leaderboard\n- It is easy to build a model based on current session data, and everyone will do that. However, the winners will be decided by those who can actually predict the hard problem: sessions truncated at the start.\n\n\n[EDITS]\n- as confirmed below, a session covers all events of a particular user. Hence, introducing the custom sub-session logic might be helpful in the modeling. The hypothesis here  is that users are much more likely to purchase products withing the current \"logical\" session, compared to historical sessions.",
    "2015741": "Are you sure that test sessions are truncated randomly? I didn't see that information. If that's the case, you are right. It's really hard to predict those sessions, besides predicting the most clicked/carted/ordered product might be the best choice for them.\n\nAbout H&M competition, I don't know much about the business side of this application but recall@x makes more sense than map@x in my opinion.",
    "2015758": "I am not sure. It is my guess because:\n- it makes sense\n- scores align with the assumption :)",
    "2015773": "Yep, it makes sense. Is it even possible to predict those sessions accurately? This reminds me what Spotify does when I create a new playlist. It doesn't recommend anything when the playlist is empty. When I add a song, it recommends similar songs from the same artist. As the playlist becomes larger, it recommends different songs from different artists. Spotify probably uses a representation to query but we don't have much here. I think we can assume that people would click/cart/order similar things and learn a representation from session sequences somehow...",
    "2015813": "Yep - we could use standard worv2vec approach, where we treat session sequence like a sentence. AirBnB published a nice article about it https://medium.com/airbnb-engineering/listing-embeddings-for-similar-listing-recommendations-and-real-time-personalization-in-search-601172f7603e",
    "2015821": "The sessions are truncated randomly in that they start at an arbitrary point in time and end at an arbitrary point in time 🙂\n\nSince the sessions are a recording of \"user activity\" which can begin and end at any day, the truncation is random. It is a window into a view of user activity. I got this from doing my [EDA](https://www.kaggle.com/code/radek1/eda-an-overview-of-the-full-dataset).\n\nNow the interesting bit here also is that `test` sessions are much shorter than `train` session! (because the duration over which the `test` sessions were generated is much shorter!). I guess that is an interesting aspect of all of this to ponder as well 🙂",
    "2015825": "This is a great overview of the situation, @narsil! Thx for posting! Definitely helped me organize my thinking on this problem thx to your comment! 🙏",
    "2015897": "You are right. We dont actually need the user id  - we get all events for a user within a tracking period. It means we can create sub-sessions and use such logic in the modelling.",
    "2015898": "Thank you. It seems to be a very interesting competition",
    "2016226": "Not exactly sure if I am looking into the correct place but if so, the sessions are truncated randomly. \nthe code is here: https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py#L25",
    "2016242": "Nice find @snnclsr! 🙌 Especially assuming [the code below](https://github.com/otto-de/recsys-dataset/blob/5d9c5f6b0a82cb09f4b98542c77a774facaf51c7/src/testset.py#L34) was used to create the test set for Kaggle.\n\nThis would indeed suggest that the sessions in test were randomly split.\n\nBut the timestamps for test never exceed the 1 week period. 🤔 I wonder if the \"ground truth data\" also only comes from the 1 week period? I guess that would make sense.\n\nThe test set would then be constructed as such:\n- take session data from a one-week period (as opposed to session data that starts in the one week period but might run beyond the end date of the test week)\n- randomly truncate it",
    "2016243": "narsil just as an FYI today I came across [this piece of information](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363965), I think it might be useful to the reasoning you outline in your OP 🙂\n\nEssentially, train session data that would overlap with the test period has been thrown away, so all the sessions in test are never truncated from the left! They are always brand new sessions, though might be very short (which is the problem that you mention -- a session truncated at the start).",
    "2016253": "Hey @narsil, thanks for sharing your thoughts! We can confirm your assumptions regarding the nature of the data, and as @snnclsr mentioned, you can find the procedure of how we created the test set for the challenge in our GitHub repo [here](https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py).",
    "2016257": "Hey @radek1, the truncated test sessions, including the ground truth, which you are supposed to predict, all come from the same week after the train set and never exceed the week's end date.",
    "2016327": "Thank you @pnormann for your answer!",
    "2016499": "Thanks @pnormann for confirmation",
    "2016504": "Thanks for confirmation. No truncation from the left is a logical choice and relflecting real life.  During the inference on production, we also have all user history up until that point.",
    "2017479": "Hi narsil, \nthanks for sharing your insights, they are great! I do have a question regarding the hypothesis at the end. Could you elaborate what you mean by \"users are much more likely to purchase products withing the current \"logical\" session, compared to historical sessions.\"?",
    "2018387": "When the uses closes his browser and goes on to some other activity, her attention is reset, and she \"forgets\" about products she was considering. Abandoning the session without adding to cart or purchasing is a negative signal",
    "2018395": "Oh okay, I get it now thanks! But isnt that very speculative? This doesnt need to be the case for every person. I know someone (me lol) who clicks on a pair of shoes multiple times before he really buys them or considers putting them in the cart. Not everyone is \"impulsive\" (might be a to negative word) to put in the cart at first glance. :)",
    "2020005": "It might be untrue for many users and still provide a strong signal to the model",
    "2033197": "Addition problem, for training set, when you create label, though random split, you will most probaly get sequence feature longer than 1 week. While in test set, the longest sequence feature, you just can get 1 week, which maybe worse for test prediction...",
    "2033243": "Good analysis. Train data is 4 weeks and test data is 1 week. The `test.csv` are random uniform from length of user events in 1 week. So most test users have very little data compared with train data. Most test predictions will be \"cold start\" problem.\n\nI agree that the winning solutions will solve the \"cold start\" problem (i.e. predict test users with very little history data).",
    "2033300": "Indeed. Feature definition mismatch between training and inference is one of the worst things to happen. You can mitigate it by directly controlling the timespan of data for feature calculation.",
    "2051043": "Thanks Chris. Another interesting data property is that for every user we have at least the first product xe clicked. I ran some tests and it seems this is already a powerful context. Conceptally it resembles a bit the acquistion channel of the user (email, organic, etc), which normally is a strong feature in these types of problems."
  },
  "source": "meta"
}