{
  "id": 370384,
  "title": "Some concerns about validation",
  "url": "/competitions/otto-recommender-system/discussion/370384",
  "author_name": "",
  "post_date": "2022-12-04T08:23:46.249061400Z",
  "votes": 10,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I noticed almost everyone uses test set split code from organizers <a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py\" target=\"_blank\">repository</a>, but I think it might have some flaws.</p>\n<p>Using sessions from last week as a holdout test set makes sense however splitting sessions randomly could be problematic. For example session 11098537:<br>\n<img src=\"https://i.imgur.com/H7PKxub.png\" alt=\"sess1\"></p>\n<p>Ground-truth belongs to the highlighted event. If you are using aids from the entire session then you might have access to ground-truth click. This can be prevented by finding cutoff index and use aids before that index. It isn't efficient and it requires complex logic if there are duplicate ground-truth aids in sessions. It becomes even more complex when there are ground-truth carts and orders.</p>\n<p>Another thing I noticed is some of the ground-truth labels can be without clicks when they are split randomly. It doesn't necessarily have to be a problem but it feels more natural to have ground-truth clicks.</p>\n<p>My suggestions are</p>\n<ul>\n<li>Always keep one last click at the end and split sessions randomly before that click event so every session can have a ground-truth click</li>\n<li>Store cutoff index while splitting sessions so only events before that index can be utilized</li>\n<li>Filter unseen aids in last week sessions</li>\n</ul>",
  "messages": [
    {
      "id": "2054549",
      "postDate": "12/04/2022 08:23:46",
      "content": "<p>I noticed almost everyone uses test set split code from organizers <a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py\" target=\"_blank\">repository</a>, but I think it might have some flaws.</p>\n<p>Using sessions from last week as a holdout test set makes sense however splitting sessions randomly could be problematic. For example session 11098537:<br>\n<img src=\"https://i.imgur.com/H7PKxub.png\" alt=\"sess1\"></p>\n<p>Ground-truth belongs to the highlighted event. If you are using aids from the entire session then you might have access to ground-truth click. This can be prevented by finding cutoff index and use aids before that index. It isn't efficient and it requires complex logic if there are duplicate ground-truth aids in sessions. It becomes even more complex when there are ground-truth carts and orders.</p>\n<p>Another thing I noticed is some of the ground-truth labels can be without clicks when they are split randomly. It doesn't necessarily have to be a problem but it feels more natural to have ground-truth clicks.</p>\n<p>My suggestions are</p>\n<ul>\n<li>Always keep one last click at the end and split sessions randomly before that click event so every session can have a ground-truth click</li>\n<li>Store cutoff index while splitting sessions so only events before that index can be utilized</li>\n<li>Filter unseen aids in last week sessions</li>\n</ul>",
      "rawMarkdown": "I noticed almost everyone uses test set split code from organizers [repository](https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py), but I think it might have some flaws.\n\nUsing sessions from last week as a holdout test set makes sense however splitting sessions randomly could be problematic. For example session 11098537:\n![sess1](https://i.imgur.com/H7PKxub.png)\n\nGround-truth belongs to the highlighted event. If you are using aids from the entire session then you might have access to ground-truth click. This can be prevented by finding cutoff index and use aids before that index. It isn't efficient and it requires complex logic if there are duplicate ground-truth aids in sessions. It becomes even more complex when there are ground-truth carts and orders.\n\nAnother thing I noticed is some of the ground-truth labels can be without clicks when they are split randomly. It doesn't necessarily have to be a problem but it feels more natural to have ground-truth clicks.\n\nMy suggestions are\n* Always keep one last click at the end and split sessions randomly before that click event so every session can have a ground-truth click\n* Store cutoff index while splitting sessions so only events before that index can be utilized\n* Filter unseen aids in last week sessions",
      "votes": null
    },
    {
      "id": "2054809",
      "postDate": "12/04/2022 13:25:29",
      "content": "<p>i don't think there are flaws in organizers repository, you can always add rules ( your filters ) to apply to the splinted data set. </p>",
      "rawMarkdown": "i don't think there are flaws in organizers repository, you can always add rules ( your filters ) to apply to the splinted data set.",
      "votes": null
    },
    {
      "id": "2055721",
      "postDate": "12/05/2022 11:05:03",
      "content": "<p>Another useful way to think about this is from the perspective of the company behind the online store. You want to be able to serve as useful recommendations as you can, even on the first event!</p>\n<p>In fact, there is probably a lot of business value in being able to do something valuable for our prospective customer on the first, second and third event, to make sure they stay for a while longer 🙂 Offering them something of value could go a long way for improving retention and improving the bottom line!</p>\n<p>In this sense, random splits probably express the needs of OTTO quite well. Additionally, <code>being creative</code> about splits is an awesome way to shoot yourself in the foot in some way you can't foresee but the tens of Kaggle Grandmasters who are likely to part take in this competition would be likely able to identify 🙂</p>\n<p>Interesting discussion, thanks for starting it! 🙌 Not necessarily arguing against your point, just wanted to share a slightly different perspective on why random splitting might have some advantages 🙂 In general, I find randomness can often be a great ally in our ML pursuits!</p>",
      "rawMarkdown": "Another useful way to think about this is from the perspective of the company behind the online store. You want to be able to serve as useful recommendations as you can, even on the first event!\n\nIn fact, there is probably a lot of business value in being able to do something valuable for our prospective customer on the first, second and third event, to make sure they stay for a while longer 🙂 Offering them something of value could go a long way for improving retention and improving the bottom line!\n\nIn this sense, random splits probably express the needs of OTTO quite well. Additionally, `being creative` about splits is an awesome way to shoot yourself in the foot in some way you can't foresee but the tens of Kaggle Grandmasters who are likely to part take in this competition would be likely able to identify 🙂\n\nInteresting discussion, thanks for starting it! 🙌 Not necessarily arguing against your point, just wanted to share a slightly different perspective on why random splitting might have some advantages 🙂 In general, I find randomness can often be a great ally in our ML pursuits!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2054809,
      "author_name": "tyeestudio",
      "author_url": "",
      "post_date": "12/04/2022 13:25:29",
      "content": "<p>i don't think there are flaws in organizers repository, you can always add rules ( your filters ) to apply to the splinted data set. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2055721,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "12/05/2022 11:05:03",
      "content": "<p>Another useful way to think about this is from the perspective of the company behind the online store. You want to be able to serve as useful recommendations as you can, even on the first event!</p>\n<p>In fact, there is probably a lot of business value in being able to do something valuable for our prospective customer on the first, second and third event, to make sure they stay for a while longer 🙂 Offering them something of value could go a long way for improving retention and improving the bottom line!</p>\n<p>In this sense, random splits probably express the needs of OTTO quite well. Additionally, <code>being creative</code> about splits is an awesome way to shoot yourself in the foot in some way you can't foresee but the tens of Kaggle Grandmasters who are likely to part take in this competition would be likely able to identify 🙂</p>\n<p>Interesting discussion, thanks for starting it! 🙌 Not necessarily arguing against your point, just wanted to share a slightly different perspective on why random splitting might have some advantages 🙂 In general, I find randomness can often be a great ally in our ML pursuits!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2054549": "I noticed almost everyone uses test set split code from organizers [repository](https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py), but I think it might have some flaws.\n\nUsing sessions from last week as a holdout test set makes sense however splitting sessions randomly could be problematic. For example session 11098537:\n![sess1](https://i.imgur.com/H7PKxub.png)\n\nGround-truth belongs to the highlighted event. If you are using aids from the entire session then you might have access to ground-truth click. This can be prevented by finding cutoff index and use aids before that index. It isn't efficient and it requires complex logic if there are duplicate ground-truth aids in sessions. It becomes even more complex when there are ground-truth carts and orders.\n\nAnother thing I noticed is some of the ground-truth labels can be without clicks when they are split randomly. It doesn't necessarily have to be a problem but it feels more natural to have ground-truth clicks.\n\nMy suggestions are\n* Always keep one last click at the end and split sessions randomly before that click event so every session can have a ground-truth click\n* Store cutoff index while splitting sessions so only events before that index can be utilized\n* Filter unseen aids in last week sessions",
    "2054809": "i don't think there are flaws in organizers repository, you can always add rules ( your filters ) to apply to the splinted data set.",
    "2055721": "Another useful way to think about this is from the perspective of the company behind the online store. You want to be able to serve as useful recommendations as you can, even on the first event!\n\nIn fact, there is probably a lot of business value in being able to do something valuable for our prospective customer on the first, second and third event, to make sure they stay for a while longer 🙂 Offering them something of value could go a long way for improving retention and improving the bottom line!\n\nIn this sense, random splits probably express the needs of OTTO quite well. Additionally, `being creative` about splits is an awesome way to shoot yourself in the foot in some way you can't foresee but the tens of Kaggle Grandmasters who are likely to part take in this competition would be likely able to identify 🙂\n\nInteresting discussion, thanks for starting it! 🙌 Not necessarily arguing against your point, just wanted to share a slightly different perspective on why random splitting might have some advantages 🙂 In general, I find randomness can often be a great ally in our ML pursuits!"
  },
  "source": "meta"
}