{
  "id": 363869,
  "title": "Question about the definition of the ground truth",
  "url": "/competitions/otto-recommender-system/discussion/363869",
  "author_name": "",
  "post_date": "2022-11-03T12:56:29.816893100Z",
  "votes": 12,
  "comment_count": 4,
  "views": 0,
  "content": "<p><a href=\"https://github.com/otto-de/recsys-dataset\" target=\"_blank\">The host's github</a> provides the following explanation and diagram regarding ground truth.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640255%2F7fd6ce025145f71f6d2b1bb6a432b310%2Fground_truth.png?generation=1667478424007389&amp;alt=media\" alt=\"ground truth\"></p>\n<blockquote>\n  <p>For clicks there is only a single ground truth value for each session, which is the next aid clicked during the session (although you can still predict up to 20 aid values). The ground truth for carts and orders contains all aid values that were added to a cart and ordered respectively during the session.</p>\n</blockquote>\n<p>I understand about the ground truth for clicks, but I have a question about the ground truth for carts and orders.</p>\n<p>which of the following two would be correct for our task?</p>\n<p>(1) predict items <strong>newly</strong> added to a cart and ordered <strong>after the time listed in test.jsonl</strong></p>\n<p>(2) predict items added to a cart and ordered during the session <strong>including the time listed in test.jsonl</strong><br>\n(Therefore, all history in test.jsonl will be correct if submitted as is. Of course, some items were ordered after the time listed in test.jsonl, so that alone will not result in LB: 1.0.)</p>\n<p>LB score is high (Recall@20: 0.52), probably because the score is calculated using method (2), but on the other hand, by looking at the above figure, method (1) appears to be more correct.</p>\n<p>Let me know what you guys think.</p>",
  "messages": [
    {
      "id": "2015669",
      "postDate": "11/03/2022 12:56:29",
      "content": "<p><a href=\"https://github.com/otto-de/recsys-dataset\" target=\"_blank\">The host's github</a> provides the following explanation and diagram regarding ground truth.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640255%2F7fd6ce025145f71f6d2b1bb6a432b310%2Fground_truth.png?generation=1667478424007389&amp;alt=media\" alt=\"ground truth\"></p>\n<blockquote>\n  <p>For clicks there is only a single ground truth value for each session, which is the next aid clicked during the session (although you can still predict up to 20 aid values). The ground truth for carts and orders contains all aid values that were added to a cart and ordered respectively during the session.</p>\n</blockquote>\n<p>I understand about the ground truth for clicks, but I have a question about the ground truth for carts and orders.</p>\n<p>which of the following two would be correct for our task?</p>\n<p>(1) predict items <strong>newly</strong> added to a cart and ordered <strong>after the time listed in test.jsonl</strong></p>\n<p>(2) predict items added to a cart and ordered during the session <strong>including the time listed in test.jsonl</strong><br>\n(Therefore, all history in test.jsonl will be correct if submitted as is. Of course, some items were ordered after the time listed in test.jsonl, so that alone will not result in LB: 1.0.)</p>\n<p>LB score is high (Recall@20: 0.52), probably because the score is calculated using method (2), but on the other hand, by looking at the above figure, method (1) appears to be more correct.</p>\n<p>Let me know what you guys think.</p>",
      "rawMarkdown": "[The host's github](https://github.com/otto-de/recsys-dataset) provides the following explanation and diagram regarding ground truth.\n\n![ground truth](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640255%2F7fd6ce025145f71f6d2b1bb6a432b310%2Fground_truth.png?generation=1667478424007389&alt=media)\n\n> For clicks there is only a single ground truth value for each session, which is the next aid clicked during the session (although you can still predict up to 20 aid values). The ground truth for carts and orders contains all aid values that were added to a cart and ordered respectively during the session.\n\nI understand about the ground truth for clicks, but I have a question about the ground truth for carts and orders.\n\nwhich of the following two would be correct for our task?\n\n(1) predict items **newly** added to a cart and ordered **after the time listed in test.jsonl**\n\n(2) predict items added to a cart and ordered during the session **including the time listed in test.jsonl**\n(Therefore, all history in test.jsonl will be correct if submitted as is. Of course, some items were ordered after the time listed in test.jsonl, so that alone will not result in LB: 1.0.)\n\nLB score is high (Recall@20: 0.52), probably because the score is calculated using method (2), but on the other hand, by looking at the above figure, method (1) appears to be more correct.\n\nLet me know what you guys think.",
      "votes": null
    },
    {
      "id": "2015831",
      "postDate": "11/03/2022 15:10:25",
      "content": "<p>I think it's method 1 🙂 Not sure it would make a lot of sense for it to be method 2.</p>\n<p>The recall might be that high because there is a high correlation between what one clicks and what one buys down the road, I would imagine. There is a really interesting discussion on why the score might be what it is in this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363874\" target=\"_blank\">thread</a> by <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> </p>",
      "rawMarkdown": "I think it's method 1 🙂 Not sure it would make a lot of sense for it to be method 2.\n\nThe recall might be that high because there is a high correlation between what one clicks and what one buys down the road, I would imagine. There is a really interesting discussion on why the score might be what it is in this [thread](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363874) by @narsil",
      "votes": null
    },
    {
      "id": "2015857",
      "postDate": "11/03/2022 15:26:20",
      "content": "<p>When I try to get a validation score using train.jsonl, I get a score close to LB with method 2, and with method 1, I don't get such a good score…</p>\n<p>I agree that method 2 makes no sense.</p>",
      "rawMarkdown": "When I try to get a validation score using train.jsonl, I get a score close to LB with method 2, and with method 1, I don't get such a good score...\n\nI agree that method 2 makes no sense.",
      "votes": null
    },
    {
      "id": "2015881",
      "postDate": "11/03/2022 15:51:46",
      "content": "<p>I think it's method 1 as well, and I got similar score.<br>\nDataset Description says</p>\n<blockquote>\n  <p>your task is to predict the next aid clicked after the session truncation</p>\n</blockquote>",
      "rawMarkdown": "I think it's method 1 as well, and I got similar score.\nDataset Description says\n> your task is to predict the next aid clicked after the session truncation",
      "votes": null
    },
    {
      "id": "2015894",
      "postDate": "11/03/2022 15:57:48",
      "content": "<p>Thanks for your comment!<br>\nSo my calculation of the validation score is wrong…</p>",
      "rawMarkdown": "Thanks for your comment!\nSo my calculation of the validation score is wrong...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2015831,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "11/03/2022 15:10:25",
      "content": "<p>I think it's method 1 🙂 Not sure it would make a lot of sense for it to be method 2.</p>\n<p>The recall might be that high because there is a high correlation between what one clicks and what one buys down the road, I would imagine. There is a really interesting discussion on why the score might be what it is in this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363874\" target=\"_blank\">thread</a> by <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 2015857,
          "author_name": "shkanda",
          "author_url": "",
          "post_date": "11/03/2022 15:26:20",
          "content": "<p>When I try to get a validation score using train.jsonl, I get a score close to LB with method 2, and with method 1, I don't get such a good score…</p>\n<p>I agree that method 2 makes no sense.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2015881,
          "author_name": "onodera",
          "author_url": "",
          "post_date": "11/03/2022 15:51:46",
          "content": "<p>I think it's method 1 as well, and I got similar score.<br>\nDataset Description says</p>\n<blockquote>\n  <p>your task is to predict the next aid clicked after the session truncation</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2015894,
          "author_name": "shkanda",
          "author_url": "",
          "post_date": "11/03/2022 15:57:48",
          "content": "<p>Thanks for your comment!<br>\nSo my calculation of the validation score is wrong…</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2015669": "[The host's github](https://github.com/otto-de/recsys-dataset) provides the following explanation and diagram regarding ground truth.\n\n![ground truth](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2640255%2F7fd6ce025145f71f6d2b1bb6a432b310%2Fground_truth.png?generation=1667478424007389&alt=media)\n\n> For clicks there is only a single ground truth value for each session, which is the next aid clicked during the session (although you can still predict up to 20 aid values). The ground truth for carts and orders contains all aid values that were added to a cart and ordered respectively during the session.\n\nI understand about the ground truth for clicks, but I have a question about the ground truth for carts and orders.\n\nwhich of the following two would be correct for our task?\n\n(1) predict items **newly** added to a cart and ordered **after the time listed in test.jsonl**\n\n(2) predict items added to a cart and ordered during the session **including the time listed in test.jsonl**\n(Therefore, all history in test.jsonl will be correct if submitted as is. Of course, some items were ordered after the time listed in test.jsonl, so that alone will not result in LB: 1.0.)\n\nLB score is high (Recall@20: 0.52), probably because the score is calculated using method (2), but on the other hand, by looking at the above figure, method (1) appears to be more correct.\n\nLet me know what you guys think.",
    "2015831": "I think it's method 1 🙂 Not sure it would make a lot of sense for it to be method 2.\n\nThe recall might be that high because there is a high correlation between what one clicks and what one buys down the road, I would imagine. There is a really interesting discussion on why the score might be what it is in this [thread](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363874) by @narsil",
    "2015857": "When I try to get a validation score using train.jsonl, I get a score close to LB with method 2, and with method 1, I don't get such a good score...\n\nI agree that method 2 makes no sense.",
    "2015881": "I think it's method 1 as well, and I got similar score.\nDataset Description says\n> your task is to predict the next aid clicked after the session truncation",
    "2015894": "Thanks for your comment!\nSo my calculation of the validation score is wrong..."
  },
  "source": "meta"
}