{
  "id": 363792,
  "title": "Ground-truth?",
  "url": "/competitions/otto-recommender-system/discussion/363792",
  "author_name": "",
  "post_date": "2022-11-03T06:48:27.946348500Z",
  "votes": 16,
  "comment_count": 6,
  "views": 0,
  "content": "<p>This may sound so simple but I'm trying to understand what \"test set is truncated\" means. Is it truncated by only one event? If that's the case, can we create ground-truth for training set by simply doing</p>\n<p><code>df_train.groupby(['session', 'type'])['aid'].shift(-1)</code>?</p>\n<p>Edit:<br>\n<a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/labels.py#L4-L25\" target=\"_blank\">https://github.com/otto-de/recsys-dataset/blob/main/src/labels.py#L4-L25</a></p>\n<p>I found this code and if I understood correctly <code>df_train.groupby(['session', 'type'])['aid'].shift(-1)</code> for ground-truth is only correct for clicks, but carts and orders work like a queue. They should have the last 20 timesteps in an array. This code doesn't have that functionality so I'm trying to figure what happens when there are more than 20 carts or orders in a session.</p>",
  "messages": [
    {
      "id": "2015214",
      "postDate": "11/03/2022 06:48:27",
      "content": "<p>This may sound so simple but I'm trying to understand what \"test set is truncated\" means. Is it truncated by only one event? If that's the case, can we create ground-truth for training set by simply doing</p>\n<p><code>df_train.groupby(['session', 'type'])['aid'].shift(-1)</code>?</p>\n<p>Edit:<br>\n<a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/labels.py#L4-L25\" target=\"_blank\">https://github.com/otto-de/recsys-dataset/blob/main/src/labels.py#L4-L25</a></p>\n<p>I found this code and if I understood correctly <code>df_train.groupby(['session', 'type'])['aid'].shift(-1)</code> for ground-truth is only correct for clicks, but carts and orders work like a queue. They should have the last 20 timesteps in an array. This code doesn't have that functionality so I'm trying to figure what happens when there are more than 20 carts or orders in a session.</p>",
      "rawMarkdown": "This may sound so simple but I'm trying to understand what \"test set is truncated\" means. Is it truncated by only one event? If that's the case, can we create ground-truth for training set by simply doing\n\n`df_train.groupby(['session', 'type'])['aid'].shift(-1)`?\n\nEdit:\nhttps://github.com/otto-de/recsys-dataset/blob/main/src/labels.py#L4-L25\n\nI found this code and if I understood correctly `df_train.groupby(['session', 'type'])['aid'].shift(-1)` for ground-truth is only correct for clicks, but carts and orders work like a queue. They should have the last 20 timesteps in an array. This code doesn't have that functionality so I'm trying to figure what happens when there are more than 20 carts or orders in a session.",
      "votes": null
    },
    {
      "id": "2015321",
      "postDate": "11/03/2022 07:52:40",
      "content": "<p>From their <a href=\"https://github.com/otto-de/recsys-dataset\" target=\"_blank\">github</a>:<br>\nFor clicks there is only a single ground truth value for each session, which is the next aid clicked during the session (although you can still predict up to 20 aid values). The ground truth for carts and orders contains all aid values that were added to a cart and ordered respectively during the session. […]<br>\nOur train set consists of observations from 4 weeks, while the test set contains user sessions from the following week. Furthermore, we trimmed train sessions overlapping with the test period. <br>\n<img src=\"https://github.com/otto-de/recsys-dataset/raw/main/.readme/train_test_split.png\" alt=\"\">. </p>",
      "rawMarkdown": "From their [github](https://github.com/otto-de/recsys-dataset):\nFor clicks there is only a single ground truth value for each session, which is the next aid clicked during the session (although you can still predict up to 20 aid values). The ground truth for carts and orders contains all aid values that were added to a cart and ordered respectively during the session. [...]\nOur train set consists of observations from 4 weeks, while the test set contains user sessions from the following week. Furthermore, we trimmed train sessions overlapping with the test period. \n![](https://github.com/otto-de/recsys-dataset/raw/main/.readme/train_test_split.png).",
      "votes": null
    },
    {
      "id": "2016686",
      "postDate": "11/04/2022 07:16:40",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>!</p>\n<p>This is <a href=\"https://github.com/otto-de/recsys-dataset/blob/5d9c5f6b0a82cb09f4b98542c77a774facaf51c7/src/testset.py#L22\" target=\"_blank\">the code</a> I believe was used to split individual sessions into test data that we received and ground truth.</p>\n<p>There is also a bit more to how the ground truth data was created, from what I understand. I wrote a bit more about it <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363965\" target=\"_blank\">here</a>, maybe it can be of help.</p>",
      "rawMarkdown": "Hey @gunesevitan!\n\nThis is [the code](https://github.com/otto-de/recsys-dataset/blob/5d9c5f6b0a82cb09f4b98542c77a774facaf51c7/src/testset.py#L22) I believe was used to split individual sessions into test data that we received and ground truth.\n\nThere is also a bit more to how the ground truth data was created, from what I understand. I wrote a bit more about it [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363965), maybe it can be of help.",
      "votes": null
    },
    {
      "id": "2018512",
      "postDate": "11/05/2022 19:01:44",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> </p>\n<p>What do you think is the correct prediction for orders where the user didn't order any product during the session (like ~75% of users in the training dataset)?<br>\nShould add an empty string alongside the other 19 possible values?</p>",
      "rawMarkdown": "Hey @radek1 \n\n\nWhat do you think is the correct prediction for orders where the user didn't order any product during the session (like ~75% of users in the training dataset)?\nShould add an empty string alongside the other 19 possible values?",
      "votes": null
    },
    {
      "id": "2018697",
      "postDate": "11/05/2022 22:45:31",
      "content": "<p>It all boils down to how the evaluation metric has been implemented (it hasn't been shared with us, unfortunately), but I don't think this is necessary. But that would be extremely unusual for an empty string to make a difference!</p>\n<p>Essentially, the score is recall based, and all recall cares about is \"recalling\" the ground truth entries that are there! In the case where there is not ground truth, I don't think we would be getting any score for that entry (neither 0 nor 1, this line will just not be taken into consideration for score calculation at all). That would be my guess 🙂</p>",
      "rawMarkdown": "It all boils down to how the evaluation metric has been implemented (it hasn't been shared with us, unfortunately), but I don't think this is necessary. But that would be extremely unusual for an empty string to make a difference!\n\nEssentially, the score is recall based, and all recall cares about is \"recalling\" the ground truth entries that are there! In the case where there is not ground truth, I don't think we would be getting any score for that entry (neither 0 nor 1, this line will just not be taken into consideration for score calculation at all). That would be my guess 🙂",
      "votes": null
    },
    {
      "id": "2018873",
      "postDate": "11/06/2022 05:14:47",
      "content": "<p>It makes more sense to ignore those timesteps. Setting 0 or 1 could lead to under/over estimating the error.</p>",
      "rawMarkdown": "It makes more sense to ignore those timesteps. Setting 0 or 1 could lead to under/over estimating the error.",
      "votes": null
    },
    {
      "id": "2019194",
      "postDate": "11/06/2022 12:16:37",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> and <a href=\"https://www.kaggle.com/hamolyavlad\" target=\"_blank\">@hamolyavlad</a>, we published all the scripts you need for a local evaluation in our <a href=\"https://github.com/otto-de/recsys-dataset\" target=\"_blank\">GitHub repo</a>, which is equivalent to the logic used for the leaderboard here on Kaggle. You can create a validation set from the train data using the <a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py\" target=\"_blank\">testset.py</a> and evaluate your predictions using the <a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/evaluate.py\" target=\"_blank\">evalute.py</a>. To answer your question - <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> is right in his guess that customers who didn't make any purchases or didn't add anything to their cart during the test period don't contribute to the respective recall scores.</p>",
      "rawMarkdown": "Hey @radek1 and @hamolyavlad, we published all the scripts you need for a local evaluation in our [GitHub repo](https://github.com/otto-de/recsys-dataset), which is equivalent to the logic used for the leaderboard here on Kaggle. You can create a validation set from the train data using the [testset.py](https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py) and evaluate your predictions using the [evalute.py](https://github.com/otto-de/recsys-dataset/blob/main/src/evaluate.py). To answer your question - @radek1 is right in his guess that customers who didn't make any purchases or didn't add anything to their cart during the test period don't contribute to the respective recall scores.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2015321,
      "author_name": "tompaulat",
      "author_url": "",
      "post_date": "11/03/2022 07:52:40",
      "content": "<p>From their <a href=\"https://github.com/otto-de/recsys-dataset\" target=\"_blank\">github</a>:<br>\nFor clicks there is only a single ground truth value for each session, which is the next aid clicked during the session (although you can still predict up to 20 aid values). The ground truth for carts and orders contains all aid values that were added to a cart and ordered respectively during the session. […]<br>\nOur train set consists of observations from 4 weeks, while the test set contains user sessions from the following week. Furthermore, we trimmed train sessions overlapping with the test period. <br>\n<img src=\"https://github.com/otto-de/recsys-dataset/raw/main/.readme/train_test_split.png\" alt=\"\">. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2016686,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "11/04/2022 07:16:40",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a>!</p>\n<p>This is <a href=\"https://github.com/otto-de/recsys-dataset/blob/5d9c5f6b0a82cb09f4b98542c77a774facaf51c7/src/testset.py#L22\" target=\"_blank\">the code</a> I believe was used to split individual sessions into test data that we received and ground truth.</p>\n<p>There is also a bit more to how the ground truth data was created, from what I understand. I wrote a bit more about it <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363965\" target=\"_blank\">here</a>, maybe it can be of help.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2018512,
          "author_name": "hamolyavlad",
          "author_url": "",
          "post_date": "11/05/2022 19:01:44",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> </p>\n<p>What do you think is the correct prediction for orders where the user didn't order any product during the session (like ~75% of users in the training dataset)?<br>\nShould add an empty string alongside the other 19 possible values?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2018697,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/05/2022 22:45:31",
          "content": "<p>It all boils down to how the evaluation metric has been implemented (it hasn't been shared with us, unfortunately), but I don't think this is necessary. But that would be extremely unusual for an empty string to make a difference!</p>\n<p>Essentially, the score is recall based, and all recall cares about is \"recalling\" the ground truth entries that are there! In the case where there is not ground truth, I don't think we would be getting any score for that entry (neither 0 nor 1, this line will just not be taken into consideration for score calculation at all). That would be my guess 🙂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2018873,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "11/06/2022 05:14:47",
          "content": "<p>It makes more sense to ignore those timesteps. Setting 0 or 1 could lead to under/over estimating the error.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2019194,
          "author_name": "pnormann",
          "author_url": "",
          "post_date": "11/06/2022 12:16:37",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> and <a href=\"https://www.kaggle.com/hamolyavlad\" target=\"_blank\">@hamolyavlad</a>, we published all the scripts you need for a local evaluation in our <a href=\"https://github.com/otto-de/recsys-dataset\" target=\"_blank\">GitHub repo</a>, which is equivalent to the logic used for the leaderboard here on Kaggle. You can create a validation set from the train data using the <a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py\" target=\"_blank\">testset.py</a> and evaluate your predictions using the <a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/evaluate.py\" target=\"_blank\">evalute.py</a>. To answer your question - <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> is right in his guess that customers who didn't make any purchases or didn't add anything to their cart during the test period don't contribute to the respective recall scores.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2015214": "This may sound so simple but I'm trying to understand what \"test set is truncated\" means. Is it truncated by only one event? If that's the case, can we create ground-truth for training set by simply doing\n\n`df_train.groupby(['session', 'type'])['aid'].shift(-1)`?\n\nEdit:\nhttps://github.com/otto-de/recsys-dataset/blob/main/src/labels.py#L4-L25\n\nI found this code and if I understood correctly `df_train.groupby(['session', 'type'])['aid'].shift(-1)` for ground-truth is only correct for clicks, but carts and orders work like a queue. They should have the last 20 timesteps in an array. This code doesn't have that functionality so I'm trying to figure what happens when there are more than 20 carts or orders in a session.",
    "2015321": "From their [github](https://github.com/otto-de/recsys-dataset):\nFor clicks there is only a single ground truth value for each session, which is the next aid clicked during the session (although you can still predict up to 20 aid values). The ground truth for carts and orders contains all aid values that were added to a cart and ordered respectively during the session. [...]\nOur train set consists of observations from 4 weeks, while the test set contains user sessions from the following week. Furthermore, we trimmed train sessions overlapping with the test period. \n![](https://github.com/otto-de/recsys-dataset/raw/main/.readme/train_test_split.png).",
    "2016686": "Hey @gunesevitan!\n\nThis is [the code](https://github.com/otto-de/recsys-dataset/blob/5d9c5f6b0a82cb09f4b98542c77a774facaf51c7/src/testset.py#L22) I believe was used to split individual sessions into test data that we received and ground truth.\n\nThere is also a bit more to how the ground truth data was created, from what I understand. I wrote a bit more about it [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363965), maybe it can be of help.",
    "2018512": "Hey @radek1 \n\n\nWhat do you think is the correct prediction for orders where the user didn't order any product during the session (like ~75% of users in the training dataset)?\nShould add an empty string alongside the other 19 possible values?",
    "2018697": "It all boils down to how the evaluation metric has been implemented (it hasn't been shared with us, unfortunately), but I don't think this is necessary. But that would be extremely unusual for an empty string to make a difference!\n\nEssentially, the score is recall based, and all recall cares about is \"recalling\" the ground truth entries that are there! In the case where there is not ground truth, I don't think we would be getting any score for that entry (neither 0 nor 1, this line will just not be taken into consideration for score calculation at all). That would be my guess 🙂",
    "2018873": "It makes more sense to ignore those timesteps. Setting 0 or 1 could lead to under/over estimating the error.",
    "2019194": "Hey @radek1 and @hamolyavlad, we published all the scripts you need for a local evaluation in our [GitHub repo](https://github.com/otto-de/recsys-dataset), which is equivalent to the logic used for the leaderboard here on Kaggle. You can create a validation set from the train data using the [testset.py](https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py) and evaluate your predictions using the [evalute.py](https://github.com/otto-de/recsys-dataset/blob/main/src/evaluate.py). To answer your question - @radek1 is right in his guess that customers who didn't make any purchases or didn't add anything to their cart during the test period don't contribute to the respective recall scores."
  },
  "source": "meta"
}