{
  "id": 364064,
  "title": "The order of appearance of type in each aid",
  "url": "/competitions/otto-recommender-system/discussion/364064",
  "author_name": "",
  "post_date": "2022-11-04T10:21:27.898484500Z",
  "votes": 32,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I thought that <code>type</code> always appeared in the order: <code>click</code> -&gt; <code>cart</code> -&gt; <code>order</code>.<br>\nHowever, from my survey(<a href=\"https://www.kaggle.com/code/tj0612/eda-understand-user-behavior-via-digraph?kernelSessionId=109959389\" target=\"_blank\">notebook</a>), you can see that <code>cart</code> or <code>order</code> appear without <code>click</code>, furthermore, it is possible to have a series of the same <code>type</code> for each <code>aid</code>. </p>\n<p>This means that we need to analyze not only co-occurrence (as is currently used for baselines) but the items which are ordered multiple times.</p>",
  "messages": [
    {
      "id": "2016878",
      "postDate": "11/04/2022 10:21:27",
      "content": "<p>I thought that <code>type</code> always appeared in the order: <code>click</code> -&gt; <code>cart</code> -&gt; <code>order</code>.<br>\nHowever, from my survey(<a href=\"https://www.kaggle.com/code/tj0612/eda-understand-user-behavior-via-digraph?kernelSessionId=109959389\" target=\"_blank\">notebook</a>), you can see that <code>cart</code> or <code>order</code> appear without <code>click</code>, furthermore, it is possible to have a series of the same <code>type</code> for each <code>aid</code>. </p>\n<p>This means that we need to analyze not only co-occurrence (as is currently used for baselines) but the items which are ordered multiple times.</p>",
      "rawMarkdown": "I thought that `type` always appeared in the order: `click` -> `cart` -> `order`.\nHowever, from my survey([notebook](https://www.kaggle.com/code/tj0612/eda-understand-user-behavior-via-digraph?kernelSessionId=109959389)), you can see that `cart` or `order` appear without `click`, furthermore, it is possible to have a series of the same `type` for each `aid`. \n\nThis means that we need to analyze not only co-occurrence (as is currently used for baselines) but the items which are ordered multiple times.",
      "votes": null
    },
    {
      "id": "2017161",
      "postDate": "11/04/2022 14:49:19",
      "content": "<p>I noticed that too. Session 12391984 starts with 32 orders then do 2 clicks and 3 orders. Why there aren't any carts in that session?</p>",
      "rawMarkdown": "I noticed that too. Session 12391984 starts with 32 orders then do 2 clicks and 3 orders. Why there aren't any carts in that session?",
      "votes": null
    },
    {
      "id": "2017166",
      "postDate": "11/04/2022 15:00:19",
      "content": "<p>Thank you for comment.<br>\nThis is just my hypothesis: it may be that only the last action within a certain threshold time is recorded.</p>",
      "rawMarkdown": "Thank you for comment.\nThis is just my hypothesis: it may be that only the last action within a certain threshold time is recorded.",
      "votes": null
    },
    {
      "id": "2017172",
      "postDate": "11/04/2022 15:08:10",
      "content": "<p>No problem. Another thing is what would be the ground-truth of those kind of sessions? For example; session 12839499 has 31 unique orders. Ground-truth of the first time step would have 31 items but recall@20 would penalize missed 11 remaining false negatives. <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> </p>",
      "rawMarkdown": "No problem. Another thing is what would be the ground-truth of those kind of sessions? For example; session 12839499 has 31 unique orders. Ground-truth of the first time step would have 31 items but recall@20 would penalize missed 11 remaining false negatives. @pnormann",
      "votes": null
    },
    {
      "id": "2017181",
      "postDate": "11/04/2022 15:18:55",
      "content": "<p>Certainly. If you idea is correct, the highest value of LB is not 1.0.</p>",
      "rawMarkdown": "Certainly. If you idea is correct, the highest value of LB is not 1.0.",
      "votes": null
    },
    {
      "id": "2018475",
      "postDate": "11/05/2022 18:15:22",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> and <a href=\"https://www.kaggle.com/tj0612\" target=\"_blank\">@tj0612</a>, I can reassure you that you will not be penalized for cases where the ground truth contains more than 20 items. In other words, if you make 20 correct predictions in such a case, you will still get a score of 1.0.</p>",
      "rawMarkdown": "Hey @gunesevitan and @tj0612, I can reassure you that you will not be penalized for cases where the ground truth contains more than 20 items. In other words, if you make 20 correct predictions in such a case, you will still get a score of 1.0.",
      "votes": null
    },
    {
      "id": "2018509",
      "postDate": "11/05/2022 18:53:44",
      "content": "<p><a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> Can you share the formula you are using to compute this competition's metric. According to the classic <code>recall@K</code> formula i think you would not get LB 1.0 when the number of ground truths is greater than 20.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/recall.png\" alt=\"\"></p>\n<p>In the formula above, <code>A_i</code> is the list of ground truth for user <code>i</code>. And <code>P_i</code> is the list of prediction for user <code>i</code>. When the cardinality of <code>A_i</code> is greater than <code>k</code>, we see that the fraction become less than 1.0 regardless of the list of predictions.</p>",
      "rawMarkdown": "pnormann Can you share the formula you are using to compute this competition's metric. According to the classic `recall@K` formula i think you would not get LB 1.0 when the number of ground truths is greater than 20.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/recall.png)\n\nIn the formula above, `A_i` is the list of ground truth for user `i`. And `P_i` is the list of prediction for user `i`. When the cardinality of `A_i` is greater than `k`, we see that the fraction become less than 1.0 regardless of the list of predictions.",
      "votes": null
    },
    {
      "id": "2018510",
      "postDate": "11/05/2022 18:55:23",
      "content": "<p>I am particularly interested to know what score do i get if i make 19 correction predictions when the ground truth has 30 items. Will that be <code>19/30</code> and then if i make one more correct prediction, my score jumps to <code>30/30</code>! Or is the denominator of your fraction <code>min(20, |A_i|)</code> ?</p>",
      "rawMarkdown": "I am particularly interested to know what score do i get if i make 19 correction predictions when the ground truth has 30 items. Will that be `19/30` and then if i make one more correct prediction, my score jumps to `30/30`! Or is the denominator of your fraction `min(20, |A_i|)` ?",
      "votes": null
    },
    {
      "id": "2018522",
      "postDate": "11/05/2022 19:15:38",
      "content": "<p>I searched OTTO's GitHub, i think they modify the recall formula by adding <code>min(20, |A_i|)</code> to formula denominator. Philip, can you confirmed that this competition adds <code>min</code> to the denominator. Here is code from GitHub <a href=\"https://github.com/otto-de/recsys-dataset/blob/17ed879b6da075eb39f918cf1c34e9c92296c13f/src/evaluate.py#L68\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "I searched OTTO's GitHub, i think they modify the recall formula by adding `min(20, |A_i|)` to formula denominator. Philip, can you confirmed that this competition adds `min` to the denominator. Here is code from GitHub [here][1]\n\n[1]: https://github.com/otto-de/recsys-dataset/blob/17ed879b6da075eb39f918cf1c34e9c92296c13f/src/evaluate.py#L68",
      "votes": null
    },
    {
      "id": "2018530",
      "postDate": "11/05/2022 19:23:31",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> You are right. The denominator for the Recall@k we use for this competition is <code>min(20, |A_i|)</code>, as you have already discovered in our code :) I'll try to add a LaTeX formula to the evaluation page to prevent any further misconceptions.</p>",
      "rawMarkdown": "cdeotte You are right. The denominator for the Recall@k we use for this competition is `min(20, |A_i|)`, as you have already discovered in our code :) I'll try to add a LaTeX formula to the evaluation page to prevent any further misconceptions.",
      "votes": null
    },
    {
      "id": "2019198",
      "postDate": "11/06/2022 12:17:21",
      "content": "<p><img src=\"https://i.ibb.co/Wxp3Fbw/Screenshot-from-2022-11-06-15-13-58.png\" alt=\"sess\"></p>\n<p>What about sessions starting with carts or orders? How do you create ground-truth clicks for their first timesteps?</p>",
      "rawMarkdown": "![sess](https://i.ibb.co/Wxp3Fbw/Screenshot-from-2022-11-06-15-13-58.png)\n\nWhat about sessions starting with carts or orders? How do you create ground-truth clicks for their first timesteps?",
      "votes": null
    },
    {
      "id": "2019209",
      "postDate": "11/06/2022 12:24:36",
      "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> In that case, the first cart or order event would not be part of the ground truth and would only serve as an input feature to your model. Please have a look at the <a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/labels.py#L5\" target=\"_blank\"><code>ground_truth</code></a> function from our GitHub repository for more details.</p>",
      "rawMarkdown": "gunesevitan In that case, the first cart or order event would not be part of the ground truth and would only serve as an input feature to your model. Please have a look at the [`ground_truth`](https://github.com/otto-de/recsys-dataset/blob/main/src/labels.py#L5) function from our GitHub repository for more details.",
      "votes": null
    },
    {
      "id": "2020002",
      "postDate": "11/07/2022 06:00:38",
      "content": "<p>Great converstion  here 😊 Turns out I <strong>still</strong> am calculating the score incorrectly in my code 😄 The missing piece I believe is what <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> pointed out, what happens when <code>len(GT) &gt; 20</code></p>\n<p>Thanks for this exchange of comments 🙏</p>",
      "rawMarkdown": "Great converstion  here 😊 Turns out I **still** am calculating the score incorrectly in my code 😄 The missing piece I believe is what @cdeotte pointed out, what happens when `len(GT) > 20`\n\nThanks for this exchange of comments 🙏",
      "votes": null
    },
    {
      "id": "2020937",
      "postDate": "11/07/2022 21:34:09",
      "content": "<p>A LaTeX formula of our metric was added to the evaluation page ✔️</p>",
      "rawMarkdown": "A LaTeX formula of our metric was added to the evaluation page ✔️",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2017161,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "11/04/2022 14:49:19",
      "content": "<p>I noticed that too. Session 12391984 starts with 32 orders then do 2 clicks and 3 orders. Why there aren't any carts in that session?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2017166,
          "author_name": "tj0612",
          "author_url": "",
          "post_date": "11/04/2022 15:00:19",
          "content": "<p>Thank you for comment.<br>\nThis is just my hypothesis: it may be that only the last action within a certain threshold time is recorded.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2017172,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "11/04/2022 15:08:10",
          "content": "<p>No problem. Another thing is what would be the ground-truth of those kind of sessions? For example; session 12839499 has 31 unique orders. Ground-truth of the first time step would have 31 items but recall@20 would penalize missed 11 remaining false negatives. <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2017181,
          "author_name": "tj0612",
          "author_url": "",
          "post_date": "11/04/2022 15:18:55",
          "content": "<p>Certainly. If you idea is correct, the highest value of LB is not 1.0.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2018475,
          "author_name": "pnormann",
          "author_url": "",
          "post_date": "11/05/2022 18:15:22",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> and <a href=\"https://www.kaggle.com/tj0612\" target=\"_blank\">@tj0612</a>, I can reassure you that you will not be penalized for cases where the ground truth contains more than 20 items. In other words, if you make 20 correct predictions in such a case, you will still get a score of 1.0.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2018509,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/05/2022 18:53:44",
          "content": "<p><a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> Can you share the formula you are using to compute this competition's metric. According to the classic <code>recall@K</code> formula i think you would not get LB 1.0 when the number of ground truths is greater than 20.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/recall.png\" alt=\"\"></p>\n<p>In the formula above, <code>A_i</code> is the list of ground truth for user <code>i</code>. And <code>P_i</code> is the list of prediction for user <code>i</code>. When the cardinality of <code>A_i</code> is greater than <code>k</code>, we see that the fraction become less than 1.0 regardless of the list of predictions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2018510,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/05/2022 18:55:23",
          "content": "<p>I am particularly interested to know what score do i get if i make 19 correction predictions when the ground truth has 30 items. Will that be <code>19/30</code> and then if i make one more correct prediction, my score jumps to <code>30/30</code>! Or is the denominator of your fraction <code>min(20, |A_i|)</code> ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2018522,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/05/2022 19:15:38",
          "content": "<p>I searched OTTO's GitHub, i think they modify the recall formula by adding <code>min(20, |A_i|)</code> to formula denominator. Philip, can you confirmed that this competition adds <code>min</code> to the denominator. Here is code from GitHub <a href=\"https://github.com/otto-de/recsys-dataset/blob/17ed879b6da075eb39f918cf1c34e9c92296c13f/src/evaluate.py#L68\" target=\"_blank\">here</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2018530,
          "author_name": "pnormann",
          "author_url": "",
          "post_date": "11/05/2022 19:23:31",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> You are right. The denominator for the Recall@k we use for this competition is <code>min(20, |A_i|)</code>, as you have already discovered in our code :) I'll try to add a LaTeX formula to the evaluation page to prevent any further misconceptions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2019198,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "11/06/2022 12:17:21",
          "content": "<p><img src=\"https://i.ibb.co/Wxp3Fbw/Screenshot-from-2022-11-06-15-13-58.png\" alt=\"sess\"></p>\n<p>What about sessions starting with carts or orders? How do you create ground-truth clicks for their first timesteps?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2019209,
          "author_name": "pnormann",
          "author_url": "",
          "post_date": "11/06/2022 12:24:36",
          "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> In that case, the first cart or order event would not be part of the ground truth and would only serve as an input feature to your model. Please have a look at the <a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/labels.py#L5\" target=\"_blank\"><code>ground_truth</code></a> function from our GitHub repository for more details.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2020002,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/07/2022 06:00:38",
          "content": "<p>Great converstion  here 😊 Turns out I <strong>still</strong> am calculating the score incorrectly in my code 😄 The missing piece I believe is what <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> pointed out, what happens when <code>len(GT) &gt; 20</code></p>\n<p>Thanks for this exchange of comments 🙏</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2020937,
          "author_name": "pnormann",
          "author_url": "",
          "post_date": "11/07/2022 21:34:09",
          "content": "<p>A LaTeX formula of our metric was added to the evaluation page ✔️</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2016878": "I thought that `type` always appeared in the order: `click` -> `cart` -> `order`.\nHowever, from my survey([notebook](https://www.kaggle.com/code/tj0612/eda-understand-user-behavior-via-digraph?kernelSessionId=109959389)), you can see that `cart` or `order` appear without `click`, furthermore, it is possible to have a series of the same `type` for each `aid`. \n\nThis means that we need to analyze not only co-occurrence (as is currently used for baselines) but the items which are ordered multiple times.",
    "2017161": "I noticed that too. Session 12391984 starts with 32 orders then do 2 clicks and 3 orders. Why there aren't any carts in that session?",
    "2017166": "Thank you for comment.\nThis is just my hypothesis: it may be that only the last action within a certain threshold time is recorded.",
    "2017172": "No problem. Another thing is what would be the ground-truth of those kind of sessions? For example; session 12839499 has 31 unique orders. Ground-truth of the first time step would have 31 items but recall@20 would penalize missed 11 remaining false negatives. @pnormann",
    "2017181": "Certainly. If you idea is correct, the highest value of LB is not 1.0.",
    "2018475": "Hey @gunesevitan and @tj0612, I can reassure you that you will not be penalized for cases where the ground truth contains more than 20 items. In other words, if you make 20 correct predictions in such a case, you will still get a score of 1.0.",
    "2018509": "pnormann Can you share the formula you are using to compute this competition's metric. According to the classic `recall@K` formula i think you would not get LB 1.0 when the number of ground truths is greater than 20.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/recall.png)\n\nIn the formula above, `A_i` is the list of ground truth for user `i`. And `P_i` is the list of prediction for user `i`. When the cardinality of `A_i` is greater than `k`, we see that the fraction become less than 1.0 regardless of the list of predictions.",
    "2018510": "I am particularly interested to know what score do i get if i make 19 correction predictions when the ground truth has 30 items. Will that be `19/30` and then if i make one more correct prediction, my score jumps to `30/30`! Or is the denominator of your fraction `min(20, |A_i|)` ?",
    "2018522": "I searched OTTO's GitHub, i think they modify the recall formula by adding `min(20, |A_i|)` to formula denominator. Philip, can you confirmed that this competition adds `min` to the denominator. Here is code from GitHub [here][1]\n\n[1]: https://github.com/otto-de/recsys-dataset/blob/17ed879b6da075eb39f918cf1c34e9c92296c13f/src/evaluate.py#L68",
    "2018530": "cdeotte You are right. The denominator for the Recall@k we use for this competition is `min(20, |A_i|)`, as you have already discovered in our code :) I'll try to add a LaTeX formula to the evaluation page to prevent any further misconceptions.",
    "2019198": "![sess](https://i.ibb.co/Wxp3Fbw/Screenshot-from-2022-11-06-15-13-58.png)\n\nWhat about sessions starting with carts or orders? How do you create ground-truth clicks for their first timesteps?",
    "2019209": "gunesevitan In that case, the first cart or order event would not be part of the ground truth and would only serve as an input feature to your model. Please have a look at the [`ground_truth`](https://github.com/otto-de/recsys-dataset/blob/main/src/labels.py#L5) function from our GitHub repository for more details.",
    "2020002": "Great converstion  here 😊 Turns out I **still** am calculating the score incorrectly in my code 😄 The missing piece I believe is what @cdeotte pointed out, what happens when `len(GT) > 20`\n\nThanks for this exchange of comments 🙏",
    "2020937": "A LaTeX formula of our metric was added to the evaluation page ✔️"
  },
  "source": "meta"
}