{
  "id": 368149,
  "title": "Products in test and not in train?",
  "url": "/competitions/otto-recommender-system/discussion/368149",
  "author_name": "",
  "post_date": "2022-11-23T18:54:10.285651Z",
  "votes": 17,
  "comment_count": 7,
  "views": 0,
  "content": "<p>While further looking at the data I noticed that there are:</p>\n<ul>\n<li>951,294 unique products in the <code>train</code> set</li>\n<li>259,445 unique products in the <code>test</code> set, out of which:<ul>\n<li>211,446 products (81%) can be found in the <code>train</code> set (so they repeat there at least once)</li>\n<li>47,999 products (~19%) that <strong>cannot</strong> be found in the <code>train</code> set</li></ul></li>\n</ul>\n<p>This means that these ~19% products will <em>never</em> be seen during training.</p>\n<p>Scary 😂 So there might be some products which we'll have to predict that aren't even in the train set to begin with?</p>\n<p><img src=\"https://i.imgur.com/1j1dcxI.jpg\"></p>\n<p>PS: I am using <a href=\"https://www.kaggle.com/datasets/konradb/otto-dataset-in-dataframe\" target=\"_blank\">this dataset</a>.</p>",
  "messages": [
    {
      "id": "2041250",
      "postDate": "11/23/2022 18:54:10",
      "content": "<p>While further looking at the data I noticed that there are:</p>\n<ul>\n<li>951,294 unique products in the <code>train</code> set</li>\n<li>259,445 unique products in the <code>test</code> set, out of which:<ul>\n<li>211,446 products (81%) can be found in the <code>train</code> set (so they repeat there at least once)</li>\n<li>47,999 products (~19%) that <strong>cannot</strong> be found in the <code>train</code> set</li></ul></li>\n</ul>\n<p>This means that these ~19% products will <em>never</em> be seen during training.</p>\n<p>Scary 😂 So there might be some products which we'll have to predict that aren't even in the train set to begin with?</p>\n<p><img src=\"https://i.imgur.com/1j1dcxI.jpg\"></p>\n<p>PS: I am using <a href=\"https://www.kaggle.com/datasets/konradb/otto-dataset-in-dataframe\" target=\"_blank\">this dataset</a>.</p>",
      "rawMarkdown": "While further looking at the data I noticed that there are:\n* 951,294 unique products in the `train` set\n* 259,445 unique products in the `test` set, out of which:\n * 211,446 products (81%) can be found in the `train` set (so they repeat there at least once)\n * 47,999 products (~19%) that **cannot** be found in the `train` set\n\nThis means that these ~19% products will *never* be seen during training.\n\nScary 😂 So there might be some products which we'll have to predict that aren't even in the train set to begin with?\n\n<img src=\"https://i.imgur.com/1j1dcxI.jpg\">\n\nPS: I am using [this dataset](https://www.kaggle.com/datasets/konradb/otto-dataset-in-dataframe).",
      "votes": null
    },
    {
      "id": "2041323",
      "postDate": "11/23/2022 20:25:05",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a> . I've checked it before and based on <a href=\"https://www.kaggle.com/code/piotrekga/new-aids-in-test-set\" target=\"_blank\">this</a> notebook all the aids from the test set are present in the train set . The notebooks itself refers to this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991#2029433\" target=\"_blank\">discussion</a> where we talked about code to create train/test split fro the official repo. Am I missing something?</p>",
      "rawMarkdown": "Hi @andradaolteanu . I've checked it before and based on [this](https://www.kaggle.com/code/piotrekga/new-aids-in-test-set) notebook all the aids from the test set are present in the train set . The notebooks itself refers to this [discussion](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991#2029433) where we talked about code to create train/test split fro the official repo. Am I missing something?",
      "votes": null
    },
    {
      "id": "2041448",
      "postDate": "11/24/2022 00:29:14",
      "content": "<p>Yes, that is a great point, <a href=\"https://www.kaggle.com/piotrekga\" target=\"_blank\">@piotrekga</a>! Also, the counts of unique <code>aids</code> that <a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a> reports seem to be on the lower side then what is in the full data.</p>\n<p><a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a>, if that might be of help, I put together a dataset that I made <a href=\"https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint\" target=\"_blank\">available here</a> that a bunch of us has been using, there is a good degree of confidence that it has been processed correctly 🙂 Maybe please give it a go and see what results you get?</p>\n<p>Also, I absolutely love the visualization you created in your OP 🙂 is that your handwriting? it would be so cool if it was in the era where people are losing their ability to write by hand due to using keyboards all the time 😄 (at least I know this has been happening to me 😉) very neat visualization either way!</p>",
      "rawMarkdown": "Yes, that is a great point, @piotrekga! Also, the counts of unique `aids` that @andradaolteanu reports seem to be on the lower side then what is in the full data.\n\n@andradaolteanu, if that might be of help, I put together a dataset that I made [available here](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint) that a bunch of us has been using, there is a good degree of confidence that it has been processed correctly 🙂 Maybe please give it a go and see what results you get?\n\nAlso, I absolutely love the visualization you created in your OP 🙂 is that your handwriting? it would be so cool if it was in the era where people are losing their ability to write by hand due to using keyboards all the time 😄 (at least I know this has been happening to me 😉) very neat visualization either way!",
      "votes": null
    },
    {
      "id": "2041539",
      "postDate": "11/24/2022 03:25:58",
      "content": "<p>I recommend using the original competition dataset to calculate nunique and intersections. According the hosts <a href=\"https://github.com/otto-de/recsys-dataset\" target=\"_blank\">here</a> there are 1.855.603 unique aid just in trainset. </p>",
      "rawMarkdown": "I recommend using the original competition dataset to calculate nunique and intersections. According the hosts [here](https://github.com/otto-de/recsys-dataset) there are 1.855.603 unique aid just in trainset.",
      "votes": null
    },
    {
      "id": "2042328",
      "postDate": "11/24/2022 16:08:00",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> ! Thank you lots for replying.</p>\n<p>So just by looking at <code>aid</code> in the <code>train</code> and <code>test</code> <a href=\"https://www.kaggle.com/datasets/konradb/otto-dataset-in-dataframe\" target=\"_blank\">datasets</a>, this is what I get:<br>\n<img src=\"https://i.imgur.com/rO4G0vn.png\"></p>\n<p>But I see indeed in your dataset you have a great deal more data. What is the difference between your dataset and Konrad's dataset? Because I have been using that one so far and now I am lost completely 😂</p>\n<p>PS: yup, I usually draw when I have a bunch of information, it helps me understand better. </p>",
      "rawMarkdown": "Hi @radek1 ! Thank you lots for replying.\n\nSo just by looking at `aid` in the `train` and `test` [datasets](https://www.kaggle.com/datasets/konradb/otto-dataset-in-dataframe), this is what I get:\n<img src=\"https://i.imgur.com/rO4G0vn.png\">\n\nBut I see indeed in your dataset you have a great deal more data. What is the difference between your dataset and Konrad's dataset? Because I have been using that one so far and now I am lost completely 😂\n\nPS: yup, I usually draw when I have a bunch of information, it helps me understand better.",
      "votes": null
    },
    {
      "id": "2042404",
      "postDate": "11/24/2022 16:59:06",
      "content": "<p>Konrad said he will update his dataset <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368156#2041888\" target=\"_blank\">https://www.kaggle.com/competitions/otto-recommender-system/discussion/368156#2041888</a></p>",
      "rawMarkdown": "Konrad said he will update his dataset https://www.kaggle.com/competitions/otto-recommender-system/discussion/368156#2041888",
      "votes": null
    },
    {
      "id": "2042427",
      "postDate": "11/24/2022 17:11:51",
      "content": "<p>Shoot, missed that one 😢. Sorry for the confussion you guys, guess I'm redoing the whole thing.</p>",
      "rawMarkdown": "Shoot, missed that one 😢. Sorry for the confussion you guys, guess I'm redoing the whole thing.",
      "votes": null
    },
    {
      "id": "2042714",
      "postDate": "11/24/2022 22:56:32",
      "content": "<p>Hey, no worries at all! 🙂 Sorry you have to go through all this trouble, I hate redoing work 😄</p>\n<p>I created a whole \"ecosystem\" of these datasets, in case they might be useful to you 🙂</p>\n<p><a href=\"https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint\" target=\"_blank\">Here</a> is the main one -- essentially memory-optimized data from the competition in a parquet format.</p>\n<p>But <a href=\"https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation\" target=\"_blank\">here</a> is a particularly fun one. I created it using scripts from organizers repo! Essentially, you get a validation set, created like in the competition, with exactly the same code, one the last week of the train set. This can be very useful both for training models but also for analysis, since you get a validation set created like the test set in the competition, but also have the labels 🙂 And the format is exactly the same! (I also minimized the memory foot print and transformed it to parquet).</p>",
      "rawMarkdown": "Hey, no worries at all! 🙂 Sorry you have to go through all this trouble, I hate redoing work 😄\n\nI created a whole \"ecosystem\" of these datasets, in case they might be useful to you 🙂\n\n[Here](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint) is the main one -- essentially memory-optimized data from the competition in a parquet format.\n\nBut [here](https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation) is a particularly fun one. I created it using scripts from organizers repo! Essentially, you get a validation set, created like in the competition, with exactly the same code, one the last week of the train set. This can be very useful both for training models but also for analysis, since you get a validation set created like the test set in the competition, but also have the labels 🙂 And the format is exactly the same! (I also minimized the memory foot print and transformed it to parquet).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2041323,
      "author_name": "piotrekga",
      "author_url": "",
      "post_date": "11/23/2022 20:25:05",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a> . I've checked it before and based on <a href=\"https://www.kaggle.com/code/piotrekga/new-aids-in-test-set\" target=\"_blank\">this</a> notebook all the aids from the test set are present in the train set . The notebooks itself refers to this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991#2029433\" target=\"_blank\">discussion</a> where we talked about code to create train/test split fro the official repo. Am I missing something?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2041448,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/24/2022 00:29:14",
          "content": "<p>Yes, that is a great point, <a href=\"https://www.kaggle.com/piotrekga\" target=\"_blank\">@piotrekga</a>! Also, the counts of unique <code>aids</code> that <a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a> reports seem to be on the lower side then what is in the full data.</p>\n<p><a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a>, if that might be of help, I put together a dataset that I made <a href=\"https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint\" target=\"_blank\">available here</a> that a bunch of us has been using, there is a good degree of confidence that it has been processed correctly 🙂 Maybe please give it a go and see what results you get?</p>\n<p>Also, I absolutely love the visualization you created in your OP 🙂 is that your handwriting? it would be so cool if it was in the era where people are losing their ability to write by hand due to using keyboards all the time 😄 (at least I know this has been happening to me 😉) very neat visualization either way!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2042328,
          "author_name": "andradaolteanu",
          "author_url": "",
          "post_date": "11/24/2022 16:08:00",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> ! Thank you lots for replying.</p>\n<p>So just by looking at <code>aid</code> in the <code>train</code> and <code>test</code> <a href=\"https://www.kaggle.com/datasets/konradb/otto-dataset-in-dataframe\" target=\"_blank\">datasets</a>, this is what I get:<br>\n<img src=\"https://i.imgur.com/rO4G0vn.png\"></p>\n<p>But I see indeed in your dataset you have a great deal more data. What is the difference between your dataset and Konrad's dataset? Because I have been using that one so far and now I am lost completely 😂</p>\n<p>PS: yup, I usually draw when I have a bunch of information, it helps me understand better. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2042404,
          "author_name": "piotrekga",
          "author_url": "",
          "post_date": "11/24/2022 16:59:06",
          "content": "<p>Konrad said he will update his dataset <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368156#2041888\" target=\"_blank\">https://www.kaggle.com/competitions/otto-recommender-system/discussion/368156#2041888</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2042427,
          "author_name": "andradaolteanu",
          "author_url": "",
          "post_date": "11/24/2022 17:11:51",
          "content": "<p>Shoot, missed that one 😢. Sorry for the confussion you guys, guess I'm redoing the whole thing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2042714,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/24/2022 22:56:32",
          "content": "<p>Hey, no worries at all! 🙂 Sorry you have to go through all this trouble, I hate redoing work 😄</p>\n<p>I created a whole \"ecosystem\" of these datasets, in case they might be useful to you 🙂</p>\n<p><a href=\"https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint\" target=\"_blank\">Here</a> is the main one -- essentially memory-optimized data from the competition in a parquet format.</p>\n<p>But <a href=\"https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation\" target=\"_blank\">here</a> is a particularly fun one. I created it using scripts from organizers repo! Essentially, you get a validation set, created like in the competition, with exactly the same code, one the last week of the train set. This can be very useful both for training models but also for analysis, since you get a validation set created like the test set in the competition, but also have the labels 🙂 And the format is exactly the same! (I also minimized the memory foot print and transformed it to parquet).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2041539,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "11/24/2022 03:25:58",
      "content": "<p>I recommend using the original competition dataset to calculate nunique and intersections. According the hosts <a href=\"https://github.com/otto-de/recsys-dataset\" target=\"_blank\">here</a> there are 1.855.603 unique aid just in trainset. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2041250": "While further looking at the data I noticed that there are:\n* 951,294 unique products in the `train` set\n* 259,445 unique products in the `test` set, out of which:\n * 211,446 products (81%) can be found in the `train` set (so they repeat there at least once)\n * 47,999 products (~19%) that **cannot** be found in the `train` set\n\nThis means that these ~19% products will *never* be seen during training.\n\nScary 😂 So there might be some products which we'll have to predict that aren't even in the train set to begin with?\n\n<img src=\"https://i.imgur.com/1j1dcxI.jpg\">\n\nPS: I am using [this dataset](https://www.kaggle.com/datasets/konradb/otto-dataset-in-dataframe).",
    "2041323": "Hi @andradaolteanu . I've checked it before and based on [this](https://www.kaggle.com/code/piotrekga/new-aids-in-test-set) notebook all the aids from the test set are present in the train set . The notebooks itself refers to this [discussion](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991#2029433) where we talked about code to create train/test split fro the official repo. Am I missing something?",
    "2041448": "Yes, that is a great point, @piotrekga! Also, the counts of unique `aids` that @andradaolteanu reports seem to be on the lower side then what is in the full data.\n\n@andradaolteanu, if that might be of help, I put together a dataset that I made [available here](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint) that a bunch of us has been using, there is a good degree of confidence that it has been processed correctly 🙂 Maybe please give it a go and see what results you get?\n\nAlso, I absolutely love the visualization you created in your OP 🙂 is that your handwriting? it would be so cool if it was in the era where people are losing their ability to write by hand due to using keyboards all the time 😄 (at least I know this has been happening to me 😉) very neat visualization either way!",
    "2041539": "I recommend using the original competition dataset to calculate nunique and intersections. According the hosts [here](https://github.com/otto-de/recsys-dataset) there are 1.855.603 unique aid just in trainset.",
    "2042328": "Hi @radek1 ! Thank you lots for replying.\n\nSo just by looking at `aid` in the `train` and `test` [datasets](https://www.kaggle.com/datasets/konradb/otto-dataset-in-dataframe), this is what I get:\n<img src=\"https://i.imgur.com/rO4G0vn.png\">\n\nBut I see indeed in your dataset you have a great deal more data. What is the difference between your dataset and Konrad's dataset? Because I have been using that one so far and now I am lost completely 😂\n\nPS: yup, I usually draw when I have a bunch of information, it helps me understand better.",
    "2042404": "Konrad said he will update his dataset https://www.kaggle.com/competitions/otto-recommender-system/discussion/368156#2041888",
    "2042427": "Shoot, missed that one 😢. Sorry for the confussion you guys, guess I'm redoing the whole thing.",
    "2042714": "Hey, no worries at all! 🙂 Sorry you have to go through all this trouble, I hate redoing work 😄\n\nI created a whole \"ecosystem\" of these datasets, in case they might be useful to you 🙂\n\n[Here](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint) is the main one -- essentially memory-optimized data from the competition in a parquet format.\n\nBut [here](https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation) is a particularly fun one. I created it using scripts from organizers repo! Essentially, you get a validation set, created like in the competition, with exactly the same code, one the last week of the train set. This can be very useful both for training models but also for analysis, since you get a validation set created like the test set in the competition, but also have the labels 🙂 And the format is exactly the same! (I also minimized the memory foot print and transformed it to parquet)."
  },
  "source": "meta"
}