{
  "id": 368156,
  "title": "Unseen Aids in Test Set",
  "url": "/competitions/otto-recommender-system/discussion/368156",
  "author_name": "",
  "post_date": "2022-11-23T20:23:49.155308400Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I was checking some basic analysis on the test data and was surprised to see how many items from the test set do not appear on the train set.</p>\n<p>Specifically, I saw that:</p>\n<ul>\n<li>The test clicks have <strong>18.67%</strong> of items not in common.</li>\n<li>The test carts have <strong>38.63%</strong> of items not in common.</li>\n<li>The test orders have <strong>47.05%</strong> of items not in common.</li>\n</ul>\n<p>After knowing this I went through some EDAs in the Code section and haven't found anyone mentioning this. Only <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> mentions something similar <a href=\"https://www.kaggle.com/code/radek1/eda-an-overview-of-the-full-dataset\" target=\"_blank\">here</a>, but his results are different than mine, apparently all items on test set appeared on the train set. I am using <a href=\"https://www.kaggle.com/konradb\" target=\"_blank\">@konradb</a> <a href=\"https://www.kaggle.com/datasets/konradb/otto-dataset-in-dataframe\" target=\"_blank\">dataset</a> which is different (it's in CSV format).</p>\n<p>This represents an obvious challenge since we dont have item metadata and users in the test set did not appear also in train set. Perhaps we should recommend the most popular items in case a user has only interacted with unseen items.</p>\n<p>My notebook is available <a href=\"https://www.kaggle.com/code/alejopaullier/otto-unseen-items-in-test\" target=\"_blank\">here</a>.</p>",
  "messages": [
    {
      "id": "2041320",
      "postDate": "11/23/2022 20:23:49",
      "content": "<p>I was checking some basic analysis on the test data and was surprised to see how many items from the test set do not appear on the train set.</p>\n<p>Specifically, I saw that:</p>\n<ul>\n<li>The test clicks have <strong>18.67%</strong> of items not in common.</li>\n<li>The test carts have <strong>38.63%</strong> of items not in common.</li>\n<li>The test orders have <strong>47.05%</strong> of items not in common.</li>\n</ul>\n<p>After knowing this I went through some EDAs in the Code section and haven't found anyone mentioning this. Only <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> mentions something similar <a href=\"https://www.kaggle.com/code/radek1/eda-an-overview-of-the-full-dataset\" target=\"_blank\">here</a>, but his results are different than mine, apparently all items on test set appeared on the train set. I am using <a href=\"https://www.kaggle.com/konradb\" target=\"_blank\">@konradb</a> <a href=\"https://www.kaggle.com/datasets/konradb/otto-dataset-in-dataframe\" target=\"_blank\">dataset</a> which is different (it's in CSV format).</p>\n<p>This represents an obvious challenge since we dont have item metadata and users in the test set did not appear also in train set. Perhaps we should recommend the most popular items in case a user has only interacted with unseen items.</p>\n<p>My notebook is available <a href=\"https://www.kaggle.com/code/alejopaullier/otto-unseen-items-in-test\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "I was checking some basic analysis on the test data and was surprised to see how many items from the test set do not appear on the train set.\n\nSpecifically, I saw that:\n- The test clicks have **18.67%** of items not in common.\n- The test carts have **38.63%** of items not in common.\n- The test orders have **47.05%** of items not in common.\n\nAfter knowing this I went through some EDAs in the Code section and haven't found anyone mentioning this. Only @radek1 mentions something similar [here](https://www.kaggle.com/code/radek1/eda-an-overview-of-the-full-dataset), but his results are different than mine, apparently all items on test set appeared on the train set. I am using @konradb [dataset](https://www.kaggle.com/datasets/konradb/otto-dataset-in-dataframe) which is different (it's in CSV format).\n\nThis represents an obvious challenge since we dont have item metadata and users in the test set did not appear also in train set. Perhaps we should recommend the most popular items in case a user has only interacted with unseen items.\n\nMy notebook is available [here](https://www.kaggle.com/code/alejopaullier/otto-unseen-items-in-test).",
      "votes": null
    },
    {
      "id": "2041329",
      "postDate": "11/23/2022 20:42:52",
      "content": "<p>This is strange, but it looks like <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> data parquets contains more data… The <code>train</code> dataset has 216716096 rows (194720954 clicks) and <code>test</code> is 6928123 (6292632 clicks)… I checked it with my local parquet files. In addition we had a discussion with organisers about script for train/test split <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991#2029433\" target=\"_blank\">here</a>. As I understand that all the aids in test set should be present in the train set.</p>\n<p>I saw a similar topic <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368149\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "This is strange, but it looks like @radek1 data parquets contains more data... The `train` dataset has 216716096 rows (194720954 clicks) and `test` is 6928123 (6292632 clicks)... I checked it with my local parquet files. In addition we had a discussion with organisers about script for train/test split [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991#2029433). As I understand that all the aids in test set should be present in the train set.\n\nI saw a similar topic [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368149).",
      "votes": null
    },
    {
      "id": "2041360",
      "postDate": "11/23/2022 21:20:58",
      "content": "<p>Thanks Piotr I hadn't realized that, maybe Konrad's dataset is outdated. <a href=\"https://www.kaggle.com/konradb\" target=\"_blank\">@konradb</a> is this possible?</p>",
      "rawMarkdown": "Thanks Piotr I hadn't realized that, maybe Konrad's dataset is outdated. @konradb is this possible?",
      "votes": null
    },
    {
      "id": "2041878",
      "postDate": "11/24/2022 09:01:02",
      "content": "<p>I don't know where the difference comes from. I recommend you to create your own dataset based on original competition data. You can find the code I used <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "I don't know where the difference comes from. I recommend you to create your own dataset based on original competition data. You can find the code I used [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)",
      "votes": null
    },
    {
      "id": "2041888",
      "postDate": "11/24/2022 09:21:50",
      "content": "<p>Yes - I need to update this one. </p>",
      "rawMarkdown": "Yes - I need to update this one.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2041329,
      "author_name": "piotrekga",
      "author_url": "",
      "post_date": "11/23/2022 20:42:52",
      "content": "<p>This is strange, but it looks like <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> data parquets contains more data… The <code>train</code> dataset has 216716096 rows (194720954 clicks) and <code>test</code> is 6928123 (6292632 clicks)… I checked it with my local parquet files. In addition we had a discussion with organisers about script for train/test split <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991#2029433\" target=\"_blank\">here</a>. As I understand that all the aids in test set should be present in the train set.</p>\n<p>I saw a similar topic <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368149\" target=\"_blank\">here</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2041360,
          "author_name": "alejopaullier",
          "author_url": "",
          "post_date": "11/23/2022 21:20:58",
          "content": "<p>Thanks Piotr I hadn't realized that, maybe Konrad's dataset is outdated. <a href=\"https://www.kaggle.com/konradb\" target=\"_blank\">@konradb</a> is this possible?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2041878,
          "author_name": "piotrekga",
          "author_url": "",
          "post_date": "11/24/2022 09:01:02",
          "content": "<p>I don't know where the difference comes from. I recommend you to create your own dataset based on original competition data. You can find the code I used <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">here</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2041888,
          "author_name": "konradb",
          "author_url": "",
          "post_date": "11/24/2022 09:21:50",
          "content": "<p>Yes - I need to update this one. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2041320": "I was checking some basic analysis on the test data and was surprised to see how many items from the test set do not appear on the train set.\n\nSpecifically, I saw that:\n- The test clicks have **18.67%** of items not in common.\n- The test carts have **38.63%** of items not in common.\n- The test orders have **47.05%** of items not in common.\n\nAfter knowing this I went through some EDAs in the Code section and haven't found anyone mentioning this. Only @radek1 mentions something similar [here](https://www.kaggle.com/code/radek1/eda-an-overview-of-the-full-dataset), but his results are different than mine, apparently all items on test set appeared on the train set. I am using @konradb [dataset](https://www.kaggle.com/datasets/konradb/otto-dataset-in-dataframe) which is different (it's in CSV format).\n\nThis represents an obvious challenge since we dont have item metadata and users in the test set did not appear also in train set. Perhaps we should recommend the most popular items in case a user has only interacted with unseen items.\n\nMy notebook is available [here](https://www.kaggle.com/code/alejopaullier/otto-unseen-items-in-test).",
    "2041329": "This is strange, but it looks like @radek1 data parquets contains more data... The `train` dataset has 216716096 rows (194720954 clicks) and `test` is 6928123 (6292632 clicks)... I checked it with my local parquet files. In addition we had a discussion with organisers about script for train/test split [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991#2029433). As I understand that all the aids in test set should be present in the train set.\n\nI saw a similar topic [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368149).",
    "2041360": "Thanks Piotr I hadn't realized that, maybe Konrad's dataset is outdated. @konradb is this possible?",
    "2041878": "I don't know where the difference comes from. I recommend you to create your own dataset based on original competition data. You can find the code I used [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)",
    "2041888": "Yes - I need to update this one."
  },
  "source": "meta"
}