{
  "id": 372289,
  "title": "Test data generation",
  "url": "/competitions/otto-recommender-system/discussion/372289",
  "author_name": "",
  "post_date": "2022-12-15T08:58:13.997446500Z",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I would like to know how the test data is generated.<br>\n For a session, say cut at a point in time, then join followed by one add to cart, one purchase, but multiple clicks, is the test targeting the most recent click?</p>",
  "messages": [
    {
      "id": "2065981",
      "postDate": "12/15/2022 08:58:13",
      "content": "<p>I would like to know how the test data is generated.<br>\n For a session, say cut at a point in time, then join followed by one add to cart, one purchase, but multiple clicks, is the test targeting the most recent click?</p>",
      "rawMarkdown": "I would like to know how the test data is generated.\n For a session, say cut at a point in time, then join followed by one add to cart, one purchase, but multiple clicks, is the test targeting the most recent click?",
      "votes": null
    },
    {
      "id": "2066256",
      "postDate": "12/15/2022 14:15:52",
      "content": "<p>Here is the GitHub code to make test data <a href=\"https://github.com/otto-de/recsys-dataset/blob/5772b8a85fb49c1f91d54b615641b77ac07061b7/src/testset.py#L25\" target=\"_blank\">here</a>. Specifically the data from week 5 is split as follows:</p>\n<pre><code>split_idx = random.randint(1, len(test_events))\n</code></pre>\n<p>The first half (of the random split) of each week 5 user is put into Kaggle test, and the second half of each week 5 user is put into Kaggle leaderboard data. (Weeks 1-4 are Kaggle train).</p>",
      "rawMarkdown": "Here is the GitHub code to make test data [here][1]. Specifically the data from week 5 is split as follows:\n\n    split_idx = random.randint(1, len(test_events))\n\nThe first half (of the random split) of each week 5 user is put into Kaggle test, and the second half of each week 5 user is put into Kaggle leaderboard data. (Weeks 1-4 are Kaggle train).\n\n[1]: https://github.com/otto-de/recsys-dataset/blob/5772b8a85fb49c1f91d54b615641b77ac07061b7/src/testset.py#L25",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2066256,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "12/15/2022 14:15:52",
      "content": "<p>Here is the GitHub code to make test data <a href=\"https://github.com/otto-de/recsys-dataset/blob/5772b8a85fb49c1f91d54b615641b77ac07061b7/src/testset.py#L25\" target=\"_blank\">here</a>. Specifically the data from week 5 is split as follows:</p>\n<pre><code>split_idx = random.randint(1, len(test_events))\n</code></pre>\n<p>The first half (of the random split) of each week 5 user is put into Kaggle test, and the second half of each week 5 user is put into Kaggle leaderboard data. (Weeks 1-4 are Kaggle train).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2065981": "I would like to know how the test data is generated.\n For a session, say cut at a point in time, then join followed by one add to cart, one purchase, but multiple clicks, is the test targeting the most recent click?",
    "2066256": "Here is the GitHub code to make test data [here][1]. Specifically the data from week 5 is split as follows:\n\n    split_idx = random.randint(1, len(test_events))\n\nThe first half (of the random split) of each week 5 user is put into Kaggle test, and the second half of each week 5 user is put into Kaggle leaderboard data. (Weeks 1-4 are Kaggle train).\n\n[1]: https://github.com/otto-de/recsys-dataset/blob/5772b8a85fb49c1f91d54b615641b77ac07061b7/src/testset.py#L25"
  },
  "source": "meta"
}