{
  "id": 364991,
  "title": "local validation tracks public LB perfecty -- here is the setup",
  "url": "/competitions/otto-recommender-system/discussion/364991",
  "author_name": "Radek Osmulski",
  "post_date": "2022-11-09T10:40:43.980000",
  "votes": 208,
  "comment_count": 55,
  "views": 0,
  "content": "<p>Hey!</p>\n<p>I ran a couple of experiments and local validation tracks public LB perfectly! Here is what the situation looks like:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F457ff5950918744c841a83ac4c0ece1a%2Fdownload.png?generation=1667989795036578&amp;alt=media\" alt=\"\"></p>\n<p>The results are slightly poorer in local validation, but that is not surprising -- we are running on one week less of train data!</p>\n<p>Here's the setup:</p>\n<ul>\n<li>I generated the split of the train set using the code in <a href=\"https://github.com/otto-de/recsys-dataset\" target=\"_blank\">organizer's repo</a></li>\n<li>I am using the last 7 days of train set for validation (similarly to the test set)</li>\n<li>If you don't want to run the code on the data yourself, I <a href=\"https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation\" target=\"_blank\">uploaded the prepared data to Kaggle</a></li>\n</ul>\n<p>There has been a lot of conversation on the forums on the metric, I think we now have a good understanding of how it works 🙂 (one can also always confirm one's understanding by checking out the organizer's repo!). Interesting discussions on the forums <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364064\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364530\" target=\"_blank\">here</a>.</p>\n<p>Here is the code I am using for calculating the metric in pandas, starting from a valid submission DataFrame, in case this might be of help:</p>\n<pre><code>    submission['session'] = submission.session_type.apply(lambda x: int(x.split('_')[0]))\n    submission['type'] = submission.session_type.apply(lambda x: x.split('_')[1])\n    submission.labels = submission.labels.apply(lambda x: [int(i) for i in x.split(' ')[:20]])\n\n    test_labels = pd.read_parquet('out_processed/test_labels.parquet')\n    test_labels = test_labels.merge(submission, how='left', on=['session', 'type'])\n    test_labels['hits'] = test_labels.apply(lambda df: len(set(df.ground_truth).intersection(set(df.labels))), axis=1)\n    test_labels['gt_count'] = test_labels.ground_truth.str.len().clip(0,20)\n\n    recall_per_type = test_labels.groupby(['type'])['hits'].sum() / test_labels.groupby(['type'])['gt_count'].sum() \n\n    score = (recall_per_type * pd.Series({'clicks': 0.10, 'carts': 0.30, 'orders': 0.60})).sum()\n</code></pre>\n<p>Competitions where local validation tracks public LB are much more fun, IMHO 🙂</p>\n<p>Wishing you all the best in the challenge!<br>\nRadek</p>\n<p><strong>IMPORTANT CAVEAT:</strong> Please note that in my dataset, in order to minimize its size without losing information (unless you care about milliseconds, but I assume seconds should be enough 🙂), I divide the <code>ts</code> column by 1000. Notebooks on Kaggle by other people use unmodified data, so if you would want to run this validation with that code, you need to either make changes to the code or multiply the <code>ts</code> column by 1000.</p>\n<p><strong>A couple of related resources you might find useful:</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": 2022844,
      "postDate": "2022-11-09T10:40:43.980Z",
      "content": "<p>Hey!</p>\n<p>I ran a couple of experiments and local validation tracks public LB perfectly! Here is what the situation looks like:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F457ff5950918744c841a83ac4c0ece1a%2Fdownload.png?generation=1667989795036578&amp;alt=media\" alt=\"\"></p>\n<p>The results are slightly poorer in local validation, but that is not surprising -- we are running on one week less of train data!</p>\n<p>Here's the setup:</p>\n<ul>\n<li>I generated the split of the train set using the code in <a href=\"https://github.com/otto-de/recsys-dataset\" target=\"_blank\">organizer's repo</a></li>\n<li>I am using the last 7 days of train set for validation (similarly to the test set)</li>\n<li>If you don't want to run the code on the data yourself, I <a href=\"https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation\" target=\"_blank\">uploaded the prepared data to Kaggle</a></li>\n</ul>\n<p>There has been a lot of conversation on the forums on the metric, I think we now have a good understanding of how it works 🙂 (one can also always confirm one's understanding by checking out the organizer's repo!). Interesting discussions on the forums <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364064\" target=\"_blank\">here</a> and <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364530\" target=\"_blank\">here</a>.</p>\n<p>Here is the code I am using for calculating the metric in pandas, starting from a valid submission DataFrame, in case this might be of help:</p>\n<pre><code>    submission['session'] = submission.session_type.apply(lambda x: int(x.split('_')[0]))\n    submission['type'] = submission.session_type.apply(lambda x: x.split('_')[1])\n    submission.labels = submission.labels.apply(lambda x: [int(i) for i in x.split(' ')[:20]])\n\n    test_labels = pd.read_parquet('out_processed/test_labels.parquet')\n    test_labels = test_labels.merge(submission, how='left', on=['session', 'type'])\n    test_labels['hits'] = test_labels.apply(lambda df: len(set(df.ground_truth).intersection(set(df.labels))), axis=1)\n    test_labels['gt_count'] = test_labels.ground_truth.str.len().clip(0,20)\n\n    recall_per_type = test_labels.groupby(['type'])['hits'].sum() / test_labels.groupby(['type'])['gt_count'].sum() \n\n    score = (recall_per_type * pd.Series({'clicks': 0.10, 'carts': 0.30, 'orders': 0.60})).sum()\n</code></pre>\n<p>Competitions where local validation tracks public LB are much more fun, IMHO 🙂</p>\n<p>Wishing you all the best in the challenge!<br>\nRadek</p>\n<p><strong>IMPORTANT CAVEAT:</strong> Please note that in my dataset, in order to minimize its size without losing information (unless you care about milliseconds, but I assume seconds should be enough 🙂), I divide the <code>ts</code> column by 1000. Notebooks on Kaggle by other people use unmodified data, so if you would want to run this validation with that code, you need to either make changes to the code or multiply the <code>ts</code> column by 1000.</p>\n<p><strong>A couple of related resources you might find useful:</strong></p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "Hey!\n\nI ran a couple of experiments and local validation tracks public LB perfectly! Here is what the situation looks like:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F457ff5950918744c841a83ac4c0ece1a%2Fdownload.png?generation=1667989795036578&alt=media)\n\nThe results are slightly poorer in local validation, but that is not surprising -- we are running on one week less of train data!\n\nHere's the setup:\n* I generated the split of the train set using the code in [organizer's repo](https://github.com/otto-de/recsys-dataset)\n* I am using the last 7 days of train set for validation (similarly to the test set)\n* If you don't want to run the code on the data yourself, I [uploaded the prepared data to Kaggle](https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation)\n\nThere has been a lot of conversation on the forums on the metric, I think we now have a good understanding of how it works 🙂 (one can also always confirm one's understanding by checking out the organizer's repo!). Interesting discussions on the forums [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364064) and [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364530).\n\nHere is the code I am using for calculating the metric in pandas, starting from a valid submission DataFrame, in case this might be of help:\n\n```\n    submission['session'] = submission.session_type.apply(lambda x: int(x.split('_')[0]))\n    submission['type'] = submission.session_type.apply(lambda x: x.split('_')[1])\n    submission.labels = submission.labels.apply(lambda x: [int(i) for i in x.split(' ')[:20]])\n\n    test_labels = pd.read_parquet('out_processed/test_labels.parquet')\n    test_labels = test_labels.merge(submission, how='left', on=['session', 'type'])\n    test_labels['hits'] = test_labels.apply(lambda df: len(set(df.ground_truth).intersection(set(df.labels))), axis=1)\n    test_labels['gt_count'] = test_labels.ground_truth.str.len().clip(0,20)\n\n    recall_per_type = test_labels.groupby(['type'])['hits'].sum() / test_labels.groupby(['type'])['gt_count'].sum() \n\n    score = (recall_per_type * pd.Series({'clicks': 0.10, 'carts': 0.30, 'orders': 0.60})).sum()\n```\n\nCompetitions where local validation tracks public LB are much more fun, IMHO 🙂\n\nWishing you all the best in the challenge!\nRadek\n\n**IMPORTANT CAVEAT:** Please note that in my dataset, in order to minimize its size without losing information (unless you care about milliseconds, but I assume seconds should be enough 🙂), I divide the `ts` column by 1000. Notebooks on Kaggle by other people use unmodified data, so if you would want to run this validation with that code, you need to either make changes to the code or multiply the `ts` column by 1000.\n\n**A couple of related resources you might find useful:**\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)\n\n",
      "votes": 208
    },
    {
      "id": 2024516,
      "postDate": "2022-11-10T15:25:34.453Z",
      "content": "<p>FYI, I published a version of Radek's validation data as multiple parquets and restored columns <code>ts</code> and <code>type</code> back to their original. Radek's dataset is more memory and disk efficient, but I published this alternative dataset for anyone who has a pipeline using a Kaggle dataset with multiple parquets and original <code>ts</code> and <code>type</code> columns. </p>\n<p>Now you can run your exact pipeline with validation data and compute CV score without making any changes to your pipeline. The Kaggle dataset is <a href=\"https://www.kaggle.com/datasets/cdeotte/otto-validation\" target=\"_blank\">here</a>. I plan to publish a notebook which computes validation score for my notebook soon.</p>",
      "rawMarkdown": "FYI, I published a version of Radek's validation data as multiple parquets and restored columns `ts` and `type` back to their original. Radek's dataset is more memory and disk efficient, but I published this alternative dataset for anyone who has a pipeline using a Kaggle dataset with multiple parquets and original `ts` and `type` columns. \n\nNow you can run your exact pipeline with validation data and compute CV score without making any changes to your pipeline. The Kaggle dataset is [here][1]. I plan to publish a notebook which computes validation score for my notebook soon.\n\n[1]: https://www.kaggle.com/datasets/cdeotte/otto-validation",
      "votes": 9,
      "replies": [
        {
          "id": 2024873,
          "postDate": "2022-11-10T20:03:44.977Z",
          "content": "<p>UPDATE: I published a train notebook that achieves LB 0.573 and corresponding validation notebook \"CV\" 0.563 using Radek's validation data (in my Kaggle dataset).</p>",
          "rawMarkdown": "UPDATE: I published a train notebook that achieves LB 0.573 and corresponding validation notebook \"CV\" 0.563 using Radek's validation data (in my Kaggle dataset).",
          "votes": 5
        },
        {
          "id": 2025144,
          "postDate": "2022-11-11T02:37:13.650Z",
          "content": "<p>After publishing the notebook, i made two changes that boosted CV <code>+0.003</code> and after submission, the LB boosted <code>+0.003</code>. The CV LB correlate very well.</p>",
          "rawMarkdown": "After publishing the notebook, i made two changes that boosted CV `+0.003` and after submission, the LB boosted `+0.003`. The CV LB correlate very well.",
          "votes": 7
        }
      ]
    },
    {
      "id": 2023756,
      "postDate": "2022-11-10T02:19:52.620Z",
      "content": "<p>Your CV scheme (data and metric code) has perfect correlation for me too. Thank you for publishing your setup. The first two data points are my public notebook version 4 and 9 <a href=\"https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost\" target=\"_blank\">here</a> (with LB 0.556 and LB 0.562 respectively). Then next data point is Pietro's notebook <a href=\"https://www.kaggle.com/code/pietromaldini1/multiple-clicks-vs-latest-items\" target=\"_blank\">here</a> (when it was/is LB 0.565 version 1). The next data point is Ingvaras' notebook <a href=\"https://www.kaggle.com/code/ingvarasgalinskas/item-type-vs-multiple-clicks-vs-latest-items\" target=\"_blank\">here</a> (when it was/is LB 0.570 version 3). We see that each notebook improves both CV and LB. The last data point is my current LB score using some unpublished tricks.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/cv_lb2.png\" alt=\"\"></p>",
      "rawMarkdown": "Your CV scheme (data and metric code) has perfect correlation for me too. Thank you for publishing your setup. The first two data points are my public notebook version 4 and 9 [here][2] (with LB 0.556 and LB 0.562 respectively). Then next data point is Pietro's notebook [here][1] (when it was/is LB 0.565 version 1). The next data point is Ingvaras' notebook [here][3] (when it was/is LB 0.570 version 3). We see that each notebook improves both CV and LB. The last data point is my current LB score using some unpublished tricks.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/cv_lb2.png)\n\n[1]: https://www.kaggle.com/code/pietromaldini1/multiple-clicks-vs-latest-items\n[2]: https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost\n[3]: https://www.kaggle.com/code/ingvarasgalinskas/item-type-vs-multiple-clicks-vs-latest-items",
      "votes": 8,
      "replies": [
        {
          "id": 2023758,
          "postDate": "2022-11-10T02:21:12.023Z",
          "content": "<p>wonderful data, thank you for sharing this <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! 🙂  </p>",
          "rawMarkdown": "wonderful data, thank you for sharing this @cdeotte! 🙂  ",
          "votes": 2
        },
        {
          "id": 2075047,
          "postDate": "2022-12-25T01:16:59.220Z",
          "content": "<p>Hi,Chris.I used your version 9 notebook to calculate the local CV. I suspect something went wrong with my idea of calculating local CV scores. In the notebook, I generated candidates for val A's validation set and then calculated the local CV score, I only made the following modifications, but the CV score looks so outrageous. What went wrong with me? I am looking forward to hearing from you.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11356689%2Ffc4f117d4f91aae5da0b8f807e71ad6a%2F1.png?generation=1671930944213119&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11356689%2Fa1e7af94b592aba55f2e8df7ad1383b5%2F2.png?generation=1671930967340151&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Hi,Chris.I used your version 9 notebook to calculate the local CV. I suspect something went wrong with my idea of calculating local CV scores. In the notebook, I generated candidates for val A's validation set and then calculated the local CV score, I only made the following modifications, but the CV score looks so outrageous. What went wrong with me? I am looking forward to hearing from you.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11356689%2Ffc4f117d4f91aae5da0b8f807e71ad6a%2F1.png?generation=1671930944213119&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11356689%2Fa1e7af94b592aba55f2e8df7ad1383b5%2F2.png?generation=1671930967340151&alt=media)\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 2053831,
      "postDate": "2022-12-03T16:40:30.753Z",
      "content": "<p>Thanks for sharing, this saves a lot of tine!</p>",
      "rawMarkdown": "Thanks for sharing, this saves a lot of tine!",
      "votes": 5,
      "replies": [
        {
          "id": 2054296,
          "postDate": "2022-12-04T02:24:31.563Z",
          "content": "<p>Thank you, JFP 🙂 Really glad to be of help</p>",
          "rawMarkdown": "Thank you, JFP 🙂 Really glad to be of help",
          "votes": 1
        }
      ]
    },
    {
      "id": 2023511,
      "postDate": "2022-11-09T19:43:42.153Z",
      "content": "<p>This is very helpful thank you. I used your files and code to compute the CV score for my notebook version 4 with LB 0.556. The computed CV score is 0.474 so I think there is a bug in my code somewhere. It appears that you have a smaller gap (between LB and CV) when using your CV files. I will review my code and try again…</p>",
      "rawMarkdown": "This is very helpful thank you. I used your files and code to compute the CV score for my notebook version 4 with LB 0.556. The computed CV score is 0.474 so I think there is a bug in my code somewhere. It appears that you have a smaller gap (between LB and CV) when using your CV files. I will review my code and try again...",
      "votes": 3,
      "replies": [
        {
          "id": 2023514,
          "postDate": "2022-11-09T19:51:17.187Z",
          "content": "<p>I may have found the problem. Your parquet files have <code>ts</code> column with <code>original_ts / 1000</code>. So my notebook's pipeline which adds weight based on <code>ts</code> doesn't work correctly. I will update and check CV score.</p>",
          "rawMarkdown": "I may have found the problem. Your parquet files have `ts` column with `original_ts / 1000`. So my notebook's pipeline which adds weight based on `ts` doesn't work correctly. I will update and check CV score.",
          "votes": 4
        },
        {
          "id": 2023515,
          "postDate": "2022-11-09T19:56:11.370Z",
          "content": "<p>You should add a note to your discussion here. Anyone who uses code based on original <code>ts</code> like my Chris' notebook, or Vladimir's notebook, or Pietro's notebook, or Ingvaras's notebook will run into this problem. For example the following line of code:</p>\n<pre><code>df.query('abs(ts_x - ts_y) &lt; 24 * 60 * 60 * 1000 and aid_x != aid_y')\n</code></pre>\n<p>will waste many hours running code with your pipeline before discovering the error. Also type conversion will fail</p>\n<pre><code>.astype({\"ts\": \"datetime64[ms]\"})\n</code></pre>",
          "rawMarkdown": "You should add a note to your discussion here. Anyone who uses code based on original `ts` like my Chris' notebook, or Vladimir's notebook, or Pietro's notebook, or Ingvaras's notebook will run into this problem. For example the following line of code:\n\n    df.query('abs(ts_x - ts_y) < 24 * 60 * 60 * 1000 and aid_x != aid_y')\n\nwill waste many hours running code with your pipeline before discovering the error. Also type conversion will fail\n\n    .astype({\"ts\": \"datetime64[ms]\"})",
          "votes": 2
        },
        {
          "id": 2023559,
          "postDate": "2022-11-09T21:11:00.183Z",
          "content": "<p>Yes Chris, that is a great point! Apologies for causing all this trouble on your end. You are absolutely right.</p>\n<p>I included this bit of information in the other dataset that I published, in the description, but didn't do so here. Very sorry bout that. Will correct this in a second and add a note.</p>\n<p>My apologies for all the trouble.</p>",
          "rawMarkdown": "Yes Chris, that is a great point! Apologies for causing all this trouble on your end. You are absolutely right.\n\nI included this bit of information in the other dataset that I published, in the description, but didn't do so here. Very sorry bout that. Will correct this in a second and add a note.\n\nMy apologies for all the trouble.",
          "votes": 1
        },
        {
          "id": 2023578,
          "postDate": "2022-11-09T21:35:50.963Z",
          "content": "<p>Great. Looks good now. My computed CV is 0.548 for version 4 of my notebook with LB 0.556. That's pretty close. I will compute more CV scores for different versions and check the correlation. I will post results here.</p>\n<p>Then i will begin boosting my CV LB with improved logic! and/or trained models!</p>",
          "rawMarkdown": "Great. Looks good now. My computed CV is 0.548 for version 4 of my notebook with LB 0.556. That's pretty close. I will compute more CV scores for different versions and check the correlation. I will post results here.\n\nThen i will begin boosting my CV LB with improved logic! and/or trained models!",
          "votes": 3
        },
        {
          "id": 2023731,
          "postDate": "2022-11-10T01:43:16.273Z",
          "content": "<p>It's really cool how you are sharing openly how you plan to approach this competition, I think that is great learning material! 🙂</p>\n<p>My plan is very similar, hoping to share some stuff on Matrix Factorization maybe with using some of the Merlin stuff, have some really cool ideas, but the only problem is finding the time 😄</p>\n<p>Great there are that many weeks still left in this competition! 🙂</p>",
          "rawMarkdown": "It's really cool how you are sharing openly how you plan to approach this competition, I think that is great learning material! 🙂\n\nMy plan is very similar, hoping to share some stuff on Matrix Factorization maybe with using some of the Merlin stuff, have some really cool ideas, but the only problem is finding the time 😄\n\nGreat there are that many weeks still left in this competition! 🙂",
          "votes": 2
        }
      ]
    },
    {
      "id": 2029433,
      "postDate": "2022-11-14T16:31:57.847Z",
      "content": "<p>Thanks for the dataset. I've spotted a problem with validation dataset. There are \"new\" AIDs (18785) in the validation test set comparing to validations train set). For the original train and test set this is not the case.</p>\n<p>I've added notebook with the calculation to the dataset: <a href=\"https://www.kaggle.com/code/piotrekga/new-aids-in-test-set\" target=\"_blank\">https://www.kaggle.com/code/piotrekga/new-aids-in-test-set</a></p>\n<p>I guess there was more logic in train/test split. <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a>, could you please give us information on what to do with sessions and ground truth with new AIDs?</p>",
      "rawMarkdown": "Thanks for the dataset. I've spotted a problem with validation dataset. There are \"new\" AIDs (18785) in the validation test set comparing to validations train set). For the original train and test set this is not the case.\n\nI've added notebook with the calculation to the dataset: https://www.kaggle.com/code/piotrekga/new-aids-in-test-set\n\nI guess there was more logic in train/test split. @pnormann, could you please give us information on what to do with sessions and ground truth with new AIDs?",
      "votes": 4,
      "replies": [
        {
          "id": 2029810,
          "postDate": "2022-11-15T00:27:10.533Z",
          "content": "<p>ha, interesting observation <a href=\"https://www.kaggle.com/piotrekga\" target=\"_blank\">@piotrekga</a>!</p>\n<p>I think this might just be coincidental, new AIDs might have gotten added to the system that week 🤔 Or, alternatively, there was some additional logic in pulling down the data.</p>\n<p><a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py\" target=\"_blank\">Here is the logic used by the organizer to create the test set</a>. There is nothing there about filtering sessions based on <code>aids</code> AFAICT at quick glance.</p>",
          "rawMarkdown": "ha, interesting observation @piotrekga!\n\nI think this might just be coincidental, new AIDs might have gotten added to the system that week 🤔 Or, alternatively, there was some additional logic in pulling down the data.\n\n[Here is the logic used by the organizer to create the test set](https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py). There is nothing there about filtering sessions based on `aids` AFAICT at quick glance."
        },
        {
          "id": 2029879,
          "postDate": "2022-11-15T02:22:28.200Z",
          "content": "<p>That's a helpful discovery Piotr. Differences between CV and LB are a concern.</p>\n<p>However, I will comment that Radek's validation data is perfectly correlating with LB. Every <code>0.0002+</code> improvement on Radek's CV gives me a boost on LB!</p>",
          "rawMarkdown": "That's a helpful discovery Piotr. Differences between CV and LB are a concern.\n\nHowever, I will comment that Radek's validation data is perfectly correlating with LB. Every `0.0002+` improvement on Radek's CV gives me a boost on LB!",
          "votes": 1
        },
        {
          "id": 2030124,
          "postDate": "2022-11-15T07:59:13.497Z",
          "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> I've rerun the train-test split on <code>train.jsonl</code> before and got exactly the same results. You're right the logic is not included in the code provided by organisers.</p>\n<p>About new AIDs: taking into consideration that these are users behaviours and the dataset has long tails I find it very likely that there were no new AIDs in the original test set and thousands of them in new train-test split. Of course I may be wrong 😅. I really hope organisers will provide us with some additional clarification in that regards.</p>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, you're right that now the CV-LB correlation works. I'd really like to validate my model locally as close as possible to the final validation though.</p>",
          "rawMarkdown": "@radek1 I've rerun the train-test split on `train.jsonl` before and got exactly the same results. You're right the logic is not included in the code provided by organisers.\n\nAbout new AIDs: taking into consideration that these are users behaviours and the dataset has long tails I find it very likely that there were no new AIDs in the original test set and thousands of them in new train-test split. Of course I may be wrong 😅. I really hope organisers will provide us with some additional clarification in that regards.\n\n@cdeotte, you're right that now the CV-LB correlation works. I'd really like to validate my model locally as close as possible to the final validation though.",
          "votes": 2
        },
        {
          "id": 2032183,
          "postDate": "2022-11-16T13:19:34.117Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/piotrekga\" target=\"_blank\">@piotrekga</a>, thanks for pointing this out and sorry for the confusion! This behavior was a bug in our public repo that was inconsistent with how we created the dataset for the competition. We created the actual dataset using a Clojure job, and during the reimplementation in Python, we, unfortunately, forgot to add this preprocessing step. I just pushed an <a href=\"https://github.com/otto-de/recsys-dataset/commit/14eea3938cd9b6200bda9985f2c7ff8bb1de1b18\" target=\"_blank\">update</a> to our GitHub repo which fixes this issue. Although I'm still not entirely happy with the performance of the Python script, it should be enough for you to create your local validation set, which reflects the logic we used for the LB test set. I'll see what I can do to speed things up 😌</p>",
          "rawMarkdown": "Hey @piotrekga, thanks for pointing this out and sorry for the confusion! This behavior was a bug in our public repo that was inconsistent with how we created the dataset for the competition. We created the actual dataset using a Clojure job, and during the reimplementation in Python, we, unfortunately, forgot to add this preprocessing step. I just pushed an [update](https://github.com/otto-de/recsys-dataset/commit/14eea3938cd9b6200bda9985f2c7ff8bb1de1b18) to our GitHub repo which fixes this issue. Although I'm still not entirely happy with the performance of the Python script, it should be enough for you to create your local validation set, which reflects the logic we used for the LB test set. I'll see what I can do to speed things up 😌",
          "votes": 8
        },
        {
          "id": 2034542,
          "postDate": "2022-11-18T08:18:33.893Z",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> for the update. At this point logic is far more important than performance.</p>",
          "rawMarkdown": "Thank you @pnormann for the update. At this point logic is far more important than performance.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2093576,
      "postDate": "2023-01-10T05:35:13.883Z",
      "content": "<p>Sorry <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> what is the id2type and vice versa pkl files contain? What is the purpose of these 2 data files?</p>",
      "rawMarkdown": "Sorry @radek1 what is the id2type and vice versa pkl files contain? What is the purpose of these 2 data files?",
      "votes": 1,
      "replies": [
        {
          "id": 2094694,
          "postDate": "2023-01-11T00:46:19.597Z",
          "content": "<p>hey <a href=\"https://www.kaggle.com/ppujari\" target=\"_blank\">@ppujari</a>! They just contain the mapping of types to their ids IIRC, 0 maps to <code>clicks</code>, 1 to <code>carts</code> and  2 to <code>orders</code>. Probably shouldn't have included it, makes it all a bit confusing.</p>",
          "rawMarkdown": "hey @ppujari! They just contain the mapping of types to their ids IIRC, 0 maps to `clicks`, 1 to `carts` and  2 to `orders`. Probably shouldn't have included it, makes it all a bit confusing."
        }
      ]
    },
    {
      "id": 2076910,
      "postDate": "2022-12-27T02:03:19.867Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, I have found something weird of the validation dataset you provided. </p>\n<p>According to the organizer's script for creating the validation dataset, when we split the 4 weeks train into 3-weeks-train (part-1) and 1-week-train (part-2), the part-2's aids should have all been seen in part-1. However, the <code>test.parquet</code> in the part-2 still have more than 2% aids which are not seen by part-1.</p>\n<p>I have shown this problem in this version of the notebook <a href=\"https://www.kaggle.com/code/danielliao/reading-otto-recsys-organizer-scripts?scriptVersionId=114730145&amp;cellId=9\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Hi @radek1, I have found something weird of the validation dataset you provided. \n\nAccording to the organizer's script for creating the validation dataset, when we split the 4 weeks train into 3-weeks-train (part-1) and 1-week-train (part-2), the part-2's aids should have all been seen in part-1. However, the `test.parquet` in the part-2 still have more than 2% aids which are not seen by part-1.\n\nI have shown this problem in this version of the notebook [here](https://www.kaggle.com/code/danielliao/reading-otto-recsys-organizer-scripts?scriptVersionId=114730145&cellId=9)",
      "votes": 1
    },
    {
      "id": 2049291,
      "postDate": "2022-11-30T03:18:22.520Z",
      "content": "<p>Thank you for sharing!!</p>",
      "rawMarkdown": "Thank you for sharing!!",
      "votes": 1,
      "replies": [
        {
          "id": 2053433,
          "postDate": "2022-12-03T09:21:42.247Z",
          "content": "<p>np <a href=\"https://www.kaggle.com/yuyatasaki\" target=\"_blank\">@yuyatasaki</a>! happy to be of help 🙂</p>",
          "rawMarkdown": "np @yuyatasaki! happy to be of help 🙂"
        }
      ]
    },
    {
      "id": 2025309,
      "postDate": "2022-11-11T06:29:47.053Z",
      "content": "<p>Thank you for sharing.</p>",
      "rawMarkdown": "Thank you for sharing.",
      "votes": 1,
      "replies": [
        {
          "id": 2028983,
          "postDate": "2022-11-14T11:07:04.497Z",
          "content": "<p>np <a href=\"https://www.kaggle.com/zhenzheng0328\" target=\"_blank\">@zhenzheng0328</a>! Glad you found this useful! 🙌</p>",
          "rawMarkdown": "np @zhenzheng0328! Glad you found this useful! 🙌"
        }
      ]
    },
    {
      "id": 2024958,
      "postDate": "2022-11-10T22:35:50.763Z",
      "content": "<p>Awesome Data to work with… Thanks for sharing this…</p>",
      "rawMarkdown": "Awesome Data to work with... Thanks for sharing this...",
      "votes": 1,
      "replies": [
        {
          "id": 2024969,
          "postDate": "2022-11-10T22:50:34.307Z",
          "content": "<p>NP <a href=\"https://www.kaggle.com/saha8631\" target=\"_blank\">@saha8631</a>, glad I could be of help! 🙌</p>",
          "rawMarkdown": "NP @saha8631, glad I could be of help! 🙌",
          "votes": 1
        }
      ]
    },
    {
      "id": 2069469,
      "postDate": "2022-12-19T02:44:29.387Z",
      "content": "<p>Thanks for sharing. I have a question, after you split the first 3 weeks into the train the last week into the test. How did you further split to get the <code>test.parquet</code> and <code>test_labels.parquet</code>?</p>",
      "rawMarkdown": "Thanks for sharing. I have a question, after you split the first 3 weeks into the train the last week into the test. How did you further split to get the `test.parquet` and `test_labels.parquet`?",
      "replies": [
        {
          "id": 2069493,
          "postDate": "2022-12-19T03:10:51.733Z",
          "content": "<p>The process is explained in Otto's GitHub <a href=\"https://github.com/otto-de/recsys-dataset/tree/main/test\" target=\"_blank\">here</a>. These are the steps </p>\n<ul>\n<li>For every user in Kaggle 4 week of train data, we find their earliest event</li>\n<li>For every user with their earliest event in the first 3 weeks, they become the \"new train\" and their events get truncated to the first 3 weeks</li>\n<li>For every user whos first event begins in week 4, they become the \"test users\"</li>\n<li>Next, for each test user, we pick a random split with <code>np.random.uniform(1,len(user_events))</code>. The first random half of user activity becomes <code>test.parquet</code> and the second random half becomes <code>test_labels.parquet</code>.</li>\n<li>Lastly every event involving an item in <code>test_labels.parquet</code> that does not exist in train or test gets removed.</li>\n</ul>",
          "rawMarkdown": "The process is explained in Otto's GitHub [here][1]. These are the steps \n* For every user in Kaggle 4 week of train data, we find their earliest event\n* For every user with their earliest event in the first 3 weeks, they become the \"new train\" and their events get truncated to the first 3 weeks\n* For every user whos first event begins in week 4, they become the \"test users\"\n* Next, for each test user, we pick a random split with `np.random.uniform(1,len(user_events))`. The first random half of user activity becomes `test.parquet` and the second random half becomes `test_labels.parquet`.\n* Lastly every event involving an item in `test_labels.parquet` that does not exist in train or test gets removed.\n\n[1]: https://github.com/otto-de/recsys-dataset/tree/main/test",
          "votes": 12,
          "replies": [
            {
              "id": 2069494,
              "postDate": "2022-12-19T03:18:49.043Z",
              "content": "<p>Thanks for the explanation <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , very clear. You saved my day</p>",
              "rawMarkdown": "Thanks for the explanation @cdeotte , very clear. You saved my day",
              "votes": 1
            },
            {
              "id": 2088453,
              "postDate": "2023-01-06T10:27:22.367Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2090215,
              "postDate": "2023-01-07T05:33:11.877Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 2059672,
      "postDate": "2022-12-09T05:56:54.437Z",
      "content": "<p>Thank you for sharing validation strategy. I may be wrong because I haven't thought about this too much yet, but I would like to ask one question.<br>\nYou adopted the last 7 days of train data for validation set. This means the users are same in train and validation while LB and train is different in users. From your result, domain shift by user doesn't occur, right?</p>",
      "rawMarkdown": "Thank you for sharing validation strategy. I may be wrong because I haven't thought about this too much yet, but I would like to ask one question.\nYou adopted the last 7 days of train data for validation set. This means the users are same in train and validation while LB and train is different in users. From your result, domain shift by user doesn't occur, right?\n",
      "replies": [
        {
          "id": 2059679,
          "postDate": "2022-12-09T06:05:55.640Z",
          "content": "<p>It's a valid concern, but there is no overlap. Any overlap of the sort you mention has been removed.</p>",
          "rawMarkdown": "It's a valid concern, but there is no overlap. Any overlap of the sort you mention has been removed.",
          "votes": 1
        },
        {
          "id": 2059802,
          "postDate": "2022-12-09T08:36:39.703Z",
          "content": "<p>Thank you so much！</p>",
          "rawMarkdown": "Thank you so much！"
        }
      ]
    },
    {
      "id": 2057414,
      "postDate": "2022-12-07T04:52:09.797Z",
      "content": "<p>Thank you for sharing :)</p>",
      "rawMarkdown": "Thank you for sharing :)",
      "replies": [
        {
          "id": 2057621,
          "postDate": "2022-12-07T09:04:02.877Z",
          "content": "<p>Glad to be of help, <a href=\"https://www.kaggle.com/songwonho\" target=\"_blank\">@songwonho</a>! Thanks for your kind comment 🙂 </p>",
          "rawMarkdown": "Glad to be of help, @songwonho! Thanks for your kind comment 🙂 "
        }
      ]
    },
    {
      "id": 2053420,
      "postDate": "2022-12-03T08:47:54.650Z",
      "content": "<p>I was doing sanity check and I didn't understand how ground-truth is created. Can someone verify they are correct or not?</p>\n<p><img src=\"https://i.imgur.com/E8pJ5ZK.png\" alt=\"1\"> </p>\n<p><img src=\"https://i.imgur.com/fUYju6U.png\" alt=\"2\"></p>\n<p>Why this one doesn't have a ground-truth for carts? Does the created ground-truth belongs to a timestamp after the cart happens?</p>\n<p><img src=\"https://i.imgur.com/dIAYGi4.png\" alt=\"3\"></p>\n<p>I think ground-truth labels belong to highlighted timestamp. Is this an intended behavior?</p>\n<p>Edit: I figured why it looks like that. Sessions are also randomly split while creating ground-truth. I'm not sure whether this is prone to leakage or not.</p>",
      "rawMarkdown": "I was doing sanity check and I didn't understand how ground-truth is created. Can someone verify they are correct or not?\n\n![1](https://i.imgur.com/E8pJ5ZK.png) \n\n![2](https://i.imgur.com/fUYju6U.png)\n\nWhy this one doesn't have a ground-truth for carts? Does the created ground-truth belongs to a timestamp after the cart happens?\n\n![3](https://i.imgur.com/dIAYGi4.png)\n\nI think ground-truth labels belong to highlighted timestamp. Is this an intended behavior?\n\nEdit: I figured why it looks like that. Sessions are also randomly split while creating ground-truth. I'm not sure whether this is prone to leakage or not."
    },
    {
      "id": 2034529,
      "postDate": "2022-11-18T08:02:06.210Z",
      "content": "<blockquote>\n  <p>IMPORTANT CAVEAT: Please note that in my dataset, in order to minimize its size without losing information (unless you care about milliseconds, but I assume seconds should be enough 🙂), I divide the ts column by 1000.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> is absolutely right about the loss of information. It is only millisecond accuracy is affected, not second level accuracy.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Fc9d2ab93252b3f994932fb9e2bd38bd4%2FScreen%20Shot%202022-11-18%20at%2016.03.30.png?generation=1668758635832236&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "> IMPORTANT CAVEAT: Please note that in my dataset, in order to minimize its size without losing information (unless you care about milliseconds, but I assume seconds should be enough 🙂), I divide the ts column by 1000.\n\n@radek1 is absolutely right about the loss of information. It is only millisecond accuracy is affected, not second level accuracy.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Fc9d2ab93252b3f994932fb9e2bd38bd4%2FScreen%20Shot%202022-11-18%20at%2016.03.30.png?generation=1668758635832236&alt=media)\n\n\n"
    },
    {
      "id": 2034277,
      "postDate": "2022-11-18T02:11:09.540Z",
      "content": "<blockquote>\n  <p>IMPORTANT CAVEAT: Please note that in my dataset, in order to minimize its size without losing information (unless you care about milliseconds, but I assume seconds should be enough 🙂), I divide the ts column by 1000. Notebooks on Kaggle by other people use unmodified data, so if you would want to run this validation with that code, you need to either make changes to the code or multiply the ts column by 1000.</p>\n</blockquote>\n<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, many thanks for this great post with rich resources to dig into. Sorry to still bother you with the dividing 1000 issue.</p>\n<p>I am a little confused about which dataset you are referring to in the quotation at the top. In the dataset you said you \"divide the ts column by 1000\",   is it the dataset named \"Otto Full Optimized Memory Footprint\" which is processed by \"process_data.ipynb\"? If so, why I can't find the <code>/1000</code> in the <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843#2015604\" target=\"_blank\">source code</a> of \"process_data.ipynb\"?</p>\n<p>If you refer to other dataset, could you paste me the link to the source code where you divide ts by 1000? Thanks</p>",
      "rawMarkdown": "> IMPORTANT CAVEAT: Please note that in my dataset, in order to minimize its size without losing information (unless you care about milliseconds, but I assume seconds should be enough 🙂), I divide the ts column by 1000. Notebooks on Kaggle by other people use unmodified data, so if you would want to run this validation with that code, you need to either make changes to the code or multiply the ts column by 1000.\n\nHi @radek1, many thanks for this great post with rich resources to dig into. Sorry to still bother you with the dividing 1000 issue.\n\nI am a little confused about which dataset you are referring to in the quotation at the top. In the dataset you said you \"divide the ts column by 1000\",   is it the dataset named \"Otto Full Optimized Memory Footprint\" which is processed by \"process_data.ipynb\"? If so, why I can't find the `/1000` in the [source code](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843#2015604) of \"process_data.ipynb\"?\n\nIf you refer to other dataset, could you paste me the link to the source code where you divide ts by 1000? Thanks",
      "replies": [
        {
          "id": 2034279,
          "postDate": "2022-11-18T02:15:30.467Z",
          "content": "<p>I have experimented and can confirm that convert type from string to uint8 can save CPU RAM significantly. I don't see the code of dividing 1000 on <code>ts</code>, so I can't experiment to see whether it is really saving RAM. </p>",
          "rawMarkdown": "I have experimented and can confirm that convert type from string to uint8 can save CPU RAM significantly. I don't see the code of dividing 1000 on `ts`, so I can't experiment to see whether it is really saving RAM. "
        },
        {
          "id": 2034291,
          "postDate": "2022-11-18T02:27:39.300Z",
          "content": "<p>For validation like outlined here, I created <a href=\"https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation\" target=\"_blank\">this dataset</a> (it has test labels! 🙂)</p>\n<p>If you'd like to check what impact this processing has, you can load the data that doesn't have the <code>ts</code> column modified (you can use <a href=\"https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint/versions/1\" target=\"_blank\">an earlier version of my dataset</a>, call <code>memory_usage(deep=True)</code> on the dataframe, do <code>df.ts = (df.ts / 1000).astype(np.int32)</code> and call memory usage as well. You will notice you don't lose information (unless you care about milliseconds) and the data occupies much less memory!</p>\n<p>(btw I typed the commands out from memory, so maybe you need to change something here and there for it to work 🙂) </p>",
          "rawMarkdown": "For validation like outlined here, I created [this dataset](https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation) (it has test labels! 🙂)\n\nIf you'd like to check what impact this processing has, you can load the data that doesn't have the `ts` column modified (you can use [an earlier version of my dataset](https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint/versions/1), call `memory_usage(deep=True)` on the dataframe, do `df.ts = (df.ts / 1000).astype(np.int32)` and call memory usage as well. You will notice you don't lose information (unless you care about milliseconds) and the data occupies much less memory!\n\n(btw I typed the commands out from memory, so maybe you need to change something here and there for it to work 🙂) ",
          "votes": 1
        },
        {
          "id": 2034317,
          "postDate": "2022-11-18T03:30:40.877Z",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>!</p>\n<p>I now noticed that there are two versions of otto full, and the second version must have used <code>ts/1000</code>. I tried to add the first version in a notebook but I can only find the version 2 not the version 1. How do I add the version 1 instead of version 2 (the default version)? Thank you</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F256cc78d940aa1ccc27d66ea6647013d%2FScreen%20Shot%202022-11-18%20at%2011.28.20.png?generation=1668742201458255&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Thanks a lot @radek1!\n\nI now noticed that there are two versions of otto full, and the second version must have used `ts/1000`. I tried to add the first version in a notebook but I can only find the version 2 not the version 1. How do I add the version 1 instead of version 2 (the default version)? Thank you\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F256cc78d940aa1ccc27d66ea6647013d%2FScreen%20Shot%202022-11-18%20at%2011.28.20.png?generation=1668742201458255&alt=media)"
        },
        {
          "id": 2034424,
          "postDate": "2022-11-18T05:55:08.407Z",
          "content": "<p>Well, with a patient and careful look, the way to switch different versions is right there. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F9316d4f6815d810f3e155d169b1d4996%2FScreen%20Shot%202022-11-18%20at%2013.53.50.png?generation=1668750890542455&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Well, with a patient and careful look, the way to switch different versions is right there. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F9316d4f6815d810f3e155d169b1d4996%2FScreen%20Shot%202022-11-18%20at%2013.53.50.png?generation=1668750890542455&alt=media)",
          "votes": 1
        },
        {
          "id": 2034507,
          "postDate": "2022-11-18T07:27:35.447Z",
          "content": "<p>I just tested with the test set, and you are absolutely right that after <code>df.ts = (df.ts / 1000).astype(np.int32)</code>, the RAM usage of <code>ts</code> is halved. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Fc6c608210bab954761ebda1db0326b9d%2FScreen%20Shot%202022-11-18%20at%2014.56.11.png?generation=1668756449198493&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "I just tested with the test set, and you are absolutely right that after `df.ts = (df.ts / 1000).astype(np.int32)`, the RAM usage of `ts` is halved. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Fc6c608210bab954761ebda1db0326b9d%2FScreen%20Shot%202022-11-18%20at%2014.56.11.png?generation=1668756449198493&alt=media)",
          "votes": 2
        }
      ]
    },
    {
      "id": 2022985,
      "postDate": "2022-11-09T13:03:52.770Z",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, thank you for sharing. I am also trying to replicate your steps from scratch but get into trouble in the very first step. May I ask how did you create the train/valid sets using the host's repo? I am trying to do it with Kaggle notebook, what I did is:</p>\n<ol>\n<li><code>!git clone https://github.com/otto-de/recsys-dataset.git</code></li>\n<li>Insert this into the system paths</li>\n</ol>\n<pre><code> sys\nsys.path.insert(-, )\n</code></pre>\n<ol>\n<li>Run the script<br>\n<code>!pipenv run python -m src.testset --train-set ../input/otto-recommender-system/train.jsonl --days 7 --output-path 'out/' --seed 42</code><br>\nHowever, I constantly got the error saying <code>No module named 'src'</code>. I think I somehow put the host's repo int the wrong directory, but I don't know how to fix it. (Question from a Linux noob).</li>\n</ol>\n<p>Thank you.</p>",
      "rawMarkdown": "Hi, @radek1, thank you for sharing. I am also trying to replicate your steps from scratch but get into trouble in the very first step. May I ask how did you create the train/valid sets using the host's repo? I am trying to do it with Kaggle notebook, what I did is:\n1. `!git clone https://github.com/otto-de/recsys-dataset.git`\n2. Insert this into the system paths\n```python\nimport sys\nsys.path.insert(-1, './recsys-dataset')\n```\n3. Run the script\n`!pipenv run python -m src.testset --train-set ../input/otto-recommender-system/train.jsonl --days 7 --output-path 'out/' --seed 42`\nHowever, I constantly got the error saying `No module named 'src'`. I think I somehow put the host's repo int the wrong directory, but I don't know how to fix it. (Question from a Linux noob).\n\nThank you.",
      "replies": [
        {
          "id": 2023564,
          "postDate": "2022-11-09T21:18:40.517Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/shinomoriaoshi\" target=\"_blank\">@shinomoriaoshi</a>!</p>\n<p>I ran the script like this: <code>python -m src.testset --train-set data/train.jsonl --days 7 --output-path 'out/' --seed 42</code></p>\n<p>Hope this helps! 🙂</p>",
          "rawMarkdown": "Hey @shinomoriaoshi!\n\nI ran the script like this: `python -m src.testset --train-set data/train.jsonl --days 7 --output-path 'out/' --seed 42`\n\nHope this helps! 🙂",
          "votes": 2
        },
        {
          "id": 2057618,
          "postDate": "2022-12-07T08:59:23.913Z",
          "content": "<p><a href=\"https://www.kaggle.com/shinomoriaoshi\" target=\"_blank\">@shinomoriaoshi</a> have you run successfully the module from kaggle? If so hwo?</p>",
          "rawMarkdown": "@shinomoriaoshi have you run successfully the module from kaggle? If so hwo?"
        }
      ]
    },
    {
      "id": 2095668,
      "postDate": "2023-01-11T14:52:50.863Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2029213,
      "postDate": "2022-11-14T14:07:46.810Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 2029811,
          "postDate": "2022-11-15T00:28:33.123Z",
          "content": "<p>I outlined the method here, please see <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843#2015604\" target=\"_blank\">this comment</a></p>",
          "rawMarkdown": "I outlined the method here, please see [this comment](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843#2015604)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2109732,
      "postDate": "2023-01-21T16:01:39.520Z",
      "content": "<p>Thanks for your sharing</p>",
      "rawMarkdown": "Thanks for your sharing"
    },
    {
      "id": 2098515,
      "postDate": "2023-01-13T16:20:44.553Z",
      "content": "<p>So cool.Thanks for sharing</p>",
      "rawMarkdown": "So cool.Thanks for sharing"
    }
  ],
  "comments": [
    {
      "id": 2024516,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-11-10T15:25:34.453000",
      "content": "<p>FYI, I published a version of Radek's validation data as multiple parquets and restored columns <code>ts</code> and <code>type</code> back to their original. Radek's dataset is more memory and disk efficient, but I published this alternative dataset for anyone who has a pipeline using a Kaggle dataset with multiple parquets and original <code>ts</code> and <code>type</code> columns. </p>\n<p>Now you can run your exact pipeline with validation data and compute CV score without making any changes to your pipeline. The Kaggle dataset is <a href=\"https://www.kaggle.com/datasets/cdeotte/otto-validation\" target=\"_blank\">here</a>. I plan to publish a notebook which computes validation score for my notebook soon.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 2024873,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-11-10T20:03:44.977000",
          "content": "<p>UPDATE: I published a train notebook that achieves LB 0.573 and corresponding validation notebook \"CV\" 0.563 using Radek's validation data (in my Kaggle dataset).</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 2025144,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-11-11T02:37:13.650000",
          "content": "<p>After publishing the notebook, i made two changes that boosted CV <code>+0.003</code> and after submission, the LB boosted <code>+0.003</code>. The CV LB correlate very well.</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 2023756,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-11-10T02:19:52.620000",
      "content": "<p>Your CV scheme (data and metric code) has perfect correlation for me too. Thank you for publishing your setup. The first two data points are my public notebook version 4 and 9 <a href=\"https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost\" target=\"_blank\">here</a> (with LB 0.556 and LB 0.562 respectively). Then next data point is Pietro's notebook <a href=\"https://www.kaggle.com/code/pietromaldini1/multiple-clicks-vs-latest-items\" target=\"_blank\">here</a> (when it was/is LB 0.565 version 1). The next data point is Ingvaras' notebook <a href=\"https://www.kaggle.com/code/ingvarasgalinskas/item-type-vs-multiple-clicks-vs-latest-items\" target=\"_blank\">here</a> (when it was/is LB 0.570 version 3). We see that each notebook improves both CV and LB. The last data point is my current LB score using some unpublished tricks.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/cv_lb2.png\" alt=\"\"></p>",
      "votes": 8,
      "replies": [
        {
          "id": 2023758,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-10T02:21:12.023000",
          "content": "<p>wonderful data, thank you for sharing this <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! 🙂  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2075047,
          "author_name": "Jensen",
          "author_url": "",
          "post_date": "2022-12-25T01:16:59.220000",
          "content": "<p>Hi,Chris.I used your version 9 notebook to calculate the local CV. I suspect something went wrong with my idea of calculating local CV scores. In the notebook, I generated candidates for val A's validation set and then calculated the local CV score, I only made the following modifications, but the CV score looks so outrageous. What went wrong with me? I am looking forward to hearing from you.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11356689%2Ffc4f117d4f91aae5da0b8f807e71ad6a%2F1.png?generation=1671930944213119&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11356689%2Fa1e7af94b592aba55f2e8df7ad1383b5%2F2.png?generation=1671930967340151&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2053831,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2022-12-03T16:40:30.753000",
      "content": "<p>Thanks for sharing, this saves a lot of tine!</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2054296,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-04T02:24:31.563000",
          "content": "<p>Thank you, JFP 🙂 Really glad to be of help</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2023511,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-11-09T19:43:42.153000",
      "content": "<p>This is very helpful thank you. I used your files and code to compute the CV score for my notebook version 4 with LB 0.556. The computed CV score is 0.474 so I think there is a bug in my code somewhere. It appears that you have a smaller gap (between LB and CV) when using your CV files. I will review my code and try again…</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2023514,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-11-09T19:51:17.187000",
          "content": "<p>I may have found the problem. Your parquet files have <code>ts</code> column with <code>original_ts / 1000</code>. So my notebook's pipeline which adds weight based on <code>ts</code> doesn't work correctly. I will update and check CV score.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 2023515,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-11-09T19:56:11.370000",
          "content": "<p>You should add a note to your discussion here. Anyone who uses code based on original <code>ts</code> like my Chris' notebook, or Vladimir's notebook, or Pietro's notebook, or Ingvaras's notebook will run into this problem. For example the following line of code:</p>\n<pre><code>df.query('abs(ts_x - ts_y) &lt; 24 * 60 * 60 * 1000 and aid_x != aid_y')\n</code></pre>\n<p>will waste many hours running code with your pipeline before discovering the error. Also type conversion will fail</p>\n<pre><code>.astype({\"ts\": \"datetime64[ms]\"})\n</code></pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2023559,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-09T21:11:00.183000",
          "content": "<p>Yes Chris, that is a great point! Apologies for causing all this trouble on your end. You are absolutely right.</p>\n<p>I included this bit of information in the other dataset that I published, in the description, but didn't do so here. Very sorry bout that. Will correct this in a second and add a note.</p>\n<p>My apologies for all the trouble.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2023578,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-11-09T21:35:50.963000",
          "content": "<p>Great. Looks good now. My computed CV is 0.548 for version 4 of my notebook with LB 0.556. That's pretty close. I will compute more CV scores for different versions and check the correlation. I will post results here.</p>\n<p>Then i will begin boosting my CV LB with improved logic! and/or trained models!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2023731,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-10T01:43:16.273000",
          "content": "<p>It's really cool how you are sharing openly how you plan to approach this competition, I think that is great learning material! 🙂</p>\n<p>My plan is very similar, hoping to share some stuff on Matrix Factorization maybe with using some of the Merlin stuff, have some really cool ideas, but the only problem is finding the time 😄</p>\n<p>Great there are that many weeks still left in this competition! 🙂</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2029433,
      "author_name": "Piotr Gabrys",
      "author_url": "",
      "post_date": "2022-11-14T16:31:57.847000",
      "content": "<p>Thanks for the dataset. I've spotted a problem with validation dataset. There are \"new\" AIDs (18785) in the validation test set comparing to validations train set). For the original train and test set this is not the case.</p>\n<p>I've added notebook with the calculation to the dataset: <a href=\"https://www.kaggle.com/code/piotrekga/new-aids-in-test-set\" target=\"_blank\">https://www.kaggle.com/code/piotrekga/new-aids-in-test-set</a></p>\n<p>I guess there was more logic in train/test split. <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a>, could you please give us information on what to do with sessions and ground truth with new AIDs?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2029810,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-15T00:27:10.533000",
          "content": "<p>ha, interesting observation <a href=\"https://www.kaggle.com/piotrekga\" target=\"_blank\">@piotrekga</a>!</p>\n<p>I think this might just be coincidental, new AIDs might have gotten added to the system that week 🤔 Or, alternatively, there was some additional logic in pulling down the data.</p>\n<p><a href=\"https://github.com/otto-de/recsys-dataset/blob/main/src/testset.py\" target=\"_blank\">Here is the logic used by the organizer to create the test set</a>. There is nothing there about filtering sessions based on <code>aids</code> AFAICT at quick glance.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2029879,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-11-15T02:22:28.200000",
          "content": "<p>That's a helpful discovery Piotr. Differences between CV and LB are a concern.</p>\n<p>However, I will comment that Radek's validation data is perfectly correlating with LB. Every <code>0.0002+</code> improvement on Radek's CV gives me a boost on LB!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2030124,
          "author_name": "Piotr Gabrys",
          "author_url": "",
          "post_date": "2022-11-15T07:59:13.497000",
          "content": "<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> I've rerun the train-test split on <code>train.jsonl</code> before and got exactly the same results. You're right the logic is not included in the code provided by organisers.</p>\n<p>About new AIDs: taking into consideration that these are users behaviours and the dataset has long tails I find it very likely that there were no new AIDs in the original test set and thousands of them in new train-test split. Of course I may be wrong 😅. I really hope organisers will provide us with some additional clarification in that regards.</p>\n<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, you're right that now the CV-LB correlation works. I'd really like to validate my model locally as close as possible to the final validation though.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2032183,
          "author_name": "Philipp Normann",
          "author_url": "",
          "post_date": "2022-11-16T13:19:34.117000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/piotrekga\" target=\"_blank\">@piotrekga</a>, thanks for pointing this out and sorry for the confusion! This behavior was a bug in our public repo that was inconsistent with how we created the dataset for the competition. We created the actual dataset using a Clojure job, and during the reimplementation in Python, we, unfortunately, forgot to add this preprocessing step. I just pushed an <a href=\"https://github.com/otto-de/recsys-dataset/commit/14eea3938cd9b6200bda9985f2c7ff8bb1de1b18\" target=\"_blank\">update</a> to our GitHub repo which fixes this issue. Although I'm still not entirely happy with the performance of the Python script, it should be enough for you to create your local validation set, which reflects the logic we used for the LB test set. I'll see what I can do to speed things up 😌</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 2034542,
          "author_name": "Piotr Gabrys",
          "author_url": "",
          "post_date": "2022-11-18T08:18:33.893000",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> for the update. At this point logic is far more important than performance.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2093576,
      "author_name": "Pradeep Pujari",
      "author_url": "",
      "post_date": "2023-01-10T05:35:13.883000",
      "content": "<p>Sorry <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> what is the id2type and vice versa pkl files contain? What is the purpose of these 2 data files?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2094694,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2023-01-11T00:46:19.597000",
          "content": "<p>hey <a href=\"https://www.kaggle.com/ppujari\" target=\"_blank\">@ppujari</a>! They just contain the mapping of types to their ids IIRC, 0 maps to <code>clicks</code>, 1 to <code>carts</code> and  2 to <code>orders</code>. Probably shouldn't have included it, makes it all a bit confusing.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2076910,
      "author_name": "danielliao",
      "author_url": "",
      "post_date": "2022-12-27T02:03:19.867000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, I have found something weird of the validation dataset you provided. </p>\n<p>According to the organizer's script for creating the validation dataset, when we split the 4 weeks train into 3-weeks-train (part-1) and 1-week-train (part-2), the part-2's aids should have all been seen in part-1. However, the <code>test.parquet</code> in the part-2 still have more than 2% aids which are not seen by part-1.</p>\n<p>I have shown this problem in this version of the notebook <a href=\"https://www.kaggle.com/code/danielliao/reading-otto-recsys-organizer-scripts?scriptVersionId=114730145&amp;cellId=9\" target=\"_blank\">here</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2049291,
      "author_name": "tsk",
      "author_url": "",
      "post_date": "2022-11-30T03:18:22.520000",
      "content": "<p>Thank you for sharing!!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2053433,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-03T09:21:42.247000",
          "content": "<p>np <a href=\"https://www.kaggle.com/yuyatasaki\" target=\"_blank\">@yuyatasaki</a>! happy to be of help 🙂</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2025309,
      "author_name": "zhez017",
      "author_url": "",
      "post_date": "2022-11-11T06:29:47.053000",
      "content": "<p>Thank you for sharing.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2028983,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-14T11:07:04.497000",
          "content": "<p>np <a href=\"https://www.kaggle.com/zhenzheng0328\" target=\"_blank\">@zhenzheng0328</a>! Glad you found this useful! 🙌</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2024958,
      "author_name": "Raj Saha",
      "author_url": "",
      "post_date": "2022-11-10T22:35:50.763000",
      "content": "<p>Awesome Data to work with… Thanks for sharing this…</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2024969,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-10T22:50:34.307000",
          "content": "<p>NP <a href=\"https://www.kaggle.com/saha8631\" target=\"_blank\">@saha8631</a>, glad I could be of help! 🙌</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2069469,
      "author_name": "william.wu",
      "author_url": "",
      "post_date": "2022-12-19T02:44:29.387000",
      "content": "<p>Thanks for sharing. I have a question, after you split the first 3 weeks into the train the last week into the test. How did you further split to get the <code>test.parquet</code> and <code>test_labels.parquet</code>?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2069493,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-12-19T03:10:51.733000",
          "content": "<p>The process is explained in Otto's GitHub <a href=\"https://github.com/otto-de/recsys-dataset/tree/main/test\" target=\"_blank\">here</a>. These are the steps </p>\n<ul>\n<li>For every user in Kaggle 4 week of train data, we find their earliest event</li>\n<li>For every user with their earliest event in the first 3 weeks, they become the \"new train\" and their events get truncated to the first 3 weeks</li>\n<li>For every user whos first event begins in week 4, they become the \"test users\"</li>\n<li>Next, for each test user, we pick a random split with <code>np.random.uniform(1,len(user_events))</code>. The first random half of user activity becomes <code>test.parquet</code> and the second random half becomes <code>test_labels.parquet</code>.</li>\n<li>Lastly every event involving an item in <code>test_labels.parquet</code> that does not exist in train or test gets removed.</li>\n</ul>",
          "votes": 12,
          "replies": [
            {
              "id": 2069494,
              "author_name": "william.wu",
              "author_url": "",
              "post_date": "2022-12-19T03:18:49.043000",
              "content": "<p>Thanks for the explanation <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , very clear. You saved my day</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2088453,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-01-06T10:27:22.367000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2090215,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-01-07T05:33:11.877000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2059672,
      "author_name": "wakaka",
      "author_url": "",
      "post_date": "2022-12-09T05:56:54.437000",
      "content": "<p>Thank you for sharing validation strategy. I may be wrong because I haven't thought about this too much yet, but I would like to ask one question.<br>\nYou adopted the last 7 days of train data for validation set. This means the users are same in train and validation while LB and train is different in users. From your result, domain shift by user doesn't occur, right?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2059679,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-09T06:05:55.640000",
          "content": "<p>It's a valid concern, but there is no overlap. Any overlap of the sort you mention has been removed.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2059802,
          "author_name": "wakaka",
          "author_url": "",
          "post_date": "2022-12-09T08:36:39.703000",
          "content": "<p>Thank you so much！</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2057414,
      "author_name": "Wonho Song",
      "author_url": "",
      "post_date": "2022-12-07T04:52:09.797000",
      "content": "<p>Thank you for sharing :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2057621,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-07T09:04:02.877000",
          "content": "<p>Glad to be of help, <a href=\"https://www.kaggle.com/songwonho\" target=\"_blank\">@songwonho</a>! Thanks for your kind comment 🙂 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2053420,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2022-12-03T08:47:54.650000",
      "content": "<p>I was doing sanity check and I didn't understand how ground-truth is created. Can someone verify they are correct or not?</p>\n<p><img src=\"https://i.imgur.com/E8pJ5ZK.png\" alt=\"1\"> </p>\n<p><img src=\"https://i.imgur.com/fUYju6U.png\" alt=\"2\"></p>\n<p>Why this one doesn't have a ground-truth for carts? Does the created ground-truth belongs to a timestamp after the cart happens?</p>\n<p><img src=\"https://i.imgur.com/dIAYGi4.png\" alt=\"3\"></p>\n<p>I think ground-truth labels belong to highlighted timestamp. Is this an intended behavior?</p>\n<p>Edit: I figured why it looks like that. Sessions are also randomly split while creating ground-truth. I'm not sure whether this is prone to leakage or not.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2034529,
      "author_name": "danielliao",
      "author_url": "",
      "post_date": "2022-11-18T08:02:06.210000",
      "content": "<blockquote>\n  <p>IMPORTANT CAVEAT: Please note that in my dataset, in order to minimize its size without losing information (unless you care about milliseconds, but I assume seconds should be enough 🙂), I divide the ts column by 1000.</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> is absolutely right about the loss of information. It is only millisecond accuracy is affected, not second level accuracy.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Fc9d2ab93252b3f994932fb9e2bd38bd4%2FScreen%20Shot%202022-11-18%20at%2016.03.30.png?generation=1668758635832236&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2034277,
      "author_name": "danielliao",
      "author_url": "",
      "post_date": "2022-11-18T02:11:09.540000",
      "content": "<blockquote>\n  <p>IMPORTANT CAVEAT: Please note that in my dataset, in order to minimize its size without losing information (unless you care about milliseconds, but I assume seconds should be enough 🙂), I divide the ts column by 1000. Notebooks on Kaggle by other people use unmodified data, so if you would want to run this validation with that code, you need to either make changes to the code or multiply the ts column by 1000.</p>\n</blockquote>\n<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, many thanks for this great post with rich resources to dig into. Sorry to still bother you with the dividing 1000 issue.</p>\n<p>I am a little confused about which dataset you are referring to in the quotation at the top. In the dataset you said you \"divide the ts column by 1000\",   is it the dataset named \"Otto Full Optimized Memory Footprint\" which is processed by \"process_data.ipynb\"? If so, why I can't find the <code>/1000</code> in the <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843#2015604\" target=\"_blank\">source code</a> of \"process_data.ipynb\"?</p>\n<p>If you refer to other dataset, could you paste me the link to the source code where you divide ts by 1000? Thanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2034279,
          "author_name": "danielliao",
          "author_url": "",
          "post_date": "2022-11-18T02:15:30.467000",
          "content": "<p>I have experimented and can confirm that convert type from string to uint8 can save CPU RAM significantly. I don't see the code of dividing 1000 on <code>ts</code>, so I can't experiment to see whether it is really saving RAM. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2034291,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-18T02:27:39.300000",
          "content": "<p>For validation like outlined here, I created <a href=\"https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation\" target=\"_blank\">this dataset</a> (it has test labels! 🙂)</p>\n<p>If you'd like to check what impact this processing has, you can load the data that doesn't have the <code>ts</code> column modified (you can use <a href=\"https://www.kaggle.com/datasets/radek1/otto-full-optimized-memory-footprint/versions/1\" target=\"_blank\">an earlier version of my dataset</a>, call <code>memory_usage(deep=True)</code> on the dataframe, do <code>df.ts = (df.ts / 1000).astype(np.int32)</code> and call memory usage as well. You will notice you don't lose information (unless you care about milliseconds) and the data occupies much less memory!</p>\n<p>(btw I typed the commands out from memory, so maybe you need to change something here and there for it to work 🙂) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2034317,
          "author_name": "danielliao",
          "author_url": "",
          "post_date": "2022-11-18T03:30:40.877000",
          "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>!</p>\n<p>I now noticed that there are two versions of otto full, and the second version must have used <code>ts/1000</code>. I tried to add the first version in a notebook but I can only find the version 2 not the version 1. How do I add the version 1 instead of version 2 (the default version)? Thank you</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F256cc78d940aa1ccc27d66ea6647013d%2FScreen%20Shot%202022-11-18%20at%2011.28.20.png?generation=1668742201458255&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2034424,
          "author_name": "danielliao",
          "author_url": "",
          "post_date": "2022-11-18T05:55:08.407000",
          "content": "<p>Well, with a patient and careful look, the way to switch different versions is right there. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F9316d4f6815d810f3e155d169b1d4996%2FScreen%20Shot%202022-11-18%20at%2013.53.50.png?generation=1668750890542455&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2034507,
          "author_name": "danielliao",
          "author_url": "",
          "post_date": "2022-11-18T07:27:35.447000",
          "content": "<p>I just tested with the test set, and you are absolutely right that after <code>df.ts = (df.ts / 1000).astype(np.int32)</code>, the RAM usage of <code>ts</code> is halved. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Fc6c608210bab954761ebda1db0326b9d%2FScreen%20Shot%202022-11-18%20at%2014.56.11.png?generation=1668756449198493&amp;alt=media\" alt=\"\"></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2022985,
      "author_name": "Minh Tri Phan",
      "author_url": "",
      "post_date": "2022-11-09T13:03:52.770000",
      "content": "<p>Hi, <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a>, thank you for sharing. I am also trying to replicate your steps from scratch but get into trouble in the very first step. May I ask how did you create the train/valid sets using the host's repo? I am trying to do it with Kaggle notebook, what I did is:</p>\n<ol>\n<li><code>!git clone https://github.com/otto-de/recsys-dataset.git</code></li>\n<li>Insert this into the system paths</li>\n</ol>\n<pre><code> sys\nsys.path.insert(-, )\n</code></pre>\n<ol>\n<li>Run the script<br>\n<code>!pipenv run python -m src.testset --train-set ../input/otto-recommender-system/train.jsonl --days 7 --output-path 'out/' --seed 42</code><br>\nHowever, I constantly got the error saying <code>No module named 'src'</code>. I think I somehow put the host's repo int the wrong directory, but I don't know how to fix it. (Question from a Linux noob).</li>\n</ol>\n<p>Thank you.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2023564,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-09T21:18:40.517000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/shinomoriaoshi\" target=\"_blank\">@shinomoriaoshi</a>!</p>\n<p>I ran the script like this: <code>python -m src.testset --train-set data/train.jsonl --days 7 --output-path 'out/' --seed 42</code></p>\n<p>Hope this helps! 🙂</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2057618,
          "author_name": "Davide Stenner",
          "author_url": "",
          "post_date": "2022-12-07T08:59:23.913000",
          "content": "<p><a href=\"https://www.kaggle.com/shinomoriaoshi\" target=\"_blank\">@shinomoriaoshi</a> have you run successfully the module from kaggle? If so hwo?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2095668,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-11T14:52:50.863000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2029213,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-14T14:07:46.810000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 2029811,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-15T00:28:33.123000",
          "content": "<p>I outlined the method here, please see <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843#2015604\" target=\"_blank\">this comment</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2109732,
      "author_name": "carolineoo",
      "author_url": "",
      "post_date": "2023-01-21T16:01:39.520000",
      "content": "<p>Thanks for your sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2098515,
      "author_name": "Jason",
      "author_url": "",
      "post_date": "2023-01-13T16:20:44.553000",
      "content": "<p>So cool.Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2022844": "Hey!\n\nI ran a couple of experiments and local validation tracks public LB perfectly! Here is what the situation looks like:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F457ff5950918744c841a83ac4c0ece1a%2Fdownload.png?generation=1667989795036578&alt=media)\n\nThe results are slightly poorer in local validation, but that is not surprising -- we are running on one week less of train data!\n\nHere's the setup:\n* I generated the split of the train set using the code in [organizer's repo](https://github.com/otto-de/recsys-dataset)\n* I am using the last 7 days of train set for validation (similarly to the test set)\n* If you don't want to run the code on the data yourself, I [uploaded the prepared data to Kaggle](https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation)\n\nThere has been a lot of conversation on the forums on the metric, I think we now have a good understanding of how it works 🙂 (one can also always confirm one's understanding by checking out the organizer's repo!). Interesting discussions on the forums [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364064) and [here](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364530).\n\nHere is the code I am using for calculating the metric in pandas, starting from a valid submission DataFrame, in case this might be of help:\n\n```\n    submission['session'] = submission.session_type.apply(lambda x: int(x.split('_')[0]))\n    submission['type'] = submission.session_type.apply(lambda x: x.split('_')[1])\n    submission.labels = submission.labels.apply(lambda x: [int(i) for i in x.split(' ')[:20]])\n\n    test_labels = pd.read_parquet('out_processed/test_labels.parquet')\n    test_labels = test_labels.merge(submission, how='left', on=['session', 'type'])\n    test_labels['hits'] = test_labels.apply(lambda df: len(set(df.ground_truth).intersection(set(df.labels))), axis=1)\n    test_labels['gt_count'] = test_labels.ground_truth.str.len().clip(0,20)\n\n    recall_per_type = test_labels.groupby(['type'])['hits'].sum() / test_labels.groupby(['type'])['gt_count'].sum() \n\n    score = (recall_per_type * pd.Series({'clicks': 0.10, 'carts': 0.30, 'orders': 0.60})).sum()\n```\n\nCompetitions where local validation tracks public LB are much more fun, IMHO 🙂\n\nWishing you all the best in the challenge!\nRadek\n\n**IMPORTANT CAVEAT:** Please note that in my dataset, in order to minimize its size without losing information (unless you care about milliseconds, but I assume seconds should be enough 🙂), I divide the `ts` column by 1000. Notebooks on Kaggle by other people use unmodified data, so if you would want to run this validation with that code, you need to either make changes to the code or multiply the `ts` column by 1000.\n\n**A couple of related resources you might find useful:**\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)\n\n",
    "2024516": "FYI, I published a version of Radek's validation data as multiple parquets and restored columns `ts` and `type` back to their original. Radek's dataset is more memory and disk efficient, but I published this alternative dataset for anyone who has a pipeline using a Kaggle dataset with multiple parquets and original `ts` and `type` columns. \n\nNow you can run your exact pipeline with validation data and compute CV score without making any changes to your pipeline. The Kaggle dataset is [here][1]. I plan to publish a notebook which computes validation score for my notebook soon.\n\n[1]: https://www.kaggle.com/datasets/cdeotte/otto-validation",
    "2023756": "Your CV scheme (data and metric code) has perfect correlation for me too. Thank you for publishing your setup. The first two data points are my public notebook version 4 and 9 [here][2] (with LB 0.556 and LB 0.562 respectively). Then next data point is Pietro's notebook [here][1] (when it was/is LB 0.565 version 1). The next data point is Ingvaras' notebook [here][3] (when it was/is LB 0.570 version 3). We see that each notebook improves both CV and LB. The last data point is my current LB score using some unpublished tricks.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Nov-2022/cv_lb2.png)\n\n[1]: https://www.kaggle.com/code/pietromaldini1/multiple-clicks-vs-latest-items\n[2]: https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost\n[3]: https://www.kaggle.com/code/ingvarasgalinskas/item-type-vs-multiple-clicks-vs-latest-items",
    "2053831": "Thanks for sharing, this saves a lot of tine!",
    "2023511": "This is very helpful thank you. I used your files and code to compute the CV score for my notebook version 4 with LB 0.556. The computed CV score is 0.474 so I think there is a bug in my code somewhere. It appears that you have a smaller gap (between LB and CV) when using your CV files. I will review my code and try again...",
    "2029433": "Thanks for the dataset. I've spotted a problem with validation dataset. There are \"new\" AIDs (18785) in the validation test set comparing to validations train set). For the original train and test set this is not the case.\n\nI've added notebook with the calculation to the dataset: https://www.kaggle.com/code/piotrekga/new-aids-in-test-set\n\nI guess there was more logic in train/test split. @pnormann, could you please give us information on what to do with sessions and ground truth with new AIDs?",
    "2093576": "Sorry @radek1 what is the id2type and vice versa pkl files contain? What is the purpose of these 2 data files?",
    "2076910": "Hi @radek1, I have found something weird of the validation dataset you provided. \n\nAccording to the organizer's script for creating the validation dataset, when we split the 4 weeks train into 3-weeks-train (part-1) and 1-week-train (part-2), the part-2's aids should have all been seen in part-1. However, the `test.parquet` in the part-2 still have more than 2% aids which are not seen by part-1.\n\nI have shown this problem in this version of the notebook [here](https://www.kaggle.com/code/danielliao/reading-otto-recsys-organizer-scripts?scriptVersionId=114730145&cellId=9)",
    "2049291": "Thank you for sharing!!",
    "2025309": "Thank you for sharing.",
    "2024958": "Awesome Data to work with... Thanks for sharing this...",
    "2069469": "Thanks for sharing. I have a question, after you split the first 3 weeks into the train the last week into the test. How did you further split to get the `test.parquet` and `test_labels.parquet`?",
    "2059672": "Thank you for sharing validation strategy. I may be wrong because I haven't thought about this too much yet, but I would like to ask one question.\nYou adopted the last 7 days of train data for validation set. This means the users are same in train and validation while LB and train is different in users. From your result, domain shift by user doesn't occur, right?\n",
    "2057414": "Thank you for sharing :)",
    "2053420": "I was doing sanity check and I didn't understand how ground-truth is created. Can someone verify they are correct or not?\n\n![1](https://i.imgur.com/E8pJ5ZK.png) \n\n![2](https://i.imgur.com/fUYju6U.png)\n\nWhy this one doesn't have a ground-truth for carts? Does the created ground-truth belongs to a timestamp after the cart happens?\n\n![3](https://i.imgur.com/dIAYGi4.png)\n\nI think ground-truth labels belong to highlighted timestamp. Is this an intended behavior?\n\nEdit: I figured why it looks like that. Sessions are also randomly split while creating ground-truth. I'm not sure whether this is prone to leakage or not.",
    "2034529": "> IMPORTANT CAVEAT: Please note that in my dataset, in order to minimize its size without losing information (unless you care about milliseconds, but I assume seconds should be enough 🙂), I divide the ts column by 1000.\n\n@radek1 is absolutely right about the loss of information. It is only millisecond accuracy is affected, not second level accuracy.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2Fc9d2ab93252b3f994932fb9e2bd38bd4%2FScreen%20Shot%202022-11-18%20at%2016.03.30.png?generation=1668758635832236&alt=media)\n\n\n",
    "2034277": "> IMPORTANT CAVEAT: Please note that in my dataset, in order to minimize its size without losing information (unless you care about milliseconds, but I assume seconds should be enough 🙂), I divide the ts column by 1000. Notebooks on Kaggle by other people use unmodified data, so if you would want to run this validation with that code, you need to either make changes to the code or multiply the ts column by 1000.\n\nHi @radek1, many thanks for this great post with rich resources to dig into. Sorry to still bother you with the dividing 1000 issue.\n\nI am a little confused about which dataset you are referring to in the quotation at the top. In the dataset you said you \"divide the ts column by 1000\",   is it the dataset named \"Otto Full Optimized Memory Footprint\" which is processed by \"process_data.ipynb\"? If so, why I can't find the `/1000` in the [source code](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843#2015604) of \"process_data.ipynb\"?\n\nIf you refer to other dataset, could you paste me the link to the source code where you divide ts by 1000? Thanks",
    "2022985": "Hi, @radek1, thank you for sharing. I am also trying to replicate your steps from scratch but get into trouble in the very first step. May I ask how did you create the train/valid sets using the host's repo? I am trying to do it with Kaggle notebook, what I did is:\n1. `!git clone https://github.com/otto-de/recsys-dataset.git`\n2. Insert this into the system paths\n```python\nimport sys\nsys.path.insert(-1, './recsys-dataset')\n```\n3. Run the script\n`!pipenv run python -m src.testset --train-set ../input/otto-recommender-system/train.jsonl --days 7 --output-path 'out/' --seed 42`\nHowever, I constantly got the error saying `No module named 'src'`. I think I somehow put the host's repo int the wrong directory, but I don't know how to fix it. (Question from a Linux noob).\n\nThank you.",
    "2095668": "",
    "2029213": "",
    "2109732": "Thanks for your sharing",
    "2098515": "So cool.Thanks for sharing"
  }
}