{
  "id": 364534,
  "title": "📅 Dataset for local validation created using organizer's repository (parquet files)",
  "url": "/competitions/otto-recommender-system/discussion/364534",
  "author_name": "",
  "post_date": "2022-11-07T04:07:59.527445Z",
  "votes": 34,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hey!</p>\n<p>I created and uploaded onto Kaggle a dataset for local validation using code in <a href=\"https://github.com/otto-de/recsys-dataset\" target=\"_blank\">the organizer's repo</a>.</p>\n<p>Please find the dataset <a href=\"https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation\" target=\"_blank\">here</a>.</p>\n<p>I created the dataset passing the competition train data to the script and used the last 7 days to create the test set.</p>\n<p>Generally, the interesting bit here is that the test set created on the last week of the train set seems considerably easier to predict on than what we have in the competition.</p>\n<p>Another interesting observation s regarding the calculation of the metric. It is not a mean of per-row recall, rather it is recall calculated on hits across the entire dataset! (might be an important piece of information for anyone implementing the competition metric themselves!)</p>\n<p>Here is the relevant piece of code from <code>evaluate.py</code>:</p>\n<pre><code>def recall_by_event_type(evalutated_events: dict, total_number_events: dict):\n    clicks = 0\n    carts = 0\n    orders = 0\n    for event in evalutated_events.values():\n        if 'clicks' in event and event['clicks']:\n            clicks += event['clicks']\n        if 'carts' in event and event['carts']:\n            carts += event['carts']\n        if 'orders' in event and event['orders']:\n            orders += event['orders']\n\n    return {\n        'clicks': clicks / total_number_events['clicks'],\n        'carts': carts / total_number_events['carts'],\n        'orders': orders / total_number_events['orders']\n    }\n</code></pre>\n<p>And here is how I am calculating the metric based on the files I uploaded (and a submission).</p>\n<p>This code outputs the same result as the code in the organizer's repository! 🥳</p>\n<pre><code>submission['session'] = submission.session_type.apply(lambda x: int(x.split('_')[0]))\nsubmission['type'] = submission.session_type.apply(lambda x: x.split('_')[1])\nsubmission.labels = submission.labels.apply(lambda x: [int(i) for i in x.split(' ')[:20]])\n\ntest_labels = pd.read_parquet('out_processed/test_labels.parquet')\ntest_labels = test_labels.merge(submission, how='left', on=['session', 'type'])\ntest_labels['hits'] = test_labels.apply(lambda df: len(set(df.ground_truth).intersection(set(df.labels))), axis=1)\ntest_labels['gt_count'] = test_labels.ground_truth.str.len().clip(0,20)\n\nrecall_per_type = test_labels.groupby(['type'])['hits'].sum() / test_labels.groupby(['type'])['gt_count'].sum() \n\nscore = (recall_per_type * pd.Series({'clicks': 0.10, 'carts': 0.30, 'orders': 0.60})).sum()\n</code></pre>\n<p>Hope this can be of help 🙂 </p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": "2019895",
      "postDate": "11/07/2022 04:07:59",
      "content": "<p>Hey!</p>\n<p>I created and uploaded onto Kaggle a dataset for local validation using code in <a href=\"https://github.com/otto-de/recsys-dataset\" target=\"_blank\">the organizer's repo</a>.</p>\n<p>Please find the dataset <a href=\"https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation\" target=\"_blank\">here</a>.</p>\n<p>I created the dataset passing the competition train data to the script and used the last 7 days to create the test set.</p>\n<p>Generally, the interesting bit here is that the test set created on the last week of the train set seems considerably easier to predict on than what we have in the competition.</p>\n<p>Another interesting observation s regarding the calculation of the metric. It is not a mean of per-row recall, rather it is recall calculated on hits across the entire dataset! (might be an important piece of information for anyone implementing the competition metric themselves!)</p>\n<p>Here is the relevant piece of code from <code>evaluate.py</code>:</p>\n<pre><code>def recall_by_event_type(evalutated_events: dict, total_number_events: dict):\n    clicks = 0\n    carts = 0\n    orders = 0\n    for event in evalutated_events.values():\n        if 'clicks' in event and event['clicks']:\n            clicks += event['clicks']\n        if 'carts' in event and event['carts']:\n            carts += event['carts']\n        if 'orders' in event and event['orders']:\n            orders += event['orders']\n\n    return {\n        'clicks': clicks / total_number_events['clicks'],\n        'carts': carts / total_number_events['carts'],\n        'orders': orders / total_number_events['orders']\n    }\n</code></pre>\n<p>And here is how I am calculating the metric based on the files I uploaded (and a submission).</p>\n<p>This code outputs the same result as the code in the organizer's repository! 🥳</p>\n<pre><code>submission['session'] = submission.session_type.apply(lambda x: int(x.split('_')[0]))\nsubmission['type'] = submission.session_type.apply(lambda x: x.split('_')[1])\nsubmission.labels = submission.labels.apply(lambda x: [int(i) for i in x.split(' ')[:20]])\n\ntest_labels = pd.read_parquet('out_processed/test_labels.parquet')\ntest_labels = test_labels.merge(submission, how='left', on=['session', 'type'])\ntest_labels['hits'] = test_labels.apply(lambda df: len(set(df.ground_truth).intersection(set(df.labels))), axis=1)\ntest_labels['gt_count'] = test_labels.ground_truth.str.len().clip(0,20)\n\nrecall_per_type = test_labels.groupby(['type'])['hits'].sum() / test_labels.groupby(['type'])['gt_count'].sum() \n\nscore = (recall_per_type * pd.Series({'clicks': 0.10, 'carts': 0.30, 'orders': 0.60})).sum()\n</code></pre>\n<p>Hope this can be of help 🙂 </p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "Hey!\n\nI created and uploaded onto Kaggle a dataset for local validation using code in [the organizer's repo](https://github.com/otto-de/recsys-dataset).\n\nPlease find the dataset [here](https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation).\n\nI created the dataset passing the competition train data to the script and used the last 7 days to create the test set.\n\nGenerally, the interesting bit here is that the test set created on the last week of the train set seems considerably easier to predict on than what we have in the competition.\n\nAnother interesting observation s regarding the calculation of the metric. It is not a mean of per-row recall, rather it is recall calculated on hits across the entire dataset! (might be an important piece of information for anyone implementing the competition metric themselves!)\n\nHere is the relevant piece of code from `evaluate.py`:\n```\ndef recall_by_event_type(evalutated_events: dict, total_number_events: dict):\n    clicks = 0\n    carts = 0\n    orders = 0\n    for event in evalutated_events.values():\n        if 'clicks' in event and event['clicks']:\n            clicks += event['clicks']\n        if 'carts' in event and event['carts']:\n            carts += event['carts']\n        if 'orders' in event and event['orders']:\n            orders += event['orders']\n\n    return {\n        'clicks': clicks / total_number_events['clicks'],\n        'carts': carts / total_number_events['carts'],\n        'orders': orders / total_number_events['orders']\n    }\n```\n\nAnd here is how I am calculating the metric based on the files I uploaded (and a submission).\n\nThis code outputs the same result as the code in the organizer's repository! 🥳\n\n```\nsubmission['session'] = submission.session_type.apply(lambda x: int(x.split('_')[0]))\nsubmission['type'] = submission.session_type.apply(lambda x: x.split('_')[1])\nsubmission.labels = submission.labels.apply(lambda x: [int(i) for i in x.split(' ')[:20]])\n\ntest_labels = pd.read_parquet('out_processed/test_labels.parquet')\ntest_labels = test_labels.merge(submission, how='left', on=['session', 'type'])\ntest_labels['hits'] = test_labels.apply(lambda df: len(set(df.ground_truth).intersection(set(df.labels))), axis=1)\ntest_labels['gt_count'] = test_labels.ground_truth.str.len().clip(0,20)\n\nrecall_per_type = test_labels.groupby(['type'])['hits'].sum() / test_labels.groupby(['type'])['gt_count'].sum() \n\nscore = (recall_per_type * pd.Series({'clicks': 0.10, 'carts': 0.30, 'orders': 0.60})).sum()\n```\n\nHope this can be of help 🙂 \n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": null
    },
    {
      "id": "2020083",
      "postDate": "11/07/2022 07:18:57",
      "content": "<p>This is the code I used to convert the data to <code>parquet</code> files.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F83ae2309640454544a6377e657c9b325%2FScreenshot%202022-11-07%20171835.png?generation=1667805551208861&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "This is the code I used to convert the data to `parquet` files.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F83ae2309640454544a6377e657c9b325%2FScreenshot%202022-11-07%20171835.png?generation=1667805551208861&alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2020083,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "11/07/2022 07:18:57",
      "content": "<p>This is the code I used to convert the data to <code>parquet</code> files.<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F83ae2309640454544a6377e657c9b325%2FScreenshot%202022-11-07%20171835.png?generation=1667805551208861&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2019895": "Hey!\n\nI created and uploaded onto Kaggle a dataset for local validation using code in [the organizer's repo](https://github.com/otto-de/recsys-dataset).\n\nPlease find the dataset [here](https://www.kaggle.com/datasets/radek1/otto-train-and-test-data-for-local-validation).\n\nI created the dataset passing the competition train data to the script and used the last 7 days to create the test set.\n\nGenerally, the interesting bit here is that the test set created on the last week of the train set seems considerably easier to predict on than what we have in the competition.\n\nAnother interesting observation s regarding the calculation of the metric. It is not a mean of per-row recall, rather it is recall calculated on hits across the entire dataset! (might be an important piece of information for anyone implementing the competition metric themselves!)\n\nHere is the relevant piece of code from `evaluate.py`:\n```\ndef recall_by_event_type(evalutated_events: dict, total_number_events: dict):\n    clicks = 0\n    carts = 0\n    orders = 0\n    for event in evalutated_events.values():\n        if 'clicks' in event and event['clicks']:\n            clicks += event['clicks']\n        if 'carts' in event and event['carts']:\n            carts += event['carts']\n        if 'orders' in event and event['orders']:\n            orders += event['orders']\n\n    return {\n        'clicks': clicks / total_number_events['clicks'],\n        'carts': carts / total_number_events['carts'],\n        'orders': orders / total_number_events['orders']\n    }\n```\n\nAnd here is how I am calculating the metric based on the files I uploaded (and a submission).\n\nThis code outputs the same result as the code in the organizer's repository! 🥳\n\n```\nsubmission['session'] = submission.session_type.apply(lambda x: int(x.split('_')[0]))\nsubmission['type'] = submission.session_type.apply(lambda x: x.split('_')[1])\nsubmission.labels = submission.labels.apply(lambda x: [int(i) for i in x.split(' ')[:20]])\n\ntest_labels = pd.read_parquet('out_processed/test_labels.parquet')\ntest_labels = test_labels.merge(submission, how='left', on=['session', 'type'])\ntest_labels['hits'] = test_labels.apply(lambda df: len(set(df.ground_truth).intersection(set(df.labels))), axis=1)\ntest_labels['gt_count'] = test_labels.ground_truth.str.len().clip(0,20)\n\nrecall_per_type = test_labels.groupby(['type'])['hits'].sum() / test_labels.groupby(['type'])['gt_count'].sum() \n\nscore = (recall_per_type * pd.Series({'clicks': 0.10, 'carts': 0.30, 'orders': 0.60})).sum()\n```\n\nHope this can be of help 🙂 \n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
    "2020083": "This is the code I used to convert the data to `parquet` files.![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F83267%2F83ae2309640454544a6377e657c9b325%2FScreenshot%202022-11-07%20171835.png?generation=1667805551208861&alt=media)"
  },
  "source": "meta"
}