{
  "id": 364216,
  "title": "💡A robust local validation framework 🚀🚀🚀",
  "url": "/competitions/otto-recommender-system/discussion/364216",
  "author_name": "Radek Osmulski",
  "post_date": "2022-11-05T08:42:45.665000",
  "votes": 31,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hey!</p>\n<p>One thing that has been missing so far in this competition was a local validation framework. So I went ahead and implemented one!</p>\n<p>Please find it <a href=\"https://www.kaggle.com/radek1/a-robust-local-validation-framework\" target=\"_blank\">here</a>.</p>\n<p>Essentially, I am using the last week of the train set as test/validation week. I discard the original test data.</p>\n<p>I studied quite closely all the information about how the competition test set was created and attempted to mimic it as best as I could, quite optimistic I got it right. </p>\n<p>I am also calculating the metric in a way that I believe corresponds to the LB implementation. The score certainly seems reasonable. But would love another set of eyes on all this 🙂</p>\n<p>If you find anything that seems off to you, please let us know.</p>\n<p><strong>EDIT:</strong><br>\nI added a crucial bit (clip) in the following line (this is from the conversation with <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> here</p>\n<p>submission_with_gt.labels_y = submission_with_gt.labels_y.str.len().clip(0,20)</p>\n<p>This should now be 100% aligned with the competition metric 🙂</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": 2017946,
      "postDate": "2022-11-05T08:42:45.667Z",
      "content": "<p>Hey!</p>\n<p>One thing that has been missing so far in this competition was a local validation framework. So I went ahead and implemented one!</p>\n<p>Please find it <a href=\"https://www.kaggle.com/radek1/a-robust-local-validation-framework\" target=\"_blank\">here</a>.</p>\n<p>Essentially, I am using the last week of the train set as test/validation week. I discard the original test data.</p>\n<p>I studied quite closely all the information about how the competition test set was created and attempted to mimic it as best as I could, quite optimistic I got it right. </p>\n<p>I am also calculating the metric in a way that I believe corresponds to the LB implementation. The score certainly seems reasonable. But would love another set of eyes on all this 🙂</p>\n<p>If you find anything that seems off to you, please let us know.</p>\n<p><strong>EDIT:</strong><br>\nI added a crucial bit (clip) in the following line (this is from the conversation with <a href=\"https://www.kaggle.com/pnormann\" target=\"_blank\">@pnormann</a> here</p>\n<p>submission_with_gt.labels_y = submission_with_gt.labels_y.str.len().clip(0,20)</p>\n<p>This should now be 100% aligned with the competition metric 🙂</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "Hey!\n\nOne thing that has been missing so far in this competition was a local validation framework. So I went ahead and implemented one!\n\nPlease find it [here](https://www.kaggle.com/radek1/a-robust-local-validation-framework).\n\nEssentially, I am using the last week of the train set as test/validation week. I discard the original test data.\n\nI studied quite closely all the information about how the competition test set was created and attempted to mimic it as best as I could, quite optimistic I got it right. \n\nI am also calculating the metric in a way that I believe corresponds to the LB implementation. The score certainly seems reasonable. But would love another set of eyes on all this 🙂\n\nIf you find anything that seems off to you, please let us know.\n\n**EDIT:**\nI added a crucial bit (clip) in the following line (this is from the conversation with @pnormann here\n\nsubmission_with_gt.labels_y = submission_with_gt.labels_y.str.len().clip(0,20)\n\nThis should now be 100% aligned with the competition metric 🙂\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)\n",
      "votes": 31
    },
    {
      "id": 2057346,
      "postDate": "2022-12-07T01:58:22.993Z",
      "content": "<p>Help a lot for understanding the task! Thx!</p>",
      "rawMarkdown": "Help a lot for understanding the task! Thx!",
      "votes": 1,
      "replies": [
        {
          "id": 2057347,
          "postDate": "2022-12-07T01:59:21.383Z",
          "content": "<p>np <a href=\"https://www.kaggle.com/fire15\" target=\"_blank\">@fire15</a>! 🙂 very glad you are finding this useful! 🙌 </p>",
          "rawMarkdown": "np @fire15! 🙂 very glad you are finding this useful! 🙌 "
        }
      ]
    },
    {
      "id": 2022723,
      "postDate": "2022-11-09T08:11:47.500Z",
      "content": "<p>Good job!  </p>",
      "rawMarkdown": "Good job!  ",
      "votes": 1,
      "replies": [
        {
          "id": 2022849,
          "postDate": "2022-11-09T10:43:00.147Z",
          "content": "<p>thanks a lot, <a href=\"https://www.kaggle.com/evilpsycho42\" target=\"_blank\">@evilpsycho42</a>! 😊 </p>",
          "rawMarkdown": "thanks a lot, @evilpsycho42! 😊 "
        }
      ]
    },
    {
      "id": 2020370,
      "postDate": "2022-11-07T12:10:34.960Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2057346,
      "author_name": "Mr.Fire",
      "author_url": "",
      "post_date": "2022-12-07T01:58:22.993000",
      "content": "<p>Help a lot for understanding the task! Thx!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2057347,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-07T01:59:21.383000",
          "content": "<p>np <a href=\"https://www.kaggle.com/fire15\" target=\"_blank\">@fire15</a>! 🙂 very glad you are finding this useful! 🙌 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2022723,
      "author_name": "KKY",
      "author_url": "",
      "post_date": "2022-11-09T08:11:47.500000",
      "content": "<p>Good job!  </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2022849,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-11-09T10:43:00.147000",
          "content": "<p>thanks a lot, <a href=\"https://www.kaggle.com/evilpsycho42\" target=\"_blank\">@evilpsycho42</a>! 😊 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2020370,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-11-07T12:10:34.960000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2017946": "Hey!\n\nOne thing that has been missing so far in this competition was a local validation framework. So I went ahead and implemented one!\n\nPlease find it [here](https://www.kaggle.com/radek1/a-robust-local-validation-framework).\n\nEssentially, I am using the last week of the train set as test/validation week. I discard the original test data.\n\nI studied quite closely all the information about how the competition test set was created and attempted to mimic it as best as I could, quite optimistic I got it right. \n\nI am also calculating the metric in a way that I believe corresponds to the LB implementation. The score certainly seems reasonable. But would love another set of eyes on all this 🙂\n\nIf you find anything that seems off to you, please let us know.\n\n**EDIT:**\nI added a crucial bit (clip) in the following line (this is from the conversation with @pnormann here\n\nsubmission_with_gt.labels_y = submission_with_gt.labels_y.str.len().clip(0,20)\n\nThis should now be 100% aligned with the competition metric 🙂\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)\n",
    "2057346": "Help a lot for understanding the task! Thx!",
    "2022723": "Good job!  ",
    "2020370": ""
  }
}