{
  "id": 367058,
  "title": "from zero to 60 in 2 seconds or less 🏎️🚓🚓🚓",
  "url": "/competitions/otto-recommender-system/discussion/367058",
  "author_name": "Radek Osmulski",
  "post_date": "2022-11-18T19:44:13.021000",
  "votes": 22,
  "comment_count": 0,
  "views": 0,
  "content": "<p>These are the two resources you NEED to get started with:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-files\" target=\"_blank\">💡 [Howto] Full dataset as parquet/csv files</a></li>\n<li><a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">Candidate ReRank Model - [LB 0.575]</a></li>\n</ol>\n<p>The first one makes it much more manageable to work with the data -- you can immediately jump into <code>pandas</code> instead of having to deal with <code>jasonl</code> files. Plus memory footprint is minimized significantly, saves a lot of RAM.</p>\n<p>And the second one is an unbelievably valuable notebook maintained by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. It is really hard to beat this heuristic approach 🙂 Plus it really pays off to familiarize yourself with what is going on there. You can use chunks of this notebook for candidate generation/scoring/reranking as you work on your expanded approach 🙂</p>\n<h3>So why this thread?</h3>\n<p>Just a day or two ago Kaggle sent out an email on this competition starting 😄 I think we might see an influx of people checking this competition out over the weekend. Hopefully, this thread can help them out make sense of what is going on quickly and keep them around with us.</p>\n<p>The more the merrier! 😄</p>\n<p>So looking forward to continuing to hack on this when I get a bit more time! But these two notebooks I link to above are definitely a great way to get up and running!</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": 2035366,
      "postDate": "2022-11-18T19:44:13.020Z",
      "content": "<p>These are the two resources you NEED to get started with:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-files\" target=\"_blank\">💡 [Howto] Full dataset as parquet/csv files</a></li>\n<li><a href=\"https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575\" target=\"_blank\">Candidate ReRank Model - [LB 0.575]</a></li>\n</ol>\n<p>The first one makes it much more manageable to work with the data -- you can immediately jump into <code>pandas</code> instead of having to deal with <code>jasonl</code> files. Plus memory footprint is minimized significantly, saves a lot of RAM.</p>\n<p>And the second one is an unbelievably valuable notebook maintained by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. It is really hard to beat this heuristic approach 🙂 Plus it really pays off to familiarize yourself with what is going on there. You can use chunks of this notebook for candidate generation/scoring/reranking as you work on your expanded approach 🙂</p>\n<h3>So why this thread?</h3>\n<p>Just a day or two ago Kaggle sent out an email on this competition starting 😄 I think we might see an influx of people checking this competition out over the weekend. Hopefully, this thread can help them out make sense of what is going on quickly and keep them around with us.</p>\n<p>The more the merrier! 😄</p>\n<p>So looking forward to continuing to hack on this when I get a bit more time! But these two notebooks I link to above are definitely a great way to get up and running!</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "These are the two resources you NEED to get started with:\n\n1. [💡 [Howto] Full dataset as parquet/csv files](https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-files)\n2. [Candidate ReRank Model - [LB 0.575]](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575)\n\nThe first one makes it much more manageable to work with the data -- you can immediately jump into `pandas` instead of having to deal with `jasonl` files. Plus memory footprint is minimized significantly, saves a lot of RAM.\n\nAnd the second one is an unbelievably valuable notebook maintained by @cdeotte. It is really hard to beat this heuristic approach 🙂 Plus it really pays off to familiarize yourself with what is going on there. You can use chunks of this notebook for candidate generation/scoring/reranking as you work on your expanded approach 🙂\n\n### So why this thread?\n\nJust a day or two ago Kaggle sent out an email on this competition starting 😄 I think we might see an influx of people checking this competition out over the weekend. Hopefully, this thread can help them out make sense of what is going on quickly and keep them around with us.\n\nThe more the merrier! 😄\n\nSo looking forward to continuing to hack on this when I get a bit more time! But these two notebooks I link to above are definitely a great way to get up and running!\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": 22
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2035366": "These are the two resources you NEED to get started with:\n\n1. [💡 [Howto] Full dataset as parquet/csv files](https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-files)\n2. [Candidate ReRank Model - [LB 0.575]](https://www.kaggle.com/code/cdeotte/candidate-rerank-model-lb-0-575)\n\nThe first one makes it much more manageable to work with the data -- you can immediately jump into `pandas` instead of having to deal with `jasonl` files. Plus memory footprint is minimized significantly, saves a lot of RAM.\n\nAnd the second one is an unbelievably valuable notebook maintained by @cdeotte. It is really hard to beat this heuristic approach 🙂 Plus it really pays off to familiarize yourself with what is going on there. You can use chunks of this notebook for candidate generation/scoring/reranking as you work on your expanded approach 🙂\n\n### So why this thread?\n\nJust a day or two ago Kaggle sent out an email on this competition starting 😄 I think we might see an influx of people checking this competition out over the weekend. Hopefully, this thread can help them out make sense of what is going on quickly and keep them around with us.\n\nThe more the merrier! 😄\n\nSo looking forward to continuing to hack on this when I get a bit more time! But these two notebooks I link to above are definitely a great way to get up and running!\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)"
  }
}