{
  "id": 364210,
  "title": "Co-visitation matrix with simplified code and improved logic! 🚀 [LB 0.558]",
  "url": "/competitions/otto-recommender-system/discussion/364210",
  "author_name": "",
  "post_date": "2022-11-05T07:59:08.790002800Z",
  "votes": 31,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Please find the kernel <a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">here</a>!</p>\n<p>I simplified the code so that it is easier to understand what is happening. Also, this makes it easier to come up with improvements and they are more straightforward to implement.</p>\n<p>There is some code repetition, but I specifically didn't wrap things into functions to make the notebook more hackable 🙂</p>\n<p>For quick experimentation, a good fraction of data to run the notebook on is 1/1000. This should get you to around 0.488 and the notebook finishes in under 4 minutes.</p>\n<p>I am running on data I shared here in <a href=\"https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-files\" target=\"_blank\">[Howto] Full dataset as parquet/csv files</a> which is one of the factors that allowed for the simplification of the code 🙂</p>\n<p>We are probably only scratching the surface on what is possible with this approach and a much better score can likely be achieved with more experimentation.</p>\n<p>Happy hacking! 🙂</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": "2017905",
      "postDate": "11/05/2022 07:59:08",
      "content": "<p>Please find the kernel <a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">here</a>!</p>\n<p>I simplified the code so that it is easier to understand what is happening. Also, this makes it easier to come up with improvements and they are more straightforward to implement.</p>\n<p>There is some code repetition, but I specifically didn't wrap things into functions to make the notebook more hackable 🙂</p>\n<p>For quick experimentation, a good fraction of data to run the notebook on is 1/1000. This should get you to around 0.488 and the notebook finishes in under 4 minutes.</p>\n<p>I am running on data I shared here in <a href=\"https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-files\" target=\"_blank\">[Howto] Full dataset as parquet/csv files</a> which is one of the factors that allowed for the simplification of the code 🙂</p>\n<p>We are probably only scratching the surface on what is possible with this approach and a much better score can likely be achieved with more experimentation.</p>\n<p>Happy hacking! 🙂</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "Please find the kernel [here](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)!\n\nI simplified the code so that it is easier to understand what is happening. Also, this makes it easier to come up with improvements and they are more straightforward to implement.\n\nThere is some code repetition, but I specifically didn't wrap things into functions to make the notebook more hackable 🙂\n\nFor quick experimentation, a good fraction of data to run the notebook on is 1/1000. This should get you to around 0.488 and the notebook finishes in under 4 minutes.\n\nI am running on data I shared here in [[Howto] Full dataset as parquet/csv files](https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-files) which is one of the factors that allowed for the simplification of the code 🙂\n\nWe are probably only scratching the surface on what is possible with this approach and a much better score can likely be achieved with more experimentation.\n\nHappy hacking! 🙂\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": null
    },
    {
      "id": "2018461",
      "postDate": "11/05/2022 17:54:05",
      "content": "<p>Thanks for sharing Radek. I did a quick comparison between your code and Vladimir's (which is the code i am using too) to determine where your <code>+0.012</code> boost is coming from. For curious Kagglers, it appears that most of your boost comes from this line of code</p>\n<pre><code>AIDs = list(dict.fromkeys(AIDs[::-1]))\n</code></pre>\n<p>in code cell 10. This is a great idea. When i add this line to my notebook (which is a fork of Vladimir's notebook), our LB score is nearly the same as yours. (See my notebook version 4 <a href=\"https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost?scriptVersionId=110131985\" target=\"_blank\">here</a>).</p>",
      "rawMarkdown": "Thanks for sharing Radek. I did a quick comparison between your code and Vladimir's (which is the code i am using too) to determine where your `+0.012` boost is coming from. For curious Kagglers, it appears that most of your boost comes from this line of code\n\n    AIDs = list(dict.fromkeys(AIDs[::-1]))\n\nin code cell 10. This is a great idea. When i add this line to my notebook (which is a fork of Vladimir's notebook), our LB score is nearly the same as yours. (See my notebook version 4 [here][1]).\n\n[1]: https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost?scriptVersionId=110131985",
      "votes": null
    },
    {
      "id": "2018696",
      "postDate": "11/05/2022 22:40:15",
      "content": "<p>Thank you, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! That is a great experiment, thank you for sharing the results! 😊</p>\n<p>That line is quite interesting I must say, it is an idiom I looked up on the Internet to solve this specific problem. For anyone who might be wondering what it does exactly, it removes duplicated items from a list, but preserves list ordering!</p>\n<p>Normally we would use <code>set()</code> to remove duplicated entries but using <code>set()</code> the order is gone. </p>",
      "rawMarkdown": "Thank you, @cdeotte! That is a great experiment, thank you for sharing the results! 😊\n\nThat line is quite interesting I must say, it is an idiom I looked up on the Internet to solve this specific problem. For anyone who might be wondering what it does exactly, it removes duplicated items from a list, but preserves list ordering!\n\nNormally we would use `set()` to remove duplicated entries but using `set()` the order is gone.",
      "votes": null
    },
    {
      "id": "2023039",
      "postDate": "11/09/2022 13:41:09",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> Thanks so much for sharing!</p>\n<p>As a total beginner, it's interesting to see that Co-visitation matrix needs no gpu but only 91 minutes to train on a CPU. </p>\n<p>Thanks for making the code simple enough so that a newbie like me can understand how to change code to do the experiment on 1/1000 dataset, which only takes less than 4 minutes to run as you promised 👍 (thanks for keeping on improving it)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F9f0094afa5e84b8d08ac5957d4a760a1%2FScreen%20Shot%202022-11-09%20at%2021.42.11.png?generation=1668001375376757&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi @radek1 Thanks so much for sharing!\n\nAs a total beginner, it's interesting to see that Co-visitation matrix needs no gpu but only 91 minutes to train on a CPU. \n\nThanks for making the code simple enough so that a newbie like me can understand how to change code to do the experiment on 1/1000 dataset, which only takes less than 4 minutes to run as you promised 👍 (thanks for keeping on improving it)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F9f0094afa5e84b8d08ac5957d4a760a1%2FScreen%20Shot%202022-11-09%20at%2021.42.11.png?generation=1668001375376757&alt=media)",
      "votes": null
    },
    {
      "id": "2023727",
      "postDate": "11/10/2022 01:36:46",
      "content": "<p>No worries, really glad you are finding this useful! 🙂🙏</p>",
      "rawMarkdown": "No worries, really glad you are finding this useful! 🙂🙏",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2018461,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/05/2022 17:54:05",
      "content": "<p>Thanks for sharing Radek. I did a quick comparison between your code and Vladimir's (which is the code i am using too) to determine where your <code>+0.012</code> boost is coming from. For curious Kagglers, it appears that most of your boost comes from this line of code</p>\n<pre><code>AIDs = list(dict.fromkeys(AIDs[::-1]))\n</code></pre>\n<p>in code cell 10. This is a great idea. When i add this line to my notebook (which is a fork of Vladimir's notebook), our LB score is nearly the same as yours. (See my notebook version 4 <a href=\"https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost?scriptVersionId=110131985\" target=\"_blank\">here</a>).</p>",
      "votes": null,
      "replies": [
        {
          "id": 2018696,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/05/2022 22:40:15",
          "content": "<p>Thank you, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! That is a great experiment, thank you for sharing the results! 😊</p>\n<p>That line is quite interesting I must say, it is an idiom I looked up on the Internet to solve this specific problem. For anyone who might be wondering what it does exactly, it removes duplicated items from a list, but preserves list ordering!</p>\n<p>Normally we would use <code>set()</code> to remove duplicated entries but using <code>set()</code> the order is gone. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2023039,
      "author_name": "danielliao",
      "author_url": "",
      "post_date": "11/09/2022 13:41:09",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">@radek1</a> Thanks so much for sharing!</p>\n<p>As a total beginner, it's interesting to see that Co-visitation matrix needs no gpu but only 91 minutes to train on a CPU. </p>\n<p>Thanks for making the code simple enough so that a newbie like me can understand how to change code to do the experiment on 1/1000 dataset, which only takes less than 4 minutes to run as you promised 👍 (thanks for keeping on improving it)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F9f0094afa5e84b8d08ac5957d4a760a1%2FScreen%20Shot%202022-11-09%20at%2021.42.11.png?generation=1668001375376757&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2023727,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/10/2022 01:36:46",
          "content": "<p>No worries, really glad you are finding this useful! 🙂🙏</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2017905": "Please find the kernel [here](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)!\n\nI simplified the code so that it is easier to understand what is happening. Also, this makes it easier to come up with improvements and they are more straightforward to implement.\n\nThere is some code repetition, but I specifically didn't wrap things into functions to make the notebook more hackable 🙂\n\nFor quick experimentation, a good fraction of data to run the notebook on is 1/1000. This should get you to around 0.488 and the notebook finishes in under 4 minutes.\n\nI am running on data I shared here in [[Howto] Full dataset as parquet/csv files](https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-files) which is one of the factors that allowed for the simplification of the code 🙂\n\nWe are probably only scratching the surface on what is possible with this approach and a much better score can likely be achieved with more experimentation.\n\nHappy hacking! 🙂\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
    "2018461": "Thanks for sharing Radek. I did a quick comparison between your code and Vladimir's (which is the code i am using too) to determine where your `+0.012` boost is coming from. For curious Kagglers, it appears that most of your boost comes from this line of code\n\n    AIDs = list(dict.fromkeys(AIDs[::-1]))\n\nin code cell 10. This is a great idea. When i add this line to my notebook (which is a fork of Vladimir's notebook), our LB score is nearly the same as yours. (See my notebook version 4 [here][1]).\n\n[1]: https://www.kaggle.com/code/cdeotte/test-data-leak-lb-boost?scriptVersionId=110131985",
    "2018696": "Thank you, @cdeotte! That is a great experiment, thank you for sharing the results! 😊\n\nThat line is quite interesting I must say, it is an idiom I looked up on the Internet to solve this specific problem. For anyone who might be wondering what it does exactly, it removes duplicated items from a list, but preserves list ordering!\n\nNormally we would use `set()` to remove duplicated entries but using `set()` the order is gone.",
    "2023039": "Hi @radek1 Thanks so much for sharing!\n\nAs a total beginner, it's interesting to see that Co-visitation matrix needs no gpu but only 91 minutes to train on a CPU. \n\nThanks for making the code simple enough so that a newbie like me can understand how to change code to do the experiment on 1/1000 dataset, which only takes less than 4 minutes to run as you promised 👍 (thanks for keeping on improving it)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F850197%2F9f0094afa5e84b8d08ac5957d4a760a1%2FScreen%20Shot%202022-11-09%20at%2021.42.11.png?generation=1668001375376757&alt=media)",
    "2023727": "No worries, really glad you are finding this useful! 🙂🙏"
  },
  "source": "meta"
}