{
  "id": 366893,
  "title": "[Starter Pack] Matrix Factorization [Pytorch + Merlin Dataloader] 🚀🚀🚀",
  "url": "/competitions/otto-recommender-system/discussion/366893",
  "author_name": "",
  "post_date": "2022-11-18T06:25:26.522680700Z",
  "votes": 33,
  "comment_count": 6,
  "views": 0,
  "content": "<p>One of the ways we can improve candidate generation is via training a matrix factorization model!</p>\n<p>In fact, for best results, one should probably generate candidates in multiple ways (including using the co-visitation matrices I discuss here: <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365358\" target=\"_blank\">💡 What is the co-visitation matrix, really?</a> and pulling it all together using a ranking model like shown here: <a href=\"https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker\" target=\"_blank\">💡 [polars] Proof of concept: LGBM Ranker🧪🧪🧪</a>.</p>\n<p>Essentially, there are so many things we can do with a matrix factorization model! We can generate candidates, we can score aids based on similarity, cluster them, and so on!</p>\n<p>So it is high time in the competition we had a matrix factorization model at our disposal!</p>\n<p>I share code for how to train one here: <a href=\"https://www.kaggle.com/code/radek1/matrix-factorization-pytorch-merlin-dataloader\" target=\"_blank\">💡Matrix Factorization [PyTorch+Merlin Dataloader]</a></p>\n<p>We preprocess data using polars, obtaining 200+ million aid pairs! To streamline the code I am using data preprocessed to parquet files (no need to deal with <code>jsonl</code> anymore!) that you can find here: <a href=\"https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-files\" target=\"_blank\">💡 [Howto] Full dataset as parquet/csv files</a>.</p>\n<p>We train using <a href=\"https://twitter.com/radekosmulski/status/1593125626361180160?s=20&amp;t=QEYdqN58jKtN8kXrffUMXQ\" target=\"_blank\">Merlin Dataloader</a> and PyTorch!</p>\n<p>We then proceed to create a submission using approximate nearest neighbor search.</p>\n<p>Hoping this can open up additional avenues to improve performance!</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": "2034460",
      "postDate": "11/18/2022 06:25:26",
      "content": "<p>One of the ways we can improve candidate generation is via training a matrix factorization model!</p>\n<p>In fact, for best results, one should probably generate candidates in multiple ways (including using the co-visitation matrices I discuss here: <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/365358\" target=\"_blank\">💡 What is the co-visitation matrix, really?</a> and pulling it all together using a ranking model like shown here: <a href=\"https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker\" target=\"_blank\">💡 [polars] Proof of concept: LGBM Ranker🧪🧪🧪</a>.</p>\n<p>Essentially, there are so many things we can do with a matrix factorization model! We can generate candidates, we can score aids based on similarity, cluster them, and so on!</p>\n<p>So it is high time in the competition we had a matrix factorization model at our disposal!</p>\n<p>I share code for how to train one here: <a href=\"https://www.kaggle.com/code/radek1/matrix-factorization-pytorch-merlin-dataloader\" target=\"_blank\">💡Matrix Factorization [PyTorch+Merlin Dataloader]</a></p>\n<p>We preprocess data using polars, obtaining 200+ million aid pairs! To streamline the code I am using data preprocessed to parquet files (no need to deal with <code>jsonl</code> anymore!) that you can find here: <a href=\"https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-files\" target=\"_blank\">💡 [Howto] Full dataset as parquet/csv files</a>.</p>\n<p>We train using <a href=\"https://twitter.com/radekosmulski/status/1593125626361180160?s=20&amp;t=QEYdqN58jKtN8kXrffUMXQ\" target=\"_blank\">Merlin Dataloader</a> and PyTorch!</p>\n<p>We then proceed to create a submission using approximate nearest neighbor search.</p>\n<p>Hoping this can open up additional avenues to improve performance!</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "One of the ways we can improve candidate generation is via training a matrix factorization model!\n\nIn fact, for best results, one should probably generate candidates in multiple ways (including using the co-visitation matrices I discuss here: [💡 What is the co-visitation matrix, really?](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365358) and pulling it all together using a ranking model like shown here: [💡 [polars] Proof of concept: LGBM Ranker🧪🧪🧪](https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker).\n\nEssentially, there are so many things we can do with a matrix factorization model! We can generate candidates, we can score aids based on similarity, cluster them, and so on!\n\nSo it is high time in the competition we had a matrix factorization model at our disposal!\n\nI share code for how to train one here: [💡Matrix Factorization [PyTorch+Merlin Dataloader]](https://www.kaggle.com/code/radek1/matrix-factorization-pytorch-merlin-dataloader)\n\nWe preprocess data using polars, obtaining 200+ million aid pairs! To streamline the code I am using data preprocessed to parquet files (no need to deal with `jsonl` anymore!) that you can find here: [💡 [Howto] Full dataset as parquet/csv files](https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-files).\n\nWe train using [Merlin Dataloader](https://twitter.com/radekosmulski/status/1593125626361180160?s=20&t=QEYdqN58jKtN8kXrffUMXQ) and PyTorch!\n\nWe then proceed to create a submission using approximate nearest neighbor search.\n\nHoping this can open up additional avenues to improve performance!\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": null
    },
    {
      "id": "2034628",
      "postDate": "11/18/2022 09:44:18",
      "content": "<p>great…great……..</p>",
      "rawMarkdown": "great...great........",
      "votes": null
    },
    {
      "id": "2034739",
      "postDate": "11/18/2022 11:47:22",
      "content": "<p>tnx for all the effort!</p>",
      "rawMarkdown": "tnx for all the effort!",
      "votes": null
    },
    {
      "id": "2036335",
      "postDate": "11/19/2022 16:43:12",
      "content": "<p>Nice work 👍 </p>\n<p>The Devastator.</p>",
      "rawMarkdown": "Nice work 👍 \n\n\nThe Devastator.",
      "votes": null
    },
    {
      "id": "2036666",
      "postDate": "11/20/2022 03:52:19",
      "content": "<p>Thanks for shareing!</p>",
      "rawMarkdown": "Thanks for shareing!",
      "votes": null
    },
    {
      "id": "2043588",
      "postDate": "11/25/2022 19:22:00",
      "content": "<p>recently  i was doing readings about Matrix Factorization and i found many of the ideas relate to recommendation systems solved by this technic rather than going for Graph NN because it has more complexity for modeling the problem situation, i would like to explore this technic using same method you worked on but for instead i will try to use Random search or Greedy algorithm to find best Hyperamaters that performed good fitting , <br>\nthis is the resource : <a href=\"https://everdark.github.io/k9/notebooks/ml/matrix_factorization/matrix_factorization.nb.html#3_neural_netork_representation\" target=\"_blank\">Matrix Factorization</a></p>",
      "rawMarkdown": "recently  i was doing readings about Matrix Factorization and i found many of the ideas relate to recommendation systems solved by this technic rather than going for Graph NN because it has more complexity for modeling the problem situation, i would like to explore this technic using same method you worked on but for instead i will try to use Random search or Greedy algorithm to find best Hyperamaters that performed good fitting , \nthis is the resource : [Matrix Factorization](https://everdark.github.io/k9/notebooks/ml/matrix_factorization/matrix_factorization.nb.html#3_neural_netork_representation)",
      "votes": null
    },
    {
      "id": "2043740",
      "postDate": "11/25/2022 21:05:18",
      "content": "<p>Sounds like a plan! 🙌 Best of luck in the hyperparameter search! 🙂</p>",
      "rawMarkdown": "Sounds like a plan! 🙌 Best of luck in the hyperparameter search! 🙂",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2034628,
      "author_name": "mridulsyed",
      "author_url": "",
      "post_date": "11/18/2022 09:44:18",
      "content": "<p>great…great……..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2034739,
      "author_name": "simonveitner",
      "author_url": "",
      "post_date": "11/18/2022 11:47:22",
      "content": "<p>tnx for all the effort!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2036335,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "11/19/2022 16:43:12",
      "content": "<p>Nice work 👍 </p>\n<p>The Devastator.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2036666,
      "author_name": "yutokaggle",
      "author_url": "",
      "post_date": "11/20/2022 03:52:19",
      "content": "<p>Thanks for shareing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2043588,
      "author_name": "younesselbrag",
      "author_url": "",
      "post_date": "11/25/2022 19:22:00",
      "content": "<p>recently  i was doing readings about Matrix Factorization and i found many of the ideas relate to recommendation systems solved by this technic rather than going for Graph NN because it has more complexity for modeling the problem situation, i would like to explore this technic using same method you worked on but for instead i will try to use Random search or Greedy algorithm to find best Hyperamaters that performed good fitting , <br>\nthis is the resource : <a href=\"https://everdark.github.io/k9/notebooks/ml/matrix_factorization/matrix_factorization.nb.html#3_neural_netork_representation\" target=\"_blank\">Matrix Factorization</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2043740,
          "author_name": "radek1",
          "author_url": "",
          "post_date": "11/25/2022 21:05:18",
          "content": "<p>Sounds like a plan! 🙌 Best of luck in the hyperparameter search! 🙂</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2034460": "One of the ways we can improve candidate generation is via training a matrix factorization model!\n\nIn fact, for best results, one should probably generate candidates in multiple ways (including using the co-visitation matrices I discuss here: [💡 What is the co-visitation matrix, really?](https://www.kaggle.com/competitions/otto-recommender-system/discussion/365358) and pulling it all together using a ranking model like shown here: [💡 [polars] Proof of concept: LGBM Ranker🧪🧪🧪](https://www.kaggle.com/code/radek1/polars-proof-of-concept-lgbm-ranker).\n\nEssentially, there are so many things we can do with a matrix factorization model! We can generate candidates, we can score aids based on similarity, cluster them, and so on!\n\nSo it is high time in the competition we had a matrix factorization model at our disposal!\n\nI share code for how to train one here: [💡Matrix Factorization [PyTorch+Merlin Dataloader]](https://www.kaggle.com/code/radek1/matrix-factorization-pytorch-merlin-dataloader)\n\nWe preprocess data using polars, obtaining 200+ million aid pairs! To streamline the code I am using data preprocessed to parquet files (no need to deal with `jsonl` anymore!) that you can find here: [💡 [Howto] Full dataset as parquet/csv files](https://www.kaggle.com/code/radek1/howto-full-dataset-as-parquet-csv-files).\n\nWe train using [Merlin Dataloader](https://twitter.com/radekosmulski/status/1593125626361180160?s=20&t=QEYdqN58jKtN8kXrffUMXQ) and PyTorch!\n\nWe then proceed to create a submission using approximate nearest neighbor search.\n\nHoping this can open up additional avenues to improve performance!\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
    "2034628": "great...great........",
    "2034739": "tnx for all the effort!",
    "2036335": "Nice work 👍 \n\n\nThe Devastator.",
    "2036666": "Thanks for shareing!",
    "2043588": "recently  i was doing readings about Matrix Factorization and i found many of the ideas relate to recommendation systems solved by this technic rather than going for Graph NN because it has more complexity for modeling the problem situation, i would like to explore this technic using same method you worked on but for instead i will try to use Random search or Greedy algorithm to find best Hyperamaters that performed good fitting , \nthis is the resource : [Matrix Factorization](https://everdark.github.io/k9/notebooks/ml/matrix_factorization/matrix_factorization.nb.html#3_neural_netork_representation)",
    "2043740": "Sounds like a plan! 🙌 Best of luck in the hyperparameter search! 🙂"
  },
  "source": "meta"
}