{
  "id": 364722,
  "title": "🐘 the elephant in the room -- high cardinality of targets and what to do about this",
  "url": "/competitions/otto-recommender-system/discussion/364722",
  "author_name": "",
  "post_date": "2022-11-07T23:42:55.453217800Z",
  "votes": 20,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hey!</p>\n<p>So the cardinality of <code>AIDs</code> is nearly 2 million. That is way more than traditional methods can handle if we want to predict the next item (<code>aid</code>) in a sequence.</p>\n<p>Normally, when doing single-label classification, we would use softmax to output our predicitons. But softmax is known to not work that well beyond a certain size. Plus, it becomes very resource intensive to have over 1.5 million activations in a layer!</p>\n<p>I haven't seen deep learning models shared just yet, but one could imagine an RNN or a Transformer trained on sequences and outputting single (or multiple) labels. </p>\n<p>How are you addressing this situation? More generally, how are you working around the limitations of Softmax?</p>\n<p>One approach would certainly be sampled softmax, though I have not come across a good PyTorch implementation. Guess implementing it on your own shouldn't be too tricky.</p>\n<p>👉 now here is quite likely a very valuable idea: instead of using the covisitation matrix, train a matrix factorization model, I am not sure how I am doing on time this week but if time permits maybe I'll hack such a thing together and share 🙂</p>\n<p>With embeddings, you can use nearest neighbor search and the whole idea of cardinality goes away.</p>\n<p>What are your thoughts on this?</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
  "messages": [
    {
      "id": "2021034",
      "postDate": "11/07/2022 23:42:55",
      "content": "<p>Hey!</p>\n<p>So the cardinality of <code>AIDs</code> is nearly 2 million. That is way more than traditional methods can handle if we want to predict the next item (<code>aid</code>) in a sequence.</p>\n<p>Normally, when doing single-label classification, we would use softmax to output our predicitons. But softmax is known to not work that well beyond a certain size. Plus, it becomes very resource intensive to have over 1.5 million activations in a layer!</p>\n<p>I haven't seen deep learning models shared just yet, but one could imagine an RNN or a Transformer trained on sequences and outputting single (or multiple) labels. </p>\n<p>How are you addressing this situation? More generally, how are you working around the limitations of Softmax?</p>\n<p>One approach would certainly be sampled softmax, though I have not come across a good PyTorch implementation. Guess implementing it on your own shouldn't be too tricky.</p>\n<p>👉 now here is quite likely a very valuable idea: instead of using the covisitation matrix, train a matrix factorization model, I am not sure how I am doing on time this week but if time permits maybe I'll hack such a thing together and share 🙂</p>\n<p>With embeddings, you can use nearest neighbor search and the whole idea of cardinality goes away.</p>\n<p>What are your thoughts on this?</p>\n<h3>Other resources you might find useful:</h3>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions\" target=\"_blank\">💡 [2 methods] How-to ensemble predictions 🏅🏅🏅</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">local validation tracks public LB perfecty -- here is the setup</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560\" target=\"_blank\">💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843\" target=\"_blank\">Full dataset processed to CSV/parquet files with optimized memory footprint</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic\" target=\"_blank\">co-visitation matrix - simplified, imprvd logic 🔥</a></li>\n<li><a href=\"https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission\" target=\"_blank\">💡 Word2Vec How-to [training and submission]🚀🚀🚀</a></li>\n</ul>",
      "rawMarkdown": "Hey!\n\nSo the cardinality of `AIDs` is nearly 2 million. That is way more than traditional methods can handle if we want to predict the next item (`aid`) in a sequence.\n\nNormally, when doing single-label classification, we would use softmax to output our predicitons. But softmax is known to not work that well beyond a certain size. Plus, it becomes very resource intensive to have over 1.5 million activations in a layer!\n\nI haven't seen deep learning models shared just yet, but one could imagine an RNN or a Transformer trained on sequences and outputting single (or multiple) labels. \n\nHow are you addressing this situation? More generally, how are you working around the limitations of Softmax?\n\nOne approach would certainly be sampled softmax, though I have not come across a good PyTorch implementation. Guess implementing it on your own shouldn't be too tricky.\n\n👉 now here is quite likely a very valuable idea: instead of using the covisitation matrix, train a matrix factorization model, I am not sure how I am doing on time this week but if time permits maybe I'll hack such a thing together and share 🙂\n\nWith embeddings, you can use nearest neighbor search and the whole idea of cardinality goes away.\n\nWhat are your thoughts on this?\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
      "votes": null
    },
    {
      "id": "2021088",
      "postDate": "11/08/2022 01:44:56",
      "content": "<p>One technique that can handle high cardinality is \"candidate rerank\" models. For example, for each test session, we generate candidates. Then we merge features unto the candidates and finally predict final 20 via XGB rerank using the features.</p>",
      "rawMarkdown": "One technique that can handle high cardinality is \"candidate rerank\" models. For example, for each test session, we generate candidates. Then we merge features unto the candidates and finally predict final 20 via XGB rerank using the features.",
      "votes": null
    },
    {
      "id": "2021975",
      "postDate": "11/08/2022 15:47:07",
      "content": "<p>Ravi posted a discussion about this idea <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "Ravi posted a discussion about this idea [here][1]\n\n[1]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2021088,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/08/2022 01:44:56",
      "content": "<p>One technique that can handle high cardinality is \"candidate rerank\" models. For example, for each test session, we generate candidates. Then we merge features unto the candidates and finally predict final 20 via XGB rerank using the features.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2021975,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/08/2022 15:47:07",
          "content": "<p>Ravi posted a discussion about this idea <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721\" target=\"_blank\">here</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2021034": "Hey!\n\nSo the cardinality of `AIDs` is nearly 2 million. That is way more than traditional methods can handle if we want to predict the next item (`aid`) in a sequence.\n\nNormally, when doing single-label classification, we would use softmax to output our predicitons. But softmax is known to not work that well beyond a certain size. Plus, it becomes very resource intensive to have over 1.5 million activations in a layer!\n\nI haven't seen deep learning models shared just yet, but one could imagine an RNN or a Transformer trained on sequences and outputting single (or multiple) labels. \n\nHow are you addressing this situation? More generally, how are you working around the limitations of Softmax?\n\nOne approach would certainly be sampled softmax, though I have not come across a good PyTorch implementation. Guess implementing it on your own shouldn't be too tricky.\n\n👉 now here is quite likely a very valuable idea: instead of using the covisitation matrix, train a matrix factorization model, I am not sure how I am doing on time this week but if time permits maybe I'll hack such a thing together and share 🙂\n\nWith embeddings, you can use nearest neighbor search and the whole idea of cardinality goes away.\n\nWhat are your thoughts on this?\n\n### Other resources you might find useful:\n\n* [💡 [2 methods] How-to ensemble predictions 🏅🏅🏅](https://www.kaggle.com/code/radek1/2-methods-how-to-ensemble-predictions)\n* [local validation tracks public LB perfecty -- here is the setup](https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991)\n* [💡 For my friends from Twitter and LinkedIn -- here is how to dive into this competition 🐳](https://www.kaggle.com/competitions/otto-recommender-system/discussion/368560)\n* [Full dataset processed to CSV/parquet files with optimized memory footprint](https://www.kaggle.com/competitions/otto-recommender-system/discussion/363843)\n* [co-visitation matrix - simplified, imprvd logic 🔥](https://www.kaggle.com/code/radek1/co-visitation-matrix-simplified-imprvd-logic)\n* [💡 Word2Vec How-to [training and submission]🚀🚀🚀](https://www.kaggle.com/code/radek1/word2vec-how-to-training-and-submission)",
    "2021088": "One technique that can handle high cardinality is \"candidate rerank\" models. For example, for each test session, we generate candidates. Then we merge features unto the candidates and finally predict final 20 via XGB rerank using the features.",
    "2021975": "Ravi posted a discussion about this idea [here][1]\n\n[1]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364721"
  },
  "source": "meta"
}