{
  "id": 368885,
  "title": "Transformer + cls head is too slow, Need new train framework",
  "url": "/competitions/otto-recommender-system/discussion/368885",
  "author_name": "",
  "post_date": "2022-11-28T08:26:44.131400800Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I train a transformer model with cls head in this competition,  2 x RTX 3090 GPUs,  but it's to slow because the counts of the products is so big ,  I think we need use some other train framework for this comp.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1238483%2Fb4f5306f1757cb914c4edd22ca6bbc11%2F2731669622713_.pic.jpg?generation=1669623888556472&amp;alt=media\" alt=\"\">I</p>",
  "messages": [
    {
      "id": "2046547",
      "postDate": "11/28/2022 08:26:44",
      "content": "<p>I train a transformer model with cls head in this competition,  2 x RTX 3090 GPUs,  but it's to slow because the counts of the products is so big ,  I think we need use some other train framework for this comp.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1238483%2Fb4f5306f1757cb914c4edd22ca6bbc11%2F2731669622713_.pic.jpg?generation=1669623888556472&amp;alt=media\" alt=\"\">I</p>",
      "rawMarkdown": "I train a transformer model with cls head in this competition,  2 x RTX 3090 GPUs,  but it's to slow because the counts of the products is so big ,  I think we need use some other train framework for this comp.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1238483%2Fb4f5306f1757cb914c4edd22ca6bbc11%2F2731669622713_.pic.jpg?generation=1669623888556472&alt=media)I",
      "votes": null
    },
    {
      "id": "2046950",
      "postDate": "11/28/2022 13:43:20",
      "content": "<p>Truncate/Shorten the max session length (e.g. to 32~64) and keep only the top K (e.g. 10,000) most common products.</p>",
      "rawMarkdown": "Truncate/Shorten the max session length (e.g. to 32~64) and keep only the top K (e.g. 10,000) most common products.",
      "votes": null
    },
    {
      "id": "2046951",
      "postDate": "11/28/2022 13:43:54",
      "content": "<p>Also make sure you're running on a gpu. And if distributed really is running faster.</p>",
      "rawMarkdown": "Also make sure you're running on a gpu. And if distributed really is running faster.",
      "votes": null
    },
    {
      "id": "2046999",
      "postDate": "11/28/2022 14:26:10",
      "content": "<p>running on 2 gpus</p>",
      "rawMarkdown": "running on 2 gpus",
      "votes": null
    },
    {
      "id": "2047000",
      "postDate": "11/28/2022 14:30:34",
      "content": "<p>Yes, this will help.  I need to test the ceiling of score by using TopK most common products first. Thanks for your advice.</p>",
      "rawMarkdown": "Yes, this will help.  I need to test the ceiling of score by using TopK most common products first. Thanks for your advice.",
      "votes": null
    },
    {
      "id": "2051165",
      "postDate": "12/01/2022 09:02:18",
      "content": "<p>That is a problem with those high cardinality targets… </p>\n<p><a href=\"https://douglasorr.github.io/2021-10-training-objectives/3-sampled/article.html\" target=\"_blank\">Sampled softmax</a> has been working quite well in a lot of contexts, might be worth looking at 🙂</p>",
      "rawMarkdown": "That is a problem with those high cardinality targets... \n\n[Sampled softmax](https://douglasorr.github.io/2021-10-training-objectives/3-sampled/article.html) has been working quite well in a lot of contexts, might be worth looking at 🙂",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2046950,
      "author_name": "danofer",
      "author_url": "",
      "post_date": "11/28/2022 13:43:20",
      "content": "<p>Truncate/Shorten the max session length (e.g. to 32~64) and keep only the top K (e.g. 10,000) most common products.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2047000,
          "author_name": "evilpsycho42",
          "author_url": "",
          "post_date": "11/28/2022 14:30:34",
          "content": "<p>Yes, this will help.  I need to test the ceiling of score by using TopK most common products first. Thanks for your advice.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2046951,
      "author_name": "danofer",
      "author_url": "",
      "post_date": "11/28/2022 13:43:54",
      "content": "<p>Also make sure you're running on a gpu. And if distributed really is running faster.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2046999,
          "author_name": "evilpsycho42",
          "author_url": "",
          "post_date": "11/28/2022 14:26:10",
          "content": "<p>running on 2 gpus</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2051165,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "12/01/2022 09:02:18",
      "content": "<p>That is a problem with those high cardinality targets… </p>\n<p><a href=\"https://douglasorr.github.io/2021-10-training-objectives/3-sampled/article.html\" target=\"_blank\">Sampled softmax</a> has been working quite well in a lot of contexts, might be worth looking at 🙂</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2046547": "I train a transformer model with cls head in this competition,  2 x RTX 3090 GPUs,  but it's to slow because the counts of the products is so big ,  I think we need use some other train framework for this comp.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1238483%2Fb4f5306f1757cb914c4edd22ca6bbc11%2F2731669622713_.pic.jpg?generation=1669623888556472&alt=media)I",
    "2046950": "Truncate/Shorten the max session length (e.g. to 32~64) and keep only the top K (e.g. 10,000) most common products.",
    "2046951": "Also make sure you're running on a gpu. And if distributed really is running faster.",
    "2046999": "running on 2 gpus",
    "2047000": "Yes, this will help.  I need to test the ceiling of score by using TopK most common products first. Thanks for your advice.",
    "2051165": "That is a problem with those high cardinality targets... \n\n[Sampled softmax](https://douglasorr.github.io/2021-10-training-objectives/3-sampled/article.html) has been working quite well in a lot of contexts, might be worth looking at 🙂"
  },
  "source": "meta"
}