{
  "id": 373224,
  "title": "Otto's Talk on Transformer Recommendation Systems",
  "url": "/competitions/otto-recommender-system/discussion/373224",
  "author_name": "",
  "post_date": "2022-12-20T08:59:39.972175100Z",
  "votes": 24,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello,</p>\n<p>I found a talk by Otto, which shares how they implemented Transformer for Recommender Systems:<br>\n<a href=\"https://www.youtube.com/watch?v=kLub89faE1o\" target=\"_blank\">https://www.youtube.com/watch?v=kLub89faE1o</a></p>\n<p>The talk is in German, so I share a short summary relevant for the competition in English. The first half is about setting up AI projects.</p>\n<p>What is relevant for the competition?</p>\n<ul>\n<li>They review publications relevant for their problem - YOOCHOOSE dataset is close to their problem and they reviewed for papers using the dataset. First reproducing the results and then applying to their dataset.</li>\n<li>The SASRec paper describes +20% improvements going from GRU4Rec+ to their transformer architectures.  </li>\n</ul>\n<p>Based on the insights, they tried out SASRec Paper - however, they had to adopt multiple changes to make it work for their dataset:</p>\n<ul>\n<li>Using Sampled Negatives to speed-up training time - otherwise calculating the output layer is inefficient</li>\n<li>Using InBatch Negative Sampling (32 samples) + Adding Uniform Sampling (8192 samples) by a batch-size of 256</li>\n<li>Do not draw negatives from the same session</li>\n<li>Using GRU4Rec+ Loss Function: BPR-max loss</li>\n</ul>\n<p>Additional Improvements:</p>\n<ul>\n<li>Using Sampled Softmax Loss (Wu &amp; Wang On the Effectiveness of Sampled Softmax Loss)</li>\n<li>Sorting data by time</li>\n</ul>",
  "messages": [
    {
      "id": "2070680",
      "postDate": "12/20/2022 08:59:39",
      "content": "<p>Hello,</p>\n<p>I found a talk by Otto, which shares how they implemented Transformer for Recommender Systems:<br>\n<a href=\"https://www.youtube.com/watch?v=kLub89faE1o\" target=\"_blank\">https://www.youtube.com/watch?v=kLub89faE1o</a></p>\n<p>The talk is in German, so I share a short summary relevant for the competition in English. The first half is about setting up AI projects.</p>\n<p>What is relevant for the competition?</p>\n<ul>\n<li>They review publications relevant for their problem - YOOCHOOSE dataset is close to their problem and they reviewed for papers using the dataset. First reproducing the results and then applying to their dataset.</li>\n<li>The SASRec paper describes +20% improvements going from GRU4Rec+ to their transformer architectures.  </li>\n</ul>\n<p>Based on the insights, they tried out SASRec Paper - however, they had to adopt multiple changes to make it work for their dataset:</p>\n<ul>\n<li>Using Sampled Negatives to speed-up training time - otherwise calculating the output layer is inefficient</li>\n<li>Using InBatch Negative Sampling (32 samples) + Adding Uniform Sampling (8192 samples) by a batch-size of 256</li>\n<li>Do not draw negatives from the same session</li>\n<li>Using GRU4Rec+ Loss Function: BPR-max loss</li>\n</ul>\n<p>Additional Improvements:</p>\n<ul>\n<li>Using Sampled Softmax Loss (Wu &amp; Wang On the Effectiveness of Sampled Softmax Loss)</li>\n<li>Sorting data by time</li>\n</ul>",
      "rawMarkdown": "Hello,\n\nI found a talk by Otto, which shares how they implemented Transformer for Recommender Systems:\nhttps://www.youtube.com/watch?v=kLub89faE1o\n\nThe talk is in German, so I share a short summary relevant for the competition in English. The first half is about setting up AI projects.\n\nWhat is relevant for the competition?\n- They review publications relevant for their problem - YOOCHOOSE dataset is close to their problem and they reviewed for papers using the dataset. First reproducing the results and then applying to their dataset.\n- The SASRec paper describes +20% improvements going from GRU4Rec+ to their transformer architectures.  \n\nBased on the insights, they tried out SASRec Paper - however, they had to adopt multiple changes to make it work for their dataset:\n- Using Sampled Negatives to speed-up training time - otherwise calculating the output layer is inefficient\n- Using InBatch Negative Sampling (32 samples) + Adding Uniform Sampling (8192 samples) by a batch-size of 256\n- Do not draw negatives from the same session\n- Using GRU4Rec+ Loss Function: BPR-max loss\n\nAdditional Improvements:\n- Using Sampled Softmax Loss (Wu & Wang On the Effectiveness of Sampled Softmax Loss)\n- Sorting data by time",
      "votes": null
    },
    {
      "id": "2073898",
      "postDate": "12/23/2022 14:25:44",
      "content": "<p>Thanks for sharing, this brings my memory of <a href=\"https://www.kaggle.com/competitions/riiid-test-answer-prediction/discussion?sort=recent-comments\" target=\"_blank\">Riid competition</a> back. The host published a paper named SAINT+ and then hosted the competition. Then we said \"SAINT+ is ALL YOU NEED\". Looking forward to someone posting \"SASRec is ALL YOU NEED\" in this competition:)</p>",
      "rawMarkdown": "Thanks for sharing, this brings my memory of [Riid competition](https://www.kaggle.com/competitions/riiid-test-answer-prediction/discussion?sort=recent-comments) back. The host published a paper named SAINT+ and then hosted the competition. Then we said \"SAINT+ is ALL YOU NEED\". Looking forward to someone posting \"SASRec is ALL YOU NEED\" in this competition:)",
      "votes": null
    },
    {
      "id": "2079058",
      "postDate": "12/28/2022 23:29:56",
      "content": "<p>Super cool find, <a href=\"https://www.kaggle.com/benediktschifferer\" target=\"_blank\">@benediktschifferer</a>! Thank you for sharing! </p>",
      "rawMarkdown": "Super cool find, @benediktschifferer! Thank you for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2073898,
      "author_name": "wuwenmin",
      "author_url": "",
      "post_date": "12/23/2022 14:25:44",
      "content": "<p>Thanks for sharing, this brings my memory of <a href=\"https://www.kaggle.com/competitions/riiid-test-answer-prediction/discussion?sort=recent-comments\" target=\"_blank\">Riid competition</a> back. The host published a paper named SAINT+ and then hosted the competition. Then we said \"SAINT+ is ALL YOU NEED\". Looking forward to someone posting \"SASRec is ALL YOU NEED\" in this competition:)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2079058,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "12/28/2022 23:29:56",
      "content": "<p>Super cool find, <a href=\"https://www.kaggle.com/benediktschifferer\" target=\"_blank\">@benediktschifferer</a>! Thank you for sharing! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2070680": "Hello,\n\nI found a talk by Otto, which shares how they implemented Transformer for Recommender Systems:\nhttps://www.youtube.com/watch?v=kLub89faE1o\n\nThe talk is in German, so I share a short summary relevant for the competition in English. The first half is about setting up AI projects.\n\nWhat is relevant for the competition?\n- They review publications relevant for their problem - YOOCHOOSE dataset is close to their problem and they reviewed for papers using the dataset. First reproducing the results and then applying to their dataset.\n- The SASRec paper describes +20% improvements going from GRU4Rec+ to their transformer architectures.  \n\nBased on the insights, they tried out SASRec Paper - however, they had to adopt multiple changes to make it work for their dataset:\n- Using Sampled Negatives to speed-up training time - otherwise calculating the output layer is inefficient\n- Using InBatch Negative Sampling (32 samples) + Adding Uniform Sampling (8192 samples) by a batch-size of 256\n- Do not draw negatives from the same session\n- Using GRU4Rec+ Loss Function: BPR-max loss\n\nAdditional Improvements:\n- Using Sampled Softmax Loss (Wu & Wang On the Effectiveness of Sampled Softmax Loss)\n- Sorting data by time",
    "2073898": "Thanks for sharing, this brings my memory of [Riid competition](https://www.kaggle.com/competitions/riiid-test-answer-prediction/discussion?sort=recent-comments) back. The host published a paper named SAINT+ and then hosted the competition. Then we said \"SAINT+ is ALL YOU NEED\". Looking forward to someone posting \"SASRec is ALL YOU NEED\" in this competition:)",
    "2079058": "Super cool find, @benediktschifferer! Thank you for sharing!"
  },
  "source": "meta"
}