{
  "id": 384441,
  "title": "13th place.  BST transformer model",
  "url": "/competitions/otto-recommender-system/discussion/384441",
  "author_name": "",
  "post_date": "2023-02-07T23:02:28.937170700Z",
  "votes": 16,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Thank you very much for organizing this competition.  </p>\n<p>And thank you very much to my teammates <a href=\"https://www.kaggle.com/enric1296\" target=\"_blank\">@enric1296</a> <a href=\"https://www.kaggle.com/trasibulo\" target=\"_blank\">@trasibulo</a> and <a href=\"https://www.kaggle.com/albert2017\" target=\"_blank\">@albert2017</a>  We really enjoyed this competition!</p>\n<p>Congratulations to my teammates and other competitors who achieved their Master tier in this competition: well done! well deserved!!!</p>\n<p>This post is about one part of the final solution. <a href=\"https://www.kaggle.com/enric1296\" target=\"_blank\">@enric1296</a> posted <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/383229\" target=\"_blank\">the overall solution</a> </p>\n<p>I'd like to share here the main part of my contribution.  It was based on BST transformer model (<a href=\"https://arxiv.org/abs/1905.06874\" target=\"_blank\">arXiv:1905.06874</a>)</p>\n<p>The first time we used that solution was for the Kaggle Days Barcelona 11 hour datathon. That time Enric was the one doing the BST and I was in charge of feature engineering.</p>\n<p>For this competition I used all my teammates model inputs and my own ones to feed the model.</p>\n<p>And the first step was to translate Enric's datathon code from tensorflow to pytorch 😜</p>\n<p>Even when gbdt algorithms turned out to perform better for this problem, the transformer solution brought more diversity to the final ensemble.</p>\n<p>This is how articles positional features for the transformer looked like:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F223160%2F7a7c9a7ac0cbf9c9434a8bed41095ea1%2Ffig-1.png?generation=1675810715439219&amp;alt=media\" alt=\"\"></p>\n<p>And this was the scheme for the binary classifier reranker:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F223160%2F28c15d49f195ff344d01f4005f8e8969%2Ffig-2.png?generation=1675810742844854&amp;alt=media\" alt=\"\"></p>\n<h2>the k-fold split</h2>\n<p>If you tried to use StratifiedGroupKFold to stratify all the items from the ground truth, you easily run into a \"MemoryError: Unable to allocate 168. GiB for an array with shape …\"</p>\n<p>We solved it in an iterative way:</p>\n<pre><code>- stratify sessions containing less represented items for orders\n- stratify sessions containing the rest of the items for orders\n- stratify sessions containing less represented items for carts\n...\n</code></pre>",
  "messages": [
    {
      "id": "2134341",
      "postDate": "02/07/2023 23:02:28",
      "content": "<p>Thank you very much for organizing this competition.  </p>\n<p>And thank you very much to my teammates <a href=\"https://www.kaggle.com/enric1296\" target=\"_blank\">@enric1296</a> <a href=\"https://www.kaggle.com/trasibulo\" target=\"_blank\">@trasibulo</a> and <a href=\"https://www.kaggle.com/albert2017\" target=\"_blank\">@albert2017</a>  We really enjoyed this competition!</p>\n<p>Congratulations to my teammates and other competitors who achieved their Master tier in this competition: well done! well deserved!!!</p>\n<p>This post is about one part of the final solution. <a href=\"https://www.kaggle.com/enric1296\" target=\"_blank\">@enric1296</a> posted <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/383229\" target=\"_blank\">the overall solution</a> </p>\n<p>I'd like to share here the main part of my contribution.  It was based on BST transformer model (<a href=\"https://arxiv.org/abs/1905.06874\" target=\"_blank\">arXiv:1905.06874</a>)</p>\n<p>The first time we used that solution was for the Kaggle Days Barcelona 11 hour datathon. That time Enric was the one doing the BST and I was in charge of feature engineering.</p>\n<p>For this competition I used all my teammates model inputs and my own ones to feed the model.</p>\n<p>And the first step was to translate Enric's datathon code from tensorflow to pytorch 😜</p>\n<p>Even when gbdt algorithms turned out to perform better for this problem, the transformer solution brought more diversity to the final ensemble.</p>\n<p>This is how articles positional features for the transformer looked like:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F223160%2F7a7c9a7ac0cbf9c9434a8bed41095ea1%2Ffig-1.png?generation=1675810715439219&amp;alt=media\" alt=\"\"></p>\n<p>And this was the scheme for the binary classifier reranker:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F223160%2F28c15d49f195ff344d01f4005f8e8969%2Ffig-2.png?generation=1675810742844854&amp;alt=media\" alt=\"\"></p>\n<h2>the k-fold split</h2>\n<p>If you tried to use StratifiedGroupKFold to stratify all the items from the ground truth, you easily run into a \"MemoryError: Unable to allocate 168. GiB for an array with shape …\"</p>\n<p>We solved it in an iterative way:</p>\n<pre><code>- stratify sessions containing less represented items for orders\n- stratify sessions containing the rest of the items for orders\n- stratify sessions containing less represented items for carts\n...\n</code></pre>",
      "rawMarkdown": "Thank you very much for organizing this competition.  \n\nAnd thank you very much to my teammates @enric1296 @trasibulo and @albert2017  We really enjoyed this competition!\n\nCongratulations to my teammates and other competitors who achieved their Master tier in this competition: well done! well deserved!!!\n\nThis post is about one part of the final solution. @enric1296 posted [the overall solution](https://www.kaggle.com/competitions/otto-recommender-system/discussion/383229) \n\nI'd like to share here the main part of my contribution.  It was based on BST transformer model ([arXiv:1905.06874](https://arxiv.org/abs/1905.06874))\n\nThe first time we used that solution was for the Kaggle Days Barcelona 11 hour datathon. That time Enric was the one doing the BST and I was in charge of feature engineering.\n\nFor this competition I used all my teammates model inputs and my own ones to feed the model.\n\nAnd the first step was to translate Enric's datathon code from tensorflow to pytorch 😜\n\nEven when gbdt algorithms turned out to perform better for this problem, the transformer solution brought more diversity to the final ensemble.\n\nThis is how articles positional features for the transformer looked like:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F223160%2F7a7c9a7ac0cbf9c9434a8bed41095ea1%2Ffig-1.png?generation=1675810715439219&alt=media)\n\nAnd this was the scheme for the binary classifier reranker:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F223160%2F28c15d49f195ff344d01f4005f8e8969%2Ffig-2.png?generation=1675810742844854&alt=media)\n\n## the k-fold split\n\nIf you tried to use StratifiedGroupKFold to stratify all the items from the ground truth, you easily run into a \"MemoryError: Unable to allocate 168. GiB for an array with shape ...\"\n\nWe solved it in an iterative way:\n\n\t- stratify sessions containing less represented items for orders\n\t- stratify sessions containing the rest of the items for orders\n\t- stratify sessions containing less represented items for carts\n\t...",
      "votes": null
    },
    {
      "id": "2135706",
      "postDate": "02/08/2023 20:18:42",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/virilo\" target=\"_blank\">@virilo</a>, I am very happy with your result!</p>\n<p><a href=\"https://www.kaggle.com/enric1296\" target=\"_blank\">@enric1296</a> <a href=\"https://www.kaggle.com/trasibulo\" target=\"_blank\">@trasibulo</a> and <a href=\"https://www.kaggle.com/albert2017\" target=\"_blank\">@albert2017</a></p>\n<p>You have had an excellent competition, always at the top of the leaderboard and progressing every day with an original and high-level solution.</p>\n<p>Enjoy the success everyone  ! </p>",
      "rawMarkdown": "Congratulations @virilo, I am very happy with your result!\n\n@enric1296 @trasibulo and @albert2017\n\nYou have had an excellent competition, always at the top of the leaderboard and progressing every day with an original and high-level solution.\n\nEnjoy the success everyone  !",
      "votes": null
    },
    {
      "id": "2136146",
      "postDate": "02/09/2023 05:51:46",
      "content": "<p>Hello, can you open source the code of BST?</p>",
      "rawMarkdown": "Hello, can you open source the code of BST?",
      "votes": null
    },
    {
      "id": "2136213",
      "postDate": "02/09/2023 06:59:18",
      "content": "<p>Thank you very much for your kind words <a href=\"https://www.kaggle.com/coreacasa\" target=\"_blank\">@coreacasa</a>!  It's highly appreciated.</p>",
      "rawMarkdown": "Thank you very much for your kind words @coreacasa!  It's highly appreciated.",
      "votes": null
    },
    {
      "id": "2136269",
      "postDate": "02/09/2023 07:56:01",
      "content": "<p><a href=\"https://www.kaggle.com/virilo\" target=\"_blank\">@virilo</a> Thank you for sharing the interesting approach ! Could you share the scores of Transformer model on Public LB and Private LB ?</p>",
      "rawMarkdown": "virilo Thank you for sharing the interesting approach ! Could you share the scores of Transformer model on Public LB and Private LB ?",
      "votes": null
    },
    {
      "id": "2139963",
      "postDate": "02/11/2023 10:21:04",
      "content": "<p>I didn't do that exercise too much in order to save submissions.  I measured a boost of 0.009 over some version of our candidates generator top 20 using orders only.</p>\n<p>I've to say that this is not a fair comparison, because top 20 of our candidates generator worked better when we used 100 candidates instead of 150.  Or even better optimizing it to get 20 candidates only.</p>\n<p>The only thing that we measured is that adding it to the ensemble was producing better results that not using it.</p>\n<p>For the firsts versions of the ensemble, it started appearing as the most important variable (xgb gain).  What surprised us a lot, when the rest of the stacked models had better scores.  IMHO It must be because diversity matters for the ensemble</p>",
      "rawMarkdown": "I didn't do that exercise too much in order to save submissions.  I measured a boost of 0.009 over some version of our candidates generator top 20 using orders only.\n\nI've to say that this is not a fair comparison, because top 20 of our candidates generator worked better when we used 100 candidates instead of 150.  Or even better optimizing it to get 20 candidates only.\n\nThe only thing that we measured is that adding it to the ensemble was producing better results that not using it.\n\nFor the firsts versions of the ensemble, it started appearing as the most important variable (xgb gain).  What surprised us a lot, when the rest of the stacked models had better scores.  IMHO It must be because diversity matters for the ensemble",
      "votes": null
    },
    {
      "id": "2142336",
      "postDate": "02/13/2023 13:11:47",
      "content": "<p>Your use of the BST transformer model and translation of the code from tensorflow to pytorch is impressive. It's great to see that even though gbdt algorithms performed better, the transformer solution still added diversity to the final ensemble. Your solution for the k-fold split issue by stratifying sessions in an iterative way is a smart approach to handle the memory error.</p>\n<p>May I ask, what do you think were the biggest challenges you faced during this competition and how did you overcome them? <a href=\"https://www.kaggle.com/virilo\" target=\"_blank\">@virilo</a> </p>",
      "rawMarkdown": "Your use of the BST transformer model and translation of the code from tensorflow to pytorch is impressive. It's great to see that even though gbdt algorithms performed better, the transformer solution still added diversity to the final ensemble. Your solution for the k-fold split issue by stratifying sessions in an iterative way is a smart approach to handle the memory error.\n\nMay I ask, what do you think were the biggest challenges you faced during this competition and how did you overcome them? @virilo",
      "votes": null
    },
    {
      "id": "2142377",
      "postDate": "02/13/2023 13:40:40",
      "content": "<p>Thank you for your answer. I think you can check the score of pure Transformer by Late Submission.</p>",
      "rawMarkdown": "Thank you for your answer. I think you can check the score of pure Transformer by Late Submission.",
      "votes": null
    },
    {
      "id": "2142856",
      "postDate": "02/13/2023 20:43:56",
      "content": "<p>Yes, I'm interested the same.  I'll share in a few days (sorry I'm quite busy).  Stay tuned.</p>",
      "rawMarkdown": "Yes, I'm interested the same.  I'll share in a few days (sorry I'm quite busy).  Stay tuned.",
      "votes": null
    },
    {
      "id": "2142866",
      "postDate": "02/13/2023 20:53:22",
      "content": "<p>One of the biggest chalenges was simply getting ride of bugs… haha.  It took me a week or so to release that I made this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/383566\" target=\"_blank\">mistake with the metric implementation</a>.  In the meanwhile I was getting higher validation scores than my team mates, but having lower increases when submitting to Kaggle.</p>\n<p>I had to revise and check everything and I spended a lot of time just asking myself \"why?\".  At least it helped me to improve the rest of the code until I found the actual cause.</p>",
      "rawMarkdown": "One of the biggest chalenges was simply getting ride of bugs... haha.  It took me a week or so to release that I made this [mistake with the metric implementation](https://www.kaggle.com/competitions/otto-recommender-system/discussion/383566).  In the meanwhile I was getting higher validation scores than my team mates, but having lower increases when submitting to Kaggle.\n\nI had to revise and check everything and I spended a lot of time just asking myself \"why?\".  At least it helped me to improve the rest of the code until I found the actual cause.",
      "votes": null
    },
    {
      "id": "2147729",
      "postDate": "02/16/2023 20:34:26",
      "content": "<p>I made a late submission.  It boosts the candidates generator top20 baseline in 0.00954 in the public leaderboard, and 0.0084 in the private leaderboard.</p>\n<p>It would have achieved a top-56 in public (0.58944) and a top-54 position in the private lb (0.58948) just by ranking the orders.</p>\n<p>To measure the complete model, I'd have to train it for carts and clicks tasks, but I wasn't able with my 32GB CPU RAM</p>\n<p>Anyway, the purpose of this model is just to add diversity by using transformers to learn about the behavior in the sequence, instead of using aggregations</p>",
      "rawMarkdown": "I made a late submission.  It boosts the candidates generator top20 baseline in 0.00954 in the public leaderboard, and 0.0084 in the private leaderboard.\n\nIt would have achieved a top-56 in public (0.58944) and a top-54 position in the private lb (0.58948) just by ranking the orders.\n\nTo measure the complete model, I'd have to train it for carts and clicks tasks, but I wasn't able with my 32GB CPU RAM\n\nAnyway, the purpose of this model is just to add diversity by using transformers to learn about the behavior in the sequence, instead of using aggregations",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2135706,
      "author_name": "coreacasa",
      "author_url": "",
      "post_date": "02/08/2023 20:18:42",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/virilo\" target=\"_blank\">@virilo</a>, I am very happy with your result!</p>\n<p><a href=\"https://www.kaggle.com/enric1296\" target=\"_blank\">@enric1296</a> <a href=\"https://www.kaggle.com/trasibulo\" target=\"_blank\">@trasibulo</a> and <a href=\"https://www.kaggle.com/albert2017\" target=\"_blank\">@albert2017</a></p>\n<p>You have had an excellent competition, always at the top of the leaderboard and progressing every day with an original and high-level solution.</p>\n<p>Enjoy the success everyone  ! </p>",
      "votes": null,
      "replies": [
        {
          "id": 2136213,
          "author_name": "virilo",
          "author_url": "",
          "post_date": "02/09/2023 06:59:18",
          "content": "<p>Thank you very much for your kind words <a href=\"https://www.kaggle.com/coreacasa\" target=\"_blank\">@coreacasa</a>!  It's highly appreciated.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2136146,
      "author_name": "yasso1",
      "author_url": "",
      "post_date": "02/09/2023 05:51:46",
      "content": "<p>Hello, can you open source the code of BST?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2136269,
      "author_name": "toshik",
      "author_url": "",
      "post_date": "02/09/2023 07:56:01",
      "content": "<p><a href=\"https://www.kaggle.com/virilo\" target=\"_blank\">@virilo</a> Thank you for sharing the interesting approach ! Could you share the scores of Transformer model on Public LB and Private LB ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2139963,
          "author_name": "virilo",
          "author_url": "",
          "post_date": "02/11/2023 10:21:04",
          "content": "<p>I didn't do that exercise too much in order to save submissions.  I measured a boost of 0.009 over some version of our candidates generator top 20 using orders only.</p>\n<p>I've to say that this is not a fair comparison, because top 20 of our candidates generator worked better when we used 100 candidates instead of 150.  Or even better optimizing it to get 20 candidates only.</p>\n<p>The only thing that we measured is that adding it to the ensemble was producing better results that not using it.</p>\n<p>For the firsts versions of the ensemble, it started appearing as the most important variable (xgb gain).  What surprised us a lot, when the rest of the stacked models had better scores.  IMHO It must be because diversity matters for the ensemble</p>",
          "votes": null,
          "replies": [
            {
              "id": 2142377,
              "author_name": "toshik",
              "author_url": "",
              "post_date": "02/13/2023 13:40:40",
              "content": "<p>Thank you for your answer. I think you can check the score of pure Transformer by Late Submission.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2142856,
                  "author_name": "virilo",
                  "author_url": "",
                  "post_date": "02/13/2023 20:43:56",
                  "content": "<p>Yes, I'm interested the same.  I'll share in a few days (sorry I'm quite busy).  Stay tuned.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 2147729,
          "author_name": "virilo",
          "author_url": "",
          "post_date": "02/16/2023 20:34:26",
          "content": "<p>I made a late submission.  It boosts the candidates generator top20 baseline in 0.00954 in the public leaderboard, and 0.0084 in the private leaderboard.</p>\n<p>It would have achieved a top-56 in public (0.58944) and a top-54 position in the private lb (0.58948) just by ranking the orders.</p>\n<p>To measure the complete model, I'd have to train it for carts and clicks tasks, but I wasn't able with my 32GB CPU RAM</p>\n<p>Anyway, the purpose of this model is just to add diversity by using transformers to learn about the behavior in the sequence, instead of using aggregations</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2142336,
      "author_name": "giranntu",
      "author_url": "",
      "post_date": "02/13/2023 13:11:47",
      "content": "<p>Your use of the BST transformer model and translation of the code from tensorflow to pytorch is impressive. It's great to see that even though gbdt algorithms performed better, the transformer solution still added diversity to the final ensemble. Your solution for the k-fold split issue by stratifying sessions in an iterative way is a smart approach to handle the memory error.</p>\n<p>May I ask, what do you think were the biggest challenges you faced during this competition and how did you overcome them? <a href=\"https://www.kaggle.com/virilo\" target=\"_blank\">@virilo</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 2142866,
          "author_name": "virilo",
          "author_url": "",
          "post_date": "02/13/2023 20:53:22",
          "content": "<p>One of the biggest chalenges was simply getting ride of bugs… haha.  It took me a week or so to release that I made this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/383566\" target=\"_blank\">mistake with the metric implementation</a>.  In the meanwhile I was getting higher validation scores than my team mates, but having lower increases when submitting to Kaggle.</p>\n<p>I had to revise and check everything and I spended a lot of time just asking myself \"why?\".  At least it helped me to improve the rest of the code until I found the actual cause.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2134341": "Thank you very much for organizing this competition.  \n\nAnd thank you very much to my teammates @enric1296 @trasibulo and @albert2017  We really enjoyed this competition!\n\nCongratulations to my teammates and other competitors who achieved their Master tier in this competition: well done! well deserved!!!\n\nThis post is about one part of the final solution. @enric1296 posted [the overall solution](https://www.kaggle.com/competitions/otto-recommender-system/discussion/383229) \n\nI'd like to share here the main part of my contribution.  It was based on BST transformer model ([arXiv:1905.06874](https://arxiv.org/abs/1905.06874))\n\nThe first time we used that solution was for the Kaggle Days Barcelona 11 hour datathon. That time Enric was the one doing the BST and I was in charge of feature engineering.\n\nFor this competition I used all my teammates model inputs and my own ones to feed the model.\n\nAnd the first step was to translate Enric's datathon code from tensorflow to pytorch 😜\n\nEven when gbdt algorithms turned out to perform better for this problem, the transformer solution brought more diversity to the final ensemble.\n\nThis is how articles positional features for the transformer looked like:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F223160%2F7a7c9a7ac0cbf9c9434a8bed41095ea1%2Ffig-1.png?generation=1675810715439219&alt=media)\n\nAnd this was the scheme for the binary classifier reranker:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F223160%2F28c15d49f195ff344d01f4005f8e8969%2Ffig-2.png?generation=1675810742844854&alt=media)\n\n## the k-fold split\n\nIf you tried to use StratifiedGroupKFold to stratify all the items from the ground truth, you easily run into a \"MemoryError: Unable to allocate 168. GiB for an array with shape ...\"\n\nWe solved it in an iterative way:\n\n\t- stratify sessions containing less represented items for orders\n\t- stratify sessions containing the rest of the items for orders\n\t- stratify sessions containing less represented items for carts\n\t...",
    "2135706": "Congratulations @virilo, I am very happy with your result!\n\n@enric1296 @trasibulo and @albert2017\n\nYou have had an excellent competition, always at the top of the leaderboard and progressing every day with an original and high-level solution.\n\nEnjoy the success everyone  !",
    "2136146": "Hello, can you open source the code of BST?",
    "2136213": "Thank you very much for your kind words @coreacasa!  It's highly appreciated.",
    "2136269": "virilo Thank you for sharing the interesting approach ! Could you share the scores of Transformer model on Public LB and Private LB ?",
    "2139963": "I didn't do that exercise too much in order to save submissions.  I measured a boost of 0.009 over some version of our candidates generator top 20 using orders only.\n\nI've to say that this is not a fair comparison, because top 20 of our candidates generator worked better when we used 100 candidates instead of 150.  Or even better optimizing it to get 20 candidates only.\n\nThe only thing that we measured is that adding it to the ensemble was producing better results that not using it.\n\nFor the firsts versions of the ensemble, it started appearing as the most important variable (xgb gain).  What surprised us a lot, when the rest of the stacked models had better scores.  IMHO It must be because diversity matters for the ensemble",
    "2142336": "Your use of the BST transformer model and translation of the code from tensorflow to pytorch is impressive. It's great to see that even though gbdt algorithms performed better, the transformer solution still added diversity to the final ensemble. Your solution for the k-fold split issue by stratifying sessions in an iterative way is a smart approach to handle the memory error.\n\nMay I ask, what do you think were the biggest challenges you faced during this competition and how did you overcome them? @virilo",
    "2142377": "Thank you for your answer. I think you can check the score of pure Transformer by Late Submission.",
    "2142856": "Yes, I'm interested the same.  I'll share in a few days (sorry I'm quite busy).  Stay tuned.",
    "2142866": "One of the biggest chalenges was simply getting ride of bugs... haha.  It took me a week or so to release that I made this [mistake with the metric implementation](https://www.kaggle.com/competitions/otto-recommender-system/discussion/383566).  In the meanwhile I was getting higher validation scores than my team mates, but having lower increases when submitting to Kaggle.\n\nI had to revise and check everything and I spended a lot of time just asking myself \"why?\".  At least it helped me to improve the rest of the code until I found the actual cause.",
    "2147729": "I made a late submission.  It boosts the candidates generator top20 baseline in 0.00954 in the public leaderboard, and 0.0084 in the private leaderboard.\n\nIt would have achieved a top-56 in public (0.58944) and a top-54 position in the private lb (0.58948) just by ranking the orders.\n\nTo measure the complete model, I'd have to train it for carts and clicks tasks, but I wasn't able with my 32GB CPU RAM\n\nAnyway, the purpose of this model is just to add diversity by using transformers to learn about the behavior in the sequence, instead of using aggregations"
  },
  "source": "meta"
}