{
  "id": 154335,
  "title": "Best Roberta-XLM model",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/154335",
  "author_name": "",
  "post_date": "2020-05-28T03:32:29.548469800Z",
  "votes": 15,
  "comment_count": 38,
  "views": 0,
  "content": "<p>Just wanted to know how well single models of roberta-xlm doing for other people.\nMy team has managed to get a xlm model which does .9421. Would love to get some pointers on how to improve this score/</p>",
  "messages": [
    {
      "id": "864508",
      "postDate": "05/28/2020 03:32:29",
      "content": "<p>Just wanted to know how well single models of roberta-xlm doing for other people.\nMy team has managed to get a xlm model which does .9421. Would love to get some pointers on how to improve this score/</p>",
      "rawMarkdown": "Just wanted to know how well single models of roberta-xlm doing for other people.\nMy team has managed to get a xlm model which does .9421. Would love to get some pointers on how to improve this score/",
      "votes": null
    },
    {
      "id": "864531",
      "postDate": "05/28/2020 03:57:35",
      "content": "<p>If I understand correctly, so your best single model is 0.9421 but your LB is 0.9482? That's very interesting ...</p>",
      "rawMarkdown": "If I understand correctly, so your best single model is 0.9421 but your LB is 0.9482? That's very interesting ...",
      "votes": null
    },
    {
      "id": "864572",
      "postDate": "05/28/2020 04:30:37",
      "content": "<p>Yup, thats correct</p>",
      "rawMarkdown": "Yup, thats correct",
      "votes": null
    },
    {
      "id": "864718",
      "postDate": "05/28/2020 06:42:32",
      "content": "<p>Interesting, I got higher single model score but lower ensemble. I’m curious how you do ensemble.</p>\n\n<p>In terms of how to improve single models, I would think the suggestions in the following threads are worth reading / trying.\n<a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/141825\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/141825</a>\n<a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/151084\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/151084</a></p>",
      "rawMarkdown": "Interesting, I got higher single model score but lower ensemble. I’m curious how you do ensemble.\n\nIn terms of how to improve single models, I would think the suggestions in the following threads are worth reading / trying.\nhttps://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/141825\nhttps://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/151084",
      "votes": null
    },
    {
      "id": "864735",
      "postDate": "05/28/2020 06:52:40",
      "content": "<p>My best roberta-xlm is .9411.</p>",
      "rawMarkdown": "My best roberta-xlm is .9411.",
      "votes": null
    },
    {
      "id": "864742",
      "postDate": "05/28/2020 06:55:15",
      "content": "<p>What do you refer to as single model when you say it scores .9421?</p>",
      "rawMarkdown": "What do you refer to as single model when you say it scores .9421?",
      "votes": null
    },
    {
      "id": "864842",
      "postDate": "05/28/2020 08:14:53",
      "content": "<p>Share the same feeling and question. So far, I was trying to get a better single model score as my ensembling are not going above 9460 even if my best model is 9425.</p>\n\n<p>I guess I need to look deeper in ensembling !\nBy the way, I will be happy to team up with someone with good ensembling skills 😉   </p>",
      "rawMarkdown": "Share the same feeling and question. So far, I was trying to get a better single model score as my ensembling are not going above 9460 even if my best model is 9425.\n\nI guess I need to look deeper in ensembling !\nBy the way, I will be happy to team up with someone with good ensembling skills 😉",
      "votes": null
    },
    {
      "id": "864846",
      "postDate": "05/28/2020 08:18:32",
      "content": "<p>\"My team has managed to get a xlm model which does .9421.\"</p>\n\n<p>Why is \"your team\", not \"you\" :)</p>",
      "rawMarkdown": "\"My team has managed to get a xlm model which does .9421.\"\n\nWhy is \"your team\", not \"you\" :)",
      "votes": null
    },
    {
      "id": "864867",
      "postDate": "05/28/2020 08:31:57",
      "content": "<p>It's a typical Rick move 😄 </p>",
      "rawMarkdown": "It's a typical Rick move 😄",
      "votes": null
    },
    {
      "id": "864930",
      "postDate": "05/28/2020 09:25:11",
      "content": "<p><a href=\"/mcggood\">@mcggood</a>  our single xlmr got 0.9436 but how do you got to 4th place? large ensemble? stacking ? pseudo labeling?</p>",
      "rawMarkdown": "mcggood  our single xlmr got 0.9436 but how do you got to 4th place? large ensemble? stacking ? pseudo labeling?",
      "votes": null
    },
    {
      "id": "865596",
      "postDate": "05/28/2020 18:17:32",
      "content": "<p><a href=\"/mcggood\">@mcggood</a> Haha, nothing like that, my teammate and I haven't merged yet is all. We will merge before the final deadline :)</p>",
      "rawMarkdown": "mcggood Haha, nothing like that, my teammate and I haven't merged yet is all. We will merge before the final deadline :)",
      "votes": null
    },
    {
      "id": "865598",
      "postDate": "05/28/2020 18:18:49",
      "content": "<p><a href=\"/luohongchen1993\">@luohongchen1993</a> Thanks for the comments, I'll definitely check them out. There is really nothing special about my ensembles, me and my teammate (we haven't merged yet) just kept blending models to a particular, well performing submission</p>",
      "rawMarkdown": "luohongchen1993 Thanks for the comments, I'll definitely check them out. There is really nothing special about my ensembles, me and my teammate (we haven't merged yet) just kept blending models to a particular, well performing submission",
      "votes": null
    },
    {
      "id": "865609",
      "postDate": "05/28/2020 18:25:41",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a> I mean a submission without ensembling, stacking, blending etc. Just one model with whatever CV, KFold, training pipelines.</p>",
      "rawMarkdown": "philippsinger I mean a submission without ensembling, stacking, blending etc. Just one model with whatever CV, KFold, training pipelines.",
      "votes": null
    },
    {
      "id": "865611",
      "postDate": "05/28/2020 18:27:24",
      "content": "<p>Ok, but still k-fold blending I assume?</p>",
      "rawMarkdown": "Ok, but still k-fold blending I assume?",
      "votes": null
    },
    {
      "id": "865615",
      "postDate": "05/28/2020 18:31:28",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Not quite sure what you are trying to say...</p>",
      "rawMarkdown": "philippsinger Not quite sure what you are trying to say...",
      "votes": null
    },
    {
      "id": "865621",
      "postDate": "05/28/2020 18:35:05",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a>  how k-fold can be blending? it's  used to evaluate a single model, it's not ensemble</p>",
      "rawMarkdown": "philippsinger  how k-fold can be blending? it's  used to evaluate a single model, it's not ensemble",
      "votes": null
    },
    {
      "id": "865624",
      "postDate": "05/28/2020 18:36:50",
      "content": "<p>I think he is refering to the submission being an ensemble of the k-fold predictions (perhaps a simple average of the k predictions made by each fold). <a href=\"/amoghjrules\">@amoghjrules</a> </p>",
      "rawMarkdown": "I think he is refering to the submission being an ensemble of the k-fold predictions (perhaps a simple average of the k predictions made by each fold). @amoghjrules",
      "votes": null
    },
    {
      "id": "865627",
      "postDate": "05/28/2020 18:39:07",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "865639",
      "postDate": "05/28/2020 18:56:43",
      "content": "<p>That doesn't clarify though how it can be a single model if you are using CV. Look, if you use CV, you have for example 5-folds, so actually 5 fits. How do you generate test preds? Usually by predicting 5 times and averaging.</p>",
      "rawMarkdown": "That doesn't clarify though how it can be a single model if you are using CV. Look, if you use CV, you have for example 5-folds, so actually 5 fits. How do you generate test preds? Usually by predicting 5 times and averaging.",
      "votes": null
    },
    {
      "id": "865644",
      "postDate": "05/28/2020 19:06:31",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Shoot, yeah that's fine, I didn't think it through, my bad</p>",
      "rawMarkdown": "philippsinger Shoot, yeah that's fine, I didn't think it through, my bad",
      "votes": null
    },
    {
      "id": "865884",
      "postDate": "05/29/2020 00:33:57",
      "content": "<p>Why your best xml model get so high? Mine only gets 9.389 XD, \nany tips plz.\nLooking forward and appreciate</p>",
      "rawMarkdown": "Why your best xml model get so high? Mine only gets 9.389 XD, \nany tips plz.\nLooking forward and appreciate",
      "votes": null
    },
    {
      "id": "866042",
      "postDate": "05/29/2020 04:20:38",
      "content": "<p><a href=\"/kannelliu\">@kannelliu</a> Try to use more data, experiment with translated data, experiment with different learning rates and validation techniques, then try pseudolabelling. This should definitely give some improvement in score :)</p>",
      "rawMarkdown": "kannelliu Try to use more data, experiment with translated data, experiment with different learning rates and validation techniques, then try pseudolabelling. This should definitely give some improvement in score :)",
      "votes": null
    },
    {
      "id": "867008",
      "postDate": "05/29/2020 22:21:05",
      "content": "<p>This is private sharing then <a href=\"/amoghjrules\">@amoghjrules</a>. You are also exploiting max subs by doing that.</p>",
      "rawMarkdown": "This is private sharing then @amoghjrules. You are also exploiting max subs by doing that.",
      "votes": null
    },
    {
      "id": "867204",
      "postDate": "05/30/2020 04:54:19",
      "content": "<p>Are you doing k-fold? if you're I won't label it as \"single model\".</p>",
      "rawMarkdown": "Are you doing k-fold? if you're I won't label it as \"single model\".",
      "votes": null
    },
    {
      "id": "867358",
      "postDate": "05/30/2020 07:48:47",
      "content": "<p>My best model is .9425.\nI fixed this model <a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta</a>\n1. Change learning rate per epoch.\n2. Summing epoch2 and epoch3 validation data.</p>",
      "rawMarkdown": "My best model is .9425.\nI fixed this model https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\n1. Change learning rate per epoch.\n2. Summing epoch2 and epoch3 validation data.",
      "votes": null
    },
    {
      "id": "867499",
      "postDate": "05/30/2020 11:19:27",
      "content": "<p>Yup, thats correct, I had this discussion with psi</p>",
      "rawMarkdown": "Yup, thats correct, I had this discussion with psi",
      "votes": null
    },
    {
      "id": "867500",
      "postDate": "05/30/2020 11:19:51",
      "content": "<p>Thanks for your help, I'll try it out :)</p>",
      "rawMarkdown": "Thanks for your help, I'll try it out :)",
      "votes": null
    },
    {
      "id": "867504",
      "postDate": "05/30/2020 11:23:03",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Nope, kaggle does not allow teams to merge if they are above the max number of subs. We have merged now. The only benefit is max subs per day increases for a short amount of time\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1274202%2Ff9de04523090a2d116fa71a97e5c52d3%2Frules.png?generation=1590837772802101&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "philippsinger Nope, kaggle does not allow teams to merge if they are above the max number of subs. We have merged now. The only benefit is max subs per day increases for a short amount of time\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1274202%2Ff9de04523090a2d116fa71a97e5c52d3%2Frules.png?generation=1590837772802101&amp;alt=media)",
      "votes": null
    },
    {
      "id": "867783",
      "postDate": "05/30/2020 16:09:08",
      "content": "<p>Current score is 0.9427. I change lr.</p>",
      "rawMarkdown": "Current score is 0.9427. I change lr.",
      "votes": null
    },
    {
      "id": "868073",
      "postDate": "05/30/2020 21:53:30",
      "content": "<p>Can you give more details about which parts you have changed? Thank you.</p>",
      "rawMarkdown": "Can you give more details about which parts you have changed? Thank you.",
      "votes": null
    },
    {
      "id": "868113",
      "postDate": "05/31/2020 00:06:20",
      "content": "<p>In my understanding, this is still not allowed by the rules, however. See <a href=\"https://www.kaggle.com/c/abstraction-and-reasoning-challenge/discussion/150294\">https://www.kaggle.com/c/abstraction-and-reasoning-challenge/discussion/150294</a> for a bit of a discussion, but the two individuals involved were later removed from the leaderboard at the end of the competition, even though they teamed up after they were found out. </p>",
      "rawMarkdown": "In my understanding, this is still not allowed by the rules, however. See https://www.kaggle.com/c/abstraction-and-reasoning-challenge/discussion/150294 for a bit of a discussion, but the two individuals involved were later removed from the leaderboard at the end of the competition, even though they teamed up after they were found out.",
      "votes": null
    },
    {
      "id": "868208",
      "postDate": "05/31/2020 03:49:48",
      "content": "<p>Thanks for sharing <a href=\"/frankrosenblatt\">@frankrosenblatt</a>. What do you mean by \"2. Summing epoch2 and epoch3 validation data.\" ?\nMy best single model PL -  0.9443</p>",
      "rawMarkdown": "Thanks for sharing @frankrosenblatt. What do you mean by \"2. Summing epoch2 and epoch3 validation data.\" ?\nMy best single model PL -  0.9443",
      "votes": null
    },
    {
      "id": "868580",
      "postDate": "05/31/2020 10:59:21",
      "content": "<p>Oh wow... We totally did not have an idea about this. We both are new to competitions in fact this is one of the first competitions that I've teamed up for. This is an honest mistake, we discussed this issue and now have decided to take it up with Kaggle admins to ask them for a recourse.</p>",
      "rawMarkdown": "Oh wow... We totally did not have an idea about this. We both are new to competitions in fact this is one of the first competitions that I've teamed up for. This is an honest mistake, we discussed this issue and now have decided to take it up with Kaggle admins to ask them for a recourse.",
      "votes": null
    },
    {
      "id": "869089",
      "postDate": "05/31/2020 18:20:50",
      "content": "<p>How many epochs did you trained?</p>",
      "rawMarkdown": "How many epochs did you trained?",
      "votes": null
    },
    {
      "id": "870193",
      "postDate": "06/01/2020 14:50:15",
      "content": "<p>Hi <a href=\"/frankrosenblatt\">@frankrosenblatt</a>, thanks for sharing your insights. I'm also very curious, what do you mean with summing up epoch2 and epoch3 validation data? </p>",
      "rawMarkdown": "Hi @frankrosenblatt, thanks for sharing your insights. I'm also very curious, what do you mean with summing up epoch2 and epoch3 validation data?",
      "votes": null
    },
    {
      "id": "873439",
      "postDate": "06/04/2020 06:37:49",
      "content": "<p>My best XLM-Roberta model is 0.9418 !\nTraining on translated data with a learning rate scheduler really helped to improve my score.</p>",
      "rawMarkdown": "My best XLM-Roberta model is 0.9418 !\nTraining on translated data with a learning rate scheduler really helped to improve my score.",
      "votes": null
    },
    {
      "id": "873621",
      "postDate": "06/04/2020 10:01:52",
      "content": "<p>Predict score use epoch2 model and epoch3 and sum both.</p>",
      "rawMarkdown": "Predict score use epoch2 model and epoch3 and sum both.",
      "votes": null
    },
    {
      "id": "873718",
      "postDate": "06/04/2020 11:35:23",
      "content": "<p>Thanks <a href=\"/frankrosenblatt\">@frankrosenblatt</a>  you for explaining. Do you average each epoch's model parameters (weights) - similar to Stochastic Weight Averaging(SWA) procedure in <a href=\"https://arxiv.org/pdf/1803.05407.pdf\">https://arxiv.org/pdf/1803.05407.pdf</a>\nOR do you predict test data after epochs2 , 3 and then average predictions ?</p>\n\n<p>I have not done any of those yet. Hope to do it by  the end of competition. </p>",
      "rawMarkdown": "Thanks @frankrosenblatt  you for explaining. Do you average each epoch's model parameters (weights) - similar to Stochastic Weight Averaging(SWA) procedure in https://arxiv.org/pdf/1803.05407.pdf\nOR do you predict test data after epochs2 , 3 and then average predictions ?\n\n I have not done any of those yet. Hope to do it by  the end of competition.",
      "votes": null
    },
    {
      "id": "878614",
      "postDate": "06/08/2020 17:08:03",
      "content": "<p>My best is 0.9458 with a single model making use of translated data and 5-fold oof predictions. Cant seem to improve the score much with ensemble though, any tips?</p>",
      "rawMarkdown": "My best is 0.9458 with a single model making use of translated data and 5-fold oof predictions. Cant seem to improve the score much with ensemble though, any tips?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 864531,
      "author_name": "luohongchen1993",
      "author_url": "",
      "post_date": "05/28/2020 03:57:35",
      "content": "<p>If I understand correctly, so your best single model is 0.9421 but your LB is 0.9482? That's very interesting ...</p>",
      "votes": null,
      "replies": [
        {
          "id": 864572,
          "author_name": "amoghjrules",
          "author_url": "",
          "post_date": "05/28/2020 04:30:37",
          "content": "<p>Yup, thats correct</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 864718,
          "author_name": "luohongchen1993",
          "author_url": "",
          "post_date": "05/28/2020 06:42:32",
          "content": "<p>Interesting, I got higher single model score but lower ensemble. I’m curious how you do ensemble.</p>\n\n<p>In terms of how to improve single models, I would think the suggestions in the following threads are worth reading / trying.\n<a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/141825\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/141825</a>\n<a href=\"https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/151084\">https://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/151084</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 864842,
          "author_name": "seb6084",
          "author_url": "",
          "post_date": "05/28/2020 08:14:53",
          "content": "<p>Share the same feeling and question. So far, I was trying to get a better single model score as my ensembling are not going above 9460 even if my best model is 9425.</p>\n\n<p>I guess I need to look deeper in ensembling !\nBy the way, I will be happy to team up with someone with good ensembling skills 😉   </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 865598,
          "author_name": "amoghjrules",
          "author_url": "",
          "post_date": "05/28/2020 18:18:49",
          "content": "<p><a href=\"/luohongchen1993\">@luohongchen1993</a> Thanks for the comments, I'll definitely check them out. There is really nothing special about my ensembles, me and my teammate (we haven't merged yet) just kept blending models to a particular, well performing submission</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 864735,
      "author_name": "dandrocec",
      "author_url": "",
      "post_date": "05/28/2020 06:52:40",
      "content": "<p>My best roberta-xlm is .9411.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 864742,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "05/28/2020 06:55:15",
      "content": "<p>What do you refer to as single model when you say it scores .9421?</p>",
      "votes": null,
      "replies": [
        {
          "id": 865609,
          "author_name": "amoghjrules",
          "author_url": "",
          "post_date": "05/28/2020 18:25:41",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> I mean a submission without ensembling, stacking, blending etc. Just one model with whatever CV, KFold, training pipelines.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 865611,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "05/28/2020 18:27:24",
          "content": "<p>Ok, but still k-fold blending I assume?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 865615,
          "author_name": "amoghjrules",
          "author_url": "",
          "post_date": "05/28/2020 18:31:28",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Not quite sure what you are trying to say...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 865621,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "05/28/2020 18:35:05",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a>  how k-fold can be blending? it's  used to evaluate a single model, it's not ensemble</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 865624,
          "author_name": "luohongchen1993",
          "author_url": "",
          "post_date": "05/28/2020 18:36:50",
          "content": "<p>I think he is refering to the submission being an ensemble of the k-fold predictions (perhaps a simple average of the k predictions made by each fold). <a href=\"/amoghjrules\">@amoghjrules</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 865627,
          "author_name": "amoghjrules",
          "author_url": "",
          "post_date": "05/28/2020 18:39:07",
          "content": "",
          "votes": null,
          "replies": []
        },
        {
          "id": 865639,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "05/28/2020 18:56:43",
          "content": "<p>That doesn't clarify though how it can be a single model if you are using CV. Look, if you use CV, you have for example 5-folds, so actually 5 fits. How do you generate test preds? Usually by predicting 5 times and averaging.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 865644,
          "author_name": "amoghjrules",
          "author_url": "",
          "post_date": "05/28/2020 19:06:31",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Shoot, yeah that's fine, I didn't think it through, my bad</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 864846,
      "author_name": "mcggood",
      "author_url": "",
      "post_date": "05/28/2020 08:18:32",
      "content": "<p>\"My team has managed to get a xlm model which does .9421.\"</p>\n\n<p>Why is \"your team\", not \"you\" :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 864867,
          "author_name": "rafiko1",
          "author_url": "",
          "post_date": "05/28/2020 08:31:57",
          "content": "<p>It's a typical Rick move 😄 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 864930,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "05/28/2020 09:25:11",
          "content": "<p><a href=\"/mcggood\">@mcggood</a>  our single xlmr got 0.9436 but how do you got to 4th place? large ensemble? stacking ? pseudo labeling?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 865596,
          "author_name": "amoghjrules",
          "author_url": "",
          "post_date": "05/28/2020 18:17:32",
          "content": "<p><a href=\"/mcggood\">@mcggood</a> Haha, nothing like that, my teammate and I haven't merged yet is all. We will merge before the final deadline :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 867008,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "05/29/2020 22:21:05",
          "content": "<p>This is private sharing then <a href=\"/amoghjrules\">@amoghjrules</a>. You are also exploiting max subs by doing that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 867504,
          "author_name": "amoghjrules",
          "author_url": "",
          "post_date": "05/30/2020 11:23:03",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> Nope, kaggle does not allow teams to merge if they are above the max number of subs. We have merged now. The only benefit is max subs per day increases for a short amount of time\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1274202%2Ff9de04523090a2d116fa71a97e5c52d3%2Frules.png?generation=1590837772802101&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 868113,
          "author_name": "misfyre",
          "author_url": "",
          "post_date": "05/31/2020 00:06:20",
          "content": "<p>In my understanding, this is still not allowed by the rules, however. See <a href=\"https://www.kaggle.com/c/abstraction-and-reasoning-challenge/discussion/150294\">https://www.kaggle.com/c/abstraction-and-reasoning-challenge/discussion/150294</a> for a bit of a discussion, but the two individuals involved were later removed from the leaderboard at the end of the competition, even though they teamed up after they were found out. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 868580,
          "author_name": "amoghjrules",
          "author_url": "",
          "post_date": "05/31/2020 10:59:21",
          "content": "<p>Oh wow... We totally did not have an idea about this. We both are new to competitions in fact this is one of the first competitions that I've teamed up for. This is an honest mistake, we discussed this issue and now have decided to take it up with Kaggle admins to ask them for a recourse.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 865884,
      "author_name": "kannelliu",
      "author_url": "",
      "post_date": "05/29/2020 00:33:57",
      "content": "<p>Why your best xml model get so high? Mine only gets 9.389 XD, \nany tips plz.\nLooking forward and appreciate</p>",
      "votes": null,
      "replies": [
        {
          "id": 866042,
          "author_name": "amoghjrules",
          "author_url": "",
          "post_date": "05/29/2020 04:20:38",
          "content": "<p><a href=\"/kannelliu\">@kannelliu</a> Try to use more data, experiment with translated data, experiment with different learning rates and validation techniques, then try pseudolabelling. This should definitely give some improvement in score :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 867204,
      "author_name": "shahules",
      "author_url": "",
      "post_date": "05/30/2020 04:54:19",
      "content": "<p>Are you doing k-fold? if you're I won't label it as \"single model\".</p>",
      "votes": null,
      "replies": [
        {
          "id": 867499,
          "author_name": "amoghjrules",
          "author_url": "",
          "post_date": "05/30/2020 11:19:27",
          "content": "<p>Yup, thats correct, I had this discussion with psi</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 867358,
      "author_name": "frankrosenblatt",
      "author_url": "",
      "post_date": "05/30/2020 07:48:47",
      "content": "<p>My best model is .9425.\nI fixed this model <a href=\"https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\">https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta</a>\n1. Change learning rate per epoch.\n2. Summing epoch2 and epoch3 validation data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 867500,
          "author_name": "amoghjrules",
          "author_url": "",
          "post_date": "05/30/2020 11:19:51",
          "content": "<p>Thanks for your help, I'll try it out :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 867783,
          "author_name": "frankrosenblatt",
          "author_url": "",
          "post_date": "05/30/2020 16:09:08",
          "content": "<p>Current score is 0.9427. I change lr.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 868073,
          "author_name": "quanncore",
          "author_url": "",
          "post_date": "05/30/2020 21:53:30",
          "content": "<p>Can you give more details about which parts you have changed? Thank you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 868208,
          "author_name": "isakev",
          "author_url": "",
          "post_date": "05/31/2020 03:49:48",
          "content": "<p>Thanks for sharing <a href=\"/frankrosenblatt\">@frankrosenblatt</a>. What do you mean by \"2. Summing epoch2 and epoch3 validation data.\" ?\nMy best single model PL -  0.9443</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 869089,
          "author_name": "jay0606",
          "author_url": "",
          "post_date": "05/31/2020 18:20:50",
          "content": "<p>How many epochs did you trained?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 870193,
          "author_name": "carlolepelaars",
          "author_url": "",
          "post_date": "06/01/2020 14:50:15",
          "content": "<p>Hi <a href=\"/frankrosenblatt\">@frankrosenblatt</a>, thanks for sharing your insights. I'm also very curious, what do you mean with summing up epoch2 and epoch3 validation data? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 873621,
          "author_name": "frankrosenblatt",
          "author_url": "",
          "post_date": "06/04/2020 10:01:52",
          "content": "<p>Predict score use epoch2 model and epoch3 and sum both.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 873718,
          "author_name": "isakev",
          "author_url": "",
          "post_date": "06/04/2020 11:35:23",
          "content": "<p>Thanks <a href=\"/frankrosenblatt\">@frankrosenblatt</a>  you for explaining. Do you average each epoch's model parameters (weights) - similar to Stochastic Weight Averaging(SWA) procedure in <a href=\"https://arxiv.org/pdf/1803.05407.pdf\">https://arxiv.org/pdf/1803.05407.pdf</a>\nOR do you predict test data after epochs2 , 3 and then average predictions ?</p>\n\n<p>I have not done any of those yet. Hope to do it by  the end of competition. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 873439,
      "author_name": "alansun17904",
      "author_url": "",
      "post_date": "06/04/2020 06:37:49",
      "content": "<p>My best XLM-Roberta model is 0.9418 !\nTraining on translated data with a learning rate scheduler really helped to improve my score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 878614,
      "author_name": "hansungj",
      "author_url": "",
      "post_date": "06/08/2020 17:08:03",
      "content": "<p>My best is 0.9458 with a single model making use of translated data and 5-fold oof predictions. Cant seem to improve the score much with ensemble though, any tips?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "864508": "Just wanted to know how well single models of roberta-xlm doing for other people.\nMy team has managed to get a xlm model which does .9421. Would love to get some pointers on how to improve this score/",
    "864531": "If I understand correctly, so your best single model is 0.9421 but your LB is 0.9482? That's very interesting ...",
    "864572": "Yup, thats correct",
    "864718": "Interesting, I got higher single model score but lower ensemble. I’m curious how you do ensemble.\n\nIn terms of how to improve single models, I would think the suggestions in the following threads are worth reading / trying.\nhttps://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/141825\nhttps://www.kaggle.com/c/jigsaw-multilingual-toxic-comment-classification/discussion/151084",
    "864735": "My best roberta-xlm is .9411.",
    "864742": "What do you refer to as single model when you say it scores .9421?",
    "864842": "Share the same feeling and question. So far, I was trying to get a better single model score as my ensembling are not going above 9460 even if my best model is 9425.\n\nI guess I need to look deeper in ensembling !\nBy the way, I will be happy to team up with someone with good ensembling skills 😉",
    "864846": "\"My team has managed to get a xlm model which does .9421.\"\n\nWhy is \"your team\", not \"you\" :)",
    "864867": "It's a typical Rick move 😄",
    "864930": "mcggood  our single xlmr got 0.9436 but how do you got to 4th place? large ensemble? stacking ? pseudo labeling?",
    "865596": "mcggood Haha, nothing like that, my teammate and I haven't merged yet is all. We will merge before the final deadline :)",
    "865598": "luohongchen1993 Thanks for the comments, I'll definitely check them out. There is really nothing special about my ensembles, me and my teammate (we haven't merged yet) just kept blending models to a particular, well performing submission",
    "865609": "philippsinger I mean a submission without ensembling, stacking, blending etc. Just one model with whatever CV, KFold, training pipelines.",
    "865611": "Ok, but still k-fold blending I assume?",
    "865615": "philippsinger Not quite sure what you are trying to say...",
    "865621": "philippsinger  how k-fold can be blending? it's  used to evaluate a single model, it's not ensemble",
    "865624": "I think he is refering to the submission being an ensemble of the k-fold predictions (perhaps a simple average of the k predictions made by each fold). @amoghjrules",
    "865627": "",
    "865639": "That doesn't clarify though how it can be a single model if you are using CV. Look, if you use CV, you have for example 5-folds, so actually 5 fits. How do you generate test preds? Usually by predicting 5 times and averaging.",
    "865644": "philippsinger Shoot, yeah that's fine, I didn't think it through, my bad",
    "865884": "Why your best xml model get so high? Mine only gets 9.389 XD, \nany tips plz.\nLooking forward and appreciate",
    "866042": "kannelliu Try to use more data, experiment with translated data, experiment with different learning rates and validation techniques, then try pseudolabelling. This should definitely give some improvement in score :)",
    "867008": "This is private sharing then @amoghjrules. You are also exploiting max subs by doing that.",
    "867204": "Are you doing k-fold? if you're I won't label it as \"single model\".",
    "867358": "My best model is .9425.\nI fixed this model https://www.kaggle.com/shonenkov/tpu-training-super-fast-xlmroberta\n1. Change learning rate per epoch.\n2. Summing epoch2 and epoch3 validation data.",
    "867499": "Yup, thats correct, I had this discussion with psi",
    "867500": "Thanks for your help, I'll try it out :)",
    "867504": "philippsinger Nope, kaggle does not allow teams to merge if they are above the max number of subs. We have merged now. The only benefit is max subs per day increases for a short amount of time\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1274202%2Ff9de04523090a2d116fa71a97e5c52d3%2Frules.png?generation=1590837772802101&amp;alt=media)",
    "867783": "Current score is 0.9427. I change lr.",
    "868073": "Can you give more details about which parts you have changed? Thank you.",
    "868113": "In my understanding, this is still not allowed by the rules, however. See https://www.kaggle.com/c/abstraction-and-reasoning-challenge/discussion/150294 for a bit of a discussion, but the two individuals involved were later removed from the leaderboard at the end of the competition, even though they teamed up after they were found out.",
    "868208": "Thanks for sharing @frankrosenblatt. What do you mean by \"2. Summing epoch2 and epoch3 validation data.\" ?\nMy best single model PL -  0.9443",
    "868580": "Oh wow... We totally did not have an idea about this. We both are new to competitions in fact this is one of the first competitions that I've teamed up for. This is an honest mistake, we discussed this issue and now have decided to take it up with Kaggle admins to ask them for a recourse.",
    "869089": "How many epochs did you trained?",
    "870193": "Hi @frankrosenblatt, thanks for sharing your insights. I'm also very curious, what do you mean with summing up epoch2 and epoch3 validation data?",
    "873439": "My best XLM-Roberta model is 0.9418 !\nTraining on translated data with a learning rate scheduler really helped to improve my score.",
    "873621": "Predict score use epoch2 model and epoch3 and sum both.",
    "873718": "Thanks @frankrosenblatt  you for explaining. Do you average each epoch's model parameters (weights) - similar to Stochastic Weight Averaging(SWA) procedure in https://arxiv.org/pdf/1803.05407.pdf\nOR do you predict test data after epochs2 , 3 and then average predictions ?\n\n I have not done any of those yet. Hope to do it by  the end of competition.",
    "878614": "My best is 0.9458 with a single model making use of translated data and 5-fold oof predictions. Cant seem to improve the score much with ensemble though, any tips?"
  },
  "source": "meta"
}