{
  "id": 160961,
  "title": " 11th place solution(Hui Qin's)",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/writeups/qin-zhang-ma-qiao-meng-11th-place-solution-hui-qin",
  "author_name": "",
  "post_date": "2020-06-23T09:04:49.807898400Z",
  "votes": 24,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Congratulations all winners.  Thanks Kaggle supply this funny competition.  Thanks my teammates @Yang Zhang @MaChaogong @Hikkiiiiiiiii @Morphy</p>\n\n<p>It is my first to solve the Multi-Lingual problem. Thanks for those great starter kernels:\n<a href=\"https://www.kaggle.com/miklgr500/jigsaw-tpu-bert-two-stage-training\">https://www.kaggle.com/miklgr500/jigsaw-tpu-bert-two-stage-training</a>\n<a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta</a>\nI follow these codes at the very beginning and quickly got a lb 0.9415 single xlmr large model. </p>\n\n<p>After working 3 months for this problem, I got something useful tricks. </p>\n\n<p>1) What is the best training policy ? \nThere are three training policies : one stage training,two stage training and three stage training.\na) One stage training means that we always train the same  data  in training . <br>\nThis policy always get the lower scores.  But we can get some diversity.\n<br>\nb) Two stage training means that we will have two different data for training.  e.g. <br>\nWe firstly train on \"Toxic Comment Classification\" dataset +\"Unintended Bias in Toxicity\" dataset . \nAfter that we train on 8k validation data. <br>\nThis policy always get the higher scores. \n<br>\nc) Three stage training means that we will have three different data for training.  e.g. \nWe firstly train on \"Toxic Comment Classification\" dataset +\"Unintended Bias in Toxicity\" dataset . \nSecondly  we train on 200k open subtitle data. <br>\nFinally we train on 8k validation data. \nThis policy always get the similar scores with Two stage training. But we get the higher private scores.\nFor example, \nmy xlmr large model got  lb 0.9407, private score 0.9414.\nmy another xlmr large model got  lb 0.9420, private score 0.9424. </p>\n\n<p>Therefore , Three stage training is the best training policy.  I don't know why. \n<br></p>\n\n<p>2) How to use data? </p>\n\n<p>At the very beginning , my  teammate found that when we train on \"Toxic Comment Classification\" dataset +\"Unintended Bias in Toxicity\" dataset ,the valid auc  are always lower than training on \"Toxic Comment Classification\" dataset only.   To get a higher valid auc, we can not combine many different data to train. <br>\nBefore  9 days ago, I found that if we use our best single model to predict the training data's target and use the predicting target as the label instead of the 0/1 label.Then we can get a higher valid auc score.  e.g.  We use the our xlmr large model(lb 0.9411) to predict \"Toxic Comment Classification\" dataset and get the predicted target of them.  Something like these :\nid toxic preds\n1     0    0.341\n2     1     0.813\n3     0    0.211</p>\n\n<p>Then we use the predicted probabilities as label target to train.  Using this trick, we always get a 0.003 boost in valid auc score.  This trick can use on other datasets which has 0/1 labels.  After using this trick on both \"Toxic Comment Classification\" dataset  and \"Unintended Bias in Toxicity\" dataset , we can use them to train our model. The more data, the higher valid auc score. </p>\n\n<p>Based on it, I trained a xlmr large model with 790k training data(\"Toxic Comment Classification\" dataset  and \"Unintended Bias in Toxicity\" dataset and open subtitle dataset) . \nit got  a lb 0.9464 score and a 0.9445 private score. </p>\n\n<p>That's all .  Hope these two tricks can help you. </p>",
  "messages": [
    {
      "id": "898068",
      "postDate": "06/23/2020 09:04:49",
      "content": "<p>Congratulations all winners.  Thanks Kaggle supply this funny competition.  Thanks my teammates @Yang Zhang @MaChaogong @Hikkiiiiiiiii @Morphy</p>\n\n<p>It is my first to solve the Multi-Lingual problem. Thanks for those great starter kernels:\n<a href=\"https://www.kaggle.com/miklgr500/jigsaw-tpu-bert-two-stage-training\">https://www.kaggle.com/miklgr500/jigsaw-tpu-bert-two-stage-training</a>\n<a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta</a>\nI follow these codes at the very beginning and quickly got a lb 0.9415 single xlmr large model. </p>\n\n<p>After working 3 months for this problem, I got something useful tricks. </p>\n\n<p>1) What is the best training policy ? \nThere are three training policies : one stage training,two stage training and three stage training.\na) One stage training means that we always train the same  data  in training . <br>\nThis policy always get the lower scores.  But we can get some diversity.\n<br>\nb) Two stage training means that we will have two different data for training.  e.g. <br>\nWe firstly train on \"Toxic Comment Classification\" dataset +\"Unintended Bias in Toxicity\" dataset . \nAfter that we train on 8k validation data. <br>\nThis policy always get the higher scores. \n<br>\nc) Three stage training means that we will have three different data for training.  e.g. \nWe firstly train on \"Toxic Comment Classification\" dataset +\"Unintended Bias in Toxicity\" dataset . \nSecondly  we train on 200k open subtitle data. <br>\nFinally we train on 8k validation data. \nThis policy always get the similar scores with Two stage training. But we get the higher private scores.\nFor example, \nmy xlmr large model got  lb 0.9407, private score 0.9414.\nmy another xlmr large model got  lb 0.9420, private score 0.9424. </p>\n\n<p>Therefore , Three stage training is the best training policy.  I don't know why. \n<br></p>\n\n<p>2) How to use data? </p>\n\n<p>At the very beginning , my  teammate found that when we train on \"Toxic Comment Classification\" dataset +\"Unintended Bias in Toxicity\" dataset ,the valid auc  are always lower than training on \"Toxic Comment Classification\" dataset only.   To get a higher valid auc, we can not combine many different data to train. <br>\nBefore  9 days ago, I found that if we use our best single model to predict the training data's target and use the predicting target as the label instead of the 0/1 label.Then we can get a higher valid auc score.  e.g.  We use the our xlmr large model(lb 0.9411) to predict \"Toxic Comment Classification\" dataset and get the predicted target of them.  Something like these :\nid toxic preds\n1     0    0.341\n2     1     0.813\n3     0    0.211</p>\n\n<p>Then we use the predicted probabilities as label target to train.  Using this trick, we always get a 0.003 boost in valid auc score.  This trick can use on other datasets which has 0/1 labels.  After using this trick on both \"Toxic Comment Classification\" dataset  and \"Unintended Bias in Toxicity\" dataset , we can use them to train our model. The more data, the higher valid auc score. </p>\n\n<p>Based on it, I trained a xlmr large model with 790k training data(\"Toxic Comment Classification\" dataset  and \"Unintended Bias in Toxicity\" dataset and open subtitle dataset) . \nit got  a lb 0.9464 score and a 0.9445 private score. </p>\n\n<p>That's all .  Hope these two tricks can help you. </p>",
      "rawMarkdown": "Congratulations all winners.  Thanks Kaggle supply this funny competition.  Thanks my teammates @Yang Zhang @MaChaogong @Hikkiiiiiiiii @Morphy\n\n\n\nIt is my first to solve the Multi-Lingual problem. Thanks for those great starter kernels:\nhttps://www.kaggle.com/miklgr500/jigsaw-tpu-bert-two-stage-training\nhttps://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\nI follow these codes at the very beginning and quickly got a lb 0.9415 single xlmr large model. \n\nAfter working 3 months for this problem, I got something useful tricks. \n\n1) What is the best training policy ? \nThere are three training policies : one stage training,two stage training and three stage training.\na) One stage training means that we always train the same  data  in training .  \nThis policy always get the lower scores.  But we can get some diversity.\n<br>\nb) Two stage training means that we will have two different data for training.  e.g.  \nWe firstly train on \"Toxic Comment Classification\" dataset +\"Unintended Bias in Toxicity\" dataset . \nAfter that we train on 8k validation data.     \nThis policy always get the higher scores. \n<br>\nc) Three stage training means that we will have three different data for training.  e.g. \nWe firstly train on \"Toxic Comment Classification\" dataset +\"Unintended Bias in Toxicity\" dataset . \nSecondly  we train on 200k open subtitle data.    \nFinally we train on 8k validation data. \nThis policy always get the similar scores with Two stage training. But we get the higher private scores.\nFor example, \nmy xlmr large model got  lb 0.9407, private score 0.9414.\nmy another xlmr large model got  lb 0.9420, private score 0.9424. \n\nTherefore , Three stage training is the best training policy.  I don't know why. \n<br>\n\n2) How to use data? \n\nAt the very beginning , my  teammate found that when we train on \"Toxic Comment Classification\" dataset +\"Unintended Bias in Toxicity\" dataset ,the valid auc  are always lower than training on \"Toxic Comment Classification\" dataset only.   To get a higher valid auc, we can not combine many different data to train.   \nBefore  9 days ago, I found that if we use our best single model to predict the training data's target and use the predicting target as the label instead of the 0/1 label.Then we can get a higher valid auc score.  e.g.  We use the our xlmr large model(lb 0.9411) to predict \"Toxic Comment Classification\" dataset and get the predicted target of them.  Something like these :\nid toxic preds\n1     0    0.341\n2     1     0.813\n3     0    0.211\n\nThen we use the predicted probabilities as label target to train.  Using this trick, we always get a 0.003 boost in valid auc score.  This trick can use on other datasets which has 0/1 labels.  After using this trick on both \"Toxic Comment Classification\" dataset  and \"Unintended Bias in Toxicity\" dataset , we can use them to train our model. The more data, the higher valid auc score. \n\nBased on it, I trained a xlmr large model with 790k training data(\"Toxic Comment Classification\" dataset  and \"Unintended Bias in Toxicity\" dataset and open subtitle dataset) . \nit got  a lb 0.9464 score and a 0.9445 private score. \n\n\n\nThat's all .  Hope these two tricks can help you.",
      "votes": null
    },
    {
      "id": "898587",
      "postDate": "06/23/2020 15:56:49",
      "content": "<p>Using soft labels instead of hard labels like 0/1 gave great boost in the last Jigsaw toxic competition. But in that competition, I mainly used bert base while in this one this trick didn't gave as much boost perhaps because xlmr large is already large model and able to pick the signal in hard labels. But we used hard sampling to sample training data, i.e. training a model to predict training data and pick those with large gap from ground truth labels, which brought 0.9469 on the public leaderboard. More details can be seen in our solution.</p>",
      "rawMarkdown": "Using soft labels instead of hard labels like 0/1 gave great boost in the last Jigsaw toxic competition. But in that competition, I mainly used bert base while in this one this trick didn't gave as much boost perhaps because xlmr large is already large model and able to pick the signal in hard labels. But we used hard sampling to sample training data, i.e. training a model to predict training data and pick those with large gap from ground truth labels, which brought 0.9469 on the public leaderboard. More details can be seen in our solution.",
      "votes": null
    },
    {
      "id": "899016",
      "postDate": "06/23/2020 22:33:22",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": null
    },
    {
      "id": "899858",
      "postDate": "06/24/2020 13:42:37",
      "content": "<p>Great tricks 👍 </p>",
      "rawMarkdown": "Great tricks 👍",
      "votes": null
    },
    {
      "id": "900818",
      "postDate": "06/25/2020 04:46:50",
      "content": "<p>Congratulation to You all</p>",
      "rawMarkdown": "Congratulation to You all",
      "votes": null
    },
    {
      "id": "965954",
      "postDate": "08/11/2020 02:22:54",
      "content": "<p>As a beginner, I would like to ask if we use the predicting target as the label instead of the 0/1 label, if we do that, how should we set the loss function? Can you give me some code reference？Thank you!</p>",
      "rawMarkdown": "As a beginner, I would like to ask if we use the predicting target as the label instead of the 0/1 label, if we do that, how should we set the loss function? Can you give me some code reference？Thank you!",
      "votes": null
    },
    {
      "id": "966176",
      "postDate": "08/11/2020 08:08:18",
      "content": "<p>You are welcome.  You can use the BCE loss as your  loss function.  For example,  F.binary_cross_entropy_with_logits</p>",
      "rawMarkdown": "You are welcome.  You can use the BCE loss as your  loss function.  For example,  F.binary_cross_entropy_with_logits",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 965954,
      "author_name": "mnk812",
      "author_url": "",
      "post_date": "08/11/2020 02:22:54",
      "content": "<p>As a beginner, I would like to ask if we use the predicting target as the label instead of the 0/1 label, if we do that, how should we set the loss function? Can you give me some code reference？Thank you!</p>",
      "votes": null,
      "replies": [
        {
          "id": 966176,
          "author_name": "qinhui1999",
          "author_url": "",
          "post_date": "08/11/2020 08:08:18",
          "content": "<p>You are welcome.  You can use the BCE loss as your  loss function.  For example,  F.binary_cross_entropy_with_logits</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 898587,
      "author_name": "nzholmes",
      "author_url": "",
      "post_date": "06/23/2020 15:56:49",
      "content": "<p>Using soft labels instead of hard labels like 0/1 gave great boost in the last Jigsaw toxic competition. But in that competition, I mainly used bert base while in this one this trick didn't gave as much boost perhaps because xlmr large is already large model and able to pick the signal in hard labels. But we used hard sampling to sample training data, i.e. training a model to predict training data and pick those with large gap from ground truth labels, which brought 0.9469 on the public leaderboard. More details can be seen in our solution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 899016,
          "author_name": "qinhui1999",
          "author_url": "",
          "post_date": "06/23/2020 22:33:22",
          "content": "<p>Thanks for sharing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 899858,
      "author_name": "nidhalyounes",
      "author_url": "",
      "post_date": "06/24/2020 13:42:37",
      "content": "<p>Great tricks 👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 900818,
      "author_name": "mdselimreza",
      "author_url": "",
      "post_date": "06/25/2020 04:46:50",
      "content": "<p>Congratulation to You all</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "898068": "Congratulations all winners.  Thanks Kaggle supply this funny competition.  Thanks my teammates @Yang Zhang @MaChaogong @Hikkiiiiiiiii @Morphy\n\n\n\nIt is my first to solve the Multi-Lingual problem. Thanks for those great starter kernels:\nhttps://www.kaggle.com/miklgr500/jigsaw-tpu-bert-two-stage-training\nhttps://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\nI follow these codes at the very beginning and quickly got a lb 0.9415 single xlmr large model. \n\nAfter working 3 months for this problem, I got something useful tricks. \n\n1) What is the best training policy ? \nThere are three training policies : one stage training,two stage training and three stage training.\na) One stage training means that we always train the same  data  in training .  \nThis policy always get the lower scores.  But we can get some diversity.\n<br>\nb) Two stage training means that we will have two different data for training.  e.g.  \nWe firstly train on \"Toxic Comment Classification\" dataset +\"Unintended Bias in Toxicity\" dataset . \nAfter that we train on 8k validation data.     \nThis policy always get the higher scores. \n<br>\nc) Three stage training means that we will have three different data for training.  e.g. \nWe firstly train on \"Toxic Comment Classification\" dataset +\"Unintended Bias in Toxicity\" dataset . \nSecondly  we train on 200k open subtitle data.    \nFinally we train on 8k validation data. \nThis policy always get the similar scores with Two stage training. But we get the higher private scores.\nFor example, \nmy xlmr large model got  lb 0.9407, private score 0.9414.\nmy another xlmr large model got  lb 0.9420, private score 0.9424. \n\nTherefore , Three stage training is the best training policy.  I don't know why. \n<br>\n\n2) How to use data? \n\nAt the very beginning , my  teammate found that when we train on \"Toxic Comment Classification\" dataset +\"Unintended Bias in Toxicity\" dataset ,the valid auc  are always lower than training on \"Toxic Comment Classification\" dataset only.   To get a higher valid auc, we can not combine many different data to train.   \nBefore  9 days ago, I found that if we use our best single model to predict the training data's target and use the predicting target as the label instead of the 0/1 label.Then we can get a higher valid auc score.  e.g.  We use the our xlmr large model(lb 0.9411) to predict \"Toxic Comment Classification\" dataset and get the predicted target of them.  Something like these :\nid toxic preds\n1     0    0.341\n2     1     0.813\n3     0    0.211\n\nThen we use the predicted probabilities as label target to train.  Using this trick, we always get a 0.003 boost in valid auc score.  This trick can use on other datasets which has 0/1 labels.  After using this trick on both \"Toxic Comment Classification\" dataset  and \"Unintended Bias in Toxicity\" dataset , we can use them to train our model. The more data, the higher valid auc score. \n\nBased on it, I trained a xlmr large model with 790k training data(\"Toxic Comment Classification\" dataset  and \"Unintended Bias in Toxicity\" dataset and open subtitle dataset) . \nit got  a lb 0.9464 score and a 0.9445 private score. \n\n\n\nThat's all .  Hope these two tricks can help you.",
    "898587": "Using soft labels instead of hard labels like 0/1 gave great boost in the last Jigsaw toxic competition. But in that competition, I mainly used bert base while in this one this trick didn't gave as much boost perhaps because xlmr large is already large model and able to pick the signal in hard labels. But we used hard sampling to sample training data, i.e. training a model to predict training data and pick those with large gap from ground truth labels, which brought 0.9469 on the public leaderboard. More details can be seen in our solution.",
    "899016": "Thanks for sharing.",
    "899858": "Great tricks 👍",
    "900818": "Congratulation to You all",
    "965954": "As a beginner, I would like to ask if we use the predicting target as the label instead of the 0/1 label, if we do that, how should we set the loss function? Can you give me some code reference？Thank you!",
    "966176": "You are welcome.  You can use the BCE loss as your  loss function.  For example,  F.binary_cross_entropy_with_logits"
  },
  "source": "meta"
}