{
  "id": 140254,
  "title": "Insights on achieving 0.9383 without translating data",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/140254",
  "author_name": "xhlulu",
  "post_date": "2020-04-01T04:00:44.607000",
  "votes": 159,
  "comment_count": 74,
  "views": 0,
  "content": "<p>Hello.</p>\n\n<p>Some of you might have just noticed me releasing my <a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">XLM-Roberta notebook</a> that achieved 0.9383. Since the notebook takes a while to run, I wanted to share my insights in making this notebook as a discussion post (I'll integrate this to my notebook later):\n* <strong>Use a state of the art multilingual transformer</strong>: As you may know, mBERT has been around for a while (almost 2 years now). Since it was released, a lot of research has been done on improving the pre-training step in order to create a strong cross-lingual language model. One major paper in that subject is <a href=\"https://arxiv.org/pdf/1901.07291.pdf\">Cross-lingual Language Model Pretraining\n</a> which introduced the XLM model. It was then further improved through efficient pretraining in <a href=\"https://arxiv.org/abs/1901.07291\">Unsupervised Cross-lingual Representation Learning at Scale</a>; this last model is called <em>XLM-Roberta</em>, which is what I used in my notebook. It is available in the transformers library, so it was very easy to set up.\n* <strong>Balance your data</strong>: The training data set consists of over 2M data points, but it is heavily unbalanced since it only has around 120k positive labels. In order to limit training time, I selected all of the positive labels, and subsampled the negative labels such that I have in total 400k labels, which is a lot closer to the class frequency of the validation labels.\n* <strong>Train your model on the validation set</strong>: I trained two epochs on the 400k training samples, followed by 2 more epochs on the relatively small validation set. My hypothesis is that this two-steps process improves the model's ability to first learn the underlying structure of toxic comments (by training on 400k english comments), then adapt those learned structures to discriminate other languages from the validation set (Turkish, Italian, Portuguese, etc.).\n* <strong>Translation is not the (only) solution</strong>: Although notebooks have shown great results by translating the test set, I decided to stick to my guns and focus on training a cross-lingual model. I'm sure that people will find ways to improve this model through translation, but I think that if we were to deploy a toxicity detection model in production, we won't always have the luxury to request Google translate API calls or run two transformers in parallel. Additionally, Google translate might not yield good results for low-resource languages (e.g. Swahili), but models like XLM-R might still perform fairly well.</p>",
  "messages": [
    {
      "id": 793479,
      "postDate": "2020-04-01T04:00:44.607Z",
      "content": "<p>Hello.</p>\n\n<p>Some of you might have just noticed me releasing my <a href=\"https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta\">XLM-Roberta notebook</a> that achieved 0.9383. Since the notebook takes a while to run, I wanted to share my insights in making this notebook as a discussion post (I'll integrate this to my notebook later):\n* <strong>Use a state of the art multilingual transformer</strong>: As you may know, mBERT has been around for a while (almost 2 years now). Since it was released, a lot of research has been done on improving the pre-training step in order to create a strong cross-lingual language model. One major paper in that subject is <a href=\"https://arxiv.org/pdf/1901.07291.pdf\">Cross-lingual Language Model Pretraining\n</a> which introduced the XLM model. It was then further improved through efficient pretraining in <a href=\"https://arxiv.org/abs/1901.07291\">Unsupervised Cross-lingual Representation Learning at Scale</a>; this last model is called <em>XLM-Roberta</em>, which is what I used in my notebook. It is available in the transformers library, so it was very easy to set up.\n* <strong>Balance your data</strong>: The training data set consists of over 2M data points, but it is heavily unbalanced since it only has around 120k positive labels. In order to limit training time, I selected all of the positive labels, and subsampled the negative labels such that I have in total 400k labels, which is a lot closer to the class frequency of the validation labels.\n* <strong>Train your model on the validation set</strong>: I trained two epochs on the 400k training samples, followed by 2 more epochs on the relatively small validation set. My hypothesis is that this two-steps process improves the model's ability to first learn the underlying structure of toxic comments (by training on 400k english comments), then adapt those learned structures to discriminate other languages from the validation set (Turkish, Italian, Portuguese, etc.).\n* <strong>Translation is not the (only) solution</strong>: Although notebooks have shown great results by translating the test set, I decided to stick to my guns and focus on training a cross-lingual model. I'm sure that people will find ways to improve this model through translation, but I think that if we were to deploy a toxicity detection model in production, we won't always have the luxury to request Google translate API calls or run two transformers in parallel. Additionally, Google translate might not yield good results for low-resource languages (e.g. Swahili), but models like XLM-R might still perform fairly well.</p>",
      "rawMarkdown": "Hello.\n\nSome of you might have just noticed me releasing my [XLM-Roberta notebook](https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta) that achieved 0.9383. Since the notebook takes a while to run, I wanted to share my insights in making this notebook as a discussion post (I'll integrate this to my notebook later):\n* **Use a state of the art multilingual transformer**: As you may know, mBERT has been around for a while (almost 2 years now). Since it was released, a lot of research has been done on improving the pre-training step in order to create a strong cross-lingual language model. One major paper in that subject is [Cross-lingual Language Model Pretraining\n](https://arxiv.org/pdf/1901.07291.pdf) which introduced the XLM model. It was then further improved through efficient pretraining in [Unsupervised Cross-lingual Representation Learning at Scale](https://arxiv.org/abs/1901.07291); this last model is called *XLM-Roberta*, which is what I used in my notebook. It is available in the transformers library, so it was very easy to set up.\n* **Balance your data**: The training data set consists of over 2M data points, but it is heavily unbalanced since it only has around 120k positive labels. In order to limit training time, I selected all of the positive labels, and subsampled the negative labels such that I have in total 400k labels, which is a lot closer to the class frequency of the validation labels.\n* **Train your model on the validation set**: I trained two epochs on the 400k training samples, followed by 2 more epochs on the relatively small validation set. My hypothesis is that this two-steps process improves the model's ability to first learn the underlying structure of toxic comments (by training on 400k english comments), then adapt those learned structures to discriminate other languages from the validation set (Turkish, Italian, Portuguese, etc.).\n* **Translation is not the (only) solution**: Although notebooks have shown great results by translating the test set, I decided to stick to my guns and focus on training a cross-lingual model. I'm sure that people will find ways to improve this model through translation, but I think that if we were to deploy a toxicity detection model in production, we won't always have the luxury to request Google translate API calls or run two transformers in parallel. Additionally, Google translate might not yield good results for low-resource languages (e.g. Swahili), but models like XLM-R might still perform fairly well.",
      "votes": 159
    },
    {
      "id": 794815,
      "postDate": "2020-04-02T05:38:24.083Z",
      "content": "<p>Just FYI for those who doesn't know XLM like me.</p>\n\n<p>blog: <a href=\"https://ai.facebook.com/blog/cross-lingual-pretraining/\">https://ai.facebook.com/blog/cross-lingual-pretraining/</a></p>\n\n<p>github: <a href=\"https://github.com/facebookresearch/XLM\">https://github.com/facebookresearch/XLM</a></p>\n\n<p>and paper: <a href=\"https://arxiv.org/pdf/1911.02116.pdf\">https://arxiv.org/pdf/1911.02116.pdf</a></p>",
      "rawMarkdown": "Just FYI for those who doesn't know XLM like me.\n\nblog: https://ai.facebook.com/blog/cross-lingual-pretraining/\n\ngithub: https://github.com/facebookresearch/XLM\n\nand paper: https://arxiv.org/pdf/1911.02116.pdf",
      "votes": 10,
      "replies": [
        {
          "id": 794828,
          "postDate": "2020-04-02T05:51:17.887Z",
          "content": "<p>Thanks. Are you on the Game? 😃  </p>",
          "rawMarkdown": "Thanks. Are you on the Game? 😃  "
        },
        {
          "id": 795107,
          "postDate": "2020-04-02T12:08:54.047Z",
          "content": "<p>No 🤕 I'll be off for a few weeks.</p>",
          "rawMarkdown": "No 🤕 I'll be off for a few weeks."
        },
        {
          "id": 796540,
          "postDate": "2020-04-03T17:12:56.897Z",
          "content": "<p>this is also good reading <a href=\"https://towardsdatascience.com/xlm-enhancing-bert-for-cross-lingual-language-model-5aeed9e6f14b\">https://towardsdatascience.com/xlm-enhancing-bert-for-cross-lingual-language-model-5aeed9e6f14b</a></p>",
          "rawMarkdown": "this is also good reading https://towardsdatascience.com/xlm-enhancing-bert-for-cross-lingual-language-model-5aeed9e6f14b",
          "votes": 1
        }
      ]
    },
    {
      "id": 793925,
      "postDate": "2020-04-01T12:07:24.760Z",
      "content": "<p>Great Work! But where does the randomness come from? It seems like the code gives anywhere between 0.92 and 0.94? </p>",
      "rawMarkdown": "Great Work! But where does the randomness come from? It seems like the code gives anywhere between 0.92 and 0.94? ",
      "votes": 7,
      "replies": [
        {
          "id": 794025,
          "postDate": "2020-04-01T13:47:59.437Z",
          "content": "<p>The validation set is pretty small and randomly split, so the order in which data is fed affects how well the final model performs. When you are translating the test set into English, the model depends a lot more on the training set, which is huge, so there's very little variation in the final performance.</p>",
          "rawMarkdown": "The validation set is pretty small and randomly split, so the order in which data is fed affects how well the final model performs. When you are translating the test set into English, the model depends a lot more on the training set, which is huge, so there's very little variation in the final performance.",
          "votes": 12
        },
        {
          "id": 801764,
          "postDate": "2020-04-08T19:21:50.090Z",
          "content": "<p>+1</p>",
          "rawMarkdown": "+1",
          "votes": -1
        }
      ]
    },
    {
      "id": 815328,
      "postDate": "2020-04-21T13:10:56.103Z",
      "content": "<p>Competition be like: 'Am I a joke to you'??😄 \nGreat work though</p>",
      "rawMarkdown": "Competition be like: 'Am I a joke to you'??😄 \nGreat work though",
      "votes": 3
    },
    {
      "id": 794269,
      "postDate": "2020-04-01T17:40:26.527Z",
      "content": "<p>Amazing work.\nTweeted: <a href=\"https://twitter.com/martin_gorner/status/1245405207762046976\">https://twitter.com/martin_gorner/status/1245405207762046976</a></p>",
      "rawMarkdown": "Amazing work.\nTweeted: https://twitter.com/martin_gorner/status/1245405207762046976",
      "votes": 3,
      "replies": [
        {
          "id": 794441,
          "postDate": "2020-04-01T20:27:24.837Z",
          "content": "<p>Thanks for the shout out!</p>",
          "rawMarkdown": "Thanks for the shout out!",
          "votes": 1
        }
      ]
    },
    {
      "id": 794113,
      "postDate": "2020-04-01T15:13:01.717Z",
      "content": "<p>Great job! that last point is very interesting, it might be a little late for that, but I think the developed solution would be more interesting for the host if the competition had a hidden test set with many languages, maybe even on that only exists on that set.</p>",
      "rawMarkdown": "Great job! that last point is very interesting, it might be a little late for that, but I think the developed solution would be more interesting for the host if the competition had a hidden test set with many languages, maybe even on that only exists on that set.",
      "votes": 3
    },
    {
      "id": 804249,
      "postDate": "2020-04-11T12:00:32.293Z",
      "content": "<p>Amazing work! I took the part in the competition as well to improve upon my skills, reading the description of your model really helped!</p>",
      "rawMarkdown": "Amazing work! I took the part in the competition as well to improve upon my skills, reading the description of your model really helped!",
      "votes": 1
    },
    {
      "id": 794431,
      "postDate": "2020-04-01T20:10:46.827Z",
      "content": "<p>Great! I'm glad to know translation won't be the final solution😂 </p>",
      "rawMarkdown": "Great! I'm glad to know translation won't be the final solution😂 ",
      "votes": 1,
      "replies": [
        {
          "id": 794440,
          "postDate": "2020-04-01T20:27:12.870Z",
          "content": "<p>Who knows... I think that translation will definitely play a role in this... Maybe just in a different way than what has been done so far? ;)</p>",
          "rawMarkdown": "Who knows... I think that translation will definitely play a role in this... Maybe just in a different way than what has been done so far? ;)",
          "votes": 3
        }
      ]
    },
    {
      "id": 793514,
      "postDate": "2020-04-01T04:35:28.893Z",
      "content": "<p>Thanks for sharing the insights, they helped😄. In my observation when training on the validation data, the loss correlated more than when training on the train data, looks like they are from different distributions.</p>",
      "rawMarkdown": "Thanks for sharing the insights, they helped😄. In my observation when training on the validation data, the loss correlated more than when training on the train data, looks like they are from different distributions.",
      "votes": 1
    },
    {
      "id": 793503,
      "postDate": "2020-04-01T04:27:18.517Z",
      "content": "<p>Very nice! Just came across XLM-R myself. </p>",
      "rawMarkdown": "Very nice! Just came across XLM-R myself. ",
      "votes": 1
    },
    {
      "id": 795039,
      "postDate": "2020-04-02T10:33:47.867Z",
      "content": "<p>Great kernel. I am trying to understand how exactly the training loop is working with distributed TPU usage. I assume the code is running on 8 processes, right? What exactly happens here if I call <code>model.fit()</code>? Are the batches of an epoch distributed?</p>",
      "rawMarkdown": "Great kernel. I am trying to understand how exactly the training loop is working with distributed TPU usage. I assume the code is running on 8 processes, right? What exactly happens here if I call `model.fit()`? Are the batches of an epoch distributed?",
      "votes": 2,
      "replies": [
        {
          "id": 795282,
          "postDate": "2020-04-02T15:26:20.817Z",
          "content": "<p>I presume they are. I think <a href=\"/mgornergoogle\">@mgornergoogle</a> would be a better person to answer this :)</p>",
          "rawMarkdown": "I presume they are. I think @mgornergoogle would be a better person to answer this :)"
        },
        {
          "id": 795306,
          "postDate": "2020-04-02T15:53:47.847Z",
          "content": "<p>Yes, The model is replicated on the 8 cores. The batch you define through tf.data.Dataset will get split across the 8 cores of the TPU. Gradient updates are through an all-reduce algorithm running on the TPU board.</p>",
          "rawMarkdown": "Yes, The model is replicated on the 8 cores. The batch you define through tf.data.Dataset will get split across the 8 cores of the TPU. Gradient updates are through an all-reduce algorithm running on the TPU board.\n ",
          "votes": 4
        },
        {
          "id": 795957,
          "postDate": "2020-04-03T06:39:09.757Z",
          "content": "<p>Thanks <a href=\"/mgornergoogle\">@mgornergoogle</a>. I have two follow up questions if you don't mind:</p>\n\n<ul>\n<li>I see datasets using <code>.split()</code> method to duplicate them infinitely. What about using <code>.shard()</code>? I don't see it used in the kernels here on Kaggle, but is it useful?</li>\n<li>What exactly happens in callbacks that for example operate <code>on_batch_end</code> level and utilize fit parameters like <code>global_step</code> or <code>total_steps</code>. Are those per process only, or really globally? So let's say with 8 processes I am running 16 batches per epoch. Batch 1-8 get distributed across processes, and 9-16 get distributed. If <code>on_batch_end</code> is called for each batch, is it 1,2,3,...,16 or is it 1,1,1,1,...,2,2,2,2...? Hope it is clear what I am trying to ask :)</li>\n</ul>\n\n<p>Edit: I just figured that you said the batch gets split, so a single batch gets split? My second question is redundant then I guess, I thought separate batches get distributed.</p>\n\n<p>Thanks!</p>",
          "rawMarkdown": "Thanks @mgornergoogle. I have two follow up questions if you don't mind:\n\n- I see datasets using `.split()` method to duplicate them infinitely. What about using `.shard()`? I don't see it used in the kernels here on Kaggle, but is it useful?\n- What exactly happens in callbacks that for example operate `on_batch_end` level and utilize fit parameters like `global_step` or `total_steps`. Are those per process only, or really globally? So let's say with 8 processes I am running 16 batches per epoch. Batch 1-8 get distributed across processes, and 9-16 get distributed. If `on_batch_end` is called for each batch, is it 1,2,3,...,16 or is it 1,1,1,1,...,2,2,2,2...? Hope it is clear what I am trying to ask :)\n\nEdit: I just figured that you said the batch gets split, so a single batch gets split? My second question is redundant then I guess, I thought separate batches get distributed.\n\nThanks!"
        },
        {
          "id": 796452,
          "postDate": "2020-04-03T15:37:55.500Z",
          "content": "<p>.shard() is called by tf.data.Dataset() automatically when running a model in a TPUStrategy.\nActually, when running on a TPU pod, tf.data.Dataset will even attempt to shard your set of input files between TPU boards, then shard the batch between the 8 TPU cores of a board. If your dataset is made of a sufficient number of files, the benefit will be that only the required files will be loaded on each TPU board.</p>\n\n<p>Keras callbacks are processed on the Kaggle VM, so there are no parallelism issues there.</p>",
          "rawMarkdown": ".shard() is called by tf.data.Dataset() automatically when running a model in a TPUStrategy.\nActually, when running on a TPU pod, tf.data.Dataset will even attempt to shard your set of input files between TPU boards, then shard the batch between the 8 TPU cores of a board. If your dataset is made of a sufficient number of files, the benefit will be that only the required files will be loaded on each TPU board.\n\nKeras callbacks are processed on the Kaggle VM, so there are no parallelism issues there.",
          "votes": 5
        }
      ]
    },
    {
      "id": 794074,
      "postDate": "2020-04-01T14:33:52.437Z",
      "content": "<p>Great work guys! Learning a lot from you guys <a href=\"/abhishek\">@abhishek</a> and <a href=\"/xhlulu\">@xhlulu</a>. Thank you!  </p>",
      "rawMarkdown": "Great work guys! Learning a lot from you guys @abhishek and @xhlulu. Thank you!  ",
      "votes": 2
    },
    {
      "id": 793533,
      "postDate": "2020-04-01T05:16:32.843Z",
      "content": "<ul>\n<li>Another very strange thing i noticed is that the Batch-Size somehow effects the feature vectors wrt XLM-R and huggingface; Not sure why;</li>\n</ul>",
      "rawMarkdown": "+ Another very strange thing i noticed is that the Batch-Size somehow effects the feature vectors wrt XLM-R and huggingface; Not sure why;",
      "votes": 2
    },
    {
      "id": 793501,
      "postDate": "2020-04-01T04:25:07.263Z",
      "content": "<p>Great job! Impressive building a model without translating. Thanks for sharing.</p>",
      "rawMarkdown": "Great job! Impressive building a model without translating. Thanks for sharing.",
      "votes": 2
    },
    {
      "id": 808433,
      "postDate": "2020-04-15T12:26:46.880Z",
      "content": "<p>Were you able to get deterministic results with Tensorflow TPU? I am getting varied scores and losses every time I run and commit  <a href=\"/xhlulu\">@xhlulu</a> </p>",
      "rawMarkdown": "Were you able to get deterministic results with Tensorflow TPU? I am getting varied scores and losses every time I run and commit  @xhlulu ",
      "replies": [
        {
          "id": 808689,
          "postDate": "2020-04-15T15:00:16.497Z",
          "content": "<p><a href=\"/shahules\">@shahules</a> There is no way as far as I know.</p>",
          "rawMarkdown": "@shahules There is no way as far as I know.",
          "votes": 1
        },
        {
          "id": 808761,
          "postDate": "2020-04-15T15:42:49.427Z",
          "content": "<p>oh,I went through their documentation. Anyway,one more reasons to switch to pytorch!! <a href=\"/bamps53\">@bamps53</a> </p>",
          "rawMarkdown": "oh,I went through their documentation. Anyway,one more reasons to switch to pytorch!! @bamps53 "
        },
        {
          "id": 809261,
          "postDate": "2020-04-16T02:29:58.947Z",
          "content": "<p><a href=\"/shahules\">@shahules</a> Yes, but currently pytorch xla is soooo frustrating to debug the code. I recommend you to stick to GPU if you choose TPU☹️  </p>",
          "rawMarkdown": "@shahules Yes, but currently pytorch xla is soooo frustrating to debug the code. I recommend you to stick to GPU if you choose TPU☹️  ",
          "votes": 1
        },
        {
          "id": 809264,
          "postDate": "2020-04-16T02:35:06.150Z",
          "content": "<p>Thanks buddy <a href=\"/bamps53\">@bamps53</a> I will keep that in mind.</p>",
          "rawMarkdown": "Thanks buddy @bamps53 I will keep that in mind."
        }
      ]
    },
    {
      "id": 2781804,
      "postDate": "2024-04-29T02:41:44.457Z",
      "content": "<p>Great Work ! :3</p>",
      "rawMarkdown": "Great Work ! :3"
    },
    {
      "id": 1108133,
      "postDate": "2020-12-10T10:06:09.033Z",
      "content": "<p><a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a>: I'm <em>really</em> late to the party but nevertheless wanted to say thank you for sharing the above.  </p>",
      "rawMarkdown": "@xhlulu: I'm *really* late to the party but nevertheless wanted to say thank you for sharing the above.  "
    },
    {
      "id": 1091014,
      "postDate": "2020-11-25T18:16:35.793Z",
      "content": "<p>Great !              </p>",
      "rawMarkdown": "Great !              "
    },
    {
      "id": 842705,
      "postDate": "2020-05-11T15:26:44.870Z",
      "content": "<p>Xhlulu, congratulations again on the score.  Reading over your notes, I was wondering how you would deploy this model in the real world.  How would your model change and adjust to knew information?  How would the model receive and apply feed back?</p>",
      "rawMarkdown": "Xhlulu, congratulations again on the score.  Reading over your notes, I was wondering how you would deploy this model in the real world.  How would your model change and adjust to knew information?  How would the model receive and apply feed back?"
    },
    {
      "id": 828727,
      "postDate": "2020-05-01T08:44:07.910Z",
      "content": "<p>Thank you for sharing your wonderful insight!\nActualy, I'm a beginner in NLP so that your kernel gave me lots of help.\nI have two questions about it.</p>\n\n<p>1) You trained the validation data set. Can I interpret it as a transfer learning? You first trained the train dataset, and then trained the validation data set. Then, the second train (which trained the validation data set) could be considered as transfer learning?</p>\n\n<p>2) You said XLM-roBERTa is the best model that performs great. Then, how can I improve the performance more? If I participate in competition using tableau data, I can do lots of things such as feature engineering, hyper-parameter tunning, ensemble and so on. As I said, I'm a very beginner in NLP. How can I improve the performance in NLP competiton? Can I do feature engineering or something like that also in NLP? </p>\n\n<p>sorry for silly questions, Thanks for reading</p>",
      "rawMarkdown": "Thank you for sharing your wonderful insight!\nActualy, I'm a beginner in NLP so that your kernel gave me lots of help.\nI have two questions about it.\n\n1) You trained the validation data set. Can I interpret it as a transfer learning? You first trained the train dataset, and then trained the validation data set. Then, the second train (which trained the validation data set) could be considered as transfer learning?\n\n2) You said XLM-roBERTa is the best model that performs great. Then, how can I improve the performance more? If I participate in competition using tableau data, I can do lots of things such as feature engineering, hyper-parameter tunning, ensemble and so on. As I said, I'm a very beginner in NLP. How can I improve the performance in NLP competiton? Can I do feature engineering or something like that also in NLP? \n\nsorry for silly questions, Thanks for reading"
    },
    {
      "id": 816891,
      "postDate": "2020-04-22T17:25:07.190Z",
      "content": "<p>Great insights !</p>",
      "rawMarkdown": "Great insights !"
    },
    {
      "id": 814324,
      "postDate": "2020-04-20T15:22:59.157Z",
      "content": "<p>Amazing work!!</p>",
      "rawMarkdown": "Amazing work!!"
    },
    {
      "id": 814148,
      "postDate": "2020-04-20T11:54:25.140Z",
      "content": "<p>Hi there, you said there are around 120k positive labels, but I only found less than 30k in the provided data. Is there anything I missed? Thanks a lot.</p>",
      "rawMarkdown": "Hi there, you said there are around 120k positive labels, but I only found less than 30k in the provided data. Is there anything I missed? Thanks a lot.",
      "replies": [
        {
          "id": 814198,
          "postDate": "2020-04-20T12:45:19.783Z",
          "content": "<p>Did u use unintended data? Check the kernel’s training data part. </p>",
          "rawMarkdown": "Did u use unintended data? Check the kernel’s training data part. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 812971,
      "postDate": "2020-04-19T09:12:37.897Z",
      "content": "<p>Did you tried for SMOTE for sampling data and maintaining the ratio of positive to negative samples ?</p>",
      "rawMarkdown": "Did you tried for SMOTE for sampling data and maintaining the ratio of positive to negative samples ?\n"
    },
    {
      "id": 810028,
      "postDate": "2020-04-16T17:10:58.723Z",
      "content": "<p>Very good job 👍 . I like your method very much. Thanks a lot for sharing🙏 </p>",
      "rawMarkdown": "Very good job 👍 . I like your method very much. Thanks a lot for sharing🙏 "
    },
    {
      "id": 807896,
      "postDate": "2020-04-15T02:33:37.313Z",
      "content": "<p>i'm a newer</p>",
      "rawMarkdown": "i'm a newer"
    },
    {
      "id": 795196,
      "postDate": "2020-04-02T13:55:47.340Z",
      "content": "<p>How were the positive and negative data point identified? Or is it a more correct question ask, How ere the 120k positive data points selected? Or was it the other way arround?</p>",
      "rawMarkdown": "How were the positive and negative data point identified? Or is it a more correct question ask, How ere the 120k positive data points selected? Or was it the other way arround?",
      "replies": [
        {
          "id": 795281,
          "postDate": "2020-04-02T15:25:50.083Z",
          "content": "<p>Hi! The positive labels (toxic) and negative labels (non-toxic) are given in the CSV files. you just need to use pandas to filter them. you can find how to do it in the notebook.</p>",
          "rawMarkdown": "Hi! The positive labels (toxic) and negative labels (non-toxic) are given in the CSV files. you just need to use pandas to filter them. you can find how to do it in the notebook."
        }
      ]
    },
    {
      "id": 794357,
      "postDate": "2020-04-01T18:54:25.860Z",
      "content": "<p>Awesome insights, Great work <a href=\"/xhlulu\">@xhlulu</a> </p>",
      "rawMarkdown": "Awesome insights, Great work @xhlulu "
    },
    {
      "id": 793912,
      "postDate": "2020-04-01T11:57:49.273Z",
      "content": "<p>Great work. Love to know your thoughts on translating raw text. </p>",
      "rawMarkdown": "Great work. Love to know your thoughts on translating raw text. ",
      "replies": [
        {
          "id": 794045,
          "postDate": "2020-04-01T14:08:16.250Z",
          "content": "<p>I think we should start translating English data into other languages, and use that to train our model.</p>",
          "rawMarkdown": "I think we should start translating English data into other languages, and use that to train our model.",
          "votes": 5
        },
        {
          "id": 795038,
          "postDate": "2020-04-02T10:32:02.853Z",
          "content": "<p>Pretty hard with API limitations though.</p>",
          "rawMarkdown": "Pretty hard with API limitations though."
        }
      ]
    },
    {
      "id": 793703,
      "postDate": "2020-04-01T08:06:48.147Z",
      "content": "<p>Thanks for sharing this and educating me about XLM-Roberta.</p>",
      "rawMarkdown": "Thanks for sharing this and educating me about XLM-Roberta."
    },
    {
      "id": 793579,
      "postDate": "2020-04-01T05:59:22.630Z",
      "content": "<p>How are you validating your model if you are not using the validation set? i.e. how is your CV set up?</p>",
      "rawMarkdown": "How are you validating your model if you are not using the validation set? i.e. how is your CV set up?",
      "replies": [
        {
          "id": 794048,
          "postDate": "2020-04-01T14:09:42.033Z",
          "content": "<p>I'm validating the first two epochs, and for the last 2 epochs I'm not validating at all. You can easily implement CV on the validation set since it's so small, and I think it will greatly improve the performance if you add this.</p>",
          "rawMarkdown": "I'm validating the first two epochs, and for the last 2 epochs I'm not validating at all. You can easily implement CV on the validation set since it's so small, and I think it will greatly improve the performance if you add this."
        },
        {
          "id": 794445,
          "postDate": "2020-04-01T20:30:17.127Z",
          "content": "<p>You mean you suggest just to use the validation data for CV instead of training on it like you did here?</p>",
          "rawMarkdown": "You mean you suggest just to use the validation data for CV instead of training on it like you did here?"
        },
        {
          "id": 794450,
          "postDate": "2020-04-01T20:34:20.223Z",
          "content": "<p>I presume you mean k-fold CV; I think that first stage is fine (400k samples is large enough that CV would not be so useful). But since the second stage rely on only 8k samples, it is indeed a good idea to apply k-fold CV on those 8k samples, since you'd get more reliable results by blending the model output from each of the k iterations.</p>",
          "rawMarkdown": "I presume you mean k-fold CV; I think that first stage is fine (400k samples is large enough that CV would not be so useful). But since the second stage rely on only 8k samples, it is indeed a good idea to apply k-fold CV on those 8k samples, since you'd get more reliable results by blending the model output from each of the k iterations.",
          "votes": 4
        },
        {
          "id": 794521,
          "postDate": "2020-04-01T21:57:20.727Z",
          "content": "<p>Ah ok, thanks. This makes sense. Is this how you think participants should report their CV score? i.e. the k-fold CV score is the average fold score from training on validation set?</p>\n\n<p>In this competition, it seems like it different individuals might have different CVs. Some might train on validation set and use that as CV, while others report their score on the given validation set. So it's definitely confusing to determine which is the best method for CV. </p>",
          "rawMarkdown": "Ah ok, thanks. This makes sense. Is this how you think participants should report their CV score? i.e. the k-fold CV score is the average fold score from training on validation set?\n\nIn this competition, it seems like it different individuals might have different CVs. Some might train on validation set and use that as CV, while others report their score on the given validation set. So it's definitely confusing to determine which is the best method for CV. "
        },
        {
          "id": 794529,
          "postDate": "2020-04-01T22:02:58.600Z",
          "content": "<p>I don't really use CV for heavy models like transformers, so I unfortunately am not qualified to answer that question :/ I think that the best CV evaluation method will be the one that will be the closest to the true score (test AUC).</p>",
          "rawMarkdown": "I don't really use CV for heavy models like transformers, so I unfortunately am not qualified to answer that question :/ I think that the best CV evaluation method will be the one that will be the closest to the true score (test AUC).",
          "votes": 1
        },
        {
          "id": 794538,
          "postDate": "2020-04-01T22:10:04.687Z",
          "content": "<p>Hey <a href=\"/xhlulu\">@xhlulu</a> this is a nice point, is there a reason to do not use K-fold CV on heavy models other than the resource usage?</p>",
          "rawMarkdown": "Hey @xhlulu this is a nice point, is there a reason to do not use K-fold CV on heavy models other than the resource usage?"
        },
        {
          "id": 794542,
          "postDate": "2020-04-01T22:14:42.637Z",
          "content": "<p>This is interesting but I am not sure if it is necessarily true. In Google QUEST challenge, training and validation of transformers was done with a k-fold CV and there was a decent CB-LB correlation.</p>\n\n<p>Hence why I ask for this competition as well...</p>",
          "rawMarkdown": "This is interesting but I am not sure if it is necessarily true. In Google QUEST challenge, training and validation of transformers was done with a k-fold CV and there was a decent CB-LB correlation.\n\nHence why I ask for this competition as well..."
        },
        {
          "id": 794703,
          "postDate": "2020-04-02T03:00:15.723Z",
          "content": "<p>&gt;In Google QUEST challenge, training and validation of transformers was done with a k-fold CV and there was a decent CB-LB correlation.</p>\n\n<p>Don't forget that the #rows was nothing as compared to what we have here;</p>\n\n<p>When i tried the same thing on Collab with a 35GB VM instance and TPU-v2's, the MEM spiked to 24GB for a couple of secs (when i started the training)  and then back to ~10GB for the rest of the time</p>",
          "rawMarkdown": "&gt;In Google QUEST challenge, training and validation of transformers was done with a k-fold CV and there was a decent CB-LB correlation.\n\nDon't forget that the #rows was nothing as compared to what we have here;\n\nWhen i tried the same thing on Collab with a 35GB VM instance and TPU-v2's, the MEM spiked to 24GB for a couple of secs (when i started the training)  and then back to ~10GB for the rest of the time"
        }
      ]
    },
    {
      "id": 793517,
      "postDate": "2020-04-01T04:40:40.203Z",
      "content": "<p>I did <a href=\"https://www.kaggle.com/adityaecdrid/simple-xlmr-tpu-pytorch\">attempt</a> using this with PyTorch-XLA but it didn't give me a better score than my current LB; Maybe i did something incorrect for sure; Thanks!</p>",
      "rawMarkdown": "I did [attempt](https://www.kaggle.com/adityaecdrid/simple-xlmr-tpu-pytorch) using this with PyTorch-XLA but it didn't give me a better score than my current LB; Maybe i did something incorrect for sure; Thanks!",
      "replies": [
        {
          "id": 793802,
          "postDate": "2020-04-01T09:43:44.683Z",
          "content": "<p>Also tried it and ran into constant tpu errors, will play with tensorflow now I guess...</p>",
          "rawMarkdown": "Also tried it and ran into constant tpu errors, will play with tensorflow now I guess...",
          "votes": 1
        },
        {
          "id": 793927,
          "postDate": "2020-04-01T12:10:24.333Z",
          "content": "<p>It's running for me Psi, I have open sourced the xla pytorch script already above!\nIf you want to use nprocs as 8, we cannot, it's a heavy model(1.2gigs), we would need more RAM on kernel's!\nLet me know Psi if I can help with anything!</p>",
          "rawMarkdown": "It's running for me Psi, I have open sourced the xla pytorch script already above!\nIf you want to use nprocs as 8, we cannot, it's a heavy model(1.2gigs), we would need more RAM on kernel's!\nLet me know Psi if I can help with anything!"
        },
        {
          "id": 794148,
          "postDate": "2020-04-01T15:47:43.947Z",
          "content": "<p><em>Re: If you want to use nprocs as 8, we cannot</em>\nEven with the model instantiated outside of the mp_fn ? That should work.</p>",
          "rawMarkdown": "*Re: If you want to use nprocs as 8, we cannot*\nEven with the model instantiated outside of the mp_fn ? That should work."
        },
        {
          "id": 794711,
          "postDate": "2020-04-02T03:06:56.220Z",
          "content": "<p>Yep <a href=\"/mgornergoogle\">@mgornergoogle</a> ! Even if we instantiate it outside the mp_fn, it won't run with nprocs as 8 due to OOM;</p>\n\n<p>&gt; I have tried these values; MAX_LEN = 128,192; BS = 16,32 (each)</p>\n\n<p>It can run on <a href=\"https://colab.research.google.com/drive/1E1fCxlapBmmmO9y26wUsyDtxSknWua_a\">Colan NBS</a> if we are lucky to get a 35 GB instance with TPUv2;</p>",
          "rawMarkdown": "Yep @mgornergoogle ! Even if we instantiate it outside the mp_fn, it won't run with nprocs as 8 due to OOM;\n\n&gt; I have tried these values; MAX_LEN = 128,192; BS = 16,32 (each)\n\nIt can run on [Colan NBS](https://colab.research.google.com/drive/1E1fCxlapBmmmO9y26wUsyDtxSknWua_a) if we are lucky to get a 35 GB instance with TPUv2;"
        },
        {
          "id": 796833,
          "postDate": "2020-04-04T01:03:12.773Z",
          "content": "<p><a href=\"/mgornergoogle\">@mgornergoogle</a> <a href=\"https://github.com/pytorch/xla/issues/1870\">Here</a> is a GitHub issue I have started to discuss this.</p>",
          "rawMarkdown": "@mgornergoogle [Here](https://github.com/pytorch/xla/issues/1870) is a GitHub issue I have started to discuss this."
        }
      ]
    },
    {
      "id": 1329194,
      "postDate": "2021-05-31T00:41:52.647Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 827778,
      "postDate": "2020-04-30T14:35:04.907Z",
      "rawMarkdown": "",
      "votes": -3,
      "isDeleted": true
    },
    {
      "id": 813405,
      "postDate": "2020-04-19T16:36:17.663Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 807580,
      "postDate": "2020-04-14T18:12:39.083Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": 1
    },
    {
      "id": 793532,
      "postDate": "2020-04-01T05:15:19.570Z",
      "content": "<p>Good job! Thanks for sharing.</p>",
      "rawMarkdown": "Good job! Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 811696,
      "postDate": "2020-04-18T07:15:48.090Z",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": -1
    },
    {
      "id": 1288763,
      "postDate": "2021-04-30T10:09:31.800Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    },
    {
      "id": 1120748,
      "postDate": "2020-12-21T04:58:23.410Z",
      "content": "<p>Thanks for sharing this</p>",
      "rawMarkdown": "Thanks for sharing this"
    },
    {
      "id": 826184,
      "postDate": "2020-04-29T13:49:32.797Z",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks"
    },
    {
      "id": 802593,
      "postDate": "2020-04-09T16:34:28.443Z",
      "content": "<p>thanks!</p>",
      "rawMarkdown": "thanks!"
    },
    {
      "id": 802024,
      "postDate": "2020-04-09T03:59:07.803Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 794815,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2020-04-02T05:38:24.083000",
      "content": "<p>Just FYI for those who doesn't know XLM like me.</p>\n\n<p>blog: <a href=\"https://ai.facebook.com/blog/cross-lingual-pretraining/\">https://ai.facebook.com/blog/cross-lingual-pretraining/</a></p>\n\n<p>github: <a href=\"https://github.com/facebookresearch/XLM\">https://github.com/facebookresearch/XLM</a></p>\n\n<p>and paper: <a href=\"https://arxiv.org/pdf/1911.02116.pdf\">https://arxiv.org/pdf/1911.02116.pdf</a></p>",
      "votes": 10,
      "replies": [
        {
          "id": 794828,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-04-02T05:51:17.887000",
          "content": "<p>Thanks. Are you on the Game? 😃  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 795107,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-04-02T12:08:54.047000",
          "content": "<p>No 🤕 I'll be off for a few weeks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 796540,
          "author_name": "Michael Kalinin",
          "author_url": "",
          "post_date": "2020-04-03T17:12:56.897000",
          "content": "<p>this is also good reading <a href=\"https://towardsdatascience.com/xlm-enhancing-bert-for-cross-lingual-language-model-5aeed9e6f14b\">https://towardsdatascience.com/xlm-enhancing-bert-for-cross-lingual-language-model-5aeed9e6f14b</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 793925,
      "author_name": "Abhishek Thakur",
      "author_url": "",
      "post_date": "2020-04-01T12:07:24.760000",
      "content": "<p>Great Work! But where does the randomness come from? It seems like the code gives anywhere between 0.92 and 0.94? </p>",
      "votes": 7,
      "replies": [
        {
          "id": 794025,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-04-01T13:47:59.437000",
          "content": "<p>The validation set is pretty small and randomly split, so the order in which data is fed affects how well the final model performs. When you are translating the test set into English, the model depends a lot more on the training set, which is huge, so there's very little variation in the final performance.</p>",
          "votes": 12,
          "replies": []
        },
        {
          "id": 801764,
          "author_name": "Anshul Dhingra",
          "author_url": "",
          "post_date": "2020-04-08T19:21:50.090000",
          "content": "<p>+1</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 815328,
      "author_name": "Nischay Dhankhar",
      "author_url": "",
      "post_date": "2020-04-21T13:10:56.103000",
      "content": "<p>Competition be like: 'Am I a joke to you'??😄 \nGreat work though</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 794269,
      "author_name": "Martin Görner",
      "author_url": "",
      "post_date": "2020-04-01T17:40:26.527000",
      "content": "<p>Amazing work.\nTweeted: <a href=\"https://twitter.com/martin_gorner/status/1245405207762046976\">https://twitter.com/martin_gorner/status/1245405207762046976</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 794441,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-04-01T20:27:24.837000",
          "content": "<p>Thanks for the shout out!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 794113,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2020-04-01T15:13:01.717000",
      "content": "<p>Great job! that last point is very interesting, it might be a little late for that, but I think the developed solution would be more interesting for the host if the competition had a hidden test set with many languages, maybe even on that only exists on that set.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 804249,
      "author_name": "Ashutosh Anand",
      "author_url": "",
      "post_date": "2020-04-11T12:00:32.293000",
      "content": "<p>Amazing work! I took the part in the competition as well to improve upon my skills, reading the description of your model really helped!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 794431,
      "author_name": "Camaro",
      "author_url": "",
      "post_date": "2020-04-01T20:10:46.827000",
      "content": "<p>Great! I'm glad to know translation won't be the final solution😂 </p>",
      "votes": 1,
      "replies": [
        {
          "id": 794440,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-04-01T20:27:12.870000",
          "content": "<p>Who knows... I think that translation will definitely play a role in this... Maybe just in a different way than what has been done so far? ;)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 793514,
      "author_name": "Anshuman Singh",
      "author_url": "",
      "post_date": "2020-04-01T04:35:28.893000",
      "content": "<p>Thanks for sharing the insights, they helped😄. In my observation when training on the validation data, the loss correlated more than when training on the train data, looks like they are from different distributions.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 793503,
      "author_name": "Chun Ming Lee",
      "author_url": "",
      "post_date": "2020-04-01T04:27:18.517000",
      "content": "<p>Very nice! Just came across XLM-R myself. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 795039,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2020-04-02T10:33:47.867000",
      "content": "<p>Great kernel. I am trying to understand how exactly the training loop is working with distributed TPU usage. I assume the code is running on 8 processes, right? What exactly happens here if I call <code>model.fit()</code>? Are the batches of an epoch distributed?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 795282,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-04-02T15:26:20.817000",
          "content": "<p>I presume they are. I think <a href=\"/mgornergoogle\">@mgornergoogle</a> would be a better person to answer this :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 795306,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-04-02T15:53:47.847000",
          "content": "<p>Yes, The model is replicated on the 8 cores. The batch you define through tf.data.Dataset will get split across the 8 cores of the TPU. Gradient updates are through an all-reduce algorithm running on the TPU board.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 795957,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-04-03T06:39:09.757000",
          "content": "<p>Thanks <a href=\"/mgornergoogle\">@mgornergoogle</a>. I have two follow up questions if you don't mind:</p>\n\n<ul>\n<li>I see datasets using <code>.split()</code> method to duplicate them infinitely. What about using <code>.shard()</code>? I don't see it used in the kernels here on Kaggle, but is it useful?</li>\n<li>What exactly happens in callbacks that for example operate <code>on_batch_end</code> level and utilize fit parameters like <code>global_step</code> or <code>total_steps</code>. Are those per process only, or really globally? So let's say with 8 processes I am running 16 batches per epoch. Batch 1-8 get distributed across processes, and 9-16 get distributed. If <code>on_batch_end</code> is called for each batch, is it 1,2,3,...,16 or is it 1,1,1,1,...,2,2,2,2...? Hope it is clear what I am trying to ask :)</li>\n</ul>\n\n<p>Edit: I just figured that you said the batch gets split, so a single batch gets split? My second question is redundant then I guess, I thought separate batches get distributed.</p>\n\n<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 796452,
          "author_name": "Martin Görner",
          "author_url": "",
          "post_date": "2020-04-03T15:37:55.500000",
          "content": "<p>.shard() is called by tf.data.Dataset() automatically when running a model in a TPUStrategy.\nActually, when running on a TPU pod, tf.data.Dataset will even attempt to shard your set of input files between TPU boards, then shard the batch between the 8 TPU cores of a board. If your dataset is made of a sufficient number of files, the benefit will be that only the required files will be loaded on each TPU board.</p>\n\n<p>Keras callbacks are processed on the Kaggle VM, so there are no parallelism issues there.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 794074,
      "author_name": "Vikas Singh",
      "author_url": "",
      "post_date": "2020-04-01T14:33:52.437000",
      "content": "<p>Great work guys! Learning a lot from you guys <a href=\"/abhishek\">@abhishek</a> and <a href=\"/xhlulu\">@xhlulu</a>. Thank you!  </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 793533,
      "author_name": "Aditya Soni",
      "author_url": "",
      "post_date": "2020-04-01T05:16:32.843000",
      "content": "<ul>\n<li>Another very strange thing i noticed is that the Batch-Size somehow effects the feature vectors wrt XLM-R and huggingface; Not sure why;</li>\n</ul>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 793501,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-04-01T04:25:07.263000",
      "content": "<p>Great job! Impressive building a model without translating. Thanks for sharing.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 808433,
      "author_name": "Shahules",
      "author_url": "",
      "post_date": "2020-04-15T12:26:46.880000",
      "content": "<p>Were you able to get deterministic results with Tensorflow TPU? I am getting varied scores and losses every time I run and commit  <a href=\"/xhlulu\">@xhlulu</a> </p>",
      "votes": 0,
      "replies": [
        {
          "id": 808689,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2020-04-15T15:00:16.497000",
          "content": "<p><a href=\"/shahules\">@shahules</a> There is no way as far as I know.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 808761,
          "author_name": "Shahules",
          "author_url": "",
          "post_date": "2020-04-15T15:42:49.427000",
          "content": "<p>oh,I went through their documentation. Anyway,one more reasons to switch to pytorch!! <a href=\"/bamps53\">@bamps53</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 809261,
          "author_name": "Camaro",
          "author_url": "",
          "post_date": "2020-04-16T02:29:58.947000",
          "content": "<p><a href=\"/shahules\">@shahules</a> Yes, but currently pytorch xla is soooo frustrating to debug the code. I recommend you to stick to GPU if you choose TPU☹️  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 809264,
          "author_name": "Shahules",
          "author_url": "",
          "post_date": "2020-04-16T02:35:06.150000",
          "content": "<p>Thanks buddy <a href=\"/bamps53\">@bamps53</a> I will keep that in mind.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2781804,
      "author_name": "Md. Sadman Sakib Mahib",
      "author_url": "",
      "post_date": "2024-04-29T02:41:44.457000",
      "content": "<p>Great Work ! :3</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1108133,
      "author_name": "Eric McLachlan",
      "author_url": "",
      "post_date": "2020-12-10T10:06:09.033000",
      "content": "<p><a href=\"https://www.kaggle.com/xhlulu\" target=\"_blank\">@xhlulu</a>: I'm <em>really</em> late to the party but nevertheless wanted to say thank you for sharing the above.  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1091014,
      "author_name": "Mat",
      "author_url": "",
      "post_date": "2020-11-25T18:16:35.793000",
      "content": "<p>Great !              </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 842705,
      "author_name": "Data-Science Sean",
      "author_url": "",
      "post_date": "2020-05-11T15:26:44.870000",
      "content": "<p>Xhlulu, congratulations again on the score.  Reading over your notes, I was wondering how you would deploy this model in the real world.  How would your model change and adjust to knew information?  How would the model receive and apply feed back?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 828727,
      "author_name": "Baek Kyun Shin",
      "author_url": "",
      "post_date": "2020-05-01T08:44:07.910000",
      "content": "<p>Thank you for sharing your wonderful insight!\nActualy, I'm a beginner in NLP so that your kernel gave me lots of help.\nI have two questions about it.</p>\n\n<p>1) You trained the validation data set. Can I interpret it as a transfer learning? You first trained the train dataset, and then trained the validation data set. Then, the second train (which trained the validation data set) could be considered as transfer learning?</p>\n\n<p>2) You said XLM-roBERTa is the best model that performs great. Then, how can I improve the performance more? If I participate in competition using tableau data, I can do lots of things such as feature engineering, hyper-parameter tunning, ensemble and so on. As I said, I'm a very beginner in NLP. How can I improve the performance in NLP competiton? Can I do feature engineering or something like that also in NLP? </p>\n\n<p>sorry for silly questions, Thanks for reading</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 816891,
      "author_name": "Thomas Soler",
      "author_url": "",
      "post_date": "2020-04-22T17:25:07.190000",
      "content": "<p>Great insights !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 814324,
      "author_name": "Mayukh Sil",
      "author_url": "",
      "post_date": "2020-04-20T15:22:59.157000",
      "content": "<p>Amazing work!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 814148,
      "author_name": "Seawolf",
      "author_url": "",
      "post_date": "2020-04-20T11:54:25.140000",
      "content": "<p>Hi there, you said there are around 120k positive labels, but I only found less than 30k in the provided data. Is there anything I missed? Thanks a lot.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 814198,
          "author_name": "DHZM",
          "author_url": "",
          "post_date": "2020-04-20T12:45:19.783000",
          "content": "<p>Did u use unintended data? Check the kernel’s training data part. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 812971,
      "author_name": "deepak jindal",
      "author_url": "",
      "post_date": "2020-04-19T09:12:37.897000",
      "content": "<p>Did you tried for SMOTE for sampling data and maintaining the ratio of positive to negative samples ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 810028,
      "author_name": "Seawolf",
      "author_url": "",
      "post_date": "2020-04-16T17:10:58.723000",
      "content": "<p>Very good job 👍 . I like your method very much. Thanks a lot for sharing🙏 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 807896,
      "author_name": "ericgan2020",
      "author_url": "",
      "post_date": "2020-04-15T02:33:37.313000",
      "content": "<p>i'm a newer</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 795196,
      "author_name": "Data-Science Sean",
      "author_url": "",
      "post_date": "2020-04-02T13:55:47.340000",
      "content": "<p>How were the positive and negative data point identified? Or is it a more correct question ask, How ere the 120k positive data points selected? Or was it the other way arround?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 795281,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-04-02T15:25:50.083000",
          "content": "<p>Hi! The positive labels (toxic) and negative labels (non-toxic) are given in the CSV files. you just need to use pandas to filter them. you can find how to do it in the notebook.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 794357,
      "author_name": "Mayank",
      "author_url": "",
      "post_date": "2020-04-01T18:54:25.860000",
      "content": "<p>Awesome insights, Great work <a href=\"/xhlulu\">@xhlulu</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 793912,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-04-01T11:57:49.273000",
      "content": "<p>Great work. Love to know your thoughts on translating raw text. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 794045,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2020-04-01T14:08:16.250000",
          "content": "<p>I think we should start translating English data into other languages, and use that to train our model.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 795038,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2020-04-02T10:32:02.853000",
          "content": "<p>Pretty hard with API limitations though.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 793703,
      "author_name": "Andy Penrose",
      "author_url": "",
      "post_date": "2020-04-01T08:06:48.147000",
      "content": "<p>Thanks for sharing this and educating me about XLM-Roberta.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 793579,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-01T05:59:22.630000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 794048,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-01T14:09:42.033000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794445,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-01T20:30:17.127000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794450,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-01T20:34:20.223000",
          "content": "",
          "votes": 4,
          "replies": []
        },
        {
          "id": 794521,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-01T21:57:20.727000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794529,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-01T22:02:58.600000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 794538,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-01T22:10:04.687000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794542,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-01T22:14:42.637000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794703,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-02T03:00:15.723000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 793517,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-01T04:40:40.203000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 793802,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-01T09:43:44.683000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 793927,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-01T12:10:24.333000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794148,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-01T15:47:43.947000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794711,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-02T03:06:56.220000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 796833,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-04T01:03:12.773000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1329194,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-31T00:41:52.647000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 827778,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-30T14:35:04.907000",
      "content": "",
      "votes": -3,
      "replies": []
    },
    {
      "id": 813405,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-19T16:36:17.663000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 807580,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-14T18:12:39.083000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 793532,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-01T05:15:19.570000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 811696,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-18T07:15:48.090000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 1288763,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-04-30T10:09:31.800000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1120748,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-21T04:58:23.410000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 826184,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-29T13:49:32.797000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 802593,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-09T16:34:28.443000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 802024,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-09T03:59:07.803000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "793479": "Hello.\n\nSome of you might have just noticed me releasing my [XLM-Roberta notebook](https://www.kaggle.com/xhlulu/jigsaw-tpu-xlm-roberta) that achieved 0.9383. Since the notebook takes a while to run, I wanted to share my insights in making this notebook as a discussion post (I'll integrate this to my notebook later):\n* **Use a state of the art multilingual transformer**: As you may know, mBERT has been around for a while (almost 2 years now). Since it was released, a lot of research has been done on improving the pre-training step in order to create a strong cross-lingual language model. One major paper in that subject is [Cross-lingual Language Model Pretraining\n](https://arxiv.org/pdf/1901.07291.pdf) which introduced the XLM model. It was then further improved through efficient pretraining in [Unsupervised Cross-lingual Representation Learning at Scale](https://arxiv.org/abs/1901.07291); this last model is called *XLM-Roberta*, which is what I used in my notebook. It is available in the transformers library, so it was very easy to set up.\n* **Balance your data**: The training data set consists of over 2M data points, but it is heavily unbalanced since it only has around 120k positive labels. In order to limit training time, I selected all of the positive labels, and subsampled the negative labels such that I have in total 400k labels, which is a lot closer to the class frequency of the validation labels.\n* **Train your model on the validation set**: I trained two epochs on the 400k training samples, followed by 2 more epochs on the relatively small validation set. My hypothesis is that this two-steps process improves the model's ability to first learn the underlying structure of toxic comments (by training on 400k english comments), then adapt those learned structures to discriminate other languages from the validation set (Turkish, Italian, Portuguese, etc.).\n* **Translation is not the (only) solution**: Although notebooks have shown great results by translating the test set, I decided to stick to my guns and focus on training a cross-lingual model. I'm sure that people will find ways to improve this model through translation, but I think that if we were to deploy a toxicity detection model in production, we won't always have the luxury to request Google translate API calls or run two transformers in parallel. Additionally, Google translate might not yield good results for low-resource languages (e.g. Swahili), but models like XLM-R might still perform fairly well.",
    "794815": "Just FYI for those who doesn't know XLM like me.\n\nblog: https://ai.facebook.com/blog/cross-lingual-pretraining/\n\ngithub: https://github.com/facebookresearch/XLM\n\nand paper: https://arxiv.org/pdf/1911.02116.pdf",
    "793925": "Great Work! But where does the randomness come from? It seems like the code gives anywhere between 0.92 and 0.94? ",
    "815328": "Competition be like: 'Am I a joke to you'??😄 \nGreat work though",
    "794269": "Amazing work.\nTweeted: https://twitter.com/martin_gorner/status/1245405207762046976",
    "794113": "Great job! that last point is very interesting, it might be a little late for that, but I think the developed solution would be more interesting for the host if the competition had a hidden test set with many languages, maybe even on that only exists on that set.",
    "804249": "Amazing work! I took the part in the competition as well to improve upon my skills, reading the description of your model really helped!",
    "794431": "Great! I'm glad to know translation won't be the final solution😂 ",
    "793514": "Thanks for sharing the insights, they helped😄. In my observation when training on the validation data, the loss correlated more than when training on the train data, looks like they are from different distributions.",
    "793503": "Very nice! Just came across XLM-R myself. ",
    "795039": "Great kernel. I am trying to understand how exactly the training loop is working with distributed TPU usage. I assume the code is running on 8 processes, right? What exactly happens here if I call `model.fit()`? Are the batches of an epoch distributed?",
    "794074": "Great work guys! Learning a lot from you guys @abhishek and @xhlulu. Thank you!  ",
    "793533": "+ Another very strange thing i noticed is that the Batch-Size somehow effects the feature vectors wrt XLM-R and huggingface; Not sure why;",
    "793501": "Great job! Impressive building a model without translating. Thanks for sharing.",
    "808433": "Were you able to get deterministic results with Tensorflow TPU? I am getting varied scores and losses every time I run and commit  @xhlulu ",
    "2781804": "Great Work ! :3",
    "1108133": "@xhlulu: I'm *really* late to the party but nevertheless wanted to say thank you for sharing the above.  ",
    "1091014": "Great !              ",
    "842705": "Xhlulu, congratulations again on the score.  Reading over your notes, I was wondering how you would deploy this model in the real world.  How would your model change and adjust to knew information?  How would the model receive and apply feed back?",
    "828727": "Thank you for sharing your wonderful insight!\nActualy, I'm a beginner in NLP so that your kernel gave me lots of help.\nI have two questions about it.\n\n1) You trained the validation data set. Can I interpret it as a transfer learning? You first trained the train dataset, and then trained the validation data set. Then, the second train (which trained the validation data set) could be considered as transfer learning?\n\n2) You said XLM-roBERTa is the best model that performs great. Then, how can I improve the performance more? If I participate in competition using tableau data, I can do lots of things such as feature engineering, hyper-parameter tunning, ensemble and so on. As I said, I'm a very beginner in NLP. How can I improve the performance in NLP competiton? Can I do feature engineering or something like that also in NLP? \n\nsorry for silly questions, Thanks for reading",
    "816891": "Great insights !",
    "814324": "Amazing work!!",
    "814148": "Hi there, you said there are around 120k positive labels, but I only found less than 30k in the provided data. Is there anything I missed? Thanks a lot.",
    "812971": "Did you tried for SMOTE for sampling data and maintaining the ratio of positive to negative samples ?\n",
    "810028": "Very good job 👍 . I like your method very much. Thanks a lot for sharing🙏 ",
    "807896": "i'm a newer",
    "795196": "How were the positive and negative data point identified? Or is it a more correct question ask, How ere the 120k positive data points selected? Or was it the other way arround?",
    "794357": "Awesome insights, Great work @xhlulu ",
    "793912": "Great work. Love to know your thoughts on translating raw text. ",
    "793703": "Thanks for sharing this and educating me about XLM-Roberta.",
    "793579": "How are you validating your model if you are not using the validation set? i.e. how is your CV set up?",
    "793517": "I did [attempt](https://www.kaggle.com/adityaecdrid/simple-xlmr-tpu-pytorch) using this with PyTorch-XLA but it didn't give me a better score than my current LB; Maybe i did something incorrect for sure; Thanks!",
    "1329194": "",
    "827778": "",
    "813405": "",
    "807580": "Thanks for sharing",
    "793532": "Good job! Thanks for sharing.",
    "811696": "Thanks!",
    "1288763": "Thanks for sharing",
    "1120748": "Thanks for sharing this",
    "826184": "Thanks",
    "802593": "thanks!",
    "802024": "Thanks for sharing!"
  }
}