{
  "id": 160856,
  "title": "Thread to share things that didnt work",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/160856",
  "author_name": "",
  "post_date": "2020-06-23T00:22:12.807695200Z",
  "votes": 13,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Congrats to everyone who made it through to the end. Throughout these competitions all of us try a whole bunch of experiments and many of them fail. It would be nice to start a thread where we talk about the things that didn't work and potentially discuss why they failed or ways to make them work. </p>\n\n<p>To start off I will list things off and then fill in more details at a later time. </p>\n\n<p>Things we tried, but failed:\n- LASER embeddings\n  - These looked reasonably promising based on the <a href=\"https://arxiv.org/abs/1812.10464\">paper</a>, not necessarily as embeddings for our primary model but maybe as some auxiliary signal to use in an ensemble. Simply didn't have enough discriminative power</p>\n\n<ul>\n<li>Sampling data from our sets based on similarity to test set samples\n<ul><li>We tried to find the best training points to use by outputting the XLM embeddings and then doing a similarity search with FAISS. In theory we could pull out the samples closest to the test set and train to accommodate that set. Turns out these embeddings don't really line up properly across languages. Maybe LASER embeddings would've been better here.</li></ul></li>\n<li>Training monolingual models and throwing into our ensemble\n<ul><li>There are lots of readily available monolingual BERT models and other variants in the huggingface model collection, but it was unclear how exactly we blend these together with the other models because their distributions would not line up across models so turkish predictions might show up in a significantly different place than italian for example and that would screw up the AUC. There might be some way to align the ones we had validation samples for but some of the languages we did not have non-translated validation samples for.</li></ul></li>\n<li>Training on the unintended bias data\n<ul><li>This yielded a consistently worse model. I believe this to be because the labeling process might have been slightly different between the sets. It would have been nice to be able to utilize this better, but we did not find a way to make this bring us a boost. </li></ul></li>\n<li>Training on more than just the toxic label\n<ul><li>in past competitions we have had good luck training on the whole array of labels (toxic, severe toxic, threatening, etc.), but in this case with these massive XLM models it did not seem to make a difference because our models already have so much language understanding baked in. We did have some models that trained on multiple outputs but it did not yield big gains like in previous competitions. </li></ul></li>\n<li>Generating our own translations with models from FAIR seq.\n<ul><li>In the last couple days we tried doing translations into russian and french based on the models that were available. It seemed like our performance was somewhat capped by the quality of translations we used so new translations beyond the google translate ones seemed useful, but it was difficult to measure any gain because we did not have russian or french validation samples and hadn't previously set aside any. </li></ul></li>\n<li>Dropping mislabeled samples\n<ul><li>One of the things we noticed was even after training on the validation set for finetuning it would still get certain ones wrong with a high degree of confidence. We viewed these points and they looked like mislabels. We thought it might be good to clean those out of the training set as well. We did this by training on the data. Finding samples that had a very high error even after several epochs and then dropping those samples and retraining a model from scratch just on those samples. This led to amazing hyperconvergence of the model, achieving almost exactly the same performance of the previous model in the first epoch, but it did not improve upon previous results because it was essentially just trained to agree with the prior model. </li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "897537",
      "postDate": "06/23/2020 00:22:12",
      "content": "<p>Congrats to everyone who made it through to the end. Throughout these competitions all of us try a whole bunch of experiments and many of them fail. It would be nice to start a thread where we talk about the things that didn't work and potentially discuss why they failed or ways to make them work. </p>\n\n<p>To start off I will list things off and then fill in more details at a later time. </p>\n\n<p>Things we tried, but failed:\n- LASER embeddings\n  - These looked reasonably promising based on the <a href=\"https://arxiv.org/abs/1812.10464\">paper</a>, not necessarily as embeddings for our primary model but maybe as some auxiliary signal to use in an ensemble. Simply didn't have enough discriminative power</p>\n\n<ul>\n<li>Sampling data from our sets based on similarity to test set samples\n<ul><li>We tried to find the best training points to use by outputting the XLM embeddings and then doing a similarity search with FAISS. In theory we could pull out the samples closest to the test set and train to accommodate that set. Turns out these embeddings don't really line up properly across languages. Maybe LASER embeddings would've been better here.</li></ul></li>\n<li>Training monolingual models and throwing into our ensemble\n<ul><li>There are lots of readily available monolingual BERT models and other variants in the huggingface model collection, but it was unclear how exactly we blend these together with the other models because their distributions would not line up across models so turkish predictions might show up in a significantly different place than italian for example and that would screw up the AUC. There might be some way to align the ones we had validation samples for but some of the languages we did not have non-translated validation samples for.</li></ul></li>\n<li>Training on the unintended bias data\n<ul><li>This yielded a consistently worse model. I believe this to be because the labeling process might have been slightly different between the sets. It would have been nice to be able to utilize this better, but we did not find a way to make this bring us a boost. </li></ul></li>\n<li>Training on more than just the toxic label\n<ul><li>in past competitions we have had good luck training on the whole array of labels (toxic, severe toxic, threatening, etc.), but in this case with these massive XLM models it did not seem to make a difference because our models already have so much language understanding baked in. We did have some models that trained on multiple outputs but it did not yield big gains like in previous competitions. </li></ul></li>\n<li>Generating our own translations with models from FAIR seq.\n<ul><li>In the last couple days we tried doing translations into russian and french based on the models that were available. It seemed like our performance was somewhat capped by the quality of translations we used so new translations beyond the google translate ones seemed useful, but it was difficult to measure any gain because we did not have russian or french validation samples and hadn't previously set aside any. </li></ul></li>\n<li>Dropping mislabeled samples\n<ul><li>One of the things we noticed was even after training on the validation set for finetuning it would still get certain ones wrong with a high degree of confidence. We viewed these points and they looked like mislabels. We thought it might be good to clean those out of the training set as well. We did this by training on the data. Finding samples that had a very high error even after several epochs and then dropping those samples and retraining a model from scratch just on those samples. This led to amazing hyperconvergence of the model, achieving almost exactly the same performance of the previous model in the first epoch, but it did not improve upon previous results because it was essentially just trained to agree with the prior model. </li></ul></li>\n</ul>",
      "rawMarkdown": "Congrats to everyone who made it through to the end. Throughout these competitions all of us try a whole bunch of experiments and many of them fail. It would be nice to start a thread where we talk about the things that didn't work and potentially discuss why they failed or ways to make them work. \n\nTo start off I will list things off and then fill in more details at a later time. \n\nThings we tried, but failed:\n- LASER embeddings\n  - These looked reasonably promising based on the [paper](https://arxiv.org/abs/1812.10464), not necessarily as embeddings for our primary model but maybe as some auxiliary signal to use in an ensemble. Simply didn't have enough discriminative power\n\n- Sampling data from our sets based on similarity to test set samples\n  -  We tried to find the best training points to use by outputting the XLM embeddings and then doing a similarity search with FAISS. In theory we could pull out the samples closest to the test set and train to accommodate that set. Turns out these embeddings don't really line up properly across languages. Maybe LASER embeddings would've been better here.\n- Training monolingual models and throwing into our ensemble\n  - There are lots of readily available monolingual BERT models and other variants in the huggingface model collection, but it was unclear how exactly we blend these together with the other models because their distributions would not line up across models so turkish predictions might show up in a significantly different place than italian for example and that would screw up the AUC. There might be some way to align the ones we had validation samples for but some of the languages we did not have non-translated validation samples for.\n- Training on the unintended bias data\n  - This yielded a consistently worse model. I believe this to be because the labeling process might have been slightly different between the sets. It would have been nice to be able to utilize this better, but we did not find a way to make this bring us a boost. \n- Training on more than just the toxic label\n  - in past competitions we have had good luck training on the whole array of labels (toxic, severe toxic, threatening, etc.), but in this case with these massive XLM models it did not seem to make a difference because our models already have so much language understanding baked in. We did have some models that trained on multiple outputs but it did not yield big gains like in previous competitions. \n- Generating our own translations with models from FAIR seq.\n  - In the last couple days we tried doing translations into russian and french based on the models that were available. It seemed like our performance was somewhat capped by the quality of translations we used so new translations beyond the google translate ones seemed useful, but it was difficult to measure any gain because we did not have russian or french validation samples and hadn't previously set aside any. \n- Dropping mislabeled samples\n  - One of the things we noticed was even after training on the validation set for finetuning it would still get certain ones wrong with a high degree of confidence. We viewed these points and they looked like mislabels. We thought it might be good to clean those out of the training set as well. We did this by training on the data. Finding samples that had a very high error even after several epochs and then dropping those samples and retraining a model from scratch just on those samples. This led to amazing hyperconvergence of the model, achieving almost exactly the same performance of the previous model in the first epoch, but it did not improve upon previous results because it was essentially just trained to agree with the prior model.",
      "votes": null
    },
    {
      "id": "897551",
      "postDate": "06/23/2020 00:31:26",
      "content": "<ul>\n<li><p>NLP augmentations (although I suspect we didn't try enough)</p></li>\n<li><p>Multi-sample dropout</p></li>\n<li><p>Weighted mean pooling of XLM-Roberta layers</p></li>\n<li><p>Gradual unfreezing</p></li>\n<li><p>Training first on toxic then on unintended bias data</p></li>\n<li><p>Keeping validation.csv for validation set</p></li>\n</ul>",
      "rawMarkdown": "NLP augmentations (although I suspect we didn't try enough)\n\n- Multi-sample dropout\n\n- Weighted mean pooling of XLM-Roberta layers\n\n- Gradual unfreezing\n\n- Training first on toxic then on unintended bias data\n\n- Keeping validation.csv for validation set",
      "votes": null
    },
    {
      "id": "897554",
      "postDate": "06/23/2020 00:33:56",
      "content": "<p>Similar experiences with all of those. What did you do for validation instead of using the validation file?</p>",
      "rawMarkdown": "Similar experiences with all of those. What did you do for validation instead of using the validation file?",
      "votes": null
    },
    {
      "id": "897557",
      "postDate": "06/23/2020 00:38:24",
      "content": "<ol>\n<li>NLP augmentations</li>\n<li>Ensemble translated test data</li>\n<li>Pseudo labeling</li>\n<li>KNN of Doc2vec to find the most similar comment of external data and test data</li>\n</ol>",
      "rawMarkdown": "1. NLP augmentations\n2. Ensemble translated test data\n3. Pseudo labeling\n4. KNN of Doc2vec to find the most similar comment of external data and test data",
      "votes": null
    },
    {
      "id": "897572",
      "postDate": "06/23/2020 00:54:16",
      "content": "<ul>\n<li>Unsupervised Data Augmentation\nThis also looks promising as seen in <a href=\"https://arxiv.org/abs/1904.12848\">the original paper</a> and <a href=\"https://arxiv.org/abs/1909.07009\">the one on cross-lingual text classification</a>. The idea also makes sense. The loss objective is to minimize the difference between the output of 2 languages. I implement one using <code>tf.keras.losses.KLDivergence</code> and using opensubs parallel data. After wasting so many hours trying to make it work, it doesn't bring any significant value to my performance :( </li>\n</ul>\n\n<p>or maybe it's just my bad implementation ¯_(ツ)_/¯ I'll check later after the test set is released.</p>",
      "rawMarkdown": "* Unsupervised Data Augmentation\nThis also looks promising as seen in [the original paper](https://arxiv.org/abs/1904.12848) and [the one on cross-lingual text classification](https://arxiv.org/abs/1909.07009). The idea also makes sense. The loss objective is to minimize the difference between the output of 2 languages. I implement one using `tf.keras.losses.KLDivergence` and using opensubs parallel data. After wasting so many hours trying to make it work, it doesn't bring any significant value to my performance :( \n\nor maybe it's just my bad implementation ¯\\_(ツ)_/¯ I'll check later after the test set is released.",
      "votes": null
    },
    {
      "id": "897576",
      "postDate": "06/23/2020 01:00:36",
      "content": "<p>For me did not work fancy head on top of the transformers, I probably wasted too many TPU hours here.</p>",
      "rawMarkdown": "For me did not work fancy head on top of the transformers, I probably wasted too many TPU hours here.",
      "votes": null
    },
    {
      "id": "897586",
      "postDate": "06/23/2020 01:11:16",
      "content": "<p>Even we tried it. In our case it gave little boost on single models and made results consistent. </p>",
      "rawMarkdown": "Even we tried it. In our case it gave little boost on single models and made results consistent.",
      "votes": null
    },
    {
      "id": "897698",
      "postDate": "06/23/2020 03:29:45",
      "content": "<p>Hey, nice to know someone else tried it too! Perhaps it's just my bad implementation. </p>\n\n<p>I would like to know. How do you combine the UDA loss with BCE loss? Do you just sum it up? Do you also implement the prediction probability thresholding and Training Signal Annealing (TSA)?</p>",
      "rawMarkdown": "Hey, nice to know someone else tried it too! Perhaps it's just my bad implementation. \n\nI would like to know. How do you combine the UDA loss with BCE loss? Do you just sum it up? Do you also implement the prediction probability thresholding and Training Signal Annealing (TSA)?",
      "votes": null
    },
    {
      "id": "897705",
      "postDate": "06/23/2020 03:40:42",
      "content": "<p>Yes we just sumed it up with a coefficient to KL divergence. No we didn't do tsa or probability thresholding.</p>",
      "rawMarkdown": "Yes we just sumed it up with a coefficient to KL divergence. No we didn't do tsa or probability thresholding.",
      "votes": null
    },
    {
      "id": "898018",
      "postDate": "06/23/2020 08:28:53",
      "content": "<p>We trust the LB, and in the beginning we only trained for one epoch for fear of overfitting. It worked quite well since we achieved scores on the LB for a single model of around 0.9425.</p>",
      "rawMarkdown": "We trust the LB, and in the beginning we only trained for one epoch for fear of overfitting. It worked quite well since we achieved scores on the LB for a single model of around 0.9425.",
      "votes": null
    },
    {
      "id": "898023",
      "postDate": "06/23/2020 08:30:50",
      "content": "<p>Pseudo labeling did work for one of our model. It gave a small boost, but we didn't select it in the end since we found a better solution.</p>",
      "rawMarkdown": "Pseudo labeling did work for one of our model. It gave a small boost, but we didn't select it in the end since we found a better solution.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 897551,
      "author_name": "rftexas",
      "author_url": "",
      "post_date": "06/23/2020 00:31:26",
      "content": "<ul>\n<li><p>NLP augmentations (although I suspect we didn't try enough)</p></li>\n<li><p>Multi-sample dropout</p></li>\n<li><p>Weighted mean pooling of XLM-Roberta layers</p></li>\n<li><p>Gradual unfreezing</p></li>\n<li><p>Training first on toxic then on unintended bias data</p></li>\n<li><p>Keeping validation.csv for validation set</p></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 897554,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "06/23/2020 00:33:56",
          "content": "<p>Similar experiences with all of those. What did you do for validation instead of using the validation file?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 898018,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "06/23/2020 08:28:53",
          "content": "<p>We trust the LB, and in the beginning we only trained for one epoch for fear of overfitting. It worked quite well since we achieved scores on the LB for a single model of around 0.9425.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 897557,
      "author_name": "medrau",
      "author_url": "",
      "post_date": "06/23/2020 00:38:24",
      "content": "<ol>\n<li>NLP augmentations</li>\n<li>Ensemble translated test data</li>\n<li>Pseudo labeling</li>\n<li>KNN of Doc2vec to find the most similar comment of external data and test data</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 898023,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "06/23/2020 08:30:50",
          "content": "<p>Pseudo labeling did work for one of our model. It gave a small boost, but we didn't select it in the end since we found a better solution.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 897572,
      "author_name": "ilhamfp31",
      "author_url": "",
      "post_date": "06/23/2020 00:54:16",
      "content": "<ul>\n<li>Unsupervised Data Augmentation\nThis also looks promising as seen in <a href=\"https://arxiv.org/abs/1904.12848\">the original paper</a> and <a href=\"https://arxiv.org/abs/1909.07009\">the one on cross-lingual text classification</a>. The idea also makes sense. The loss objective is to minimize the difference between the output of 2 languages. I implement one using <code>tf.keras.losses.KLDivergence</code> and using opensubs parallel data. After wasting so many hours trying to make it work, it doesn't bring any significant value to my performance :( </li>\n</ul>\n\n<p>or maybe it's just my bad implementation ¯_(ツ)_/¯ I'll check later after the test set is released.</p>",
      "votes": null,
      "replies": [
        {
          "id": 897586,
          "author_name": "ubamba98",
          "author_url": "",
          "post_date": "06/23/2020 01:11:16",
          "content": "<p>Even we tried it. In our case it gave little boost on single models and made results consistent. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 897698,
          "author_name": "ilhamfp31",
          "author_url": "",
          "post_date": "06/23/2020 03:29:45",
          "content": "<p>Hey, nice to know someone else tried it too! Perhaps it's just my bad implementation. </p>\n\n<p>I would like to know. How do you combine the UDA loss with BCE loss? Do you just sum it up? Do you also implement the prediction probability thresholding and Training Signal Annealing (TSA)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 897705,
          "author_name": "ubamba98",
          "author_url": "",
          "post_date": "06/23/2020 03:40:42",
          "content": "<p>Yes we just sumed it up with a coefficient to KL divergence. No we didn't do tsa or probability thresholding.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 897576,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "06/23/2020 01:00:36",
      "content": "<p>For me did not work fancy head on top of the transformers, I probably wasted too many TPU hours here.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "897537": "Congrats to everyone who made it through to the end. Throughout these competitions all of us try a whole bunch of experiments and many of them fail. It would be nice to start a thread where we talk about the things that didn't work and potentially discuss why they failed or ways to make them work. \n\nTo start off I will list things off and then fill in more details at a later time. \n\nThings we tried, but failed:\n- LASER embeddings\n  - These looked reasonably promising based on the [paper](https://arxiv.org/abs/1812.10464), not necessarily as embeddings for our primary model but maybe as some auxiliary signal to use in an ensemble. Simply didn't have enough discriminative power\n\n- Sampling data from our sets based on similarity to test set samples\n  -  We tried to find the best training points to use by outputting the XLM embeddings and then doing a similarity search with FAISS. In theory we could pull out the samples closest to the test set and train to accommodate that set. Turns out these embeddings don't really line up properly across languages. Maybe LASER embeddings would've been better here.\n- Training monolingual models and throwing into our ensemble\n  - There are lots of readily available monolingual BERT models and other variants in the huggingface model collection, but it was unclear how exactly we blend these together with the other models because their distributions would not line up across models so turkish predictions might show up in a significantly different place than italian for example and that would screw up the AUC. There might be some way to align the ones we had validation samples for but some of the languages we did not have non-translated validation samples for.\n- Training on the unintended bias data\n  - This yielded a consistently worse model. I believe this to be because the labeling process might have been slightly different between the sets. It would have been nice to be able to utilize this better, but we did not find a way to make this bring us a boost. \n- Training on more than just the toxic label\n  - in past competitions we have had good luck training on the whole array of labels (toxic, severe toxic, threatening, etc.), but in this case with these massive XLM models it did not seem to make a difference because our models already have so much language understanding baked in. We did have some models that trained on multiple outputs but it did not yield big gains like in previous competitions. \n- Generating our own translations with models from FAIR seq.\n  - In the last couple days we tried doing translations into russian and french based on the models that were available. It seemed like our performance was somewhat capped by the quality of translations we used so new translations beyond the google translate ones seemed useful, but it was difficult to measure any gain because we did not have russian or french validation samples and hadn't previously set aside any. \n- Dropping mislabeled samples\n  - One of the things we noticed was even after training on the validation set for finetuning it would still get certain ones wrong with a high degree of confidence. We viewed these points and they looked like mislabels. We thought it might be good to clean those out of the training set as well. We did this by training on the data. Finding samples that had a very high error even after several epochs and then dropping those samples and retraining a model from scratch just on those samples. This led to amazing hyperconvergence of the model, achieving almost exactly the same performance of the previous model in the first epoch, but it did not improve upon previous results because it was essentially just trained to agree with the prior model.",
    "897551": "NLP augmentations (although I suspect we didn't try enough)\n\n- Multi-sample dropout\n\n- Weighted mean pooling of XLM-Roberta layers\n\n- Gradual unfreezing\n\n- Training first on toxic then on unintended bias data\n\n- Keeping validation.csv for validation set",
    "897554": "Similar experiences with all of those. What did you do for validation instead of using the validation file?",
    "897557": "1. NLP augmentations\n2. Ensemble translated test data\n3. Pseudo labeling\n4. KNN of Doc2vec to find the most similar comment of external data and test data",
    "897572": "* Unsupervised Data Augmentation\nThis also looks promising as seen in [the original paper](https://arxiv.org/abs/1904.12848) and [the one on cross-lingual text classification](https://arxiv.org/abs/1909.07009). The idea also makes sense. The loss objective is to minimize the difference between the output of 2 languages. I implement one using `tf.keras.losses.KLDivergence` and using opensubs parallel data. After wasting so many hours trying to make it work, it doesn't bring any significant value to my performance :( \n\nor maybe it's just my bad implementation ¯\\_(ツ)_/¯ I'll check later after the test set is released.",
    "897576": "For me did not work fancy head on top of the transformers, I probably wasted too many TPU hours here.",
    "897586": "Even we tried it. In our case it gave little boost on single models and made results consistent.",
    "897698": "Hey, nice to know someone else tried it too! Perhaps it's just my bad implementation. \n\nI would like to know. How do you combine the UDA loss with BCE loss? Do you just sum it up? Do you also implement the prediction probability thresholding and Training Signal Annealing (TSA)?",
    "897705": "Yes we just sumed it up with a coefficient to KL divergence. No we didn't do tsa or probability thresholding.",
    "898018": "We trust the LB, and in the beginning we only trained for one epoch for fear of overfitting. It worked quite well since we achieved scores on the LB for a single model of around 0.9425.",
    "898023": "Pseudo labeling did work for one of our model. It gave a small boost, but we didn't select it in the end since we found a better solution."
  },
  "source": "meta"
}