{
  "id": 80790,
  "title": "Transfer Learning BERT, ULMFIT, etc.",
  "url": "/competitions/quora-insincere-questions-classification/discussion/80790",
  "author_name": "",
  "post_date": "2019-02-16T16:30:36.099364200Z",
  "votes": 16,
  "comment_count": 8,
  "views": 0,
  "content": "<p>After competition finish, I have played a bit around with trying to fine-tune some of the hip NLP transfer models like BERT and ULMFIT to our data at hand. Kernels like <a href=\"https://www.kaggle.com/ishitori/what-if-we-could-finetune-bert-from-apache-mxnet\">https://www.kaggle.com/ishitori/what-if-we-could-finetune-bert-from-apache-mxnet</a> suggest huge improvement, but they only use 2% test data, so can't really trust it. </p>\n\n<p>Until now, I have tried out a few libraries/implementations mostly with 10% test split and even though results look reasonable they are nowhere close to really outperforming the final solutions posted here. So I am curious if someone experimented as well with this?</p>",
  "messages": [
    {
      "id": "472766",
      "postDate": "02/16/2019 16:30:36",
      "content": "<p>After competition finish, I have played a bit around with trying to fine-tune some of the hip NLP transfer models like BERT and ULMFIT to our data at hand. Kernels like <a href=\"https://www.kaggle.com/ishitori/what-if-we-could-finetune-bert-from-apache-mxnet\">https://www.kaggle.com/ishitori/what-if-we-could-finetune-bert-from-apache-mxnet</a> suggest huge improvement, but they only use 2% test data, so can't really trust it. </p>\n\n<p>Until now, I have tried out a few libraries/implementations mostly with 10% test split and even though results look reasonable they are nowhere close to really outperforming the final solutions posted here. So I am curious if someone experimented as well with this?</p>",
      "rawMarkdown": "After competition finish, I have played a bit around with trying to fine-tune some of the hip NLP transfer models like BERT and ULMFIT to our data at hand. Kernels like https://www.kaggle.com/ishitori/what-if-we-could-finetune-bert-from-apache-mxnet suggest huge improvement, but they only use 2% test data, so can't really trust it. \n\nUntil now, I have tried out a few libraries/implementations mostly with 10% test split and even though results look reasonable they are nowhere close to really outperforming the final solutions posted here. So I am curious if someone experimented as well with this?",
      "votes": null
    },
    {
      "id": "472782",
      "postDate": "02/16/2019 16:50:07",
      "content": "<p>Maybe you found the reason for hoding this competition。。。</p>",
      "rawMarkdown": "Maybe you found the reason for hoding this competition。。。",
      "votes": null
    },
    {
      "id": "472786",
      "postDate": "02/16/2019 16:57:24",
      "content": "<p>Maybe :)</p>",
      "rawMarkdown": "Maybe :)",
      "votes": null
    },
    {
      "id": "472985",
      "postDate": "02/17/2019 03:08:00",
      "content": "<p>:)</p>",
      "rawMarkdown": ":)",
      "votes": null
    },
    {
      "id": "473380",
      "postDate": "02/17/2019 22:07:19",
      "content": "<p>I tried <a href=\"https://github.com/huggingface/pytorch-pretrained-BERT\">pytorch-pretrained-BERT</a>, and got good results.</p>\n\n<p>20% for test, 100,000 for determining threshold and the rest for training. \nMost hyperparameter values are the default in <a href=\"https://github.com/huggingface/pytorch-pretrained-BERT/blob/master/examples/run_classifier.py\">run_classifier.py</a>.</p>\n\n<p>Here are the changed hyperparameter values and scores.</p>\n\n<p>'</p>\n\n<pre><code>args.bert_model = 'bert-base-uncased'\nargs.max_seq_length = 50\nargs.num_train_epochs = 2\nargs.train_batch_size = 64\nargs.gradient_accumulation_steps = 8\nargs.eval_batch_size = 160\n\nSingle\nepoch       F1      AUC\n1      0.71944  0.97505\n2      0.72185  0.97531\n\nAvg ensemble of 3 models\nepoch       F1      AUC\n1      0.72365  0.97634\n2      0.72747  0.97664\n</code></pre>",
      "rawMarkdown": "I tried [pytorch-pretrained-BERT](https://github.com/huggingface/pytorch-pretrained-BERT), and got good results.\n\n20% for test, 100,000 for determining threshold and the rest for training. \nMost hyperparameter values are the default in [run_classifier.py](https://github.com/huggingface/pytorch-pretrained-BERT/blob/master/examples/run_classifier.py).\n\nHere are the changed hyperparameter values and scores.\n\n\n'\n\n    args.bert_model = 'bert-base-uncased'\n    args.max_seq_length = 50\n    args.num_train_epochs = 2\n    args.train_batch_size = 64\n    args.gradient_accumulation_steps = 8\n    args.eval_batch_size = 160\n\n    Single\n    epoch       F1      AUC\n    1      0.71944  0.97505\n    2      0.72185  0.97531\n\n    Avg ensemble of 3 models\n    epoch       F1      AUC\n    1      0.72365  0.97634\n    2      0.72747  0.97664",
      "votes": null
    },
    {
      "id": "473922",
      "postDate": "02/18/2019 17:33:08",
      "content": "<p>Interesting, thanks for sharing, this looks much better. I also tried the same package and got worse, I might have to retry it. I have to say though that I ran it with mixed precision training, so this might be one reason. I also ran it with the multilingual-cased model. My test split was 10% with seed 42. Do you mind sharing the script you are running?</p>",
      "rawMarkdown": "Interesting, thanks for sharing, this looks much better. I also tried the same package and got worse, I might have to retry it. I have to say though that I ran it with mixed precision training, so this might be one reason. I also ran it with the multilingual-cased model. My test split was 10% with seed 42. Do you mind sharing the script you are running?",
      "votes": null
    },
    {
      "id": "475251",
      "postDate": "02/20/2019 14:01:54",
      "content": "<p>Ok. <a href=\"https://github.com/tks0123456789/kaggle-Quora\">Here</a> is the script. It doesn't contains  fp16 related code.</p>",
      "rawMarkdown": "Ok. [Here](https://github.com/tks0123456789/kaggle-Quora) is the script. It doesn't contains  fp16 related code.",
      "votes": null
    },
    {
      "id": "475487",
      "postDate": "02/20/2019 19:52:00",
      "content": "<p>Thanks a lot. Results look quite worse compared to those you had before if they are correct in the log file, right? Talking about single model, not ensemble.</p>",
      "rawMarkdown": "Thanks a lot. Results look quite worse compared to those you had before if they are correct in the log file, right? Talking about single model, not ensemble.",
      "votes": null
    },
    {
      "id": "475783",
      "postDate": "02/21/2019 07:39:36",
      "content": "<p>I got the score above over a month ago and the current code is somewhat different. The new feature in the code is varying the number of hidden layers(--n_bertlayers option). This time I started experiment with smaller model, then increase #layers. That is why you saw worse scores. I've just finished a full model experiment(exp019) and the score is similar to the above.   Please see updated logs.</p>",
      "rawMarkdown": "I got the score above over a month ago and the current code is somewhat different. The new feature in the code is varying the number of hidden layers(--n_bertlayers option). This time I started experiment with smaller model, then increase #layers. That is why you saw worse scores. I've just finished a full model experiment(exp019) and the score is similar to the above.   Please see updated logs.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 472782,
      "author_name": "noxuslol",
      "author_url": "",
      "post_date": "02/16/2019 16:50:07",
      "content": "<p>Maybe you found the reason for hoding this competition。。。</p>",
      "votes": null,
      "replies": [
        {
          "id": 472786,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/16/2019 16:57:24",
          "content": "<p>Maybe :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 472985,
      "author_name": "dhanushk2105",
      "author_url": "",
      "post_date": "02/17/2019 03:08:00",
      "content": "<p>:)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 473380,
      "author_name": "tks0123456789",
      "author_url": "",
      "post_date": "02/17/2019 22:07:19",
      "content": "<p>I tried <a href=\"https://github.com/huggingface/pytorch-pretrained-BERT\">pytorch-pretrained-BERT</a>, and got good results.</p>\n\n<p>20% for test, 100,000 for determining threshold and the rest for training. \nMost hyperparameter values are the default in <a href=\"https://github.com/huggingface/pytorch-pretrained-BERT/blob/master/examples/run_classifier.py\">run_classifier.py</a>.</p>\n\n<p>Here are the changed hyperparameter values and scores.</p>\n\n<p>'</p>\n\n<pre><code>args.bert_model = 'bert-base-uncased'\nargs.max_seq_length = 50\nargs.num_train_epochs = 2\nargs.train_batch_size = 64\nargs.gradient_accumulation_steps = 8\nargs.eval_batch_size = 160\n\nSingle\nepoch       F1      AUC\n1      0.71944  0.97505\n2      0.72185  0.97531\n\nAvg ensemble of 3 models\nepoch       F1      AUC\n1      0.72365  0.97634\n2      0.72747  0.97664\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 473922,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/18/2019 17:33:08",
          "content": "<p>Interesting, thanks for sharing, this looks much better. I also tried the same package and got worse, I might have to retry it. I have to say though that I ran it with mixed precision training, so this might be one reason. I also ran it with the multilingual-cased model. My test split was 10% with seed 42. Do you mind sharing the script you are running?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 475251,
          "author_name": "tks0123456789",
          "author_url": "",
          "post_date": "02/20/2019 14:01:54",
          "content": "<p>Ok. <a href=\"https://github.com/tks0123456789/kaggle-Quora\">Here</a> is the script. It doesn't contains  fp16 related code.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 475487,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/20/2019 19:52:00",
          "content": "<p>Thanks a lot. Results look quite worse compared to those you had before if they are correct in the log file, right? Talking about single model, not ensemble.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 475783,
          "author_name": "tks0123456789",
          "author_url": "",
          "post_date": "02/21/2019 07:39:36",
          "content": "<p>I got the score above over a month ago and the current code is somewhat different. The new feature in the code is varying the number of hidden layers(--n_bertlayers option). This time I started experiment with smaller model, then increase #layers. That is why you saw worse scores. I've just finished a full model experiment(exp019) and the score is similar to the above.   Please see updated logs.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "472766": "After competition finish, I have played a bit around with trying to fine-tune some of the hip NLP transfer models like BERT and ULMFIT to our data at hand. Kernels like https://www.kaggle.com/ishitori/what-if-we-could-finetune-bert-from-apache-mxnet suggest huge improvement, but they only use 2% test data, so can't really trust it. \n\nUntil now, I have tried out a few libraries/implementations mostly with 10% test split and even though results look reasonable they are nowhere close to really outperforming the final solutions posted here. So I am curious if someone experimented as well with this?",
    "472782": "Maybe you found the reason for hoding this competition。。。",
    "472786": "Maybe :)",
    "472985": ":)",
    "473380": "I tried [pytorch-pretrained-BERT](https://github.com/huggingface/pytorch-pretrained-BERT), and got good results.\n\n20% for test, 100,000 for determining threshold and the rest for training. \nMost hyperparameter values are the default in [run_classifier.py](https://github.com/huggingface/pytorch-pretrained-BERT/blob/master/examples/run_classifier.py).\n\nHere are the changed hyperparameter values and scores.\n\n\n'\n\n    args.bert_model = 'bert-base-uncased'\n    args.max_seq_length = 50\n    args.num_train_epochs = 2\n    args.train_batch_size = 64\n    args.gradient_accumulation_steps = 8\n    args.eval_batch_size = 160\n\n    Single\n    epoch       F1      AUC\n    1      0.71944  0.97505\n    2      0.72185  0.97531\n\n    Avg ensemble of 3 models\n    epoch       F1      AUC\n    1      0.72365  0.97634\n    2      0.72747  0.97664",
    "473922": "Interesting, thanks for sharing, this looks much better. I also tried the same package and got worse, I might have to retry it. I have to say though that I ran it with mixed precision training, so this might be one reason. I also ran it with the multilingual-cased model. My test split was 10% with seed 42. Do you mind sharing the script you are running?",
    "475251": "Ok. [Here](https://github.com/tks0123456789/kaggle-Quora) is the script. It doesn't contains  fp16 related code.",
    "475487": "Thanks a lot. Results look quite worse compared to those you had before if they are correct in the log file, right? Talking about single model, not ensemble.",
    "475783": "I got the score above over a month ago and the current code is somewhat different. The new feature in the code is varying the number of hidden layers(--n_bertlayers option). This time I started experiment with smaller model, then increase #layers. That is why you saw worse scores. I've just finished a full model experiment(exp019) and the score is similar to the above.   Please see updated logs."
  },
  "source": "meta"
}