{
  "id": 152133,
  "title": "Model doesn't learn when committing",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/152133",
  "author_name": "",
  "post_date": "2020-05-18T14:48:18.685919500Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello everyone, </p>\n\n<p>I have posted several weeks ago a message saying that my model didn't learn when fed with more than 200,000 samples. It turns out the problem is solved.</p>\n\n<p>However, I'm running into another problem now with Tensorflow. I have a XLM-Roberta model training for 8 epochs, the model learns perfectly when I edit it (val_loss decreases, val_auc increases) but when I commit it, the model no longer learns.</p>\n\n<p>Has anyone faced the same issue? Does anyone have a clue of what might change between editing and committing? </p>\n\n<p>For info, I seed everything, so there shouldn't be an issue with that. </p>",
  "messages": [
    {
      "id": "852598",
      "postDate": "05/18/2020 14:48:18",
      "content": "<p>Hello everyone, </p>\n\n<p>I have posted several weeks ago a message saying that my model didn't learn when fed with more than 200,000 samples. It turns out the problem is solved.</p>\n\n<p>However, I'm running into another problem now with Tensorflow. I have a XLM-Roberta model training for 8 epochs, the model learns perfectly when I edit it (val_loss decreases, val_auc increases) but when I commit it, the model no longer learns.</p>\n\n<p>Has anyone faced the same issue? Does anyone have a clue of what might change between editing and committing? </p>\n\n<p>For info, I seed everything, so there shouldn't be an issue with that. </p>",
      "rawMarkdown": "Hello everyone, \n\nI have posted several weeks ago a message saying that my model didn't learn when fed with more than 200,000 samples. It turns out the problem is solved.\n\nHowever, I'm running into another problem now with Tensorflow. I have a XLM-Roberta model training for 8 epochs, the model learns perfectly when I edit it (val_loss decreases, val_auc increases) but when I commit it, the model no longer learns.\n\nHas anyone faced the same issue? Does anyone have a clue of what might change between editing and committing? \n\nFor info, I seed everything, so there shouldn't be an issue with that.",
      "votes": null
    },
    {
      "id": "852888",
      "postDate": "05/18/2020 18:56:18",
      "content": "<p>Hi PAB,</p>\n\n<p>I have been running into this problem too! Here you can see 2 different commits, training XLM-R using TPU. Exactly the same configuration, everything seeded, and you can see the line diff is 0.\nFirst commit:<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1394721%2F010e1e45ccce54c388b45ea78b4e2fa3%2Fbad_run.PNG?generation=1589827948518965&amp;alt=media\" alt=\"\">  </p>\n\n<p>Second commit:<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1394721%2F32243b71dd260f0c51afd3863c9d0862%2Fgood_run.PNG?generation=1589827993829410&amp;alt=media\" alt=\"\"></p>\n\n<p>For some reason, in the second commit, it learns! The loss decreases and F1-score improved. I can't explain what happened, but have you tried running your experiment for the second time? I also noticed that lowering my learning rate reduces the prevalence of this error. All the subsequent commit (top right, C_1.5 until C_3) was a success when I lower the learning rate from 1e-5 to 5e-6.</p>",
      "rawMarkdown": "Hi PAB,\n\nI have been running into this problem too! Here you can see 2 different commits, training XLM-R using TPU. Exactly the same configuration, everything seeded, and you can see the line diff is 0.\nFirst commit:![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1394721%2F010e1e45ccce54c388b45ea78b4e2fa3%2Fbad_run.PNG?generation=1589827948518965&amp;alt=media)  \n\nSecond commit:![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1394721%2F32243b71dd260f0c51afd3863c9d0862%2Fgood_run.PNG?generation=1589827993829410&amp;alt=media)\n\nFor some reason, in the second commit, it learns! The loss decreases and F1-score improved. I can't explain what happened, but have you tried running your experiment for the second time? I also noticed that lowering my learning rate reduces the prevalence of this error. All the subsequent commit (top right, C_1.5 until C_3) was a success when I lower the learning rate from 1e-5 to 5e-6.",
      "votes": null
    },
    {
      "id": "853502",
      "postDate": "05/19/2020 08:31:27",
      "content": "<p>Thanks for your response! I'll try it out and come back to you.</p>",
      "rawMarkdown": "Thanks for your response! I'll try it out and come back to you.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 852888,
      "author_name": "ilhamfp31",
      "author_url": "",
      "post_date": "05/18/2020 18:56:18",
      "content": "<p>Hi PAB,</p>\n\n<p>I have been running into this problem too! Here you can see 2 different commits, training XLM-R using TPU. Exactly the same configuration, everything seeded, and you can see the line diff is 0.\nFirst commit:<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1394721%2F010e1e45ccce54c388b45ea78b4e2fa3%2Fbad_run.PNG?generation=1589827948518965&amp;alt=media\" alt=\"\">  </p>\n\n<p>Second commit:<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1394721%2F32243b71dd260f0c51afd3863c9d0862%2Fgood_run.PNG?generation=1589827993829410&amp;alt=media\" alt=\"\"></p>\n\n<p>For some reason, in the second commit, it learns! The loss decreases and F1-score improved. I can't explain what happened, but have you tried running your experiment for the second time? I also noticed that lowering my learning rate reduces the prevalence of this error. All the subsequent commit (top right, C_1.5 until C_3) was a success when I lower the learning rate from 1e-5 to 5e-6.</p>",
      "votes": null,
      "replies": [
        {
          "id": 853502,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "05/19/2020 08:31:27",
          "content": "<p>Thanks for your response! I'll try it out and come back to you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "852598": "Hello everyone, \n\nI have posted several weeks ago a message saying that my model didn't learn when fed with more than 200,000 samples. It turns out the problem is solved.\n\nHowever, I'm running into another problem now with Tensorflow. I have a XLM-Roberta model training for 8 epochs, the model learns perfectly when I edit it (val_loss decreases, val_auc increases) but when I commit it, the model no longer learns.\n\nHas anyone faced the same issue? Does anyone have a clue of what might change between editing and committing? \n\nFor info, I seed everything, so there shouldn't be an issue with that.",
    "852888": "Hi PAB,\n\nI have been running into this problem too! Here you can see 2 different commits, training XLM-R using TPU. Exactly the same configuration, everything seeded, and you can see the line diff is 0.\nFirst commit:![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1394721%2F010e1e45ccce54c388b45ea78b4e2fa3%2Fbad_run.PNG?generation=1589827948518965&amp;alt=media)  \n\nSecond commit:![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1394721%2F32243b71dd260f0c51afd3863c9d0862%2Fgood_run.PNG?generation=1589827993829410&amp;alt=media)\n\nFor some reason, in the second commit, it learns! The loss decreases and F1-score improved. I can't explain what happened, but have you tried running your experiment for the second time? I also noticed that lowering my learning rate reduces the prevalence of this error. All the subsequent commit (top right, C_1.5 until C_3) was a success when I lower the learning rate from 1e-5 to 5e-6.",
    "853502": "Thanks for your response! I'll try it out and come back to you."
  },
  "source": "meta"
}