{
  "id": 146337,
  "title": "Issues when training with more than 200,000 samples",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/146337",
  "author_name": "",
  "post_date": "2020-04-26T19:47:27.831267Z",
  "votes": 11,
  "comment_count": 34,
  "views": 0,
  "content": "<p>Hello everyone, </p>\n\n<p>I'm using Tensorflow and have been training my models with 200,000 samples so far. I wanted to try to train on more samples (e.g. 400,000 samples). I don't run into memory issues, but oddly enough my model doesn't learn...</p>\n\n<p>The val_auc is stuck either at 0 or 0.5 while the loss is exploding. </p>\n\n<p>Has anyone faced the same issue and come up with a solution? </p>\n\n<p>Any help is appreciated. Many thanks in advance ;)</p>",
  "messages": [
    {
      "id": "822279",
      "postDate": "04/26/2020 19:47:27",
      "content": "<p>Hello everyone, </p>\n\n<p>I'm using Tensorflow and have been training my models with 200,000 samples so far. I wanted to try to train on more samples (e.g. 400,000 samples). I don't run into memory issues, but oddly enough my model doesn't learn...</p>\n\n<p>The val_auc is stuck either at 0 or 0.5 while the loss is exploding. </p>\n\n<p>Has anyone faced the same issue and come up with a solution? </p>\n\n<p>Any help is appreciated. Many thanks in advance ;)</p>",
      "rawMarkdown": "Hello everyone, \n\nI'm using Tensorflow and have been training my models with 200,000 samples so far. I wanted to try to train on more samples (e.g. 400,000 samples). I don't run into memory issues, but oddly enough my model doesn't learn...\n\nThe val_auc is stuck either at 0 or 0.5 while the loss is exploding. \n\nHas anyone faced the same issue and come up with a solution? \n\nAny help is appreciated. Many thanks in advance ;)",
      "votes": null
    },
    {
      "id": "822483",
      "postDate": "04/27/2020 00:18:02",
      "content": "<p><a href=\"/rftexas\">@rftexas</a> I faced the same issue, and for recent runs, it is going OOM even for small batch sizes. Which model are your using? </p>",
      "rawMarkdown": "rftexas I faced the same issue, and for recent runs, it is going OOM even for small batch sizes. Which model are your using?",
      "votes": null
    },
    {
      "id": "822484",
      "postDate": "04/27/2020 00:19:07",
      "content": "<p>if you have changed batch size, you should to change learning rate proportionally. </p>\n\n<p><code>\nbs=128, lr=0.4*1e-5  ---&gt; bs=64 lr=0.2*1e-5\n</code></p>\n\n<p>Maybe It helps for you. High value for learning rate can crush pretrained weights.</p>",
      "rawMarkdown": "if you have changed batch size, you should to change learning rate proportionally. \n\n```\nbs=128, lr=0.4*1e-5  ---&gt; bs=64 lr=0.2*1e-5\n```\n\nMaybe It helps for you. High value for learning rate can crush pretrained weights.",
      "votes": null
    },
    {
      "id": "822510",
      "postDate": "04/27/2020 00:54:31",
      "content": "<p>XLM-Roberta</p>",
      "rawMarkdown": "XLM-Roberta",
      "votes": null
    },
    {
      "id": "822511",
      "postDate": "04/27/2020 00:55:14",
      "content": "<p>I haven't... But I'll try ! Thanks!</p>",
      "rawMarkdown": "I haven't... But I'll try ! Thanks!",
      "votes": null
    },
    {
      "id": "822528",
      "postDate": "04/27/2020 01:12:56",
      "content": "<p>I have the same problem too . But mine happens when I use pytorch from Abhishek when training on both training set combined. </p>",
      "rawMarkdown": "I have the same problem too . But mine happens when I use pytorch from Abhishek when training on both training set combined.",
      "votes": null
    },
    {
      "id": "822768",
      "postDate": "04/27/2020 05:50:08",
      "content": "<p>Are you sampling the dataset? 200k each class? <a href=\"/rftexas\">@rftexas</a> </p>",
      "rawMarkdown": "Are you sampling the dataset? 200k each class? @rftexas",
      "votes": null
    },
    {
      "id": "822975",
      "postDate": "04/27/2020 09:44:05",
      "content": "<p>I'm training on tensorflow and pytorch with 1.5 million of records, no issue.\n0.5 means random guess : either you model is overfitting to your train data, either, there's an issue in the validation loop. Certainly not the case if you are using model.fit().</p>\n\n<p>By the way, if you val loss is 0, be happy : predict the total opposite :)</p>",
      "rawMarkdown": "I'm training on tensorflow and pytorch with 1.5 million of records, no issue.\n0.5 means random guess : either you model is overfitting to your train data, either, there's an issue in the validation loop. Certainly not the case if you are using model.fit().\n\nBy the way, if you val loss is 0, be happy : predict the total opposite :)",
      "votes": null
    },
    {
      "id": "823004",
      "postDate": "04/27/2020 10:18:16",
      "content": "<p>Thanks for your insight!</p>",
      "rawMarkdown": "Thanks for your insight!",
      "votes": null
    },
    {
      "id": "823005",
      "postDate": "04/27/2020 10:18:22",
      "content": "<p>Yes</p>",
      "rawMarkdown": "Yes",
      "votes": null
    },
    {
      "id": "823076",
      "postDate": "04/27/2020 11:35:11",
      "content": "<p><a href=\"/rftexas\">@rftexas</a> The same thing happens to me when I try to balance the data with 200k each.Any idea what is happening?</p>",
      "rawMarkdown": "rftexas The same thing happens to me when I try to balance the data with 200k each.Any idea what is happening?",
      "votes": null
    },
    {
      "id": "823110",
      "postDate": "04/27/2020 12:14:53",
      "content": "<p>Something similar happened to me once when I forgot to shuffle train dataset  - maybe because whole batches consisted of only negative or only positive samples.\nMy impression is that (pls correct if wrong) tensorflow dataset.shuffle(2048) takes the next 2048 samples and shuffles only those rows, not the whole dataset. If all those samples are from the same class then it might cause problem</p>",
      "rawMarkdown": "Something similar happened to me once when I forgot to shuffle train dataset  - maybe because whole batches consisted of only negative or only positive samples.\nMy impression is that (pls correct if wrong) tensorflow dataset.shuffle(2048) takes the next 2048 samples and shuffles only those rows, not the whole dataset. If all those samples are from the same class then it might cause problem",
      "votes": null
    },
    {
      "id": "823222",
      "postDate": "04/27/2020 13:50:55",
      "content": "<p><a href=\"/shonenkov\">@shonenkov</a> Does the number of tpu devices matter too? I'm using pytorch if that helps...\nI am using 0.4*1e-5*8 because I am using the 8-core TPU for training</p>",
      "rawMarkdown": "shonenkov Does the number of tpu devices matter too? I'm using pytorch if that helps...\nI am using 0.4*1e-5*8 because I am using the 8-core TPU for training",
      "votes": null
    },
    {
      "id": "823409",
      "postDate": "04/27/2020 16:03:22",
      "content": "<p><a href=\"/amoghjrules\">@amoghjrules</a>  yes, if you use 8-core tpu you shoud multiply by 8 times</p>",
      "rawMarkdown": "amoghjrules  yes, if you use 8-core tpu you shoud multiply by 8 times",
      "votes": null
    },
    {
      "id": "823412",
      "postDate": "04/27/2020 16:06:01",
      "content": "<p><a href=\"https://arxiv.org/pdf/1706.02677.pdf\">https://arxiv.org/pdf/1706.02677.pdf</a></p>",
      "rawMarkdown": "https://arxiv.org/pdf/1706.02677.pdf",
      "votes": null
    },
    {
      "id": "823560",
      "postDate": "04/27/2020 18:17:42",
      "content": "<p><a href=\"/isakev\">@isakev</a> wow! I think you're right. I solved my issue by shuffling the data before feeding it to the generator.</p>\n\n<p><code>from sklearn.utils import shuffle</code></p>\n\n<p><code>train = shuffle(train).reset_index(drop=True)</code>\n<a href=\"/rftexas\">@rftexas</a> can you try and tell if this works ? Solved in my case.</p>",
      "rawMarkdown": "isakev wow! I think you're right. I solved my issue by shuffling the data before feeding it to the generator.\n\n`from sklearn.utils import shuffle`\n\n`train = shuffle(train).reset_index(drop=True)`\n@rftexas can you try and tell if this works ? Solved in my case.",
      "votes": null
    },
    {
      "id": "823669",
      "postDate": "04/27/2020 19:49:19",
      "content": "<p>You may need to implement your own Data generator to shuffle the entire data at the end of each epoch, as the default TF datagenerator will shuffle only the next batch . </p>",
      "rawMarkdown": "You may need to implement your own Data generator to shuffle the entire data at the end of each epoch, as the default TF datagenerator will shuffle only the next batch .",
      "votes": null
    },
    {
      "id": "824196",
      "postDate": "04/28/2020 08:08:29",
      "content": "<p>Thanks for the tips! I will!</p>",
      "rawMarkdown": "Thanks for the tips! I will!",
      "votes": null
    },
    {
      "id": "824200",
      "postDate": "04/28/2020 08:11:01",
      "content": "<p>I really don't know why. I will try <a href=\"/serigne\">@serigne</a> approach, and tell you!</p>",
      "rawMarkdown": "I really don't know why. I will try @serigne approach, and tell you!",
      "votes": null
    },
    {
      "id": "824202",
      "postDate": "04/28/2020 08:13:24",
      "content": "<p>I'll try it today! I'll tell you!</p>",
      "rawMarkdown": "I'll try it today! I'll tell you!",
      "votes": null
    },
    {
      "id": "824310",
      "postDate": "04/28/2020 09:44:56",
      "content": "<p>True, 2048 is the buffer size, meaning that the next 2048 samples are shuffled; not the entire dataset. I will work on that and tell you.</p>",
      "rawMarkdown": "True, 2048 is the buffer size, meaning that the next 2048 samples are shuffled; not the entire dataset. I will work on that and tell you.",
      "votes": null
    },
    {
      "id": "824742",
      "postDate": "04/28/2020 15:12:21",
      "content": "<p>hey encountering the same issue have you <a href=\"/rftexas\">@rftexas</a>  solved it .  i am trying stratifiedkfold and it has shuffle argument with random_state but still facing the same issue</p>",
      "rawMarkdown": "hey encountering the same issue have you @rftexas  solved it .  i am trying stratifiedkfold and it has shuffle argument with random_state but still facing the same issue",
      "votes": null
    },
    {
      "id": "824890",
      "postDate": "04/28/2020 16:43:14",
      "content": "<p>Are you shuffling the entire dataset ? Or just the validation? <a href=\"/pranshu29\">@pranshu29</a> </p>",
      "rawMarkdown": "Are you shuffling the entire dataset ? Or just the validation? @pranshu29",
      "votes": null
    },
    {
      "id": "824907",
      "postDate": "04/28/2020 16:57:03",
      "content": "<p>Well i am using random state 42 and shuffle true in stratified k fold and validating the model on full validation data with no shuffle</p>",
      "rawMarkdown": "Well i am using random state 42 and shuffle true in stratified k fold and validating the model on full validation data with no shuffle",
      "votes": null
    },
    {
      "id": "825644",
      "postDate": "04/29/2020 06:32:35",
      "content": "<p>How you guys sampling the dataset for 200k each class. Drop some idea, I didn't understand!   </p>",
      "rawMarkdown": "How you guys sampling the dataset for 200k each class. Drop some idea, I didn't understand!",
      "votes": null
    },
    {
      "id": "827549",
      "postDate": "04/30/2020 11:28:01",
      "content": "<p>I was facing the same issue a few days back (PyTorch). I thought maybe I had worked out something wrong with the implementation. Although my training loss did not explode, it stayed more or less constant and the val_auc was stuck at 0.5 as you've mentioned.</p>\n\n<p>Well, thanks for posting this :)\nI'll try out the suggested approaches and let you know!</p>",
      "rawMarkdown": "I was facing the same issue a few days back (PyTorch). I thought maybe I had worked out something wrong with the implementation. Although my training loss did not explode, it stayed more or less constant and the val_auc was stuck at 0.5 as you've mentioned.\n\nWell, thanks for posting this :)\nI'll try out the suggested approaches and let you know!",
      "votes": null
    },
    {
      "id": "827583",
      "postDate": "04/30/2020 11:52:56",
      "content": "<p>Sorry for the late answer....</p>\n\n<p>I am using the Tensorflow Data API. I tried to increase the buffer in the shuffle method, but nothing happens... </p>\n\n<p>I think given what <a href=\"/serigne\">@serigne</a> has suggested, I must write my own code generator. I'll work on that. </p>",
      "rawMarkdown": "Sorry for the late answer....\n\nI am using the Tensorflow Data API. I tried to increase the buffer in the shuffle method, but nothing happens... \n\nI think given what @serigne has suggested, I must write my own code generator. I'll work on that.",
      "votes": null
    },
    {
      "id": "828754",
      "postDate": "05/01/2020 09:08:15",
      "content": "<p><a href=\"/shahules\">@shahules</a>  I think what is happening is that sampling the data will create dataset where first 200k will be 1 and next 200k will be 0 so , NN will see first see 200k 0s and then 200k 1s and not able to learn, and  dataset.shuffle(2048) is not working because acc to tf documentation buffer_size must be greater then length of whole dataset.\n <a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset#shuffle\">https://www.tensorflow.org/api_docs/python/tf/data/Dataset#shuffle</a></p>\n\n<p>Although this is what I think.\nplease let me know If I misunderstood the question.</p>",
      "rawMarkdown": "shahules  I think what is happening is that sampling the data will create dataset where first 200k will be 1 and next 200k will be 0 so , NN will see first see 200k 0s and then 200k 1s and not able to learn, and  dataset.shuffle(2048) is not working because acc to tf documentation buffer_size must be greater then length of whole dataset.\n https://www.tensorflow.org/api_docs/python/tf/data/Dataset#shuffle\n\nAlthough this is what I think.\nplease let me know If I misunderstood the question.",
      "votes": null
    },
    {
      "id": "831622",
      "postDate": "05/03/2020 13:46:22",
      "content": "<p>For those who might be wondering, I solved my issues by using a custom training loop function. </p>\n\n<p>Here is a great starting point:\n- <a href=\"https://colab.research.google.com/drive/169PfzM0kvtA5UP4k6Sl1yCG9tsE2MLia\">https://colab.research.google.com/drive/169PfzM0kvtA5UP4k6Sl1yCG9tsE2MLia</a></p>",
      "rawMarkdown": "For those who might be wondering, I solved my issues by using a custom training loop function. \n\nHere is a great starting point:\n- https://colab.research.google.com/drive/169PfzM0kvtA5UP4k6Sl1yCG9tsE2MLia",
      "votes": null
    },
    {
      "id": "832396",
      "postDate": "05/04/2020 05:14:02",
      "content": "<p>Thanks for sharing the resources with us. \nI'm trying to access it and got this error \n<code>There was an error loading this notebook. Ensure that the file is accessible and try again.\nInvalid Credentials</code></p>\n\n<p>Just wondering  because the colab notebook is not publicly accessible ?</p>",
      "rawMarkdown": "Thanks for sharing the resources with us. \nI'm trying to access it and got this error \n`There was an error loading this notebook. Ensure that the file is accessible and try again.\nInvalid Credentials`\n\nJust wondering  because the colab notebook is not publicly accessible ?",
      "votes": null
    },
    {
      "id": "832866",
      "postDate": "05/04/2020 13:25:33",
      "content": "<p>That's weird, I can access it... Are you logged in ?\nIt is publicly shared by François Chollet, Keras founder.</p>",
      "rawMarkdown": "That's weird, I can access it... Are you logged in ?\nIt is publicly shared by François Chollet, Keras founder.",
      "votes": null
    },
    {
      "id": "833082",
      "postDate": "05/04/2020 15:33:04",
      "content": "<p>I can access it now. My internet is weird sometimes :) </p>",
      "rawMarkdown": "I can access it now. My internet is weird sometimes :)",
      "votes": null
    },
    {
      "id": "836847",
      "postDate": "05/07/2020 09:48:08",
      "content": "<p>Ok great! </p>",
      "rawMarkdown": "Ok great!",
      "votes": null
    },
    {
      "id": "843003",
      "postDate": "05/11/2020 18:44:11",
      "content": "<p>Hello <a href=\"/rftexas\">@rftexas</a>, for an unknown reason, i'm facing now the issue 🤔 : training is fine and suddently, the auc drop to 0.5.\nI've check that my labels are correctly balanced with toxic and none toxic example.</p>\n\n<p>I'm already using tf custom loop. Do you have your own data generator ? built on top of tf.data.dataset ?\nThx</p>",
      "rawMarkdown": "Hello @rftexas, for an unknown reason, i'm facing now the issue 🤔 : training is fine and suddently, the auc drop to 0.5.\nI've check that my labels are correctly balanced with toxic and none toxic example.\n\nI'm already using tf custom loop. Do you have your own data generator ? built on top of tf.data.dataset ?\nThx",
      "votes": null
    },
    {
      "id": "843753",
      "postDate": "05/12/2020 08:36:38",
      "content": "<p>Hey! I was using my own data generator and used fit_generator. I don't know why sometimes it stops working. Now I'm using PyTorch because I have much more control over the pipeline.</p>",
      "rawMarkdown": "Hey! I was using my own data generator and used fit_generator. I don't know why sometimes it stops working. Now I'm using PyTorch because I have much more control over the pipeline.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 822483,
      "author_name": "rrqqmm",
      "author_url": "",
      "post_date": "04/27/2020 00:18:02",
      "content": "<p><a href=\"/rftexas\">@rftexas</a> I faced the same issue, and for recent runs, it is going OOM even for small batch sizes. Which model are your using? </p>",
      "votes": null,
      "replies": [
        {
          "id": 822510,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "04/27/2020 00:54:31",
          "content": "<p>XLM-Roberta</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 822484,
      "author_name": "shonenkov",
      "author_url": "",
      "post_date": "04/27/2020 00:19:07",
      "content": "<p>if you have changed batch size, you should to change learning rate proportionally. </p>\n\n<p><code>\nbs=128, lr=0.4*1e-5  ---&gt; bs=64 lr=0.2*1e-5\n</code></p>\n\n<p>Maybe It helps for you. High value for learning rate can crush pretrained weights.</p>",
      "votes": null,
      "replies": [
        {
          "id": 822511,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "04/27/2020 00:55:14",
          "content": "<p>I haven't... But I'll try ! Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 823222,
          "author_name": "amoghjrules",
          "author_url": "",
          "post_date": "04/27/2020 13:50:55",
          "content": "<p><a href=\"/shonenkov\">@shonenkov</a> Does the number of tpu devices matter too? I'm using pytorch if that helps...\nI am using 0.4*1e-5*8 because I am using the 8-core TPU for training</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 823409,
          "author_name": "shonenkov",
          "author_url": "",
          "post_date": "04/27/2020 16:03:22",
          "content": "<p><a href=\"/amoghjrules\">@amoghjrules</a>  yes, if you use 8-core tpu you shoud multiply by 8 times</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 823412,
          "author_name": "shonenkov",
          "author_url": "",
          "post_date": "04/27/2020 16:06:01",
          "content": "<p><a href=\"https://arxiv.org/pdf/1706.02677.pdf\">https://arxiv.org/pdf/1706.02677.pdf</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 822528,
      "author_name": "dttung2905",
      "author_url": "",
      "post_date": "04/27/2020 01:12:56",
      "content": "<p>I have the same problem too . But mine happens when I use pytorch from Abhishek when training on both training set combined. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 822768,
      "author_name": "shahules",
      "author_url": "",
      "post_date": "04/27/2020 05:50:08",
      "content": "<p>Are you sampling the dataset? 200k each class? <a href=\"/rftexas\">@rftexas</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 823005,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "04/27/2020 10:18:22",
          "content": "<p>Yes</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 823076,
          "author_name": "shahules",
          "author_url": "",
          "post_date": "04/27/2020 11:35:11",
          "content": "<p><a href=\"/rftexas\">@rftexas</a> The same thing happens to me when I try to balance the data with 200k each.Any idea what is happening?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 824200,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "04/28/2020 08:11:01",
          "content": "<p>I really don't know why. I will try <a href=\"/serigne\">@serigne</a> approach, and tell you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 825644,
          "author_name": "vikassingh1996",
          "author_url": "",
          "post_date": "04/29/2020 06:32:35",
          "content": "<p>How you guys sampling the dataset for 200k each class. Drop some idea, I didn't understand!   </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 828754,
          "author_name": "maunish",
          "author_url": "",
          "post_date": "05/01/2020 09:08:15",
          "content": "<p><a href=\"/shahules\">@shahules</a>  I think what is happening is that sampling the data will create dataset where first 200k will be 1 and next 200k will be 0 so , NN will see first see 200k 0s and then 200k 1s and not able to learn, and  dataset.shuffle(2048) is not working because acc to tf documentation buffer_size must be greater then length of whole dataset.\n <a href=\"https://www.tensorflow.org/api_docs/python/tf/data/Dataset#shuffle\">https://www.tensorflow.org/api_docs/python/tf/data/Dataset#shuffle</a></p>\n\n<p>Although this is what I think.\nplease let me know If I misunderstood the question.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 822975,
      "author_name": "seb6084",
      "author_url": "",
      "post_date": "04/27/2020 09:44:05",
      "content": "<p>I'm training on tensorflow and pytorch with 1.5 million of records, no issue.\n0.5 means random guess : either you model is overfitting to your train data, either, there's an issue in the validation loop. Certainly not the case if you are using model.fit().</p>\n\n<p>By the way, if you val loss is 0, be happy : predict the total opposite :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 823004,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "04/27/2020 10:18:16",
          "content": "<p>Thanks for your insight!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 823110,
      "author_name": "isakev",
      "author_url": "",
      "post_date": "04/27/2020 12:14:53",
      "content": "<p>Something similar happened to me once when I forgot to shuffle train dataset  - maybe because whole batches consisted of only negative or only positive samples.\nMy impression is that (pls correct if wrong) tensorflow dataset.shuffle(2048) takes the next 2048 samples and shuffles only those rows, not the whole dataset. If all those samples are from the same class then it might cause problem</p>",
      "votes": null,
      "replies": [
        {
          "id": 824310,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "04/28/2020 09:44:56",
          "content": "<p>True, 2048 is the buffer size, meaning that the next 2048 samples are shuffled; not the entire dataset. I will work on that and tell you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 823560,
      "author_name": "shahules",
      "author_url": "",
      "post_date": "04/27/2020 18:17:42",
      "content": "<p><a href=\"/isakev\">@isakev</a> wow! I think you're right. I solved my issue by shuffling the data before feeding it to the generator.</p>\n\n<p><code>from sklearn.utils import shuffle</code></p>\n\n<p><code>train = shuffle(train).reset_index(drop=True)</code>\n<a href=\"/rftexas\">@rftexas</a> can you try and tell if this works ? Solved in my case.</p>",
      "votes": null,
      "replies": [
        {
          "id": 824202,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "04/28/2020 08:13:24",
          "content": "<p>I'll try it today! I'll tell you!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 824742,
          "author_name": "pranshu29",
          "author_url": "",
          "post_date": "04/28/2020 15:12:21",
          "content": "<p>hey encountering the same issue have you <a href=\"/rftexas\">@rftexas</a>  solved it .  i am trying stratifiedkfold and it has shuffle argument with random_state but still facing the same issue</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 824890,
          "author_name": "shahules",
          "author_url": "",
          "post_date": "04/28/2020 16:43:14",
          "content": "<p>Are you shuffling the entire dataset ? Or just the validation? <a href=\"/pranshu29\">@pranshu29</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 824907,
          "author_name": "pranshu29",
          "author_url": "",
          "post_date": "04/28/2020 16:57:03",
          "content": "<p>Well i am using random state 42 and shuffle true in stratified k fold and validating the model on full validation data with no shuffle</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 827583,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "04/30/2020 11:52:56",
          "content": "<p>Sorry for the late answer....</p>\n\n<p>I am using the Tensorflow Data API. I tried to increase the buffer in the shuffle method, but nothing happens... </p>\n\n<p>I think given what <a href=\"/serigne\">@serigne</a> has suggested, I must write my own code generator. I'll work on that. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 823669,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "04/27/2020 19:49:19",
      "content": "<p>You may need to implement your own Data generator to shuffle the entire data at the end of each epoch, as the default TF datagenerator will shuffle only the next batch . </p>",
      "votes": null,
      "replies": [
        {
          "id": 824196,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "04/28/2020 08:08:29",
          "content": "<p>Thanks for the tips! I will!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 827549,
      "author_name": "fuzzywizard",
      "author_url": "",
      "post_date": "04/30/2020 11:28:01",
      "content": "<p>I was facing the same issue a few days back (PyTorch). I thought maybe I had worked out something wrong with the implementation. Although my training loss did not explode, it stayed more or less constant and the val_auc was stuck at 0.5 as you've mentioned.</p>\n\n<p>Well, thanks for posting this :)\nI'll try out the suggested approaches and let you know!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 831622,
      "author_name": "rftexas",
      "author_url": "",
      "post_date": "05/03/2020 13:46:22",
      "content": "<p>For those who might be wondering, I solved my issues by using a custom training loop function. </p>\n\n<p>Here is a great starting point:\n- <a href=\"https://colab.research.google.com/drive/169PfzM0kvtA5UP4k6Sl1yCG9tsE2MLia\">https://colab.research.google.com/drive/169PfzM0kvtA5UP4k6Sl1yCG9tsE2MLia</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 832396,
          "author_name": "dttung2905",
          "author_url": "",
          "post_date": "05/04/2020 05:14:02",
          "content": "<p>Thanks for sharing the resources with us. \nI'm trying to access it and got this error \n<code>There was an error loading this notebook. Ensure that the file is accessible and try again.\nInvalid Credentials</code></p>\n\n<p>Just wondering  because the colab notebook is not publicly accessible ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 832866,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "05/04/2020 13:25:33",
          "content": "<p>That's weird, I can access it... Are you logged in ?\nIt is publicly shared by François Chollet, Keras founder.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 833082,
          "author_name": "dttung2905",
          "author_url": "",
          "post_date": "05/04/2020 15:33:04",
          "content": "<p>I can access it now. My internet is weird sometimes :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 836847,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "05/07/2020 09:48:08",
          "content": "<p>Ok great! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 843003,
      "author_name": "seb6084",
      "author_url": "",
      "post_date": "05/11/2020 18:44:11",
      "content": "<p>Hello <a href=\"/rftexas\">@rftexas</a>, for an unknown reason, i'm facing now the issue 🤔 : training is fine and suddently, the auc drop to 0.5.\nI've check that my labels are correctly balanced with toxic and none toxic example.</p>\n\n<p>I'm already using tf custom loop. Do you have your own data generator ? built on top of tf.data.dataset ?\nThx</p>",
      "votes": null,
      "replies": [
        {
          "id": 843753,
          "author_name": "rftexas",
          "author_url": "",
          "post_date": "05/12/2020 08:36:38",
          "content": "<p>Hey! I was using my own data generator and used fit_generator. I don't know why sometimes it stops working. Now I'm using PyTorch because I have much more control over the pipeline.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "822279": "Hello everyone, \n\nI'm using Tensorflow and have been training my models with 200,000 samples so far. I wanted to try to train on more samples (e.g. 400,000 samples). I don't run into memory issues, but oddly enough my model doesn't learn...\n\nThe val_auc is stuck either at 0 or 0.5 while the loss is exploding. \n\nHas anyone faced the same issue and come up with a solution? \n\nAny help is appreciated. Many thanks in advance ;)",
    "822483": "rftexas I faced the same issue, and for recent runs, it is going OOM even for small batch sizes. Which model are your using?",
    "822484": "if you have changed batch size, you should to change learning rate proportionally. \n\n```\nbs=128, lr=0.4*1e-5  ---&gt; bs=64 lr=0.2*1e-5\n```\n\nMaybe It helps for you. High value for learning rate can crush pretrained weights.",
    "822510": "XLM-Roberta",
    "822511": "I haven't... But I'll try ! Thanks!",
    "822528": "I have the same problem too . But mine happens when I use pytorch from Abhishek when training on both training set combined.",
    "822768": "Are you sampling the dataset? 200k each class? @rftexas",
    "822975": "I'm training on tensorflow and pytorch with 1.5 million of records, no issue.\n0.5 means random guess : either you model is overfitting to your train data, either, there's an issue in the validation loop. Certainly not the case if you are using model.fit().\n\nBy the way, if you val loss is 0, be happy : predict the total opposite :)",
    "823004": "Thanks for your insight!",
    "823005": "Yes",
    "823076": "rftexas The same thing happens to me when I try to balance the data with 200k each.Any idea what is happening?",
    "823110": "Something similar happened to me once when I forgot to shuffle train dataset  - maybe because whole batches consisted of only negative or only positive samples.\nMy impression is that (pls correct if wrong) tensorflow dataset.shuffle(2048) takes the next 2048 samples and shuffles only those rows, not the whole dataset. If all those samples are from the same class then it might cause problem",
    "823222": "shonenkov Does the number of tpu devices matter too? I'm using pytorch if that helps...\nI am using 0.4*1e-5*8 because I am using the 8-core TPU for training",
    "823409": "amoghjrules  yes, if you use 8-core tpu you shoud multiply by 8 times",
    "823412": "https://arxiv.org/pdf/1706.02677.pdf",
    "823560": "isakev wow! I think you're right. I solved my issue by shuffling the data before feeding it to the generator.\n\n`from sklearn.utils import shuffle`\n\n`train = shuffle(train).reset_index(drop=True)`\n@rftexas can you try and tell if this works ? Solved in my case.",
    "823669": "You may need to implement your own Data generator to shuffle the entire data at the end of each epoch, as the default TF datagenerator will shuffle only the next batch .",
    "824196": "Thanks for the tips! I will!",
    "824200": "I really don't know why. I will try @serigne approach, and tell you!",
    "824202": "I'll try it today! I'll tell you!",
    "824310": "True, 2048 is the buffer size, meaning that the next 2048 samples are shuffled; not the entire dataset. I will work on that and tell you.",
    "824742": "hey encountering the same issue have you @rftexas  solved it .  i am trying stratifiedkfold and it has shuffle argument with random_state but still facing the same issue",
    "824890": "Are you shuffling the entire dataset ? Or just the validation? @pranshu29",
    "824907": "Well i am using random state 42 and shuffle true in stratified k fold and validating the model on full validation data with no shuffle",
    "825644": "How you guys sampling the dataset for 200k each class. Drop some idea, I didn't understand!",
    "827549": "I was facing the same issue a few days back (PyTorch). I thought maybe I had worked out something wrong with the implementation. Although my training loss did not explode, it stayed more or less constant and the val_auc was stuck at 0.5 as you've mentioned.\n\nWell, thanks for posting this :)\nI'll try out the suggested approaches and let you know!",
    "827583": "Sorry for the late answer....\n\nI am using the Tensorflow Data API. I tried to increase the buffer in the shuffle method, but nothing happens... \n\nI think given what @serigne has suggested, I must write my own code generator. I'll work on that.",
    "828754": "shahules  I think what is happening is that sampling the data will create dataset where first 200k will be 1 and next 200k will be 0 so , NN will see first see 200k 0s and then 200k 1s and not able to learn, and  dataset.shuffle(2048) is not working because acc to tf documentation buffer_size must be greater then length of whole dataset.\n https://www.tensorflow.org/api_docs/python/tf/data/Dataset#shuffle\n\nAlthough this is what I think.\nplease let me know If I misunderstood the question.",
    "831622": "For those who might be wondering, I solved my issues by using a custom training loop function. \n\nHere is a great starting point:\n- https://colab.research.google.com/drive/169PfzM0kvtA5UP4k6Sl1yCG9tsE2MLia",
    "832396": "Thanks for sharing the resources with us. \nI'm trying to access it and got this error \n`There was an error loading this notebook. Ensure that the file is accessible and try again.\nInvalid Credentials`\n\nJust wondering  because the colab notebook is not publicly accessible ?",
    "832866": "That's weird, I can access it... Are you logged in ?\nIt is publicly shared by François Chollet, Keras founder.",
    "833082": "I can access it now. My internet is weird sometimes :)",
    "836847": "Ok great!",
    "843003": "Hello @rftexas, for an unknown reason, i'm facing now the issue 🤔 : training is fine and suddently, the auc drop to 0.5.\nI've check that my labels are correctly balanced with toxic and none toxic example.\n\nI'm already using tf custom loop. Do you have your own data generator ? built on top of tf.data.dataset ?\nThx",
    "843753": "Hey! I was using my own data generator and used fit_generator. I don't know why sometimes it stops working. Now I'm using PyTorch because I have much more control over the pipeline."
  },
  "source": "meta"
}