{
  "id": 43713,
  "title": "Sudden accuracy drop when training on all labels",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/43713",
  "author_name": "",
  "post_date": "2017-11-18T09:53:54.428214500Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello Everyone</p>\n\n<p>Instead of training on the required commands only, I thought it would be better if the model is capable of differentiating between the various commands (all 31 of them) and then I would replace the labels we don't need during post-processing with the 'unknown' label.</p>\n\n<p>Nonetheless, whilst I am trying to train a Model on this set, the loss simply blows up all of a sudden and continues to fluctuate minutely. (stuck in a possible local minima??? but all the time!?)</p>\n\n<p>I have ruled out the following possibilities:</p>\n\n<ol>\n<li>blowing gradients: used gradient clipping, used lr in range from [1e-5, 0.1]</li>\n<li>presence of RNNs: the same happens with and without a Recurrent Layer in the model.</li>\n<li>the same model that gave my best result shows the same behaviour, so does the current worst model</li>\n<li>Bug in data pipeline: generates correct label for the correct file, used in a lot of projects</li>\n<li>I have even used different optimizers, still the same.</li>\n<li>API: Same thing happens on Keras and Pytorch</li>\n</ol>\n\n<p>Any clues what might be happening?</p>",
  "messages": [
    {
      "id": "245392",
      "postDate": "11/18/2017 09:53:54",
      "content": "<p>Hello Everyone</p>\n\n<p>Instead of training on the required commands only, I thought it would be better if the model is capable of differentiating between the various commands (all 31 of them) and then I would replace the labels we don't need during post-processing with the 'unknown' label.</p>\n\n<p>Nonetheless, whilst I am trying to train a Model on this set, the loss simply blows up all of a sudden and continues to fluctuate minutely. (stuck in a possible local minima??? but all the time!?)</p>\n\n<p>I have ruled out the following possibilities:</p>\n\n<ol>\n<li>blowing gradients: used gradient clipping, used lr in range from [1e-5, 0.1]</li>\n<li>presence of RNNs: the same happens with and without a Recurrent Layer in the model.</li>\n<li>the same model that gave my best result shows the same behaviour, so does the current worst model</li>\n<li>Bug in data pipeline: generates correct label for the correct file, used in a lot of projects</li>\n<li>I have even used different optimizers, still the same.</li>\n<li>API: Same thing happens on Keras and Pytorch</li>\n</ol>\n\n<p>Any clues what might be happening?</p>",
      "rawMarkdown": "Hello Everyone\n\nInstead of training on the required commands only, I thought it would be better if the model is capable of differentiating between the various commands (all 31 of them) and then I would replace the labels we don't need during post-processing with the 'unknown' label.\n\nNonetheless, whilst I am trying to train a Model on this set, the loss simply blows up all of a sudden and continues to fluctuate minutely. (stuck in a possible local minima??? but all the time!?)\n\nI have ruled out the following possibilities:\n\n 1. blowing gradients: used gradient clipping, used lr in range from [1e-5, 0.1]\n 2. presence of RNNs: the same happens with and without a Recurrent Layer in the model.\n 3. the same model that gave my best result shows the same behaviour, so does the current worst model\n 4. Bug in data pipeline: generates correct label for the correct file, used in a lot of projects\n 5. I have even used different optimizers, still the same.\n 6. API: Same thing happens on Keras and Pytorch\n\nAny clues what might be happening?",
      "votes": null
    },
    {
      "id": "245477",
      "postDate": "11/18/2017 15:42:39",
      "content": "<p>Which loss, local on a cv/holdout set? Or leaderboard loss. multiclass loss locally could go up because more opportunity to misclassify between classes that don’t matter. </p>",
      "rawMarkdown": "Which loss, local on a cv/holdout set? Or leaderboard loss. multiclass loss locally could go up because more opportunity to misclassify between classes that don’t matter.",
      "votes": null
    },
    {
      "id": "245490",
      "postDate": "11/18/2017 16:12:38",
      "content": "<p>First i assume there is no bug in the code.</p>\n\n<p>Next you may want to find out where the error is from. For example the confusion of class may comes from 'three' and 'cat', which will not affect the 'unknown' class in the earlier setting of 12 classes.</p>\n\n<p>Another thing you can try is to increase the model parameters. See if you can overfit and drive the train error to zero.</p>\n\n<p>Finally, fluctuations are due to:</p>\n\n<p>Numerical instability </p>\n\n<p>Wrong label ( i remembered reading some where that there is some wrong label in the google command set. One of the unknown is labelled as slient. But i am not very sure)</p>\n\n<p>Outlier</p>\n\n<p>Too small batch size or too large learning rate</p>\n\n<p>Bug</p>",
      "rawMarkdown": "First i assume there is no bug in the code.\n\nNext you may want to find out where the error is from. For example the confusion of class may comes from 'three' and 'cat', which will not affect the 'unknown' class in the earlier setting of 12 classes.\n\nAnother thing you can try is to increase the model parameters. See if you can overfit and drive the train error to zero.\n\nFinally, fluctuations are due to:\n\nNumerical instability \n\nWrong label ( i remembered reading some where that there is some wrong label in the google command set. One of the unknown is labelled as slient. But i am not very sure)\n\nOutlier\n\nToo small batch size or too large learning rate\n\n\nBug",
      "votes": null
    },
    {
      "id": "245656",
      "postDate": "11/19/2017 06:48:08",
      "content": "<p>Local loss while training. As a matter of fact on the first epoch itself. I am not submitting it to the lb.</p>",
      "rawMarkdown": "Local loss while training. As a matter of fact on the first epoch itself. I am not submitting it to the lb.",
      "votes": null
    },
    {
      "id": "260383",
      "postDate": "12/20/2017 05:16:56",
      "content": "<p>Could it be because of ctc loss ? if you are using it. Also, the network needs to discriminate more classes with same amount of data</p>",
      "rawMarkdown": "Could it be because of ctc loss ? if you are using it. Also, the network needs to discriminate more classes with same amount of data",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 245477,
      "author_name": "thoolihan",
      "author_url": "",
      "post_date": "11/18/2017 15:42:39",
      "content": "<p>Which loss, local on a cv/holdout set? Or leaderboard loss. multiclass loss locally could go up because more opportunity to misclassify between classes that don’t matter. </p>",
      "votes": null,
      "replies": [
        {
          "id": 245656,
          "author_name": "yadavsarthak",
          "author_url": "",
          "post_date": "11/19/2017 06:48:08",
          "content": "<p>Local loss while training. As a matter of fact on the first epoch itself. I am not submitting it to the lb.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 245490,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/18/2017 16:12:38",
      "content": "<p>First i assume there is no bug in the code.</p>\n\n<p>Next you may want to find out where the error is from. For example the confusion of class may comes from 'three' and 'cat', which will not affect the 'unknown' class in the earlier setting of 12 classes.</p>\n\n<p>Another thing you can try is to increase the model parameters. See if you can overfit and drive the train error to zero.</p>\n\n<p>Finally, fluctuations are due to:</p>\n\n<p>Numerical instability </p>\n\n<p>Wrong label ( i remembered reading some where that there is some wrong label in the google command set. One of the unknown is labelled as slient. But i am not very sure)</p>\n\n<p>Outlier</p>\n\n<p>Too small batch size or too large learning rate</p>\n\n<p>Bug</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 260383,
      "author_name": "rteja1113",
      "author_url": "",
      "post_date": "12/20/2017 05:16:56",
      "content": "<p>Could it be because of ctc loss ? if you are using it. Also, the network needs to discriminate more classes with same amount of data</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "245392": "Hello Everyone\n\nInstead of training on the required commands only, I thought it would be better if the model is capable of differentiating between the various commands (all 31 of them) and then I would replace the labels we don't need during post-processing with the 'unknown' label.\n\nNonetheless, whilst I am trying to train a Model on this set, the loss simply blows up all of a sudden and continues to fluctuate minutely. (stuck in a possible local minima??? but all the time!?)\n\nI have ruled out the following possibilities:\n\n 1. blowing gradients: used gradient clipping, used lr in range from [1e-5, 0.1]\n 2. presence of RNNs: the same happens with and without a Recurrent Layer in the model.\n 3. the same model that gave my best result shows the same behaviour, so does the current worst model\n 4. Bug in data pipeline: generates correct label for the correct file, used in a lot of projects\n 5. I have even used different optimizers, still the same.\n 6. API: Same thing happens on Keras and Pytorch\n\nAny clues what might be happening?",
    "245477": "Which loss, local on a cv/holdout set? Or leaderboard loss. multiclass loss locally could go up because more opportunity to misclassify between classes that don’t matter.",
    "245490": "First i assume there is no bug in the code.\n\nNext you may want to find out where the error is from. For example the confusion of class may comes from 'three' and 'cat', which will not affect the 'unknown' class in the earlier setting of 12 classes.\n\nAnother thing you can try is to increase the model parameters. See if you can overfit and drive the train error to zero.\n\nFinally, fluctuations are due to:\n\nNumerical instability \n\nWrong label ( i remembered reading some where that there is some wrong label in the google command set. One of the unknown is labelled as slient. But i am not very sure)\n\nOutlier\n\nToo small batch size or too large learning rate\n\n\nBug",
    "245656": "Local loss while training. As a matter of fact on the first epoch itself. I am not submitting it to the lb.",
    "260383": "Could it be because of ctc loss ? if you are using it. Also, the network needs to discriminate more classes with same amount of data"
  },
  "source": "meta"
}