{
  "id": 215954,
  "title": "Why retraining the model with pre-trained weights stuck at some level ?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/215954",
  "author_name": "Vatsal Mavani",
  "post_date": "2021-02-01T02:54:05.947000",
  "votes": 5,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Hello, </p>\n<p>I'm using pre-trained models with a <strong>timm</strong> library. I used <strong>Adam</strong> optimizer with <code>lr=1e-4</code> and <strong>TylorCrossEntropy</strong> Loss for 15 epochs but I didn't see any improvement in <code>val_accuracy</code> and <code>val_loss</code>. If you have any idea why is that happening or if there is something which I missing then give your suggestion and thanks in advance.</p>",
  "messages": [
    {
      "id": 1179956,
      "postDate": "2021-02-01T02:54:05.947Z",
      "content": "<p>Hello, </p>\n<p>I'm using pre-trained models with a <strong>timm</strong> library. I used <strong>Adam</strong> optimizer with <code>lr=1e-4</code> and <strong>TylorCrossEntropy</strong> Loss for 15 epochs but I didn't see any improvement in <code>val_accuracy</code> and <code>val_loss</code>. If you have any idea why is that happening or if there is something which I missing then give your suggestion and thanks in advance.</p>",
      "rawMarkdown": "Hello, \n\nI'm using pre-trained models with a **timm** library. I used **Adam** optimizer with `lr=1e-4` and **TylorCrossEntropy** Loss for 15 epochs but I didn't see any improvement in `val_accuracy` and `val_loss`. If you have any idea why is that happening or if there is something which I missing then give your suggestion and thanks in advance.",
      "votes": 4
    },
    {
      "id": 1182849,
      "postDate": "2021-02-02T16:09:00.467Z",
      "content": "<p>At such point your model is already converged/stuck at local minima(very close to global minima), there won't be much improvement after such point. You can try to converge your model using SGD if it helps but it'll take many more epochs and you might not see any improvements. Better is to train multiple models and ensemble.</p>",
      "rawMarkdown": "At such point your model is already converged/stuck at local minima(very close to global minima), there won't be much improvement after such point. You can try to converge your model using SGD if it helps but it'll take many more epochs and you might not see any improvements. Better is to train multiple models and ensemble.",
      "votes": 1
    },
    {
      "id": 1180670,
      "postDate": "2021-02-01T12:07:18.670Z",
      "content": "<p>as <a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a>  pointed out, most probably your code is broken, your weights may not be updating or there are many other possibilities as well. You need to at least share some code to know what is messed up. using sgd is a good option too, although adam works fine for me. use 0.001 as your initial learning rate and use cosine decay or reduceLRonplateu, to adjust LR.</p>",
      "rawMarkdown": "as @alexanderriedel  pointed out, most probably your code is broken, your weights may not be updating or there are many other possibilities as well. You need to at least share some code to know what is messed up. using sgd is a good option too, although adam works fine for me. use 0.001 as your initial learning rate and use cosine decay or reduceLRonplateu, to adjust LR.",
      "votes": 1,
      "replies": [
        {
          "id": 1183786,
          "postDate": "2021-02-03T08:07:39.870Z",
          "rawMarkdown": "",
          "isDeleted": true,
          "replies": [
            {
              "id": 1183811,
              "postDate": "2021-02-03T08:27:00.667Z",
              "content": "<p>Your training pipeline seems okay as per my understanding. I don't have much idea about the custom loss functions you have used. Also the utils script is not visible; so if the new df created is as expected, then there's no mistake I guess. You can try lr_schedulers or different optimizer as suggested and try tuning the different hyperparameters to get better score. </p>",
              "rawMarkdown": "Your training pipeline seems okay as per my understanding. I don't have much idea about the custom loss functions you have used. Also the utils script is not visible; so if the new df created is as expected, then there's no mistake I guess. You can try lr_schedulers or different optimizer as suggested and try tuning the different hyperparameters to get better score. \n"
            }
          ]
        }
      ]
    },
    {
      "id": 1180569,
      "postDate": "2021-02-01T10:35:41.157Z",
      "content": "<p>It really depends on so many factors! Maybe your model isn't capable of holding more information, maybe your model is underfitting or overfitting, maybe your code is broken.<br>\nPlease provide more information about your model, your augmentation strategies, you actual validation accuracy, your code, etc, to receive help</p>",
      "rawMarkdown": "It really depends on so many factors! Maybe your model isn't capable of holding more information, maybe your model is underfitting or overfitting, maybe your code is broken.\nPlease provide more information about your model, your augmentation strategies, you actual validation accuracy, your code, etc, to receive help",
      "votes": 2,
      "replies": [
        {
          "id": 1181081,
          "postDate": "2021-02-01T16:29:29.190Z",
          "content": "<p><a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a>, I used ResNext50 with lr=0.0001, but I experienced more than once and in other models also. And the code is also not broken because the model works well for 5-7 epochs, loss also decreases but after then it will start overfitting. I didn't use lr_scheduler but I will. Other then that is there anything that can cause this issue? </p>\n<p>And thank you again 😊.</p>",
          "rawMarkdown": "@alexanderriedel, I used ResNext50 with lr=0.0001, but I experienced more than once and in other models also. And the code is also not broken because the model works well for 5-7 epochs, loss also decreases but after then it will start overfitting. I didn't use lr_scheduler but I will. Other then that is there anything that can cause this issue? \n\nAnd thank you again 😊.",
          "votes": 1,
          "replies": [
            {
              "id": 1181932,
              "postDate": "2021-02-02T07:54:41.863Z",
              "content": "<p>As others pointed out, it'll be easier if you can share more details. Like train loss, val loss curves. If possible, share your training pipeline. Also what is the validation strategy you're using? If you're working with just 1 fold, its possible that your model has overfit. The lr_scheduler might also be an issue. Generally generally after training for sometime, its better to reduce learning rate. Else your model might be jumping between local minima. something like the 3rd image in this <a href=\"https://www.google.com/url?sa=i&amp;url=https%3A%2F%2Fwww.jeremyjordan.me%2Fnn-learning-rate%2F&amp;psig=AOvVaw1LnTIxVP4bw_MCJprjPsQ7&amp;ust=1612338749355000&amp;source=images&amp;cd=vfe&amp;ved=0CAIQjRxqFwoTCPjPjIbcyu4CFQAAAAAdAAAAABAD\" target=\"_blank\">link</a></p>",
              "rawMarkdown": "As others pointed out, it'll be easier if you can share more details. Like train loss, val loss curves. If possible, share your training pipeline. Also what is the validation strategy you're using? If you're working with just 1 fold, its possible that your model has overfit. The lr_scheduler might also be an issue. Generally generally after training for sometime, its better to reduce learning rate. Else your model might be jumping between local minima. something like the 3rd image in this [link](https://www.google.com/url?sa=i&url=https%3A%2F%2Fwww.jeremyjordan.me%2Fnn-learning-rate%2F&psig=AOvVaw1LnTIxVP4bw_MCJprjPsQ7&ust=1612338749355000&source=images&cd=vfe&ved=0CAIQjRxqFwoTCPjPjIbcyu4CFQAAAAAdAAAAABAD)"
            }
          ]
        },
        {
          "id": 1181112,
          "postDate": "2021-02-01T16:50:41.100Z",
          "content": "<p>What is your actual accuracy? Your validation accuracy will hardly get higher than 0.90x.</p>",
          "rawMarkdown": "What is your actual accuracy? Your validation accuracy will hardly get higher than 0.90x.\n"
        },
        {
          "id": 1183516,
          "postDate": "2021-02-03T04:06:09.647Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1183712,
          "postDate": "2021-02-03T06:59:01.603Z",
          "content": "<p>The important part, your training loop and function is missing…</p>",
          "rawMarkdown": "The important part, your training loop and function is missing..."
        },
        {
          "id": 1183757,
          "postDate": "2021-02-03T07:48:21.683Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1183816,
          "postDate": "2021-02-03T08:31:57.437Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1183824,
          "postDate": "2021-02-03T08:36:01.977Z",
          "content": "<p>I don't know how your <code>utils.create_folds</code> is working but maybe it's not stratifying your folds, so you should check out <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html\" target=\"_blank\">StratifiedKFold</a>. </p>\n<p>Also something about that whole folding process looks fishy, as this function doesn't take the amount of folds as an argument?</p>\n<p>But again, what is your actual accuracy? Maybe it's alright already reagrding your pipeline?</p>",
          "rawMarkdown": "I don't know how your `utils.create_folds` is working but maybe it's not stratifying your folds, so you should check out [StratifiedKFold](https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html). \n\nAlso something about that whole folding process looks fishy, as this function doesn't take the amount of folds as an argument?\n\nBut again, what is your actual accuracy? Maybe it's alright already reagrding your pipeline?",
          "replies": [
            {
              "id": 1183941,
              "postDate": "2021-02-03T10:02:21.827Z",
              "content": "<p>I'm already using <strong>StratifiedKFold</strong> in <code>utils.create_folds</code>.</p>\n<p>I don't understand this,</p>\n<blockquote>\n  <p>Also something about that whole folding process looks fishy, as this function doesn't take the amount of folds as an argument?</p>\n</blockquote>\n<p>After 15 epochs,<br>\nTrain accuracy: 0.93% to 0.95%<br>\nVal accuracy: 0.87% to 0.88%</p>",
              "rawMarkdown": "I'm already using **StratifiedKFold** in `utils.create_folds`.\n\nI don't understand this,\n> Also something about that whole folding process looks fishy, as this function doesn't take the amount of folds as an argument?\n\nAfter 15 epochs,\nTrain accuracy: 0.93% to 0.95%\nVal accuracy: 0.87% to 0.88%",
              "votes": 1
            }
          ]
        },
        {
          "id": 1183981,
          "postDate": "2021-02-03T10:31:02.270Z",
          "content": "<p>Ah ok! So you hardcoded the five folds in your <code>utils.create_folds</code>?</p>\n<p>You values for val accuracy are normal for this competition and what i experienced myself and read from others. If you don't do any fancy stuff, this is roughly what you will get per fold and model</p>",
          "rawMarkdown": "Ah ok! So you hardcoded the five folds in your `utils.create_folds`?\n\nYou values for val accuracy are normal for this competition and what i experienced myself and read from others. If you don't do any fancy stuff, this is roughly what you will get per fold and model"
        },
        {
          "id": 1183983,
          "postDate": "2021-02-03T10:33:06.103Z",
          "content": "<pre><code>import numpy as np\nimport pandas as pd\n\nfrom sklearn import model_selection\n\n\ndef create_folds(data, target_col_name='target'):\n\n    # create new column 'kfold' and fill it with '-1'\n    data['kfold'] = -1\n\n    # randomize the rows of the data\n    data = data.sample(frac=1).reset_index(drop=True)\n\n    # calculate number of bins\n    num_bins = int(np.floor(1 + np.log2(len(data))))\n\n    # set bin targets\n    data.loc[:, 'bins'] = pd.cut(\n        data[target_col_name], bins=num_bins, labels=False\n    )\n\n    # initiate the kfold class\n    kf = model_selection.StratifiedKFold(n_splits=5)\n\n    # fill the new 'kfold' column\n    for f, (t_,v_) in enumerate(kf.split(X=data, y=data.bins.values)):\n        data.loc[v_, 'kfold'] = f\n\n    # drop the bins column\n    data = data.drop('bins', axis=1)\n\n    return data\n</code></pre>\n<p>And to increase accuracy, any tips?</p>",
          "rawMarkdown": "```\nimport numpy as np\nimport pandas as pd\n\nfrom sklearn import model_selection\n\n\ndef create_folds(data, target_col_name='target'):\n\n    # create new column 'kfold' and fill it with '-1'\n    data['kfold'] = -1\n\n    # randomize the rows of the data\n    data = data.sample(frac=1).reset_index(drop=True)\n\n    # calculate number of bins\n    num_bins = int(np.floor(1 + np.log2(len(data))))\n\n    # set bin targets\n    data.loc[:, 'bins'] = pd.cut(\n        data[target_col_name], bins=num_bins, labels=False\n    )\n\n    # initiate the kfold class\n    kf = model_selection.StratifiedKFold(n_splits=5)\n\n    # fill the new 'kfold' column\n    for f, (t_,v_) in enumerate(kf.split(X=data, y=data.bins.values)):\n        data.loc[v_, 'kfold'] = f\n\n    # drop the bins column\n    data = data.drop('bins', axis=1)\n\n    return data\n```\n\nAnd to increase accuracy, any tips?"
        },
        {
          "id": 1183991,
          "postDate": "2021-02-03T10:40:48.873Z",
          "content": "<p>Read some discussions here, check out some noteboks, i can't give you much more advise than that</p>",
          "rawMarkdown": "Read some discussions here, check out some noteboks, i can't give you much more advise than that"
        }
      ]
    },
    {
      "id": 1190287,
      "postDate": "2021-02-07T16:09:39.750Z",
      "content": "<p>Do you train the whole layers of the pre-trained model from the scratch on the data or just the output layer? Do you think 15 epochs are enough to train and converge ? </p>",
      "rawMarkdown": "Do you train the whole layers of the pre-trained model from the scratch on the data or just the output layer? Do you think 15 epochs are enough to train and converge ? ",
      "replies": [
        {
          "id": 1190333,
          "postDate": "2021-02-07T16:43:37.433Z",
          "content": "<p>Thank you for taking interest, I retrained the model and now I'm using scheduler and problem is also almost solved.</p>",
          "rawMarkdown": "Thank you for taking interest, I retrained the model and now I'm using scheduler and problem is also almost solved.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1188011,
      "postDate": "2021-02-05T21:47:12.173Z",
      "content": "<p>interesting topic</p>",
      "rawMarkdown": "interesting topic"
    },
    {
      "id": 1181311,
      "postDate": "2021-02-01T19:30:44.277Z",
      "content": "<p>Are you sure that is an appropriate learning rate?</p>",
      "rawMarkdown": "Are you sure that is an appropriate learning rate?"
    },
    {
      "id": 1180020,
      "postDate": "2021-02-01T03:43:02.393Z",
      "content": "<p>Try SGD instead of Adam with a Lr scheduler like CosineDecayRestarts </p>",
      "rawMarkdown": "Try SGD instead of Adam with a Lr scheduler like CosineDecayRestarts ",
      "replies": [
        {
          "id": 1180174,
          "postDate": "2021-02-01T06:22:55.063Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1180173,
          "postDate": "2021-02-01T06:22:55.063Z",
          "content": "<p>Sure I'll try but can you also tell me the reason behind this.</p>",
          "rawMarkdown": "Sure I'll try but can you also tell me the reason behind this.",
          "votes": 1
        },
        {
          "id": 1180546,
          "postDate": "2021-02-01T10:16:26.023Z",
          "content": "<p>The model must be stuck on a point because of a low or a very high lr. Please SGD is better in my case,</p>",
          "rawMarkdown": "The model must be stuck on a point because of a low or a very high lr. Please SGD is better in my case,"
        }
      ]
    },
    {
      "id": 1210528,
      "postDate": "2021-02-19T13:53:14.250Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1182849,
      "author_name": "Ankush kuwar",
      "author_url": "",
      "post_date": "2021-02-02T16:09:00.467000",
      "content": "<p>At such point your model is already converged/stuck at local minima(very close to global minima), there won't be much improvement after such point. You can try to converge your model using SGD if it helps but it'll take many more epochs and you might not see any improvements. Better is to train multiple models and ensemble.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1180670,
      "author_name": "Mohneesh_Sreegirisetty",
      "author_url": "",
      "post_date": "2021-02-01T12:07:18.670000",
      "content": "<p>as <a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a>  pointed out, most probably your code is broken, your weights may not be updating or there are many other possibilities as well. You need to at least share some code to know what is messed up. using sgd is a good option too, although adam works fine for me. use 0.001 as your initial learning rate and use cosine decay or reduceLRonplateu, to adjust LR.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1183786,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-03T08:07:39.870000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 1183811,
              "author_name": "SuryaJR_Rafl",
              "author_url": "",
              "post_date": "2021-02-03T08:27:00.667000",
              "content": "<p>Your training pipeline seems okay as per my understanding. I don't have much idea about the custom loss functions you have used. Also the utils script is not visible; so if the new df created is as expected, then there's no mistake I guess. You can try lr_schedulers or different optimizer as suggested and try tuning the different hyperparameters to get better score. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1180569,
      "author_name": "Alexander Riedel",
      "author_url": "",
      "post_date": "2021-02-01T10:35:41.157000",
      "content": "<p>It really depends on so many factors! Maybe your model isn't capable of holding more information, maybe your model is underfitting or overfitting, maybe your code is broken.<br>\nPlease provide more information about your model, your augmentation strategies, you actual validation accuracy, your code, etc, to receive help</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1181081,
          "author_name": "Vatsal Mavani",
          "author_url": "",
          "post_date": "2021-02-01T16:29:29.190000",
          "content": "<p><a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a>, I used ResNext50 with lr=0.0001, but I experienced more than once and in other models also. And the code is also not broken because the model works well for 5-7 epochs, loss also decreases but after then it will start overfitting. I didn't use lr_scheduler but I will. Other then that is there anything that can cause this issue? </p>\n<p>And thank you again 😊.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 1181932,
              "author_name": "SuryaJR_Rafl",
              "author_url": "",
              "post_date": "2021-02-02T07:54:41.863000",
              "content": "<p>As others pointed out, it'll be easier if you can share more details. Like train loss, val loss curves. If possible, share your training pipeline. Also what is the validation strategy you're using? If you're working with just 1 fold, its possible that your model has overfit. The lr_scheduler might also be an issue. Generally generally after training for sometime, its better to reduce learning rate. Else your model might be jumping between local minima. something like the 3rd image in this <a href=\"https://www.google.com/url?sa=i&amp;url=https%3A%2F%2Fwww.jeremyjordan.me%2Fnn-learning-rate%2F&amp;psig=AOvVaw1LnTIxVP4bw_MCJprjPsQ7&amp;ust=1612338749355000&amp;source=images&amp;cd=vfe&amp;ved=0CAIQjRxqFwoTCPjPjIbcyu4CFQAAAAAdAAAAABAD\" target=\"_blank\">link</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 1181112,
          "author_name": "Alexander Riedel",
          "author_url": "",
          "post_date": "2021-02-01T16:50:41.100000",
          "content": "<p>What is your actual accuracy? Your validation accuracy will hardly get higher than 0.90x.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1183516,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-03T04:06:09.647000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1183712,
          "author_name": "Alexander Riedel",
          "author_url": "",
          "post_date": "2021-02-03T06:59:01.603000",
          "content": "<p>The important part, your training loop and function is missing…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1183757,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-03T07:48:21.683000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1183816,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-03T08:31:57.437000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1183824,
          "author_name": "Alexander Riedel",
          "author_url": "",
          "post_date": "2021-02-03T08:36:01.977000",
          "content": "<p>I don't know how your <code>utils.create_folds</code> is working but maybe it's not stratifying your folds, so you should check out <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.StratifiedKFold.html\" target=\"_blank\">StratifiedKFold</a>. </p>\n<p>Also something about that whole folding process looks fishy, as this function doesn't take the amount of folds as an argument?</p>\n<p>But again, what is your actual accuracy? Maybe it's alright already reagrding your pipeline?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 1183941,
              "author_name": "Vatsal Mavani",
              "author_url": "",
              "post_date": "2021-02-03T10:02:21.827000",
              "content": "<p>I'm already using <strong>StratifiedKFold</strong> in <code>utils.create_folds</code>.</p>\n<p>I don't understand this,</p>\n<blockquote>\n  <p>Also something about that whole folding process looks fishy, as this function doesn't take the amount of folds as an argument?</p>\n</blockquote>\n<p>After 15 epochs,<br>\nTrain accuracy: 0.93% to 0.95%<br>\nVal accuracy: 0.87% to 0.88%</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 1183981,
          "author_name": "Alexander Riedel",
          "author_url": "",
          "post_date": "2021-02-03T10:31:02.270000",
          "content": "<p>Ah ok! So you hardcoded the five folds in your <code>utils.create_folds</code>?</p>\n<p>You values for val accuracy are normal for this competition and what i experienced myself and read from others. If you don't do any fancy stuff, this is roughly what you will get per fold and model</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1183983,
          "author_name": "Vatsal Mavani",
          "author_url": "",
          "post_date": "2021-02-03T10:33:06.103000",
          "content": "<pre><code>import numpy as np\nimport pandas as pd\n\nfrom sklearn import model_selection\n\n\ndef create_folds(data, target_col_name='target'):\n\n    # create new column 'kfold' and fill it with '-1'\n    data['kfold'] = -1\n\n    # randomize the rows of the data\n    data = data.sample(frac=1).reset_index(drop=True)\n\n    # calculate number of bins\n    num_bins = int(np.floor(1 + np.log2(len(data))))\n\n    # set bin targets\n    data.loc[:, 'bins'] = pd.cut(\n        data[target_col_name], bins=num_bins, labels=False\n    )\n\n    # initiate the kfold class\n    kf = model_selection.StratifiedKFold(n_splits=5)\n\n    # fill the new 'kfold' column\n    for f, (t_,v_) in enumerate(kf.split(X=data, y=data.bins.values)):\n        data.loc[v_, 'kfold'] = f\n\n    # drop the bins column\n    data = data.drop('bins', axis=1)\n\n    return data\n</code></pre>\n<p>And to increase accuracy, any tips?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1183991,
          "author_name": "Alexander Riedel",
          "author_url": "",
          "post_date": "2021-02-03T10:40:48.873000",
          "content": "<p>Read some discussions here, check out some noteboks, i can't give you much more advise than that</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1190287,
      "author_name": "Omar Mohamed",
      "author_url": "",
      "post_date": "2021-02-07T16:09:39.750000",
      "content": "<p>Do you train the whole layers of the pre-trained model from the scratch on the data or just the output layer? Do you think 15 epochs are enough to train and converge ? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1190333,
          "author_name": "Vatsal Mavani",
          "author_url": "",
          "post_date": "2021-02-07T16:43:37.433000",
          "content": "<p>Thank you for taking interest, I retrained the model and now I'm using scheduler and problem is also almost solved.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1188011,
      "author_name": "Mishra Rayn",
      "author_url": "",
      "post_date": "2021-02-05T21:47:12.173000",
      "content": "<p>interesting topic</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1181311,
      "author_name": "Gurhar Khalsa",
      "author_url": "",
      "post_date": "2021-02-01T19:30:44.277000",
      "content": "<p>Are you sure that is an appropriate learning rate?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1180020,
      "author_name": "Mithil Salunkhe",
      "author_url": "",
      "post_date": "2021-02-01T03:43:02.393000",
      "content": "<p>Try SGD instead of Adam with a Lr scheduler like CosineDecayRestarts </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1180174,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-01T06:22:55.063000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1180173,
          "author_name": "Vatsal Mavani",
          "author_url": "",
          "post_date": "2021-02-01T06:22:55.063000",
          "content": "<p>Sure I'll try but can you also tell me the reason behind this.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1180546,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-02-01T10:16:26.023000",
          "content": "<p>The model must be stuck on a point because of a low or a very high lr. Please SGD is better in my case,</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1210528,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-19T13:53:14.250000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1179956": "Hello, \n\nI'm using pre-trained models with a **timm** library. I used **Adam** optimizer with `lr=1e-4` and **TylorCrossEntropy** Loss for 15 epochs but I didn't see any improvement in `val_accuracy` and `val_loss`. If you have any idea why is that happening or if there is something which I missing then give your suggestion and thanks in advance.",
    "1182849": "At such point your model is already converged/stuck at local minima(very close to global minima), there won't be much improvement after such point. You can try to converge your model using SGD if it helps but it'll take many more epochs and you might not see any improvements. Better is to train multiple models and ensemble.",
    "1180670": "as @alexanderriedel  pointed out, most probably your code is broken, your weights may not be updating or there are many other possibilities as well. You need to at least share some code to know what is messed up. using sgd is a good option too, although adam works fine for me. use 0.001 as your initial learning rate and use cosine decay or reduceLRonplateu, to adjust LR.",
    "1180569": "It really depends on so many factors! Maybe your model isn't capable of holding more information, maybe your model is underfitting or overfitting, maybe your code is broken.\nPlease provide more information about your model, your augmentation strategies, you actual validation accuracy, your code, etc, to receive help",
    "1190287": "Do you train the whole layers of the pre-trained model from the scratch on the data or just the output layer? Do you think 15 epochs are enough to train and converge ? ",
    "1188011": "interesting topic",
    "1181311": "Are you sure that is an appropriate learning rate?",
    "1180020": "Try SGD instead of Adam with a Lr scheduler like CosineDecayRestarts ",
    "1210528": ""
  }
}