{
  "id": 204784,
  "title": "Why is my model not improving during training time?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/204784",
  "author_name": "",
  "post_date": "2020-12-16T19:25:23.207436900Z",
  "votes": 2,
  "comment_count": 14,
  "views": 0,
  "content": "<p>when I call fit() on my model it starts training. In the beginning of the training its already at a pretty hight accuracy around 0.81. But after that it doesn't really improve and stays around 0.82 accuracy after 50 epochs.<br>\nShouldn't it improve further during training?</p>",
  "messages": [
    {
      "id": "1116042",
      "postDate": "12/16/2020 19:25:23",
      "content": "<p>when I call fit() on my model it starts training. In the beginning of the training its already at a pretty hight accuracy around 0.81. But after that it doesn't really improve and stays around 0.82 accuracy after 50 epochs.<br>\nShouldn't it improve further during training?</p>",
      "rawMarkdown": "when I call fit() on my model it starts training. In the beginning of the training its already at a pretty hight accuracy around 0.81. But after that it doesn't really improve and stays around 0.82 accuracy after 50 epochs.\nShouldn't it improve further during training?",
      "votes": null
    },
    {
      "id": "1116125",
      "postDate": "12/16/2020 21:42:48",
      "content": "<p>Is that the accuracy on the training data or a validation set that you have not used for training? The former can certainly (unless you regularize) improve up to 1.0, for the latter there is going to be some limit that depends on various things (e.g. type of model - see e.g. <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203111\" target=\"_blank\">this thread</a> what validation accuracy people are seeing with single models).  It's plausible that some set-up might just max out at 0.82 (either too small a model, not using pretrained weights, not a good set of image augmentations, not a good learning rate schedule etc.).</p>",
      "rawMarkdown": "Is that the accuracy on the training data or a validation set that you have not used for training? The former can certainly (unless you regularize) improve up to 1.0, for the latter there is going to be some limit that depends on various things (e.g. type of model - see e.g. [this thread](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203111) what validation accuracy people are seeing with single models).  It's plausible that some set-up might just max out at 0.82 (either too small a model, not using pretrained weights, not a good set of image augmentations, not a good learning rate schedule etc.).",
      "votes": null
    },
    {
      "id": "1116692",
      "postDate": "12/17/2020 11:57:07",
      "content": "<p>According to my limited knowledge, You need to go deeper with the model</p>",
      "rawMarkdown": "According to my limited knowledge, You need to go deeper with the model",
      "votes": null
    },
    {
      "id": "1116985",
      "postDate": "12/17/2020 16:06:56",
      "content": "<p>I'm using a pertained Exception network I would say thats pretty deep already</p>",
      "rawMarkdown": "I'm using a pertained Exception network I would say thats pretty deep already",
      "votes": null
    },
    {
      "id": "1116986",
      "postDate": "12/17/2020 16:09:07",
      "content": "<p>I'm using a pertained Exception network and using different kinds of augmentation methods (cropping, flipping etc.). When I have my lr around 0.001 thats where I get the best results by far everything else is around 0.6. </p>",
      "rawMarkdown": "I'm using a pertained Exception network and using different kinds of augmentation methods (cropping, flipping etc.). When I have my lr around 0.001 thats where I get the best results by far everything else is around 0.6.",
      "votes": null
    },
    {
      "id": "1116987",
      "postDate": "12/17/2020 16:09:36",
      "content": "<p>Which model is it and input size?</p>",
      "rawMarkdown": "Which model is it and input size?",
      "votes": null
    },
    {
      "id": "1116992",
      "postDate": "12/17/2020 16:13:37",
      "content": "<p>Training longer does not mean that it will learn more, you probably need more regulization in your pipeline. Maybe try resnets or efficientnets, they are usually the way to go</p>",
      "rawMarkdown": "Training longer does not mean that it will learn more, you probably need more regulization in your pipeline. Maybe try resnets or efficientnets, they are usually the way to go",
      "votes": null
    },
    {
      "id": "1116996",
      "postDate": "12/17/2020 16:17:55",
      "content": "<p>*an pretrained Xception network + global average + dense + output and the input size = [512,512,3]</p>",
      "rawMarkdown": "*an pretrained Xception network + global average + dense + output and the input size = [512,512,3]",
      "votes": null
    },
    {
      "id": "1116997",
      "postDate": "12/17/2020 16:19:05",
      "content": "<p>I thought that since resnets achieved lower accuracy on imagined that xception-nets would perform better</p>",
      "rawMarkdown": "I thought that since resnets achieved lower accuracy on imagined that xception-nets would perform better",
      "votes": null
    },
    {
      "id": "1117022",
      "postDate": "12/17/2020 16:45:21",
      "content": "<p>Yes it might be true but different architechture can work better on different datasets. There are newer versions of resnets like resnext,seresnet,seresnext,wideresnet,resnest and more. These are probably faster and better than xception. Also, with newer augmentation techniques some architechtures are better than the others. I would recommend to always try resnets first, in my experience they usually give good results</p>",
      "rawMarkdown": "Yes it might be true but different architechture can work better on different datasets. There are newer versions of resnets like resnext,seresnet,seresnext,wideresnet,resnest and more. These are probably faster and better than xception. Also, with newer augmentation techniques some architechtures are better than the others. I would recommend to always try resnets first, in my experience they usually give good results",
      "votes": null
    },
    {
      "id": "1117038",
      "postDate": "12/17/2020 17:07:09",
      "content": "<p>ok thank you for the detailed reply 🙌🏻</p>",
      "rawMarkdown": "ok thank you for the detailed reply 🙌🏻",
      "votes": null
    },
    {
      "id": "1117135",
      "postDate": "12/17/2020 18:44:53",
      "content": "<p>Also some others tips i can give you: <br>\n-Start with a lower resolution with samller architechtures and no augmentations to have a baseline. <br>\n-Find some good parameters with this set up<br>\n-Add some augmentations and find the best ones<br>\n-Tune some hyperparameters like epochs,weight decay, dropout, batchsize etc<br>\n-Then increase resolution and retune a bit<br>\n-Then increase architechture size</p>\n<p>This way you iterate faster and find better augs that work for this dataset. Increasing the depths of you network will only increase by a couple of % the performance, your pipeline is much more important.</p>",
      "rawMarkdown": "Also some others tips i can give you: \n-Start with a lower resolution with samller architechtures and no augmentations to have a baseline. \n-Find some good parameters with this set up\n-Add some augmentations and find the best ones\n-Tune some hyperparameters like epochs,weight decay, dropout, batchsize etc\n-Then increase resolution and retune a bit\n-Then increase architechture size\n\nThis way you iterate faster and find better augs that work for this dataset. Increasing the depths of you network will only increase by a couple of % the performance, your pipeline is much more important.",
      "votes": null
    },
    {
      "id": "1117869",
      "postDate": "12/18/2020 14:39:18",
      "content": "<p>how can I start with a lower resolution? Do you mean cropping the down the image in size? (say instead of [512,512,3] cropping to [300,300,3]) or how do you lower the resolution?</p>",
      "rawMarkdown": "how can I start with a lower resolution? Do you mean cropping the down the image in size? (say instead of [512,512,3] cropping to [300,300,3]) or how do you lower the resolution?",
      "votes": null
    },
    {
      "id": "1117882",
      "postDate": "12/18/2020 14:52:01",
      "content": "<p>In your image augmentation you would resize it to a desired resolution, maybe start with 256x256 or 384x384. Make sure to resize for training images and validation images!</p>",
      "rawMarkdown": "In your image augmentation you would resize it to a desired resolution, maybe start with 256x256 or 384x384. Make sure to resize for training images and validation images!",
      "votes": null
    },
    {
      "id": "1118147",
      "postDate": "12/18/2020 19:16:20",
      "content": "<p>alright I'll check it out.. thanks</p>",
      "rawMarkdown": "alright I'll check it out.. thanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1116125,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "12/16/2020 21:42:48",
      "content": "<p>Is that the accuracy on the training data or a validation set that you have not used for training? The former can certainly (unless you regularize) improve up to 1.0, for the latter there is going to be some limit that depends on various things (e.g. type of model - see e.g. <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203111\" target=\"_blank\">this thread</a> what validation accuracy people are seeing with single models).  It's plausible that some set-up might just max out at 0.82 (either too small a model, not using pretrained weights, not a good set of image augmentations, not a good learning rate schedule etc.).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1116986,
          "author_name": "docphilipp",
          "author_url": "",
          "post_date": "12/17/2020 16:09:07",
          "content": "<p>I'm using a pertained Exception network and using different kinds of augmentation methods (cropping, flipping etc.). When I have my lr around 0.001 thats where I get the best results by far everything else is around 0.6. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1116692,
      "author_name": "deepdreamx",
      "author_url": "",
      "post_date": "12/17/2020 11:57:07",
      "content": "<p>According to my limited knowledge, You need to go deeper with the model</p>",
      "votes": null,
      "replies": [
        {
          "id": 1116985,
          "author_name": "docphilipp",
          "author_url": "",
          "post_date": "12/17/2020 16:06:56",
          "content": "<p>I'm using a pertained Exception network I would say thats pretty deep already</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1116987,
          "author_name": "deepdreamx",
          "author_url": "",
          "post_date": "12/17/2020 16:09:36",
          "content": "<p>Which model is it and input size?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1116996,
          "author_name": "docphilipp",
          "author_url": "",
          "post_date": "12/17/2020 16:17:55",
          "content": "<p>*an pretrained Xception network + global average + dense + output and the input size = [512,512,3]</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1116992,
      "author_name": "yannmajewski",
      "author_url": "",
      "post_date": "12/17/2020 16:13:37",
      "content": "<p>Training longer does not mean that it will learn more, you probably need more regulization in your pipeline. Maybe try resnets or efficientnets, they are usually the way to go</p>",
      "votes": null,
      "replies": [
        {
          "id": 1116997,
          "author_name": "docphilipp",
          "author_url": "",
          "post_date": "12/17/2020 16:19:05",
          "content": "<p>I thought that since resnets achieved lower accuracy on imagined that xception-nets would perform better</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1117022,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "12/17/2020 16:45:21",
          "content": "<p>Yes it might be true but different architechture can work better on different datasets. There are newer versions of resnets like resnext,seresnet,seresnext,wideresnet,resnest and more. These are probably faster and better than xception. Also, with newer augmentation techniques some architechtures are better than the others. I would recommend to always try resnets first, in my experience they usually give good results</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1117038,
          "author_name": "docphilipp",
          "author_url": "",
          "post_date": "12/17/2020 17:07:09",
          "content": "<p>ok thank you for the detailed reply 🙌🏻</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1117135,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "12/17/2020 18:44:53",
          "content": "<p>Also some others tips i can give you: <br>\n-Start with a lower resolution with samller architechtures and no augmentations to have a baseline. <br>\n-Find some good parameters with this set up<br>\n-Add some augmentations and find the best ones<br>\n-Tune some hyperparameters like epochs,weight decay, dropout, batchsize etc<br>\n-Then increase resolution and retune a bit<br>\n-Then increase architechture size</p>\n<p>This way you iterate faster and find better augs that work for this dataset. Increasing the depths of you network will only increase by a couple of % the performance, your pipeline is much more important.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1117869,
          "author_name": "docphilipp",
          "author_url": "",
          "post_date": "12/18/2020 14:39:18",
          "content": "<p>how can I start with a lower resolution? Do you mean cropping the down the image in size? (say instead of [512,512,3] cropping to [300,300,3]) or how do you lower the resolution?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1117882,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "12/18/2020 14:52:01",
          "content": "<p>In your image augmentation you would resize it to a desired resolution, maybe start with 256x256 or 384x384. Make sure to resize for training images and validation images!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1118147,
          "author_name": "docphilipp",
          "author_url": "",
          "post_date": "12/18/2020 19:16:20",
          "content": "<p>alright I'll check it out.. thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1116042": "when I call fit() on my model it starts training. In the beginning of the training its already at a pretty hight accuracy around 0.81. But after that it doesn't really improve and stays around 0.82 accuracy after 50 epochs.\nShouldn't it improve further during training?",
    "1116125": "Is that the accuracy on the training data or a validation set that you have not used for training? The former can certainly (unless you regularize) improve up to 1.0, for the latter there is going to be some limit that depends on various things (e.g. type of model - see e.g. [this thread](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/203111) what validation accuracy people are seeing with single models).  It's plausible that some set-up might just max out at 0.82 (either too small a model, not using pretrained weights, not a good set of image augmentations, not a good learning rate schedule etc.).",
    "1116692": "According to my limited knowledge, You need to go deeper with the model",
    "1116985": "I'm using a pertained Exception network I would say thats pretty deep already",
    "1116986": "I'm using a pertained Exception network and using different kinds of augmentation methods (cropping, flipping etc.). When I have my lr around 0.001 thats where I get the best results by far everything else is around 0.6.",
    "1116987": "Which model is it and input size?",
    "1116992": "Training longer does not mean that it will learn more, you probably need more regulization in your pipeline. Maybe try resnets or efficientnets, they are usually the way to go",
    "1116996": "*an pretrained Xception network + global average + dense + output and the input size = [512,512,3]",
    "1116997": "I thought that since resnets achieved lower accuracy on imagined that xception-nets would perform better",
    "1117022": "Yes it might be true but different architechture can work better on different datasets. There are newer versions of resnets like resnext,seresnet,seresnext,wideresnet,resnest and more. These are probably faster and better than xception. Also, with newer augmentation techniques some architechtures are better than the others. I would recommend to always try resnets first, in my experience they usually give good results",
    "1117038": "ok thank you for the detailed reply 🙌🏻",
    "1117135": "Also some others tips i can give you: \n-Start with a lower resolution with samller architechtures and no augmentations to have a baseline. \n-Find some good parameters with this set up\n-Add some augmentations and find the best ones\n-Tune some hyperparameters like epochs,weight decay, dropout, batchsize etc\n-Then increase resolution and retune a bit\n-Then increase architechture size\n\nThis way you iterate faster and find better augs that work for this dataset. Increasing the depths of you network will only increase by a couple of % the performance, your pipeline is much more important.",
    "1117869": "how can I start with a lower resolution? Do you mean cropping the down the image in size? (say instead of [512,512,3] cropping to [300,300,3]) or how do you lower the resolution?",
    "1117882": "In your image augmentation you would resize it to a desired resolution, maybe start with 256x256 or 384x384. Make sure to resize for training images and validation images!",
    "1118147": "alright I'll check it out.. thanks"
  },
  "source": "meta"
}