{
  "id": 157383,
  "title": "[Advice Needed] Hyperparameter/Architecture tuning",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/157383",
  "author_name": "",
  "post_date": "2020-06-10T13:15:25.261301600Z",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello everyone,\nI was wondering if anyone can guide me to some decent resources or give some advice on hyperparameter tuning or network tuning in general.\nIt seems like the architecture changes that I make while making the validation accuracy drop, does not give a better score on the leaderboard. Also, I am using a fixed seed but the results seem to vary a lot which I am unsure how to tackle. \nI apologize if this post doesn't belong here but I am relatively new and any help would be much appreciated!</p>",
  "messages": [
    {
      "id": "880646",
      "postDate": "06/10/2020 13:15:25",
      "content": "<p>Hello everyone,\nI was wondering if anyone can guide me to some decent resources or give some advice on hyperparameter tuning or network tuning in general.\nIt seems like the architecture changes that I make while making the validation accuracy drop, does not give a better score on the leaderboard. Also, I am using a fixed seed but the results seem to vary a lot which I am unsure how to tackle. \nI apologize if this post doesn't belong here but I am relatively new and any help would be much appreciated!</p>",
      "rawMarkdown": "Hello everyone,\nI was wondering if anyone can guide me to some decent resources or give some advice on hyperparameter tuning or network tuning in general.\nIt seems like the architecture changes that I make while making the validation accuracy drop, does not give a better score on the leaderboard. Also, I am using a fixed seed but the results seem to vary a lot which I am unsure how to tackle. \nI apologize if this post doesn't belong here but I am relatively new and any help would be much appreciated!",
      "votes": null
    },
    {
      "id": "880769",
      "postDate": "06/10/2020 14:56:30",
      "content": "<p>just a few thoughts:\n- most times tuning the learing rate is very important. When you only have a limited time budget, start with tuning the learing rate.\n- implement a performant validation procedure. The faster the better. When you can make more round trips per time, you can search more paramters/architectures. Also validating only by submitting to the leaderboard limits you to only 5 evaluations per day. You want a stable and fast cross validation using only the train data.\n- You can use different approaches for selection of hyperparameters, grid search, random search.\n- If your search space is huge, i.e. you want to search on many different combinations, you will have to use a heuristic like genetic algorithms, sinnflood search, simulated annealing\n- regarding the architecture of neural networks, there are a lot of interesting papers and concepts. Have a look at\n\"Learning Transferable Architectures for Scalable Image Recognition\" -&gt; <a href=\"https://arxiv.org/pdf/1707.07012.pdf\">https://arxiv.org/pdf/1707.07012.pdf</a>\nor \"Designing Network Design Spaces\" -&gt; <a href=\"https://arxiv.org/pdf/2003.13678.pdf\">https://arxiv.org/pdf/2003.13678.pdf</a></p>",
      "rawMarkdown": "just a few thoughts:\n- most times tuning the learing rate is very important. When you only have a limited time budget, start with tuning the learing rate.\n- implement a performant validation procedure. The faster the better. When you can make more round trips per time, you can search more paramters/architectures. Also validating only by submitting to the leaderboard limits you to only 5 evaluations per day. You want a stable and fast cross validation using only the train data.\n- You can use different approaches for selection of hyperparameters, grid search, random search.\n- If your search space is huge, i.e. you want to search on many different combinations, you will have to use a heuristic like genetic algorithms, sinnflood search, simulated annealing\n- regarding the architecture of neural networks, there are a lot of interesting papers and concepts. Have a look at\n\"Learning Transferable Architectures for Scalable Image Recognition\" -&gt; https://arxiv.org/pdf/1707.07012.pdf\nor \"Designing Network Design Spaces\" -&gt; https://arxiv.org/pdf/2003.13678.pdf",
      "votes": null
    },
    {
      "id": "881615",
      "postDate": "06/11/2020 07:47:46",
      "content": "<p>Super cool tips!</p>\n\n<p>Just some random comments:</p>\n\n<ul>\n<li>Learning rate might turn on to be less important if you use adaptive solvers (like Adam). However, spending some time playing with it and also with learning rate schedules might be a good investment</li>\n<li>Regarding the validation, this competition is really tricky in this regard. Scores vary crazily across folds even when stratification is used, so you have to come up with something smarter than just the default KFold scheme. One possible way is to sacrifice a part of the competition data for validation while using external data for training. This will likely make your validation score lower, but you would observe less variance between models/runs.</li>\n<li>If you have access to decent compute (for this competition, I would call 4-8 high-end GPUs decent enough), you could probably afford some grid search. Otherwise, tune only the essential parameters. Which parameters are essential is a good question and it's actually your task as a developer to identify those.</li>\n<li>Regarding the architecture search, I wouldn't expect much gains from it. Basically everyone uses already existing networks pretrained on ImageNet and you can't tweak them. For the parameters you can tune (number of layers in the head, learning rate, number of epochs), I would expect regular random/grid search to work just fine.</li>\n</ul>",
      "rawMarkdown": "Super cool tips!\n\nJust some random comments:\n\n* Learning rate might turn on to be less important if you use adaptive solvers (like Adam). However, spending some time playing with it and also with learning rate schedules might be a good investment\n* Regarding the validation, this competition is really tricky in this regard. Scores vary crazily across folds even when stratification is used, so you have to come up with something smarter than just the default KFold scheme. One possible way is to sacrifice a part of the competition data for validation while using external data for training. This will likely make your validation score lower, but you would observe less variance between models/runs.\n* If you have access to decent compute (for this competition, I would call 4-8 high-end GPUs decent enough), you could probably afford some grid search. Otherwise, tune only the essential parameters. Which parameters are essential is a good question and it's actually your task as a developer to identify those.\n* Regarding the architecture search, I wouldn't expect much gains from it. Basically everyone uses already existing networks pretrained on ImageNet and you can't tweak them. For the parameters you can tune (number of layers in the head, learning rate, number of epochs), I would expect regular random/grid search to work just fine.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 880769,
      "author_name": "agentauers",
      "author_url": "",
      "post_date": "06/10/2020 14:56:30",
      "content": "<p>just a few thoughts:\n- most times tuning the learing rate is very important. When you only have a limited time budget, start with tuning the learing rate.\n- implement a performant validation procedure. The faster the better. When you can make more round trips per time, you can search more paramters/architectures. Also validating only by submitting to the leaderboard limits you to only 5 evaluations per day. You want a stable and fast cross validation using only the train data.\n- You can use different approaches for selection of hyperparameters, grid search, random search.\n- If your search space is huge, i.e. you want to search on many different combinations, you will have to use a heuristic like genetic algorithms, sinnflood search, simulated annealing\n- regarding the architecture of neural networks, there are a lot of interesting papers and concepts. Have a look at\n\"Learning Transferable Architectures for Scalable Image Recognition\" -&gt; <a href=\"https://arxiv.org/pdf/1707.07012.pdf\">https://arxiv.org/pdf/1707.07012.pdf</a>\nor \"Designing Network Design Spaces\" -&gt; <a href=\"https://arxiv.org/pdf/2003.13678.pdf\">https://arxiv.org/pdf/2003.13678.pdf</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 881615,
          "author_name": "ddanevskyi",
          "author_url": "",
          "post_date": "06/11/2020 07:47:46",
          "content": "<p>Super cool tips!</p>\n\n<p>Just some random comments:</p>\n\n<ul>\n<li>Learning rate might turn on to be less important if you use adaptive solvers (like Adam). However, spending some time playing with it and also with learning rate schedules might be a good investment</li>\n<li>Regarding the validation, this competition is really tricky in this regard. Scores vary crazily across folds even when stratification is used, so you have to come up with something smarter than just the default KFold scheme. One possible way is to sacrifice a part of the competition data for validation while using external data for training. This will likely make your validation score lower, but you would observe less variance between models/runs.</li>\n<li>If you have access to decent compute (for this competition, I would call 4-8 high-end GPUs decent enough), you could probably afford some grid search. Otherwise, tune only the essential parameters. Which parameters are essential is a good question and it's actually your task as a developer to identify those.</li>\n<li>Regarding the architecture search, I wouldn't expect much gains from it. Basically everyone uses already existing networks pretrained on ImageNet and you can't tweak them. For the parameters you can tune (number of layers in the head, learning rate, number of epochs), I would expect regular random/grid search to work just fine.</li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "880646": "Hello everyone,\nI was wondering if anyone can guide me to some decent resources or give some advice on hyperparameter tuning or network tuning in general.\nIt seems like the architecture changes that I make while making the validation accuracy drop, does not give a better score on the leaderboard. Also, I am using a fixed seed but the results seem to vary a lot which I am unsure how to tackle. \nI apologize if this post doesn't belong here but I am relatively new and any help would be much appreciated!",
    "880769": "just a few thoughts:\n- most times tuning the learing rate is very important. When you only have a limited time budget, start with tuning the learing rate.\n- implement a performant validation procedure. The faster the better. When you can make more round trips per time, you can search more paramters/architectures. Also validating only by submitting to the leaderboard limits you to only 5 evaluations per day. You want a stable and fast cross validation using only the train data.\n- You can use different approaches for selection of hyperparameters, grid search, random search.\n- If your search space is huge, i.e. you want to search on many different combinations, you will have to use a heuristic like genetic algorithms, sinnflood search, simulated annealing\n- regarding the architecture of neural networks, there are a lot of interesting papers and concepts. Have a look at\n\"Learning Transferable Architectures for Scalable Image Recognition\" -&gt; https://arxiv.org/pdf/1707.07012.pdf\nor \"Designing Network Design Spaces\" -&gt; https://arxiv.org/pdf/2003.13678.pdf",
    "881615": "Super cool tips!\n\nJust some random comments:\n\n* Learning rate might turn on to be less important if you use adaptive solvers (like Adam). However, spending some time playing with it and also with learning rate schedules might be a good investment\n* Regarding the validation, this competition is really tricky in this regard. Scores vary crazily across folds even when stratification is used, so you have to come up with something smarter than just the default KFold scheme. One possible way is to sacrifice a part of the competition data for validation while using external data for training. This will likely make your validation score lower, but you would observe less variance between models/runs.\n* If you have access to decent compute (for this competition, I would call 4-8 high-end GPUs decent enough), you could probably afford some grid search. Otherwise, tune only the essential parameters. Which parameters are essential is a good question and it's actually your task as a developer to identify those.\n* Regarding the architecture search, I wouldn't expect much gains from it. Basically everyone uses already existing networks pretrained on ImageNet and you can't tweak them. For the parameters you can tune (number of layers in the head, learning rate, number of epochs), I would expect regular random/grid search to work just fine."
  },
  "source": "meta"
}