{
  "id": 100198,
  "title": "Hyper-parameters tuning ",
  "url": "/competitions/aptos2019-blindness-detection/discussion/100198",
  "author_name": "",
  "post_date": "2019-07-17T06:24:31.088564200Z",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi \nHow do you tune your hyper-parameters?\nFirst, I assume that the main hyper-parameters to tune is the learning rate, The weight decay, and maybe the batch size. Perhaps we can add the image size as a hyper-parameters. \nI did some reading, and there are  three options for this task  : \n  1.  Grid Search \n  2. Random Search \n  3. Bayesian optimization </p>\n\n<p>Grid search is costly as you need to train multiple models, and if you are like me and limited with the GPU resources, this is probably not practical.\nRandom search is what I am doing now, but  I am doing it in a less arrange manner (I would like to improve it and learn a bit more).\nBayesian optimization  - The idea is that the hyper parameters and the cost or objective function (that the Deep learning network tries to optimize) can comply to a certain probability or distribution. Using the data, you have already collected and the Bayesian equation you can tune the hyper-parameters. You can refine the selection on every time you collect more data .\nHere is a tool that can use this approach \n<a href=\"https://ax.dev/versions/latest/tutorials/tune_cnn.html\">https://ax.dev/versions/latest/tutorials/tune_cnn.html</a> </p>\n\n<p>So what do you do for this task ? is it helpful to increase your LB or is it marginal ?</p>",
  "messages": [
    {
      "id": "577858",
      "postDate": "07/17/2019 06:24:31",
      "content": "<p>Hi \nHow do you tune your hyper-parameters?\nFirst, I assume that the main hyper-parameters to tune is the learning rate, The weight decay, and maybe the batch size. Perhaps we can add the image size as a hyper-parameters. \nI did some reading, and there are  three options for this task  : \n  1.  Grid Search \n  2. Random Search \n  3. Bayesian optimization </p>\n\n<p>Grid search is costly as you need to train multiple models, and if you are like me and limited with the GPU resources, this is probably not practical.\nRandom search is what I am doing now, but  I am doing it in a less arrange manner (I would like to improve it and learn a bit more).\nBayesian optimization  - The idea is that the hyper parameters and the cost or objective function (that the Deep learning network tries to optimize) can comply to a certain probability or distribution. Using the data, you have already collected and the Bayesian equation you can tune the hyper-parameters. You can refine the selection on every time you collect more data .\nHere is a tool that can use this approach \n<a href=\"https://ax.dev/versions/latest/tutorials/tune_cnn.html\">https://ax.dev/versions/latest/tutorials/tune_cnn.html</a> </p>\n\n<p>So what do you do for this task ? is it helpful to increase your LB or is it marginal ?</p>",
      "rawMarkdown": "Hi \nHow do you tune your hyper-parameters?\nFirst, I assume that the main hyper-parameters to tune is the learning rate, The weight decay, and maybe the batch size. Perhaps we can add the image size as a hyper-parameters. \nI did some reading, and there are  three options for this task  : \n  1.  Grid Search \n  2. Random Search \n  3. Bayesian optimization \n\nGrid search is costly as you need to train multiple models, and if you are like me and limited with the GPU resources, this is probably not practical.\nRandom search is what I am doing now, but  I am doing it in a less arrange manner (I would like to improve it and learn a bit more).\nBayesian optimization  - The idea is that the hyper parameters and the cost or objective function (that the Deep learning network tries to optimize) can comply to a certain probability or distribution. Using the data, you have already collected and the Bayesian equation you can tune the hyper-parameters. You can refine the selection on every time you collect more data .\nHere is a tool that can use this approach \nhttps://ax.dev/versions/latest/tutorials/tune_cnn.html \n\nSo what do you do for this task ? is it helpful to increase your LB or is it marginal ?",
      "votes": null
    },
    {
      "id": "583121",
      "postDate": "07/24/2019 04:10:56",
      "content": "<p>Hi, some things I may contribute:\n1. For learning rate tuning, I suggest playing around with cyclic learning rate and check-pointing best weights, and also adaptive optimizers e.g. adam, etc\n2. batchsize is a \"hyper-parameter\", but since it is dependent on model/GPU, do not care about it too much (though larger batch may help in case of noisy data). The discussion here so far seems to indicate that focusing on larger image size may be beneficial.\n3. Since we are quite early on the competition, it may be more worth your while to experiment different architectures as a way to spot-check the one you would like to stick with and further fine-tune via approaches you mentioned above. A lot of fine-tuning, ensembling, post-processing people get into a lot when the end is near. </p>\n\n<p>Not much else I could say, hope this helps.</p>",
      "rawMarkdown": "Hi, some things I may contribute:\n1. For learning rate tuning, I suggest playing around with cyclic learning rate and check-pointing best weights, and also adaptive optimizers e.g. adam, etc\n2. batchsize is a \"hyper-parameter\", but since it is dependent on model/GPU, do not care about it too much (though larger batch may help in case of noisy data). The discussion here so far seems to indicate that focusing on larger image size may be beneficial.\n3. Since we are quite early on the competition, it may be more worth your while to experiment different architectures as a way to spot-check the one you would like to stick with and further fine-tune via approaches you mentioned above. A lot of fine-tuning, ensembling, post-processing people get into a lot when the end is near. \n\nNot much else I could say, hope this helps.",
      "votes": null
    },
    {
      "id": "585932",
      "postDate": "07/28/2019 11:08:49",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 583121,
      "author_name": "joonl04",
      "author_url": "",
      "post_date": "07/24/2019 04:10:56",
      "content": "<p>Hi, some things I may contribute:\n1. For learning rate tuning, I suggest playing around with cyclic learning rate and check-pointing best weights, and also adaptive optimizers e.g. adam, etc\n2. batchsize is a \"hyper-parameter\", but since it is dependent on model/GPU, do not care about it too much (though larger batch may help in case of noisy data). The discussion here so far seems to indicate that focusing on larger image size may be beneficial.\n3. Since we are quite early on the competition, it may be more worth your while to experiment different architectures as a way to spot-check the one you would like to stick with and further fine-tune via approaches you mentioned above. A lot of fine-tuning, ensembling, post-processing people get into a lot when the end is near. </p>\n\n<p>Not much else I could say, hope this helps.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 585932,
      "author_name": "omershect",
      "author_url": "",
      "post_date": "07/28/2019 11:08:49",
      "content": "<p>Thanks</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "577858": "Hi \nHow do you tune your hyper-parameters?\nFirst, I assume that the main hyper-parameters to tune is the learning rate, The weight decay, and maybe the batch size. Perhaps we can add the image size as a hyper-parameters. \nI did some reading, and there are  three options for this task  : \n  1.  Grid Search \n  2. Random Search \n  3. Bayesian optimization \n\nGrid search is costly as you need to train multiple models, and if you are like me and limited with the GPU resources, this is probably not practical.\nRandom search is what I am doing now, but  I am doing it in a less arrange manner (I would like to improve it and learn a bit more).\nBayesian optimization  - The idea is that the hyper parameters and the cost or objective function (that the Deep learning network tries to optimize) can comply to a certain probability or distribution. Using the data, you have already collected and the Bayesian equation you can tune the hyper-parameters. You can refine the selection on every time you collect more data .\nHere is a tool that can use this approach \nhttps://ax.dev/versions/latest/tutorials/tune_cnn.html \n\nSo what do you do for this task ? is it helpful to increase your LB or is it marginal ?",
    "583121": "Hi, some things I may contribute:\n1. For learning rate tuning, I suggest playing around with cyclic learning rate and check-pointing best weights, and also adaptive optimizers e.g. adam, etc\n2. batchsize is a \"hyper-parameter\", but since it is dependent on model/GPU, do not care about it too much (though larger batch may help in case of noisy data). The discussion here so far seems to indicate that focusing on larger image size may be beneficial.\n3. Since we are quite early on the competition, it may be more worth your while to experiment different architectures as a way to spot-check the one you would like to stick with and further fine-tune via approaches you mentioned above. A lot of fine-tuning, ensembling, post-processing people get into a lot when the end is near. \n\nNot much else I could say, hope this helps.",
    "585932": "Thanks"
  },
  "source": "meta"
}