{
  "id": 546556,
  "title": "Help with staying in GPU memory during hyperparameter optimization",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/546556",
  "author_name": "",
  "post_date": "2024-11-16T15:15:15.510208400Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hey everyone,</p>\n<p>One of my goals of this competition is to learn how to iterate quicker and more effectively. One place I often spend a lot of time is fiddling with batch sizes and the number of filters for different architechtures to make my models fit in GPU memory.</p>\n<p>At home I have an RTX 3060 with 12 GB of VRAM and I also use the Kaggle GPUs (which I'm really grateful they provide).</p>\n<p>I'd like to do more Bayesian Optimization of the hyperparameters and model depth/width, but it's not always clear to me what the max number of filters will be or what kinds of batch sizes I can fit on my GPU without a lot of trial and error prior to initiating the optimization.</p>\n<p>Does anyone have advice on how to quickly determine the max size of a tensor I could pass into a specified model given a GPU's VRAM? Being able to do this quickly or in an automated way would help me figure out the range of hyperparameters I can work with in a given grid search or Bayesian Optimization run with the hardware I have available.</p>",
  "messages": [
    {
      "id": "3047335",
      "postDate": "11/16/2024 15:15:15",
      "content": "<p>Hey everyone,</p>\n<p>One of my goals of this competition is to learn how to iterate quicker and more effectively. One place I often spend a lot of time is fiddling with batch sizes and the number of filters for different architechtures to make my models fit in GPU memory.</p>\n<p>At home I have an RTX 3060 with 12 GB of VRAM and I also use the Kaggle GPUs (which I'm really grateful they provide).</p>\n<p>I'd like to do more Bayesian Optimization of the hyperparameters and model depth/width, but it's not always clear to me what the max number of filters will be or what kinds of batch sizes I can fit on my GPU without a lot of trial and error prior to initiating the optimization.</p>\n<p>Does anyone have advice on how to quickly determine the max size of a tensor I could pass into a specified model given a GPU's VRAM? Being able to do this quickly or in an automated way would help me figure out the range of hyperparameters I can work with in a given grid search or Bayesian Optimization run with the hardware I have available.</p>",
      "rawMarkdown": "Hey everyone,\n\nOne of my goals of this competition is to learn how to iterate quicker and more effectively. One place I often spend a lot of time is fiddling with batch sizes and the number of filters for different architechtures to make my models fit in GPU memory.\n\nAt home I have an RTX 3060 with 12 GB of VRAM and I also use the Kaggle GPUs (which I'm really grateful they provide).\n\nI'd like to do more Bayesian Optimization of the hyperparameters and model depth/width, but it's not always clear to me what the max number of filters will be or what kinds of batch sizes I can fit on my GPU without a lot of trial and error prior to initiating the optimization.\n\nDoes anyone have advice on how to quickly determine the max size of a tensor I could pass into a specified model given a GPU's VRAM? Being able to do this quickly or in an automated way would help me figure out the range of hyperparameters I can work with in a given grid search or Bayesian Optimization run with the hardware I have available.",
      "votes": null
    },
    {
      "id": "3049590",
      "postDate": "11/19/2024 09:54:29",
      "content": "<p>Hi, I am quite new and I cannot really help you with your problem. However, I do have some questions regarding your hyperparameter optimization and bayesian approach. What surrogate model are you using to update your beliefs on the new hyperparameters? </p>",
      "rawMarkdown": "Hi, I am quite new and I cannot really help you with your problem. However, I do have some questions regarding your hyperparameter optimization and bayesian approach. What surrogate model are you using to update your beliefs on the new hyperparameters?",
      "votes": null
    },
    {
      "id": "3052009",
      "postDate": "11/21/2024 22:49:32",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ignasialemany\" target=\"_blank\">@ignasialemany</a> , I'm using Optuna to tune the hyperparameters. It's a nice framework for doing this sort of thing :). <a href=\"https://optuna.readthedocs.io/en/stable/tutorial/10_key_features/003_efficient_optimization_algorithms.html\" target=\"_blank\">Optuna Docs</a> the docs at this link indicate that a \"Tree Structured Parzen Estimator\" is used as default, but there are other options.</p>",
      "rawMarkdown": "Hi @ignasialemany , I'm using Optuna to tune the hyperparameters. It's a nice framework for doing this sort of thing :). [Optuna Docs](https://optuna.readthedocs.io/en/stable/tutorial/10_key_features/003_efficient_optimization_algorithms.html) the docs at this link indicate that a \"Tree Structured Parzen Estimator\" is used as default, but there are other options.",
      "votes": null
    },
    {
      "id": "3052326",
      "postDate": "11/22/2024 09:34:02",
      "content": "<p><a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> thank you! I will have a look </p>",
      "rawMarkdown": "chemdatafarmer thank you! I will have a look",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3049590,
      "author_name": "ignasialemany",
      "author_url": "",
      "post_date": "11/19/2024 09:54:29",
      "content": "<p>Hi, I am quite new and I cannot really help you with your problem. However, I do have some questions regarding your hyperparameter optimization and bayesian approach. What surrogate model are you using to update your beliefs on the new hyperparameters? </p>",
      "votes": null,
      "replies": [
        {
          "id": 3052009,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "11/21/2024 22:49:32",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/ignasialemany\" target=\"_blank\">@ignasialemany</a> , I'm using Optuna to tune the hyperparameters. It's a nice framework for doing this sort of thing :). <a href=\"https://optuna.readthedocs.io/en/stable/tutorial/10_key_features/003_efficient_optimization_algorithms.html\" target=\"_blank\">Optuna Docs</a> the docs at this link indicate that a \"Tree Structured Parzen Estimator\" is used as default, but there are other options.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3052326,
              "author_name": "ignasialemany",
              "author_url": "",
              "post_date": "11/22/2024 09:34:02",
              "content": "<p><a href=\"https://www.kaggle.com/chemdatafarmer\" target=\"_blank\">@chemdatafarmer</a> thank you! I will have a look </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3047335": "Hey everyone,\n\nOne of my goals of this competition is to learn how to iterate quicker and more effectively. One place I often spend a lot of time is fiddling with batch sizes and the number of filters for different architechtures to make my models fit in GPU memory.\n\nAt home I have an RTX 3060 with 12 GB of VRAM and I also use the Kaggle GPUs (which I'm really grateful they provide).\n\nI'd like to do more Bayesian Optimization of the hyperparameters and model depth/width, but it's not always clear to me what the max number of filters will be or what kinds of batch sizes I can fit on my GPU without a lot of trial and error prior to initiating the optimization.\n\nDoes anyone have advice on how to quickly determine the max size of a tensor I could pass into a specified model given a GPU's VRAM? Being able to do this quickly or in an automated way would help me figure out the range of hyperparameters I can work with in a given grid search or Bayesian Optimization run with the hardware I have available.",
    "3049590": "Hi, I am quite new and I cannot really help you with your problem. However, I do have some questions regarding your hyperparameter optimization and bayesian approach. What surrogate model are you using to update your beliefs on the new hyperparameters?",
    "3052009": "Hi @ignasialemany , I'm using Optuna to tune the hyperparameters. It's a nice framework for doing this sort of thing :). [Optuna Docs](https://optuna.readthedocs.io/en/stable/tutorial/10_key_features/003_efficient_optimization_algorithms.html) the docs at this link indicate that a \"Tree Structured Parzen Estimator\" is used as default, but there are other options.",
    "3052326": "chemdatafarmer thank you! I will have a look"
  },
  "source": "meta"
}