{
  "id": 482140,
  "title": "Beginner question：How to optimize DeepLearning model parameters",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/482140",
  "author_name": "",
  "post_date": "2024-03-06T14:00:16.787727200Z",
  "votes": 10,
  "comment_count": 10,
  "views": 0,
  "content": "<p>As a beginner, I am eager to learn how pros optimize their parameters. Are there any specific strategies or is it a matter of trial and error? I would greatly appreciate it if you could share any valuable resources or reading materials. 😃</p>",
  "messages": [
    {
      "id": "2684245",
      "postDate": "03/06/2024 14:00:16",
      "content": "<p>As a beginner, I am eager to learn how pros optimize their parameters. Are there any specific strategies or is it a matter of trial and error? I would greatly appreciate it if you could share any valuable resources or reading materials. 😃</p>",
      "rawMarkdown": "As a beginner, I am eager to learn how pros optimize their parameters. Are there any specific strategies or is it a matter of trial and error? I would greatly appreciate it if you could share any valuable resources or reading materials. 😃",
      "votes": null
    },
    {
      "id": "2684306",
      "postDate": "03/06/2024 14:44:37",
      "content": "<p>I start with a fixed learning rate and try values like 1e-3, 5e-4, 1e-4, etc and then get familiar with the loss landscape of each fold i.e. seeing how tweaking changes losses and convergence. It's basically trial and error but with an intuition that you build over time. </p>",
      "rawMarkdown": "I start with a fixed learning rate and try values like 1e-3, 5e-4, 1e-4, etc and then get familiar with the loss landscape of each fold i.e. seeing how tweaking changes losses and convergence. It's basically trial and error but with an intuition that you build over time.",
      "votes": null
    },
    {
      "id": "2685093",
      "postDate": "03/07/2024 01:23:28",
      "content": "<p>Thank you for you reply！</p>",
      "rawMarkdown": "Thank you for you reply！",
      "votes": null
    },
    {
      "id": "2685250",
      "postDate": "03/07/2024 05:28:37",
      "content": "<p>I think that people who share the training notebook do a lot of trial and errors in their local machine.<br>\nKaggle notebook is not suitable for parameter tuning because of not enough gpu time.<br>\nTo run a lot of experiments efficiently, setting hyper parameters only by config file like yaml may be important</p>\n<p>There is also an official guide written by google researchers.<br>\n<a href=\"https://github.com/google-research/tuning_playbook\" target=\"_blank\">https://github.com/google-research/tuning_playbook</a></p>",
      "rawMarkdown": "I think that people who share the training notebook do a lot of trial and errors in their local machine.\nKaggle notebook is not suitable for parameter tuning because of not enough gpu time.\nTo run a lot of experiments efficiently, setting hyper parameters only by config file like yaml may be important\n\nThere is also an official guide written by google researchers.\n[https://github.com/google-research/tuning_playbook](https://github.com/google-research/tuning_playbook)",
      "votes": null
    },
    {
      "id": "2685636",
      "postDate": "03/07/2024 10:09:55",
      "content": "<p>I will read this， thank you！</p>",
      "rawMarkdown": "I will read this， thank you！",
      "votes": null
    },
    {
      "id": "2685956",
      "postDate": "03/07/2024 14:40:22",
      "content": "<p>In addition to to what others say, it is also important to make your experiment pipeline as fast as possible. This includes fast preprocess, fast dataloader, fast training, fast inference, fast metric computation. Additionally it helps if we can experiment successfully with subsets of entire data and/or small models and/or short train schedules.</p>\n<p>Utilizing multiple GPU and removing all bottle necks is crucial. We want to minimize disk reading, minimize data dtypes, remove unnecessary features and processes, and avoid performing the exact same process procedure more than once. Everything needs to be optimized and accelerated.</p>\n<p>After this is established we then begin the art of experimentation. Experimentation is certainly an art form. We are guided by experience, intuition, and the knowledge of how things work. It is important to reflect after each experiment to understand the models and data better. We constantly redefine plans for future experiments based on previous results and gained understanding. </p>\n<p>Finally, remember to enjoy the journey. Experimentation is a fun way to learn data science. We learn how models work and behave and we learn about data. Experimentation is also exciting, waking up to see new results, and rewards of discoveries!</p>",
      "rawMarkdown": "In addition to to what others say, it is also important to make your experiment pipeline as fast as possible. This includes fast preprocess, fast dataloader, fast training, fast inference, fast metric computation. Additionally it helps if we can experiment successfully with subsets of entire data and/or small models and/or short train schedules.\n\nUtilizing multiple GPU and removing all bottle necks is crucial. We want to minimize disk reading, minimize data dtypes, remove unnecessary features and processes, and avoid performing the exact same process procedure more than once. Everything needs to be optimized and accelerated.\n\nAfter this is established we then begin the art of experimentation. Experimentation is certainly an art form. We are guided by experience, intuition, and the knowledge of how things work. It is important to reflect after each experiment to understand the models and data better. We constantly redefine plans for future experiments based on previous results and gained understanding. \n\nFinally, remember to enjoy the journey. Experimentation is a fun way to learn data science. We learn how models work and behave and we learn about data. Experimentation is also exciting, waking up to see new results, and rewards of discoveries!",
      "votes": null
    },
    {
      "id": "2686649",
      "postDate": "03/08/2024 00:58:34",
      "content": "<p>I'm very glad to see your reply! I have a question about how to accelerate these processes and remove unnecessary features in the code, how can I implement that? Every time I participate in a competition, I normalize the data and directly input it, but this seems like a significant issue.🤣</p>",
      "rawMarkdown": "I'm very glad to see your reply! I have a question about how to accelerate these processes and remove unnecessary features in the code, how can I implement that? Every time I participate in a competition, I normalize the data and directly input it, but this seems like a significant issue.🤣",
      "votes": null
    },
    {
      "id": "2686657",
      "postDate": "03/08/2024 01:20:02",
      "content": "<p>Here are a three specific tips:<br>\n<strong>Remove Features</strong><br>\nWhen i say \"remove features\", i'm referring to input data. It is not always the case that using all the provided data is best. For example, <code>train.csv</code> has 100k rows and EEG parquets have 20 columns and Spectrogram parquets have 400 columns. (That's approx 40 million elements!) Maybe we can build a great model using only 10k rows and a subset of columns. If we do this, then we speed up experiments by 10x or more! After we tune our model, we can then add more features and more data (but experimentation is best using small stuff).</p>\n<p><strong>Accelerate</strong><br>\nOne of the most overlooked speed components is preprocess and dataloader. Try running your dataloader without training a model. For example, time how long does the following take:</p>\n<pre><code>%%\n batch  loader:\n    pass\n</code></pre>\n<p>This should be lightning fast. If it is not, find the bottlenecks (disk reading, repeated operations, inefficient operations, etc) and remove them. Also utilize CPU/GPU multiprocessing in dataloader. When we train a model on GPU we want to feed large batch sizes the fastest we can because GPUs want lots of data quickly. If we run <code>nvidia-smi</code> we should see <code>100%</code> GPU utilization. Anything less indicates a dataloader bottleneck and unnecessary slow down.</p>\n<p><strong>Training</strong><br>\nUse mixed precision, multiple GPU, small models (at first), efficiently designed short train schedules (i.e. the largest possible learning rates without overfitting instead of many epochs of low learning rates), and other tricks to accelerate model training like freezing layers, PEFT, loading pretrained models, using diverse subsets of train data (instead of all train data)</p>",
      "rawMarkdown": "Here are a three specific tips:\n**Remove Features**\nWhen i say \"remove features\", i'm referring to input data. It is not always the case that using all the provided data is best. For example, `train.csv` has 100k rows and EEG parquets have 20 columns and Spectrogram parquets have 400 columns. (That's approx 40 million elements!) Maybe we can build a great model using only 10k rows and a subset of columns. If we do this, then we speed up experiments by 10x or more! After we tune our model, we can then add more features and more data (but experimentation is best using small stuff).\n\n**Accelerate**\nOne of the most overlooked speed components is preprocess and dataloader. Try running your dataloader without training a model. For example, time how long does the following take:\n\n    %%time\n    for batch in loader:\n        pass\n\nThis should be lightning fast. If it is not, find the bottlenecks (disk reading, repeated operations, inefficient operations, etc) and remove them. Also utilize CPU/GPU multiprocessing in dataloader. When we train a model on GPU we want to feed large batch sizes the fastest we can because GPUs want lots of data quickly. If we run `nvidia-smi` we should see `100%` GPU utilization. Anything less indicates a dataloader bottleneck and unnecessary slow down.\n\n**Training**\nUse mixed precision, multiple GPU, small models (at first), efficiently designed short train schedules (i.e. the largest possible learning rates without overfitting instead of many epochs of low learning rates), and other tricks to accelerate model training like freezing layers, PEFT, loading pretrained models, using diverse subsets of train data (instead of all train data)",
      "votes": null
    },
    {
      "id": "2689791",
      "postDate": "03/10/2024 05:55:10",
      "content": "<p>This is just information from another world. You are truly a genius. Thank you for sharing. You are an inspiration. </p>",
      "rawMarkdown": "This is just information from another world. You are truly a genius. Thank you for sharing. You are an inspiration.",
      "votes": null
    },
    {
      "id": "2689889",
      "postDate": "03/10/2024 07:24:23",
      "content": "<p>Thank you for your instructions！</p>",
      "rawMarkdown": "Thank you for your instructions！",
      "votes": null
    },
    {
      "id": "2704359",
      "postDate": "03/18/2024 18:04:57",
      "content": "<p>Thank for sharing Chris! </p>",
      "rawMarkdown": "Thank for sharing Chris!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2684306,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "03/06/2024 14:44:37",
      "content": "<p>I start with a fixed learning rate and try values like 1e-3, 5e-4, 1e-4, etc and then get familiar with the loss landscape of each fold i.e. seeing how tweaking changes losses and convergence. It's basically trial and error but with an intuition that you build over time. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2685093,
          "author_name": "seeingtimes",
          "author_url": "",
          "post_date": "03/07/2024 01:23:28",
          "content": "<p>Thank you for you reply！</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2685250,
      "author_name": "clearwaterkzk",
      "author_url": "",
      "post_date": "03/07/2024 05:28:37",
      "content": "<p>I think that people who share the training notebook do a lot of trial and errors in their local machine.<br>\nKaggle notebook is not suitable for parameter tuning because of not enough gpu time.<br>\nTo run a lot of experiments efficiently, setting hyper parameters only by config file like yaml may be important</p>\n<p>There is also an official guide written by google researchers.<br>\n<a href=\"https://github.com/google-research/tuning_playbook\" target=\"_blank\">https://github.com/google-research/tuning_playbook</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2685636,
          "author_name": "seeingtimes",
          "author_url": "",
          "post_date": "03/07/2024 10:09:55",
          "content": "<p>I will read this， thank you！</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2685956,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/07/2024 14:40:22",
      "content": "<p>In addition to to what others say, it is also important to make your experiment pipeline as fast as possible. This includes fast preprocess, fast dataloader, fast training, fast inference, fast metric computation. Additionally it helps if we can experiment successfully with subsets of entire data and/or small models and/or short train schedules.</p>\n<p>Utilizing multiple GPU and removing all bottle necks is crucial. We want to minimize disk reading, minimize data dtypes, remove unnecessary features and processes, and avoid performing the exact same process procedure more than once. Everything needs to be optimized and accelerated.</p>\n<p>After this is established we then begin the art of experimentation. Experimentation is certainly an art form. We are guided by experience, intuition, and the knowledge of how things work. It is important to reflect after each experiment to understand the models and data better. We constantly redefine plans for future experiments based on previous results and gained understanding. </p>\n<p>Finally, remember to enjoy the journey. Experimentation is a fun way to learn data science. We learn how models work and behave and we learn about data. Experimentation is also exciting, waking up to see new results, and rewards of discoveries!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2686649,
          "author_name": "seeingtimes",
          "author_url": "",
          "post_date": "03/08/2024 00:58:34",
          "content": "<p>I'm very glad to see your reply! I have a question about how to accelerate these processes and remove unnecessary features in the code, how can I implement that? Every time I participate in a competition, I normalize the data and directly input it, but this seems like a significant issue.🤣</p>",
          "votes": null,
          "replies": [
            {
              "id": 2686657,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "03/08/2024 01:20:02",
              "content": "<p>Here are a three specific tips:<br>\n<strong>Remove Features</strong><br>\nWhen i say \"remove features\", i'm referring to input data. It is not always the case that using all the provided data is best. For example, <code>train.csv</code> has 100k rows and EEG parquets have 20 columns and Spectrogram parquets have 400 columns. (That's approx 40 million elements!) Maybe we can build a great model using only 10k rows and a subset of columns. If we do this, then we speed up experiments by 10x or more! After we tune our model, we can then add more features and more data (but experimentation is best using small stuff).</p>\n<p><strong>Accelerate</strong><br>\nOne of the most overlooked speed components is preprocess and dataloader. Try running your dataloader without training a model. For example, time how long does the following take:</p>\n<pre><code>%%\n batch  loader:\n    pass\n</code></pre>\n<p>This should be lightning fast. If it is not, find the bottlenecks (disk reading, repeated operations, inefficient operations, etc) and remove them. Also utilize CPU/GPU multiprocessing in dataloader. When we train a model on GPU we want to feed large batch sizes the fastest we can because GPUs want lots of data quickly. If we run <code>nvidia-smi</code> we should see <code>100%</code> GPU utilization. Anything less indicates a dataloader bottleneck and unnecessary slow down.</p>\n<p><strong>Training</strong><br>\nUse mixed precision, multiple GPU, small models (at first), efficiently designed short train schedules (i.e. the largest possible learning rates without overfitting instead of many epochs of low learning rates), and other tricks to accelerate model training like freezing layers, PEFT, loading pretrained models, using diverse subsets of train data (instead of all train data)</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2689791,
                  "author_name": "devbilalkhan",
                  "author_url": "",
                  "post_date": "03/10/2024 05:55:10",
                  "content": "<p>This is just information from another world. You are truly a genius. Thank you for sharing. You are an inspiration. </p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2689889,
                  "author_name": "seeingtimes",
                  "author_url": "",
                  "post_date": "03/10/2024 07:24:23",
                  "content": "<p>Thank you for your instructions！</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2704359,
                  "author_name": "arindamroy23",
                  "author_url": "",
                  "post_date": "03/18/2024 18:04:57",
                  "content": "<p>Thank for sharing Chris! </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2684245": "As a beginner, I am eager to learn how pros optimize their parameters. Are there any specific strategies or is it a matter of trial and error? I would greatly appreciate it if you could share any valuable resources or reading materials. 😃",
    "2684306": "I start with a fixed learning rate and try values like 1e-3, 5e-4, 1e-4, etc and then get familiar with the loss landscape of each fold i.e. seeing how tweaking changes losses and convergence. It's basically trial and error but with an intuition that you build over time.",
    "2685093": "Thank you for you reply！",
    "2685250": "I think that people who share the training notebook do a lot of trial and errors in their local machine.\nKaggle notebook is not suitable for parameter tuning because of not enough gpu time.\nTo run a lot of experiments efficiently, setting hyper parameters only by config file like yaml may be important\n\nThere is also an official guide written by google researchers.\n[https://github.com/google-research/tuning_playbook](https://github.com/google-research/tuning_playbook)",
    "2685636": "I will read this， thank you！",
    "2685956": "In addition to to what others say, it is also important to make your experiment pipeline as fast as possible. This includes fast preprocess, fast dataloader, fast training, fast inference, fast metric computation. Additionally it helps if we can experiment successfully with subsets of entire data and/or small models and/or short train schedules.\n\nUtilizing multiple GPU and removing all bottle necks is crucial. We want to minimize disk reading, minimize data dtypes, remove unnecessary features and processes, and avoid performing the exact same process procedure more than once. Everything needs to be optimized and accelerated.\n\nAfter this is established we then begin the art of experimentation. Experimentation is certainly an art form. We are guided by experience, intuition, and the knowledge of how things work. It is important to reflect after each experiment to understand the models and data better. We constantly redefine plans for future experiments based on previous results and gained understanding. \n\nFinally, remember to enjoy the journey. Experimentation is a fun way to learn data science. We learn how models work and behave and we learn about data. Experimentation is also exciting, waking up to see new results, and rewards of discoveries!",
    "2686649": "I'm very glad to see your reply! I have a question about how to accelerate these processes and remove unnecessary features in the code, how can I implement that? Every time I participate in a competition, I normalize the data and directly input it, but this seems like a significant issue.🤣",
    "2686657": "Here are a three specific tips:\n**Remove Features**\nWhen i say \"remove features\", i'm referring to input data. It is not always the case that using all the provided data is best. For example, `train.csv` has 100k rows and EEG parquets have 20 columns and Spectrogram parquets have 400 columns. (That's approx 40 million elements!) Maybe we can build a great model using only 10k rows and a subset of columns. If we do this, then we speed up experiments by 10x or more! After we tune our model, we can then add more features and more data (but experimentation is best using small stuff).\n\n**Accelerate**\nOne of the most overlooked speed components is preprocess and dataloader. Try running your dataloader without training a model. For example, time how long does the following take:\n\n    %%time\n    for batch in loader:\n        pass\n\nThis should be lightning fast. If it is not, find the bottlenecks (disk reading, repeated operations, inefficient operations, etc) and remove them. Also utilize CPU/GPU multiprocessing in dataloader. When we train a model on GPU we want to feed large batch sizes the fastest we can because GPUs want lots of data quickly. If we run `nvidia-smi` we should see `100%` GPU utilization. Anything less indicates a dataloader bottleneck and unnecessary slow down.\n\n**Training**\nUse mixed precision, multiple GPU, small models (at first), efficiently designed short train schedules (i.e. the largest possible learning rates without overfitting instead of many epochs of low learning rates), and other tricks to accelerate model training like freezing layers, PEFT, loading pretrained models, using diverse subsets of train data (instead of all train data)",
    "2689791": "This is just information from another world. You are truly a genius. Thank you for sharing. You are an inspiration.",
    "2689889": "Thank you for your instructions！",
    "2704359": "Thank for sharing Chris!"
  },
  "source": "meta"
}