{
  "id": 346143,
  "title": "Ensembling Neural Networks",
  "url": "/competitions/amex-default-prediction/discussion/346143",
  "author_name": "Adam W",
  "post_date": "2022-08-18T05:44:13.004000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>Joined this competition recently and it seems like a fun project.</p>\n<p>As you probably know, most of the top solutions on Kaggle/modern data science competitions rely on ML ensembling methods. Personally I am very interested in deep learning due to its theoretical properties, but for it to be competitive with a well-designed ensemble on a single metric would require extensive hyperparameter tuning and architectural considerations.</p>\n<p>Anyways there is an interesting paper <a href=\"https://arxiv.org/pdf/1704.00109.pdf\" target=\"_blank\">Huang et al. (2017)</a> that proposes the use of cyclical learning rate schedules to generate an ensemble of trained neural networks while training only once. Loosely speaking, we can think of this as letting the neural network rapidly converge into a local minima, saving the model, resetting the learning rate, and repeat. We then take all of those saved models and use it as an ensemble without any additional computational cost. It's a very simple idea but it surprisingly works pretty well.</p>\n<p>If you're working with neural networks in TensorFlow Keras, I've created a small library <a href=\"https://github.com/adamvvu/snapshot_ensemble\" target=\"_blank\">here</a> that may be helpful. Let me know if you've gotten some interesting results with it!</p>",
  "messages": [
    {
      "id": 1904316,
      "postDate": "2022-08-18T05:44:13.003Z",
      "content": "<p>Hi everyone,</p>\n<p>Joined this competition recently and it seems like a fun project.</p>\n<p>As you probably know, most of the top solutions on Kaggle/modern data science competitions rely on ML ensembling methods. Personally I am very interested in deep learning due to its theoretical properties, but for it to be competitive with a well-designed ensemble on a single metric would require extensive hyperparameter tuning and architectural considerations.</p>\n<p>Anyways there is an interesting paper <a href=\"https://arxiv.org/pdf/1704.00109.pdf\" target=\"_blank\">Huang et al. (2017)</a> that proposes the use of cyclical learning rate schedules to generate an ensemble of trained neural networks while training only once. Loosely speaking, we can think of this as letting the neural network rapidly converge into a local minima, saving the model, resetting the learning rate, and repeat. We then take all of those saved models and use it as an ensemble without any additional computational cost. It's a very simple idea but it surprisingly works pretty well.</p>\n<p>If you're working with neural networks in TensorFlow Keras, I've created a small library <a href=\"https://github.com/adamvvu/snapshot_ensemble\" target=\"_blank\">here</a> that may be helpful. Let me know if you've gotten some interesting results with it!</p>",
      "rawMarkdown": "Hi everyone,\n\nJoined this competition recently and it seems like a fun project.\n\nAs you probably know, most of the top solutions on Kaggle/modern data science competitions rely on ML ensembling methods. Personally I am very interested in deep learning due to its theoretical properties, but for it to be competitive with a well-designed ensemble on a single metric would require extensive hyperparameter tuning and architectural considerations.\n\nAnyways there is an interesting paper [Huang et al. (2017)](https://arxiv.org/pdf/1704.00109.pdf) that proposes the use of cyclical learning rate schedules to generate an ensemble of trained neural networks while training only once. Loosely speaking, we can think of this as letting the neural network rapidly converge into a local minima, saving the model, resetting the learning rate, and repeat. We then take all of those saved models and use it as an ensemble without any additional computational cost. It's a very simple idea but it surprisingly works pretty well.\n\nIf you're working with neural networks in TensorFlow Keras, I've created a small library [here](https://github.com/adamvvu/snapshot_ensemble) that may be helpful. Let me know if you've gotten some interesting results with it!",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1904316": "Hi everyone,\n\nJoined this competition recently and it seems like a fun project.\n\nAs you probably know, most of the top solutions on Kaggle/modern data science competitions rely on ML ensembling methods. Personally I am very interested in deep learning due to its theoretical properties, but for it to be competitive with a well-designed ensemble on a single metric would require extensive hyperparameter tuning and architectural considerations.\n\nAnyways there is an interesting paper [Huang et al. (2017)](https://arxiv.org/pdf/1704.00109.pdf) that proposes the use of cyclical learning rate schedules to generate an ensemble of trained neural networks while training only once. Loosely speaking, we can think of this as letting the neural network rapidly converge into a local minima, saving the model, resetting the learning rate, and repeat. We then take all of those saved models and use it as an ensemble without any additional computational cost. It's a very simple idea but it surprisingly works pretty well.\n\nIf you're working with neural networks in TensorFlow Keras, I've created a small library [here](https://github.com/adamvvu/snapshot_ensemble) that may be helpful. Let me know if you've gotten some interesting results with it!"
  }
}