{
  "id": 111531,
  "title": "My machine learning template",
  "url": "/competitions/understanding_cloud_organization/discussion/111531",
  "author_name": "bilal2vec",
  "post_date": "2019-10-06T17:40:19.070000",
  "votes": 3,
  "comment_count": 1,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1524032%2F0b74a80cc47ef0cfa1926bbc2f64ed96%2Fcarbon.png?generation=1570382921456799&amp;alt=media\" alt=\"screenshot of kaggle-siim quickstart\"></p>\n\n<p><em>TL;DR</em>\nThis is my <a href=\"https://github.com/bkkaggle/kaggle-siim\">template</a>, it helps you write less code, but isn't opinionated to a certain workflow and is easily modified and extended.</p>\n\n<p>Hi,</p>\n\n<p>I saw @artgor 's <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/111375\">post</a> on his modular kaggle framework and wanted to share my own template for kaggle and other machine learning projects. </p>\n\n<p>Since most of the competitions I've worked on in the last two years are image segmentation competitions, I've found that a lot of the core parts of the code, - the training loop, the calculation of val metrics, tensorboard logging, FP16 training, and gradient accumulation - are reused from competiton to competition.</p>\n\n<p>I made <a href=\"https://github.com/bkkaggle/kaggle-siim\">kaggle-siim</a> while I was working on the SIIM pneumothorax competition because I wanted a way to be able to have a framework that would work just as well for trying out new ideas and architectures as for reducing the amount of boilerplate code I have to write at the beginning of every competition.</p>\n\n<p>To make it both reusable and extendable, I've implemented a config [file] (<a href=\"https://github.com/bkkaggle/kaggle-siim/blob/master/config.py\">https://github.com/bkkaggle/kaggle-siim/blob/master/config.py</a>) that you can populate with reasonable defaults for commonly used hyperparameters like batch size, n_epochs, and learning rates. At the same time, you can override these defaults by passing command line parameters to the training script. </p>\n\n<p>Here's an example to show you how this could be done:</p>\n\n<p>Set reasonable defaults in the config file</p>\n\n<p><code>config.py</code>:\n```\nBATCH_SIZE = 4\nGRADIENT_ACCUMULATION_STEPS = 2</p>\n\n<p>WORKERS = 8</p>\n\n<p>EPOCHS = 10\n```</p>\n\n<p>but have the ability to override them from the command line\n<code>\npython train.py --batch_size=32 --epochs=100\n</code></p>\n\n<hr>\n\n<p>I also made a pypi package, <a href=\"https://github.com/bkkaggle/pytorch_zoo\">pytorch_zoo</a>, that I used in this project. I converted it into a utility <a href=\"https://www.kaggle.com/bkkaggle/pytorch-zoo\">script</a> for the utility script <a href=\"https://www.kaggle.com/general/109651639219\">competition</a>. Take a look at the demo <a href=\"https://www.kaggle.com/bkkaggle/pytorch-zoo-demo\">here</a>.</p>",
  "messages": [
    {
      "id": 642819,
      "postDate": "2019-10-06T17:40:19.070Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1524032%2F0b74a80cc47ef0cfa1926bbc2f64ed96%2Fcarbon.png?generation=1570382921456799&amp;alt=media\" alt=\"screenshot of kaggle-siim quickstart\"></p>\n\n<p><em>TL;DR</em>\nThis is my <a href=\"https://github.com/bkkaggle/kaggle-siim\">template</a>, it helps you write less code, but isn't opinionated to a certain workflow and is easily modified and extended.</p>\n\n<p>Hi,</p>\n\n<p>I saw @artgor 's <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/111375\">post</a> on his modular kaggle framework and wanted to share my own template for kaggle and other machine learning projects. </p>\n\n<p>Since most of the competitions I've worked on in the last two years are image segmentation competitions, I've found that a lot of the core parts of the code, - the training loop, the calculation of val metrics, tensorboard logging, FP16 training, and gradient accumulation - are reused from competiton to competition.</p>\n\n<p>I made <a href=\"https://github.com/bkkaggle/kaggle-siim\">kaggle-siim</a> while I was working on the SIIM pneumothorax competition because I wanted a way to be able to have a framework that would work just as well for trying out new ideas and architectures as for reducing the amount of boilerplate code I have to write at the beginning of every competition.</p>\n\n<p>To make it both reusable and extendable, I've implemented a config [file] (<a href=\"https://github.com/bkkaggle/kaggle-siim/blob/master/config.py\">https://github.com/bkkaggle/kaggle-siim/blob/master/config.py</a>) that you can populate with reasonable defaults for commonly used hyperparameters like batch size, n_epochs, and learning rates. At the same time, you can override these defaults by passing command line parameters to the training script. </p>\n\n<p>Here's an example to show you how this could be done:</p>\n\n<p>Set reasonable defaults in the config file</p>\n\n<p><code>config.py</code>:\n```\nBATCH_SIZE = 4\nGRADIENT_ACCUMULATION_STEPS = 2</p>\n\n<p>WORKERS = 8</p>\n\n<p>EPOCHS = 10\n```</p>\n\n<p>but have the ability to override them from the command line\n<code>\npython train.py --batch_size=32 --epochs=100\n</code></p>\n\n<hr>\n\n<p>I also made a pypi package, <a href=\"https://github.com/bkkaggle/pytorch_zoo\">pytorch_zoo</a>, that I used in this project. I converted it into a utility <a href=\"https://www.kaggle.com/bkkaggle/pytorch-zoo\">script</a> for the utility script <a href=\"https://www.kaggle.com/general/109651639219\">competition</a>. Take a look at the demo <a href=\"https://www.kaggle.com/bkkaggle/pytorch-zoo-demo\">here</a>.</p>",
      "rawMarkdown": "![screenshot of kaggle-siim quickstart](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1524032%2F0b74a80cc47ef0cfa1926bbc2f64ed96%2Fcarbon.png?generation=1570382921456799&amp;alt=media)\n\n*TL;DR*\nThis is my [template](https://github.com/bkkaggle/kaggle-siim), it helps you write less code, but isn't opinionated to a certain workflow and is easily modified and extended.\n\nHi,\n\nI saw @artgor 's [post](https://www.kaggle.com/c/understanding_cloud_organization/discussion/111375) on his modular kaggle framework and wanted to share my own template for kaggle and other machine learning projects. \n\nSince most of the competitions I've worked on in the last two years are image segmentation competitions, I've found that a lot of the core parts of the code, - the training loop, the calculation of val metrics, tensorboard logging, FP16 training, and gradient accumulation - are reused from competiton to competition.\n\nI made [kaggle-siim](https://github.com/bkkaggle/kaggle-siim) while I was working on the SIIM pneumothorax competition because I wanted a way to be able to have a framework that would work just as well for trying out new ideas and architectures as for reducing the amount of boilerplate code I have to write at the beginning of every competition.\n\nTo make it both reusable and extendable, I've implemented a config [file] (https://github.com/bkkaggle/kaggle-siim/blob/master/config.py) that you can populate with reasonable defaults for commonly used hyperparameters like batch size, n_epochs, and learning rates. At the same time, you can override these defaults by passing command line parameters to the training script. \n\nHere's an example to show you how this could be done:\n\nSet reasonable defaults in the config file\n\n`config.py`:\n```\nBATCH_SIZE = 4\nGRADIENT_ACCUMULATION_STEPS = 2\n\nWORKERS = 8\n\nEPOCHS = 10\n```\n\nbut have the ability to override them from the command line\n```\npython train.py --batch_size=32 --epochs=100\n```\n\n------\n\nI also made a pypi package, [pytorch_zoo](https://github.com/bkkaggle/pytorch_zoo), that I used in this project. I converted it into a utility [script](https://www.kaggle.com/bkkaggle/pytorch-zoo) for the utility script [competition](https://www.kaggle.com/general/109651639219). Take a look at the demo [here](https://www.kaggle.com/bkkaggle/pytorch-zoo-demo).",
      "votes": 3
    },
    {
      "id": 643115,
      "postDate": "2019-10-07T06:11:16.487Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 643115,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-10-07T06:11:16.487000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "642819": "![screenshot of kaggle-siim quickstart](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1524032%2F0b74a80cc47ef0cfa1926bbc2f64ed96%2Fcarbon.png?generation=1570382921456799&amp;alt=media)\n\n*TL;DR*\nThis is my [template](https://github.com/bkkaggle/kaggle-siim), it helps you write less code, but isn't opinionated to a certain workflow and is easily modified and extended.\n\nHi,\n\nI saw @artgor 's [post](https://www.kaggle.com/c/understanding_cloud_organization/discussion/111375) on his modular kaggle framework and wanted to share my own template for kaggle and other machine learning projects. \n\nSince most of the competitions I've worked on in the last two years are image segmentation competitions, I've found that a lot of the core parts of the code, - the training loop, the calculation of val metrics, tensorboard logging, FP16 training, and gradient accumulation - are reused from competiton to competition.\n\nI made [kaggle-siim](https://github.com/bkkaggle/kaggle-siim) while I was working on the SIIM pneumothorax competition because I wanted a way to be able to have a framework that would work just as well for trying out new ideas and architectures as for reducing the amount of boilerplate code I have to write at the beginning of every competition.\n\nTo make it both reusable and extendable, I've implemented a config [file] (https://github.com/bkkaggle/kaggle-siim/blob/master/config.py) that you can populate with reasonable defaults for commonly used hyperparameters like batch size, n_epochs, and learning rates. At the same time, you can override these defaults by passing command line parameters to the training script. \n\nHere's an example to show you how this could be done:\n\nSet reasonable defaults in the config file\n\n`config.py`:\n```\nBATCH_SIZE = 4\nGRADIENT_ACCUMULATION_STEPS = 2\n\nWORKERS = 8\n\nEPOCHS = 10\n```\n\nbut have the ability to override them from the command line\n```\npython train.py --batch_size=32 --epochs=100\n```\n\n------\n\nI also made a pypi package, [pytorch_zoo](https://github.com/bkkaggle/pytorch_zoo), that I used in this project. I converted it into a utility [script](https://www.kaggle.com/bkkaggle/pytorch-zoo) for the utility script [competition](https://www.kaggle.com/general/109651639219). Take a look at the demo [here](https://www.kaggle.com/bkkaggle/pytorch-zoo-demo).",
    "643115": ""
  }
}