{
  "id": 160388,
  "title": "What does your Kaggle Pipeline look like?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/160388",
  "author_name": "",
  "post_date": "2020-06-21T05:08:48.848599500Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello Kagglers -</p>\n\n<p>This post has nothing do with this competition as such and may not help most people, but this would benefit people who would want to build better competition pipelines to experiment quickly and improve scores.  </p>\n\n<p>So, how does your kaggle pipeline look like for competitions in general?</p>\n\n<p>1) Do you have a DL system of your own or do you use only kaggle or any other cloud platform to compete at the highest level?</p>\n\n<p>2) How do you normally structure the pipeline/code folders to experiment and build models and submit them? and what other tech(apart from py, R) do you use to build the pipeline if in case there are any other. </p>\n\n<p>3) Would be great if competition Grandmasters, masters comment as I think that would benefit the interested kagglers in general.</p>\n\n<p>Happy Kaggling!! </p>",
  "messages": [
    {
      "id": "895083",
      "postDate": "06/21/2020 05:08:48",
      "content": "<p>Hello Kagglers -</p>\n\n<p>This post has nothing do with this competition as such and may not help most people, but this would benefit people who would want to build better competition pipelines to experiment quickly and improve scores.  </p>\n\n<p>So, how does your kaggle pipeline look like for competitions in general?</p>\n\n<p>1) Do you have a DL system of your own or do you use only kaggle or any other cloud platform to compete at the highest level?</p>\n\n<p>2) How do you normally structure the pipeline/code folders to experiment and build models and submit them? and what other tech(apart from py, R) do you use to build the pipeline if in case there are any other. </p>\n\n<p>3) Would be great if competition Grandmasters, masters comment as I think that would benefit the interested kagglers in general.</p>\n\n<p>Happy Kaggling!! </p>",
      "rawMarkdown": "Hello Kagglers -\n\nThis post has nothing do with this competition as such and may not help most people, but this would benefit people who would want to build better competition pipelines to experiment quickly and improve scores.  \n\nSo, how does your kaggle pipeline look like for competitions in general?\n\n1) Do you have a DL system of your own or do you use only kaggle or any other cloud platform to compete at the highest level?\n\n2) How do you normally structure the pipeline/code folders to experiment and build models and submit them? and what other tech(apart from py, R) do you use to build the pipeline if in case there are any other. \n\n3) Would be great if competition Grandmasters, masters comment as I think that would benefit the interested kagglers in general.\n\n\nHappy Kaggling!!",
      "votes": null
    },
    {
      "id": "895137",
      "postDate": "06/21/2020 06:02:03",
      "content": "<p>1) I have my own. There's also a work machine I sometimes use. If I'm really desperate for some more GPU power then I'll get some V100s, typically through AWS, but that's not common. I pretty much never use Kaggle notebooks unless there's something I want to show publicly. </p>\n\n<p>2) This a little hard to answer because it is always changing and can be dependent on the competition</p>\n\n<p>Typical folders are:\n- data -- stores competition data, external data, and any other preprocessed data. In some competitions it makes a lot more sense to do feature engineering or augmentations only one time and save them. It can save you loads of time if you're preprocessing things the same way every time you train a model.\n- models/weights -- stores saved model files and/or model weights\n- predictions -- contains predictions on the test set and CV predictions. If the model has multiple layers then there will be subdirectories for the layers.</p>\n\n<p>Other folders I keep are usually just other github repos that are competition specific and/or that I've made changes to.</p>",
      "rawMarkdown": "1) I have my own. There's also a work machine I sometimes use. If I'm really desperate for some more GPU power then I'll get some V100s, typically through AWS, but that's not common. I pretty much never use Kaggle notebooks unless there's something I want to show publicly. \n\n2) This a little hard to answer because it is always changing and can be dependent on the competition\n\nTypical folders are:\n- data -- stores competition data, external data, and any other preprocessed data. In some competitions it makes a lot more sense to do feature engineering or augmentations only one time and save them. It can save you loads of time if you're preprocessing things the same way every time you train a model.\n- models/weights -- stores saved model files and/or model weights\n- predictions -- contains predictions on the test set and CV predictions. If the model has multiple layers then there will be subdirectories for the layers.\n\nOther folders I keep are usually just other github repos that are competition specific and/or that I've made changes to.",
      "votes": null
    },
    {
      "id": "895151",
      "postDate": "06/21/2020 06:20:54",
      "content": "<p>Thanks <a href=\"/brandenkmurray\">@brandenkmurray</a> </p>\n\n<p>\"predictions -- contains predictions on the test set and CV predictions. If the model has multiple layers then there will be subdirectories for the layers.\" -----  Do you mean subdirectories to save weights for each layer?</p>\n\n<p>\"  I have my own. There's also a work machine I sometimes use. If I'm really desperate for some more GPU power then I'll get some V100s, typically through AWS, but that's not common. I pretty much never use Kaggle notebooks unless there's something I want to show publicly. \" -----  Curious to know how you submit predictions when everything is done outside the kaggle environment. Thats why asked on the tech stack. Do you use git to pull and push submissions? Would be great to know!! </p>",
      "rawMarkdown": "Thanks @brandenkmurray \n\n\"predictions -- contains predictions on the test set and CV predictions. If the model has multiple layers then there will be subdirectories for the layers.\" -----  Do you mean subdirectories to save weights for each layer?\n\n\"  I have my own. There's also a work machine I sometimes use. If I'm really desperate for some more GPU power then I'll get some V100s, typically through AWS, but that's not common. I pretty much never use Kaggle notebooks unless there's something I want to show publicly. \" -----  Curious to know how you submit predictions when everything is done outside the kaggle environment. Thats why asked on the tech stack. Do you use git to pull and push submissions? Would be great to know!!",
      "votes": null
    },
    {
      "id": "895188",
      "postDate": "06/21/2020 06:59:06",
      "content": "<ol>\n<li><p>I use my own computer for simple/light-weight experiments. For longer training, we switch to my server. </p></li>\n<li><p>Here is what my pipeline looks like. <strong>It works for all DL competitions (CV, NLP)</strong> </p></li>\n</ol>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1938879%2Fbe7ea424229516701c161db54434a649%2F2020-06-21_15-46.png?generation=1592722052362619&amp;alt=media\" alt=\"\"></p>\n\n<p>My pipeline is based on the Catalyst. I collect the most common functions, modules from my previous competitions into a <code>hub</code>. \nWhat I do when joining a new competition is: create a new configure file \n```\nmodel_params:\n  model: TIMMModels\n  model_name: \"resnet34\"\n  num_classes: 1 </p>\n\n<p>criterion_params:\n    criterion: BCEWithLogitsLoss</p>\n\n<p>optimizer_params:\n    optimizer: AdamW\n    lr: 0.0001\n    weight_decay: 0.0001 </p>\n\n<p>.... \n``` \nMy pipeline will know what it needs to do automatically when I pass this configuration file</p>",
      "rawMarkdown": "1.  I use my own computer for simple/light-weight experiments. For longer training, we switch to my server. \n\n2. Here is what my pipeline looks like. **It works for all DL competitions (CV, NLP)** \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1938879%2Fbe7ea424229516701c161db54434a649%2F2020-06-21_15-46.png?generation=1592722052362619&amp;alt=media)\n\nMy pipeline is based on the Catalyst. I collect the most common functions, modules from my previous competitions into a `hub`. \nWhat I do when joining a new competition is: create a new configure file \n```\nmodel_params:\n  model: TIMMModels\n  model_name: \"resnet34\"\n  num_classes: 1 \n\n criterion_params:\n    criterion: BCEWithLogitsLoss\n\n optimizer_params:\n    optimizer: AdamW\n    lr: 0.0001\n    weight_decay: 0.0001 \n\n.... \n``` \nMy pipeline will know what it needs to do automatically when I pass this configuration file",
      "votes": null
    },
    {
      "id": "896961",
      "postDate": "06/22/2020 14:41:33",
      "content": "<p>Honestly, I used to download data on my computer, then have nicely sorted folders. But I quickly switched to simple Kaggle notebooks. It forces you to organize your notebook in a clever way and to streamline as much useless code. If I reuse a lot a snippet of code, I might upload it as a Python script and then reuse it when needed as <a href=\"/abhishek\">@abhishek</a> suggested.</p>",
      "rawMarkdown": "Honestly, I used to download data on my computer, then have nicely sorted folders. But I quickly switched to simple Kaggle notebooks. It forces you to organize your notebook in a clever way and to streamline as much useless code. If I reuse a lot a snippet of code, I might upload it as a Python script and then reuse it when needed as @abhishek suggested.",
      "votes": null
    },
    {
      "id": "897173",
      "postDate": "06/22/2020 17:07:13",
      "content": "<p><a href=\"/rftexas\">@rftexas</a>  - hmm , how do you experiment using notebooks? is there a sample look-alike structure you can share? Thanks! </p>",
      "rawMarkdown": "rftexas  - hmm , how do you experiment using notebooks? is there a sample look-alike structure you can share? Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 895137,
      "author_name": "brandenkmurray",
      "author_url": "",
      "post_date": "06/21/2020 06:02:03",
      "content": "<p>1) I have my own. There's also a work machine I sometimes use. If I'm really desperate for some more GPU power then I'll get some V100s, typically through AWS, but that's not common. I pretty much never use Kaggle notebooks unless there's something I want to show publicly. </p>\n\n<p>2) This a little hard to answer because it is always changing and can be dependent on the competition</p>\n\n<p>Typical folders are:\n- data -- stores competition data, external data, and any other preprocessed data. In some competitions it makes a lot more sense to do feature engineering or augmentations only one time and save them. It can save you loads of time if you're preprocessing things the same way every time you train a model.\n- models/weights -- stores saved model files and/or model weights\n- predictions -- contains predictions on the test set and CV predictions. If the model has multiple layers then there will be subdirectories for the layers.</p>\n\n<p>Other folders I keep are usually just other github repos that are competition specific and/or that I've made changes to.</p>",
      "votes": null,
      "replies": [
        {
          "id": 895151,
          "author_name": "akashram",
          "author_url": "",
          "post_date": "06/21/2020 06:20:54",
          "content": "<p>Thanks <a href=\"/brandenkmurray\">@brandenkmurray</a> </p>\n\n<p>\"predictions -- contains predictions on the test set and CV predictions. If the model has multiple layers then there will be subdirectories for the layers.\" -----  Do you mean subdirectories to save weights for each layer?</p>\n\n<p>\"  I have my own. There's also a work machine I sometimes use. If I'm really desperate for some more GPU power then I'll get some V100s, typically through AWS, but that's not common. I pretty much never use Kaggle notebooks unless there's something I want to show publicly. \" -----  Curious to know how you submit predictions when everything is done outside the kaggle environment. Thats why asked on the tech stack. Do you use git to pull and push submissions? Would be great to know!! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 895188,
      "author_name": "backaggle",
      "author_url": "",
      "post_date": "06/21/2020 06:59:06",
      "content": "<ol>\n<li><p>I use my own computer for simple/light-weight experiments. For longer training, we switch to my server. </p></li>\n<li><p>Here is what my pipeline looks like. <strong>It works for all DL competitions (CV, NLP)</strong> </p></li>\n</ol>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1938879%2Fbe7ea424229516701c161db54434a649%2F2020-06-21_15-46.png?generation=1592722052362619&amp;alt=media\" alt=\"\"></p>\n\n<p>My pipeline is based on the Catalyst. I collect the most common functions, modules from my previous competitions into a <code>hub</code>. \nWhat I do when joining a new competition is: create a new configure file \n```\nmodel_params:\n  model: TIMMModels\n  model_name: \"resnet34\"\n  num_classes: 1 </p>\n\n<p>criterion_params:\n    criterion: BCEWithLogitsLoss</p>\n\n<p>optimizer_params:\n    optimizer: AdamW\n    lr: 0.0001\n    weight_decay: 0.0001 </p>\n\n<p>.... \n``` \nMy pipeline will know what it needs to do automatically when I pass this configuration file</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 896961,
      "author_name": "rftexas",
      "author_url": "",
      "post_date": "06/22/2020 14:41:33",
      "content": "<p>Honestly, I used to download data on my computer, then have nicely sorted folders. But I quickly switched to simple Kaggle notebooks. It forces you to organize your notebook in a clever way and to streamline as much useless code. If I reuse a lot a snippet of code, I might upload it as a Python script and then reuse it when needed as <a href=\"/abhishek\">@abhishek</a> suggested.</p>",
      "votes": null,
      "replies": [
        {
          "id": 897173,
          "author_name": "akashram",
          "author_url": "",
          "post_date": "06/22/2020 17:07:13",
          "content": "<p><a href=\"/rftexas\">@rftexas</a>  - hmm , how do you experiment using notebooks? is there a sample look-alike structure you can share? Thanks! </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "895083": "Hello Kagglers -\n\nThis post has nothing do with this competition as such and may not help most people, but this would benefit people who would want to build better competition pipelines to experiment quickly and improve scores.  \n\nSo, how does your kaggle pipeline look like for competitions in general?\n\n1) Do you have a DL system of your own or do you use only kaggle or any other cloud platform to compete at the highest level?\n\n2) How do you normally structure the pipeline/code folders to experiment and build models and submit them? and what other tech(apart from py, R) do you use to build the pipeline if in case there are any other. \n\n3) Would be great if competition Grandmasters, masters comment as I think that would benefit the interested kagglers in general.\n\n\nHappy Kaggling!!",
    "895137": "1) I have my own. There's also a work machine I sometimes use. If I'm really desperate for some more GPU power then I'll get some V100s, typically through AWS, but that's not common. I pretty much never use Kaggle notebooks unless there's something I want to show publicly. \n\n2) This a little hard to answer because it is always changing and can be dependent on the competition\n\nTypical folders are:\n- data -- stores competition data, external data, and any other preprocessed data. In some competitions it makes a lot more sense to do feature engineering or augmentations only one time and save them. It can save you loads of time if you're preprocessing things the same way every time you train a model.\n- models/weights -- stores saved model files and/or model weights\n- predictions -- contains predictions on the test set and CV predictions. If the model has multiple layers then there will be subdirectories for the layers.\n\nOther folders I keep are usually just other github repos that are competition specific and/or that I've made changes to.",
    "895151": "Thanks @brandenkmurray \n\n\"predictions -- contains predictions on the test set and CV predictions. If the model has multiple layers then there will be subdirectories for the layers.\" -----  Do you mean subdirectories to save weights for each layer?\n\n\"  I have my own. There's also a work machine I sometimes use. If I'm really desperate for some more GPU power then I'll get some V100s, typically through AWS, but that's not common. I pretty much never use Kaggle notebooks unless there's something I want to show publicly. \" -----  Curious to know how you submit predictions when everything is done outside the kaggle environment. Thats why asked on the tech stack. Do you use git to pull and push submissions? Would be great to know!!",
    "895188": "1.  I use my own computer for simple/light-weight experiments. For longer training, we switch to my server. \n\n2. Here is what my pipeline looks like. **It works for all DL competitions (CV, NLP)** \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1938879%2Fbe7ea424229516701c161db54434a649%2F2020-06-21_15-46.png?generation=1592722052362619&amp;alt=media)\n\nMy pipeline is based on the Catalyst. I collect the most common functions, modules from my previous competitions into a `hub`. \nWhat I do when joining a new competition is: create a new configure file \n```\nmodel_params:\n  model: TIMMModels\n  model_name: \"resnet34\"\n  num_classes: 1 \n\n criterion_params:\n    criterion: BCEWithLogitsLoss\n\n optimizer_params:\n    optimizer: AdamW\n    lr: 0.0001\n    weight_decay: 0.0001 \n\n.... \n``` \nMy pipeline will know what it needs to do automatically when I pass this configuration file",
    "896961": "Honestly, I used to download data on my computer, then have nicely sorted folders. But I quickly switched to simple Kaggle notebooks. It forces you to organize your notebook in a clever way and to streamline as much useless code. If I reuse a lot a snippet of code, I might upload it as a Python script and then reuse it when needed as @abhishek suggested.",
    "897173": "rftexas  - hmm , how do you experiment using notebooks? is there a sample look-alike structure you can share? Thanks!"
  },
  "source": "meta"
}