{
  "id": 111375,
  "title": "Improving code quality with utility scripts",
  "url": "/competitions/understanding_cloud_organization/discussion/111375",
  "author_name": "Andrey Lukyanenko",
  "post_date": "2019-10-05T07:07:33.140000",
  "votes": 53,
  "comment_count": 14,
  "views": 0,
  "content": "<p>When we take part in kaggle competitions, we usually run a lot of experiments. If it is necessary to manually change the core code for each experiment, we'll spent a lot of unnecessary time, also chances are that we'll make some mistakes.</p>\n\n<p>So it is more efficient to write a modular code which can be easily reused and which would allow to try various parameters without changing the code. Of course, if you have some completely new idea, it would be necessary to write a code for it. But after this you should be able to use it easily.</p>\n\n<p>Some time ago I wrote this <a href=\"https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools\">kernel</a> for the competition. It is okay for kaggle, but when I run experiments locally, I prefer to write the code in Pycharm, use Jupyter notebooks for visualization and run the experiments from the command line.</p>\n\n<p>Here is my github repo for this competition: <a href=\"https://github.com/Erlemar/Understanding-Clouds-from-Satellite-Images\">https://github.com/Erlemar/Understanding-Clouds-from-Satellite-Images</a></p>\n\n<p>The work is still in progress, but it is already easier to run experiments for me.</p>\n\n<p>For example, I can run a segmenation model with this command:\n<code>\npython train.py --encoder resnet50 --bs 20 --num_epochs 100 --train True --optimize_postprocess True --make_prediction True\n</code></p>\n\n<p>I'm still thinking about better ways to run experiments - maybe using config files.</p>\n\n<p>Considering there is currently a competition for useful utility scripts, I have decided to write such a script based on my code, here is it:\n<a href=\"https://www.kaggle.com/artgor/pytorch-utils-for-images\">https://www.kaggle.com/artgor/pytorch-utils-for-images</a></p>\n\n<p>I hope it would be useful to you! By the way, it can be easily used for example for severstal competition - you'd only need to modify the <code>prepare_loaders</code> function ;)</p>\n\n<p>My post in a thread with competition for utility scripts: <a href=\"https://www.kaggle.com/general/109651#641828\">https://www.kaggle.com/general/109651#641828</a></p>\n\n<p>To demonstrate the possibilities, I wrote this kernel:\n<a href=\"https://www.kaggle.com/artgor/classification-in-catalyst-with-utility-scripts\">https://www.kaggle.com/artgor/classification-in-catalyst-with-utility-scripts</a></p>\n\n<p>These are the main libraries used in my code:\n* <a href=\"https://github.com/albu/albumentations\">albumentations</a>: this is a great library for image augmentation which makes it easier and more convenient\n* <a href=\"https://github.com/catalyst-team/catalyst\">catalyst</a>: this is a great library which makes using PyTorch easier, helps with reprodicibility and contains a lot of useful utils\n* <a href=\"https://github.com/qubvel/segmentation_models.pytorch\">segmentation_models_pytorch</a>: this is a great library with convenient wrappers for models, losses and other useful things\n* <a href=\"https://github.com/BloodAxe/pytorch-toolbelt\">pytorch-toolbelt</a>: this is a great library with many useful shortcuts for building pytorch models</p>\n\n<p>Some examples:\nYou can get loaders like this:\n<code>\npreprocessing_fn = smp.encoders.get_preprocessing_fn(encoder, encoder_weights)\nloaders = prepare_loaders(path=path, bs=batch_size, num_workers=num_workers, preprocessing_fn=preprocessing_fn, preload=False, task=task, image_size=(224, 224))\n</code></p>\n\n<p>There are two dataset classes - one for segmentation and one for classification. The main difference is in the way the target is created.</p>\n\n<p>Switching between classification and segmentation tasks and choosing encoders is also simple:\n<code>\nmodel = get_model(model_type=segm_type, encoder=encoder, encoder_weights=encoder_weights, activation=activation, task=task, n_classes=n_classes, head='simple')\n</code></p>",
  "messages": [
    {
      "id": 641826,
      "postDate": "2019-10-05T07:07:33.140Z",
      "content": "<p>When we take part in kaggle competitions, we usually run a lot of experiments. If it is necessary to manually change the core code for each experiment, we'll spent a lot of unnecessary time, also chances are that we'll make some mistakes.</p>\n\n<p>So it is more efficient to write a modular code which can be easily reused and which would allow to try various parameters without changing the code. Of course, if you have some completely new idea, it would be necessary to write a code for it. But after this you should be able to use it easily.</p>\n\n<p>Some time ago I wrote this <a href=\"https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools\">kernel</a> for the competition. It is okay for kaggle, but when I run experiments locally, I prefer to write the code in Pycharm, use Jupyter notebooks for visualization and run the experiments from the command line.</p>\n\n<p>Here is my github repo for this competition: <a href=\"https://github.com/Erlemar/Understanding-Clouds-from-Satellite-Images\">https://github.com/Erlemar/Understanding-Clouds-from-Satellite-Images</a></p>\n\n<p>The work is still in progress, but it is already easier to run experiments for me.</p>\n\n<p>For example, I can run a segmenation model with this command:\n<code>\npython train.py --encoder resnet50 --bs 20 --num_epochs 100 --train True --optimize_postprocess True --make_prediction True\n</code></p>\n\n<p>I'm still thinking about better ways to run experiments - maybe using config files.</p>\n\n<p>Considering there is currently a competition for useful utility scripts, I have decided to write such a script based on my code, here is it:\n<a href=\"https://www.kaggle.com/artgor/pytorch-utils-for-images\">https://www.kaggle.com/artgor/pytorch-utils-for-images</a></p>\n\n<p>I hope it would be useful to you! By the way, it can be easily used for example for severstal competition - you'd only need to modify the <code>prepare_loaders</code> function ;)</p>\n\n<p>My post in a thread with competition for utility scripts: <a href=\"https://www.kaggle.com/general/109651#641828\">https://www.kaggle.com/general/109651#641828</a></p>\n\n<p>To demonstrate the possibilities, I wrote this kernel:\n<a href=\"https://www.kaggle.com/artgor/classification-in-catalyst-with-utility-scripts\">https://www.kaggle.com/artgor/classification-in-catalyst-with-utility-scripts</a></p>\n\n<p>These are the main libraries used in my code:\n* <a href=\"https://github.com/albu/albumentations\">albumentations</a>: this is a great library for image augmentation which makes it easier and more convenient\n* <a href=\"https://github.com/catalyst-team/catalyst\">catalyst</a>: this is a great library which makes using PyTorch easier, helps with reprodicibility and contains a lot of useful utils\n* <a href=\"https://github.com/qubvel/segmentation_models.pytorch\">segmentation_models_pytorch</a>: this is a great library with convenient wrappers for models, losses and other useful things\n* <a href=\"https://github.com/BloodAxe/pytorch-toolbelt\">pytorch-toolbelt</a>: this is a great library with many useful shortcuts for building pytorch models</p>\n\n<p>Some examples:\nYou can get loaders like this:\n<code>\npreprocessing_fn = smp.encoders.get_preprocessing_fn(encoder, encoder_weights)\nloaders = prepare_loaders(path=path, bs=batch_size, num_workers=num_workers, preprocessing_fn=preprocessing_fn, preload=False, task=task, image_size=(224, 224))\n</code></p>\n\n<p>There are two dataset classes - one for segmentation and one for classification. The main difference is in the way the target is created.</p>\n\n<p>Switching between classification and segmentation tasks and choosing encoders is also simple:\n<code>\nmodel = get_model(model_type=segm_type, encoder=encoder, encoder_weights=encoder_weights, activation=activation, task=task, n_classes=n_classes, head='simple')\n</code></p>",
      "rawMarkdown": "When we take part in kaggle competitions, we usually run a lot of experiments. If it is necessary to manually change the core code for each experiment, we'll spent a lot of unnecessary time, also chances are that we'll make some mistakes.\n\nSo it is more efficient to write a modular code which can be easily reused and which would allow to try various parameters without changing the code. Of course, if you have some completely new idea, it would be necessary to write a code for it. But after this you should be able to use it easily.\n\nSome time ago I wrote this [kernel](https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools) for the competition. It is okay for kaggle, but when I run experiments locally, I prefer to write the code in Pycharm, use Jupyter notebooks for visualization and run the experiments from the command line.\n\nHere is my github repo for this competition: https://github.com/Erlemar/Understanding-Clouds-from-Satellite-Images\n\nThe work is still in progress, but it is already easier to run experiments for me.\n\nFor example, I can run a segmenation model with this command:\n```\npython train.py --encoder resnet50 --bs 20 --num_epochs 100 --train True --optimize_postprocess True --make_prediction True\n```\n\nI'm still thinking about better ways to run experiments - maybe using config files.\n\nConsidering there is currently a competition for useful utility scripts, I have decided to write such a script based on my code, here is it:\nhttps://www.kaggle.com/artgor/pytorch-utils-for-images\n\nI hope it would be useful to you! By the way, it can be easily used for example for severstal competition - you'd only need to modify the `prepare_loaders` function ;)\n\nMy post in a thread with competition for utility scripts: https://www.kaggle.com/general/109651#641828\n\nTo demonstrate the possibilities, I wrote this kernel:\nhttps://www.kaggle.com/artgor/classification-in-catalyst-with-utility-scripts\n\nThese are the main libraries used in my code:\n* [albumentations](https://github.com/albu/albumentations): this is a great library for image augmentation which makes it easier and more convenient\n* [catalyst](https://github.com/catalyst-team/catalyst): this is a great library which makes using PyTorch easier, helps with reprodicibility and contains a lot of useful utils\n* [segmentation_models_pytorch](https://github.com/qubvel/segmentation_models.pytorch): this is a great library with convenient wrappers for models, losses and other useful things\n* [pytorch-toolbelt](https://github.com/BloodAxe/pytorch-toolbelt): this is a great library with many useful shortcuts for building pytorch models\n\nSome examples:\nYou can get loaders like this:\n```\npreprocessing_fn = smp.encoders.get_preprocessing_fn(encoder, encoder_weights)\nloaders = prepare_loaders(path=path, bs=batch_size, num_workers=num_workers, preprocessing_fn=preprocessing_fn, preload=False, task=task, image_size=(224, 224))\n```\n\nThere are two dataset classes - one for segmentation and one for classification. The main difference is in the way the target is created.\n\nSwitching between classification and segmentation tasks and choosing encoders is also simple:\n```\nmodel = get_model(model_type=segm_type, encoder=encoder, encoder_weights=encoder_weights, activation=activation, task=task, n_classes=n_classes, head='simple')\n```",
      "votes": 53
    },
    {
      "id": 642360,
      "postDate": "2019-10-06T00:32:41.023Z",
      "content": "<p>Amazing work! Thanks for sharing! \nI recently heard that the pipeline offered by QuantumBlack is really useful.</p>",
      "rawMarkdown": "Amazing work! Thanks for sharing! \nI recently heard that the pipeline offered by QuantumBlack is really useful.",
      "votes": 1,
      "replies": [
        {
          "id": 642417,
          "postDate": "2019-10-06T03:44:15.707Z",
          "content": "<p>I have heard about Kedro, but this tool is quite new - there is little information about it.</p>",
          "rawMarkdown": "I have heard about Kedro, but this tool is quite new - there is little information about it.",
          "votes": 1
        },
        {
          "id": 642420,
          "postDate": "2019-10-06T03:50:30.880Z",
          "content": "<p>Actually, I talked with one of the manages of Kedro last week. He said that it is suited for a big project and can reduce the gap between data engineers and data analysists. </p>",
          "rawMarkdown": "Actually, I talked with one of the manages of Kedro last week. He said that it is suited for a big project and can reduce the gap between data engineers and data analysists. ",
          "votes": 1
        },
        {
          "id": 642421,
          "postDate": "2019-10-06T03:57:49.360Z",
          "content": "<p>I see, thanks! Then I suppose it is more suited for the real job and not for Kaggle 😄 </p>",
          "rawMarkdown": "I see, thanks! Then I suppose it is more suited for the real job and not for Kaggle 😄 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 641891,
      "postDate": "2019-10-05T09:16:40.547Z",
      "content": "<p>could be great assets for researchers in this field,thank you <a href=\"/artgor\">@artgor</a> </p>",
      "rawMarkdown": "could be great assets for researchers in this field,thank you @artgor ",
      "votes": 1
    },
    {
      "id": 641876,
      "postDate": "2019-10-05T08:59:43.523Z",
      "content": "<p>Great...\nVery Helpful...\nMany Thanks for Sharing....!! <a href=\"/artgor\">@artgor</a> </p>",
      "rawMarkdown": "Great...\nVery Helpful...\nMany Thanks for Sharing....!! @artgor ",
      "votes": 1
    },
    {
      "id": 641857,
      "postDate": "2019-10-05T08:40:51.090Z",
      "content": "<p>disposed old comment</p>",
      "rawMarkdown": "disposed old comment",
      "votes": 2,
      "replies": [
        {
          "id": 642229,
          "postDate": "2019-10-05T18:18:52.863Z",
          "content": "<p>Thanks, I'm glad it is useful :)</p>\n\n<p>Sharing is also useful to me - it helps me to better structure my code and thoughts as well as receive a valuable feedback.</p>\n\n<p>Having separate notebooks for different tasks is also a good idea.</p>",
          "rawMarkdown": "Thanks, I'm glad it is useful :)\n\nSharing is also useful to me - it helps me to better structure my code and thoughts as well as receive a valuable feedback.\n\nHaving separate notebooks for different tasks is also a good idea."
        }
      ]
    },
    {
      "id": 1310824,
      "postDate": "2021-05-17T02:10:16.620Z",
      "content": "<p>Thanks for sharing!  Wow! You're amazing!  I went to your GitHub, your own website, your portfolio, your blog and saw your interview.  Such a great accomplishment! </p>",
      "rawMarkdown": "Thanks for sharing!  Wow! You're amazing!  I went to your GitHub, your own website, your portfolio, your blog and saw your interview.  Such a great accomplishment! "
    },
    {
      "id": 668283,
      "postDate": "2019-11-08T08:11:35.530Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 642227,
      "postDate": "2019-10-05T18:15:29.083Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 642230,
          "postDate": "2019-10-05T18:19:31.017Z",
          "content": "<p>Yes, I saw these config files several times already. I suppose after some time I'll also use something similar.</p>",
          "rawMarkdown": "Yes, I saw these config files several times already. I suppose after some time I'll also use something similar."
        },
        {
          "id": 642391,
          "postDate": "2019-10-06T01:54:57.343Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 642419,
          "postDate": "2019-10-06T03:48:14.317Z",
          "content": "<p>I didn't try Catalyst's segmentation models yet, but I think they should also work well.</p>",
          "rawMarkdown": "I didn't try Catalyst's segmentation models yet, but I think they should also work well."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 642360,
      "author_name": "Hideaki Takahashi",
      "author_url": "",
      "post_date": "2019-10-06T00:32:41.023000",
      "content": "<p>Amazing work! Thanks for sharing! \nI recently heard that the pipeline offered by QuantumBlack is really useful.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 642417,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2019-10-06T03:44:15.707000",
          "content": "<p>I have heard about Kedro, but this tool is quite new - there is little information about it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 642420,
          "author_name": "Hideaki Takahashi",
          "author_url": "",
          "post_date": "2019-10-06T03:50:30.880000",
          "content": "<p>Actually, I talked with one of the manages of Kedro last week. He said that it is suited for a big project and can reduce the gap between data engineers and data analysists. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 642421,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2019-10-06T03:57:49.360000",
          "content": "<p>I see, thanks! Then I suppose it is more suited for the real job and not for Kaggle 😄 </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 641891,
      "author_name": "Mobassir",
      "author_url": "",
      "post_date": "2019-10-05T09:16:40.547000",
      "content": "<p>could be great assets for researchers in this field,thank you <a href=\"/artgor\">@artgor</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 641876,
      "author_name": "Ailurophile",
      "author_url": "",
      "post_date": "2019-10-05T08:59:43.523000",
      "content": "<p>Great...\nVery Helpful...\nMany Thanks for Sharing....!! <a href=\"/artgor\">@artgor</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 641857,
      "author_name": "robga",
      "author_url": "",
      "post_date": "2019-10-05T08:40:51.090000",
      "content": "<p>disposed old comment</p>",
      "votes": 2,
      "replies": [
        {
          "id": 642229,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2019-10-05T18:18:52.863000",
          "content": "<p>Thanks, I'm glad it is useful :)</p>\n\n<p>Sharing is also useful to me - it helps me to better structure my code and thoughts as well as receive a valuable feedback.</p>\n\n<p>Having separate notebooks for different tasks is also a good idea.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1310824,
      "author_name": "Sau Kha",
      "author_url": "",
      "post_date": "2021-05-17T02:10:16.620000",
      "content": "<p>Thanks for sharing!  Wow! You're amazing!  I went to your GitHub, your own website, your portfolio, your blog and saw your interview.  Such a great accomplishment! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 668283,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-08T08:11:35.530000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 642227,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-10-05T18:15:29.083000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 642230,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2019-10-05T18:19:31.017000",
          "content": "<p>Yes, I saw these config files several times already. I suppose after some time I'll also use something similar.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 642391,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-10-06T01:54:57.343000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 642419,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2019-10-06T03:48:14.317000",
          "content": "<p>I didn't try Catalyst's segmentation models yet, but I think they should also work well.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "641826": "When we take part in kaggle competitions, we usually run a lot of experiments. If it is necessary to manually change the core code for each experiment, we'll spent a lot of unnecessary time, also chances are that we'll make some mistakes.\n\nSo it is more efficient to write a modular code which can be easily reused and which would allow to try various parameters without changing the code. Of course, if you have some completely new idea, it would be necessary to write a code for it. But after this you should be able to use it easily.\n\nSome time ago I wrote this [kernel](https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools) for the competition. It is okay for kaggle, but when I run experiments locally, I prefer to write the code in Pycharm, use Jupyter notebooks for visualization and run the experiments from the command line.\n\nHere is my github repo for this competition: https://github.com/Erlemar/Understanding-Clouds-from-Satellite-Images\n\nThe work is still in progress, but it is already easier to run experiments for me.\n\nFor example, I can run a segmenation model with this command:\n```\npython train.py --encoder resnet50 --bs 20 --num_epochs 100 --train True --optimize_postprocess True --make_prediction True\n```\n\nI'm still thinking about better ways to run experiments - maybe using config files.\n\nConsidering there is currently a competition for useful utility scripts, I have decided to write such a script based on my code, here is it:\nhttps://www.kaggle.com/artgor/pytorch-utils-for-images\n\nI hope it would be useful to you! By the way, it can be easily used for example for severstal competition - you'd only need to modify the `prepare_loaders` function ;)\n\nMy post in a thread with competition for utility scripts: https://www.kaggle.com/general/109651#641828\n\nTo demonstrate the possibilities, I wrote this kernel:\nhttps://www.kaggle.com/artgor/classification-in-catalyst-with-utility-scripts\n\nThese are the main libraries used in my code:\n* [albumentations](https://github.com/albu/albumentations): this is a great library for image augmentation which makes it easier and more convenient\n* [catalyst](https://github.com/catalyst-team/catalyst): this is a great library which makes using PyTorch easier, helps with reprodicibility and contains a lot of useful utils\n* [segmentation_models_pytorch](https://github.com/qubvel/segmentation_models.pytorch): this is a great library with convenient wrappers for models, losses and other useful things\n* [pytorch-toolbelt](https://github.com/BloodAxe/pytorch-toolbelt): this is a great library with many useful shortcuts for building pytorch models\n\nSome examples:\nYou can get loaders like this:\n```\npreprocessing_fn = smp.encoders.get_preprocessing_fn(encoder, encoder_weights)\nloaders = prepare_loaders(path=path, bs=batch_size, num_workers=num_workers, preprocessing_fn=preprocessing_fn, preload=False, task=task, image_size=(224, 224))\n```\n\nThere are two dataset classes - one for segmentation and one for classification. The main difference is in the way the target is created.\n\nSwitching between classification and segmentation tasks and choosing encoders is also simple:\n```\nmodel = get_model(model_type=segm_type, encoder=encoder, encoder_weights=encoder_weights, activation=activation, task=task, n_classes=n_classes, head='simple')\n```",
    "642360": "Amazing work! Thanks for sharing! \nI recently heard that the pipeline offered by QuantumBlack is really useful.",
    "641891": "could be great assets for researchers in this field,thank you @artgor ",
    "641876": "Great...\nVery Helpful...\nMany Thanks for Sharing....!! @artgor ",
    "641857": "disposed old comment",
    "1310824": "Thanks for sharing!  Wow! You're amazing!  I went to your GitHub, your own website, your portfolio, your blog and saw your interview.  Such a great accomplishment! ",
    "668283": "",
    "642227": ""
  }
}