{
  "id": 126054,
  "title": "Optimization tips, eliminate Kaggle failures",
  "url": "/competitions/bengaliai-cv19/discussion/126054",
  "author_name": "",
  "post_date": "2020-01-15T11:38:59.459257600Z",
  "votes": 44,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I saw lots of topics/comments about submission failures. Most of the cases are related to memory and time limit issues. I gather a few optimization tips and tricks. I hope these tips help you to solve your issues.</p>\n\n<h3>Explanation notebook</h3>\n\n<p><a href=\"https://www.kaggle.com/pestipeti/optimization-tips-faster-train-faster-inference\">https://www.kaggle.com/pestipeti/optimization-tips-faster-train-faster-inference</a></p>\n\n<h3>Implementation notebook</h3>\n\n<p><a href=\"https://www.kaggle.com/pestipeti/fast-ensemble-5-folds-20-minutes\">https://www.kaggle.com/pestipeti/fast-ensemble-5-folds-20-minutes</a></p>\n\n<h3>Summary</h3>\n\n<ul>\n<li>Use script instead of notebook</li>\n<li>Import only things you need</li>\n<li>Use logs instead of tqdm</li>\n<li>Cleanup after usage</li>\n<li>Load parquet files once</li>\n<li>Do not load data you don't need</li>\n<li>Check your dtypes</li>\n<li>Preprocess your images once</li>\n<li>Use CUDA for preprocessing</li>\n<li>Do not use albumentation (inference)</li>\n<li>Only use 3 channels if you really need it</li>\n<li>Process in batches</li>\n<li>Optimized TTA</li>\n</ul>",
  "messages": [
    {
      "id": "719331",
      "postDate": "01/15/2020 11:38:59",
      "content": "<p>I saw lots of topics/comments about submission failures. Most of the cases are related to memory and time limit issues. I gather a few optimization tips and tricks. I hope these tips help you to solve your issues.</p>\n\n<h3>Explanation notebook</h3>\n\n<p><a href=\"https://www.kaggle.com/pestipeti/optimization-tips-faster-train-faster-inference\">https://www.kaggle.com/pestipeti/optimization-tips-faster-train-faster-inference</a></p>\n\n<h3>Implementation notebook</h3>\n\n<p><a href=\"https://www.kaggle.com/pestipeti/fast-ensemble-5-folds-20-minutes\">https://www.kaggle.com/pestipeti/fast-ensemble-5-folds-20-minutes</a></p>\n\n<h3>Summary</h3>\n\n<ul>\n<li>Use script instead of notebook</li>\n<li>Import only things you need</li>\n<li>Use logs instead of tqdm</li>\n<li>Cleanup after usage</li>\n<li>Load parquet files once</li>\n<li>Do not load data you don't need</li>\n<li>Check your dtypes</li>\n<li>Preprocess your images once</li>\n<li>Use CUDA for preprocessing</li>\n<li>Do not use albumentation (inference)</li>\n<li>Only use 3 channels if you really need it</li>\n<li>Process in batches</li>\n<li>Optimized TTA</li>\n</ul>",
      "rawMarkdown": "I saw lots of topics/comments about submission failures. Most of the cases are related to memory and time limit issues. I gather a few optimization tips and tricks. I hope these tips help you to solve your issues.\n\n\n### Explanation notebook\n[https://www.kaggle.com/pestipeti/optimization-tips-faster-train-faster-inference](https://www.kaggle.com/pestipeti/optimization-tips-faster-train-faster-inference)\n\n### Implementation notebook\n[https://www.kaggle.com/pestipeti/fast-ensemble-5-folds-20-minutes](https://www.kaggle.com/pestipeti/fast-ensemble-5-folds-20-minutes)\n\n### Summary\n- Use script instead of notebook\n- Import only things you need\n- Use logs instead of tqdm\n- Cleanup after usage\n- Load parquet files once\n- Do not load data you don't need\n- Check your dtypes\n- Preprocess your images once\n- Use CUDA for preprocessing\n- Do not use albumentation (inference)\n- Only use 3 channels if you really need it\n- Process in batches\n- Optimized TTA",
      "votes": null
    },
    {
      "id": "719339",
      "postDate": "01/15/2020 11:53:00",
      "content": "<p>Or: don't write your own submission script, clone someone else, which works for sure and then just replace model and pre/post processing for yours)</p>",
      "rawMarkdown": "Or: don't write your own submission script, clone someone else, which works for sure and then just replace model and pre/post processing for yours)",
      "votes": null
    },
    {
      "id": "720176",
      "postDate": "01/16/2020 08:14:12",
      "content": "<p>\"Load parquet files once\" makes me OOM(Out of Memory). I guess it would be much better if dataset is provided by image files and corresponding csv file.</p>",
      "rawMarkdown": "\"Load parquet files once\" makes me OOM(Out of Memory). I guess it would be much better if dataset is provided by image files and corresponding csv file.",
      "votes": null
    },
    {
      "id": "720188",
      "postDate": "01/16/2020 08:22:19",
      "content": "<p>Try to store one channel only (uint8) and convert it to your format in your train/inference loop.</p>",
      "rawMarkdown": "Try to store one channel only (uint8) and convert it to your format in your train/inference loop.",
      "votes": null
    },
    {
      "id": "721769",
      "postDate": "01/17/2020 17:13:10",
      "content": "<p><a href=\"/pestipeti\">@pestipeti</a> thanks for sharing. This is very useful.</p>",
      "rawMarkdown": "pestipeti thanks for sharing. This is very useful.",
      "votes": null
    },
    {
      "id": "722781",
      "postDate": "01/19/2020 05:23:56",
      "content": "<p>Hi,\n were you able to implement TTA ? If yes, can you please advice how to do that with Multiout</p>",
      "rawMarkdown": "Hi,\n were you able to implement TTA ? If yes, can you please advice how to do that with Multiout",
      "votes": null
    },
    {
      "id": "750984",
      "postDate": "02/19/2020 22:16:07",
      "content": "<p>Can you provide some details/references for these two below? Thanks!\n- Use CUDA for preprocessing\n- Optimized TTA</p>",
      "rawMarkdown": "Can you provide some details/references for these two below? Thanks!\n- Use CUDA for preprocessing\n- Optimized TTA",
      "votes": null
    },
    {
      "id": "751017",
      "postDate": "02/19/2020 23:43:01",
      "content": "<p><a href=\"/zhangyang\">@zhangyang</a> It turned out I can do a 5-fold ensemble without those two tricks, so I haven't implemented it.</p>",
      "rawMarkdown": "zhangyang It turned out I can do a 5-fold ensemble without those two tricks, so I haven't implemented it.",
      "votes": null
    },
    {
      "id": "751046",
      "postDate": "02/20/2020 00:42:38",
      "content": "<p>During inference, you are feeding images into your GPU models. If you need to crop and resize test images (and/or other TTA) beforehand it may be that the CPU preprocess is slower than the GPU inference. In that case, you can speed up your inference by doing preprocess (and/or TTA) on GPU instead of CPU.</p>\n\n<p>If you must preprocess on CPU, remember that Kaggle GPU vm has 2 core CPU so you should use at least 2 threads.</p>",
      "rawMarkdown": "During inference, you are feeding images into your GPU models. If you need to crop and resize test images (and/or other TTA) beforehand it may be that the CPU preprocess is slower than the GPU inference. In that case, you can speed up your inference by doing preprocess (and/or TTA) on GPU instead of CPU.\n\nIf you must preprocess on CPU, remember that Kaggle GPU vm has 2 core CPU so you should use at least 2 threads.",
      "votes": null
    },
    {
      "id": "751130",
      "postDate": "02/20/2020 02:35:01",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "761503",
      "postDate": "03/02/2020 15:49:14",
      "content": "<p>i would like to know the clustering analysis on mixed data type in python</p>",
      "rawMarkdown": "i would like to know the clustering analysis on mixed data type in python",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 719339,
      "author_name": "oldufo",
      "author_url": "",
      "post_date": "01/15/2020 11:53:00",
      "content": "<p>Or: don't write your own submission script, clone someone else, which works for sure and then just replace model and pre/post processing for yours)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 720176,
      "author_name": "ildoonet",
      "author_url": "",
      "post_date": "01/16/2020 08:14:12",
      "content": "<p>\"Load parquet files once\" makes me OOM(Out of Memory). I guess it would be much better if dataset is provided by image files and corresponding csv file.</p>",
      "votes": null,
      "replies": [
        {
          "id": 720188,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "01/16/2020 08:22:19",
          "content": "<p>Try to store one channel only (uint8) and convert it to your format in your train/inference loop.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 721769,
      "author_name": "sabinhashmi",
      "author_url": "",
      "post_date": "01/17/2020 17:13:10",
      "content": "<p><a href=\"/pestipeti\">@pestipeti</a> thanks for sharing. This is very useful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 722781,
      "author_name": "mayank17",
      "author_url": "",
      "post_date": "01/19/2020 05:23:56",
      "content": "<p>Hi,\n were you able to implement TTA ? If yes, can you please advice how to do that with Multiout</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 750984,
      "author_name": "zhangyang",
      "author_url": "",
      "post_date": "02/19/2020 22:16:07",
      "content": "<p>Can you provide some details/references for these two below? Thanks!\n- Use CUDA for preprocessing\n- Optimized TTA</p>",
      "votes": null,
      "replies": [
        {
          "id": 751017,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "02/19/2020 23:43:01",
          "content": "<p><a href=\"/zhangyang\">@zhangyang</a> It turned out I can do a 5-fold ensemble without those two tricks, so I haven't implemented it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 751046,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/20/2020 00:42:38",
          "content": "<p>During inference, you are feeding images into your GPU models. If you need to crop and resize test images (and/or other TTA) beforehand it may be that the CPU preprocess is slower than the GPU inference. In that case, you can speed up your inference by doing preprocess (and/or TTA) on GPU instead of CPU.</p>\n\n<p>If you must preprocess on CPU, remember that Kaggle GPU vm has 2 core CPU so you should use at least 2 threads.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 751130,
          "author_name": "zhangyang",
          "author_url": "",
          "post_date": "02/20/2020 02:35:01",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 761503,
      "author_name": "sumathikanna",
      "author_url": "",
      "post_date": "03/02/2020 15:49:14",
      "content": "<p>i would like to know the clustering analysis on mixed data type in python</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "719331": "I saw lots of topics/comments about submission failures. Most of the cases are related to memory and time limit issues. I gather a few optimization tips and tricks. I hope these tips help you to solve your issues.\n\n\n### Explanation notebook\n[https://www.kaggle.com/pestipeti/optimization-tips-faster-train-faster-inference](https://www.kaggle.com/pestipeti/optimization-tips-faster-train-faster-inference)\n\n### Implementation notebook\n[https://www.kaggle.com/pestipeti/fast-ensemble-5-folds-20-minutes](https://www.kaggle.com/pestipeti/fast-ensemble-5-folds-20-minutes)\n\n### Summary\n- Use script instead of notebook\n- Import only things you need\n- Use logs instead of tqdm\n- Cleanup after usage\n- Load parquet files once\n- Do not load data you don't need\n- Check your dtypes\n- Preprocess your images once\n- Use CUDA for preprocessing\n- Do not use albumentation (inference)\n- Only use 3 channels if you really need it\n- Process in batches\n- Optimized TTA",
    "719339": "Or: don't write your own submission script, clone someone else, which works for sure and then just replace model and pre/post processing for yours)",
    "720176": "\"Load parquet files once\" makes me OOM(Out of Memory). I guess it would be much better if dataset is provided by image files and corresponding csv file.",
    "720188": "Try to store one channel only (uint8) and convert it to your format in your train/inference loop.",
    "721769": "pestipeti thanks for sharing. This is very useful.",
    "722781": "Hi,\n were you able to implement TTA ? If yes, can you please advice how to do that with Multiout",
    "750984": "Can you provide some details/references for these two below? Thanks!\n- Use CUDA for preprocessing\n- Optimized TTA",
    "751017": "zhangyang It turned out I can do a 5-fold ensemble without those two tricks, so I haven't implemented it.",
    "751046": "During inference, you are feeding images into your GPU models. If you need to crop and resize test images (and/or other TTA) beforehand it may be that the CPU preprocess is slower than the GPU inference. In that case, you can speed up your inference by doing preprocess (and/or TTA) on GPU instead of CPU.\n\nIf you must preprocess on CPU, remember that Kaggle GPU vm has 2 core CPU so you should use at least 2 threads.",
    "751130": "Thank you!",
    "761503": "i would like to know the clustering analysis on mixed data type in python"
  },
  "source": "meta"
}