{
  "id": 393969,
  "title": "A few questions regarding disk i/o bottleneck",
  "url": "/competitions/asl-signs/discussion/393969",
  "author_name": "",
  "post_date": "2023-03-11T13:51:31.597077500Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi there, guys.</p>\n<p>Recently, I've started with a baseline model and got inspired by basic models which involve using only mean/std as features (like this <a href=\"https://www.kaggle.com/code/masterofdeception/isolated-sign-language-recognition-with-dnn\" target=\"_blank\">cool notebook</a> by <a href=\"https://www.kaggle.com/masterofdeception\" target=\"_blank\">@masterofdeception</a>) Those models are focused on the preprocessed dataset. It seems clear that such naive features aren't enough to train powerful enough models.</p>\n<p>So, started preparing a full dataset via tf.data.Dataset and original parquet files. It took a few minutes for me to realize that reading data from the original parquet is slow enough, even if using some advanced techniques like ParquetFile described in this discussion. </p>\n<p>So, I drift to TfRecords instead. I've built a .tfrecords file containing the whole dataset without any compression or reduction. In fact, the size of *.tfrecord is three times lower (which made me and y machine with ~40GB disk space left happy :D). But I got disappointed in half an hour, once the very first pipeline GPU utilization drops under 1%. </p>\n<p>I've tried a lot - tf records parallelization, caching, prefetching, and so on. Nothing helped. After some advanced inspection, I realized that the real bottleneck is not a CPU but disk i/o utilization. Even with one thread, it is 100%. With a very basic model, one epoch takes about 5 minutes, while the same model needs about 12 seconds if the whole dataset could fit into RAM.</p>\n<p>So, I am a little bit confused about what strategy Kaggle's guru prefers in such a case. I see a few of variants:</p>\n<ol>\n<li>Reduce a dataset (like what <a href=\"https://www.kaggle.com/lonnieqin\" target=\"_blank\">@lonnieqin</a> did in <a href=\"https://www.kaggle.com/code/lonnieqin/islr-create-tfrecord\" target=\"_blank\">his notebook</a>)</li>\n<li>Start with an inflated model, so no disk /io would not be a bottleneck anymore</li>\n<li>Migrate to TPU. I heard that TPU supports direct loading of the .tfrecords to TPU, thus making it super fast. </li>\n<li>But a few SSD disks and try to distribute data across disks</li>\n</ol>\n<p>I had no intention to create this post, but I've made a simple calculation. With 543 landmarks and 3 coordinates, let's assume that the average sequence length is 300 (btw, does anyone perform EDA toward this?). It turns out that the total number of floats is equal to 488700. Let's do the same calculation for a 480p image: 720<em>480</em>3 = 1036800, which a more than 2 times bigger. I have an intuition that there would not be a problem to load such an ima fast enough. I suspect my SSD disk should be replaced soon (I would check a bottleneck on another machine and will let you know in comments by findings). I also tried to run my notebook in Kaggle's environment, but the result was almost the same in terms of GPU utilization.</p>\n<p>I would love to hear any though on a data pipeline approach in this competition or in general. Thank you in advance</p>",
  "messages": [
    {
      "id": "2177449",
      "postDate": "03/11/2023 13:51:31",
      "content": "<p>Hi there, guys.</p>\n<p>Recently, I've started with a baseline model and got inspired by basic models which involve using only mean/std as features (like this <a href=\"https://www.kaggle.com/code/masterofdeception/isolated-sign-language-recognition-with-dnn\" target=\"_blank\">cool notebook</a> by <a href=\"https://www.kaggle.com/masterofdeception\" target=\"_blank\">@masterofdeception</a>) Those models are focused on the preprocessed dataset. It seems clear that such naive features aren't enough to train powerful enough models.</p>\n<p>So, started preparing a full dataset via tf.data.Dataset and original parquet files. It took a few minutes for me to realize that reading data from the original parquet is slow enough, even if using some advanced techniques like ParquetFile described in this discussion. </p>\n<p>So, I drift to TfRecords instead. I've built a .tfrecords file containing the whole dataset without any compression or reduction. In fact, the size of *.tfrecord is three times lower (which made me and y machine with ~40GB disk space left happy :D). But I got disappointed in half an hour, once the very first pipeline GPU utilization drops under 1%. </p>\n<p>I've tried a lot - tf records parallelization, caching, prefetching, and so on. Nothing helped. After some advanced inspection, I realized that the real bottleneck is not a CPU but disk i/o utilization. Even with one thread, it is 100%. With a very basic model, one epoch takes about 5 minutes, while the same model needs about 12 seconds if the whole dataset could fit into RAM.</p>\n<p>So, I am a little bit confused about what strategy Kaggle's guru prefers in such a case. I see a few of variants:</p>\n<ol>\n<li>Reduce a dataset (like what <a href=\"https://www.kaggle.com/lonnieqin\" target=\"_blank\">@lonnieqin</a> did in <a href=\"https://www.kaggle.com/code/lonnieqin/islr-create-tfrecord\" target=\"_blank\">his notebook</a>)</li>\n<li>Start with an inflated model, so no disk /io would not be a bottleneck anymore</li>\n<li>Migrate to TPU. I heard that TPU supports direct loading of the .tfrecords to TPU, thus making it super fast. </li>\n<li>But a few SSD disks and try to distribute data across disks</li>\n</ol>\n<p>I had no intention to create this post, but I've made a simple calculation. With 543 landmarks and 3 coordinates, let's assume that the average sequence length is 300 (btw, does anyone perform EDA toward this?). It turns out that the total number of floats is equal to 488700. Let's do the same calculation for a 480p image: 720<em>480</em>3 = 1036800, which a more than 2 times bigger. I have an intuition that there would not be a problem to load such an ima fast enough. I suspect my SSD disk should be replaced soon (I would check a bottleneck on another machine and will let you know in comments by findings). I also tried to run my notebook in Kaggle's environment, but the result was almost the same in terms of GPU utilization.</p>\n<p>I would love to hear any though on a data pipeline approach in this competition or in general. Thank you in advance</p>",
      "rawMarkdown": "Hi there, guys.\n\nRecently, I've started with a baseline model and got inspired by basic models which involve using only mean/std as features (like this [cool notebook](https://www.kaggle.com/code/masterofdeception/isolated-sign-language-recognition-with-dnn) by @masterofdeception) Those models are focused on the preprocessed dataset. It seems clear that such naive features aren't enough to train powerful enough models.\n\nSo, started preparing a full dataset via tf.data.Dataset and original parquet files. It took a few minutes for me to realize that reading data from the original parquet is slow enough, even if using some advanced techniques like ParquetFile described in this discussion. \n\nSo, I drift to TfRecords instead. I've built a .tfrecords file containing the whole dataset without any compression or reduction. In fact, the size of *.tfrecord is three times lower (which made me and y machine with ~40GB disk space left happy :D). But I got disappointed in half an hour, once the very first pipeline GPU utilization drops under 1%. \n\nI've tried a lot - tf records parallelization, caching, prefetching, and so on. Nothing helped. After some advanced inspection, I realized that the real bottleneck is not a CPU but disk i/o utilization. Even with one thread, it is 100%. With a very basic model, one epoch takes about 5 minutes, while the same model needs about 12 seconds if the whole dataset could fit into RAM.\n\nSo, I am a little bit confused about what strategy Kaggle's guru prefers in such a case. I see a few of variants:\n\n1. Reduce a dataset (like what @lonnieqin did in [his notebook](https://www.kaggle.com/code/lonnieqin/islr-create-tfrecord))\n2. Start with an inflated model, so no disk /io would not be a bottleneck anymore\n3. Migrate to TPU. I heard that TPU supports direct loading of the .tfrecords to TPU, thus making it super fast. \n4. But a few SSD disks and try to distribute data across disks\n\nI had no intention to create this post, but I've made a simple calculation. With 543 landmarks and 3 coordinates, let's assume that the average sequence length is 300 (btw, does anyone perform EDA toward this?). It turns out that the total number of floats is equal to 488700. Let's do the same calculation for a 480p image: 720*480*3 = 1036800, which a more than 2 times bigger. I have an intuition that there would not be a problem to load such an ima fast enough. I suspect my SSD disk should be replaced soon (I would check a bottleneck on another machine and will let you know in comments by findings). I also tried to run my notebook in Kaggle's environment, but the result was almost the same in terms of GPU utilization.\n\nI would love to hear any though on a data pipeline approach in this competition or in general. Thank you in advance",
      "votes": null
    },
    {
      "id": "2177722",
      "postDate": "03/11/2023 18:07:33",
      "content": "<p>If you're using pre-fetching then I suspect the bottleneck is GPU I/O latency rather than disk I/O.  Have you tried a much larger batch size (which should help to confirm this, even if you don't stick with it)?</p>\n<p>Also, the consensus on the discussion forums is that most of the face points aren't any use to you.  If you keep the hands, the lips and the pose, that reduces 543 landmarks to 115 landmarks (an 80% reduction in size).  Not only does it make training much faster, you'll achieve higher accuracy too because it won't be overfitting to irrelevant features.  You could also consider down/up-sampling the number of frames.  Robert's 0.63 notebook (which was leading at the time he published it) uses a fixed 15 frames per sample.  I haven't seen any studies on the impact of number of frames on model accuracy yet.</p>",
      "rawMarkdown": "If you're using pre-fetching then I suspect the bottleneck is GPU I/O latency rather than disk I/O.  Have you tried a much larger batch size (which should help to confirm this, even if you don't stick with it)?\n\nAlso, the consensus on the discussion forums is that most of the face points aren't any use to you.  If you keep the hands, the lips and the pose, that reduces 543 landmarks to 115 landmarks (an 80% reduction in size).  Not only does it make training much faster, you'll achieve higher accuracy too because it won't be overfitting to irrelevant features.  You could also consider down/up-sampling the number of frames.  Robert's 0.63 notebook (which was leading at the time he published it) uses a fixed 15 frames per sample.  I haven't seen any studies on the impact of number of frames on model accuracy yet.",
      "votes": null
    },
    {
      "id": "2177743",
      "postDate": "03/11/2023 18:32:16",
      "content": "<p>the total data is not very large (e.g. if you are using ,mean shape, using only lip and hands). you can store all in RAM</p>",
      "rawMarkdown": "the total data is not very large (e.g. if you are using ,mean shape, using only lip and hands). you can store all in RAM",
      "votes": null
    },
    {
      "id": "2177767",
      "postDate": "03/11/2023 18:41:22",
      "content": "<p>average number of fame is 38</p>\n<pre><code>mean        37.935021\nstd         44.177069\nmin          2.000000\n25%         12.000000\n50%         22.000000\n75%         44.000000\nmax        537.000000\nName: num_frame, dtype: float64\n</code></pre>\n<pre><code>np.percentile(kaggle_df.num_frame.values,90)\nOut[12]: 92.0\nnp.percentile(kaggle_df.num_frame.values,95)\nOut[13]: 135.0\nnp.percentile(kaggle_df.num_frame.values,98)\nOut[14]: 189.0\nnp.percentile(kaggle_df.num_frame.values,99)\nOut[15]: 219.0\n</code></pre>",
      "rawMarkdown": "average number of fame is 38\n\n```\nmean        37.935021\nstd         44.177069\nmin          2.000000\n25%         12.000000\n50%         22.000000\n75%         44.000000\nmax        537.000000\nName: num_frame, dtype: float64\n\n```\n\n```\nnp.percentile(kaggle_df.num_frame.values,90)\nOut[12]: 92.0\nnp.percentile(kaggle_df.num_frame.values,95)\nOut[13]: 135.0\nnp.percentile(kaggle_df.num_frame.values,98)\nOut[14]: 189.0\nnp.percentile(kaggle_df.num_frame.values,99)\nOut[15]: 219.0\n```",
      "votes": null
    },
    {
      "id": "2178125",
      "postDate": "03/12/2023 06:21:52",
      "content": "<p>Thank you very much!</p>",
      "rawMarkdown": "Thank you very much!",
      "votes": null
    },
    {
      "id": "2178138",
      "postDate": "03/12/2023 06:30:06",
      "content": "<p>Yes, I've tried a relatively big batch size with prefetching (prefetching over batch size, not individual samples). Nothing didn't help. <br>\nThe interesting thing I noticed, is that at the beginning of training (when the prefetching buffer is full) training is very fast, but once the buffer cleared, it dramatically slows down. This gave me the idea that I/O is botlneck. </p>\n<p>I inspected disk i/o throughput and it was 100%, while the CPU and GPU were idling.</p>\n<p>Finally, I solved the mystery. It seems, my SSD is down. I moved to another machine with exactly the same GPU and relatively the same hardware. The same notebook produces about 30% GPU utilization for the first epoch and about 77% for the rest (thanks to tf.data.Dataset.cache())</p>",
      "rawMarkdown": "Yes, I've tried a relatively big batch size with prefetching (prefetching over batch size, not individual samples). Nothing didn't help. \nThe interesting thing I noticed, is that at the beginning of training (when the prefetching buffer is full) training is very fast, but once the buffer cleared, it dramatically slows down. This gave me the idea that I/O is botlneck. \n\nI inspected disk i/o throughput and it was 100%, while the CPU and GPU were idling.\n\nFinally, I solved the mystery. It seems, my SSD is down. I moved to another machine with exactly the same GPU and relatively the same hardware. The same notebook produces about 30% GPU utilization for the first epoch and about 77% for the rest (thanks to tf.data.Dataset.cache())",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2177722,
      "author_name": "andrewrrose",
      "author_url": "",
      "post_date": "03/11/2023 18:07:33",
      "content": "<p>If you're using pre-fetching then I suspect the bottleneck is GPU I/O latency rather than disk I/O.  Have you tried a much larger batch size (which should help to confirm this, even if you don't stick with it)?</p>\n<p>Also, the consensus on the discussion forums is that most of the face points aren't any use to you.  If you keep the hands, the lips and the pose, that reduces 543 landmarks to 115 landmarks (an 80% reduction in size).  Not only does it make training much faster, you'll achieve higher accuracy too because it won't be overfitting to irrelevant features.  You could also consider down/up-sampling the number of frames.  Robert's 0.63 notebook (which was leading at the time he published it) uses a fixed 15 frames per sample.  I haven't seen any studies on the impact of number of frames on model accuracy yet.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2178138,
          "author_name": "meowmeowmeowmeowmeow",
          "author_url": "",
          "post_date": "03/12/2023 06:30:06",
          "content": "<p>Yes, I've tried a relatively big batch size with prefetching (prefetching over batch size, not individual samples). Nothing didn't help. <br>\nThe interesting thing I noticed, is that at the beginning of training (when the prefetching buffer is full) training is very fast, but once the buffer cleared, it dramatically slows down. This gave me the idea that I/O is botlneck. </p>\n<p>I inspected disk i/o throughput and it was 100%, while the CPU and GPU were idling.</p>\n<p>Finally, I solved the mystery. It seems, my SSD is down. I moved to another machine with exactly the same GPU and relatively the same hardware. The same notebook produces about 30% GPU utilization for the first epoch and about 77% for the rest (thanks to tf.data.Dataset.cache())</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2177743,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/11/2023 18:32:16",
      "content": "<p>the total data is not very large (e.g. if you are using ,mean shape, using only lip and hands). you can store all in RAM</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2177767,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/11/2023 18:41:22",
      "content": "<p>average number of fame is 38</p>\n<pre><code>mean        37.935021\nstd         44.177069\nmin          2.000000\n25%         12.000000\n50%         22.000000\n75%         44.000000\nmax        537.000000\nName: num_frame, dtype: float64\n</code></pre>\n<pre><code>np.percentile(kaggle_df.num_frame.values,90)\nOut[12]: 92.0\nnp.percentile(kaggle_df.num_frame.values,95)\nOut[13]: 135.0\nnp.percentile(kaggle_df.num_frame.values,98)\nOut[14]: 189.0\nnp.percentile(kaggle_df.num_frame.values,99)\nOut[15]: 219.0\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2178125,
          "author_name": "meowmeowmeowmeowmeow",
          "author_url": "",
          "post_date": "03/12/2023 06:21:52",
          "content": "<p>Thank you very much!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2177449": "Hi there, guys.\n\nRecently, I've started with a baseline model and got inspired by basic models which involve using only mean/std as features (like this [cool notebook](https://www.kaggle.com/code/masterofdeception/isolated-sign-language-recognition-with-dnn) by @masterofdeception) Those models are focused on the preprocessed dataset. It seems clear that such naive features aren't enough to train powerful enough models.\n\nSo, started preparing a full dataset via tf.data.Dataset and original parquet files. It took a few minutes for me to realize that reading data from the original parquet is slow enough, even if using some advanced techniques like ParquetFile described in this discussion. \n\nSo, I drift to TfRecords instead. I've built a .tfrecords file containing the whole dataset without any compression or reduction. In fact, the size of *.tfrecord is three times lower (which made me and y machine with ~40GB disk space left happy :D). But I got disappointed in half an hour, once the very first pipeline GPU utilization drops under 1%. \n\nI've tried a lot - tf records parallelization, caching, prefetching, and so on. Nothing helped. After some advanced inspection, I realized that the real bottleneck is not a CPU but disk i/o utilization. Even with one thread, it is 100%. With a very basic model, one epoch takes about 5 minutes, while the same model needs about 12 seconds if the whole dataset could fit into RAM.\n\nSo, I am a little bit confused about what strategy Kaggle's guru prefers in such a case. I see a few of variants:\n\n1. Reduce a dataset (like what @lonnieqin did in [his notebook](https://www.kaggle.com/code/lonnieqin/islr-create-tfrecord))\n2. Start with an inflated model, so no disk /io would not be a bottleneck anymore\n3. Migrate to TPU. I heard that TPU supports direct loading of the .tfrecords to TPU, thus making it super fast. \n4. But a few SSD disks and try to distribute data across disks\n\nI had no intention to create this post, but I've made a simple calculation. With 543 landmarks and 3 coordinates, let's assume that the average sequence length is 300 (btw, does anyone perform EDA toward this?). It turns out that the total number of floats is equal to 488700. Let's do the same calculation for a 480p image: 720*480*3 = 1036800, which a more than 2 times bigger. I have an intuition that there would not be a problem to load such an ima fast enough. I suspect my SSD disk should be replaced soon (I would check a bottleneck on another machine and will let you know in comments by findings). I also tried to run my notebook in Kaggle's environment, but the result was almost the same in terms of GPU utilization.\n\nI would love to hear any though on a data pipeline approach in this competition or in general. Thank you in advance",
    "2177722": "If you're using pre-fetching then I suspect the bottleneck is GPU I/O latency rather than disk I/O.  Have you tried a much larger batch size (which should help to confirm this, even if you don't stick with it)?\n\nAlso, the consensus on the discussion forums is that most of the face points aren't any use to you.  If you keep the hands, the lips and the pose, that reduces 543 landmarks to 115 landmarks (an 80% reduction in size).  Not only does it make training much faster, you'll achieve higher accuracy too because it won't be overfitting to irrelevant features.  You could also consider down/up-sampling the number of frames.  Robert's 0.63 notebook (which was leading at the time he published it) uses a fixed 15 frames per sample.  I haven't seen any studies on the impact of number of frames on model accuracy yet.",
    "2177743": "the total data is not very large (e.g. if you are using ,mean shape, using only lip and hands). you can store all in RAM",
    "2177767": "average number of fame is 38\n\n```\nmean        37.935021\nstd         44.177069\nmin          2.000000\n25%         12.000000\n50%         22.000000\n75%         44.000000\nmax        537.000000\nName: num_frame, dtype: float64\n\n```\n\n```\nnp.percentile(kaggle_df.num_frame.values,90)\nOut[12]: 92.0\nnp.percentile(kaggle_df.num_frame.values,95)\nOut[13]: 135.0\nnp.percentile(kaggle_df.num_frame.values,98)\nOut[14]: 189.0\nnp.percentile(kaggle_df.num_frame.values,99)\nOut[15]: 219.0\n```",
    "2178125": "Thank you very much!",
    "2178138": "Yes, I've tried a relatively big batch size with prefetching (prefetching over batch size, not individual samples). Nothing didn't help. \nThe interesting thing I noticed, is that at the beginning of training (when the prefetching buffer is full) training is very fast, but once the buffer cleared, it dramatically slows down. This gave me the idea that I/O is botlneck. \n\nI inspected disk i/o throughput and it was 100%, while the CPU and GPU were idling.\n\nFinally, I solved the mystery. It seems, my SSD is down. I moved to another machine with exactly the same GPU and relatively the same hardware. The same notebook produces about 30% GPU utilization for the first epoch and about 77% for the rest (thanks to tf.data.Dataset.cache())"
  },
  "source": "meta"
}