{
  "id": 375718,
  "title": "Overcome the bottleneck?",
  "url": "/competitions/nfl-player-contact-detection/discussion/375718",
  "author_name": "",
  "post_date": "2023-01-03T03:17:26.993568200Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I tried to write my trainining pipeline similar to this <a href=\"https://www.kaggle.com/code/zzy990106/nfl-2-5d-cnn-baseline-inference\" target=\"_blank\">notebook</a>, created by <a href=\"https://www.kaggle.com/zzy990106\" target=\"_blank\">@zzy990106 </a>, you can have a look in the <code>Data object</code><br>\nBut as the bottleneck in the process of the cropping image step, I tried to dump all cropped image at once time, and then in the training I can load it back again.<br>\nThis still raises a problem, it's too slow. I tried to using for loop and joblib to iter through all instance in the training dataset, this will cost me <strong>45 days on a machine with 32 cores and 60Gb of RAM</strong>.<br>\nIs there anything I can do to solve this problem, because the bottleneck in the training pipeline so the idea of dumping all the cropped images is all I can do.</p>",
  "messages": [
    {
      "id": "2083969",
      "postDate": "01/03/2023 03:17:26",
      "content": "<p>I tried to write my trainining pipeline similar to this <a href=\"https://www.kaggle.com/code/zzy990106/nfl-2-5d-cnn-baseline-inference\" target=\"_blank\">notebook</a>, created by <a href=\"https://www.kaggle.com/zzy990106\" target=\"_blank\">@zzy990106 </a>, you can have a look in the <code>Data object</code><br>\nBut as the bottleneck in the process of the cropping image step, I tried to dump all cropped image at once time, and then in the training I can load it back again.<br>\nThis still raises a problem, it's too slow. I tried to using for loop and joblib to iter through all instance in the training dataset, this will cost me <strong>45 days on a machine with 32 cores and 60Gb of RAM</strong>.<br>\nIs there anything I can do to solve this problem, because the bottleneck in the training pipeline so the idea of dumping all the cropped images is all I can do.</p>",
      "rawMarkdown": "I tried to write my trainining pipeline similar to this [notebook](https://www.kaggle.com/code/zzy990106/nfl-2-5d-cnn-baseline-inference), created by [@zzy990106 ](https://www.kaggle.com/zzy990106 ), you can have a look in the `Data object`\nBut as the bottleneck in the process of the cropping image step, I tried to dump all cropped image at once time, and then in the training I can load it back again.\nThis still raises a problem, it's too slow. I tried to using for loop and joblib to iter through all instance in the training dataset, this will cost me **45 days on a machine with 32 cores and 60Gb of RAM**.\nIs there anything I can do to solve this problem, because the bottleneck in the training pipeline so the idea of dumping all the cropped images is all I can do.",
      "votes": null
    },
    {
      "id": "2083982",
      "postDate": "01/03/2023 03:33:12",
      "content": "<ol>\n<li>The bottleneck is reading images, not cropping. If you save them, you still need to read a lot.</li>\n<li>This is just a get-started idea. I suggest you change to a better method, which may be based on the whole image.</li>\n</ol>",
      "rawMarkdown": "1. The bottleneck is reading images, not cropping. If you save them, you still need to read a lot.\n2. This is just a get-started idea. I suggest you change to a better method, which may be based on the whole image.",
      "votes": null
    },
    {
      "id": "2083989",
      "postDate": "01/03/2023 03:50:36",
      "content": "<p>Thanks for your commend, I trying to optimize and reproduce your score. I think I'll try another method.</p>",
      "rawMarkdown": "Thanks for your commend, I trying to optimize and reproduce your score. I think I'll try another method.",
      "votes": null
    },
    {
      "id": "2087203",
      "postDate": "01/05/2023 12:16:38",
      "content": "<p>There are several strategies you can try to overcome the bottleneck in your training pipeline:</p>\n<p>Use a faster hardware setup: One way to improve the performance of your training pipeline is to use faster hardware, such as a more powerful CPU or a GPU. This can help to reduce the time it takes to crop and preprocess the images.</p>\n<p>Optimize your code: Another way to improve the performance of your training pipeline is to optimize your code. This may involve optimizing the algorithms you are using, reducing the number of unnecessary computations, and minimizing the use of memory-intensive operations.</p>\n<p>Use data augmentation: Data augmentation is a technique that involves generating new training data by applying transformations to existing data. This can help to increase the size of your training dataset, and may also improve the generalization performance of your model.</p>\n<p>Use a distributed training approach: If you are using a machine with multiple cores or GPUs, you can use a distributed training approach to parallelize the training process. This can help to reduce the training time by distributing the workload across multiple cores or GPUs.</p>\n<p>Use a pre-trained model: Another option is to use a pre-trained model as a starting point, and fine-tune it on your dataset. This can help to reduce the amount of training data required, and may also improve the performance of your model.</p>",
      "rawMarkdown": "There are several strategies you can try to overcome the bottleneck in your training pipeline:\n\nUse a faster hardware setup: One way to improve the performance of your training pipeline is to use faster hardware, such as a more powerful CPU or a GPU. This can help to reduce the time it takes to crop and preprocess the images.\n\nOptimize your code: Another way to improve the performance of your training pipeline is to optimize your code. This may involve optimizing the algorithms you are using, reducing the number of unnecessary computations, and minimizing the use of memory-intensive operations.\n\nUse data augmentation: Data augmentation is a technique that involves generating new training data by applying transformations to existing data. This can help to increase the size of your training dataset, and may also improve the generalization performance of your model.\n\nUse a distributed training approach: If you are using a machine with multiple cores or GPUs, you can use a distributed training approach to parallelize the training process. This can help to reduce the training time by distributing the workload across multiple cores or GPUs.\n\nUse a pre-trained model: Another option is to use a pre-trained model as a starting point, and fine-tune it on your dataset. This can help to reduce the amount of training data required, and may also improve the performance of your model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2083982,
      "author_name": "zzy990106",
      "author_url": "",
      "post_date": "01/03/2023 03:33:12",
      "content": "<ol>\n<li>The bottleneck is reading images, not cropping. If you save them, you still need to read a lot.</li>\n<li>This is just a get-started idea. I suggest you change to a better method, which may be based on the whole image.</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 2083989,
          "author_name": "locbaop",
          "author_url": "",
          "post_date": "01/03/2023 03:50:36",
          "content": "<p>Thanks for your commend, I trying to optimize and reproduce your score. I think I'll try another method.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2087203,
      "author_name": "aviralmishra1998",
      "author_url": "",
      "post_date": "01/05/2023 12:16:38",
      "content": "<p>There are several strategies you can try to overcome the bottleneck in your training pipeline:</p>\n<p>Use a faster hardware setup: One way to improve the performance of your training pipeline is to use faster hardware, such as a more powerful CPU or a GPU. This can help to reduce the time it takes to crop and preprocess the images.</p>\n<p>Optimize your code: Another way to improve the performance of your training pipeline is to optimize your code. This may involve optimizing the algorithms you are using, reducing the number of unnecessary computations, and minimizing the use of memory-intensive operations.</p>\n<p>Use data augmentation: Data augmentation is a technique that involves generating new training data by applying transformations to existing data. This can help to increase the size of your training dataset, and may also improve the generalization performance of your model.</p>\n<p>Use a distributed training approach: If you are using a machine with multiple cores or GPUs, you can use a distributed training approach to parallelize the training process. This can help to reduce the training time by distributing the workload across multiple cores or GPUs.</p>\n<p>Use a pre-trained model: Another option is to use a pre-trained model as a starting point, and fine-tune it on your dataset. This can help to reduce the amount of training data required, and may also improve the performance of your model.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2083969": "I tried to write my trainining pipeline similar to this [notebook](https://www.kaggle.com/code/zzy990106/nfl-2-5d-cnn-baseline-inference), created by [@zzy990106 ](https://www.kaggle.com/zzy990106 ), you can have a look in the `Data object`\nBut as the bottleneck in the process of the cropping image step, I tried to dump all cropped image at once time, and then in the training I can load it back again.\nThis still raises a problem, it's too slow. I tried to using for loop and joblib to iter through all instance in the training dataset, this will cost me **45 days on a machine with 32 cores and 60Gb of RAM**.\nIs there anything I can do to solve this problem, because the bottleneck in the training pipeline so the idea of dumping all the cropped images is all I can do.",
    "2083982": "1. The bottleneck is reading images, not cropping. If you save them, you still need to read a lot.\n2. This is just a get-started idea. I suggest you change to a better method, which may be based on the whole image.",
    "2083989": "Thanks for your commend, I trying to optimize and reproduce your score. I think I'll try another method.",
    "2087203": "There are several strategies you can try to overcome the bottleneck in your training pipeline:\n\nUse a faster hardware setup: One way to improve the performance of your training pipeline is to use faster hardware, such as a more powerful CPU or a GPU. This can help to reduce the time it takes to crop and preprocess the images.\n\nOptimize your code: Another way to improve the performance of your training pipeline is to optimize your code. This may involve optimizing the algorithms you are using, reducing the number of unnecessary computations, and minimizing the use of memory-intensive operations.\n\nUse data augmentation: Data augmentation is a technique that involves generating new training data by applying transformations to existing data. This can help to increase the size of your training dataset, and may also improve the generalization performance of your model.\n\nUse a distributed training approach: If you are using a machine with multiple cores or GPUs, you can use a distributed training approach to parallelize the training process. This can help to reduce the training time by distributing the workload across multiple cores or GPUs.\n\nUse a pre-trained model: Another option is to use a pre-trained model as a starting point, and fine-tune it on your dataset. This can help to reduce the amount of training data required, and may also improve the performance of your model."
  },
  "source": "meta"
}