{
  "id": 392279,
  "title": "How to load multi-frame images efficiently?",
  "url": "/competitions/nfl-player-contact-detection/discussion/392279",
  "author_name": "",
  "post_date": "2023-03-04T14:07:09.238429Z",
  "votes": 12,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Many top solutions have been shared and I'm learning from each team's unique approach. Thank you.</p>\n<p>According to the top solutions, one key was to input many multi-frame images to 2.5D and 3D CNN, but the image loading speed is the bottleneck to do so.</p>\n<p>A naive approach would be to use OpenCV or TurboJPEG (a bit faster) to load images one by one, as <a href=\"https://www.kaggle.com/code/zzy990106/nfl-2-5d-cnn-baseline-inference\" target=\"_blank\">PublicNotebook</a> does, but this would take too much learning time.</p>\n<p>What are some tips for efficient learning? Do you use <a href=\"https://github.com/dmlc/decord\" target=\"_blank\">decord</a> or something similar to process video by video?</p>\n<p>I would like to learn for the future, If you don't mind could you tell me.</p>\n<p>Best regards.</p>",
  "messages": [
    {
      "id": "2168779",
      "postDate": "03/04/2023 14:07:09",
      "content": "<p>Many top solutions have been shared and I'm learning from each team's unique approach. Thank you.</p>\n<p>According to the top solutions, one key was to input many multi-frame images to 2.5D and 3D CNN, but the image loading speed is the bottleneck to do so.</p>\n<p>A naive approach would be to use OpenCV or TurboJPEG (a bit faster) to load images one by one, as <a href=\"https://www.kaggle.com/code/zzy990106/nfl-2-5d-cnn-baseline-inference\" target=\"_blank\">PublicNotebook</a> does, but this would take too much learning time.</p>\n<p>What are some tips for efficient learning? Do you use <a href=\"https://github.com/dmlc/decord\" target=\"_blank\">decord</a> or something similar to process video by video?</p>\n<p>I would like to learn for the future, If you don't mind could you tell me.</p>\n<p>Best regards.</p>",
      "rawMarkdown": "Many top solutions have been shared and I'm learning from each team's unique approach. Thank you.\n\nAccording to the top solutions, one key was to input many multi-frame images to 2.5D and 3D CNN, but the image loading speed is the bottleneck to do so.\n\nA naive approach would be to use OpenCV or TurboJPEG (a bit faster) to load images one by one, as [PublicNotebook](https://www.kaggle.com/code/zzy990106/nfl-2-5d-cnn-baseline-inference) does, but this would take too much learning time.\n\nWhat are some tips for efficient learning? Do you use [decord](https://github.com/dmlc/decord) or something similar to process video by video?\n\nI would like to learn for the future, If you don't mind could you tell me.\n\nBest regards.",
      "votes": null
    },
    {
      "id": "2168845",
      "postDate": "03/04/2023 15:19:00",
      "content": "<p>Files load faster in npy format than in jpeg.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1377145%2Fc13ec936cb918fddec0983b8079d013d%2FScreenshot%20from%202023-03-05%2000-12-58.png?generation=1677943049469449&amp;alt=media\" alt=\"\"><br>\nHowever, npy has a very large file size compared to jpeg.</p>\n<p>I used to cache images in memory for learning and inference whenever possible.</p>",
      "rawMarkdown": "Files load faster in npy format than in jpeg.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1377145%2Fc13ec936cb918fddec0983b8079d013d%2FScreenshot%20from%202023-03-05%2000-12-58.png?generation=1677943049469449&alt=media)\nHowever, npy has a very large file size compared to jpeg.\n\nI used to cache images in memory for learning and inference whenever possible.",
      "votes": null
    },
    {
      "id": "2168849",
      "postDate": "03/04/2023 15:23:57",
      "content": "<p>Our team used multi-frame images and inferred them all at once.<br>\nThis speeded up training and inference time and improved scores with utilizing time-series information.<br>\nFYI, <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392290\" target=\"_blank\">our solution</a>. </p>",
      "rawMarkdown": "Our team used multi-frame images and inferred them all at once.\nThis speeded up training and inference time and improved scores with utilizing time-series information.\nFYI, [our solution](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392290).",
      "votes": null
    },
    {
      "id": "2168905",
      "postDate": "03/04/2023 16:12:05",
      "content": "<p>I’m also interested in this topic because we didn't have much time to optimize the pipeline for sequential inputs.</p>\n<p>As the tentative solution, our team used lycon instead of cv2. This modification reduced 20% of data preprocessing time.<br>\ngithub: <a href=\"https://github.com/ethereon/lycon\" target=\"_blank\">https://github.com/ethereon/lycon</a></p>\n<p>But augmentation process (albumentation) is also heavy. So maybe we should consider this part also, e.g. using kornia.</p>",
      "rawMarkdown": "I’m also interested in this topic because we didn't have much time to optimize the pipeline for sequential inputs.\n\nAs the tentative solution, our team used lycon instead of cv2. This modification reduced 20% of data preprocessing time.\ngithub: https://github.com/ethereon/lycon\n\nBut augmentation process (albumentation) is also heavy. So maybe we should consider this part also, e.g. using kornia.",
      "votes": null
    },
    {
      "id": "2169293",
      "postDate": "03/05/2023 02:02:53",
      "content": "<p>During inference, I used the <a href=\"https://www.kaggle.com/lru\" target=\"_blank\">@lru</a>_cache() decorator to keep quite a large number of frames in RAM since the same frame is used multiple times for different players or steps when predicting sequentially.</p>\n<p>During training I calculated the bounding box for the target crop (after all geometry augmentations) and loaded only the necessary crop, something like:</p>\n<pre><code>img = PIL.Image(fn) # data is not decoded at this stage yet\nimg_crop = np.array(img.crop((left, upper, right, lower)))\n</code></pre>\n<p>This was a few times faster compared to loading the full resolution image with opencv.</p>",
      "rawMarkdown": "During inference, I used the @lru_cache() decorator to keep quite a large number of frames in RAM since the same frame is used multiple times for different players or steps when predicting sequentially.\n\nDuring training I calculated the bounding box for the target crop (after all geometry augmentations) and loaded only the necessary crop, something like:\n\n```\nimg = PIL.Image(fn) # data is not decoded at this stage yet\nimg_crop = np.array(img.crop((left, upper, right, lower)))\n```\n\nThis was a few times faster compared to loading the full resolution image with opencv.",
      "votes": null
    },
    {
      "id": "2169470",
      "postDate": "03/05/2023 06:26:17",
      "content": "<p>Thanks for letting me know!<br>\nnpy is indeed fast, but I had given up because of the lack of memory for the total number of frames…<br>\nAs others have said, caching seems to be effective. I'll look into the implementation!</p>",
      "rawMarkdown": "Thanks for letting me know!\nnpy is indeed fast, but I had given up because of the lack of memory for the total number of frames...\nAs others have said, caching seems to be effective. I'll look into the implementation!",
      "votes": null
    },
    {
      "id": "2169475",
      "postDate": "03/05/2023 06:35:49",
      "content": "<p>Thanks for letting me know !<br>\nI learned a lot from reading your solution.<br>\nI now know that not only TurboJPEG but also cache by lru_cache is important.</p>",
      "rawMarkdown": "Thanks for letting me know !\nI learned a lot from reading your solution.\nI now know that not only TurboJPEG but also cache by lru_cache is important.",
      "votes": null
    },
    {
      "id": "2169488",
      "postDate": "03/05/2023 06:45:32",
      "content": "<p>Thanks for letting me know !<br>\nI didn’t know about the lycon library.<br>\nAs you mentioned, augmentation also has a significant impact on learning speed. Our team used TorchVision as an alternative.</p>",
      "rawMarkdown": "Thanks for letting me know !\nI didn’t know about the lycon library.\nAs you mentioned, augmentation also has a significant impact on learning speed. Our team used TorchVision as an alternative.",
      "votes": null
    },
    {
      "id": "2169493",
      "postDate": "03/05/2023 06:55:27",
      "content": "<p>Thanks for letting me know !<br>\nIt seems that lru_cache is an important feature, as some participants have responded to me about it.<br>\nAlso, thank you for sharing the implementation of the PIL. This is the first time I learned that it is possible to crop before decoding.</p>\n<p>I would like to check for correct understanding, but the PIL.Image class does not seem to be able to read the file as it is, is the following implementation correct?</p>\n<pre><code> PIL  Image\nimg = Image.()\nimg_crop = np.array(img.crop((left, upper, right, lower)))\n</code></pre>",
      "rawMarkdown": "Thanks for letting me know !\nIt seems that lru_cache is an important feature, as some participants have responded to me about it.\nAlso, thank you for sharing the implementation of the PIL. This is the first time I learned that it is possible to crop before decoding.\n\nI would like to check for correct understanding, but the PIL.Image class does not seem to be able to read the file as it is, is the following implementation correct?\n```python\nfrom PIL import Image\nimg = Image.open('xxx.jpeg')\nimg_crop = np.array(img.crop((left, upper, right, lower)))\n```",
      "votes": null
    },
    {
      "id": "2169685",
      "postDate": "03/05/2023 10:34:26",
      "content": "<p>Thanks for your information. Did you test the processing speed of other augmentation libraries?</p>\n<p>I used albumentations to apply the same augmentations for each sequence (images and masks) as shown in the following url. But torchvision seems easier to implement the same augmentations for the sequence because it supports batch inputs.<br>\nreference: <a href=\"https://albumentations.ai/docs/examples/example_multi_target/\" target=\"_blank\">https://albumentations.ai/docs/examples/example_multi_target/</a></p>",
      "rawMarkdown": "Thanks for your information. Did you test the processing speed of other augmentation libraries?\n\nI used albumentations to apply the same augmentations for each sequence (images and masks) as shown in the following url. But torchvision seems easier to implement the same augmentations for the sequence because it supports batch inputs.\nreference: https://albumentations.ai/docs/examples/example_multi_target/",
      "votes": null
    },
    {
      "id": "2169703",
      "postDate": "03/05/2023 10:59:39",
      "content": "<p>Yes, your version looks correct.</p>\n<p>The speed up is not as large as I'd expect from the full optimal crop reading, but it's still faster, here is a benchmark on my system:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2Fb50e10737bed2571e512fa2609f5c27e%2Fimg_load_benchmark.png?generation=1678013879901151&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Yes, your version looks correct.\n\nThe speed up is not as large as I'd expect from the full optimal crop reading, but it's still faster, here is a benchmark on my system:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2Fb50e10737bed2571e512fa2609f5c27e%2Fimg_load_benchmark.png?generation=1678013879901151&alt=media)",
      "votes": null
    },
    {
      "id": "2169893",
      "postDate": "03/05/2023 14:28:34",
      "content": "<p>Loading image formats like JPEG or PNG can be slow, especially for large images. Instead, consider using optimized image formats like WebP or AVIF, which can be loaded much faster.</p>",
      "rawMarkdown": "Loading image formats like JPEG or PNG can be slow, especially for large images. Instead, consider using optimized image formats like WebP or AVIF, which can be loaded much faster.",
      "votes": null
    },
    {
      "id": "2169896",
      "postDate": "03/05/2023 14:30:57",
      "content": "<p>You may also use preprocessed data as preprocessing data ahead of time can also speed up image loading. This can include resizing images to a smaller size or precomputing feature representations using techniques like PCA or SIFT.</p>",
      "rawMarkdown": "You may also use preprocessed data as preprocessing data ahead of time can also speed up image loading. This can include resizing images to a smaller size or precomputing feature representations using techniques like PCA or SIFT.",
      "votes": null
    },
    {
      "id": "2169994",
      "postDate": "03/05/2023 15:39:53",
      "content": "<p>Thanks for the additional information.<br>\nTo be honest, I have not been able to compare processing speeds.<br>\nMy team was using TorchVision because it is GPU compatible, as we were pre-stack images and masks in the channel direction to single images.<br>\nAs you say, albumentations seems to be easier to use for multiple inputs.</p>",
      "rawMarkdown": "Thanks for the additional information.\nTo be honest, I have not been able to compare processing speeds.\nMy team was using TorchVision because it is GPU compatible, as we were pre-stack images and masks in the channel direction to single images.\nAs you say, albumentations seems to be easier to use for multiple inputs.",
      "votes": null
    },
    {
      "id": "2170044",
      "postDate": "03/05/2023 16:32:16",
      "content": "<p>Thanks for the answer and the specific sample code!<br>\nI tried the code you gave me in Kaggle notebook and got the result of ' normal PIL &gt; cropped PIL &gt; OpenCV ' (same result in my environment created based on Kaggle Docker). I appreciate your very useful information, but it may be affected by the library version, machine specs, etc….<br>\n<a href=\"https://www.kaggle.com/code/anyai28/load-image-benchmark\" target=\"_blank\">Benchmark notebook</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3381868%2F4a116ce7239dfcaa575794c2dafdd03f%2Fload_benchmark.png?generation=1678033586120894&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thanks for the answer and the specific sample code!\nI tried the code you gave me in Kaggle notebook and got the result of ' normal PIL > cropped PIL > OpenCV ' (same result in my environment created based on Kaggle Docker). I appreciate your very useful information, but it may be affected by the library version, machine specs, etc....\n[Benchmark notebook](https://www.kaggle.com/code/anyai28/load-image-benchmark)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3381868%2F4a116ce7239dfcaa575794c2dafdd03f%2Fload_benchmark.png?generation=1678033586120894&alt=media)",
      "votes": null
    },
    {
      "id": "2170048",
      "postDate": "03/05/2023 16:36:32",
      "content": "<p>Thanks for letting me know !<br>\nI’m not familiar with WebP and AVIF formats, I will look into it when I try in the future.</p>",
      "rawMarkdown": "Thanks for letting me know !\nI’m not familiar with WebP and AVIF formats, I will look into it when I try in the future.",
      "votes": null
    },
    {
      "id": "2171083",
      "postDate": "03/06/2023 14:18:54",
      "content": "<p>Never knew that about PIL tbh. Thanks for sharing 🙏</p>",
      "rawMarkdown": "Never knew that about PIL tbh. Thanks for sharing 🙏",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2168845,
      "author_name": "yururoi",
      "author_url": "",
      "post_date": "03/04/2023 15:19:00",
      "content": "<p>Files load faster in npy format than in jpeg.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1377145%2Fc13ec936cb918fddec0983b8079d013d%2FScreenshot%20from%202023-03-05%2000-12-58.png?generation=1677943049469449&amp;alt=media\" alt=\"\"><br>\nHowever, npy has a very large file size compared to jpeg.</p>\n<p>I used to cache images in memory for learning and inference whenever possible.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2169470,
          "author_name": "anyai28",
          "author_url": "",
          "post_date": "03/05/2023 06:26:17",
          "content": "<p>Thanks for letting me know!<br>\nnpy is indeed fast, but I had given up because of the lack of memory for the total number of frames…<br>\nAs others have said, caching seems to be effective. I'll look into the implementation!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2168849,
      "author_name": "takashisomeya",
      "author_url": "",
      "post_date": "03/04/2023 15:23:57",
      "content": "<p>Our team used multi-frame images and inferred them all at once.<br>\nThis speeded up training and inference time and improved scores with utilizing time-series information.<br>\nFYI, <a href=\"https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392290\" target=\"_blank\">our solution</a>. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2169475,
          "author_name": "anyai28",
          "author_url": "",
          "post_date": "03/05/2023 06:35:49",
          "content": "<p>Thanks for letting me know !<br>\nI learned a lot from reading your solution.<br>\nI now know that not only TurboJPEG but also cache by lru_cache is important.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2168905,
      "author_name": "koheiumeki",
      "author_url": "",
      "post_date": "03/04/2023 16:12:05",
      "content": "<p>I’m also interested in this topic because we didn't have much time to optimize the pipeline for sequential inputs.</p>\n<p>As the tentative solution, our team used lycon instead of cv2. This modification reduced 20% of data preprocessing time.<br>\ngithub: <a href=\"https://github.com/ethereon/lycon\" target=\"_blank\">https://github.com/ethereon/lycon</a></p>\n<p>But augmentation process (albumentation) is also heavy. So maybe we should consider this part also, e.g. using kornia.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2169488,
          "author_name": "anyai28",
          "author_url": "",
          "post_date": "03/05/2023 06:45:32",
          "content": "<p>Thanks for letting me know !<br>\nI didn’t know about the lycon library.<br>\nAs you mentioned, augmentation also has a significant impact on learning speed. Our team used TorchVision as an alternative.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2169685,
              "author_name": "koheiumeki",
              "author_url": "",
              "post_date": "03/05/2023 10:34:26",
              "content": "<p>Thanks for your information. Did you test the processing speed of other augmentation libraries?</p>\n<p>I used albumentations to apply the same augmentations for each sequence (images and masks) as shown in the following url. But torchvision seems easier to implement the same augmentations for the sequence because it supports batch inputs.<br>\nreference: <a href=\"https://albumentations.ai/docs/examples/example_multi_target/\" target=\"_blank\">https://albumentations.ai/docs/examples/example_multi_target/</a></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2169994,
                  "author_name": "anyai28",
                  "author_url": "",
                  "post_date": "03/05/2023 15:39:53",
                  "content": "<p>Thanks for the additional information.<br>\nTo be honest, I have not been able to compare processing speeds.<br>\nMy team was using TorchVision because it is GPU compatible, as we were pre-stack images and masks in the channel direction to single images.<br>\nAs you say, albumentations seems to be easier to use for multiple inputs.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2169293,
      "author_name": "dmytropoplavskiy",
      "author_url": "",
      "post_date": "03/05/2023 02:02:53",
      "content": "<p>During inference, I used the <a href=\"https://www.kaggle.com/lru\" target=\"_blank\">@lru</a>_cache() decorator to keep quite a large number of frames in RAM since the same frame is used multiple times for different players or steps when predicting sequentially.</p>\n<p>During training I calculated the bounding box for the target crop (after all geometry augmentations) and loaded only the necessary crop, something like:</p>\n<pre><code>img = PIL.Image(fn) # data is not decoded at this stage yet\nimg_crop = np.array(img.crop((left, upper, right, lower)))\n</code></pre>\n<p>This was a few times faster compared to loading the full resolution image with opencv.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2169493,
          "author_name": "anyai28",
          "author_url": "",
          "post_date": "03/05/2023 06:55:27",
          "content": "<p>Thanks for letting me know !<br>\nIt seems that lru_cache is an important feature, as some participants have responded to me about it.<br>\nAlso, thank you for sharing the implementation of the PIL. This is the first time I learned that it is possible to crop before decoding.</p>\n<p>I would like to check for correct understanding, but the PIL.Image class does not seem to be able to read the file as it is, is the following implementation correct?</p>\n<pre><code> PIL  Image\nimg = Image.()\nimg_crop = np.array(img.crop((left, upper, right, lower)))\n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 2169703,
              "author_name": "dmytropoplavskiy",
              "author_url": "",
              "post_date": "03/05/2023 10:59:39",
              "content": "<p>Yes, your version looks correct.</p>\n<p>The speed up is not as large as I'd expect from the full optimal crop reading, but it's still faster, here is a benchmark on my system:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2Fb50e10737bed2571e512fa2609f5c27e%2Fimg_load_benchmark.png?generation=1678013879901151&amp;alt=media\" alt=\"\"></p>",
              "votes": null,
              "replies": [
                {
                  "id": 2170044,
                  "author_name": "anyai28",
                  "author_url": "",
                  "post_date": "03/05/2023 16:32:16",
                  "content": "<p>Thanks for the answer and the specific sample code!<br>\nI tried the code you gave me in Kaggle notebook and got the result of ' normal PIL &gt; cropped PIL &gt; OpenCV ' (same result in my environment created based on Kaggle Docker). I appreciate your very useful information, but it may be affected by the library version, machine specs, etc….<br>\n<a href=\"https://www.kaggle.com/code/anyai28/load-image-benchmark\" target=\"_blank\">Benchmark notebook</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3381868%2F4a116ce7239dfcaa575794c2dafdd03f%2Fload_benchmark.png?generation=1678033586120894&amp;alt=media\" alt=\"\"></p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        },
        {
          "id": 2171083,
          "author_name": "samir95",
          "author_url": "",
          "post_date": "03/06/2023 14:18:54",
          "content": "<p>Never knew that about PIL tbh. Thanks for sharing 🙏</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2169893,
      "author_name": "paikerhussain",
      "author_url": "",
      "post_date": "03/05/2023 14:28:34",
      "content": "<p>Loading image formats like JPEG or PNG can be slow, especially for large images. Instead, consider using optimized image formats like WebP or AVIF, which can be loaded much faster.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2170048,
          "author_name": "anyai28",
          "author_url": "",
          "post_date": "03/05/2023 16:36:32",
          "content": "<p>Thanks for letting me know !<br>\nI’m not familiar with WebP and AVIF formats, I will look into it when I try in the future.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2169896,
      "author_name": "paikerhussain",
      "author_url": "",
      "post_date": "03/05/2023 14:30:57",
      "content": "<p>You may also use preprocessed data as preprocessing data ahead of time can also speed up image loading. This can include resizing images to a smaller size or precomputing feature representations using techniques like PCA or SIFT.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2168779": "Many top solutions have been shared and I'm learning from each team's unique approach. Thank you.\n\nAccording to the top solutions, one key was to input many multi-frame images to 2.5D and 3D CNN, but the image loading speed is the bottleneck to do so.\n\nA naive approach would be to use OpenCV or TurboJPEG (a bit faster) to load images one by one, as [PublicNotebook](https://www.kaggle.com/code/zzy990106/nfl-2-5d-cnn-baseline-inference) does, but this would take too much learning time.\n\nWhat are some tips for efficient learning? Do you use [decord](https://github.com/dmlc/decord) or something similar to process video by video?\n\nI would like to learn for the future, If you don't mind could you tell me.\n\nBest regards.",
    "2168845": "Files load faster in npy format than in jpeg.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1377145%2Fc13ec936cb918fddec0983b8079d013d%2FScreenshot%20from%202023-03-05%2000-12-58.png?generation=1677943049469449&alt=media)\nHowever, npy has a very large file size compared to jpeg.\n\nI used to cache images in memory for learning and inference whenever possible.",
    "2168849": "Our team used multi-frame images and inferred them all at once.\nThis speeded up training and inference time and improved scores with utilizing time-series information.\nFYI, [our solution](https://www.kaggle.com/competitions/nfl-player-contact-detection/discussion/392290).",
    "2168905": "I’m also interested in this topic because we didn't have much time to optimize the pipeline for sequential inputs.\n\nAs the tentative solution, our team used lycon instead of cv2. This modification reduced 20% of data preprocessing time.\ngithub: https://github.com/ethereon/lycon\n\nBut augmentation process (albumentation) is also heavy. So maybe we should consider this part also, e.g. using kornia.",
    "2169293": "During inference, I used the @lru_cache() decorator to keep quite a large number of frames in RAM since the same frame is used multiple times for different players or steps when predicting sequentially.\n\nDuring training I calculated the bounding box for the target crop (after all geometry augmentations) and loaded only the necessary crop, something like:\n\n```\nimg = PIL.Image(fn) # data is not decoded at this stage yet\nimg_crop = np.array(img.crop((left, upper, right, lower)))\n```\n\nThis was a few times faster compared to loading the full resolution image with opencv.",
    "2169470": "Thanks for letting me know!\nnpy is indeed fast, but I had given up because of the lack of memory for the total number of frames...\nAs others have said, caching seems to be effective. I'll look into the implementation!",
    "2169475": "Thanks for letting me know !\nI learned a lot from reading your solution.\nI now know that not only TurboJPEG but also cache by lru_cache is important.",
    "2169488": "Thanks for letting me know !\nI didn’t know about the lycon library.\nAs you mentioned, augmentation also has a significant impact on learning speed. Our team used TorchVision as an alternative.",
    "2169493": "Thanks for letting me know !\nIt seems that lru_cache is an important feature, as some participants have responded to me about it.\nAlso, thank you for sharing the implementation of the PIL. This is the first time I learned that it is possible to crop before decoding.\n\nI would like to check for correct understanding, but the PIL.Image class does not seem to be able to read the file as it is, is the following implementation correct?\n```python\nfrom PIL import Image\nimg = Image.open('xxx.jpeg')\nimg_crop = np.array(img.crop((left, upper, right, lower)))\n```",
    "2169685": "Thanks for your information. Did you test the processing speed of other augmentation libraries?\n\nI used albumentations to apply the same augmentations for each sequence (images and masks) as shown in the following url. But torchvision seems easier to implement the same augmentations for the sequence because it supports batch inputs.\nreference: https://albumentations.ai/docs/examples/example_multi_target/",
    "2169703": "Yes, your version looks correct.\n\nThe speed up is not as large as I'd expect from the full optimal crop reading, but it's still faster, here is a benchmark on my system:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F743064%2Fb50e10737bed2571e512fa2609f5c27e%2Fimg_load_benchmark.png?generation=1678013879901151&alt=media)",
    "2169893": "Loading image formats like JPEG or PNG can be slow, especially for large images. Instead, consider using optimized image formats like WebP or AVIF, which can be loaded much faster.",
    "2169896": "You may also use preprocessed data as preprocessing data ahead of time can also speed up image loading. This can include resizing images to a smaller size or precomputing feature representations using techniques like PCA or SIFT.",
    "2169994": "Thanks for the additional information.\nTo be honest, I have not been able to compare processing speeds.\nMy team was using TorchVision because it is GPU compatible, as we were pre-stack images and masks in the channel direction to single images.\nAs you say, albumentations seems to be easier to use for multiple inputs.",
    "2170044": "Thanks for the answer and the specific sample code!\nI tried the code you gave me in Kaggle notebook and got the result of ' normal PIL > cropped PIL > OpenCV ' (same result in my environment created based on Kaggle Docker). I appreciate your very useful information, but it may be affected by the library version, machine specs, etc....\n[Benchmark notebook](https://www.kaggle.com/code/anyai28/load-image-benchmark)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3381868%2F4a116ce7239dfcaa575794c2dafdd03f%2Fload_benchmark.png?generation=1678033586120894&alt=media)",
    "2170048": "Thanks for letting me know !\nI’m not familiar with WebP and AVIF formats, I will look into it when I try in the future.",
    "2171083": "Never knew that about PIL tbh. Thanks for sharing 🙏"
  },
  "source": "meta"
}