{
  "id": 132251,
  "title": "NVIDIA DALI fails at submission",
  "url": "/competitions/deepfake-detection-challenge/discussion/132251",
  "author_name": "Iafoss",
  "post_date": "2020-02-25T05:10:15.355000",
  "votes": 7,
  "comment_count": 12,
  "views": 0,
  "content": "<p>The thing that takes the most of the time during submission is reading the video, especially if GPU kernel is used. One thing that could resolve the problem is using NVIDIA DALI.\nThe kernel I wrote works perfectly fine on sample test video, but fails on submission. Specifically, building a video pipe leads to an error (I ran several experiments with putting this part into a separate cell). This error could be caused either by an error during installation of the package <code>!pip install /kaggle/input/nvidia-dali/nvidia_dali-0.18.0-1062351-cp36-cp36m-manylinux1_x86_64.whl</code> or by a different environment where the submission is run. You understand that the submission run cannot return any exception messages... It's just a black box.</p>\n\n<p>Did any of you build successfully a submission pipeline with DALI or have an idea that could cause this problem? I expect that installation from *.whl file uploaded to a dataset shouldn't be an issue for this competition.</p>",
  "messages": [
    {
      "id": 755754,
      "postDate": "2020-02-25T05:10:15.357Z",
      "content": "<p>The thing that takes the most of the time during submission is reading the video, especially if GPU kernel is used. One thing that could resolve the problem is using NVIDIA DALI.\nThe kernel I wrote works perfectly fine on sample test video, but fails on submission. Specifically, building a video pipe leads to an error (I ran several experiments with putting this part into a separate cell). This error could be caused either by an error during installation of the package <code>!pip install /kaggle/input/nvidia-dali/nvidia_dali-0.18.0-1062351-cp36-cp36m-manylinux1_x86_64.whl</code> or by a different environment where the submission is run. You understand that the submission run cannot return any exception messages... It's just a black box.</p>\n\n<p>Did any of you build successfully a submission pipeline with DALI or have an idea that could cause this problem? I expect that installation from *.whl file uploaded to a dataset shouldn't be an issue for this competition.</p>",
      "rawMarkdown": "The thing that takes the most of the time during submission is reading the video, especially if GPU kernel is used. One thing that could resolve the problem is using NVIDIA DALI.\nThe kernel I wrote works perfectly fine on sample test video, but fails on submission. Specifically, building a video pipe leads to an error (I ran several experiments with putting this part into a separate cell). This error could be caused either by an error during installation of the package `!pip install /kaggle/input/nvidia-dali/nvidia_dali-0.18.0-1062351-cp36-cp36m-manylinux1_x86_64.whl` or by a different environment where the submission is run. You understand that the submission run cannot return any exception messages... It's just a black box.\n\nDid any of you build successfully a submission pipeline with DALI or have an idea that could cause this problem? I expect that installation from *.whl file uploaded to a dataset shouldn't be an issue for this competition.",
      "votes": 7
    },
    {
      "id": 759909,
      "postDate": "2020-02-29T16:01:30.163Z",
      "content": "<p>Could we receive some guidance on the topic? It is not the first post about issues with kernel submissions. Is it possible to make the error reports more verbose to somehow simplify the debugging and installation of custom packages? Spending hours of time and dozens of submits to debug opaque kernel environment is less exciting (and much less useful) than coming up with a working model :)</p>",
      "rawMarkdown": "Could we receive some guidance on the topic? It is not the first post about issues with kernel submissions. Is it possible to make the error reports more verbose to somehow simplify the debugging and installation of custom packages? Spending hours of time and dozens of submits to debug opaque kernel environment is less exciting (and much less useful) than coming up with a working model :)",
      "votes": 1
    },
    {
      "id": 755832,
      "postDate": "2020-02-25T06:50:39.990Z",
      "content": "<p>The sample testset has 400 videos that are perfectly fine.</p>\n\n<p>The hidden public test set has 4000 videos and 27 corrupt videos in it. (the one your submission gets tested against) if your code crashes only on submission ... I would read the frames with something like this: \n<code>\ntry:\n      v_cap = cv2.VideoCapture(video_Path)\n      for j in range(v_len):\n         success, vframe = v_cap.read()\n         if(success):\n            #Add vframe to your batch or whatever pipelines steps you have\nexcept:\n    raise Exception(\"Stopped at \"+video_Path) \n</code></p>",
      "rawMarkdown": "The sample testset has 400 videos that are perfectly fine.\n\nThe hidden public test set has 4000 videos and 27 corrupt videos in it. (the one your submission gets tested against) if your code crashes only on submission ... I would read the frames with something like this: \n```\ntry:\n      v_cap = cv2.VideoCapture(video_Path)\n      for j in range(v_len):\n         success, vframe = v_cap.read()\n         if(success):\n            #Add vframe to your batch or whatever pipelines steps you have\nexcept:\n    raise Exception(\"Stopped at \"+video_Path) \n```",
      "votes": 2,
      "replies": [
        {
          "id": 755838,
          "postDate": "2020-02-25T07:01:23.240Z",
          "content": "<p>Thanks. Yes, I added exception blocks into the inference part, so that if any exception is risen the code predicts 0.5. But with using several subs I figured out that the issue is that DALI video Pipeline is not created, and when I create the corresponding object and call <code>build</code> on it, I get an exception. Though, it is very difficult to extract any further details from a submission run.</p>",
          "rawMarkdown": "Thanks. Yes, I added exception blocks into the inference part, so that if any exception is risen the code predicts 0.5. But with using several subs I figured out that the issue is that DALI video Pipeline is not created, and when I create the corresponding object and call `build` on it, I get an exception. Though, it is very difficult to extract any further details from a submission run."
        },
        {
          "id": 755842,
          "postDate": "2020-02-25T07:11:54.517Z",
          "content": "<p>I'm not familiar with DALI ... but the only difference I have personally experienced between the test sample environment and the submission environment is the videos ... it seems DALI crashes on corrupt data? you can spend submissions debugging the problem by setting all videos to 1 for instance and when it crashes to 0 .. then have multiple values to troubleshoot further ... Also you can email kaggle support to ask specifically.\n<a href=\"/juliaelliott\">@juliaelliott</a> might be able to point you out to the right direction?</p>",
          "rawMarkdown": "I'm not familiar with DALI ... but the only difference I have personally experienced between the test sample environment and the submission environment is the videos ... it seems DALI crashes on corrupt data? you can spend submissions debugging the problem by setting all videos to 1 for instance and when it crashes to 0 .. then have multiple values to troubleshoot further ... Also you can email kaggle support to ask specifically.\n@juliaelliott might be able to point you out to the right direction?",
          "votes": 1
        },
        {
          "id": 755848,
          "postDate": "2020-02-25T07:18:44.180Z",
          "content": "<p>Thank you so much for your help, I really appreciate it.</p>",
          "rawMarkdown": "Thank you so much for your help, I really appreciate it."
        },
        {
          "id": 756029,
          "postDate": "2020-02-25T11:03:01.583Z",
          "content": "<p>Can I ask you how much videos you load per pipe?</p>",
          "rawMarkdown": "Can I ask you how much videos you load per pipe?"
        },
        {
          "id": 756318,
          "postDate": "2020-02-25T15:57:57.953Z",
          "content": "<p>It loads just one video each time keeping every 3-d frame, however when the pipe is constructed I need to pass the list of all video.</p>",
          "rawMarkdown": "It loads just one video each time keeping every 3-d frame, however when the pipe is constructed I need to pass the list of all video.",
          "votes": 1
        },
        {
          "id": 756371,
          "postDate": "2020-02-25T17:02:07.340Z",
          "content": "<p>For me it crashes while I start iteration. So I can safely run <code>pipe.build()</code>. The notebook dies when I execute <code>pipe.run()</code> and there aren't any available samples. So does the creation of a pytorch iterator because it prefetches the first batch.\nAfter adding <code>pipe.epoch_size('reader') &gt; 0</code> before fetching any samples I can avoid crashes. However I haven't tried running it on kaggle yet.</p>",
          "rawMarkdown": "For me it crashes while I start iteration. So I can safely run `pipe.build()`. The notebook dies when I execute `pipe.run()` and there aren't any available samples. So does the creation of a pytorch iterator because it prefetches the first batch.\nAfter adding `pipe.epoch_size('reader') &gt; 0` before fetching any samples I can avoid crashes. However I haven't tried running it on kaggle yet."
        },
        {
          "id": 756433,
          "postDate": "2020-02-25T18:06:07.637Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> , it would be quite helpful for debugging to have a possibility, somehow, to see the exception description thrown during submission (instead of spending tens of submissions to track what is going on line by line, especially for this competition which counts incorrect subs and has 2 subs per day limit). However, I understand that some people could start abusing it to prob the data. Probably, if kaggle representative could review the generated exception list under specific request and disclose it if nothing related to probing is found there. But it is extra load for people working for kaggle(</p>",
          "rawMarkdown": "@juliaelliott , it would be quite helpful for debugging to have a possibility, somehow, to see the exception description thrown during submission (instead of spending tens of submissions to track what is going on line by line, especially for this competition which counts incorrect subs and has 2 subs per day limit). However, I understand that some people could start abusing it to prob the data. Probably, if kaggle representative could review the generated exception list under specific request and disclose it if nothing related to probing is found there. But it is extra load for people working for kaggle(",
          "votes": 1
        },
        {
          "id": 765957,
          "postDate": "2020-03-07T12:42:46.650Z",
          "content": "<p><a href=\"/iafoss\">@iafoss</a> I've tried it and noticed:\n1. DALI crashes if you try to load more frames than available (i.e. 400 frames but 300 frames only in video).\nThe only solution I've found is to use OpenCV to get frame count before creating the Pipeline.\n2. DALI crashes if your sequence length is higher than 100.</p>\n\n<p>I've created a kernel to test it:\n<a href=\"https://www.kaggle.com/mpware/videoreaders-benchmark\">https://www.kaggle.com/mpware/videoreaders-benchmark</a></p>\n\n<p>My DALIVideoReader implementation with around 90 frames crashes on test set (4000 videos) and works fine on the 400 videos. No idea why, the only error is Notebook Exceeded allowed compute which should be a problem with memory or GPU.</p>",
          "rawMarkdown": "@iafoss I've tried it and noticed:\n1. DALI crashes if you try to load more frames than available (i.e. 400 frames but 300 frames only in video).\nThe only solution I've found is to use OpenCV to get frame count before creating the Pipeline.\n2. DALI crashes if your sequence length is higher than 100.\n\nI've created a kernel to test it:\nhttps://www.kaggle.com/mpware/videoreaders-benchmark\n\nMy DALIVideoReader implementation with around 90 frames crashes on test set (4000 videos) and works fine on the 400 videos. No idea why, the only error is Notebook Exceeded allowed compute which should be a problem with memory or GPU."
        },
        {
          "id": 766084,
          "postDate": "2020-03-07T16:38:24.933Z",
          "content": "<p>I couldn't find the issue since I didn't want to spend too many subs on that( So right now I just use one of the publically available pipelines, and each sub just takes 8 hours from kaggle GPU time and ~1 hour from my 30 hour quota( \nIn my setup I was reading every 3-d frame, so 100 frames per video, which works fine to me on sample data. After rearranging the code I found that creation of a video pipe fails on private LB sub: If I keep it in the same cell as inference loop, I get an error in sub (no submission.csv file), while if I put it into another cell, I get a result corresponding to 0.5 for each image submission. In my code if I get an exception during inference, I return 0.5 as a prediction. So, I think the creation of a video pipe fails, and then the next cell with inference is executed with exception thrown for each video (because cannot call any methods on None), that generates a file with 0.5 predictions. I couldn't figure it out why the video pipe cannot be created here: it doesn't involve any data, just creation of an object... There could be several issues for that: DALI is not installed properly at LB submission, or some environmental things are different from committing the kernel... But I don't know any other details.</p>",
          "rawMarkdown": "I couldn't find the issue since I didn't want to spend too many subs on that( So right now I just use one of the publically available pipelines, and each sub just takes 8 hours from kaggle GPU time and ~1 hour from my 30 hour quota( \nIn my setup I was reading every 3-d frame, so 100 frames per video, which works fine to me on sample data. After rearranging the code I found that creation of a video pipe fails on private LB sub: If I keep it in the same cell as inference loop, I get an error in sub (no submission.csv file), while if I put it into another cell, I get a result corresponding to 0.5 for each image submission. In my code if I get an exception during inference, I return 0.5 as a prediction. So, I think the creation of a video pipe fails, and then the next cell with inference is executed with exception thrown for each video (because cannot call any methods on None), that generates a file with 0.5 predictions. I couldn't figure it out why the video pipe cannot be created here: it doesn't involve any data, just creation of an object... There could be several issues for that: DALI is not installed properly at LB submission, or some environmental things are different from committing the kernel... But I don't know any other details.",
          "votes": 1
        }
      ]
    },
    {
      "id": 758584,
      "postDate": "2020-02-28T00:16:55.320Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 759909,
      "author_name": "Ilia Zaitsev",
      "author_url": "",
      "post_date": "2020-02-29T16:01:30.163000",
      "content": "<p>Could we receive some guidance on the topic? It is not the first post about issues with kernel submissions. Is it possible to make the error reports more verbose to somehow simplify the debugging and installation of custom packages? Spending hours of time and dozens of submits to debug opaque kernel environment is less exciting (and much less useful) than coming up with a working model :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 755832,
      "author_name": "mjsML",
      "author_url": "",
      "post_date": "2020-02-25T06:50:39.990000",
      "content": "<p>The sample testset has 400 videos that are perfectly fine.</p>\n\n<p>The hidden public test set has 4000 videos and 27 corrupt videos in it. (the one your submission gets tested against) if your code crashes only on submission ... I would read the frames with something like this: \n<code>\ntry:\n      v_cap = cv2.VideoCapture(video_Path)\n      for j in range(v_len):\n         success, vframe = v_cap.read()\n         if(success):\n            #Add vframe to your batch or whatever pipelines steps you have\nexcept:\n    raise Exception(\"Stopped at \"+video_Path) \n</code></p>",
      "votes": 2,
      "replies": [
        {
          "id": 755838,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-02-25T07:01:23.240000",
          "content": "<p>Thanks. Yes, I added exception blocks into the inference part, so that if any exception is risen the code predicts 0.5. But with using several subs I figured out that the issue is that DALI video Pipeline is not created, and when I create the corresponding object and call <code>build</code> on it, I get an exception. Though, it is very difficult to extract any further details from a submission run.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755842,
          "author_name": "mjsML",
          "author_url": "",
          "post_date": "2020-02-25T07:11:54.517000",
          "content": "<p>I'm not familiar with DALI ... but the only difference I have personally experienced between the test sample environment and the submission environment is the videos ... it seems DALI crashes on corrupt data? you can spend submissions debugging the problem by setting all videos to 1 for instance and when it crashes to 0 .. then have multiple values to troubleshoot further ... Also you can email kaggle support to ask specifically.\n<a href=\"/juliaelliott\">@juliaelliott</a> might be able to point you out to the right direction?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 755848,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-02-25T07:18:44.180000",
          "content": "<p>Thank you so much for your help, I really appreciate it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756029,
          "author_name": "Dmitry Vorobiev",
          "author_url": "",
          "post_date": "2020-02-25T11:03:01.583000",
          "content": "<p>Can I ask you how much videos you load per pipe?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756318,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-02-25T15:57:57.953000",
          "content": "<p>It loads just one video each time keeping every 3-d frame, however when the pipe is constructed I need to pass the list of all video.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 756371,
          "author_name": "Dmitry Vorobiev",
          "author_url": "",
          "post_date": "2020-02-25T17:02:07.340000",
          "content": "<p>For me it crashes while I start iteration. So I can safely run <code>pipe.build()</code>. The notebook dies when I execute <code>pipe.run()</code> and there aren't any available samples. So does the creation of a pytorch iterator because it prefetches the first batch.\nAfter adding <code>pipe.epoch_size('reader') &gt; 0</code> before fetching any samples I can avoid crashes. However I haven't tried running it on kaggle yet.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 756433,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-02-25T18:06:07.637000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> , it would be quite helpful for debugging to have a possibility, somehow, to see the exception description thrown during submission (instead of spending tens of submissions to track what is going on line by line, especially for this competition which counts incorrect subs and has 2 subs per day limit). However, I understand that some people could start abusing it to prob the data. Probably, if kaggle representative could review the generated exception list under specific request and disclose it if nothing related to probing is found there. But it is extra load for people working for kaggle(</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 765957,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-03-07T12:42:46.650000",
          "content": "<p><a href=\"/iafoss\">@iafoss</a> I've tried it and noticed:\n1. DALI crashes if you try to load more frames than available (i.e. 400 frames but 300 frames only in video).\nThe only solution I've found is to use OpenCV to get frame count before creating the Pipeline.\n2. DALI crashes if your sequence length is higher than 100.</p>\n\n<p>I've created a kernel to test it:\n<a href=\"https://www.kaggle.com/mpware/videoreaders-benchmark\">https://www.kaggle.com/mpware/videoreaders-benchmark</a></p>\n\n<p>My DALIVideoReader implementation with around 90 frames crashes on test set (4000 videos) and works fine on the 400 videos. No idea why, the only error is Notebook Exceeded allowed compute which should be a problem with memory or GPU.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766084,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-03-07T16:38:24.933000",
          "content": "<p>I couldn't find the issue since I didn't want to spend too many subs on that( So right now I just use one of the publically available pipelines, and each sub just takes 8 hours from kaggle GPU time and ~1 hour from my 30 hour quota( \nIn my setup I was reading every 3-d frame, so 100 frames per video, which works fine to me on sample data. After rearranging the code I found that creation of a video pipe fails on private LB sub: If I keep it in the same cell as inference loop, I get an error in sub (no submission.csv file), while if I put it into another cell, I get a result corresponding to 0.5 for each image submission. In my code if I get an exception during inference, I return 0.5 as a prediction. So, I think the creation of a video pipe fails, and then the next cell with inference is executed with exception thrown for each video (because cannot call any methods on None), that generates a file with 0.5 predictions. I couldn't figure it out why the video pipe cannot be created here: it doesn't involve any data, just creation of an object... There could be several issues for that: DALI is not installed properly at LB submission, or some environmental things are different from committing the kernel... But I don't know any other details.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 758584,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-28T00:16:55.320000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "755754": "The thing that takes the most of the time during submission is reading the video, especially if GPU kernel is used. One thing that could resolve the problem is using NVIDIA DALI.\nThe kernel I wrote works perfectly fine on sample test video, but fails on submission. Specifically, building a video pipe leads to an error (I ran several experiments with putting this part into a separate cell). This error could be caused either by an error during installation of the package `!pip install /kaggle/input/nvidia-dali/nvidia_dali-0.18.0-1062351-cp36-cp36m-manylinux1_x86_64.whl` or by a different environment where the submission is run. You understand that the submission run cannot return any exception messages... It's just a black box.\n\nDid any of you build successfully a submission pipeline with DALI or have an idea that could cause this problem? I expect that installation from *.whl file uploaded to a dataset shouldn't be an issue for this competition.",
    "759909": "Could we receive some guidance on the topic? It is not the first post about issues with kernel submissions. Is it possible to make the error reports more verbose to somehow simplify the debugging and installation of custom packages? Spending hours of time and dozens of submits to debug opaque kernel environment is less exciting (and much less useful) than coming up with a working model :)",
    "755832": "The sample testset has 400 videos that are perfectly fine.\n\nThe hidden public test set has 4000 videos and 27 corrupt videos in it. (the one your submission gets tested against) if your code crashes only on submission ... I would read the frames with something like this: \n```\ntry:\n      v_cap = cv2.VideoCapture(video_Path)\n      for j in range(v_len):\n         success, vframe = v_cap.read()\n         if(success):\n            #Add vframe to your batch or whatever pipelines steps you have\nexcept:\n    raise Exception(\"Stopped at \"+video_Path) \n```",
    "758584": ""
  }
}