{
  "id": 145719,
  "title": "Reasons some kernels might have failed",
  "url": "/competitions/deepfake-detection-challenge/discussion/145719",
  "author_name": "ryches",
  "post_date": "2020-04-24T08:53:58.532000",
  "votes": 6,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I will update this list more tomorrow, but I audited some people's work after they made their kernels public after the submission deadline and I noticed a couple of subtle errors that concerned me. </p>\n\n<ol>\n<li><p>writing to file and not cleaning up between predictions</p>\n\n<ul><li>In this kernel, the videos are rewritten, but then never cleared until the very end with the rm-r videos call. <a href=\"https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference/data\">https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference/data</a>\nThis can accumulate in disk space and eventually past a certain number of videos it will fail. </li></ul></li>\n</ol>\n\n<p>2.Not putting boundaries on input size. </p>\n\n<ul>\n<li>Some kernels did such things as sampling frames from the expected 300. if you sample every 10 frames it is only 30 frames and not that large in memory or to push to gpu, but we were given no guarantee of the length of the test set videos. Since they were in the wild it's entirely possible they were 60 fps and a minute long resulting in your input being a much larger size and potentially OOM'ing the entire kernel and killing it. It is better to put some bounds on the max size recorded from the videos. </li>\n</ul>\n\n<p>...to be continued when it's not 2am. </p>",
  "messages": [
    {
      "id": 818973,
      "postDate": "2020-04-24T08:53:58.533Z",
      "content": "<p>I will update this list more tomorrow, but I audited some people's work after they made their kernels public after the submission deadline and I noticed a couple of subtle errors that concerned me. </p>\n\n<ol>\n<li><p>writing to file and not cleaning up between predictions</p>\n\n<ul><li>In this kernel, the videos are rewritten, but then never cleared until the very end with the rm-r videos call. <a href=\"https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference/data\">https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference/data</a>\nThis can accumulate in disk space and eventually past a certain number of videos it will fail. </li></ul></li>\n</ol>\n\n<p>2.Not putting boundaries on input size. </p>\n\n<ul>\n<li>Some kernels did such things as sampling frames from the expected 300. if you sample every 10 frames it is only 30 frames and not that large in memory or to push to gpu, but we were given no guarantee of the length of the test set videos. Since they were in the wild it's entirely possible they were 60 fps and a minute long resulting in your input being a much larger size and potentially OOM'ing the entire kernel and killing it. It is better to put some bounds on the max size recorded from the videos. </li>\n</ul>\n\n<p>...to be continued when it's not 2am. </p>",
      "rawMarkdown": "I will update this list more tomorrow, but I audited some people's work after they made their kernels public after the submission deadline and I noticed a couple of subtle errors that concerned me. \n\n1. writing to file and not cleaning up between predictions\n\n- In this kernel, the videos are rewritten, but then never cleared until the very end with the rm-r videos call. https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference/data\nThis can accumulate in disk space and eventually past a certain number of videos it will fail. \n\n2.Not putting boundaries on input size. \n\n- Some kernels did such things as sampling frames from the expected 300. if you sample every 10 frames it is only 30 frames and not that large in memory or to push to gpu, but we were given no guarantee of the length of the test set videos. Since they were in the wild it's entirely possible they were 60 fps and a minute long resulting in your input being a much larger size and potentially OOM'ing the entire kernel and killing it. It is better to put some bounds on the max size recorded from the videos. \n\n...to be continued when it's not 2am. \n\n",
      "votes": 6
    },
    {
      "id": 819029,
      "postDate": "2020-04-24T09:46:01.657Z",
      "content": "<p>Good points. I think also as <a href=\"/vaillant\">@vaillant</a> raised <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145643#818587\">here </a>people might not have been filling completely failed predictions with a default prediction like 0.5 if their inference pipeline had never failed to parse a video in the public dataset. This is relevant giving their phrasing:</p>\n\n<p><code>There was a ~6.5% failure rate on re-runs, due mostly to incomplete submissions (submissions which failed to create predictions for all samples)</code></p>",
      "rawMarkdown": "Good points. I think also as @vaillant raised [here ](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145643#818587)people might not have been filling completely failed predictions with a default prediction like 0.5 if their inference pipeline had never failed to parse a video in the public dataset. This is relevant giving their phrasing:\n\n`There was a ~6.5% failure rate on re-runs, due mostly to incomplete submissions (submissions which failed to create predictions for all samples)`",
      "votes": 2,
      "replies": [
        {
          "id": 819038,
          "postDate": "2020-04-24T09:55:43.637Z",
          "content": "<p>That looks like a plausible reason. After checking I noticed not all my error paths made sure the skipped file received a 0.5 prediction. 🤕 </p>",
          "rawMarkdown": "That looks like a plausible reason. After checking I noticed not all my error paths made sure the skipped file received a 0.5 prediction. 🤕 "
        },
        {
          "id": 819401,
          "postDate": "2020-04-24T15:13:44.400Z",
          "content": "<p><a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122841#701781\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122841#701781</a></p>\n\n<p>I listed some of these previously. It is very unfortunate that these subtle gotchas were revealed on the test set. I am unsure of how kaggle will handle this. Some in the past have said it's the fault of the competitor for not properly hardening their code, but it's exceptionally difficult to make a perfect pipeline on completely unseen data. Don't think it should be a defensive programming competition, we're here for machine learning. </p>",
          "rawMarkdown": "https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122841#701781\n\nI listed some of these previously. It is very unfortunate that these subtle gotchas were revealed on the test set. I am unsure of how kaggle will handle this. Some in the past have said it's the fault of the competitor for not properly hardening their code, but it's exceptionally difficult to make a perfect pipeline on completely unseen data. Don't think it should be a defensive programming competition, we're here for machine learning. "
        }
      ]
    },
    {
      "id": 823338,
      "postDate": "2020-04-27T15:20:13.360Z",
      "content": "<p>variable resolution and variable frame rate videos (as described in the paper)</p>",
      "rawMarkdown": "variable resolution and variable frame rate videos (as described in the paper)"
    },
    {
      "id": 820303,
      "postDate": "2020-04-25T10:19:31.790Z",
      "content": "<p>Interesting points! There sure was a lot that could go wrong in pipelines for this competition.</p>",
      "rawMarkdown": "Interesting points! There sure was a lot that could go wrong in pipelines for this competition."
    },
    {
      "id": 819156,
      "postDate": "2020-04-24T11:47:40.353Z",
      "content": "<p>Here is mine. Submission failed on private LB. Interested to see if you can spot anything out of the ordinary.</p>\n\n<p><a href=\"https://www.kaggle.com/maralski/fork-of-ensemble-of-5-networks?scriptVersionId=31169240\">https://www.kaggle.com/maralski/fork-of-ensemble-of-5-networks?scriptVersionId=31169240</a></p>",
      "rawMarkdown": "Here is mine. Submission failed on private LB. Interested to see if you can spot anything out of the ordinary.\n\nhttps://www.kaggle.com/maralski/fork-of-ensemble-of-5-networks?scriptVersionId=31169240",
      "replies": [
        {
          "id": 819429,
          "postDate": "2020-04-24T15:31:24.330Z",
          "content": "<p><code>get_frameids = np.linspace(0, frame_count - 1, grab_frames, endpoint=True, dtype=np.int)</code></p>\n\n<p>this is the line that concerns me. Your grab_frames is saying to grab every 60 frames, but your frame_count is just the number of frames in the video so you are potentially putting an unbounded amount of video there. Unless I am missing somewhere else where this is protected against or I am misunderstanding the code, I think it might fall under the second point I raised. </p>",
          "rawMarkdown": "`get_frameids = np.linspace(0, frame_count - 1, grab_frames, endpoint=True, dtype=np.int)`\n\nthis is the line that concerns me. Your grab_frames is saying to grab every 60 frames, but your frame_count is just the number of frames in the video so you are potentially putting an unbounded amount of video there. Unless I am missing somewhere else where this is protected against or I am misunderstanding the code, I think it might fall under the second point I raised. ",
          "votes": 1
        },
        {
          "id": 819983,
          "postDate": "2020-04-25T04:26:20.640Z",
          "content": "<p>Thanks for looking. That just grabs a total of 60 frames across the frame count. So if frame count is 300 it will grab:</p>\n\n<blockquote>\n  <p>[ins] In [6]: np.linspace(0, 300 - 1, 60, endpoint=True, dtype=np.int) <br>\n  Out[6]: \n  array([  0,   5,  10,  15,  20,  25,  30,  35,  40,  45,  50,  55,  60,\n          65,  70,  76,  81,  86,  91,  96, 101, 106, 111, 116, 121, 126,\n         131, 136, 141, 146, 152, 157, 162, 167, 172, 177, 182, 187, 192,\n         197, 202, 207, 212, 217, 222, 228, 233, 238, 243, 248, 253, 258,\n         263, 268, 273, 278, 283, 288, 293, 299])</p>\n</blockquote>\n\n<p>If it's 600 it will still grab 60.</p>\n\n<blockquote>\n  <p>[ins] In [7]: np.linspace(0, 600 - 1, 60, endpoint=True, dtype=np.int) <br>\n  Out[7]: \n  array([  0,  10,  20,  30,  40,  50,  60,  71,  81,  91, 101, 111, 121,\n         131, 142, 152, 162, 172, 182, 192, 203, 213, 223, 233, 243, 253,\n         263, 274, 284, 294, 304, 314, 324, 335, 345, 355, 365, 375, 385,\n         395, 406, 416, 426, 436, 446, 456, 467, 477, 487, 497, 507, 517,\n         527, 538, 548, 558, 568, 578, 588, 599])</p>\n</blockquote>\n\n<p>I'll keep looking.</p>",
          "rawMarkdown": "Thanks for looking. That just grabs a total of 60 frames across the frame count. So if frame count is 300 it will grab:\n\n&gt; [ins] In [6]: np.linspace(0, 300 - 1, 60, endpoint=True, dtype=np.int)                                                   \nOut[6]: \narray([  0,   5,  10,  15,  20,  25,  30,  35,  40,  45,  50,  55,  60,\n        65,  70,  76,  81,  86,  91,  96, 101, 106, 111, 116, 121, 126,\n       131, 136, 141, 146, 152, 157, 162, 167, 172, 177, 182, 187, 192,\n       197, 202, 207, 212, 217, 222, 228, 233, 238, 243, 248, 253, 258,\n       263, 268, 273, 278, 283, 288, 293, 299])\n\nIf it's 600 it will still grab 60.\n\n&gt; [ins] In [7]: np.linspace(0, 600 - 1, 60, endpoint=True, dtype=np.int)                                                   \nOut[7]: \narray([  0,  10,  20,  30,  40,  50,  60,  71,  81,  91, 101, 111, 121,\n       131, 142, 152, 162, 172, 182, 192, 203, 213, 223, 233, 243, 253,\n       263, 274, 284, 294, 304, 314, 324, 335, 345, 355, 365, 375, 385,\n       395, 406, 416, 426, 436, 446, 456, 467, 477, 487, 497, 507, 517,\n       527, 538, 548, 558, 568, 578, 588, 599])\n\nI'll keep looking."
        },
        {
          "id": 819989,
          "postDate": "2020-04-25T04:35:20.827Z",
          "content": "<p>That's true, I was thinking of the range function. Linspace is doing what you described</p>",
          "rawMarkdown": "That's true, I was thinking of the range function. Linspace is doing what you described"
        },
        {
          "id": 823350,
          "postDate": "2020-04-27T15:28:09.933Z",
          "content": "<p><code>\nframe_count = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n        if frame_count &lt; 1:\n            frame_count = 300\n</code></p>\n\n<p>The paper said there will be degraded videos so fixing the framecount is not a good idea ... what happens if you get a 150 frames total frames (the paper specifically said that the private test set will have artificial downgrading of video quality both on resolution and frame rates, I believe the numbers that were mentioned were 480p and 15 fps) so if the meta is missing and it had -1 or 0 in the fps property and you sampled based on the 300 constant , you would overflow ...</p>",
          "rawMarkdown": "```\nframe_count = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n        if frame_count &lt; 1:\n            frame_count = 300\n```\n\nThe paper said there will be degraded videos so fixing the framecount is not a good idea ... what happens if you get a 150 frames total frames (the paper specifically said that the private test set will have artificial downgrading of video quality both on resolution and frame rates, I believe the numbers that were mentioned were 480p and 15 fps) so if the meta is missing and it had -1 or 0 in the fps property and you sampled based on the 300 constant , you would overflow ..."
        },
        {
          "id": 823413,
          "postDate": "2020-04-27T16:07:15.613Z",
          "content": "<p>Thanks for looking <a href=\"/ma7moudkhalid\">@ma7moudkhalid</a> </p>\n\n<p>The video processing is in a try/except block so even if it overflows it will just get skipped and receive the default 0.5 probability.</p>",
          "rawMarkdown": "Thanks for looking @ma7moudkhalid \n\nThe video processing is in a try/except block so even if it overflows it will just get skipped and receive the default 0.5 probability.",
          "votes": 1
        }
      ]
    },
    {
      "id": 820112,
      "postDate": "2020-04-25T06:31:53.213Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 819029,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2020-04-24T09:46:01.657000",
      "content": "<p>Good points. I think also as <a href=\"/vaillant\">@vaillant</a> raised <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145643#818587\">here </a>people might not have been filling completely failed predictions with a default prediction like 0.5 if their inference pipeline had never failed to parse a video in the public dataset. This is relevant giving their phrasing:</p>\n\n<p><code>There was a ~6.5% failure rate on re-runs, due mostly to incomplete submissions (submissions which failed to create predictions for all samples)</code></p>",
      "votes": 2,
      "replies": [
        {
          "id": 819038,
          "author_name": "DinoJuice",
          "author_url": "",
          "post_date": "2020-04-24T09:55:43.637000",
          "content": "<p>That looks like a plausible reason. After checking I noticed not all my error paths made sure the skipped file received a 0.5 prediction. 🤕 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 819401,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-04-24T15:13:44.400000",
          "content": "<p><a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122841#701781\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122841#701781</a></p>\n\n<p>I listed some of these previously. It is very unfortunate that these subtle gotchas were revealed on the test set. I am unsure of how kaggle will handle this. Some in the past have said it's the fault of the competitor for not properly hardening their code, but it's exceptionally difficult to make a perfect pipeline on completely unseen data. Don't think it should be a defensive programming competition, we're here for machine learning. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 823338,
      "author_name": "mjsML",
      "author_url": "",
      "post_date": "2020-04-27T15:20:13.360000",
      "content": "<p>variable resolution and variable frame rate videos (as described in the paper)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 820303,
      "author_name": "Carlo",
      "author_url": "",
      "post_date": "2020-04-25T10:19:31.790000",
      "content": "<p>Interesting points! There sure was a lot that could go wrong in pipelines for this competition.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 819156,
      "author_name": "maralski",
      "author_url": "",
      "post_date": "2020-04-24T11:47:40.353000",
      "content": "<p>Here is mine. Submission failed on private LB. Interested to see if you can spot anything out of the ordinary.</p>\n\n<p><a href=\"https://www.kaggle.com/maralski/fork-of-ensemble-of-5-networks?scriptVersionId=31169240\">https://www.kaggle.com/maralski/fork-of-ensemble-of-5-networks?scriptVersionId=31169240</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 819429,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-04-24T15:31:24.330000",
          "content": "<p><code>get_frameids = np.linspace(0, frame_count - 1, grab_frames, endpoint=True, dtype=np.int)</code></p>\n\n<p>this is the line that concerns me. Your grab_frames is saying to grab every 60 frames, but your frame_count is just the number of frames in the video so you are potentially putting an unbounded amount of video there. Unless I am missing somewhere else where this is protected against or I am misunderstanding the code, I think it might fall under the second point I raised. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 819983,
          "author_name": "maralski",
          "author_url": "",
          "post_date": "2020-04-25T04:26:20.640000",
          "content": "<p>Thanks for looking. That just grabs a total of 60 frames across the frame count. So if frame count is 300 it will grab:</p>\n\n<blockquote>\n  <p>[ins] In [6]: np.linspace(0, 300 - 1, 60, endpoint=True, dtype=np.int) <br>\n  Out[6]: \n  array([  0,   5,  10,  15,  20,  25,  30,  35,  40,  45,  50,  55,  60,\n          65,  70,  76,  81,  86,  91,  96, 101, 106, 111, 116, 121, 126,\n         131, 136, 141, 146, 152, 157, 162, 167, 172, 177, 182, 187, 192,\n         197, 202, 207, 212, 217, 222, 228, 233, 238, 243, 248, 253, 258,\n         263, 268, 273, 278, 283, 288, 293, 299])</p>\n</blockquote>\n\n<p>If it's 600 it will still grab 60.</p>\n\n<blockquote>\n  <p>[ins] In [7]: np.linspace(0, 600 - 1, 60, endpoint=True, dtype=np.int) <br>\n  Out[7]: \n  array([  0,  10,  20,  30,  40,  50,  60,  71,  81,  91, 101, 111, 121,\n         131, 142, 152, 162, 172, 182, 192, 203, 213, 223, 233, 243, 253,\n         263, 274, 284, 294, 304, 314, 324, 335, 345, 355, 365, 375, 385,\n         395, 406, 416, 426, 436, 446, 456, 467, 477, 487, 497, 507, 517,\n         527, 538, 548, 558, 568, 578, 588, 599])</p>\n</blockquote>\n\n<p>I'll keep looking.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 819989,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-04-25T04:35:20.827000",
          "content": "<p>That's true, I was thinking of the range function. Linspace is doing what you described</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 823350,
          "author_name": "mjsML",
          "author_url": "",
          "post_date": "2020-04-27T15:28:09.933000",
          "content": "<p><code>\nframe_count = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))\n        if frame_count &lt; 1:\n            frame_count = 300\n</code></p>\n\n<p>The paper said there will be degraded videos so fixing the framecount is not a good idea ... what happens if you get a 150 frames total frames (the paper specifically said that the private test set will have artificial downgrading of video quality both on resolution and frame rates, I believe the numbers that were mentioned were 480p and 15 fps) so if the meta is missing and it had -1 or 0 in the fps property and you sampled based on the 300 constant , you would overflow ...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 823413,
          "author_name": "maralski",
          "author_url": "",
          "post_date": "2020-04-27T16:07:15.613000",
          "content": "<p>Thanks for looking <a href=\"/ma7moudkhalid\">@ma7moudkhalid</a> </p>\n\n<p>The video processing is in a try/except block so even if it overflows it will just get skipped and receive the default 0.5 probability.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 820112,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-25T06:31:53.213000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "818973": "I will update this list more tomorrow, but I audited some people's work after they made their kernels public after the submission deadline and I noticed a couple of subtle errors that concerned me. \n\n1. writing to file and not cleaning up between predictions\n\n- In this kernel, the videos are rewritten, but then never cleared until the very end with the rm-r videos call. https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference/data\nThis can accumulate in disk space and eventually past a certain number of videos it will fail. \n\n2.Not putting boundaries on input size. \n\n- Some kernels did such things as sampling frames from the expected 300. if you sample every 10 frames it is only 30 frames and not that large in memory or to push to gpu, but we were given no guarantee of the length of the test set videos. Since they were in the wild it's entirely possible they were 60 fps and a minute long resulting in your input being a much larger size and potentially OOM'ing the entire kernel and killing it. It is better to put some bounds on the max size recorded from the videos. \n\n...to be continued when it's not 2am. \n\n",
    "819029": "Good points. I think also as @vaillant raised [here ](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145643#818587)people might not have been filling completely failed predictions with a default prediction like 0.5 if their inference pipeline had never failed to parse a video in the public dataset. This is relevant giving their phrasing:\n\n`There was a ~6.5% failure rate on re-runs, due mostly to incomplete submissions (submissions which failed to create predictions for all samples)`",
    "823338": "variable resolution and variable frame rate videos (as described in the paper)",
    "820303": "Interesting points! There sure was a lot that could go wrong in pipelines for this competition.",
    "819156": "Here is mine. Submission failed on private LB. Interested to see if you can spot anything out of the ordinary.\n\nhttps://www.kaggle.com/maralski/fork-of-ensemble-of-5-networks?scriptVersionId=31169240",
    "820112": ""
  }
}