{
  "id": 133875,
  "title": "CV2 and MTCNN takes ~10 hours on kaggle kernel?",
  "url": "/competitions/deepfake-detection-challenge/discussion/133875",
  "author_name": "OptimalFit",
  "post_date": "2020-03-04T18:39:08.097000",
  "votes": 0,
  "comment_count": 11,
  "views": 0,
  "content": "<p>CV2 and MTCNN pipeline will take me close to the maximum allowed time for predicting the test set. I am extracting all 300 frames from every video and would like to continue doing this preferably.</p>\n\n<p>Should I be optimising my pipeline or do I need to extract fewer frames?</p>",
  "messages": [
    {
      "id": 763863,
      "postDate": "2020-03-04T23:03:49.133Z",
      "content": "<p><code>\npd.pivot_table(ALL_VIDEOS_SUMMARY_DF, index=\"frame_count\", columns=\"label\"\n               , values=\"file_name\",aggfunc=len, margins=True, fill_value=0\n              ).astype(\"int\")\n</code></p>\n\n<p>|  FrameCount / Label   | FAKE  |REAL   |All|\n| --- | --- | --- | --- |\n| 0 | 8 | 0 | 8 |\n| 83 | 45 | 5 | 50 |\n| 99 | 98 |     7 |     105 | \n| 100 |     32 |    5 |     37 | \n| 109 |     3 |     3 |     6 | \n| 125 |     15 |    5 |     20 | \n| 134 |     10 |    5 |     15 | \n| 136 |     126 |   9 |     135 | \n| 143 |     34 |    5 |     39  | \n| 146 |     35 |    5 |     40 | \n| 147 |     108 |   9 |     117 | \n| 167 |     90 |    5 |     95 | \n| 193 |     13 |    5 |     18 | \n| 236 |     36 |    6 |     42 | \n| 237 |     35 |    5 |     40 | \n| 239 |     139 |   31 |    170 | \n| 240 |     535 |   45 |    580 | \n| 241 |     1832 |  152  | 1984 | \n| 242 |     17 |    5 |     22 | \n| 249 |     119 |   15 |    134 | \n| 250 |     18 |    5 |     23 | \n| 255 |     74 |    4 |     78 | \n| 269 |     144 |   16 |    160 | \n| 270 |     23 |    5 |     28 | \n| 288 |     215 |   11 |    226 | \n| 290 |     15 |    7 |     22 | \n| 292 |     95 |    5 |     100 | \n| 294 |     52 |    5 |     57 | \n| 295 |     25 |    6 |     31 | \n| 296 |     380 |   22 |    402 | \n| 297 |     903 |   69 |    972 | \n| 298 |     3598 |  299 |   3897 | \n| 299 |     2847 |  239 |   3086 | \n| 300 |     80514 |     16998 |     97512 | \n| 301 |     6920 |  981 |   7901 | \n| 302 |     777 |   132 |   909 | \n| 303 |     20 |    13 |    33 | \n| 600 | 35 |    5 |     40\n| 601 | 15 |    5 |     20 | \n| All   |  100000 |     19154 |     119154 |  </p>",
      "rawMarkdown": "```\npd.pivot_table(ALL_VIDEOS_SUMMARY_DF, index=\"frame_count\", columns=\"label\"\n               , values=\"file_name\",aggfunc=len, margins=True, fill_value=0\n              ).astype(\"int\")\n```\n\n\n|  FrameCount / Label \t| FAKE\t|REAL\t|All|\n| --- | --- | --- | --- |\n| 0\t| 8\t| 0\t| 8 |\n| 83 | 45 |\t5 |\t50 |\n| 99 | 98 | \t7 | \t105 | \n| 100 | \t32 | \t5 | \t37 | \n| 109 | \t3 | \t3 | \t6 | \n| 125 | \t15 | \t5 | \t20 | \n| 134 | \t10 | \t5 | \t15 | \n| 136 | \t126 | \t9 | \t135 | \n| 143 | \t34 | \t5 | \t39  | \n| 146 | \t35 | \t5 | \t40 | \n| 147 | \t108 |  \t9 | \t117 | \n| 167 | \t90 | \t5 | \t95 | \n| 193 | \t13 | \t5 | \t18 | \n| 236 | \t36 | \t6 | \t42 | \n| 237 | \t35 | \t5 | \t40 | \n| 239 | \t139 | \t31 | \t170 | \n| 240 | \t535 | \t45 | \t580 | \n| 241 | \t1832 | \t152\t | 1984 | \n| 242 | \t17 | \t5 | \t22 | \n| 249 | \t119 | \t15 | \t134 | \n| 250 | \t18 | \t5 | \t23 | \n| 255 | \t74 | \t4 | \t78 | \n| 269 | \t144 | \t16 | \t160 | \n| 270 | \t23 |  \t5 | \t28 | \n| 288 | \t215 | \t11 | \t226 | \n| 290 | \t15 | \t7 | \t22 | \n| 292 | \t95 | \t5 | \t100 | \n| 294 | \t52 | \t5 | \t57 | \n| 295 | \t25 | \t6 | \t31 | \n| 296 | \t380 | \t22 | \t402 | \n| 297 | \t903 | \t69 | \t972 | \n| 298 | \t3598 | \t299 | \t3897 | \n| 299 | \t2847 | \t239 | \t3086 | \n| 300 | \t80514 | \t16998 | \t97512 | \n| 301 | \t6920 | \t981 | \t7901 | \n| 302 | \t777 | \t132 | \t909 | \n| 303 | \t20 | \t13 | \t33 | \n| 600 | 35 | \t5 | \t40\n| 601 | 15 | \t5 | \t20 | \n| All\t|  100000 | \t19154 | \t119154 |  ",
      "votes": 3,
      "replies": [
        {
          "id": 764781,
          "postDate": "2020-03-05T22:22:54.117Z",
          "content": "<p>Thanks for this table, very interesting.</p>",
          "rawMarkdown": "Thanks for this table, very interesting."
        }
      ]
    },
    {
      "id": 764258,
      "postDate": "2020-03-05T09:22:26.550Z",
      "content": "<p>There's also no point doing your face detection on full res images. You can reduce the resolution and then upsample the bounding boxes.</p>",
      "rawMarkdown": "There's also no point doing your face detection on full res images. You can reduce the resolution and then upsample the bounding boxes.",
      "votes": 4,
      "replies": [
        {
          "id": 764783,
          "postDate": "2020-03-05T22:24:30.827Z",
          "content": "<p>great point</p>",
          "rawMarkdown": "great point"
        }
      ]
    },
    {
      "id": 763870,
      "postDate": "2020-03-04T23:17:03.077Z",
      "content": "<p>Even in the public training set, only 97 K out of 119 K videos have exactly 300 frames (see the pivot table below). Some have as low as 83 frames, while some others have 600+ frames. There is no way to know how many frames your algorithm will have to deal with while processing private test set. So, instead of hard-coding the number of frames I would recommend to treat this # as a parameter of your model and to see how you can fine-tune this parameter to reach a balance between performance and prediction quality.</p>",
      "rawMarkdown": "Even in the public training set, only 97 K out of 119 K videos have exactly 300 frames (see the pivot table below). Some have as low as 83 frames, while some others have 600+ frames. There is no way to know how many frames your algorithm will have to deal with while processing private test set. So, instead of hard-coding the number of frames I would recommend to treat this # as a parameter of your model and to see how you can fine-tune this parameter to reach a balance between performance and prediction quality.",
      "votes": 4
    },
    {
      "id": 763744,
      "postDate": "2020-03-04T20:04:55.973Z",
      "content": "<p>As suggested by <a href=\"/akashnandi\">@akashnandi</a>, try getting less frames when predicting. Also, have you tried running your code in parallel? Finally, if your function run time can be optimized by say 10% or more, then you might get under the limit.\nBest of luck! </p>",
      "rawMarkdown": "As suggested by @akashnandi, try getting less frames when predicting. Also, have you tried running your code in parallel? Finally, if your function run time can be optimized by say 10% or more, then you might get under the limit.\nBest of luck! ",
      "votes": 1,
      "replies": [
        {
          "id": 763751,
          "postDate": "2020-03-04T20:15:39.410Z",
          "content": "<p>cheers</p>",
          "rawMarkdown": "cheers"
        }
      ]
    },
    {
      "id": 763711,
      "postDate": "2020-03-04T19:10:44.817Z",
      "content": "<p>Well, don't extract 300 frames. Try 1 FPS. Should be decent.</p>",
      "rawMarkdown": "Well, don't extract 300 frames. Try 1 FPS. Should be decent.",
      "votes": 2,
      "replies": [
        {
          "id": 763724,
          "postDate": "2020-03-04T19:32:53.863Z",
          "content": "<p>might as well, at least to begin with :p</p>",
          "rawMarkdown": "might as well, at least to begin with :p"
        }
      ]
    },
    {
      "id": 764167,
      "postDate": "2020-03-05T07:41:23.240Z",
      "content": "<p>I bet even top10 guys don't use 300 frames for inference. 17 Frames can get you to 0.46lb. </p>",
      "rawMarkdown": "I bet even top10 guys don't use 300 frames for inference. 17 Frames can get you to 0.46lb. "
    },
    {
      "id": 763690,
      "postDate": "2020-03-04T18:39:08.097Z",
      "content": "<p>CV2 and MTCNN pipeline will take me close to the maximum allowed time for predicting the test set. I am extracting all 300 frames from every video and would like to continue doing this preferably.</p>\n\n<p>Should I be optimising my pipeline or do I need to extract fewer frames?</p>",
      "rawMarkdown": "CV2 and MTCNN pipeline will take me close to the maximum allowed time for predicting the test set. I am extracting all 300 frames from every video and would like to continue doing this preferably.\n\nShould I be optimising my pipeline or do I need to extract fewer frames?"
    },
    {
      "id": 763834,
      "postDate": "2020-03-04T22:25:02.217Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 763863,
      "author_name": "Vlad Pavlov",
      "author_url": "",
      "post_date": "2020-03-04T23:03:49.133000",
      "content": "<p><code>\npd.pivot_table(ALL_VIDEOS_SUMMARY_DF, index=\"frame_count\", columns=\"label\"\n               , values=\"file_name\",aggfunc=len, margins=True, fill_value=0\n              ).astype(\"int\")\n</code></p>\n\n<p>|  FrameCount / Label   | FAKE  |REAL   |All|\n| --- | --- | --- | --- |\n| 0 | 8 | 0 | 8 |\n| 83 | 45 | 5 | 50 |\n| 99 | 98 |     7 |     105 | \n| 100 |     32 |    5 |     37 | \n| 109 |     3 |     3 |     6 | \n| 125 |     15 |    5 |     20 | \n| 134 |     10 |    5 |     15 | \n| 136 |     126 |   9 |     135 | \n| 143 |     34 |    5 |     39  | \n| 146 |     35 |    5 |     40 | \n| 147 |     108 |   9 |     117 | \n| 167 |     90 |    5 |     95 | \n| 193 |     13 |    5 |     18 | \n| 236 |     36 |    6 |     42 | \n| 237 |     35 |    5 |     40 | \n| 239 |     139 |   31 |    170 | \n| 240 |     535 |   45 |    580 | \n| 241 |     1832 |  152  | 1984 | \n| 242 |     17 |    5 |     22 | \n| 249 |     119 |   15 |    134 | \n| 250 |     18 |    5 |     23 | \n| 255 |     74 |    4 |     78 | \n| 269 |     144 |   16 |    160 | \n| 270 |     23 |    5 |     28 | \n| 288 |     215 |   11 |    226 | \n| 290 |     15 |    7 |     22 | \n| 292 |     95 |    5 |     100 | \n| 294 |     52 |    5 |     57 | \n| 295 |     25 |    6 |     31 | \n| 296 |     380 |   22 |    402 | \n| 297 |     903 |   69 |    972 | \n| 298 |     3598 |  299 |   3897 | \n| 299 |     2847 |  239 |   3086 | \n| 300 |     80514 |     16998 |     97512 | \n| 301 |     6920 |  981 |   7901 | \n| 302 |     777 |   132 |   909 | \n| 303 |     20 |    13 |    33 | \n| 600 | 35 |    5 |     40\n| 601 | 15 |    5 |     20 | \n| All   |  100000 |     19154 |     119154 |  </p>",
      "votes": 3,
      "replies": [
        {
          "id": 764781,
          "author_name": "OptimalFit",
          "author_url": "",
          "post_date": "2020-03-05T22:22:54.117000",
          "content": "<p>Thanks for this table, very interesting.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 764258,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2020-03-05T09:22:26.550000",
      "content": "<p>There's also no point doing your face detection on full res images. You can reduce the resolution and then upsample the bounding boxes.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 764783,
          "author_name": "OptimalFit",
          "author_url": "",
          "post_date": "2020-03-05T22:24:30.827000",
          "content": "<p>great point</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 763870,
      "author_name": "Vlad Pavlov",
      "author_url": "",
      "post_date": "2020-03-04T23:17:03.077000",
      "content": "<p>Even in the public training set, only 97 K out of 119 K videos have exactly 300 frames (see the pivot table below). Some have as low as 83 frames, while some others have 600+ frames. There is no way to know how many frames your algorithm will have to deal with while processing private test set. So, instead of hard-coding the number of frames I would recommend to treat this # as a parameter of your model and to see how you can fine-tune this parameter to reach a balance between performance and prediction quality.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 763744,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-03-04T20:04:55.973000",
      "content": "<p>As suggested by <a href=\"/akashnandi\">@akashnandi</a>, try getting less frames when predicting. Also, have you tried running your code in parallel? Finally, if your function run time can be optimized by say 10% or more, then you might get under the limit.\nBest of luck! </p>",
      "votes": 1,
      "replies": [
        {
          "id": 763751,
          "author_name": "OptimalFit",
          "author_url": "",
          "post_date": "2020-03-04T20:15:39.410000",
          "content": "<p>cheers</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 763711,
      "author_name": "Akash",
      "author_url": "",
      "post_date": "2020-03-04T19:10:44.817000",
      "content": "<p>Well, don't extract 300 frames. Try 1 FPS. Should be decent.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 763724,
          "author_name": "OptimalFit",
          "author_url": "",
          "post_date": "2020-03-04T19:32:53.863000",
          "content": "<p>might as well, at least to begin with :p</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 764167,
      "author_name": "petya",
      "author_url": "",
      "post_date": "2020-03-05T07:41:23.240000",
      "content": "<p>I bet even top10 guys don't use 300 frames for inference. 17 Frames can get you to 0.46lb. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 763834,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-04T22:25:02.217000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "763863": "```\npd.pivot_table(ALL_VIDEOS_SUMMARY_DF, index=\"frame_count\", columns=\"label\"\n               , values=\"file_name\",aggfunc=len, margins=True, fill_value=0\n              ).astype(\"int\")\n```\n\n\n|  FrameCount / Label \t| FAKE\t|REAL\t|All|\n| --- | --- | --- | --- |\n| 0\t| 8\t| 0\t| 8 |\n| 83 | 45 |\t5 |\t50 |\n| 99 | 98 | \t7 | \t105 | \n| 100 | \t32 | \t5 | \t37 | \n| 109 | \t3 | \t3 | \t6 | \n| 125 | \t15 | \t5 | \t20 | \n| 134 | \t10 | \t5 | \t15 | \n| 136 | \t126 | \t9 | \t135 | \n| 143 | \t34 | \t5 | \t39  | \n| 146 | \t35 | \t5 | \t40 | \n| 147 | \t108 |  \t9 | \t117 | \n| 167 | \t90 | \t5 | \t95 | \n| 193 | \t13 | \t5 | \t18 | \n| 236 | \t36 | \t6 | \t42 | \n| 237 | \t35 | \t5 | \t40 | \n| 239 | \t139 | \t31 | \t170 | \n| 240 | \t535 | \t45 | \t580 | \n| 241 | \t1832 | \t152\t | 1984 | \n| 242 | \t17 | \t5 | \t22 | \n| 249 | \t119 | \t15 | \t134 | \n| 250 | \t18 | \t5 | \t23 | \n| 255 | \t74 | \t4 | \t78 | \n| 269 | \t144 | \t16 | \t160 | \n| 270 | \t23 |  \t5 | \t28 | \n| 288 | \t215 | \t11 | \t226 | \n| 290 | \t15 | \t7 | \t22 | \n| 292 | \t95 | \t5 | \t100 | \n| 294 | \t52 | \t5 | \t57 | \n| 295 | \t25 | \t6 | \t31 | \n| 296 | \t380 | \t22 | \t402 | \n| 297 | \t903 | \t69 | \t972 | \n| 298 | \t3598 | \t299 | \t3897 | \n| 299 | \t2847 | \t239 | \t3086 | \n| 300 | \t80514 | \t16998 | \t97512 | \n| 301 | \t6920 | \t981 | \t7901 | \n| 302 | \t777 | \t132 | \t909 | \n| 303 | \t20 | \t13 | \t33 | \n| 600 | 35 | \t5 | \t40\n| 601 | 15 | \t5 | \t20 | \n| All\t|  100000 | \t19154 | \t119154 |  ",
    "764258": "There's also no point doing your face detection on full res images. You can reduce the resolution and then upsample the bounding boxes.",
    "763870": "Even in the public training set, only 97 K out of 119 K videos have exactly 300 frames (see the pivot table below). Some have as low as 83 frames, while some others have 600+ frames. There is no way to know how many frames your algorithm will have to deal with while processing private test set. So, instead of hard-coding the number of frames I would recommend to treat this # as a parameter of your model and to see how you can fine-tune this parameter to reach a balance between performance and prediction quality.",
    "763744": "As suggested by @akashnandi, try getting less frames when predicting. Also, have you tried running your code in parallel? Finally, if your function run time can be optimized by say 10% or more, then you might get under the limit.\nBest of luck! ",
    "763711": "Well, don't extract 300 frames. Try 1 FPS. Should be decent.",
    "764167": "I bet even top10 guys don't use 300 frames for inference. 17 Frames can get you to 0.46lb. ",
    "763690": "CV2 and MTCNN pipeline will take me close to the maximum allowed time for predicting the test set. I am extracting all 300 frames from every video and would like to continue doing this preferably.\n\nShould I be optimising my pipeline or do I need to extract fewer frames?",
    "763834": ""
  }
}