{
  "id": 131843,
  "title": "[Taking Request] Notebook Publish",
  "url": "/competitions/deepfake-detection-challenge/discussion/131843",
  "author_name": "",
  "post_date": "2020-02-22T04:34:50.313296600Z",
  "votes": 25,
  "comment_count": 48,
  "views": 0,
  "content": "<p>Taking request for:\n1. Better face extractor(similar as blazeface but better accuracy, a balance between accuracy and speed)\n2. vanilla LRCN\n3. Underbalancing technique that will re-choose corresponding fake video every epoch(This way it can cover more data)(similar to overbalancing but better).</p>",
  "messages": [
    {
      "id": "753363",
      "postDate": "02/22/2020 04:34:50",
      "content": "<p>Taking request for:\n1. Better face extractor(similar as blazeface but better accuracy, a balance between accuracy and speed)\n2. vanilla LRCN\n3. Underbalancing technique that will re-choose corresponding fake video every epoch(This way it can cover more data)(similar to overbalancing but better).</p>",
      "rawMarkdown": "Taking request for:\n1. Better face extractor(similar as blazeface but better accuracy, a balance between accuracy and speed)\n2. vanilla LRCN\n3. Underbalancing technique that will re-choose corresponding fake video every epoch(This way it can cover more data)(similar to overbalancing but better).",
      "votes": null
    },
    {
      "id": "753395",
      "postDate": "02/22/2020 05:44:50",
      "content": "<p>Approved. 😃 </p>",
      "rawMarkdown": "Approved. 😃",
      "votes": null
    },
    {
      "id": "753752",
      "postDate": "02/22/2020 16:13:13",
      "content": "<p>All of the above mentioned seems interesting...</p>",
      "rawMarkdown": "All of the above mentioned seems interesting...",
      "votes": null
    },
    {
      "id": "753827",
      "postDate": "02/22/2020 17:32:22",
      "content": "<p>you know I am not much interested in 3 😁 . But eagerly waiting for 1 and 2.</p>",
      "rawMarkdown": "you know I am not much interested in 3 😁 . But eagerly waiting for 1 and 2.",
      "votes": null
    },
    {
      "id": "753873",
      "postDate": "02/22/2020 18:58:05",
      "content": "<p>1 is probably the most useful in practice, since the quality of the face recognition directly affects both training and final results.</p>\n\n<p>From an educative POV, 2 is certainly the most interesting. I don't even know what it is, if it's worth it, how it compares, etc.</p>\n\n<p>3 is IMHO not that necessary. I guess everyone has his/her own way of doing something similar.</p>\n\n<p>For 1,  a simple link would do. So I vote for 2 ;)  + the link for 1 ;)</p>",
      "rawMarkdown": "1 is probably the most useful in practice, since the quality of the face recognition directly affects both training and final results.\n\nFrom an educative POV, 2 is certainly the most interesting. I don't even know what it is, if it's worth it, how it compares, etc.\n\n3 is IMHO not that necessary. I guess everyone has his/her own way of doing something similar.\n\nFor 1,  a simple link would do. So I vote for 2 ;)  + the link for 1 ;)",
      "votes": null
    },
    {
      "id": "753968",
      "postDate": "02/22/2020 21:48:32",
      "content": "<p>Upvote this comment if you want notebook on 1st point. Choose 1 only. The first comment to get 12 upvotes wins the race and kernel will be published on that topic within few days.</p>",
      "rawMarkdown": "Upvote this comment if you want notebook on 1st point. Choose 1 only. The first comment to get 12 upvotes wins the race and kernel will be published on that topic within few days.",
      "votes": null
    },
    {
      "id": "753969",
      "postDate": "02/22/2020 21:48:40",
      "content": "<p>Upvote this comment if you want notebook on 2nd point. Choose 1 only. The first comment to get 12 upvotes wins the race and kernel will be published on that topic within few days.</p>",
      "rawMarkdown": "Upvote this comment if you want notebook on 2nd point. Choose 1 only. The first comment to get 12 upvotes wins the race and kernel will be published on that topic within few days.",
      "votes": null
    },
    {
      "id": "753970",
      "postDate": "02/22/2020 21:48:47",
      "content": "<p>Upvote this comment if you want notebook on 3rd point. Choose 1 only. The first comment to get 12 upvotes wins the race and kernel will be published on that topic within few days.</p>",
      "rawMarkdown": "Upvote this comment if you want notebook on 3rd point. Choose 1 only. The first comment to get 12 upvotes wins the race and kernel will be published on that topic within few days.",
      "votes": null
    },
    {
      "id": "754026",
      "postDate": "02/23/2020 00:09:35",
      "content": "<p>1 \nDid you try facenet-PyTorch?  I think it is better than blazeface. \nI found that submitting time using blazeface is an hour longer than using MTCNN.\nand according to <a href=\"https://www.kaggle.com/basharallabadi/yolov2-vs-faced-vs-blazeface-vs-mtcnn\"></a> , mtcnn has better accuracy.</p>",
      "rawMarkdown": "1 \nDid you try facenet-PyTorch?  I think it is better than blazeface. \nI found that submitting time using blazeface is an hour longer than using MTCNN.\nand according to [](https://www.kaggle.com/basharallabadi/yolov2-vs-faced-vs-blazeface-vs-mtcnn) , mtcnn has better accuracy.",
      "votes": null
    },
    {
      "id": "754061",
      "postDate": "02/23/2020 02:00:21",
      "content": "<p>Mine have better accuracy than MTCNN, facenet-pytorch and also have a competitive speed(about 30 seconds to process 300 frames on CPU).</p>",
      "rawMarkdown": "Mine have better accuracy than MTCNN, facenet-pytorch and also have a competitive speed(about 30 seconds to process 300 frames on CPU).",
      "votes": null
    },
    {
      "id": "754148",
      "postDate": "02/23/2020 05:41:06",
      "content": "<p>Ahem <a href=\"https://github.com/sfzhang15/FaceBoxes\">https://github.com/sfzhang15/FaceBoxes</a></p>",
      "rawMarkdown": "Ahem https://github.com/sfzhang15/FaceBoxes",
      "votes": null
    },
    {
      "id": "754199",
      "postDate": "02/23/2020 07:45:40",
      "content": "<p>Not this one <a href=\"/maralski\">@maralski</a>, it took me days to find out after knowing the concept of how it works.</p>",
      "rawMarkdown": "Not this one @maralski, it took me days to find out after knowing the concept of how it works.",
      "votes": null
    },
    {
      "id": "754323",
      "postDate": "02/23/2020 12:01:21",
      "content": "<blockquote>\n  <p>I found that submitting time using blazeface is an hour longer than using MTCNN</p>\n</blockquote>\n\n<p>Are you batching multiple images into the same request or are you just doing one image at a time? I haven't used MTCNN but for me BlazeFace is not the bottleneck (reading the videos is). Since BlazeFace is small, the key to getting good speed is to use large batches. (It is true that other face detection models may be better than BlazeFace, but that is usually the trade-off for more speed.)</p>",
      "rawMarkdown": "&gt; I found that submitting time using blazeface is an hour longer than using MTCNN\n\nAre you batching multiple images into the same request or are you just doing one image at a time? I haven't used MTCNN but for me BlazeFace is not the bottleneck (reading the videos is). Since BlazeFace is small, the key to getting good speed is to use large batches. (It is true that other face detection models may be better than BlazeFace, but that is usually the trade-off for more speed.)",
      "votes": null
    },
    {
      "id": "754325",
      "postDate": "02/23/2020 12:07:30",
      "content": "<p>On my pc, with a 1080 ti :- \nRead 10 frames equally gapped from a particular video, crops the faces, make them into one image, then, save the image.\nI can do this step ~2.3 it/s on my pc which is pretty good, you can consider the speed from here</p>",
      "rawMarkdown": "On my pc, with a 1080 ti :- \nRead 10 frames equally gapped from a particular video, crops the faces, make them into one image, then, save the image.\nI can do this step ~2.3 it/s on my pc which is pretty good, you can consider the speed from here",
      "votes": null
    },
    {
      "id": "754699",
      "postDate": "02/24/2020 00:14:36",
      "content": "<p>Hi, I jusr wonder how to judge the accuracy?</p>",
      "rawMarkdown": "Hi, I jusr wonder how to judge the accuracy?",
      "votes": null
    },
    {
      "id": "754732",
      "postDate": "02/24/2020 01:41:51",
      "content": "<p>Wow, I feel like I'm a small son who's looking for mommy's present :D</p>",
      "rawMarkdown": "Wow, I feel like I'm a small son who's looking for mommy's present :D",
      "votes": null
    },
    {
      "id": "754923",
      "postDate": "02/24/2020 08:13:40",
      "content": "<p>Unfortunately, I saw both of them having 12 upvotes at the same time, I guess the good thing is both of the kernels are coming out at once.</p>",
      "rawMarkdown": "Unfortunately, I saw both of them having 12 upvotes at the same time, I guess the good thing is both of the kernels are coming out at once.",
      "votes": null
    },
    {
      "id": "755553",
      "postDate": "02/24/2020 22:01:00",
      "content": "<p>Well I have switched to RetinaFace w/ ResNet backbone. Results look very good. Just need to add it to my pipeline now.</p>",
      "rawMarkdown": "Well I have switched to RetinaFace w/ ResNet backbone. Results look very good. Just need to add it to my pipeline now.",
      "votes": null
    },
    {
      "id": "755577",
      "postDate": "02/24/2020 22:45:30",
      "content": "<p>asides from clustering on perhaps features / embedding ... any ideas on how to separate identities in the train set (to create good validation set)?</p>",
      "rawMarkdown": "asides from clustering on perhaps features / embedding ... any ideas on how to separate identities in the train set (to create good validation set)?",
      "votes": null
    },
    {
      "id": "756425",
      "postDate": "02/25/2020 18:01:38",
      "content": "<p>UPDATE: finished draft notebook of number one.</p>",
      "rawMarkdown": "UPDATE: finished draft notebook of number one.",
      "votes": null
    },
    {
      "id": "756623",
      "postDate": "02/25/2020 22:42:40",
      "content": "<p>last 10 folders have majority actors different than the rest of the folders.</p>",
      "rawMarkdown": "last 10 folders have majority actors different than the rest of the folders.",
      "votes": null
    },
    {
      "id": "757404",
      "postDate": "02/26/2020 18:26:52",
      "content": "<p><a href=\"/yzcwansui\">@yzcwansui</a> I didn't use a metric to judge accuracy BUT you can obviously see that blazeface have a bad accuracy(it often have a long bounding box); and mtcnn, mobilenet are pretty close. BTW blazeface often unable to detect face.</p>",
      "rawMarkdown": "yzcwansui I didn't use a metric to judge accuracy BUT you can obviously see that blazeface have a bad accuracy(it often have a long bounding box); and mtcnn, mobilenet are pretty close. BTW blazeface often unable to detect face.",
      "votes": null
    },
    {
      "id": "757776",
      "postDate": "02/27/2020 05:04:10",
      "content": "<p>now I'm comparing blazeface, mtcnn and that face extractor. Any request for other networks for comparison?</p>",
      "rawMarkdown": "now I'm comparing blazeface, mtcnn and that face extractor. Any request for other networks for comparison?",
      "votes": null
    },
    {
      "id": "757816",
      "postDate": "02/27/2020 06:04:20",
      "content": "<p>Maybe you can add DSFD and YoloV3 to the comparison</p>",
      "rawMarkdown": "Maybe you can add DSFD and YoloV3 to the comparison",
      "votes": null
    },
    {
      "id": "758067",
      "postDate": "02/27/2020 12:07:34",
      "content": "<p>We know just because of how our models work that DSFD have the best accuracy to even any single model, I mean it is similar to our model, but a lot slower, I had tested it a month or two ago, when I was searching an optimal face detector. That is why I can not recall exact numbers for now. We might be implementing YoloV3, Thank You. <a href=\"/ngcferreira\">@ngcferreira</a> </p>",
      "rawMarkdown": "We know just because of how our models work that DSFD have the best accuracy to even any single model, I mean it is similar to our model, but a lot slower, I had tested it a month or two ago, when I was searching an optimal face detector. That is why I can not recall exact numbers for now. We might be implementing YoloV3, Thank You. @ngcferreira",
      "votes": null
    },
    {
      "id": "759358",
      "postDate": "02/28/2020 23:33:49",
      "content": "<p>Final Draft and Published. Please let me know if there's any bug in the code. Thanks. \nComparisons: <a href=\"https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-comparison\">link</a>\nClean Helper Code: <a href=\"https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-helper-code\">link</a></p>",
      "rawMarkdown": "Final Draft and Published. Please let me know if there's any bug in the code. Thanks. \nComparisons: [link](https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-comparison)\nClean Helper Code: [link](https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-helper-code)",
      "votes": null
    },
    {
      "id": "759373",
      "postDate": "02/29/2020 00:21:49",
      "content": "<p>the real thing is only the third part is the one which makes the difference which is actually considerable, I guess if these 2 notebooks are out and we give the third one too, we will loose all our kernels, as they all are made with slight changes, only 1 was meant to come out, but I guess now it is these 2.</p>",
      "rawMarkdown": "the real thing is only the third part is the one which makes the difference which is actually considerable, I guess if these 2 notebooks are out and we give the third one too, we will loose all our kernels, as they all are made with slight changes, only 1 was meant to come out, but I guess now it is these 2.",
      "votes": null
    },
    {
      "id": "763720",
      "postDate": "03/04/2020 19:31:19",
      "content": "<p>Well done, it looks good. Thanks for the kernel.</p>",
      "rawMarkdown": "Well done, it looks good. Thanks for the kernel.",
      "votes": null
    },
    {
      "id": "763915",
      "postDate": "03/05/2020 01:09:12",
      "content": "<p>May I ask just a question: the face detector works better on real face. For fake videos, it may not detect  blurred (fake) faces. So using a face detector to generate training samples has this issue: the extracted faces for fake videos don't correspond to real faces in the real video. So how did you do it? I think you fine-tune the face detector with correct fake faces (by translating bounding box from real video to fake video)?</p>",
      "rawMarkdown": "May I ask just a question: the face detector works better on real face. For fake videos, it may not detect  blurred (fake) faces. So using a face detector to generate training samples has this issue: the extracted faces for fake videos don't correspond to real faces in the real video. So how did you do it? I think you fine-tune the face detector with correct fake faces (by translating bounding box from real video to fake video)?",
      "votes": null
    },
    {
      "id": "763949",
      "postDate": "03/05/2020 02:31:39",
      "content": "<p><a href=\"/khahuras\">@khahuras</a> \nIts not even complicated at all, I am not currently doing anything to bypass this weakness intentionally, cuz I do want some noise and wrongs in the dataset so the model will always be learning, but to overcome the issue, take same frame from real video and same from fake video, and get the bounding boxes from real frame and apply the same on the fake frame, with a slight random distortion to make model think that they are actually made by face detector and are completely random, take the face from the fake frame, and you are done.</p>",
      "rawMarkdown": "khahuras \nIts not even complicated at all, I am not currently doing anything to bypass this weakness intentionally, cuz I do want some noise and wrongs in the dataset so the model will always be learning, but to overcome the issue, take same frame from real video and same from fake video, and get the bounding boxes from real frame and apply the same on the fake frame, with a slight random distortion to make model think that they are actually made by face detector and are completely random, take the face from the fake frame, and you are done.",
      "votes": null
    },
    {
      "id": "764000",
      "postDate": "03/05/2020 04:08:31",
      "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> Thanks for your kindness! Honestly, given your advice, I think it's not a major issue that comes from face detector, even if we introduce some randomness/distortion as you suggested. I think the problem comes from my validation. Every time a trained classifier produces more aggressive predictions towards 0/1, the worse LB score, even though better CV. That means my CV has leakage or something like that. I should find a better validation scheme. Now my validation is folder split with last 10 folders as validation, and validation of those last 10 folders is around 0.25.</p>",
      "rawMarkdown": "harshitsheoran Thanks for your kindness! Honestly, given your advice, I think it's not a major issue that comes from face detector, even if we introduce some randomness/distortion as you suggested. I think the problem comes from my validation. Every time a trained classifier produces more aggressive predictions towards 0/1, the worse LB score, even though better CV. That means my CV has leakage or something like that. I should find a better validation scheme. Now my validation is folder split with last 10 folders as validation, and validation of those last 10 folders is around 0.25.",
      "votes": null
    },
    {
      "id": "764107",
      "postDate": "03/05/2020 06:33:00",
      "content": "<p>The problem is not the CV, it is about that your model can not learn manage loss while predicting, note, going towards 1 or 0 only helps when the classifier have a very good accuracy change your model arch to get the model just right, train 1 epoch, in my case, I can just check with training one epoch, as if start is good, I know that the final result after training more epochs will also be good.</p>\n\n<p>EDIT, We do have same data for validation and trust me there was at time when I also had validation \naround 0.25, changing data and checking it on the bench (this bench is trustful), this bench does not tell how much the different is exactly, but it surely tells me that which model is better for example, my 0.318 model have cv on bench 0.16, and my 0.340 had cv on bench 0.18*.</p>",
      "rawMarkdown": "The problem is not the CV, it is about that your model can not learn manage loss while predicting, note, going towards 1 or 0 only helps when the classifier have a very good accuracy change your model arch to get the model just right, train 1 epoch, in my case, I can just check with training one epoch, as if start is good, I know that the final result after training more epochs will also be good.\n\nEDIT, We do have same data for validation and trust me there was at time when I also had validation \naround 0.25, changing data and checking it on the bench (this bench is trustful), this bench does not tell how much the different is exactly, but it surely tells me that which model is better for example, my 0.318 model have cv on bench 0.16, and my 0.340 had cv on bench 0.18*.",
      "votes": null
    },
    {
      "id": "764134",
      "postDate": "03/05/2020 06:53:03",
      "content": "<p>Thanks again so much :). I will try to surpass your team!</p>",
      "rawMarkdown": "Thanks again so much :). I will try to surpass your team!",
      "votes": null
    },
    {
      "id": "764140",
      "postDate": "03/05/2020 07:00:45",
      "content": "<p>Good Luck, cuz definately we are not even close to done improving yet, LRCN is taking our score to new heights.</p>",
      "rawMarkdown": "Good Luck, cuz definately we are not even close to done improving yet, LRCN is taking our score to new heights.",
      "votes": null
    },
    {
      "id": "764143",
      "postDate": "03/05/2020 07:16:00",
      "content": "<p>Given the amount of things you share in this competition, I totally am rooting for your team to get the possibly best result!</p>",
      "rawMarkdown": "Given the amount of things you share in this competition, I totally am rooting for your team to get the possibly best result!",
      "votes": null
    },
    {
      "id": "774896",
      "postDate": "03/16/2020 03:29:23",
      "content": "<p>I made publish <a href=\"https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference\">this</a> notebook. It is an inference of LRCN using the face extractor I shared before. \nI didn't make publish of the weights. I'm sure people(including me) doesn't like high LB notebook. But I'm OK with making it public if people want me to. I'm taking request. </p>",
      "rawMarkdown": "I made publish [this](https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference) notebook. It is an inference of LRCN using the face extractor I shared before. \nI didn't make publish of the weights. I'm sure people(including me) doesn't like high LB notebook. But I'm OK with making it public if people want me to. I'm taking request.",
      "votes": null
    },
    {
      "id": "774900",
      "postDate": "03/16/2020 03:35:25",
      "content": "<p>Thanks a lot! I believe, what you shared is good for people to try LRCN. Sharing weights will defeat the purpose as competition is coming closer to the end.</p>",
      "rawMarkdown": "Thanks a lot! I believe, what you shared is good for people to try LRCN. Sharing weights will defeat the purpose as competition is coming closer to the end.",
      "votes": null
    },
    {
      "id": "774939",
      "postDate": "03/16/2020 04:45:11",
      "content": "<p>Well, the weights will only get you to 0.38 so, we dont have a problem to make it publically available, the reason not to is the same as <a href=\"/debanga\">@debanga</a> 's.</p>",
      "rawMarkdown": "Well, the weights will only get you to 0.38 so, we dont have a problem to make it publically available, the reason not to is the same as @debanga 's.",
      "votes": null
    },
    {
      "id": "775040",
      "postDate": "03/16/2020 08:14:27",
      "content": "<p>Thanks for sharing the notebook!</p>\n\n<p>With regard to the weights, I agree with Zenify. While sharing the weights would not affect the top solutions, it would certainly remove all those better than 0.43 from medal zone. Apart from that, I doubt that there is much to learn from the weights.</p>",
      "rawMarkdown": "Thanks for sharing the notebook!\n\nWith regard to the weights, I agree with Zenify. While sharing the weights would not affect the top solutions, it would certainly remove all those better than 0.43 from medal zone. Apart from that, I doubt that there is much to learn from the weights.",
      "votes": null
    },
    {
      "id": "775157",
      "postDate": "03/16/2020 11:11:32",
      "content": "<p>If you have to learn, we have summaried our exact model which we used until 0.31, there.</p>",
      "rawMarkdown": "If you have to learn, we have summaried our exact model which we used until 0.31, there.",
      "votes": null
    },
    {
      "id": "775366",
      "postDate": "03/16/2020 15:52:07",
      "content": "<p>Some people are requesting the training code. I published <a href=\"https://www.kaggle.com/unkownhihi/dfdc-lrcn-training/\">this</a> notebook. Its really nothing special, just load data and train. </p>\n\n<p>I removed the dataset because I want to avoid people just clone, run, submit. \nI'm OK with providing the weight/dataset, because we already started doing a new thing.</p>\n\n<p>I'm OK with making public the training dataset, making public the weights, but still, I'm taking request.</p>",
      "rawMarkdown": "Some people are requesting the training code. I published [this](https://www.kaggle.com/unkownhihi/dfdc-lrcn-training/) notebook. Its really nothing special, just load data and train. \n\nI removed the dataset because I want to avoid people just clone, run, submit. \nI'm OK with providing the weight/dataset, because we already started doing a new thing.\n\nI'm OK with making public the training dataset, making public the weights, but still, I'm taking request.",
      "votes": null
    },
    {
      "id": "775475",
      "postDate": "03/16/2020 18:03:08",
      "content": "<p>Please don't publish the weights. It will undermine the 3 month effort of a lot of people who worked hard to get to wherever they got in the Leaderboard. The whole idea of the Leaderboard is to give some reward to the people competing. Don't take that away from them. Imagine how you would you feel if someone above you did the same thing and your team ends up in 2000th place at the end of the competition?</p>",
      "rawMarkdown": "Please don't publish the weights. It will undermine the 3 month effort of a lot of people who worked hard to get to wherever they got in the Leaderboard. The whole idea of the Leaderboard is to give some reward to the people competing. Don't take that away from them. Imagine how you would you feel if someone above you did the same thing and your team ends up in 2000th place at the end of the competition?",
      "votes": null
    },
    {
      "id": "775476",
      "postDate": "03/16/2020 18:05:04",
      "content": "<p><a href=\"/cesb45\">@cesb45</a> I agree with you. I felt the same when there's a high scoring public notebook and our score is below it. We worked extra hard to get out of the \"bubble\". </p>",
      "rawMarkdown": "cesb45 I agree with you. I felt the same when there's a high scoring public notebook and our score is below it. We worked extra hard to get out of the \"bubble\".",
      "votes": null
    },
    {
      "id": "778941",
      "postDate": "03/18/2020 21:45:20",
      "content": "<p>Maybe you can upload only partial of your dataset in a separate Kaggle Dataset folder, then we will get a better idea on how the dataset is been prepare while no one can simply just clone, run, submit without putting in any effort.</p>",
      "rawMarkdown": "Maybe you can upload only partial of your dataset in a separate Kaggle Dataset folder, then we will get a better idea on how the dataset is been prepare while no one can simply just clone, run, submit without putting in any effort.",
      "votes": null
    },
    {
      "id": "778948",
      "postDate": "03/18/2020 21:56:26",
      "content": "<p>Well, you know, I don't want people to just like copy our method. All I wanted is to give them an idea how its achieved. I can give you a demo image:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2F168e1beaac474a9757fb3d0d6b45b171%2Faaqaifqrwn.jpg?generation=1584568583604380&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Well, you know, I don't want people to just like copy our method. All I wanted is to give them an idea how its achieved. I can give you a demo image:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2F168e1beaac474a9757fb3d0d6b45b171%2Faaqaifqrwn.jpg?generation=1584568583604380&amp;alt=media)",
      "votes": null
    },
    {
      "id": "781319",
      "postDate": "03/21/2020 06:28:05",
      "content": "<p>Really appreciate for your constantly sharing in this completion, it inspired me a lot. </p>\n\n<p>May I ask how do you deal with multiple faces in one frame? I tried to use tensor like this [Batch, Frames, Faces, C, H, W], and reshape it to [Batch x Frames x Faces, C, H, W] to backbone-CNN then  [Batch, Frames x Face, CNN-features] to Sequential model. But it seems not working and looks like just to add more noise to the model.</p>",
      "rawMarkdown": "Really appreciate for your constantly sharing in this completion, it inspired me a lot. \n\nMay I ask how do you deal with multiple faces in one frame? I tried to use tensor like this [Batch, Frames, Faces, C, H, W], and reshape it to [Batch x Frames x Faces, C, H, W] to backbone-CNN then  [Batch, Frames x Face, CNN-features] to Sequential model. But it seems not working and looks like just to add more noise to the model.",
      "votes": null
    },
    {
      "id": "781673",
      "postDate": "03/21/2020 14:59:25",
      "content": "<p>We ignore it.... It doesn't affect that much, especially its 10 frames data, and will very likely include the second person as well. The model should be able to learn about 2 people situations.</p>",
      "rawMarkdown": "We ignore it.... It doesn't affect that much, especially its 10 frames data, and will very likely include the second person as well. The model should be able to learn about 2 people situations.",
      "votes": null
    },
    {
      "id": "784780",
      "postDate": "03/24/2020 13:59:12",
      "content": "<p>So did you finally publish the dataset? <a href=\"/harshitsheoran\">@harshitsheoran</a> </p>",
      "rawMarkdown": "So did you finally publish the dataset? @harshitsheoran",
      "votes": null
    },
    {
      "id": "784789",
      "postDate": "03/24/2020 14:03:01",
      "content": "<p>which dataset, and if you are talking about something you think I am only doing cuz I do not have private space left, dont disclose it on kaggle. We are only making some of our datas public cuz first we do not have enough space left to get our models in and data in for training and all, so please dont access them cuz they are waste of time, not gonna make you score to our best sub.</p>",
      "rawMarkdown": "which dataset, and if you are talking about something you think I am only doing cuz I do not have private space left, dont disclose it on kaggle. We are only making some of our datas public cuz first we do not have enough space left to get our models in and data in for training and all, so please dont access them cuz they are waste of time, not gonna make you score to our best sub.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 753395,
      "author_name": "khahuras",
      "author_url": "",
      "post_date": "02/22/2020 05:44:50",
      "content": "<p>Approved. 😃 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 753752,
      "author_name": "dandrocec",
      "author_url": "",
      "post_date": "02/22/2020 16:13:13",
      "content": "<p>All of the above mentioned seems interesting...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 753827,
      "author_name": "ankitsainiankit",
      "author_url": "",
      "post_date": "02/22/2020 17:32:22",
      "content": "<p>you know I am not much interested in 3 😁 . But eagerly waiting for 1 and 2.</p>",
      "votes": null,
      "replies": [
        {
          "id": 759373,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/29/2020 00:21:49",
          "content": "<p>the real thing is only the third part is the one which makes the difference which is actually considerable, I guess if these 2 notebooks are out and we give the third one too, we will loose all our kernels, as they all are made with slight changes, only 1 was meant to come out, but I guess now it is these 2.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 753873,
      "author_name": "dagnelies",
      "author_url": "",
      "post_date": "02/22/2020 18:58:05",
      "content": "<p>1 is probably the most useful in practice, since the quality of the face recognition directly affects both training and final results.</p>\n\n<p>From an educative POV, 2 is certainly the most interesting. I don't even know what it is, if it's worth it, how it compares, etc.</p>\n\n<p>3 is IMHO not that necessary. I guess everyone has his/her own way of doing something similar.</p>\n\n<p>For 1,  a simple link would do. So I vote for 2 ;)  + the link for 1 ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 753968,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "02/22/2020 21:48:32",
      "content": "<p>Upvote this comment if you want notebook on 1st point. Choose 1 only. The first comment to get 12 upvotes wins the race and kernel will be published on that topic within few days.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 753969,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "02/22/2020 21:48:40",
      "content": "<p>Upvote this comment if you want notebook on 2nd point. Choose 1 only. The first comment to get 12 upvotes wins the race and kernel will be published on that topic within few days.</p>",
      "votes": null,
      "replies": [
        {
          "id": 754732,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "02/24/2020 01:41:51",
          "content": "<p>Wow, I feel like I'm a small son who's looking for mommy's present :D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 784780,
          "author_name": "",
          "author_url": "",
          "post_date": "03/24/2020 13:59:12",
          "content": "<p>So did you finally publish the dataset? <a href=\"/harshitsheoran\">@harshitsheoran</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 784789,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "03/24/2020 14:03:01",
          "content": "<p>which dataset, and if you are talking about something you think I am only doing cuz I do not have private space left, dont disclose it on kaggle. We are only making some of our datas public cuz first we do not have enough space left to get our models in and data in for training and all, so please dont access them cuz they are waste of time, not gonna make you score to our best sub.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 753970,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "02/22/2020 21:48:47",
      "content": "<p>Upvote this comment if you want notebook on 3rd point. Choose 1 only. The first comment to get 12 upvotes wins the race and kernel will be published on that topic within few days.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 754026,
      "author_name": "",
      "author_url": "",
      "post_date": "02/23/2020 00:09:35",
      "content": "<p>1 \nDid you try facenet-PyTorch?  I think it is better than blazeface. \nI found that submitting time using blazeface is an hour longer than using MTCNN.\nand according to <a href=\"https://www.kaggle.com/basharallabadi/yolov2-vs-faced-vs-blazeface-vs-mtcnn\"></a> , mtcnn has better accuracy.</p>",
      "votes": null,
      "replies": [
        {
          "id": 754061,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "02/23/2020 02:00:21",
          "content": "<p>Mine have better accuracy than MTCNN, facenet-pytorch and also have a competitive speed(about 30 seconds to process 300 frames on CPU).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 754148,
          "author_name": "maralski",
          "author_url": "",
          "post_date": "02/23/2020 05:41:06",
          "content": "<p>Ahem <a href=\"https://github.com/sfzhang15/FaceBoxes\">https://github.com/sfzhang15/FaceBoxes</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 754199,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/23/2020 07:45:40",
          "content": "<p>Not this one <a href=\"/maralski\">@maralski</a>, it took me days to find out after knowing the concept of how it works.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 754323,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "02/23/2020 12:01:21",
          "content": "<blockquote>\n  <p>I found that submitting time using blazeface is an hour longer than using MTCNN</p>\n</blockquote>\n\n<p>Are you batching multiple images into the same request or are you just doing one image at a time? I haven't used MTCNN but for me BlazeFace is not the bottleneck (reading the videos is). Since BlazeFace is small, the key to getting good speed is to use large batches. (It is true that other face detection models may be better than BlazeFace, but that is usually the trade-off for more speed.)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 754325,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/23/2020 12:07:30",
          "content": "<p>On my pc, with a 1080 ti :- \nRead 10 frames equally gapped from a particular video, crops the faces, make them into one image, then, save the image.\nI can do this step ~2.3 it/s on my pc which is pretty good, you can consider the speed from here</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 754699,
          "author_name": "",
          "author_url": "",
          "post_date": "02/24/2020 00:14:36",
          "content": "<p>Hi, I jusr wonder how to judge the accuracy?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755553,
          "author_name": "maralski",
          "author_url": "",
          "post_date": "02/24/2020 22:01:00",
          "content": "<p>Well I have switched to RetinaFace w/ ResNet backbone. Results look very good. Just need to add it to my pipeline now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 757404,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "02/26/2020 18:26:52",
          "content": "<p><a href=\"/yzcwansui\">@yzcwansui</a> I didn't use a metric to judge accuracy BUT you can obviously see that blazeface have a bad accuracy(it often have a long bounding box); and mtcnn, mobilenet are pretty close. BTW blazeface often unable to detect face.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 754923,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "02/24/2020 08:13:40",
      "content": "<p>Unfortunately, I saw both of them having 12 upvotes at the same time, I guess the good thing is both of the kernels are coming out at once.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 755577,
      "author_name": "ma7moud",
      "author_url": "",
      "post_date": "02/24/2020 22:45:30",
      "content": "<p>asides from clustering on perhaps features / embedding ... any ideas on how to separate identities in the train set (to create good validation set)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 756623,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/25/2020 22:42:40",
          "content": "<p>last 10 folders have majority actors different than the rest of the folders.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 756425,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "02/25/2020 18:01:38",
      "content": "<p>UPDATE: finished draft notebook of number one.</p>",
      "votes": null,
      "replies": [
        {
          "id": 757776,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "02/27/2020 05:04:10",
          "content": "<p>now I'm comparing blazeface, mtcnn and that face extractor. Any request for other networks for comparison?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 757816,
          "author_name": "ngcferreira",
          "author_url": "",
          "post_date": "02/27/2020 06:04:20",
          "content": "<p>Maybe you can add DSFD and YoloV3 to the comparison</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 758067,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/27/2020 12:07:34",
          "content": "<p>We know just because of how our models work that DSFD have the best accuracy to even any single model, I mean it is similar to our model, but a lot slower, I had tested it a month or two ago, when I was searching an optimal face detector. That is why I can not recall exact numbers for now. We might be implementing YoloV3, Thank You. <a href=\"/ngcferreira\">@ngcferreira</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 759358,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "02/28/2020 23:33:49",
      "content": "<p>Final Draft and Published. Please let me know if there's any bug in the code. Thanks. \nComparisons: <a href=\"https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-comparison\">link</a>\nClean Helper Code: <a href=\"https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-helper-code\">link</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 763720,
          "author_name": "ngcferreira",
          "author_url": "",
          "post_date": "03/04/2020 19:31:19",
          "content": "<p>Well done, it looks good. Thanks for the kernel.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 763915,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "03/05/2020 01:09:12",
          "content": "<p>May I ask just a question: the face detector works better on real face. For fake videos, it may not detect  blurred (fake) faces. So using a face detector to generate training samples has this issue: the extracted faces for fake videos don't correspond to real faces in the real video. So how did you do it? I think you fine-tune the face detector with correct fake faces (by translating bounding box from real video to fake video)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 763949,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "03/05/2020 02:31:39",
          "content": "<p><a href=\"/khahuras\">@khahuras</a> \nIts not even complicated at all, I am not currently doing anything to bypass this weakness intentionally, cuz I do want some noise and wrongs in the dataset so the model will always be learning, but to overcome the issue, take same frame from real video and same from fake video, and get the bounding boxes from real frame and apply the same on the fake frame, with a slight random distortion to make model think that they are actually made by face detector and are completely random, take the face from the fake frame, and you are done.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764000,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "03/05/2020 04:08:31",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> Thanks for your kindness! Honestly, given your advice, I think it's not a major issue that comes from face detector, even if we introduce some randomness/distortion as you suggested. I think the problem comes from my validation. Every time a trained classifier produces more aggressive predictions towards 0/1, the worse LB score, even though better CV. That means my CV has leakage or something like that. I should find a better validation scheme. Now my validation is folder split with last 10 folders as validation, and validation of those last 10 folders is around 0.25.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764107,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "03/05/2020 06:33:00",
          "content": "<p>The problem is not the CV, it is about that your model can not learn manage loss while predicting, note, going towards 1 or 0 only helps when the classifier have a very good accuracy change your model arch to get the model just right, train 1 epoch, in my case, I can just check with training one epoch, as if start is good, I know that the final result after training more epochs will also be good.</p>\n\n<p>EDIT, We do have same data for validation and trust me there was at time when I also had validation \naround 0.25, changing data and checking it on the bench (this bench is trustful), this bench does not tell how much the different is exactly, but it surely tells me that which model is better for example, my 0.318 model have cv on bench 0.16, and my 0.340 had cv on bench 0.18*.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764134,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "03/05/2020 06:53:03",
          "content": "<p>Thanks again so much :). I will try to surpass your team!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764140,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "03/05/2020 07:00:45",
          "content": "<p>Good Luck, cuz definately we are not even close to done improving yet, LRCN is taking our score to new heights.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764143,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "03/05/2020 07:16:00",
          "content": "<p>Given the amount of things you share in this competition, I totally am rooting for your team to get the possibly best result!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 774896,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "03/16/2020 03:29:23",
      "content": "<p>I made publish <a href=\"https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference\">this</a> notebook. It is an inference of LRCN using the face extractor I shared before. \nI didn't make publish of the weights. I'm sure people(including me) doesn't like high LB notebook. But I'm OK with making it public if people want me to. I'm taking request. </p>",
      "votes": null,
      "replies": [
        {
          "id": 774900,
          "author_name": "debanga",
          "author_url": "",
          "post_date": "03/16/2020 03:35:25",
          "content": "<p>Thanks a lot! I believe, what you shared is good for people to try LRCN. Sharing weights will defeat the purpose as competition is coming closer to the end.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 774939,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "03/16/2020 04:45:11",
          "content": "<p>Well, the weights will only get you to 0.38 so, we dont have a problem to make it publically available, the reason not to is the same as <a href=\"/debanga\">@debanga</a> 's.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 775040,
          "author_name": "catochris",
          "author_url": "",
          "post_date": "03/16/2020 08:14:27",
          "content": "<p>Thanks for sharing the notebook!</p>\n\n<p>With regard to the weights, I agree with Zenify. While sharing the weights would not affect the top solutions, it would certainly remove all those better than 0.43 from medal zone. Apart from that, I doubt that there is much to learn from the weights.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 775157,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "03/16/2020 11:11:32",
          "content": "<p>If you have to learn, we have summaried our exact model which we used until 0.31, there.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 775475,
          "author_name": "cesb45",
          "author_url": "",
          "post_date": "03/16/2020 18:03:08",
          "content": "<p>Please don't publish the weights. It will undermine the 3 month effort of a lot of people who worked hard to get to wherever they got in the Leaderboard. The whole idea of the Leaderboard is to give some reward to the people competing. Don't take that away from them. Imagine how you would you feel if someone above you did the same thing and your team ends up in 2000th place at the end of the competition?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 775476,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "03/16/2020 18:05:04",
          "content": "<p><a href=\"/cesb45\">@cesb45</a> I agree with you. I felt the same when there's a high scoring public notebook and our score is below it. We worked extra hard to get out of the \"bubble\". </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 775366,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "03/16/2020 15:52:07",
      "content": "<p>Some people are requesting the training code. I published <a href=\"https://www.kaggle.com/unkownhihi/dfdc-lrcn-training/\">this</a> notebook. Its really nothing special, just load data and train. </p>\n\n<p>I removed the dataset because I want to avoid people just clone, run, submit. \nI'm OK with providing the weight/dataset, because we already started doing a new thing.</p>\n\n<p>I'm OK with making public the training dataset, making public the weights, but still, I'm taking request.</p>",
      "votes": null,
      "replies": [
        {
          "id": 778941,
          "author_name": "chewkokwahibrainai",
          "author_url": "",
          "post_date": "03/18/2020 21:45:20",
          "content": "<p>Maybe you can upload only partial of your dataset in a separate Kaggle Dataset folder, then we will get a better idea on how the dataset is been prepare while no one can simply just clone, run, submit without putting in any effort.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 778948,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "03/18/2020 21:56:26",
          "content": "<p>Well, you know, I don't want people to just like copy our method. All I wanted is to give them an idea how its achieved. I can give you a demo image:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2F168e1beaac474a9757fb3d0d6b45b171%2Faaqaifqrwn.jpg?generation=1584568583604380&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 781319,
          "author_name": "kentchun33333",
          "author_url": "",
          "post_date": "03/21/2020 06:28:05",
          "content": "<p>Really appreciate for your constantly sharing in this completion, it inspired me a lot. </p>\n\n<p>May I ask how do you deal with multiple faces in one frame? I tried to use tensor like this [Batch, Frames, Faces, C, H, W], and reshape it to [Batch x Frames x Faces, C, H, W] to backbone-CNN then  [Batch, Frames x Face, CNN-features] to Sequential model. But it seems not working and looks like just to add more noise to the model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 781673,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "03/21/2020 14:59:25",
          "content": "<p>We ignore it.... It doesn't affect that much, especially its 10 frames data, and will very likely include the second person as well. The model should be able to learn about 2 people situations.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "753363": "Taking request for:\n1. Better face extractor(similar as blazeface but better accuracy, a balance between accuracy and speed)\n2. vanilla LRCN\n3. Underbalancing technique that will re-choose corresponding fake video every epoch(This way it can cover more data)(similar to overbalancing but better).",
    "753395": "Approved. 😃",
    "753752": "All of the above mentioned seems interesting...",
    "753827": "you know I am not much interested in 3 😁 . But eagerly waiting for 1 and 2.",
    "753873": "1 is probably the most useful in practice, since the quality of the face recognition directly affects both training and final results.\n\nFrom an educative POV, 2 is certainly the most interesting. I don't even know what it is, if it's worth it, how it compares, etc.\n\n3 is IMHO not that necessary. I guess everyone has his/her own way of doing something similar.\n\nFor 1,  a simple link would do. So I vote for 2 ;)  + the link for 1 ;)",
    "753968": "Upvote this comment if you want notebook on 1st point. Choose 1 only. The first comment to get 12 upvotes wins the race and kernel will be published on that topic within few days.",
    "753969": "Upvote this comment if you want notebook on 2nd point. Choose 1 only. The first comment to get 12 upvotes wins the race and kernel will be published on that topic within few days.",
    "753970": "Upvote this comment if you want notebook on 3rd point. Choose 1 only. The first comment to get 12 upvotes wins the race and kernel will be published on that topic within few days.",
    "754026": "1 \nDid you try facenet-PyTorch?  I think it is better than blazeface. \nI found that submitting time using blazeface is an hour longer than using MTCNN.\nand according to [](https://www.kaggle.com/basharallabadi/yolov2-vs-faced-vs-blazeface-vs-mtcnn) , mtcnn has better accuracy.",
    "754061": "Mine have better accuracy than MTCNN, facenet-pytorch and also have a competitive speed(about 30 seconds to process 300 frames on CPU).",
    "754148": "Ahem https://github.com/sfzhang15/FaceBoxes",
    "754199": "Not this one @maralski, it took me days to find out after knowing the concept of how it works.",
    "754323": "&gt; I found that submitting time using blazeface is an hour longer than using MTCNN\n\nAre you batching multiple images into the same request or are you just doing one image at a time? I haven't used MTCNN but for me BlazeFace is not the bottleneck (reading the videos is). Since BlazeFace is small, the key to getting good speed is to use large batches. (It is true that other face detection models may be better than BlazeFace, but that is usually the trade-off for more speed.)",
    "754325": "On my pc, with a 1080 ti :- \nRead 10 frames equally gapped from a particular video, crops the faces, make them into one image, then, save the image.\nI can do this step ~2.3 it/s on my pc which is pretty good, you can consider the speed from here",
    "754699": "Hi, I jusr wonder how to judge the accuracy?",
    "754732": "Wow, I feel like I'm a small son who's looking for mommy's present :D",
    "754923": "Unfortunately, I saw both of them having 12 upvotes at the same time, I guess the good thing is both of the kernels are coming out at once.",
    "755553": "Well I have switched to RetinaFace w/ ResNet backbone. Results look very good. Just need to add it to my pipeline now.",
    "755577": "asides from clustering on perhaps features / embedding ... any ideas on how to separate identities in the train set (to create good validation set)?",
    "756425": "UPDATE: finished draft notebook of number one.",
    "756623": "last 10 folders have majority actors different than the rest of the folders.",
    "757404": "yzcwansui I didn't use a metric to judge accuracy BUT you can obviously see that blazeface have a bad accuracy(it often have a long bounding box); and mtcnn, mobilenet are pretty close. BTW blazeface often unable to detect face.",
    "757776": "now I'm comparing blazeface, mtcnn and that face extractor. Any request for other networks for comparison?",
    "757816": "Maybe you can add DSFD and YoloV3 to the comparison",
    "758067": "We know just because of how our models work that DSFD have the best accuracy to even any single model, I mean it is similar to our model, but a lot slower, I had tested it a month or two ago, when I was searching an optimal face detector. That is why I can not recall exact numbers for now. We might be implementing YoloV3, Thank You. @ngcferreira",
    "759358": "Final Draft and Published. Please let me know if there's any bug in the code. Thanks. \nComparisons: [link](https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-comparison)\nClean Helper Code: [link](https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-helper-code)",
    "759373": "the real thing is only the third part is the one which makes the difference which is actually considerable, I guess if these 2 notebooks are out and we give the third one too, we will loose all our kernels, as they all are made with slight changes, only 1 was meant to come out, but I guess now it is these 2.",
    "763720": "Well done, it looks good. Thanks for the kernel.",
    "763915": "May I ask just a question: the face detector works better on real face. For fake videos, it may not detect  blurred (fake) faces. So using a face detector to generate training samples has this issue: the extracted faces for fake videos don't correspond to real faces in the real video. So how did you do it? I think you fine-tune the face detector with correct fake faces (by translating bounding box from real video to fake video)?",
    "763949": "khahuras \nIts not even complicated at all, I am not currently doing anything to bypass this weakness intentionally, cuz I do want some noise and wrongs in the dataset so the model will always be learning, but to overcome the issue, take same frame from real video and same from fake video, and get the bounding boxes from real frame and apply the same on the fake frame, with a slight random distortion to make model think that they are actually made by face detector and are completely random, take the face from the fake frame, and you are done.",
    "764000": "harshitsheoran Thanks for your kindness! Honestly, given your advice, I think it's not a major issue that comes from face detector, even if we introduce some randomness/distortion as you suggested. I think the problem comes from my validation. Every time a trained classifier produces more aggressive predictions towards 0/1, the worse LB score, even though better CV. That means my CV has leakage or something like that. I should find a better validation scheme. Now my validation is folder split with last 10 folders as validation, and validation of those last 10 folders is around 0.25.",
    "764107": "The problem is not the CV, it is about that your model can not learn manage loss while predicting, note, going towards 1 or 0 only helps when the classifier have a very good accuracy change your model arch to get the model just right, train 1 epoch, in my case, I can just check with training one epoch, as if start is good, I know that the final result after training more epochs will also be good.\n\nEDIT, We do have same data for validation and trust me there was at time when I also had validation \naround 0.25, changing data and checking it on the bench (this bench is trustful), this bench does not tell how much the different is exactly, but it surely tells me that which model is better for example, my 0.318 model have cv on bench 0.16, and my 0.340 had cv on bench 0.18*.",
    "764134": "Thanks again so much :). I will try to surpass your team!",
    "764140": "Good Luck, cuz definately we are not even close to done improving yet, LRCN is taking our score to new heights.",
    "764143": "Given the amount of things you share in this competition, I totally am rooting for your team to get the possibly best result!",
    "774896": "I made publish [this](https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference) notebook. It is an inference of LRCN using the face extractor I shared before. \nI didn't make publish of the weights. I'm sure people(including me) doesn't like high LB notebook. But I'm OK with making it public if people want me to. I'm taking request.",
    "774900": "Thanks a lot! I believe, what you shared is good for people to try LRCN. Sharing weights will defeat the purpose as competition is coming closer to the end.",
    "774939": "Well, the weights will only get you to 0.38 so, we dont have a problem to make it publically available, the reason not to is the same as @debanga 's.",
    "775040": "Thanks for sharing the notebook!\n\nWith regard to the weights, I agree with Zenify. While sharing the weights would not affect the top solutions, it would certainly remove all those better than 0.43 from medal zone. Apart from that, I doubt that there is much to learn from the weights.",
    "775157": "If you have to learn, we have summaried our exact model which we used until 0.31, there.",
    "775366": "Some people are requesting the training code. I published [this](https://www.kaggle.com/unkownhihi/dfdc-lrcn-training/) notebook. Its really nothing special, just load data and train. \n\nI removed the dataset because I want to avoid people just clone, run, submit. \nI'm OK with providing the weight/dataset, because we already started doing a new thing.\n\nI'm OK with making public the training dataset, making public the weights, but still, I'm taking request.",
    "775475": "Please don't publish the weights. It will undermine the 3 month effort of a lot of people who worked hard to get to wherever they got in the Leaderboard. The whole idea of the Leaderboard is to give some reward to the people competing. Don't take that away from them. Imagine how you would you feel if someone above you did the same thing and your team ends up in 2000th place at the end of the competition?",
    "775476": "cesb45 I agree with you. I felt the same when there's a high scoring public notebook and our score is below it. We worked extra hard to get out of the \"bubble\".",
    "778941": "Maybe you can upload only partial of your dataset in a separate Kaggle Dataset folder, then we will get a better idea on how the dataset is been prepare while no one can simply just clone, run, submit without putting in any effort.",
    "778948": "Well, you know, I don't want people to just like copy our method. All I wanted is to give them an idea how its achieved. I can give you a demo image:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2F168e1beaac474a9757fb3d0d6b45b171%2Faaqaifqrwn.jpg?generation=1584568583604380&amp;alt=media)",
    "781319": "Really appreciate for your constantly sharing in this completion, it inspired me a lot. \n\nMay I ask how do you deal with multiple faces in one frame? I tried to use tensor like this [Batch, Frames, Faces, C, H, W], and reshape it to [Batch x Frames x Faces, C, H, W] to backbone-CNN then  [Batch, Frames x Face, CNN-features] to Sequential model. But it seems not working and looks like just to add more noise to the model.",
    "781673": "We ignore it.... It doesn't affect that much, especially its 10 frames data, and will very likely include the second person as well. The model should be able to learn about 2 people situations.",
    "784780": "So did you finally publish the dataset? @harshitsheoran",
    "784789": "which dataset, and if you are talking about something you think I am only doing cuz I do not have private space left, dont disclose it on kaggle. We are only making some of our datas public cuz first we do not have enough space left to get our models in and data in for training and all, so please dont access them cuz they are waste of time, not gonna make you score to our best sub."
  },
  "source": "meta"
}