{
  "id": 132101,
  "title": "Ensemble suggestions?",
  "url": "/competitions/deepfake-detection-challenge/discussion/132101",
  "author_name": "",
  "post_date": "2020-02-24T05:18:29.825964400Z",
  "votes": 10,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Currently, I am ensembling two models available in the public kernels: ResNeXt and Xception, with equal weights. It gives me a 0.01 LB boost. I know that my Xception model performs slightly better in LB. Given this, what strategy should I take to determine appropriate weights? Any insights will be highly appreciated! Thanks in advance.</p>",
  "messages": [
    {
      "id": "754826",
      "postDate": "02/24/2020 05:18:29",
      "content": "<p>Currently, I am ensembling two models available in the public kernels: ResNeXt and Xception, with equal weights. It gives me a 0.01 LB boost. I know that my Xception model performs slightly better in LB. Given this, what strategy should I take to determine appropriate weights? Any insights will be highly appreciated! Thanks in advance.</p>",
      "rawMarkdown": "Currently, I am ensembling two models available in the public kernels: ResNeXt and Xception, with equal weights. It gives me a 0.01 LB boost. I know that my Xception model performs slightly better in LB. Given this, what strategy should I take to determine appropriate weights? Any insights will be highly appreciated! Thanks in advance.",
      "votes": null
    },
    {
      "id": "754848",
      "postDate": "02/24/2020 05:48:48",
      "content": "<p>I haven't tried ensembling myself, but have you done an analysis of time required for an extra model.predict() to run vs the extra number of frames you could have processed instead with only 1 model? </p>",
      "rawMarkdown": "I haven't tried ensembling myself, but have you done an analysis of time required for an extra model.predict() to run vs the extra number of frames you could have processed instead with only 1 model?",
      "votes": null
    },
    {
      "id": "754922",
      "postDate": "02/24/2020 08:13:32",
      "content": "<p>In your situation you probably should weight the Xception model slightly higher, but how much higher, who knows? You know the performance now weighing 1:1, so just try 2:1.</p>\n\n<p>However, things are much less clear when you have &gt; 2 models, because then it comes down to how well those models' predictions are correlated.</p>\n\n<p>For example, if you had 3 models, A, B and C. You might find A and B performed better than model C, but A and B also tended to produce very similar predictions, whereas C produced very different predictions but was only slightly worse. This implies that models A (or B) and C are working in fundamentally different ways.</p>\n\n<p>In this situation, it would probably be unwise to weight A and B more than C, because B is providing very little incremental data beyond having model A, alone. Instead, it is very useful to know if model C agrees or disagrees with model A. If anything, more weighting should be given to the worst performing model, C (though less than A and B combined), so its predictions are better able to influence the other two models.</p>\n\n<p>In my limited experience, due to these complexities, it's better to just experiment. It's a bit tedious with only 2 submissions a day, but it's just too hard to predict, otherwise.</p>",
      "rawMarkdown": "In your situation you probably should weight the Xception model slightly higher, but how much higher, who knows? You know the performance now weighing 1:1, so just try 2:1.\n\nHowever, things are much less clear when you have &gt; 2 models, because then it comes down to how well those models' predictions are correlated.\n\nFor example, if you had 3 models, A, B and C. You might find A and B performed better than model C, but A and B also tended to produce very similar predictions, whereas C produced very different predictions but was only slightly worse. This implies that models A (or B) and C are working in fundamentally different ways.\n\nIn this situation, it would probably be unwise to weight A and B more than C, because B is providing very little incremental data beyond having model A, alone. Instead, it is very useful to know if model C agrees or disagrees with model A. If anything, more weighting should be given to the worst performing model, C (though less than A and B combined), so its predictions are better able to influence the other two models.\n\nIn my limited experience, due to these complexities, it's better to just experiment. It's a bit tedious with only 2 submissions a day, but it's just too hard to predict, otherwise.",
      "votes": null
    },
    {
      "id": "754955",
      "postDate": "02/24/2020 09:15:43",
      "content": "<p>In this kernel you can see a technique to find a good weiht for 3 models. One can use this technique independend from the programming language - just look at the charts and the Outcome if you are not familiar with R. For two models it is equivalent.</p>\n\n<p><a href=\"https://www.kaggle.com/frankmollard/bag3models\">mixing 3 models</a></p>",
      "rawMarkdown": "In this kernel you can see a technique to find a good weiht for 3 models. One can use this technique independend from the programming language - just look at the charts and the Outcome if you are not familiar with R. For two models it is equivalent.\n\n[mixing 3 models](https://www.kaggle.com/frankmollard/bag3models)",
      "votes": null
    },
    {
      "id": "755006",
      "postDate": "02/24/2020 10:31:42",
      "content": "<p>you can try temp sharpening method also ,take average of sq roots of predictions. \n<a href=\"/debanga\">@debanga</a> \nBetween could u point me which Xception model u tried ?</p>",
      "rawMarkdown": "you can try temp sharpening method also ,take average of sq roots of predictions. \n@debanga \nBetween could u point me which Xception model u tried ?",
      "votes": null
    },
    {
      "id": "755009",
      "postDate": "02/24/2020 10:37:55",
      "content": "<p>Can you tell me more about this method? I can't see why square rooting before averaging helps? Also I assume you're square-rooting the distance from 0.5 rather than the actual prediction?</p>",
      "rawMarkdown": "Can you tell me more about this method? I can't see why square rooting before averaging helps? Also I assume you're square-rooting the distance from 0.5 rather than the actual prediction?",
      "votes": null
    },
    {
      "id": "755022",
      "postDate": "02/24/2020 10:56:36",
      "content": "<p>This helps in smoothing of  far away predictions  taken an eg. your model  A predicts 0.1 and but B  predicts 0.9  if u do simple avg u will get 0.5  but using temp sharpening u will 0.6 to 0.65 .So it tends to be bit cautious in making shift towards the lowest side of prediction...\nsome times this approach helps but some times not.. better try it out once. \nSecondly ,If you can point me to xception model which works well .. i tried the public kernel xception model but it seems to be not performing well compared to resnext. not sure if i try the right one. \nAlso if u can provide some cool suggestion to improve the rank further,may be in a 1-1 message doctor Just an optional ask :)..</p>",
      "rawMarkdown": "This helps in smoothing of  far away predictions  taken an eg. your model  A predicts 0.1 and but B  predicts 0.9  if u do simple avg u will get 0.5  but using temp sharpening u will 0.6 to 0.65 .So it tends to be bit cautious in making shift towards the lowest side of prediction...\nsome times this approach helps but some times not.. better try it out once. \nSecondly ,If you can point me to xception model which works well .. i tried the public kernel xception model but it seems to be not performing well compared to resnext. not sure if i try the right one. \nAlso if u can provide some cool suggestion to improve the rank further,may be in a 1-1 message doctor Just an optional ask :)..",
      "votes": null
    },
    {
      "id": "755050",
      "postDate": "02/24/2020 11:37:42",
      "content": "<p>Hi I don’t personally use an Xception model but I believe there’s a public kernel which ensembles one with ResNeXT that is popular?</p>\n\n<p>Happy to help with specific questions but I think privately giving advice in the competition 1on1 is against competition rules...</p>",
      "rawMarkdown": "Hi I don’t personally use an Xception model but I believe there’s a public kernel which ensembles one with ResNeXT that is popular?\n\nHappy to help with specific questions but I think privately giving advice in the competition 1on1 is against competition rules...",
      "votes": null
    },
    {
      "id": "755059",
      "postDate": "02/24/2020 11:45:52",
      "content": "<p>No problem.. i ask it here only\nHow many frames u used per video for training and which face detector ?\nHave u used audio also for detection.. \nAny augs u used ? models used is  resnext or any other one.. </p>",
      "rawMarkdown": "No problem.. i ask it here only\nHow many frames u used per video for training and which face detector ?\nHave u used audio also for detection.. \nAny augs u used ? models used is  resnext or any other one..",
      "votes": null
    },
    {
      "id": "755063",
      "postDate": "02/24/2020 11:50:34",
      "content": "<p>I do some slightly atypical stuff with frames that I can’t really go into but my network that processes videos with a CNN -&gt; LSTM uses efficientnet-b3 as the CNN. I use between 30 and 120 frames per video for this model.\nFor training I use faces randomly pulled from all videos, 1 per face.\nAug I use very little - hflip, random crop and brightness only.</p>",
      "rawMarkdown": "I do some slightly atypical stuff with frames that I can’t really go into but my network that processes videos with a CNN -&gt; LSTM uses efficientnet-b3 as the CNN. I use between 30 and 120 frames per video for this model.\nFor training I use faces randomly pulled from all videos, 1 per face.\nAug I use very little - hflip, random crop and brightness only.",
      "votes": null
    },
    {
      "id": "755143",
      "postDate": "02/24/2020 13:41:49",
      "content": "<p>Just to clarify, you are using 30-120 frames for the purpose of prediction in the kaggle submissions?</p>",
      "rawMarkdown": "Just to clarify, you are using 30-120 frames for the purpose of prediction in the kaggle submissions?",
      "votes": null
    },
    {
      "id": "755146",
      "postDate": "02/24/2020 13:43:15",
      "content": "<p>Hi thank you! I am out of home now but will get back in the evening with details, public suggestions are good as I believe it’s  a general topic and will help everyone...!</p>",
      "rawMarkdown": "Hi thank you! I am out of home now but will get back in the evening with details, public suggestions are good as I believe it’s  a general topic and will help everyone...!",
      "votes": null
    },
    {
      "id": "755180",
      "postDate": "02/24/2020 14:23:30",
      "content": "<p>Thanks for sharing!😀 </p>",
      "rawMarkdown": "Thanks for sharing!😀",
      "votes": null
    },
    {
      "id": "755212",
      "postDate": "02/24/2020 14:56:15",
      "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> so your model architecture is CNN+LSTM? not just CNN?</p>",
      "rawMarkdown": "jamesphoward so your model architecture is CNN+LSTM? not just CNN?",
      "votes": null
    },
    {
      "id": "755306",
      "postDate": "02/24/2020 16:48:43",
      "content": "<p>Thanks! Really useful comment.</p>",
      "rawMarkdown": "Thanks! Really useful comment.",
      "votes": null
    },
    {
      "id": "755309",
      "postDate": "02/24/2020 16:50:46",
      "content": "<p>For me each additional model (CNN-based) adds around 100-200ms per video I think. Also, consider I am reading multiple frames per video as shown in <a href=\"/humananalog\">@humananalog</a> public kernel.</p>",
      "rawMarkdown": "For me each additional model (CNN-based) adds around 100-200ms per video I think. Also, consider I am reading multiple frames per video as shown in @humananalog public kernel.",
      "votes": null
    },
    {
      "id": "755310",
      "postDate": "02/24/2020 16:51:51",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "755314",
      "postDate": "02/24/2020 16:53:01",
      "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> This kernel: <a href=\"https://www.kaggle.com/greatgamedota/xception-classifier-w-ffhq-training-lb-537\">https://www.kaggle.com/greatgamedota/xception-classifier-w-ffhq-training-lb-537</a>. I think my data is helping me to get a better score.</p>",
      "rawMarkdown": "jaideepvalani This kernel: https://www.kaggle.com/greatgamedota/xception-classifier-w-ffhq-training-lb-537. I think my data is helping me to get a better score.",
      "votes": null
    },
    {
      "id": "755351",
      "postDate": "02/24/2020 17:35:57",
      "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a>\n\"For training I use faces randomly pulled from all videos, 1 per face.\" \nI did not get this line, Can you please explain. Thank You!, And this might sound weird to you, but just curious to know your batch size, That would help alot if I get a score close to you, Its gonna be a team-up request from my Side, Thank You, Again!</p>",
      "rawMarkdown": "jamesphoward\n\"For training I use faces randomly pulled from all videos, 1 per face.\" \nI did not get this line, Can you please explain. Thank You!, And this might sound weird to you, but just curious to know your batch size, That would help alot if I get a score close to you, Its gonna be a team-up request from my Side, Thank You, Again!",
      "votes": null
    },
    {
      "id": "755384",
      "postDate": "02/24/2020 18:17:27",
      "content": "<p><a href=\"/yangsaewon\">@yangsaewon</a> That's one of my models, yes.</p>\n\n<p><a href=\"/harshitsheoran\">@harshitsheoran</a> I really can't go into too much detail as I suspect my position is largely due to my pre-processing. But I process all videos, and then I find unique faces in those videos, and only use each face once. My batch size usually 32.</p>",
      "rawMarkdown": "yangsaewon That's one of my models, yes.\n\n@harshitsheoran I really can't go into too much detail as I suspect my position is largely due to my pre-processing. But I process all videos, and then I find unique faces in those videos, and only use each face once. My batch size usually 32.",
      "votes": null
    },
    {
      "id": "755399",
      "postDate": "02/24/2020 18:37:44",
      "content": "<p>can I know on what kinda gpu you are training, specifically how much ram in them? <a href=\"/jamesphoward\">@jamesphoward</a> </p>",
      "rawMarkdown": "can I know on what kinda gpu you are training, specifically how much ram in them? @jamesphoward",
      "votes": null
    },
    {
      "id": "755405",
      "postDate": "02/24/2020 18:42:16",
      "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> Nice! So you got better results with CNN + LSTM since your previous attempt:\n<a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/129663#742527\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/129663#742527</a></p>\n\n<p>What's new since then? I've tried 20 frames per video but behavior did not improve for me, it always underfits. I've trained a backbone alone, then I've frozen the backbone and trained the LSTM only.</p>",
      "rawMarkdown": "jamesphoward Nice! So you got better results with CNN + LSTM since your previous attempt:\nhttps://www.kaggle.com/c/deepfake-detection-challenge/discussion/129663#742527\n\nWhat's new since then? I've tried 20 frames per video but behavior did not improve for me, it always underfits. I've trained a backbone alone, then I've frozen the backbone and trained the LSTM only.",
      "votes": null
    },
    {
      "id": "755432",
      "postDate": "02/24/2020 19:17:19",
      "content": "<p><a href=\"/mpware\">@mpware</a> Yeah I'm not sure, it's not MUCH better than a CNN, just slightly. I tried to use more basic models, and yes I also did freeze them. I also tried various things like 1 cycle training, various levels of augmentation, 1 or 2 levels of LSTM, various amounts of dropout etc. etc.\nIn the end it seems the simpler the model (in every way), the better...</p>",
      "rawMarkdown": "mpware Yeah I'm not sure, it's not MUCH better than a CNN, just slightly. I tried to use more basic models, and yes I also did freeze them. I also tried various things like 1 cycle training, various levels of augmentation, 1 or 2 levels of LSTM, various amounts of dropout etc. etc.\nIn the end it seems the simpler the model (in every way), the better...",
      "votes": null
    },
    {
      "id": "755552",
      "postDate": "02/24/2020 21:59:41",
      "content": "<p>Did you try random erasing?</p>",
      "rawMarkdown": "Did you try random erasing?",
      "votes": null
    },
    {
      "id": "755554",
      "postDate": "02/24/2020 22:05:16",
      "content": "<p>I tried Albumentations' <code>CoarseDropout</code> and it did not help - in fact it makes things considerably worse.</p>",
      "rawMarkdown": "I tried Albumentations' `CoarseDropout` and it did not help - in fact it makes things considerably worse.",
      "votes": null
    },
    {
      "id": "758715",
      "postDate": "02/28/2020 04:34:39",
      "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> is it possible to let us know the input face size?</p>",
      "rawMarkdown": "jamesphoward is it possible to let us know the input face size?",
      "votes": null
    },
    {
      "id": "759873",
      "postDate": "02/29/2020 15:08:37",
      "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> I don't quiet understand your words \"find unique faces in those videos, and only use each face once\". so you are saying that you are using one face per video for training? what if you capture a face in the frame and that face happens to be fake-like face image in the real labeled video and vice versa?  wouldn't that cause a problem?\nreally thanks for the reply :)</p>",
      "rawMarkdown": "jamesphoward I don't quiet understand your words \"find unique faces in those videos, and only use each face once\". so you are saying that you are using one face per video for training? what if you capture a face in the frame and that face happens to be fake-like face image in the real labeled video and vice versa?  wouldn't that cause a problem?\nreally thanks for the reply :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 754848,
      "author_name": "akashnandi",
      "author_url": "",
      "post_date": "02/24/2020 05:48:48",
      "content": "<p>I haven't tried ensembling myself, but have you done an analysis of time required for an extra model.predict() to run vs the extra number of frames you could have processed instead with only 1 model? </p>",
      "votes": null,
      "replies": [
        {
          "id": 755309,
          "author_name": "debanga",
          "author_url": "",
          "post_date": "02/24/2020 16:50:46",
          "content": "<p>For me each additional model (CNN-based) adds around 100-200ms per video I think. Also, consider I am reading multiple frames per video as shown in <a href=\"/humananalog\">@humananalog</a> public kernel.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 754922,
      "author_name": "jamesphoward",
      "author_url": "",
      "post_date": "02/24/2020 08:13:32",
      "content": "<p>In your situation you probably should weight the Xception model slightly higher, but how much higher, who knows? You know the performance now weighing 1:1, so just try 2:1.</p>\n\n<p>However, things are much less clear when you have &gt; 2 models, because then it comes down to how well those models' predictions are correlated.</p>\n\n<p>For example, if you had 3 models, A, B and C. You might find A and B performed better than model C, but A and B also tended to produce very similar predictions, whereas C produced very different predictions but was only slightly worse. This implies that models A (or B) and C are working in fundamentally different ways.</p>\n\n<p>In this situation, it would probably be unwise to weight A and B more than C, because B is providing very little incremental data beyond having model A, alone. Instead, it is very useful to know if model C agrees or disagrees with model A. If anything, more weighting should be given to the worst performing model, C (though less than A and B combined), so its predictions are better able to influence the other two models.</p>\n\n<p>In my limited experience, due to these complexities, it's better to just experiment. It's a bit tedious with only 2 submissions a day, but it's just too hard to predict, otherwise.</p>",
      "votes": null,
      "replies": [
        {
          "id": 755306,
          "author_name": "debanga",
          "author_url": "",
          "post_date": "02/24/2020 16:48:43",
          "content": "<p>Thanks! Really useful comment.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 754955,
      "author_name": "frankmollard",
      "author_url": "",
      "post_date": "02/24/2020 09:15:43",
      "content": "<p>In this kernel you can see a technique to find a good weiht for 3 models. One can use this technique independend from the programming language - just look at the charts and the Outcome if you are not familiar with R. For two models it is equivalent.</p>\n\n<p><a href=\"https://www.kaggle.com/frankmollard/bag3models\">mixing 3 models</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 755310,
          "author_name": "debanga",
          "author_url": "",
          "post_date": "02/24/2020 16:51:51",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 755006,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "02/24/2020 10:31:42",
      "content": "<p>you can try temp sharpening method also ,take average of sq roots of predictions. \n<a href=\"/debanga\">@debanga</a> \nBetween could u point me which Xception model u tried ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 755009,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "02/24/2020 10:37:55",
          "content": "<p>Can you tell me more about this method? I can't see why square rooting before averaging helps? Also I assume you're square-rooting the distance from 0.5 rather than the actual prediction?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755022,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "02/24/2020 10:56:36",
          "content": "<p>This helps in smoothing of  far away predictions  taken an eg. your model  A predicts 0.1 and but B  predicts 0.9  if u do simple avg u will get 0.5  but using temp sharpening u will 0.6 to 0.65 .So it tends to be bit cautious in making shift towards the lowest side of prediction...\nsome times this approach helps but some times not.. better try it out once. \nSecondly ,If you can point me to xception model which works well .. i tried the public kernel xception model but it seems to be not performing well compared to resnext. not sure if i try the right one. \nAlso if u can provide some cool suggestion to improve the rank further,may be in a 1-1 message doctor Just an optional ask :)..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755050,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "02/24/2020 11:37:42",
          "content": "<p>Hi I don’t personally use an Xception model but I believe there’s a public kernel which ensembles one with ResNeXT that is popular?</p>\n\n<p>Happy to help with specific questions but I think privately giving advice in the competition 1on1 is against competition rules...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755059,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "02/24/2020 11:45:52",
          "content": "<p>No problem.. i ask it here only\nHow many frames u used per video for training and which face detector ?\nHave u used audio also for detection.. \nAny augs u used ? models used is  resnext or any other one.. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755063,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "02/24/2020 11:50:34",
          "content": "<p>I do some slightly atypical stuff with frames that I can’t really go into but my network that processes videos with a CNN -&gt; LSTM uses efficientnet-b3 as the CNN. I use between 30 and 120 frames per video for this model.\nFor training I use faces randomly pulled from all videos, 1 per face.\nAug I use very little - hflip, random crop and brightness only.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755143,
          "author_name": "akashnandi",
          "author_url": "",
          "post_date": "02/24/2020 13:41:49",
          "content": "<p>Just to clarify, you are using 30-120 frames for the purpose of prediction in the kaggle submissions?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755146,
          "author_name": "debanga",
          "author_url": "",
          "post_date": "02/24/2020 13:43:15",
          "content": "<p>Hi thank you! I am out of home now but will get back in the evening with details, public suggestions are good as I believe it’s  a general topic and will help everyone...!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755180,
          "author_name": "chmhzc",
          "author_url": "",
          "post_date": "02/24/2020 14:23:30",
          "content": "<p>Thanks for sharing!😀 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755212,
          "author_name": "yangsaewon",
          "author_url": "",
          "post_date": "02/24/2020 14:56:15",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> so your model architecture is CNN+LSTM? not just CNN?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755314,
          "author_name": "debanga",
          "author_url": "",
          "post_date": "02/24/2020 16:53:01",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> This kernel: <a href=\"https://www.kaggle.com/greatgamedota/xception-classifier-w-ffhq-training-lb-537\">https://www.kaggle.com/greatgamedota/xception-classifier-w-ffhq-training-lb-537</a>. I think my data is helping me to get a better score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755351,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/24/2020 17:35:57",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a>\n\"For training I use faces randomly pulled from all videos, 1 per face.\" \nI did not get this line, Can you please explain. Thank You!, And this might sound weird to you, but just curious to know your batch size, That would help alot if I get a score close to you, Its gonna be a team-up request from my Side, Thank You, Again!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755384,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "02/24/2020 18:17:27",
          "content": "<p><a href=\"/yangsaewon\">@yangsaewon</a> That's one of my models, yes.</p>\n\n<p><a href=\"/harshitsheoran\">@harshitsheoran</a> I really can't go into too much detail as I suspect my position is largely due to my pre-processing. But I process all videos, and then I find unique faces in those videos, and only use each face once. My batch size usually 32.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755399,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/24/2020 18:37:44",
          "content": "<p>can I know on what kinda gpu you are training, specifically how much ram in them? <a href=\"/jamesphoward\">@jamesphoward</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755405,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "02/24/2020 18:42:16",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> Nice! So you got better results with CNN + LSTM since your previous attempt:\n<a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/129663#742527\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/129663#742527</a></p>\n\n<p>What's new since then? I've tried 20 frames per video but behavior did not improve for me, it always underfits. I've trained a backbone alone, then I've frozen the backbone and trained the LSTM only.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755432,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "02/24/2020 19:17:19",
          "content": "<p><a href=\"/mpware\">@mpware</a> Yeah I'm not sure, it's not MUCH better than a CNN, just slightly. I tried to use more basic models, and yes I also did freeze them. I also tried various things like 1 cycle training, various levels of augmentation, 1 or 2 levels of LSTM, various amounts of dropout etc. etc.\nIn the end it seems the simpler the model (in every way), the better...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755552,
          "author_name": "maralski",
          "author_url": "",
          "post_date": "02/24/2020 21:59:41",
          "content": "<p>Did you try random erasing?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755554,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "02/24/2020 22:05:16",
          "content": "<p>I tried Albumentations' <code>CoarseDropout</code> and it did not help - in fact it makes things considerably worse.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 758715,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "02/28/2020 04:34:39",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> is it possible to let us know the input face size?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759873,
          "author_name": "yangsaewon",
          "author_url": "",
          "post_date": "02/29/2020 15:08:37",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> I don't quiet understand your words \"find unique faces in those videos, and only use each face once\". so you are saying that you are using one face per video for training? what if you capture a face in the frame and that face happens to be fake-like face image in the real labeled video and vice versa?  wouldn't that cause a problem?\nreally thanks for the reply :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "754826": "Currently, I am ensembling two models available in the public kernels: ResNeXt and Xception, with equal weights. It gives me a 0.01 LB boost. I know that my Xception model performs slightly better in LB. Given this, what strategy should I take to determine appropriate weights? Any insights will be highly appreciated! Thanks in advance.",
    "754848": "I haven't tried ensembling myself, but have you done an analysis of time required for an extra model.predict() to run vs the extra number of frames you could have processed instead with only 1 model?",
    "754922": "In your situation you probably should weight the Xception model slightly higher, but how much higher, who knows? You know the performance now weighing 1:1, so just try 2:1.\n\nHowever, things are much less clear when you have &gt; 2 models, because then it comes down to how well those models' predictions are correlated.\n\nFor example, if you had 3 models, A, B and C. You might find A and B performed better than model C, but A and B also tended to produce very similar predictions, whereas C produced very different predictions but was only slightly worse. This implies that models A (or B) and C are working in fundamentally different ways.\n\nIn this situation, it would probably be unwise to weight A and B more than C, because B is providing very little incremental data beyond having model A, alone. Instead, it is very useful to know if model C agrees or disagrees with model A. If anything, more weighting should be given to the worst performing model, C (though less than A and B combined), so its predictions are better able to influence the other two models.\n\nIn my limited experience, due to these complexities, it's better to just experiment. It's a bit tedious with only 2 submissions a day, but it's just too hard to predict, otherwise.",
    "754955": "In this kernel you can see a technique to find a good weiht for 3 models. One can use this technique independend from the programming language - just look at the charts and the Outcome if you are not familiar with R. For two models it is equivalent.\n\n[mixing 3 models](https://www.kaggle.com/frankmollard/bag3models)",
    "755006": "you can try temp sharpening method also ,take average of sq roots of predictions. \n@debanga \nBetween could u point me which Xception model u tried ?",
    "755009": "Can you tell me more about this method? I can't see why square rooting before averaging helps? Also I assume you're square-rooting the distance from 0.5 rather than the actual prediction?",
    "755022": "This helps in smoothing of  far away predictions  taken an eg. your model  A predicts 0.1 and but B  predicts 0.9  if u do simple avg u will get 0.5  but using temp sharpening u will 0.6 to 0.65 .So it tends to be bit cautious in making shift towards the lowest side of prediction...\nsome times this approach helps but some times not.. better try it out once. \nSecondly ,If you can point me to xception model which works well .. i tried the public kernel xception model but it seems to be not performing well compared to resnext. not sure if i try the right one. \nAlso if u can provide some cool suggestion to improve the rank further,may be in a 1-1 message doctor Just an optional ask :)..",
    "755050": "Hi I don’t personally use an Xception model but I believe there’s a public kernel which ensembles one with ResNeXT that is popular?\n\nHappy to help with specific questions but I think privately giving advice in the competition 1on1 is against competition rules...",
    "755059": "No problem.. i ask it here only\nHow many frames u used per video for training and which face detector ?\nHave u used audio also for detection.. \nAny augs u used ? models used is  resnext or any other one..",
    "755063": "I do some slightly atypical stuff with frames that I can’t really go into but my network that processes videos with a CNN -&gt; LSTM uses efficientnet-b3 as the CNN. I use between 30 and 120 frames per video for this model.\nFor training I use faces randomly pulled from all videos, 1 per face.\nAug I use very little - hflip, random crop and brightness only.",
    "755143": "Just to clarify, you are using 30-120 frames for the purpose of prediction in the kaggle submissions?",
    "755146": "Hi thank you! I am out of home now but will get back in the evening with details, public suggestions are good as I believe it’s  a general topic and will help everyone...!",
    "755180": "Thanks for sharing!😀",
    "755212": "jamesphoward so your model architecture is CNN+LSTM? not just CNN?",
    "755306": "Thanks! Really useful comment.",
    "755309": "For me each additional model (CNN-based) adds around 100-200ms per video I think. Also, consider I am reading multiple frames per video as shown in @humananalog public kernel.",
    "755310": "Thanks!",
    "755314": "jaideepvalani This kernel: https://www.kaggle.com/greatgamedota/xception-classifier-w-ffhq-training-lb-537. I think my data is helping me to get a better score.",
    "755351": "jamesphoward\n\"For training I use faces randomly pulled from all videos, 1 per face.\" \nI did not get this line, Can you please explain. Thank You!, And this might sound weird to you, but just curious to know your batch size, That would help alot if I get a score close to you, Its gonna be a team-up request from my Side, Thank You, Again!",
    "755384": "yangsaewon That's one of my models, yes.\n\n@harshitsheoran I really can't go into too much detail as I suspect my position is largely due to my pre-processing. But I process all videos, and then I find unique faces in those videos, and only use each face once. My batch size usually 32.",
    "755399": "can I know on what kinda gpu you are training, specifically how much ram in them? @jamesphoward",
    "755405": "jamesphoward Nice! So you got better results with CNN + LSTM since your previous attempt:\nhttps://www.kaggle.com/c/deepfake-detection-challenge/discussion/129663#742527\n\nWhat's new since then? I've tried 20 frames per video but behavior did not improve for me, it always underfits. I've trained a backbone alone, then I've frozen the backbone and trained the LSTM only.",
    "755432": "mpware Yeah I'm not sure, it's not MUCH better than a CNN, just slightly. I tried to use more basic models, and yes I also did freeze them. I also tried various things like 1 cycle training, various levels of augmentation, 1 or 2 levels of LSTM, various amounts of dropout etc. etc.\nIn the end it seems the simpler the model (in every way), the better...",
    "755552": "Did you try random erasing?",
    "755554": "I tried Albumentations' `CoarseDropout` and it did not help - in fact it makes things considerably worse.",
    "758715": "jamesphoward is it possible to let us know the input face size?",
    "759873": "jamesphoward I don't quiet understand your words \"find unique faces in those videos, and only use each face once\". so you are saying that you are using one face per video for training? what if you capture a face in the frame and that face happens to be fake-like face image in the real labeled video and vice versa?  wouldn't that cause a problem?\nreally thanks for the reply :)"
  },
  "source": "meta"
}