{
  "id": 131060,
  "title": "Simple Classifier Baseline (LB .489)",
  "url": "/competitions/deepfake-detection-challenge/discussion/131060",
  "author_name": "",
  "post_date": "2020-02-18T02:36:44.463087600Z",
  "votes": 28,
  "comment_count": 30,
  "views": 0,
  "content": "<p>Notebook showing the training and inference of a simple Xception binary classifier:\n<a href=\"https://www.kaggle.com/greatgamedota/xception-classifier-w-ffhq-training-lb-555\">https://www.kaggle.com/greatgamedota/xception-classifier-w-ffhq-training-lb-555</a>\nThe notebook shows training for LB of .555 however with minimal changes you can train a model that scores .489 and possibly less!</p>",
  "messages": [
    {
      "id": "748809",
      "postDate": "02/18/2020 02:36:44",
      "content": "<p>Notebook showing the training and inference of a simple Xception binary classifier:\n<a href=\"https://www.kaggle.com/greatgamedota/xception-classifier-w-ffhq-training-lb-555\">https://www.kaggle.com/greatgamedota/xception-classifier-w-ffhq-training-lb-555</a>\nThe notebook shows training for LB of .555 however with minimal changes you can train a model that scores .489 and possibly less!</p>",
      "rawMarkdown": "Notebook showing the training and inference of a simple Xception binary classifier:\nhttps://www.kaggle.com/greatgamedota/xception-classifier-w-ffhq-training-lb-555\nThe notebook shows training for LB of .555 however with minimal changes you can train a model that scores .489 and possibly less!",
      "votes": null
    },
    {
      "id": "748831",
      "postDate": "02/18/2020 03:11:08",
      "content": "<p>Thanks! It's great.</p>",
      "rawMarkdown": "Thanks! It's great.",
      "votes": null
    },
    {
      "id": "748845",
      "postDate": "02/18/2020 03:33:43",
      "content": "<p>Nice! Great to know someone used AND ACHIEVED A PRETTY GOOD SCORE with my dataset I've abandoned for quite long(I discovered a better face detector(I will share it soon 😉) to achieve a new score! BTW, a hint: you might want to replace the backbone(xception) with another network😉 . My teammates will be mad at me if I give more details.</p>\n\n<p>EDIT: *My teammate discovered(Happy now? <a href=\"/harshitsheoran\">@harshitsheoran</a> )</p>",
      "rawMarkdown": "Nice! Great to know someone used AND ACHIEVED A PRETTY GOOD SCORE with my dataset I've abandoned for quite long(I discovered a better face detector(I will share it soon 😉) to achieve a new score! BTW, a hint: you might want to replace the backbone(xception) with another network😉 . My teammates will be mad at me if I give more details.\n\nEDIT: *My teammate discovered(Happy now? @harshitsheoran )",
      "votes": null
    },
    {
      "id": "748846",
      "postDate": "02/18/2020 03:36:44",
      "content": "<p>Thanks very much for the dataset and all the hints you've dropped 🙂 Xception was my first model I've experimented with and will definitely try others!</p>",
      "rawMarkdown": "Thanks very much for the dataset and all the hints you've dropped 🙂 Xception was my first model I've experimented with and will definitely try others!",
      "votes": null
    },
    {
      "id": "748977",
      "postDate": "02/18/2020 06:06:06",
      "content": "<p><a href=\"/unkownhihi\">@unkownhihi</a> Thanks again for your dataset, my first one and a half month in this competition (without any GCP, AWS credits, and too poor to pay for GPU hours) your dataset was the only source of experimentation!</p>",
      "rawMarkdown": "unkownhihi Thanks again for your dataset, my first one and a half month in this competition (without any GCP, AWS credits, and too poor to pay for GPU hours) your dataset was the only source of experimentation!",
      "votes": null
    },
    {
      "id": "749033",
      "postDate": "02/18/2020 07:57:07",
      "content": "<p>I would be real mad if you even talk about that face detector xD. You can disclose it later in the competition, not in this week.</p>\n\n<p>EDIT : Those downvotes, I guess you guys really needs it, okay I guess I will allow that btw where he says that he discovered a better face detector, He means that it was me who gave it to him in the first place when we teamed up, it makes no sense to downvote, if you are really good and need to know that or need to know how our baseline is scoring more than most people, what tricks are we applying, cuz I guess we have got a lot of those to make our baseline score that, You can join with us, I am willing to take people who have a leaderboard score of less than 0.4.</p>",
      "rawMarkdown": "I would be real mad if you even talk about that face detector xD. You can disclose it later in the competition, not in this week.\n\nEDIT : Those downvotes, I guess you guys really needs it, okay I guess I will allow that btw where he says that he discovered a better face detector, He means that it was me who gave it to him in the first place when we teamed up, it makes no sense to downvote, if you are really good and need to know that or need to know how our baseline is scoring more than most people, what tricks are we applying, cuz I guess we have got a lot of those to make our baseline score that, You can join with us, I am willing to take people who have a leaderboard score of less than 0.4.",
      "votes": null
    },
    {
      "id": "749075",
      "postDate": "02/18/2020 09:16:18",
      "content": "<p>You guys are really creating suspense. May I ask this that have you changed your dataset or used the same for the new submission. <a href=\"/harshitsheoran\">@harshitsheoran</a> <a href=\"/unkownhihi\">@unkownhihi</a> </p>",
      "rawMarkdown": "You guys are really creating suspense. May I ask this that have you changed your dataset or used the same for the new submission. @harshitsheoran @unkownhihi",
      "votes": null
    },
    {
      "id": "749081",
      "postDate": "02/18/2020 09:28:11",
      "content": "<p>If I am not mistaken, We are using a new dataset for all of our new submissions, the current score is an ensemble of models scoring very similar. ~0.38 ensembling gave ~0.34.</p>",
      "rawMarkdown": "If I am not mistaken, We are using a new dataset for all of our new submissions, the current score is an ensemble of models scoring very similar. ~0.38 ensembling gave ~0.34.",
      "votes": null
    },
    {
      "id": "749263",
      "postDate": "02/18/2020 14:00:24",
      "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a> Thanks, man! I just tried your training kernel with my personal data + some augmentation and scored my best so far with a single Xception model LB 0.37222 😁</p>\n\n<p>PS: My logs for the above model:</p>\n\n<ul>\n<li>Training loss: 0.1899</li>\n<li>Validation loss: 0.3593</li>\n<li>Validation Acc: 0.835837</li>\n</ul>\n\n<p>I &lt;3 my CV (folder-wise validation)!</p>",
      "rawMarkdown": "greatgamedota Thanks, man! I just tried your training kernel with my personal data + some augmentation and scored my best so far with a single Xception model LB 0.37222 😁\n\nPS: My logs for the above model:\n\n- Training loss: 0.1899\n- Validation loss: 0.3593\n- Validation Acc: 0.835837\n\nI &lt;3 my CV (folder-wise validation)!",
      "votes": null
    },
    {
      "id": "749277",
      "postDate": "02/18/2020 14:16:09",
      "content": "<p>Nice result! I haven't even attempted a CV method yet.</p>",
      "rawMarkdown": "Nice result! I haven't even attempted a CV method yet.",
      "votes": null
    },
    {
      "id": "749330",
      "postDate": "02/18/2020 15:37:11",
      "content": "<p>Weird thing, I am using same folders for the validation I get a very good validation loss but it does not track the leaderboard well, I can except leaderboard to be 0.1-0.2 above the validation loss in my case.</p>",
      "rawMarkdown": "Weird thing, I am using same folders for the validation I get a very good validation loss but it does not track the leaderboard well, I can except leaderboard to be 0.1-0.2 above the validation loss in my case.",
      "votes": null
    },
    {
      "id": "749345",
      "postDate": "02/18/2020 16:01:18",
      "content": "<p>Hmm... <a href=\"/harshitsheoran\">@harshitsheoran</a> also to add, this time I tried different folderwise folds, and used folders 0-9 for the final model. You can check may be? If you try let me know if that makes any change, I’m also curious! For this competition CV is quite a mystery!</p>",
      "rawMarkdown": "Hmm... @harshitsheoran also to add, this time I tried different folderwise folds, and used folders 0-9 for the final model. You can check may be? If you try let me know if that makes any change, I’m also curious! For this competition CV is quite a mystery!",
      "votes": null
    },
    {
      "id": "749464",
      "postDate": "02/18/2020 18:10:49",
      "content": "<p>My single frame models always score around 0.44. Looks like I missed the batch normalization and dropout part!</p>",
      "rawMarkdown": "My single frame models always score around 0.44. Looks like I missed the batch normalization and dropout part!",
      "votes": null
    },
    {
      "id": "749650",
      "postDate": "02/18/2020 20:43:52",
      "content": "<p>Great!</p>",
      "rawMarkdown": "Great!",
      "votes": null
    },
    {
      "id": "750459",
      "postDate": "02/19/2020 11:58:30",
      "content": "<p><a href=\"/debanga\">@debanga</a>  i get very good Val set results but get very less score. \nWhy should not get good score</p>\n\n<p><code>0.058996   0.174389    0.950241</code>\nMy validation strategy was based on split of original video ids  and image size 128/128.\nPlease suggest me some thing good narrow down the gap between val and test results..</p>",
      "rawMarkdown": "debanga  i get very good Val set results but get very less score. \nWhy should not get good score\n\n`0.058996\t0.174389\t0.950241`\nMy validation strategy was based on split of original video ids  and image size 128/128.\nPlease suggest me some thing good narrow down the gap between val and test results..",
      "votes": null
    },
    {
      "id": "750673",
      "postDate": "02/19/2020 15:47:19",
      "content": "<p>That's a really good LB score for that validation score!</p>",
      "rawMarkdown": "That's a really good LB score for that validation score!",
      "votes": null
    },
    {
      "id": "750712",
      "postDate": "02/19/2020 16:24:58",
      "content": "<p>but i dont know why i dont get eqw score .. :(\nlooking for some help </p>",
      "rawMarkdown": "but i dont know why i dont get eqw score .. :(\nlooking for some help",
      "votes": null
    },
    {
      "id": "750820",
      "postDate": "02/19/2020 18:01:18",
      "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> Your validation split is not good, I guess you should go for the last 10 folders as those are the folders with the faces which mostly are not included in folders 0-40.</p>",
      "rawMarkdown": "jaideepvalani Your validation split is not good, I guess you should go for the last 10 folders as those are the folders with the faces which mostly are not included in folders 0-40.",
      "votes": null
    },
    {
      "id": "752652",
      "postDate": "02/21/2020 09:04:27",
      "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a> <a href=\"/debanga\">@debanga</a> Usin <a href=\"https://www.kaggle.com/greatgamedota/ffhq-face-data-set\">FFHQ-FACE</a> might be violating the competition rules. Look <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203#744289\">here</a> Using external data which is for non-commercial use only, seems to be in violation of the rules. </p>",
      "rawMarkdown": "greatgamedota @debanga Usin [FFHQ-FACE](https://www.kaggle.com/greatgamedota/ffhq-face-data-set) might be violating the competition rules. Look [here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203#744289) Using external data which is for non-commercial use only, seems to be in violation of the rules.",
      "votes": null
    },
    {
      "id": "752854",
      "postDate": "02/21/2020 13:28:28",
      "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a> thanks for sharing. I tried to download the ffhq dataset, but the link does not seem to work.\n<a href=\"https://www.kaggle.com/greatgamedota/ffhq-face-data-set/download\">https://www.kaggle.com/greatgamedota/ffhq-face-data-set/download</a>, getting 404 error.</p>",
      "rawMarkdown": "greatgamedota thanks for sharing. I tried to download the ffhq dataset, but the link does not seem to work.\nhttps://www.kaggle.com/greatgamedota/ffhq-face-data-set/download, getting 404 error.",
      "votes": null
    },
    {
      "id": "752922",
      "postDate": "02/21/2020 14:45:08",
      "content": "<p>I don't know why that doesn't work, use the kaggle api I guess</p>",
      "rawMarkdown": "I don't know why that doesn't work, use the kaggle api I guess",
      "votes": null
    },
    {
      "id": "752987",
      "postDate": "02/21/2020 15:45:07",
      "content": "<p>Thanks for the heads up, I looked into it and think you're right.</p>",
      "rawMarkdown": "Thanks for the heads up, I looked into it and think you're right.",
      "votes": null
    },
    {
      "id": "758720",
      "postDate": "02/28/2020 04:38:16",
      "content": "<p><a href=\"/debanga\">@debanga</a> How do you calculate accuracy? I use the 0-39 folder as the training set and 40-48 as the validation set. I took 600 videos (300 real, 300 fake) from each folder for validation, and got the following results:</p>\n\n<p>Format: folder, BCE, abs (y_true-y_pred)&gt; 0.3, abs (y_true-y_pred)&gt; 0.5, abs (y_true-y_pred)&gt; 0.7, abs (y_true-y_pred)&gt; 0.9</p>\n\n<p>part_40 0.4773 194 128 82 24</p>\n\n<p>part_41 0.3328 191 95 39 3</p>\n\n<p>part_42 0.2235 138 42 5 0</p>\n\n<p>part_43 0.3498 156 84 33 10</p>\n\n<p>part_44 0.2157 124 50 14 0</p>\n\n<p>part_45 0.4283 158 89 58 28</p>\n\n<p>part_46 0.3213 184 90 40 0</p>\n\n<p>part_47 0.2247 132 45 10 3</p>\n\n<p>part_48 0.3299 178 92 28 7</p>\n\n<p>But my loss on the public LB is only 0.56. I don’t know if this is reasonable. I feel there is a big gap between the public LB and my local CV. Do you have any suggestions? Thanks very much!</p>",
      "rawMarkdown": "debanga How do you calculate accuracy? I use the 0-39 folder as the training set and 40-48 as the validation set. I took 600 videos (300 real, 300 fake) from each folder for validation, and got the following results:\n\nFormat: folder, BCE, abs (y_true-y_pred)&gt; 0.3, abs (y_true-y_pred)&gt; 0.5, abs (y_true-y_pred)&gt; 0.7, abs (y_true-y_pred)&gt; 0.9\n\npart_40 0.4773 194 128 82 24\n\npart_41 0.3328 191 95 39 3\n\npart_42 0.2235 138 42 5 0\n\npart_43 0.3498 156 84 33 10\n\npart_44 0.2157 124 50 14 0\n\npart_45 0.4283 158 89 58 28\n\npart_46 0.3213 184 90 40 0\n\npart_47 0.2247 132 45 10 3\n\npart_48 0.3299 178 92 28 7\n\n\nBut my loss on the public LB is only 0.56. I don’t know if this is reasonable. I feel there is a big gap between the public LB and my local CV. Do you have any suggestions? Thanks very much!",
      "votes": null
    },
    {
      "id": "758778",
      "postDate": "02/28/2020 06:25:05",
      "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a>:</p>\n\n<p>See my suggestions below. Also, try higher resolution, which may improve your results.</p>\n\n<p><a href=\"/beeaware\">@beeaware</a>:\nIn my case, accuracy is a by-product. I only focus on validation BCE.</p>\n\n<p>In my current validation folds, if I get validation BCE = x during training, for a single model with <a href=\"/humananalog\">@humananalog</a> inference kernel with some tweaks, I get LB = x + (0.01 to 0.03).</p>\n\n<p>To be honest, I cannot pinpoint why it works, as I spent 1-2 weeks creating my dataset incorporating a lot of small bits of hints from the discussions. After that, when I started using the dataset, it just worked!</p>\n\n<p>Try hints from different discussions to produce a better dataset with a lot of images (e.g. 500k~1m). Try different networks- some networks work well for us and some not. Try ensembling to see how it affects your score.</p>\n\n<p>After teaming up, my current score has improved a lot due to the expert tips and access to more GPUs, thanks to my amazing teammate! </p>\n\n<p>Hope it helps!</p>",
      "rawMarkdown": "jaideepvalani:\n\nSee my suggestions below. Also, try higher resolution, which may improve your results.\n\n@beeaware:\nIn my case, accuracy is a by-product. I only focus on validation BCE.\n\nIn my current validation folds, if I get validation BCE = x during training, for a single model with @humananalog inference kernel with some tweaks, I get LB = x + (0.01 to 0.03).\n\nTo be honest, I cannot pinpoint why it works, as I spent 1-2 weeks creating my dataset incorporating a lot of small bits of hints from the discussions. After that, when I started using the dataset, it just worked!\n\nTry hints from different discussions to produce a better dataset with a lot of images (e.g. 500k~1m). Try different networks- some networks work well for us and some not. Try ensembling to see how it affects your score.\n\nAfter teaming up, my current score has improved a lot due to the expert tips and access to more GPUs, thanks to my amazing teammate! \n\nHope it helps!",
      "votes": null
    },
    {
      "id": "758817",
      "postDate": "02/28/2020 07:24:52",
      "content": "<p><a href=\"/debanga\">@debanga</a> Thanks very much, that helps a lot. My training dataset is about 800k images(10 images/video,80k videos from 0-40 folders), but I have not used data augmentation, I still doubt that data augmentation can significantly reduce my loss, Your cv and LB can match so well, my cv looks not bad but i don't know why LB is so different, I will try data augmentation to see what happens. </p>\n\n<p>Thanks again!</p>",
      "rawMarkdown": "debanga Thanks very much, that helps a lot. My training dataset is about 800k images(10 images/video,80k videos from 0-40 folders), but I have not used data augmentation, I still doubt that data augmentation can significantly reduce my loss, Your cv and LB can match so well, my cv looks not bad but i don't know why LB is so different, I will try data augmentation to see what happens. \n\nThanks again!",
      "votes": null
    },
    {
      "id": "759064",
      "postDate": "02/28/2020 14:25:30",
      "content": "<p>I am using a few, e.g hflip, downscale; I think they help a bit to prevent overfitting.</p>",
      "rawMarkdown": "I am using a few, e.g hflip, downscale; I think they help a bit to prevent overfitting.",
      "votes": null
    },
    {
      "id": "759349",
      "postDate": "02/28/2020 23:20:31",
      "content": "<p>Thanks, I will try it, by the way, would you share your image size？My current is\n299*299.</p>",
      "rawMarkdown": "Thanks, I will try it, by the way, would you share your image size？My current is\n299*299.",
      "votes": null
    },
    {
      "id": "759356",
      "postDate": "02/28/2020 23:29:25",
      "content": "<p>224 :)</p>",
      "rawMarkdown": "224 :)",
      "votes": null
    },
    {
      "id": "759820",
      "postDate": "02/29/2020 13:54:05",
      "content": "<p><a href=\"/debanga\">@debanga</a>  thanks \nHow many frames from each video did u take \nIf more than one ,how u consolidated loss for multiple frames </p>",
      "rawMarkdown": "debanga  thanks \nHow many frames from each video did u take \nIf more than one ,how u consolidated loss for multiple frames",
      "votes": null
    },
    {
      "id": "761418",
      "postDate": "03/02/2020 13:50:11",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "767945",
      "postDate": "03/10/2020 09:36:04",
      "content": "<p>Was the original input size for xception model 224x224? Why it works with 150x150 then?\nAre CNNs not sensitive to different input? Or are they are, but with sufficient training are able to relearn to a different image size?</p>",
      "rawMarkdown": "Was the original input size for xception model 224x224? Why it works with 150x150 then?\nAre CNNs not sensitive to different input? Or are they are, but with sufficient training are able to relearn to a different image size?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 748831,
      "author_name": "debanga",
      "author_url": "",
      "post_date": "02/18/2020 03:11:08",
      "content": "<p>Thanks! It's great.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 748845,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "02/18/2020 03:33:43",
      "content": "<p>Nice! Great to know someone used AND ACHIEVED A PRETTY GOOD SCORE with my dataset I've abandoned for quite long(I discovered a better face detector(I will share it soon 😉) to achieve a new score! BTW, a hint: you might want to replace the backbone(xception) with another network😉 . My teammates will be mad at me if I give more details.</p>\n\n<p>EDIT: *My teammate discovered(Happy now? <a href=\"/harshitsheoran\">@harshitsheoran</a> )</p>",
      "votes": null,
      "replies": [
        {
          "id": 748846,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "02/18/2020 03:36:44",
          "content": "<p>Thanks very much for the dataset and all the hints you've dropped 🙂 Xception was my first model I've experimented with and will definitely try others!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 748977,
          "author_name": "debanga",
          "author_url": "",
          "post_date": "02/18/2020 06:06:06",
          "content": "<p><a href=\"/unkownhihi\">@unkownhihi</a> Thanks again for your dataset, my first one and a half month in this competition (without any GCP, AWS credits, and too poor to pay for GPU hours) your dataset was the only source of experimentation!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749033,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/18/2020 07:57:07",
          "content": "<p>I would be real mad if you even talk about that face detector xD. You can disclose it later in the competition, not in this week.</p>\n\n<p>EDIT : Those downvotes, I guess you guys really needs it, okay I guess I will allow that btw where he says that he discovered a better face detector, He means that it was me who gave it to him in the first place when we teamed up, it makes no sense to downvote, if you are really good and need to know that or need to know how our baseline is scoring more than most people, what tricks are we applying, cuz I guess we have got a lot of those to make our baseline score that, You can join with us, I am willing to take people who have a leaderboard score of less than 0.4.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749075,
          "author_name": "navjotsingh47",
          "author_url": "",
          "post_date": "02/18/2020 09:16:18",
          "content": "<p>You guys are really creating suspense. May I ask this that have you changed your dataset or used the same for the new submission. <a href=\"/harshitsheoran\">@harshitsheoran</a> <a href=\"/unkownhihi\">@unkownhihi</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749081,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/18/2020 09:28:11",
          "content": "<p>If I am not mistaken, We are using a new dataset for all of our new submissions, the current score is an ensemble of models scoring very similar. ~0.38 ensembling gave ~0.34.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 749263,
      "author_name": "debanga",
      "author_url": "",
      "post_date": "02/18/2020 14:00:24",
      "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a> Thanks, man! I just tried your training kernel with my personal data + some augmentation and scored my best so far with a single Xception model LB 0.37222 😁</p>\n\n<p>PS: My logs for the above model:</p>\n\n<ul>\n<li>Training loss: 0.1899</li>\n<li>Validation loss: 0.3593</li>\n<li>Validation Acc: 0.835837</li>\n</ul>\n\n<p>I &lt;3 my CV (folder-wise validation)!</p>",
      "votes": null,
      "replies": [
        {
          "id": 749277,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "02/18/2020 14:16:09",
          "content": "<p>Nice result! I haven't even attempted a CV method yet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749330,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/18/2020 15:37:11",
          "content": "<p>Weird thing, I am using same folders for the validation I get a very good validation loss but it does not track the leaderboard well, I can except leaderboard to be 0.1-0.2 above the validation loss in my case.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749345,
          "author_name": "debanga",
          "author_url": "",
          "post_date": "02/18/2020 16:01:18",
          "content": "<p>Hmm... <a href=\"/harshitsheoran\">@harshitsheoran</a> also to add, this time I tried different folderwise folds, and used folders 0-9 for the final model. You can check may be? If you try let me know if that makes any change, I’m also curious! For this competition CV is quite a mystery!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749464,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "02/18/2020 18:10:49",
          "content": "<p>My single frame models always score around 0.44. Looks like I missed the batch normalization and dropout part!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 750459,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "02/19/2020 11:58:30",
          "content": "<p><a href=\"/debanga\">@debanga</a>  i get very good Val set results but get very less score. \nWhy should not get good score</p>\n\n<p><code>0.058996   0.174389    0.950241</code>\nMy validation strategy was based on split of original video ids  and image size 128/128.\nPlease suggest me some thing good narrow down the gap between val and test results..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 750673,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "02/19/2020 15:47:19",
          "content": "<p>That's a really good LB score for that validation score!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 750712,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "02/19/2020 16:24:58",
          "content": "<p>but i dont know why i dont get eqw score .. :(\nlooking for some help </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 750820,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/19/2020 18:01:18",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> Your validation split is not good, I guess you should go for the last 10 folders as those are the folders with the faces which mostly are not included in folders 0-40.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 758720,
          "author_name": "beeaware",
          "author_url": "",
          "post_date": "02/28/2020 04:38:16",
          "content": "<p><a href=\"/debanga\">@debanga</a> How do you calculate accuracy? I use the 0-39 folder as the training set and 40-48 as the validation set. I took 600 videos (300 real, 300 fake) from each folder for validation, and got the following results:</p>\n\n<p>Format: folder, BCE, abs (y_true-y_pred)&gt; 0.3, abs (y_true-y_pred)&gt; 0.5, abs (y_true-y_pred)&gt; 0.7, abs (y_true-y_pred)&gt; 0.9</p>\n\n<p>part_40 0.4773 194 128 82 24</p>\n\n<p>part_41 0.3328 191 95 39 3</p>\n\n<p>part_42 0.2235 138 42 5 0</p>\n\n<p>part_43 0.3498 156 84 33 10</p>\n\n<p>part_44 0.2157 124 50 14 0</p>\n\n<p>part_45 0.4283 158 89 58 28</p>\n\n<p>part_46 0.3213 184 90 40 0</p>\n\n<p>part_47 0.2247 132 45 10 3</p>\n\n<p>part_48 0.3299 178 92 28 7</p>\n\n<p>But my loss on the public LB is only 0.56. I don’t know if this is reasonable. I feel there is a big gap between the public LB and my local CV. Do you have any suggestions? Thanks very much!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 758778,
          "author_name": "debanga",
          "author_url": "",
          "post_date": "02/28/2020 06:25:05",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a>:</p>\n\n<p>See my suggestions below. Also, try higher resolution, which may improve your results.</p>\n\n<p><a href=\"/beeaware\">@beeaware</a>:\nIn my case, accuracy is a by-product. I only focus on validation BCE.</p>\n\n<p>In my current validation folds, if I get validation BCE = x during training, for a single model with <a href=\"/humananalog\">@humananalog</a> inference kernel with some tweaks, I get LB = x + (0.01 to 0.03).</p>\n\n<p>To be honest, I cannot pinpoint why it works, as I spent 1-2 weeks creating my dataset incorporating a lot of small bits of hints from the discussions. After that, when I started using the dataset, it just worked!</p>\n\n<p>Try hints from different discussions to produce a better dataset with a lot of images (e.g. 500k~1m). Try different networks- some networks work well for us and some not. Try ensembling to see how it affects your score.</p>\n\n<p>After teaming up, my current score has improved a lot due to the expert tips and access to more GPUs, thanks to my amazing teammate! </p>\n\n<p>Hope it helps!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 758817,
          "author_name": "beeaware",
          "author_url": "",
          "post_date": "02/28/2020 07:24:52",
          "content": "<p><a href=\"/debanga\">@debanga</a> Thanks very much, that helps a lot. My training dataset is about 800k images(10 images/video,80k videos from 0-40 folders), but I have not used data augmentation, I still doubt that data augmentation can significantly reduce my loss, Your cv and LB can match so well, my cv looks not bad but i don't know why LB is so different, I will try data augmentation to see what happens. </p>\n\n<p>Thanks again!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759064,
          "author_name": "debanga",
          "author_url": "",
          "post_date": "02/28/2020 14:25:30",
          "content": "<p>I am using a few, e.g hflip, downscale; I think they help a bit to prevent overfitting.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759349,
          "author_name": "beeaware",
          "author_url": "",
          "post_date": "02/28/2020 23:20:31",
          "content": "<p>Thanks, I will try it, by the way, would you share your image size？My current is\n299*299.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759356,
          "author_name": "debanga",
          "author_url": "",
          "post_date": "02/28/2020 23:29:25",
          "content": "<p>224 :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759820,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "02/29/2020 13:54:05",
          "content": "<p><a href=\"/debanga\">@debanga</a>  thanks \nHow many frames from each video did u take \nIf more than one ,how u consolidated loss for multiple frames </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761418,
          "author_name": "beeaware",
          "author_url": "",
          "post_date": "03/02/2020 13:50:11",
          "content": "",
          "votes": null,
          "replies": []
        },
        {
          "id": 767945,
          "author_name": "mxmka87",
          "author_url": "",
          "post_date": "03/10/2020 09:36:04",
          "content": "<p>Was the original input size for xception model 224x224? Why it works with 150x150 then?\nAre CNNs not sensitive to different input? Or are they are, but with sufficient training are able to relearn to a different image size?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 749650,
      "author_name": "",
      "author_url": "",
      "post_date": "02/18/2020 20:43:52",
      "content": "<p>Great!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 752652,
      "author_name": "yuval6967",
      "author_url": "",
      "post_date": "02/21/2020 09:04:27",
      "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a> <a href=\"/debanga\">@debanga</a> Usin <a href=\"https://www.kaggle.com/greatgamedota/ffhq-face-data-set\">FFHQ-FACE</a> might be violating the competition rules. Look <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203#744289\">here</a> Using external data which is for non-commercial use only, seems to be in violation of the rules. </p>",
      "votes": null,
      "replies": [
        {
          "id": 752987,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "02/21/2020 15:45:07",
          "content": "<p>Thanks for the heads up, I looked into it and think you're right.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 752854,
      "author_name": "zungmann",
      "author_url": "",
      "post_date": "02/21/2020 13:28:28",
      "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a> thanks for sharing. I tried to download the ffhq dataset, but the link does not seem to work.\n<a href=\"https://www.kaggle.com/greatgamedota/ffhq-face-data-set/download\">https://www.kaggle.com/greatgamedota/ffhq-face-data-set/download</a>, getting 404 error.</p>",
      "votes": null,
      "replies": [
        {
          "id": 752922,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "02/21/2020 14:45:08",
          "content": "<p>I don't know why that doesn't work, use the kaggle api I guess</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "748809": "Notebook showing the training and inference of a simple Xception binary classifier:\nhttps://www.kaggle.com/greatgamedota/xception-classifier-w-ffhq-training-lb-555\nThe notebook shows training for LB of .555 however with minimal changes you can train a model that scores .489 and possibly less!",
    "748831": "Thanks! It's great.",
    "748845": "Nice! Great to know someone used AND ACHIEVED A PRETTY GOOD SCORE with my dataset I've abandoned for quite long(I discovered a better face detector(I will share it soon 😉) to achieve a new score! BTW, a hint: you might want to replace the backbone(xception) with another network😉 . My teammates will be mad at me if I give more details.\n\nEDIT: *My teammate discovered(Happy now? @harshitsheoran )",
    "748846": "Thanks very much for the dataset and all the hints you've dropped 🙂 Xception was my first model I've experimented with and will definitely try others!",
    "748977": "unkownhihi Thanks again for your dataset, my first one and a half month in this competition (without any GCP, AWS credits, and too poor to pay for GPU hours) your dataset was the only source of experimentation!",
    "749033": "I would be real mad if you even talk about that face detector xD. You can disclose it later in the competition, not in this week.\n\nEDIT : Those downvotes, I guess you guys really needs it, okay I guess I will allow that btw where he says that he discovered a better face detector, He means that it was me who gave it to him in the first place when we teamed up, it makes no sense to downvote, if you are really good and need to know that or need to know how our baseline is scoring more than most people, what tricks are we applying, cuz I guess we have got a lot of those to make our baseline score that, You can join with us, I am willing to take people who have a leaderboard score of less than 0.4.",
    "749075": "You guys are really creating suspense. May I ask this that have you changed your dataset or used the same for the new submission. @harshitsheoran @unkownhihi",
    "749081": "If I am not mistaken, We are using a new dataset for all of our new submissions, the current score is an ensemble of models scoring very similar. ~0.38 ensembling gave ~0.34.",
    "749263": "greatgamedota Thanks, man! I just tried your training kernel with my personal data + some augmentation and scored my best so far with a single Xception model LB 0.37222 😁\n\nPS: My logs for the above model:\n\n- Training loss: 0.1899\n- Validation loss: 0.3593\n- Validation Acc: 0.835837\n\nI &lt;3 my CV (folder-wise validation)!",
    "749277": "Nice result! I haven't even attempted a CV method yet.",
    "749330": "Weird thing, I am using same folders for the validation I get a very good validation loss but it does not track the leaderboard well, I can except leaderboard to be 0.1-0.2 above the validation loss in my case.",
    "749345": "Hmm... @harshitsheoran also to add, this time I tried different folderwise folds, and used folders 0-9 for the final model. You can check may be? If you try let me know if that makes any change, I’m also curious! For this competition CV is quite a mystery!",
    "749464": "My single frame models always score around 0.44. Looks like I missed the batch normalization and dropout part!",
    "749650": "Great!",
    "750459": "debanga  i get very good Val set results but get very less score. \nWhy should not get good score\n\n`0.058996\t0.174389\t0.950241`\nMy validation strategy was based on split of original video ids  and image size 128/128.\nPlease suggest me some thing good narrow down the gap between val and test results..",
    "750673": "That's a really good LB score for that validation score!",
    "750712": "but i dont know why i dont get eqw score .. :(\nlooking for some help",
    "750820": "jaideepvalani Your validation split is not good, I guess you should go for the last 10 folders as those are the folders with the faces which mostly are not included in folders 0-40.",
    "752652": "greatgamedota @debanga Usin [FFHQ-FACE](https://www.kaggle.com/greatgamedota/ffhq-face-data-set) might be violating the competition rules. Look [here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203#744289) Using external data which is for non-commercial use only, seems to be in violation of the rules.",
    "752854": "greatgamedota thanks for sharing. I tried to download the ffhq dataset, but the link does not seem to work.\nhttps://www.kaggle.com/greatgamedota/ffhq-face-data-set/download, getting 404 error.",
    "752922": "I don't know why that doesn't work, use the kaggle api I guess",
    "752987": "Thanks for the heads up, I looked into it and think you're right.",
    "758720": "debanga How do you calculate accuracy? I use the 0-39 folder as the training set and 40-48 as the validation set. I took 600 videos (300 real, 300 fake) from each folder for validation, and got the following results:\n\nFormat: folder, BCE, abs (y_true-y_pred)&gt; 0.3, abs (y_true-y_pred)&gt; 0.5, abs (y_true-y_pred)&gt; 0.7, abs (y_true-y_pred)&gt; 0.9\n\npart_40 0.4773 194 128 82 24\n\npart_41 0.3328 191 95 39 3\n\npart_42 0.2235 138 42 5 0\n\npart_43 0.3498 156 84 33 10\n\npart_44 0.2157 124 50 14 0\n\npart_45 0.4283 158 89 58 28\n\npart_46 0.3213 184 90 40 0\n\npart_47 0.2247 132 45 10 3\n\npart_48 0.3299 178 92 28 7\n\n\nBut my loss on the public LB is only 0.56. I don’t know if this is reasonable. I feel there is a big gap between the public LB and my local CV. Do you have any suggestions? Thanks very much!",
    "758778": "jaideepvalani:\n\nSee my suggestions below. Also, try higher resolution, which may improve your results.\n\n@beeaware:\nIn my case, accuracy is a by-product. I only focus on validation BCE.\n\nIn my current validation folds, if I get validation BCE = x during training, for a single model with @humananalog inference kernel with some tweaks, I get LB = x + (0.01 to 0.03).\n\nTo be honest, I cannot pinpoint why it works, as I spent 1-2 weeks creating my dataset incorporating a lot of small bits of hints from the discussions. After that, when I started using the dataset, it just worked!\n\nTry hints from different discussions to produce a better dataset with a lot of images (e.g. 500k~1m). Try different networks- some networks work well for us and some not. Try ensembling to see how it affects your score.\n\nAfter teaming up, my current score has improved a lot due to the expert tips and access to more GPUs, thanks to my amazing teammate! \n\nHope it helps!",
    "758817": "debanga Thanks very much, that helps a lot. My training dataset is about 800k images(10 images/video,80k videos from 0-40 folders), but I have not used data augmentation, I still doubt that data augmentation can significantly reduce my loss, Your cv and LB can match so well, my cv looks not bad but i don't know why LB is so different, I will try data augmentation to see what happens. \n\nThanks again!",
    "759064": "I am using a few, e.g hflip, downscale; I think they help a bit to prevent overfitting.",
    "759349": "Thanks, I will try it, by the way, would you share your image size？My current is\n299*299.",
    "759356": "224 :)",
    "759820": "debanga  thanks \nHow many frames from each video did u take \nIf more than one ,how u consolidated loss for multiple frames",
    "761418": "",
    "767945": "Was the original input size for xception model 224x224? Why it works with 150x150 then?\nAre CNNs not sensitive to different input? Or are they are, but with sufficient training are able to relearn to a different image size?"
  },
  "source": "meta"
}