{
  "id": 135648,
  "title": "Leaderboard groupings",
  "url": "/competitions/deepfake-detection-challenge/discussion/135648",
  "author_name": "ryches",
  "post_date": "2020-03-15T06:06:06.893000",
  "votes": 21,
  "comment_count": 31,
  "views": 0,
  "content": "<p>It's always interesting looking at the kaggle leaderboard stats and seeing how the distribution lands. Sometimes there is a gradual trailing off and everyone is clustering towards the top. In this competition, it has been quite a bit different. There are clear groupings of teams who have overcome various hurdles. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F42ed785e6460c04be97bceca5ec2922b%2FScreenshot%20from%202020-03-14%2022-59-42.png?generation=1584252006114093&amp;alt=media\" alt=\"\"></p>\n\n<p>These are just some hypotheses as to what might be the trick to get to these points. Looking at the top 250 teams we can see quite a lot are bunched up on the best public kernel around .43 and then a small group who have made some minor improvements beyond that, likely with ensembling or some model/data improvements. Beyond that, I hypothesize the next group has done something smart in terms of removing the noise in the labels, either removing frames from videos that were labeled fake that weren't really fake or removing poorly formed videos. </p>\n\n<p>The next step up is probably just some better method to do this than the previous grouping. They may have also found some way to make validation and test line up a bit better or make their model more robust. </p>\n\n<p>Beyond that, I have to assume there are some more clever methods being used. Alternative representations of the data, additional features, non-standard model architectures or massive ensembles. </p>",
  "messages": [
    {
      "id": 772177,
      "postDate": "2020-03-15T06:06:06.893Z",
      "content": "<p>It's always interesting looking at the kaggle leaderboard stats and seeing how the distribution lands. Sometimes there is a gradual trailing off and everyone is clustering towards the top. In this competition, it has been quite a bit different. There are clear groupings of teams who have overcome various hurdles. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F42ed785e6460c04be97bceca5ec2922b%2FScreenshot%20from%202020-03-14%2022-59-42.png?generation=1584252006114093&amp;alt=media\" alt=\"\"></p>\n\n<p>These are just some hypotheses as to what might be the trick to get to these points. Looking at the top 250 teams we can see quite a lot are bunched up on the best public kernel around .43 and then a small group who have made some minor improvements beyond that, likely with ensembling or some model/data improvements. Beyond that, I hypothesize the next group has done something smart in terms of removing the noise in the labels, either removing frames from videos that were labeled fake that weren't really fake or removing poorly formed videos. </p>\n\n<p>The next step up is probably just some better method to do this than the previous grouping. They may have also found some way to make validation and test line up a bit better or make their model more robust. </p>\n\n<p>Beyond that, I have to assume there are some more clever methods being used. Alternative representations of the data, additional features, non-standard model architectures or massive ensembles. </p>",
      "rawMarkdown": "It's always interesting looking at the kaggle leaderboard stats and seeing how the distribution lands. Sometimes there is a gradual trailing off and everyone is clustering towards the top. In this competition, it has been quite a bit different. There are clear groupings of teams who have overcome various hurdles. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F42ed785e6460c04be97bceca5ec2922b%2FScreenshot%20from%202020-03-14%2022-59-42.png?generation=1584252006114093&amp;alt=media)\n\nThese are just some hypotheses as to what might be the trick to get to these points. Looking at the top 250 teams we can see quite a lot are bunched up on the best public kernel around .43 and then a small group who have made some minor improvements beyond that, likely with ensembling or some model/data improvements. Beyond that, I hypothesize the next group has done something smart in terms of removing the noise in the labels, either removing frames from videos that were labeled fake that weren't really fake or removing poorly formed videos. \n\nThe next step up is probably just some better method to do this than the previous grouping. They may have also found some way to make validation and test line up a bit better or make their model more robust. \n\nBeyond that, I have to assume there are some more clever methods being used. Alternative representations of the data, additional features, non-standard model architectures or massive ensembles. \n",
      "votes": 21
    },
    {
      "id": 772713,
      "postDate": "2020-03-15T20:06:36.500Z",
      "content": "<p>Our 0.30003LB: No cleaning (simple uniform N frame extraction from video), ensembles of several single CNN models, multi-frame inference, good folder-wise validation.</p>",
      "rawMarkdown": "Our 0.30003LB: No cleaning (simple uniform N frame extraction from video), ensembles of several single CNN models, multi-frame inference, good folder-wise validation.",
      "votes": 14,
      "replies": [
        {
          "id": 784369,
          "postDate": "2020-03-24T06:53:54.310Z",
          "content": "<p>How many frames per video?and how mang CNN models do you use to ensamble? </p>",
          "rawMarkdown": "How many frames per video?and how mang CNN models do you use to ensamble? \n"
        }
      ]
    },
    {
      "id": 776667,
      "postDate": "2020-03-17T14:43:35.483Z",
      "content": "<p>we got 0.31 with single image classifier\nnot use ensemble or cleaning</p>",
      "rawMarkdown": "we got 0.31 with single image classifier\nnot use ensemble or cleaning",
      "votes": 10
    },
    {
      "id": 776443,
      "postDate": "2020-03-17T11:56:42.017Z",
      "content": "<p>I think leaders have focused on cropped models of eyes, nose and mouth. Reason being if we visualise the difference between real and fake bounding boxes the eyes, nose and mouth are the most frequently changed regions. Below are samples from easiest (top) to hardest/false positives (bottom).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F389345%2Fcda15c9344bf5f9aeb36fd76b497e5d1%2Fdownload%20(2\" alt=\"\">.jpg?generation=1584445924227726&amp;alt=media)</p>",
      "rawMarkdown": "I think leaders have focused on cropped models of eyes, nose and mouth. Reason being if we visualise the difference between real and fake bounding boxes the eyes, nose and mouth are the most frequently changed regions. Below are samples from easiest (top) to hardest/false positives (bottom).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F389345%2Fcda15c9344bf5f9aeb36fd76b497e5d1%2Fdownload%20(2).jpg?generation=1584445924227726&amp;alt=media)\n",
      "votes": 7,
      "replies": [
        {
          "id": 783151,
          "postDate": "2020-03-23T03:30:04.107Z",
          "content": "<p>Good analysis, I want to know how this difference map is generated, can you share the processing code?  Direct subtraction of pictures introduces a lot of background noise</p>",
          "rawMarkdown": "Good analysis, I want to know how this difference map is generated, can you share the processing code?  Direct subtraction of pictures introduces a lot of background noise",
          "votes": 2
        },
        {
          "id": 784130,
          "postDate": "2020-03-24T00:34:00.327Z",
          "content": "<p>Here is a code snippet.</p>\n\n<p><code>from PIL import Image, ImageOps, ImageChops, ImageEnhance</code>\n<code>real_img = Image.open(real_file)</code>\n<code>fake_img = Image.open(fake_file)</code>\n<code>diff_img = ImageEnhance.Brightness(ImageChops.difference(real_img, fake_img)).enhance(4.0)</code></p>",
          "rawMarkdown": "Here is a code snippet.\n\n`from PIL import Image, ImageOps, ImageChops, ImageEnhance`\n`real_img = Image.open(real_file)`\n`fake_img = Image.open(fake_file)`\n`diff_img = ImageEnhance.Brightness(ImageChops.difference(real_img, fake_img)).enhance(4.0)`",
          "votes": 2
        },
        {
          "id": 784203,
          "postDate": "2020-03-24T02:11:56.020Z",
          "content": "<p>Thanks~</p>",
          "rawMarkdown": "Thanks~"
        }
      ]
    },
    {
      "id": 772330,
      "postDate": "2020-03-15T11:00:58.397Z",
      "content": "<p><a href=\"/ryches\">@ryches</a> Interesting analysis about this competition!</p>\n\n<p>We're in 0.30 group and we didn't do any real cleaning other than selecting videos with less than 3 faces. I mean if our face detector finds more than 3 faces we keep only the best 3. If your assumption is correct then we should start to do some labels cleaning. Any thread here with known bad labels? It would be heplful for everyone.</p>\n\n<p>Also for the 0.25 group, maybe they're using audio (we're not because there is an additional work to get labels), if they're using images only then they've found something, I would bet it's based on frames sequence as such methods have good results in State Of The Art. However their implementation and tuning is not obvious, we've tried a few models with sequences and it always fails (underfit or overfit).</p>",
      "rawMarkdown": "@ryches Interesting analysis about this competition!\n\nWe're in 0.30 group and we didn't do any real cleaning other than selecting videos with less than 3 faces. I mean if our face detector finds more than 3 faces we keep only the best 3. If your assumption is correct then we should start to do some labels cleaning. Any thread here with known bad labels? It would be heplful for everyone.\n\nAlso for the 0.25 group, maybe they're using audio (we're not because there is an additional work to get labels), if they're using images only then they've found something, I would bet it's based on frames sequence as such methods have good results in State Of The Art. However their implementation and tuning is not obvious, we've tried a few models with sequences and it always fails (underfit or overfit).\n",
      "votes": 6,
      "replies": [
        {
          "id": 772613,
          "postDate": "2020-03-15T17:24:13.063Z",
          "content": "<p>Very interesting. My attempts with audio have not been particularly useful as some others have mentioned. Maybe that is the case though. </p>\n\n<p>And in terms of label cleaning I meant stuff like exactly what you did. Not using the labels and applying them across all frames, but rather making sure you only have relevant images associated with the label. </p>",
          "rawMarkdown": "Very interesting. My attempts with audio have not been particularly useful as some others have mentioned. Maybe that is the case though. \n\nAnd in terms of label cleaning I meant stuff like exactly what you did. Not using the labels and applying them across all frames, but rather making sure you only have relevant images associated with the label. ",
          "votes": 2
        },
        {
          "id": 786996,
          "postDate": "2020-03-26T12:39:30.627Z",
          "content": "<p>We've tried to remove noisy labels (assuming that some REAL were labelled as FAKE) and it did not help, score was not really better. Noisy labels detected (with high confidence) was about 1%-2% of our training set. Even, we've watched a few videos to confirm such noisy labels and it looks there are not FAKE at all (all frames seems REAL). Maybe what we consider as bad video labels are just great fake even human's eye cannot detect such as GAN ones at: <a href=\"https://www.thispersondoesnotexist.com/\">https://www.thispersondoesnotexist.com/</a>\nOrganizers should have included such high fake quality videos in dataset otherwise 1 Million price to detect bad fakes would not be worth.</p>",
          "rawMarkdown": "We've tried to remove noisy labels (assuming that some REAL were labelled as FAKE) and it did not help, score was not really better. Noisy labels detected (with high confidence) was about 1%-2% of our training set. Even, we've watched a few videos to confirm such noisy labels and it looks there are not FAKE at all (all frames seems REAL). Maybe what we consider as bad video labels are just great fake even human's eye cannot detect such as GAN ones at: https://www.thispersondoesnotexist.com/\nOrganizers should have included such high fake quality videos in dataset otherwise 1 Million price to detect bad fakes would not be worth.\n"
        }
      ]
    },
    {
      "id": 772706,
      "postDate": "2020-03-15T19:46:58.443Z",
      "content": "<p>For us, no cleaning label involved to get to 0.31, just 4 ensemble good single models.</p>",
      "rawMarkdown": "For us, no cleaning label involved to get to 0.31, just 4 ensemble good single models.",
      "votes": 4,
      "replies": [
        {
          "id": 775161,
          "postDate": "2020-03-16T11:21:37.133Z",
          "content": "<p>only 3 models not 4 lol</p>",
          "rawMarkdown": "only 3 models not 4 lol"
        },
        {
          "id": 779350,
          "postDate": "2020-03-19T08:47:51.137Z",
          "content": "<p>I am just happy that finally <a href=\"/unkownhihi\">@unkownhihi</a>  found a good name . </p>",
          "rawMarkdown": "I am just happy that finally @unkownhihi  found a good name . ",
          "votes": 1
        },
        {
          "id": 779670,
          "postDate": "2020-03-19T15:26:00.300Z",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> I was trying to find a good nickname, but I decided that I'm going to use my real name instead of a nickname. I feel like kaggle is somewhere I can trust, to not leak/use/make money of my information. </p>",
          "rawMarkdown": "@phoenix9032 I was trying to find a good nickname, but I decided that I'm going to use my real name instead of a nickname. I feel like kaggle is somewhere I can trust, to not leak/use/make money of my information. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 772336,
      "postDate": "2020-03-15T11:16:32.183Z",
      "content": "<p>I'm really curious to see what methods the top 5 are using. Vladislav (currently #2) posted on his LinkedIn that he used \"a simple solution\". It's probably not a neural network for image classification?!</p>",
      "rawMarkdown": "I'm really curious to see what methods the top 5 are using. Vladislav (currently #2) posted on his LinkedIn that he used \"a simple solution\". It's probably not a neural network for image classification?!",
      "votes": 1,
      "replies": [
        {
          "id": 772569,
          "postDate": "2020-03-15T16:31:37.247Z",
          "content": "<p>I guess he meant a 4 layers cnn? DL is not necessary at all</p>",
          "rawMarkdown": "I guess he meant a 4 layers cnn? DL is not necessary at all"
        },
        {
          "id": 772608,
          "postDate": "2020-03-15T17:21:45.443Z",
          "content": "<p>I would be intrigued by a simple method, I think there are likely some simple features that could be calculated to check continuity between frames and also extract various things like histogram manipulation and noise patterns. My bet is that the primary score driver is going to be the data prep/cleaning and the representation that the model receives. I think the specific NN architecture is likely not as important. </p>",
          "rawMarkdown": "I would be intrigued by a simple method, I think there are likely some simple features that could be calculated to check continuity between frames and also extract various things like histogram manipulation and noise patterns. My bet is that the primary score driver is going to be the data prep/cleaning and the representation that the model receives. I think the specific NN architecture is likely not as important. ",
          "votes": 1
        },
        {
          "id": 772641,
          "postDate": "2020-03-15T18:04:52.657Z",
          "content": "<blockquote>\n  <p>I guess he meant a 4 layers cnn? DL is not necessary at all</p>\n</blockquote>\n\n<p>I don't know what he meant but simple may mean different things to different people. He appears to work with face data in his job, so he probably knows more about this stuff than I do.</p>\n\n<p>Whether DL is not necessary or not... I'd love to see a solution that doesn't use it. (Technically speaking, a 4 layer CNN is also DL.)</p>",
          "rawMarkdown": "&gt; I guess he meant a 4 layers cnn? DL is not necessary at all\n\nI don't know what he meant but simple may mean different things to different people. He appears to work with face data in his job, so he probably knows more about this stuff than I do.\n\nWhether DL is not necessary or not... I'd love to see a solution that doesn't use it. (Technically speaking, a 4 layer CNN is also DL.)",
          "votes": 1
        },
        {
          "id": 775654,
          "postDate": "2020-03-16T22:59:53.450Z",
          "content": "<p>Maybe they turned it into a facial recognition problem. Firstly, they memorized the real characters (limited amount）with a vector. The rest is to detect which character it corresponds to and if the eyes, nose, mouth are modified by comparing the vectors. <br>\nFrom my experiments, the less params, the better the performance. </p>",
          "rawMarkdown": "Maybe they turned it into a facial recognition problem. Firstly, they memorized the real characters (limited amount）with a vector. The rest is to detect which character it corresponds to and if the eyes, nose, mouth are modified by comparing the vectors.  \nFrom my experiments, the less params, the better the performance. \n"
        },
        {
          "id": 775656,
          "postDate": "2020-03-16T23:14:18.473Z",
          "content": "<p>I utilized feature vectors created from the pretrained facenet neural network and wasnt able to get any measurable boost from using those features. </p>",
          "rawMarkdown": "I utilized feature vectors created from the pretrained facenet neural network and wasnt able to get any measurable boost from using those features. ",
          "votes": 2
        },
        {
          "id": 777542,
          "postDate": "2020-03-17T18:55:33.067Z",
          "content": "<p>Did you try to reduce the dimension of the vectors, I think vector from lower level of facenet should be more helpful.\nAnother idea is studying the distribution of noise and noise pattern of camera. However, I didn't have my own dataset and enough knowledge to test them out...</p>",
          "rawMarkdown": "Did you try to reduce the dimension of the vectors, I think vector from lower level of facenet should be more helpful.\nAnother idea is studying the distribution of noise and noise pattern of camera. However, I didn't have my own dataset and enough knowledge to test them out..."
        }
      ]
    },
    {
      "id": 786927,
      "postDate": "2020-03-26T10:57:57.543Z",
      "content": "<p>We only have ~500 faces here.</p>\n\n<p>The obvious way to improve one's score is to get more faces (for example, download a bunch of YouTube videos) and apply forgery methods similar to what the organizers used (or any publicly available methods).</p>\n\n<p>This is questionable from the legality perspective and has little value as far as improving the art of forgery detection.</p>\n\n<p>However, the organizers had many opportunities to state that this is not allowed and didn't take them.</p>",
      "rawMarkdown": "We only have ~500 faces here.\n\nThe obvious way to improve one's score is to get more faces (for example, download a bunch of YouTube videos) and apply forgery methods similar to what the organizers used (or any publicly available methods).\n\nThis is questionable from the legality perspective and has little value as far as improving the art of forgery detection.\n\nHowever, the organizers had many opportunities to state that this is not allowed and didn't take them.\n",
      "votes": 2,
      "replies": [
        {
          "id": 786955,
          "postDate": "2020-03-26T11:40:58.517Z",
          "content": "<p>Okay, Mr. but I did train a model with 8000 unique faces, more than tripling my training data, and no improvement, if you think generalization is a problem, well, guess what, it is not.</p>",
          "rawMarkdown": "Okay, Mr. but I did train a model with 8000 unique faces, more than tripling my training data, and no improvement, if you think generalization is a problem, well, guess what, it is not."
        },
        {
          "id": 786972,
          "postDate": "2020-03-26T12:13:23.713Z",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> Where did you get the faces, and which forgery software did you use?</p>",
          "rawMarkdown": "@harshitsheoran Where did you get the faces, and which forgery software did you use?"
        }
      ]
    },
    {
      "id": 775097,
      "postDate": "2020-03-16T09:47:43.387Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F75f2420699a5257913c4002bba92653c%2FHere.png?generation=1584352035205130&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F75f2420699a5257913c4002bba92653c%2FHere.png?generation=1584352035205130&amp;alt=media)\n",
      "votes": 2
    },
    {
      "id": 772203,
      "postDate": "2020-03-15T06:40:36.120Z",
      "content": "<p>Massive ensambles are not really feasible in this competition because of the time and space constraints </p>",
      "rawMarkdown": "Massive ensambles are not really feasible in this competition because of the time and space constraints ",
      "replies": [
        {
          "id": 772206,
          "postDate": "2020-03-15T06:43:08.950Z",
          "content": "<p>I've found that it is easily possible to fit 25+ models in 4 hours runtime since the bottleneck is the image loading and not the actual neural net inference</p>",
          "rawMarkdown": "I've found that it is easily possible to fit 25+ models in 4 hours runtime since the bottleneck is the image loading and not the actual neural net inference",
          "votes": 10
        },
        {
          "id": 775089,
          "postDate": "2020-03-16T09:41:06.583Z",
          "content": "<p><a href=\"/ryches\">@ryches</a>  This was quite important comment .</p>",
          "rawMarkdown": "@ryches  This was quite important comment ."
        }
      ]
    },
    {
      "id": 783775,
      "postDate": "2020-03-23T16:48:32.953Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 783960,
          "postDate": "2020-03-23T20:29:07.477Z",
          "content": "<p>Pretrained model is always more robust and converges faster. </p>",
          "rawMarkdown": "Pretrained model is always more robust and converges faster. ",
          "votes": 2
        },
        {
          "id": 784804,
          "postDate": "2020-03-24T14:10:05.807Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 772713,
      "author_name": "Debanga Raj Neog",
      "author_url": "",
      "post_date": "2020-03-15T20:06:36.500000",
      "content": "<p>Our 0.30003LB: No cleaning (simple uniform N frame extraction from video), ensembles of several single CNN models, multi-frame inference, good folder-wise validation.</p>",
      "votes": 14,
      "replies": [
        {
          "id": 784369,
          "author_name": "BokingChen",
          "author_url": "",
          "post_date": "2020-03-24T06:53:54.310000",
          "content": "<p>How many frames per video?and how mang CNN models do you use to ensamble? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 776667,
      "author_name": "あほ",
      "author_url": "",
      "post_date": "2020-03-17T14:43:35.483000",
      "content": "<p>we got 0.31 with single image classifier\nnot use ensemble or cleaning</p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 776443,
      "author_name": "maralski",
      "author_url": "",
      "post_date": "2020-03-17T11:56:42.017000",
      "content": "<p>I think leaders have focused on cropped models of eyes, nose and mouth. Reason being if we visualise the difference between real and fake bounding boxes the eyes, nose and mouth are the most frequently changed regions. Below are samples from easiest (top) to hardest/false positives (bottom).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F389345%2Fcda15c9344bf5f9aeb36fd76b497e5d1%2Fdownload%20(2\" alt=\"\">.jpg?generation=1584445924227726&amp;alt=media)</p>",
      "votes": 7,
      "replies": [
        {
          "id": 783151,
          "author_name": "sndl",
          "author_url": "",
          "post_date": "2020-03-23T03:30:04.107000",
          "content": "<p>Good analysis, I want to know how this difference map is generated, can you share the processing code?  Direct subtraction of pictures introduces a lot of background noise</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 784130,
          "author_name": "maralski",
          "author_url": "",
          "post_date": "2020-03-24T00:34:00.327000",
          "content": "<p>Here is a code snippet.</p>\n\n<p><code>from PIL import Image, ImageOps, ImageChops, ImageEnhance</code>\n<code>real_img = Image.open(real_file)</code>\n<code>fake_img = Image.open(fake_file)</code>\n<code>diff_img = ImageEnhance.Brightness(ImageChops.difference(real_img, fake_img)).enhance(4.0)</code></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 784203,
          "author_name": "sndl",
          "author_url": "",
          "post_date": "2020-03-24T02:11:56.020000",
          "content": "<p>Thanks~</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 772330,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2020-03-15T11:00:58.397000",
      "content": "<p><a href=\"/ryches\">@ryches</a> Interesting analysis about this competition!</p>\n\n<p>We're in 0.30 group and we didn't do any real cleaning other than selecting videos with less than 3 faces. I mean if our face detector finds more than 3 faces we keep only the best 3. If your assumption is correct then we should start to do some labels cleaning. Any thread here with known bad labels? It would be heplful for everyone.</p>\n\n<p>Also for the 0.25 group, maybe they're using audio (we're not because there is an additional work to get labels), if they're using images only then they've found something, I would bet it's based on frames sequence as such methods have good results in State Of The Art. However their implementation and tuning is not obvious, we've tried a few models with sequences and it always fails (underfit or overfit).</p>",
      "votes": 6,
      "replies": [
        {
          "id": 772613,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-03-15T17:24:13.063000",
          "content": "<p>Very interesting. My attempts with audio have not been particularly useful as some others have mentioned. Maybe that is the case though. </p>\n\n<p>And in terms of label cleaning I meant stuff like exactly what you did. Not using the labels and applying them across all frames, but rather making sure you only have relevant images associated with the label. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 786996,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-03-26T12:39:30.627000",
          "content": "<p>We've tried to remove noisy labels (assuming that some REAL were labelled as FAKE) and it did not help, score was not really better. Noisy labels detected (with high confidence) was about 1%-2% of our training set. Even, we've watched a few videos to confirm such noisy labels and it looks there are not FAKE at all (all frames seems REAL). Maybe what we consider as bad video labels are just great fake even human's eye cannot detect such as GAN ones at: <a href=\"https://www.thispersondoesnotexist.com/\">https://www.thispersondoesnotexist.com/</a>\nOrganizers should have included such high fake quality videos in dataset otherwise 1 Million price to detect bad fakes would not be worth.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 772706,
      "author_name": "Shangqiu Li",
      "author_url": "",
      "post_date": "2020-03-15T19:46:58.443000",
      "content": "<p>For us, no cleaning label involved to get to 0.31, just 4 ensemble good single models.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 775161,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-03-16T11:21:37.133000",
          "content": "<p>only 3 models not 4 lol</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 779350,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-03-19T08:47:51.137000",
          "content": "<p>I am just happy that finally <a href=\"/unkownhihi\">@unkownhihi</a>  found a good name . </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 779670,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-03-19T15:26:00.300000",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> I was trying to find a good nickname, but I decided that I'm going to use my real name instead of a nickname. I feel like kaggle is somewhere I can trust, to not leak/use/make money of my information. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 772336,
      "author_name": "Human Analog",
      "author_url": "",
      "post_date": "2020-03-15T11:16:32.183000",
      "content": "<p>I'm really curious to see what methods the top 5 are using. Vladislav (currently #2) posted on his LinkedIn that he used \"a simple solution\". It's probably not a neural network for image classification?!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 772569,
          "author_name": "yuanzhe zhou",
          "author_url": "",
          "post_date": "2020-03-15T16:31:37.247000",
          "content": "<p>I guess he meant a 4 layers cnn? DL is not necessary at all</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 772608,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-03-15T17:21:45.443000",
          "content": "<p>I would be intrigued by a simple method, I think there are likely some simple features that could be calculated to check continuity between frames and also extract various things like histogram manipulation and noise patterns. My bet is that the primary score driver is going to be the data prep/cleaning and the representation that the model receives. I think the specific NN architecture is likely not as important. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 772641,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2020-03-15T18:04:52.657000",
          "content": "<blockquote>\n  <p>I guess he meant a 4 layers cnn? DL is not necessary at all</p>\n</blockquote>\n\n<p>I don't know what he meant but simple may mean different things to different people. He appears to work with face data in his job, so he probably knows more about this stuff than I do.</p>\n\n<p>Whether DL is not necessary or not... I'd love to see a solution that doesn't use it. (Technically speaking, a 4 layer CNN is also DL.)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 775654,
          "author_name": "yuanzhe zhou",
          "author_url": "",
          "post_date": "2020-03-16T22:59:53.450000",
          "content": "<p>Maybe they turned it into a facial recognition problem. Firstly, they memorized the real characters (limited amount）with a vector. The rest is to detect which character it corresponds to and if the eyes, nose, mouth are modified by comparing the vectors. <br>\nFrom my experiments, the less params, the better the performance. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 775656,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-03-16T23:14:18.473000",
          "content": "<p>I utilized feature vectors created from the pretrained facenet neural network and wasnt able to get any measurable boost from using those features. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 777542,
          "author_name": "yuanzhe zhou",
          "author_url": "",
          "post_date": "2020-03-17T18:55:33.067000",
          "content": "<p>Did you try to reduce the dimension of the vectors, I think vector from lower level of facenet should be more helpful.\nAnother idea is studying the distribution of noise and noise pattern of camera. However, I didn't have my own dataset and enough knowledge to test them out...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 786927,
      "author_name": "Oleg Trott",
      "author_url": "",
      "post_date": "2020-03-26T10:57:57.543000",
      "content": "<p>We only have ~500 faces here.</p>\n\n<p>The obvious way to improve one's score is to get more faces (for example, download a bunch of YouTube videos) and apply forgery methods similar to what the organizers used (or any publicly available methods).</p>\n\n<p>This is questionable from the legality perspective and has little value as far as improving the art of forgery detection.</p>\n\n<p>However, the organizers had many opportunities to state that this is not allowed and didn't take them.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 786955,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-03-26T11:40:58.517000",
          "content": "<p>Okay, Mr. but I did train a model with 8000 unique faces, more than tripling my training data, and no improvement, if you think generalization is a problem, well, guess what, it is not.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 786972,
          "author_name": "Oleg Trott",
          "author_url": "",
          "post_date": "2020-03-26T12:13:23.713000",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> Where did you get the faces, and which forgery software did you use?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 775097,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2020-03-16T09:47:43.387000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F75f2420699a5257913c4002bba92653c%2FHere.png?generation=1584352035205130&amp;alt=media\" alt=\"\"></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 772203,
      "author_name": "Moshel",
      "author_url": "",
      "post_date": "2020-03-15T06:40:36.120000",
      "content": "<p>Massive ensambles are not really feasible in this competition because of the time and space constraints </p>",
      "votes": 0,
      "replies": [
        {
          "id": 772206,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-03-15T06:43:08.950000",
          "content": "<p>I've found that it is easily possible to fit 25+ models in 4 hours runtime since the bottleneck is the image loading and not the actual neural net inference</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 775089,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-03-16T09:41:06.583000",
          "content": "<p><a href=\"/ryches\">@ryches</a>  This was quite important comment .</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 783775,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-23T16:48:32.953000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 783960,
          "author_name": "yuanzhe zhou",
          "author_url": "",
          "post_date": "2020-03-23T20:29:07.477000",
          "content": "<p>Pretrained model is always more robust and converges faster. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 784804,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-24T14:10:05.807000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "772177": "It's always interesting looking at the kaggle leaderboard stats and seeing how the distribution lands. Sometimes there is a gradual trailing off and everyone is clustering towards the top. In this competition, it has been quite a bit different. There are clear groupings of teams who have overcome various hurdles. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F42ed785e6460c04be97bceca5ec2922b%2FScreenshot%20from%202020-03-14%2022-59-42.png?generation=1584252006114093&amp;alt=media)\n\nThese are just some hypotheses as to what might be the trick to get to these points. Looking at the top 250 teams we can see quite a lot are bunched up on the best public kernel around .43 and then a small group who have made some minor improvements beyond that, likely with ensembling or some model/data improvements. Beyond that, I hypothesize the next group has done something smart in terms of removing the noise in the labels, either removing frames from videos that were labeled fake that weren't really fake or removing poorly formed videos. \n\nThe next step up is probably just some better method to do this than the previous grouping. They may have also found some way to make validation and test line up a bit better or make their model more robust. \n\nBeyond that, I have to assume there are some more clever methods being used. Alternative representations of the data, additional features, non-standard model architectures or massive ensembles. \n",
    "772713": "Our 0.30003LB: No cleaning (simple uniform N frame extraction from video), ensembles of several single CNN models, multi-frame inference, good folder-wise validation.",
    "776667": "we got 0.31 with single image classifier\nnot use ensemble or cleaning",
    "776443": "I think leaders have focused on cropped models of eyes, nose and mouth. Reason being if we visualise the difference between real and fake bounding boxes the eyes, nose and mouth are the most frequently changed regions. Below are samples from easiest (top) to hardest/false positives (bottom).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F389345%2Fcda15c9344bf5f9aeb36fd76b497e5d1%2Fdownload%20(2).jpg?generation=1584445924227726&amp;alt=media)\n",
    "772330": "@ryches Interesting analysis about this competition!\n\nWe're in 0.30 group and we didn't do any real cleaning other than selecting videos with less than 3 faces. I mean if our face detector finds more than 3 faces we keep only the best 3. If your assumption is correct then we should start to do some labels cleaning. Any thread here with known bad labels? It would be heplful for everyone.\n\nAlso for the 0.25 group, maybe they're using audio (we're not because there is an additional work to get labels), if they're using images only then they've found something, I would bet it's based on frames sequence as such methods have good results in State Of The Art. However their implementation and tuning is not obvious, we've tried a few models with sequences and it always fails (underfit or overfit).\n",
    "772706": "For us, no cleaning label involved to get to 0.31, just 4 ensemble good single models.",
    "772336": "I'm really curious to see what methods the top 5 are using. Vladislav (currently #2) posted on his LinkedIn that he used \"a simple solution\". It's probably not a neural network for image classification?!",
    "786927": "We only have ~500 faces here.\n\nThe obvious way to improve one's score is to get more faces (for example, download a bunch of YouTube videos) and apply forgery methods similar to what the organizers used (or any publicly available methods).\n\nThis is questionable from the legality perspective and has little value as far as improving the art of forgery detection.\n\nHowever, the organizers had many opportunities to state that this is not allowed and didn't take them.\n",
    "775097": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F75f2420699a5257913c4002bba92653c%2FHere.png?generation=1584352035205130&amp;alt=media)\n",
    "772203": "Massive ensambles are not really feasible in this competition because of the time and space constraints ",
    "783775": ""
  }
}