{
  "id": 139377,
  "title": "No-holdout validation and potential surprises",
  "url": "/competitions/deepfake-detection-challenge/discussion/139377",
  "author_name": "",
  "post_date": "2020-03-28T15:48:21.251000700Z",
  "votes": 10,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Well, it has been fun but I think I've played enough with this competition by now.  </p>\n\n<p>I'd like to share some thoughts on the nature of overfitting here and possible implications for private set.</p>\n\n<p>This is what I did weeks ago:  No matter how carefully preparing datasets there was some source of leakage to validation; folder based splits, cluster based splits, a bunch of different tricks and still overfitting. What I found quite weird. So I decided to try a no-holdout validation approach, as described here <a href=\"https://openreview.net/pdf?id=B1lKtjA9FQ\">no holdout cv</a></p>\n\n<p>The idea is simple, you want to find out the point in which your model is memorizing the train set, dependent on architecture and hyperparameters. So you destroy/randomize labels and train normally. You can then have a general idea of the speed the model memorizes training. The experiment showed that it doesnt take long to memorize this dataset but -as expected- no validation leakage at all.</p>\n\n<p>So... what does this mean? It means that discarding some shameful bug in my dataset preparation <strong>there seems to be a label related source of overfitting</strong>.</p>\n\n<p><strong>My best explanation is related to fake videos generation process. Specifically, fake faces are chosen in each case be \"similar\" to the real ones. This means that -in the train set- the model has an easy way to overfit: memorize fake faces</strong>. </p>\n\n<p>The obvious problem is,  a face that has been used as a fake might be precisely the real one in the test set, leading to the worst case of overfitting/concept shift; when a train positive is a test negative. Or, of course, fake faces might be just new ones.</p>\n\n<p>Will this interpretation be correct? The extent of final shakeup will say...</p>",
  "messages": [
    {
      "id": "789373",
      "postDate": "03/28/2020 15:48:21",
      "content": "<p>Well, it has been fun but I think I've played enough with this competition by now.  </p>\n\n<p>I'd like to share some thoughts on the nature of overfitting here and possible implications for private set.</p>\n\n<p>This is what I did weeks ago:  No matter how carefully preparing datasets there was some source of leakage to validation; folder based splits, cluster based splits, a bunch of different tricks and still overfitting. What I found quite weird. So I decided to try a no-holdout validation approach, as described here <a href=\"https://openreview.net/pdf?id=B1lKtjA9FQ\">no holdout cv</a></p>\n\n<p>The idea is simple, you want to find out the point in which your model is memorizing the train set, dependent on architecture and hyperparameters. So you destroy/randomize labels and train normally. You can then have a general idea of the speed the model memorizes training. The experiment showed that it doesnt take long to memorize this dataset but -as expected- no validation leakage at all.</p>\n\n<p>So... what does this mean? It means that discarding some shameful bug in my dataset preparation <strong>there seems to be a label related source of overfitting</strong>.</p>\n\n<p><strong>My best explanation is related to fake videos generation process. Specifically, fake faces are chosen in each case be \"similar\" to the real ones. This means that -in the train set- the model has an easy way to overfit: memorize fake faces</strong>. </p>\n\n<p>The obvious problem is,  a face that has been used as a fake might be precisely the real one in the test set, leading to the worst case of overfitting/concept shift; when a train positive is a test negative. Or, of course, fake faces might be just new ones.</p>\n\n<p>Will this interpretation be correct? The extent of final shakeup will say...</p>",
      "rawMarkdown": "Well, it has been fun but I think I've played enough with this competition by now.  \n\nI'd like to share some thoughts on the nature of overfitting here and possible implications for private set.\n\nThis is what I did weeks ago:  No matter how carefully preparing datasets there was some source of leakage to validation; folder based splits, cluster based splits, a bunch of different tricks and still overfitting. What I found quite weird. So I decided to try a no-holdout validation approach, as described here [no holdout cv](https://openreview.net/pdf?id=B1lKtjA9FQ)\n\nThe idea is simple, you want to find out the point in which your model is memorizing the train set, dependent on architecture and hyperparameters. So you destroy/randomize labels and train normally. You can then have a general idea of the speed the model memorizes training. The experiment showed that it doesnt take long to memorize this dataset but -as expected- no validation leakage at all.\n\nSo... what does this mean? It means that discarding some shameful bug in my dataset preparation **there seems to be a label related source of overfitting**.\n\n**My best explanation is related to fake videos generation process. Specifically, fake faces are chosen in each case be \"similar\" to the real ones. This means that -in the train set- the model has an easy way to overfit: memorize fake faces**. \n\nThe obvious problem is,  a face that has been used as a fake might be precisely the real one in the test set, leading to the worst case of overfitting/concept shift; when a train positive is a test negative. Or, of course, fake faces might be just new ones.\n\nWill this interpretation be correct? The extent of final shakeup will say...",
      "votes": null
    },
    {
      "id": "789397",
      "postDate": "03/28/2020 16:10:16",
      "content": "<p>Overfitting is the curse of this competition. So that's an interesting point of view. Let's see how true this is at the end of the competition. </p>",
      "rawMarkdown": "Overfitting is the curse of this competition. So that's an interesting point of view. Let's see how true this is at the end of the competition.",
      "votes": null
    },
    {
      "id": "789739",
      "postDate": "03/29/2020 00:04:30",
      "content": "<p>An interesting fact is that mislabeled faces can prevent overfitting. And data cleaning will make things worse?  </p>",
      "rawMarkdown": "An interesting fact is that mislabeled faces can prevent overfitting. And data cleaning will make things worse?",
      "votes": null
    },
    {
      "id": "789820",
      "postDate": "03/29/2020 02:44:26",
      "content": "<p>I kind of understand what you are saying, but I'm not sure. My understanding is two fold: 1) there are a small set of actors relative to the number of videos 2) fake videos have a corresponding real video and real videos are tied to an average of 10 fake videos if I remember correctly. This dataset has overfitting written all over it. A dataset of random youtube videos would have been much better.</p>",
      "rawMarkdown": "I kind of understand what you are saying, but I'm not sure. My understanding is two fold: 1) there are a small set of actors relative to the number of videos 2) fake videos have a corresponding real video and real videos are tied to an average of 10 fake videos if I remember correctly. This dataset has overfitting written all over it. A dataset of random youtube videos would have been much better.",
      "votes": null
    },
    {
      "id": "790187",
      "postDate": "03/29/2020 11:46:49",
      "content": "<p>Both 1) and 2) are true here, \nbut I'm pointing to: 3) In the -automated- fakes generation process for this dataset another face is used from the group of actors based on similarity with the original. This means that \"fakeness\" of a video can be detected by recognizing those particular actors' faces. This explains the amount of overfitting even when 1) and 2) are taken care off.</p>",
      "rawMarkdown": "Both 1) and 2) are true here, \nbut I'm pointing to: 3) In the -automated- fakes generation process for this dataset another face is used from the group of actors based on similarity with the original. This means that \"fakeness\" of a video can be detected by recognizing those particular actors' faces. This explains the amount of overfitting even when 1) and 2) are taken care off.",
      "votes": null
    },
    {
      "id": "790194",
      "postDate": "03/29/2020 11:55:18",
      "content": "<p>So your saying:</p>\n\n<p>Actor A real video\nActor A' fake video\nActor A' fake face replaced with Actor A real face</p>\n\n<p>?</p>",
      "rawMarkdown": "So your saying:\n\nActor A real video\nActor A' fake video\nActor A' fake face replaced with Actor A real face\n\n?",
      "votes": null
    },
    {
      "id": "790275",
      "postDate": "03/29/2020 13:25:48",
      "content": "<p>All I'm saying is that in <a href=\"https://arxiv.org/pdf/1910.08854.pdf\">this paper</a> explanation you can read:</p>\n\n<p>&gt; A number of face swaps were computed <strong>across subjects with similar appearances</strong>, where each appearance was inferred from facial attributes (skin tone, facial hair, glasses, etc.)</p>\n\n<p>What the way I see it implies there are two actors involved in face swap fakes, the real one and the \"fake\" one chosen based on its similarity. And this \"fake\" faces can be memorized by the model, thus overfitting.</p>",
      "rawMarkdown": "All I'm saying is that in [this paper](https://arxiv.org/pdf/1910.08854.pdf) explanation you can read:\n\n&gt; A number of face swaps were computed **across subjects with similar appearances**, where each appearance was inferred from facial attributes (skin tone, facial hair, glasses, etc.)\n\nWhat the way I see it implies there are two actors involved in face swap fakes, the real one and the \"fake\" one chosen based on its similarity. And this \"fake\" faces can be memorized by the model, thus overfitting.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 789397,
      "author_name": "yassinealouini",
      "author_url": "",
      "post_date": "03/28/2020 16:10:16",
      "content": "<p>Overfitting is the curse of this competition. So that's an interesting point of view. Let's see how true this is at the end of the competition. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 789739,
      "author_name": "yuanzhezhou",
      "author_url": "",
      "post_date": "03/29/2020 00:04:30",
      "content": "<p>An interesting fact is that mislabeled faces can prevent overfitting. And data cleaning will make things worse?  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 789820,
      "author_name": "sherkt1",
      "author_url": "",
      "post_date": "03/29/2020 02:44:26",
      "content": "<p>I kind of understand what you are saying, but I'm not sure. My understanding is two fold: 1) there are a small set of actors relative to the number of videos 2) fake videos have a corresponding real video and real videos are tied to an average of 10 fake videos if I remember correctly. This dataset has overfitting written all over it. A dataset of random youtube videos would have been much better.</p>",
      "votes": null,
      "replies": [
        {
          "id": 790187,
          "author_name": "miguelpm",
          "author_url": "",
          "post_date": "03/29/2020 11:46:49",
          "content": "<p>Both 1) and 2) are true here, \nbut I'm pointing to: 3) In the -automated- fakes generation process for this dataset another face is used from the group of actors based on similarity with the original. This means that \"fakeness\" of a video can be detected by recognizing those particular actors' faces. This explains the amount of overfitting even when 1) and 2) are taken care off.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 790194,
          "author_name": "maralski",
          "author_url": "",
          "post_date": "03/29/2020 11:55:18",
          "content": "<p>So your saying:</p>\n\n<p>Actor A real video\nActor A' fake video\nActor A' fake face replaced with Actor A real face</p>\n\n<p>?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 790275,
          "author_name": "miguelpm",
          "author_url": "",
          "post_date": "03/29/2020 13:25:48",
          "content": "<p>All I'm saying is that in <a href=\"https://arxiv.org/pdf/1910.08854.pdf\">this paper</a> explanation you can read:</p>\n\n<p>&gt; A number of face swaps were computed <strong>across subjects with similar appearances</strong>, where each appearance was inferred from facial attributes (skin tone, facial hair, glasses, etc.)</p>\n\n<p>What the way I see it implies there are two actors involved in face swap fakes, the real one and the \"fake\" one chosen based on its similarity. And this \"fake\" faces can be memorized by the model, thus overfitting.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "789373": "Well, it has been fun but I think I've played enough with this competition by now.  \n\nI'd like to share some thoughts on the nature of overfitting here and possible implications for private set.\n\nThis is what I did weeks ago:  No matter how carefully preparing datasets there was some source of leakage to validation; folder based splits, cluster based splits, a bunch of different tricks and still overfitting. What I found quite weird. So I decided to try a no-holdout validation approach, as described here [no holdout cv](https://openreview.net/pdf?id=B1lKtjA9FQ)\n\nThe idea is simple, you want to find out the point in which your model is memorizing the train set, dependent on architecture and hyperparameters. So you destroy/randomize labels and train normally. You can then have a general idea of the speed the model memorizes training. The experiment showed that it doesnt take long to memorize this dataset but -as expected- no validation leakage at all.\n\nSo... what does this mean? It means that discarding some shameful bug in my dataset preparation **there seems to be a label related source of overfitting**.\n\n**My best explanation is related to fake videos generation process. Specifically, fake faces are chosen in each case be \"similar\" to the real ones. This means that -in the train set- the model has an easy way to overfit: memorize fake faces**. \n\nThe obvious problem is,  a face that has been used as a fake might be precisely the real one in the test set, leading to the worst case of overfitting/concept shift; when a train positive is a test negative. Or, of course, fake faces might be just new ones.\n\nWill this interpretation be correct? The extent of final shakeup will say...",
    "789397": "Overfitting is the curse of this competition. So that's an interesting point of view. Let's see how true this is at the end of the competition.",
    "789739": "An interesting fact is that mislabeled faces can prevent overfitting. And data cleaning will make things worse?",
    "789820": "I kind of understand what you are saying, but I'm not sure. My understanding is two fold: 1) there are a small set of actors relative to the number of videos 2) fake videos have a corresponding real video and real videos are tied to an average of 10 fake videos if I remember correctly. This dataset has overfitting written all over it. A dataset of random youtube videos would have been much better.",
    "790187": "Both 1) and 2) are true here, \nbut I'm pointing to: 3) In the -automated- fakes generation process for this dataset another face is used from the group of actors based on similarity with the original. This means that \"fakeness\" of a video can be detected by recognizing those particular actors' faces. This explains the amount of overfitting even when 1) and 2) are taken care off.",
    "790194": "So your saying:\n\nActor A real video\nActor A' fake video\nActor A' fake face replaced with Actor A real face\n\n?",
    "790275": "All I'm saying is that in [this paper](https://arxiv.org/pdf/1910.08854.pdf) explanation you can read:\n\n&gt; A number of face swaps were computed **across subjects with similar appearances**, where each appearance was inferred from facial attributes (skin tone, facial hair, glasses, etc.)\n\nWhat the way I see it implies there are two actors involved in face swap fakes, the real one and the \"fake\" one chosen based on its similarity. And this \"fake\" faces can be memorized by the model, thus overfitting."
  },
  "source": "meta"
}