{
  "id": 129537,
  "title": "Will NeuralTextures be on the test? :-)",
  "url": "/competitions/deepfake-detection-challenge/discussion/129537",
  "author_name": "",
  "post_date": "2020-02-08T18:55:44.033291200Z",
  "votes": 2,
  "comment_count": 12,
  "views": 0,
  "content": "<p>NeuralTextures look almost perceptually authentic to a human, let alone a machine.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F351627%2Fb1bf131960c78613401f55740349248c%2Ffake.png?generation=1581187572634792&amp;alt=media\" alt=\"\"></p>\n\n<p>Suppose the test set contains <code>fraction_real</code> and <code>fraction_neuraltextures</code>, and your algorithm can't tell the difference[1].</p>\n\n<p>Then when the algorithm sees something authentic-looking, and it doesn't predict</p>\n\n<pre><code>fraction_neuraltextures / (fraction_neuraltextures + fraction_real)\n</code></pre>\n\n<p>it will lose to similar algorithms that do.</p>\n\n<p>Since we don't know these numbers[2], a large part of this contest seems to be guessing them, making this contest something of a lottery.</p>\n\n<p>This must demotivate serious competitors: you could expend a lot of effort to submit a sound entry, but your odds of winning are lottery-like (Even less, if another competitor learned <code>fraction_real</code> and <code>fraction_neuraltextures</code> through proximity to the organizers, for example, or by asking them privately)</p>\n\n<p>To my mind, stating these two numbers publicly would make this contest infinitely better.</p>\n\n<hr>\n\n<p>[1] Even if your algorithm is not entirely useless on NeuralTextures, I hope it's easy to see that knowing these numbers is critically important to what predictions it should make.</p>\n\n<p>[2] We know that <code>fraction_real</code> is 0.5 on the public test, but not on the private test set.</p>",
  "messages": [
    {
      "id": "740046",
      "postDate": "02/08/2020 18:55:44",
      "content": "<p>NeuralTextures look almost perceptually authentic to a human, let alone a machine.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F351627%2Fb1bf131960c78613401f55740349248c%2Ffake.png?generation=1581187572634792&amp;alt=media\" alt=\"\"></p>\n\n<p>Suppose the test set contains <code>fraction_real</code> and <code>fraction_neuraltextures</code>, and your algorithm can't tell the difference[1].</p>\n\n<p>Then when the algorithm sees something authentic-looking, and it doesn't predict</p>\n\n<pre><code>fraction_neuraltextures / (fraction_neuraltextures + fraction_real)\n</code></pre>\n\n<p>it will lose to similar algorithms that do.</p>\n\n<p>Since we don't know these numbers[2], a large part of this contest seems to be guessing them, making this contest something of a lottery.</p>\n\n<p>This must demotivate serious competitors: you could expend a lot of effort to submit a sound entry, but your odds of winning are lottery-like (Even less, if another competitor learned <code>fraction_real</code> and <code>fraction_neuraltextures</code> through proximity to the organizers, for example, or by asking them privately)</p>\n\n<p>To my mind, stating these two numbers publicly would make this contest infinitely better.</p>\n\n<hr>\n\n<p>[1] Even if your algorithm is not entirely useless on NeuralTextures, I hope it's easy to see that knowing these numbers is critically important to what predictions it should make.</p>\n\n<p>[2] We know that <code>fraction_real</code> is 0.5 on the public test, but not on the private test set.</p>",
      "rawMarkdown": "NeuralTextures look almost perceptually authentic to a human, let alone a machine.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F351627%2Fb1bf131960c78613401f55740349248c%2Ffake.png?generation=1581187572634792&amp;alt=media)\n\n\n\nSuppose the test set contains `fraction_real` and `fraction_neuraltextures`, and your algorithm can't tell the difference[1].\n\nThen when the algorithm sees something authentic-looking, and it doesn't predict\n\n    fraction_neuraltextures / (fraction_neuraltextures + fraction_real)\n\nit will lose to similar algorithms that do.\n\nSince we don't know these numbers[2], a large part of this contest seems to be guessing them, making this contest something of a lottery.\n\nThis must demotivate serious competitors: you could expend a lot of effort to submit a sound entry, but your odds of winning are lottery-like (Even less, if another competitor learned `fraction_real` and `fraction_neuraltextures` through proximity to the organizers, for example, or by asking them privately)\n\nTo my mind, stating these two numbers publicly would make this contest infinitely better.\n\n---\n\n[1] Even if your algorithm is not entirely useless on NeuralTextures, I hope it's easy to see that knowing these numbers is critically important to what predictions it should make.\n\n[2] We know that `fraction_real` is 0.5 on the public test, but not on the private test set.",
      "votes": null
    },
    {
      "id": "740245",
      "postDate": "02/09/2020 04:37:01",
      "content": "<p>Is data distribution different in training and test sets? Isn't the basic assumption of ML that they are sampled from the same distribution?</p>",
      "rawMarkdown": "Is data distribution different in training and test sets? Isn't the basic assumption of ML that they are sampled from the same distribution?",
      "votes": null
    },
    {
      "id": "740429",
      "postDate": "02/09/2020 12:13:50",
      "content": "<blockquote>\n  <p>Isn't the basic assumption of ML that they are sampled from the same distribution?</p>\n</blockquote>\n\n<p>Haha, not on Kaggle. ;-) Part of building a solution is to think of ways that the competition organizer may have changed the private test set to make it harder. This why people whose models don't generalize enough will get great scores on the public leaderboard but then do much worse on the private LB.</p>",
      "rawMarkdown": "&gt; Isn't the basic assumption of ML that they are sampled from the same distribution?\n\nHaha, not on Kaggle. ;-) Part of building a solution is to think of ways that the competition organizer may have changed the private test set to make it harder. This why people whose models don't generalize enough will get great scores on the public leaderboard but then do much worse on the private LB.",
      "votes": null
    },
    {
      "id": "740557",
      "postDate": "02/09/2020 14:27:03",
      "content": "<p>She looks like minecraft steve though..</p>",
      "rawMarkdown": "She looks like minecraft steve though..",
      "votes": null
    },
    {
      "id": "740682",
      "postDate": "02/09/2020 16:49:54",
      "content": "<blockquote>\n  <p><strong>Human Analog wrote:</strong></p>\n  \n  <p>Haha, not on Kaggle. ;-) Part of building a solution is to think of ways that the competition organizer may have changed the private test set to make it harder. This why people whose models don't generalize enough will get great scores on the public leaderboard but then do much worse on the private LB.</p>\n</blockquote>\n\n<p>Generalization is good performance on novel samples from the <em>same</em> distribution.</p>\n\n<p>Which competitions changed the distribution? I think it's rare.</p>",
      "rawMarkdown": "&gt; **Human Analog wrote:**\n&gt; \n&gt; Haha, not on Kaggle. ;-) Part of building a solution is to think of ways that the competition organizer may have changed the private test set to make it harder. This why people whose models don't generalize enough will get great scores on the public leaderboard but then do much worse on the private LB.\n\nGeneralization is good performance on novel samples from the *same* distribution.\n\nWhich competitions changed the distribution? I think it's rare.",
      "votes": null
    },
    {
      "id": "740705",
      "postDate": "02/09/2020 17:15:53",
      "content": "<p>&gt; <strong>Sumit Mishra wrote:</strong>\n&gt; \n&gt; She looks like minecraft steve though..</p>\n\n<p>Really? Here is frame 35 from a real and forged video (face only) :</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F351627%2F65a8d30537b376a8b1a810ea3194eb59%2F1.png?generation=1581272587789257&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F351627%2F6a5e584b7cd316f4e8ddd23119e786db%2F11.png?generation=1581272602552317&amp;alt=media\" alt=\"\"></p>\n\n<p>Which one is fake?</p>",
      "rawMarkdown": "&gt; **Sumit Mishra wrote:**\n&gt; \n&gt; She looks like minecraft steve though..\n\nReally? Here is frame 35 from a real and forged video (face only) :\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F351627%2F65a8d30537b376a8b1a810ea3194eb59%2F1.png?generation=1581272587789257&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F351627%2F6a5e584b7cd316f4e8ddd23119e786db%2F11.png?generation=1581272602552317&amp;alt=media)\n\n\n\nWhich one is fake?",
      "votes": null
    },
    {
      "id": "740751",
      "postDate": "02/09/2020 18:31:07",
      "content": "<blockquote>\n  <p>Which one is fake?</p>\n</blockquote>\n\n<p>Just because it's hard to see for a human doesn't mean it's hard for a computer. The classifier may detect patterns in the fake image that are not present in the real image (or vice versa). </p>",
      "rawMarkdown": "&gt; Which one is fake?\n\nJust because it's hard to see for a human doesn't mean it's hard for a computer. The classifier may detect patterns in the fake image that are not present in the real image (or vice versa).",
      "votes": null
    },
    {
      "id": "740752",
      "postDate": "02/09/2020 18:32:02",
      "content": "<p>The same frame from a real and forged video:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F351627%2F97960a3b7b03650adf84446e538e44a1%2F3.png?generation=1581273034862445&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F351627%2F8d83c25be7f4e790f40aa39c010323b6%2F33.png?generation=1581273059942158&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "The same frame from a real and forged video:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F351627%2F97960a3b7b03650adf84446e538e44a1%2F3.png?generation=1581273034862445&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F351627%2F8d83c25be7f4e790f40aa39c010323b6%2F33.png?generation=1581273059942158&amp;alt=media)",
      "votes": null
    },
    {
      "id": "740758",
      "postDate": "02/09/2020 18:46:14",
      "content": "<p>&gt; <strong>Human Analog wrote:</strong>\n&gt; \n&gt; Just because it's hard to see for a human doesn't mean it's hard for a computer.</p>\n\n<p>My point is that some forgery methods will be much harder to detect than others. The probabilities you predict must depend on the fractions of these methods in the test set. If the organizers changed these fractions, you are forced to <strong>guess</strong> what they did or didn't do (See the original post, which outlines a somewhat extreme example of when your algorithm is completely useless against some method, NeuralTextures. But similar considerations apply when it's not useless, but just performs worse than against other methods).</p>",
      "rawMarkdown": "&gt; **Human Analog wrote:**\n&gt; \n&gt; Just because it's hard to see for a human doesn't mean it's hard for a computer.\n\nMy point is that some forgery methods will be much harder to detect than others. The probabilities you predict must depend on the fractions of these methods in the test set. If the organizers changed these fractions, you are forced to **guess** what they did or didn't do (See the original post, which outlines a somewhat extreme example of when your algorithm is completely useless against some method, NeuralTextures. But similar considerations apply when it's not useless, but just performs worse than against other methods).",
      "votes": null
    },
    {
      "id": "740795",
      "postDate": "02/09/2020 19:55:59",
      "content": "<blockquote>\n  <p>you are forced to guess what they did or didn't do</p>\n</blockquote>\n\n<p>This has been the case for all the competitions I've been in (which isn't that many so perhaps I just picked the ones where this happened).</p>",
      "rawMarkdown": "&gt; you are forced to guess what they did or didn't do\n\nThis has been the case for all the competitions I've been in (which isn't that many so perhaps I just picked the ones where this happened).",
      "votes": null
    },
    {
      "id": "741207",
      "postDate": "02/10/2020 10:15:58",
      "content": "<p>Case in point: my best scoring model is one that I didn't spend any time tweaking the training process on. I just picked something reasonable and let it train. </p>\n\n<p>Over the weekend I improved the training process (1-cycle learning policy, AdamW optimizer, fine-tuning the earlier layers with different LRs for different groups of layers, etc). The model now learned a lot better (according to the train &amp; val losses), but its leaderboard score was worse. </p>\n\n<p>I don't think the model was overfitting per se (the validation loss kept going down). Even though it was learning better than before, it was just learning the wrong things.</p>\n\n<p>The results is that the model that didn't learn as well, actually generalized better on the test set used for the leaderboard.</p>\n\n<p>So if your model does better on the training set, but worse on the leaderboard, it's either overfitting -- or the distribution of the leaderboard test set is different.</p>",
      "rawMarkdown": "Case in point: my best scoring model is one that I didn't spend any time tweaking the training process on. I just picked something reasonable and let it train. \n\nOver the weekend I improved the training process (1-cycle learning policy, AdamW optimizer, fine-tuning the earlier layers with different LRs for different groups of layers, etc). The model now learned a lot better (according to the train &amp; val losses), but its leaderboard score was worse. \n\nI don't think the model was overfitting per se (the validation loss kept going down). Even though it was learning better than before, it was just learning the wrong things.\n\nThe results is that the model that didn't learn as well, actually generalized better on the test set used for the leaderboard.\n\nSo if your model does better on the training set, but worse on the leaderboard, it's either overfitting -- or the distribution of the leaderboard test set is different.",
      "votes": null
    },
    {
      "id": "741634",
      "postDate": "02/10/2020 21:03:26",
      "content": "<p>I think you are right. In order for the network to generalise we’ll cross-domain it must be free of biases towards a particular deepfake method. That is why I think 80% of the work is in training set selection to avoid such bias.</p>",
      "rawMarkdown": "I think you are right. In order for the network to generalise we’ll cross-domain it must be free of biases towards a particular deepfake method. That is why I think 80% of the work is in training set selection to avoid such bias.",
      "votes": null
    },
    {
      "id": "741685",
      "postDate": "02/10/2020 22:04:39",
      "content": "<p>I have the same problem. I have models that performed much better on validation (and i made sure there are no cross links as much as i could between train and validation set).\nMy conclusion is that the test set is not made the same way as the training set. Probably real world videos? And the models overfit to something specific to the training set that we can't easily comprehend. Having said that, i might still have some other stupid mistakes... </p>",
      "rawMarkdown": "I have the same problem. I have models that performed much better on validation (and i made sure there are no cross links as much as i could between train and validation set).\nMy conclusion is that the test set is not made the same way as the training set. Probably real world videos? And the models overfit to something specific to the training set that we can't easily comprehend. Having said that, i might still have some other stupid mistakes...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 740245,
      "author_name": "uraler",
      "author_url": "",
      "post_date": "02/09/2020 04:37:01",
      "content": "<p>Is data distribution different in training and test sets? Isn't the basic assumption of ML that they are sampled from the same distribution?</p>",
      "votes": null,
      "replies": [
        {
          "id": 740429,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "02/09/2020 12:13:50",
          "content": "<blockquote>\n  <p>Isn't the basic assumption of ML that they are sampled from the same distribution?</p>\n</blockquote>\n\n<p>Haha, not on Kaggle. ;-) Part of building a solution is to think of ways that the competition organizer may have changed the private test set to make it harder. This why people whose models don't generalize enough will get great scores on the public leaderboard but then do much worse on the private LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 740682,
          "author_name": "olegtrott",
          "author_url": "",
          "post_date": "02/09/2020 16:49:54",
          "content": "<blockquote>\n  <p><strong>Human Analog wrote:</strong></p>\n  \n  <p>Haha, not on Kaggle. ;-) Part of building a solution is to think of ways that the competition organizer may have changed the private test set to make it harder. This why people whose models don't generalize enough will get great scores on the public leaderboard but then do much worse on the private LB.</p>\n</blockquote>\n\n<p>Generalization is good performance on novel samples from the <em>same</em> distribution.</p>\n\n<p>Which competitions changed the distribution? I think it's rare.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 741634,
          "author_name": "maralski",
          "author_url": "",
          "post_date": "02/10/2020 21:03:26",
          "content": "<p>I think you are right. In order for the network to generalise we’ll cross-domain it must be free of biases towards a particular deepfake method. That is why I think 80% of the work is in training set selection to avoid such bias.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 740557,
      "author_name": "sumitm004",
      "author_url": "",
      "post_date": "02/09/2020 14:27:03",
      "content": "<p>She looks like minecraft steve though..</p>",
      "votes": null,
      "replies": [
        {
          "id": 740705,
          "author_name": "olegtrott",
          "author_url": "",
          "post_date": "02/09/2020 17:15:53",
          "content": "<p>&gt; <strong>Sumit Mishra wrote:</strong>\n&gt; \n&gt; She looks like minecraft steve though..</p>\n\n<p>Really? Here is frame 35 from a real and forged video (face only) :</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F351627%2F65a8d30537b376a8b1a810ea3194eb59%2F1.png?generation=1581272587789257&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F351627%2F6a5e584b7cd316f4e8ddd23119e786db%2F11.png?generation=1581272602552317&amp;alt=media\" alt=\"\"></p>\n\n<p>Which one is fake?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 740751,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "02/09/2020 18:31:07",
          "content": "<blockquote>\n  <p>Which one is fake?</p>\n</blockquote>\n\n<p>Just because it's hard to see for a human doesn't mean it's hard for a computer. The classifier may detect patterns in the fake image that are not present in the real image (or vice versa). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 740752,
          "author_name": "olegtrott",
          "author_url": "",
          "post_date": "02/09/2020 18:32:02",
          "content": "<p>The same frame from a real and forged video:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F351627%2F97960a3b7b03650adf84446e538e44a1%2F3.png?generation=1581273034862445&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F351627%2F8d83c25be7f4e790f40aa39c010323b6%2F33.png?generation=1581273059942158&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 740758,
          "author_name": "olegtrott",
          "author_url": "",
          "post_date": "02/09/2020 18:46:14",
          "content": "<p>&gt; <strong>Human Analog wrote:</strong>\n&gt; \n&gt; Just because it's hard to see for a human doesn't mean it's hard for a computer.</p>\n\n<p>My point is that some forgery methods will be much harder to detect than others. The probabilities you predict must depend on the fractions of these methods in the test set. If the organizers changed these fractions, you are forced to <strong>guess</strong> what they did or didn't do (See the original post, which outlines a somewhat extreme example of when your algorithm is completely useless against some method, NeuralTextures. But similar considerations apply when it's not useless, but just performs worse than against other methods).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 740795,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "02/09/2020 19:55:59",
          "content": "<blockquote>\n  <p>you are forced to guess what they did or didn't do</p>\n</blockquote>\n\n<p>This has been the case for all the competitions I've been in (which isn't that many so perhaps I just picked the ones where this happened).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 741207,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "02/10/2020 10:15:58",
          "content": "<p>Case in point: my best scoring model is one that I didn't spend any time tweaking the training process on. I just picked something reasonable and let it train. </p>\n\n<p>Over the weekend I improved the training process (1-cycle learning policy, AdamW optimizer, fine-tuning the earlier layers with different LRs for different groups of layers, etc). The model now learned a lot better (according to the train &amp; val losses), but its leaderboard score was worse. </p>\n\n<p>I don't think the model was overfitting per se (the validation loss kept going down). Even though it was learning better than before, it was just learning the wrong things.</p>\n\n<p>The results is that the model that didn't learn as well, actually generalized better on the test set used for the leaderboard.</p>\n\n<p>So if your model does better on the training set, but worse on the leaderboard, it's either overfitting -- or the distribution of the leaderboard test set is different.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 741685,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "02/10/2020 22:04:39",
          "content": "<p>I have the same problem. I have models that performed much better on validation (and i made sure there are no cross links as much as i could between train and validation set).\nMy conclusion is that the test set is not made the same way as the training set. Probably real world videos? And the models overfit to something specific to the training set that we can't easily comprehend. Having said that, i might still have some other stupid mistakes... </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "740046": "NeuralTextures look almost perceptually authentic to a human, let alone a machine.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F351627%2Fb1bf131960c78613401f55740349248c%2Ffake.png?generation=1581187572634792&amp;alt=media)\n\n\n\nSuppose the test set contains `fraction_real` and `fraction_neuraltextures`, and your algorithm can't tell the difference[1].\n\nThen when the algorithm sees something authentic-looking, and it doesn't predict\n\n    fraction_neuraltextures / (fraction_neuraltextures + fraction_real)\n\nit will lose to similar algorithms that do.\n\nSince we don't know these numbers[2], a large part of this contest seems to be guessing them, making this contest something of a lottery.\n\nThis must demotivate serious competitors: you could expend a lot of effort to submit a sound entry, but your odds of winning are lottery-like (Even less, if another competitor learned `fraction_real` and `fraction_neuraltextures` through proximity to the organizers, for example, or by asking them privately)\n\nTo my mind, stating these two numbers publicly would make this contest infinitely better.\n\n---\n\n[1] Even if your algorithm is not entirely useless on NeuralTextures, I hope it's easy to see that knowing these numbers is critically important to what predictions it should make.\n\n[2] We know that `fraction_real` is 0.5 on the public test, but not on the private test set.",
    "740245": "Is data distribution different in training and test sets? Isn't the basic assumption of ML that they are sampled from the same distribution?",
    "740429": "&gt; Isn't the basic assumption of ML that they are sampled from the same distribution?\n\nHaha, not on Kaggle. ;-) Part of building a solution is to think of ways that the competition organizer may have changed the private test set to make it harder. This why people whose models don't generalize enough will get great scores on the public leaderboard but then do much worse on the private LB.",
    "740557": "She looks like minecraft steve though..",
    "740682": "&gt; **Human Analog wrote:**\n&gt; \n&gt; Haha, not on Kaggle. ;-) Part of building a solution is to think of ways that the competition organizer may have changed the private test set to make it harder. This why people whose models don't generalize enough will get great scores on the public leaderboard but then do much worse on the private LB.\n\nGeneralization is good performance on novel samples from the *same* distribution.\n\nWhich competitions changed the distribution? I think it's rare.",
    "740705": "&gt; **Sumit Mishra wrote:**\n&gt; \n&gt; She looks like minecraft steve though..\n\nReally? Here is frame 35 from a real and forged video (face only) :\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F351627%2F65a8d30537b376a8b1a810ea3194eb59%2F1.png?generation=1581272587789257&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F351627%2F6a5e584b7cd316f4e8ddd23119e786db%2F11.png?generation=1581272602552317&amp;alt=media)\n\n\n\nWhich one is fake?",
    "740751": "&gt; Which one is fake?\n\nJust because it's hard to see for a human doesn't mean it's hard for a computer. The classifier may detect patterns in the fake image that are not present in the real image (or vice versa).",
    "740752": "The same frame from a real and forged video:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F351627%2F97960a3b7b03650adf84446e538e44a1%2F3.png?generation=1581273034862445&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F351627%2F8d83c25be7f4e790f40aa39c010323b6%2F33.png?generation=1581273059942158&amp;alt=media)",
    "740758": "&gt; **Human Analog wrote:**\n&gt; \n&gt; Just because it's hard to see for a human doesn't mean it's hard for a computer.\n\nMy point is that some forgery methods will be much harder to detect than others. The probabilities you predict must depend on the fractions of these methods in the test set. If the organizers changed these fractions, you are forced to **guess** what they did or didn't do (See the original post, which outlines a somewhat extreme example of when your algorithm is completely useless against some method, NeuralTextures. But similar considerations apply when it's not useless, but just performs worse than against other methods).",
    "740795": "&gt; you are forced to guess what they did or didn't do\n\nThis has been the case for all the competitions I've been in (which isn't that many so perhaps I just picked the ones where this happened).",
    "741207": "Case in point: my best scoring model is one that I didn't spend any time tweaking the training process on. I just picked something reasonable and let it train. \n\nOver the weekend I improved the training process (1-cycle learning policy, AdamW optimizer, fine-tuning the earlier layers with different LRs for different groups of layers, etc). The model now learned a lot better (according to the train &amp; val losses), but its leaderboard score was worse. \n\nI don't think the model was overfitting per se (the validation loss kept going down). Even though it was learning better than before, it was just learning the wrong things.\n\nThe results is that the model that didn't learn as well, actually generalized better on the test set used for the leaderboard.\n\nSo if your model does better on the training set, but worse on the leaderboard, it's either overfitting -- or the distribution of the leaderboard test set is different.",
    "741634": "I think you are right. In order for the network to generalise we’ll cross-domain it must be free of biases towards a particular deepfake method. That is why I think 80% of the work is in training set selection to avoid such bias.",
    "741685": "I have the same problem. I have models that performed much better on validation (and i made sure there are no cross links as much as i could between train and validation set).\nMy conclusion is that the test set is not made the same way as the training set. Probably real world videos? And the models overfit to something specific to the training set that we can't easily comprehend. Having said that, i might still have some other stupid mistakes..."
  },
  "source": "meta"
}