{
  "id": 122359,
  "title": "Good validation set?",
  "url": "/competitions/deepfake-detection-challenge/discussion/122359",
  "author_name": "",
  "post_date": "2019-12-19T16:10:04.664630Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I trained a CNN based model on the  training data which is 90% split from the main dataset and used other 10% split as the validation model. My model achieved to a very good performance on the validation set but the submission log score is very high. I have taken care of the skew and my model is not biased for one class, at least as seen from train and validation splits. I want to know if the problem is:\n(1) Overfitting as the validation and train split set are very identical OR \n(2) The test set used by Kaggle is very different from the dataset provided for training(in terms of structure as in generated videos using different methods or augmentations)? \nI am new to field, can someone help me figure out how can I improve my model from here?</p>",
  "messages": [
    {
      "id": "698715",
      "postDate": "12/19/2019 16:10:04",
      "content": "<p>I trained a CNN based model on the  training data which is 90% split from the main dataset and used other 10% split as the validation model. My model achieved to a very good performance on the validation set but the submission log score is very high. I have taken care of the skew and my model is not biased for one class, at least as seen from train and validation splits. I want to know if the problem is:\n(1) Overfitting as the validation and train split set are very identical OR \n(2) The test set used by Kaggle is very different from the dataset provided for training(in terms of structure as in generated videos using different methods or augmentations)? \nI am new to field, can someone help me figure out how can I improve my model from here?</p>",
      "rawMarkdown": "I trained a CNN based model on the  training data which is 90% split from the main dataset and used other 10% split as the validation model. My model achieved to a very good performance on the validation set but the submission log score is very high. I have taken care of the skew and my model is not biased for one class, at least as seen from train and validation splits. I want to know if the problem is:\n(1) Overfitting as the validation and train split set are very identical OR \n(2) The test set used by Kaggle is very different from the dataset provided for training(in terms of structure as in generated videos using different methods or augmentations)? \nI am new to field, can someone help me figure out how can I improve my model from here?",
      "votes": null
    },
    {
      "id": "698855",
      "postDate": "12/19/2019 19:24:39",
      "content": "<p>Make sure that the videos in your validation set are not also in the training set. For example, if real video A was used to create fake videos B, C, and D, then you shouldn't have B and C in the validation set and A and D in the training set. Either all of them should go in the training set, or all in the validation set.</p>\n\n<p>I think each of the 50 subdirectories with training videos is self-contained in this way, so you might try using subdir 0 as the validation set and 1-49 to train on. (Although I'm not sure how good that particular split is.)</p>",
      "rawMarkdown": "Make sure that the videos in your validation set are not also in the training set. For example, if real video A was used to create fake videos B, C, and D, then you shouldn't have B and C in the validation set and A and D in the training set. Either all of them should go in the training set, or all in the validation set.\n\nI think each of the 50 subdirectories with training videos is self-contained in this way, so you might try using subdir 0 as the validation set and 1-49 to train on. (Although I'm not sure how good that particular split is.)",
      "votes": null
    },
    {
      "id": "698858",
      "postDate": "12/19/2019 19:29:51",
      "content": "<p>Train -&gt; validation leak if you didn''t split by original video or even actor?</p>",
      "rawMarkdown": "Train -&gt; validation leak if you didn''t split by original video or even actor?",
      "votes": null
    },
    {
      "id": "698883",
      "postDate": "12/19/2019 20:21:34",
      "content": "<blockquote>\n  <p>The test set used by Kaggle is very different from the dataset provided for training</p>\n</blockquote>\n\n<p>\"Public Test Set\" has these augmentations:</p>\n\n<p>(1) reduce the FPS of the video to 15\n(2) reduce the resolution of the video to 1/4 of its original size\n(3) reduce the overall encoding quality</p>\n\n<p>number 3 is a big problem when a video is real and your model think it's fake because it has a lower quality</p>",
      "rawMarkdown": "&gt; The test set used by Kaggle is very different from the dataset provided for training\n\n\"Public Test Set\" has these augmentations:\n\n(1) reduce the FPS of the video to 15\n(2) reduce the resolution of the video to 1/4 of its original size\n(3) reduce the overall encoding quality\n\nnumber 3 is a big problem when a video is real and your model think it's fake because it has a lower quality",
      "votes": null
    },
    {
      "id": "698926",
      "postDate": "12/19/2019 22:34:27",
      "content": "<p>Where did you get the info from?\nAlso beware of using the first dirs (big faces) or the last dirs (only audio modification) when validating. </p>",
      "rawMarkdown": "Where did you get the info from?\nAlso beware of using the first dirs (big faces) or the last dirs (only audio modification) when validating.",
      "votes": null
    },
    {
      "id": "699091",
      "postDate": "12/20/2019 03:30:30",
      "content": "<p>Look at this paper <a href=\"https://arxiv.org/abs/1910.08854\">https://arxiv.org/abs/1910.08854</a></p>",
      "rawMarkdown": "Look at this paper https://arxiv.org/abs/1910.08854",
      "votes": null
    },
    {
      "id": "699368",
      "postDate": "12/20/2019 11:06:34",
      "content": "<p>Thank you. This was helpful. </p>",
      "rawMarkdown": "Thank you. This was helpful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 698855,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "12/19/2019 19:24:39",
      "content": "<p>Make sure that the videos in your validation set are not also in the training set. For example, if real video A was used to create fake videos B, C, and D, then you shouldn't have B and C in the validation set and A and D in the training set. Either all of them should go in the training set, or all in the validation set.</p>\n\n<p>I think each of the 50 subdirectories with training videos is self-contained in this way, so you might try using subdir 0 as the validation set and 1-49 to train on. (Although I'm not sure how good that particular split is.)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 698858,
      "author_name": "bacterio",
      "author_url": "",
      "post_date": "12/19/2019 19:29:51",
      "content": "<p>Train -&gt; validation leak if you didn''t split by original video or even actor?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 698883,
      "author_name": "behrouzp",
      "author_url": "",
      "post_date": "12/19/2019 20:21:34",
      "content": "<blockquote>\n  <p>The test set used by Kaggle is very different from the dataset provided for training</p>\n</blockquote>\n\n<p>\"Public Test Set\" has these augmentations:</p>\n\n<p>(1) reduce the FPS of the video to 15\n(2) reduce the resolution of the video to 1/4 of its original size\n(3) reduce the overall encoding quality</p>\n\n<p>number 3 is a big problem when a video is real and your model think it's fake because it has a lower quality</p>",
      "votes": null,
      "replies": [
        {
          "id": 698926,
          "author_name": "simoninparis",
          "author_url": "",
          "post_date": "12/19/2019 22:34:27",
          "content": "<p>Where did you get the info from?\nAlso beware of using the first dirs (big faces) or the last dirs (only audio modification) when validating. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 699091,
          "author_name": "behrouzp",
          "author_url": "",
          "post_date": "12/20/2019 03:30:30",
          "content": "<p>Look at this paper <a href=\"https://arxiv.org/abs/1910.08854\">https://arxiv.org/abs/1910.08854</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 699368,
          "author_name": "ameypatil",
          "author_url": "",
          "post_date": "12/20/2019 11:06:34",
          "content": "<p>Thank you. This was helpful. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "698715": "I trained a CNN based model on the  training data which is 90% split from the main dataset and used other 10% split as the validation model. My model achieved to a very good performance on the validation set but the submission log score is very high. I have taken care of the skew and my model is not biased for one class, at least as seen from train and validation splits. I want to know if the problem is:\n(1) Overfitting as the validation and train split set are very identical OR \n(2) The test set used by Kaggle is very different from the dataset provided for training(in terms of structure as in generated videos using different methods or augmentations)? \nI am new to field, can someone help me figure out how can I improve my model from here?",
    "698855": "Make sure that the videos in your validation set are not also in the training set. For example, if real video A was used to create fake videos B, C, and D, then you shouldn't have B and C in the validation set and A and D in the training set. Either all of them should go in the training set, or all in the validation set.\n\nI think each of the 50 subdirectories with training videos is self-contained in this way, so you might try using subdir 0 as the validation set and 1-49 to train on. (Although I'm not sure how good that particular split is.)",
    "698858": "Train -&gt; validation leak if you didn''t split by original video or even actor?",
    "698883": "&gt; The test set used by Kaggle is very different from the dataset provided for training\n\n\"Public Test Set\" has these augmentations:\n\n(1) reduce the FPS of the video to 15\n(2) reduce the resolution of the video to 1/4 of its original size\n(3) reduce the overall encoding quality\n\nnumber 3 is a big problem when a video is real and your model think it's fake because it has a lower quality",
    "698926": "Where did you get the info from?\nAlso beware of using the first dirs (big faces) or the last dirs (only audio modification) when validating.",
    "699091": "Look at this paper https://arxiv.org/abs/1910.08854",
    "699368": "Thank you. This was helpful."
  },
  "source": "meta"
}