{
  "id": 134895,
  "title": "Is this what overfitting looks like?",
  "url": "/competitions/deepfake-detection-challenge/discussion/134895",
  "author_name": "",
  "post_date": "2020-03-11T02:03:51.385853800Z",
  "votes": 1,
  "comment_count": 16,
  "views": 0,
  "content": "<p>I trained a Resnet50 model starting with vggface2 weights using ~40,000 videos, extracting faces from 30 frames per video with facenet_pytorch. I created a validation set of ~1,000 videos such that there were no overlapping original videos or fake versions thereof. I created separate datasets for the real and fake faces and sampled them evenly while training. I augmented the real videos with a random horizontal flip, thinking that there were enough fake videos without augmentation.</p>\n\n<p>I was able to get the validation loss down to ~0.25. I got a similar loss when I calculated it for the 400 test videos. However, when I submitted it, I ended up with a loss of 0.97 against the 4,000 public test set. In reading through some of the discussion threads, others noted differences in validation and test scores, but not nearly as large.</p>\n\n<p>Here is a histogram of my predictions on the 400 test videos.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2F5a8d702117d43d54f78454ba57438568%2F2020-03-10_18-41-49.png?generation=1583891005206513&amp;alt=media\" alt=\"\"></p>\n\n<p>For comparison sake, here is a histogram of the predictions from the <a href=\"https://www.kaggle.com/greatgamedota/xception-binary-classifier-inference\">Xception Binary Classifier - Inference notebook</a>, which had a leader board score of 0.55.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2F345fb8076ed27611b2efc6dfd38b45ea%2F2020-03-10_18-49-04.png?generation=1583891377484329&amp;alt=media\" alt=\"\"></p>\n\n<p>The predictions are obviously much more distributed throughout the range. Is the polarization in the predictions from my model indicative of over-training or have I gone wrong someplace else?</p>",
  "messages": [
    {
      "id": "768582",
      "postDate": "03/11/2020 02:03:51",
      "content": "<p>I trained a Resnet50 model starting with vggface2 weights using ~40,000 videos, extracting faces from 30 frames per video with facenet_pytorch. I created a validation set of ~1,000 videos such that there were no overlapping original videos or fake versions thereof. I created separate datasets for the real and fake faces and sampled them evenly while training. I augmented the real videos with a random horizontal flip, thinking that there were enough fake videos without augmentation.</p>\n\n<p>I was able to get the validation loss down to ~0.25. I got a similar loss when I calculated it for the 400 test videos. However, when I submitted it, I ended up with a loss of 0.97 against the 4,000 public test set. In reading through some of the discussion threads, others noted differences in validation and test scores, but not nearly as large.</p>\n\n<p>Here is a histogram of my predictions on the 400 test videos.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2F5a8d702117d43d54f78454ba57438568%2F2020-03-10_18-41-49.png?generation=1583891005206513&amp;alt=media\" alt=\"\"></p>\n\n<p>For comparison sake, here is a histogram of the predictions from the <a href=\"https://www.kaggle.com/greatgamedota/xception-binary-classifier-inference\">Xception Binary Classifier - Inference notebook</a>, which had a leader board score of 0.55.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2F345fb8076ed27611b2efc6dfd38b45ea%2F2020-03-10_18-49-04.png?generation=1583891377484329&amp;alt=media\" alt=\"\"></p>\n\n<p>The predictions are obviously much more distributed throughout the range. Is the polarization in the predictions from my model indicative of over-training or have I gone wrong someplace else?</p>",
      "rawMarkdown": "I trained a Resnet50 model starting with vggface2 weights using ~40,000 videos, extracting faces from 30 frames per video with facenet_pytorch. I created a validation set of ~1,000 videos such that there were no overlapping original videos or fake versions thereof. I created separate datasets for the real and fake faces and sampled them evenly while training. I augmented the real videos with a random horizontal flip, thinking that there were enough fake videos without augmentation.\n\nI was able to get the validation loss down to ~0.25. I got a similar loss when I calculated it for the 400 test videos. However, when I submitted it, I ended up with a loss of 0.97 against the 4,000 public test set. In reading through some of the discussion threads, others noted differences in validation and test scores, but not nearly as large.\n\nHere is a histogram of my predictions on the 400 test videos.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2F5a8d702117d43d54f78454ba57438568%2F2020-03-10_18-41-49.png?generation=1583891005206513&amp;alt=media)\n\nFor comparison sake, here is a histogram of the predictions from the [Xception Binary Classifier - Inference notebook](https://www.kaggle.com/greatgamedota/xception-binary-classifier-inference), which had a leader board score of 0.55.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2F345fb8076ed27611b2efc6dfd38b45ea%2F2020-03-10_18-49-04.png?generation=1583891377484329&amp;alt=media)\n\nThe predictions are obviously much more distributed throughout the range. Is the polarization in the predictions from my model indicative of over-training or have I gone wrong someplace else?",
      "votes": null
    },
    {
      "id": "768643",
      "postDate": "03/11/2020 03:44:44",
      "content": "<p>Not sure over-fitting versus log loss not being intuitive versus accuracy. You can try to clip your prediction e.g. max 0.05 or .95, or take square root of multiple predictions and average them.  Log loss can highly penalize confident but wrong predictions, e.g. definitely don't want 0 and 1s.  ~.69 is for guessing all .5s.</p>",
      "rawMarkdown": "Not sure over-fitting versus log loss not being intuitive versus accuracy. You can try to clip your prediction e.g. max 0.05 or .95, or take square root of multiple predictions and average them.  Log loss can highly penalize confident but wrong predictions, e.g. definitely don't want 0 and 1s.  ~.69 is for guessing all .5s.",
      "votes": null
    },
    {
      "id": "768645",
      "postDate": "03/11/2020 03:48:29",
      "content": "<p>The first one looks more\"correct\" to me.\nIs it possible that you are switching the labels in your model?</p>",
      "rawMarkdown": "The first one looks more\"correct\" to me.\nIs it possible that you are switching the labels in your model?",
      "votes": null
    },
    {
      "id": "768663",
      "postDate": "03/11/2020 04:26:54",
      "content": "<p>There is no label for 400 test videos, how did you calculate the loss? </p>",
      "rawMarkdown": "There is no label for 400 test videos, how did you calculate the loss?",
      "votes": null
    },
    {
      "id": "768665",
      "postDate": "03/11/2020 04:28:30",
      "content": "<p>Thanks - I'll go back through and check to make sure. One thing is that my model predicts real, i.e. 1 = real and 0 = fake, just as a legacy of the way I initially put the data frame together that I then used to write records to disk. After the final sigmoid activation, I was just subtracting it from 1 to get the complementary probability. I didn't think that was messing up the results, especially since my local log loss and confusion matrix were looking decent, but I'll go back and check and in any event make sure I'm not making an error in conforming the predictions.</p>\n\n<p>I've struggled with getting my prediction routine to actually produce a score, so it's also possible something is still amiss there. The difference between 0.25 locally and .97 on the leader board struck me as possible the prediction routine on submit isn't working right.</p>",
      "rawMarkdown": "Thanks - I'll go back through and check to make sure. One thing is that my model predicts real, i.e. 1 = real and 0 = fake, just as a legacy of the way I initially put the data frame together that I then used to write records to disk. After the final sigmoid activation, I was just subtracting it from 1 to get the complementary probability. I didn't think that was messing up the results, especially since my local log loss and confusion matrix were looking decent, but I'll go back and check and in any event make sure I'm not making an error in conforming the predictions.\n\nI've struggled with getting my prediction routine to actually produce a score, so it's also possible something is still amiss there. The difference between 0.25 locally and .97 on the leader board struck me as possible the prediction routine on submit isn't working right.",
      "votes": null
    },
    {
      "id": "768669",
      "postDate": "03/11/2020 04:37:39",
      "content": "<p>Thanks - I was messing around with the predictions and noted as somebody else did that the 0.69 with all 0.5s means that the public test set is perfectly balanced. That may not be the case for the leader board. My model, for whatever reason does a better on the fakes than it does on the real ones. My local 0.25 on the 50/50 test videos is basically getting all of the fakes right but missing about a quarter of the real. So that would be another possibility, that the leader board test set is skewed towards real videos.</p>\n\n<p><strong>EDIT</strong>: This isn't right - I think the leaderboard test set likely is balanced 50/50 because the submit score comes back at 0.69 with 0.5 for all videos.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2Fb54e61aaf5ccd3be886b6e09489767f1%2F2020-03-10_21-52-21.png?generation=1583902389166759&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thanks - I was messing around with the predictions and noted as somebody else did that the 0.69 with all 0.5s means that the public test set is perfectly balanced. That may not be the case for the leader board. My model, for whatever reason does a better on the fakes than it does on the real ones. My local 0.25 on the 50/50 test videos is basically getting all of the fakes right but missing about a quarter of the real. So that would be another possibility, that the leader board test set is skewed towards real videos.\n\n**EDIT**: This isn't right - I think the leaderboard test set likely is balanced 50/50 because the submit score comes back at 0.69 with 0.5 for all videos.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2Fb54e61aaf5ccd3be886b6e09489767f1%2F2020-03-10_21-52-21.png?generation=1583902389166759&amp;alt=media)",
      "votes": null
    },
    {
      "id": "768672",
      "postDate": "03/11/2020 04:39:41",
      "content": "<p>They are a subset of the training videos. I can put up the data frame I constructed from the meta data json files if you'd like it. Then you can just look up the label by the video name.</p>",
      "rawMarkdown": "They are a subset of the training videos. I can put up the data frame I constructed from the meta data json files if you'd like it. Then you can just look up the label by the video name.",
      "votes": null
    },
    {
      "id": "768676",
      "postDate": "03/11/2020 04:42:53",
      "content": "<p><a href=\"https://www.kaggle.com/calebeverett/metadata-dataframe\">https://www.kaggle.com/calebeverett/metadata-dataframe</a></p>",
      "rawMarkdown": "https://www.kaggle.com/calebeverett/metadata-dataframe",
      "votes": null
    },
    {
      "id": "768707",
      "postDate": "03/11/2020 05:22:30",
      "content": "<p>It appears your model may be too confident. If you use 20 bins it would be more evident. Notice not much going on in the middle. You also have severe bias to predicting fake. Here is mine for comparison.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F389345%2F18e71790830104f5032e05eb08f34b09%2Fmodel.png?generation=1583903823209986&amp;alt=media\" alt=\"\"></p>\n\n<p>I suggest just use 1 frame and 1 face per video to start with and under sample the fake videos.</p>",
      "rawMarkdown": "It appears your model may be too confident. If you use 20 bins it would be more evident. Notice not much going on in the middle. You also have severe bias to predicting fake. Here is mine for comparison.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F389345%2F18e71790830104f5032e05eb08f34b09%2Fmodel.png?generation=1583903823209986&amp;alt=media)\n\nI suggest just use 1 frame and 1 face per video to start with and under sample the fake videos.",
      "votes": null
    },
    {
      "id": "768749",
      "postDate": "03/11/2020 06:25:59",
      "content": "<p>I'll give that a try. Just out of curiosity, if you don't mind sharing what is the difference between your lb loss and loss on the 400?</p>\n\n<p>I noticed the bias and couldn't think of a reason why that would be the case. Every model I've trained has ended up with the same bias, even though I'm balancing the training data. Maybe it has something to do with the starting weights.</p>",
      "rawMarkdown": "I'll give that a try. Just out of curiosity, if you don't mind sharing what is the difference between your lb loss and loss on the 400?\n\nI noticed the bias and couldn't think of a reason why that would be the case. Every model I've trained has ended up with the same bias, even though I'm balancing the training data. Maybe it has something to do with the starting weights.",
      "votes": null
    },
    {
      "id": "768762",
      "postDate": "03/11/2020 06:44:41",
      "content": "<p>About 0.14 on 4000 balanced samples in validation set.</p>\n\n<p>You are balancing by class but you might also be repeating the same images to the extent that there are naturally occurring clusters of similar images that can bias the model towards the dominant clusters. This is more likely to happen if you sample multiple frames per video. To reduce the likelihood sample one frame per real video and one per fake video. Then under sample the fake videos to match the real. From there you can look at more complicated approaches.</p>",
      "rawMarkdown": "About 0.14 on 4000 balanced samples in validation set.\n\nYou are balancing by class but you might also be repeating the same images to the extent that there are naturally occurring clusters of similar images that can bias the model towards the dominant clusters. This is more likely to happen if you sample multiple frames per video. To reduce the likelihood sample one frame per real video and one per fake video. Then under sample the fake videos to match the real. From there you can look at more complicated approaches.",
      "votes": null
    },
    {
      "id": "768766",
      "postDate": "03/11/2020 06:50:52",
      "content": "<p><a href=\"/maralski\">@maralski</a> I also faced that problem (skewed predictions toward 1). But according to your reason, my predictions should skew towards 0 instead, because more real samples are used (I currently used, say, 5 frames for 1 specific real video if that real video has 5 fake versions, and each fake has 1 frame). How do you explain this (predictions still skew towards 1)?</p>",
      "rawMarkdown": "maralski I also faced that problem (skewed predictions toward 1). But according to your reason, my predictions should skew towards 0 instead, because more real samples are used (I currently used, say, 5 frames for 1 specific real video if that real video has 5 fake versions, and each fake has 1 frame). How do you explain this (predictions still skew towards 1)?",
      "votes": null
    },
    {
      "id": "768778",
      "postDate": "03/11/2020 07:08:10",
      "content": "<p><a href=\"/khahuras\">@khahuras</a> What does the distribution look like and how are you choosing frames?</p>",
      "rawMarkdown": "khahuras What does the distribution look like and how are you choosing frames?",
      "votes": null
    },
    {
      "id": "772754",
      "postDate": "03/15/2020 21:38:36",
      "content": "<p>NO. THIS IS NOT AS BAD AS YOU THINK. Too confident is a good thing for log loss. Our best subs looks just like that. Notice: If overfitted, it should be all 1.0s and 0.0s.</p>",
      "rawMarkdown": "NO. THIS IS NOT AS BAD AS YOU THINK. Too confident is a good thing for log loss. Our best subs looks just like that. Notice: If overfitted, it should be all 1.0s and 0.0s.",
      "votes": null
    },
    {
      "id": "773812",
      "postDate": "03/15/2020 23:43:54",
      "content": "<p>Just for reference, here is our 0.30003LB histogram (20 bins) for 400 test videos. I usually check it every time and gives me a rough idea about how the model will perform after submission.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2279276%2Fc242d12df997fe52374606836fde25cf%2FScreen%20Shot%202020-03-15%20at%207.42.28%20PM.png?generation=1584315794339994&amp;alt=media\" alt=\"\"></p>\n\n<p>P.S. To match your 10 bin histogram:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2279276%2F3016dbbb523dc3e916ec720b304f76da%2FScreen%20Shot%202020-03-15%20at%207.46.34%20PM.png?generation=1584316041881590&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Just for reference, here is our 0.30003LB histogram (20 bins) for 400 test videos. I usually check it every time and gives me a rough idea about how the model will perform after submission.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2279276%2Fc242d12df997fe52374606836fde25cf%2FScreen%20Shot%202020-03-15%20at%207.42.28%20PM.png?generation=1584315794339994&amp;alt=media)\n\nP.S. To match your 10 bin histogram:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2279276%2F3016dbbb523dc3e916ec720b304f76da%2FScreen%20Shot%202020-03-15%20at%207.46.34%20PM.png?generation=1584316041881590&amp;alt=media)",
      "votes": null
    },
    {
      "id": "774838",
      "postDate": "03/16/2020 01:31:02",
      "content": "<p>So you're not clipping then?</p>",
      "rawMarkdown": "So you're not clipping then?",
      "votes": null
    },
    {
      "id": "774843",
      "postDate": "03/16/2020 01:49:35",
      "content": "<p>I am clipping, but with a very small value(0.01,0.09).</p>",
      "rawMarkdown": "I am clipping, but with a very small value(0.01,0.09).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 768643,
      "author_name": "botgreet",
      "author_url": "",
      "post_date": "03/11/2020 03:44:44",
      "content": "<p>Not sure over-fitting versus log loss not being intuitive versus accuracy. You can try to clip your prediction e.g. max 0.05 or .95, or take square root of multiple predictions and average them.  Log loss can highly penalize confident but wrong predictions, e.g. definitely don't want 0 and 1s.  ~.69 is for guessing all .5s.</p>",
      "votes": null,
      "replies": [
        {
          "id": 768669,
          "author_name": "calebeverett",
          "author_url": "",
          "post_date": "03/11/2020 04:37:39",
          "content": "<p>Thanks - I was messing around with the predictions and noted as somebody else did that the 0.69 with all 0.5s means that the public test set is perfectly balanced. That may not be the case for the leader board. My model, for whatever reason does a better on the fakes than it does on the real ones. My local 0.25 on the 50/50 test videos is basically getting all of the fakes right but missing about a quarter of the real. So that would be another possibility, that the leader board test set is skewed towards real videos.</p>\n\n<p><strong>EDIT</strong>: This isn't right - I think the leaderboard test set likely is balanced 50/50 because the submit score comes back at 0.69 with 0.5 for all videos.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2Fb54e61aaf5ccd3be886b6e09489767f1%2F2020-03-10_21-52-21.png?generation=1583902389166759&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 768645,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "03/11/2020 03:48:29",
      "content": "<p>The first one looks more\"correct\" to me.\nIs it possible that you are switching the labels in your model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 768665,
          "author_name": "calebeverett",
          "author_url": "",
          "post_date": "03/11/2020 04:28:30",
          "content": "<p>Thanks - I'll go back through and check to make sure. One thing is that my model predicts real, i.e. 1 = real and 0 = fake, just as a legacy of the way I initially put the data frame together that I then used to write records to disk. After the final sigmoid activation, I was just subtracting it from 1 to get the complementary probability. I didn't think that was messing up the results, especially since my local log loss and confusion matrix were looking decent, but I'll go back and check and in any event make sure I'm not making an error in conforming the predictions.</p>\n\n<p>I've struggled with getting my prediction routine to actually produce a score, so it's also possible something is still amiss there. The difference between 0.25 locally and .97 on the leader board struck me as possible the prediction routine on submit isn't working right.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 768663,
      "author_name": "ant1ss",
      "author_url": "",
      "post_date": "03/11/2020 04:26:54",
      "content": "<p>There is no label for 400 test videos, how did you calculate the loss? </p>",
      "votes": null,
      "replies": [
        {
          "id": 768672,
          "author_name": "calebeverett",
          "author_url": "",
          "post_date": "03/11/2020 04:39:41",
          "content": "<p>They are a subset of the training videos. I can put up the data frame I constructed from the meta data json files if you'd like it. Then you can just look up the label by the video name.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 768676,
          "author_name": "calebeverett",
          "author_url": "",
          "post_date": "03/11/2020 04:42:53",
          "content": "<p><a href=\"https://www.kaggle.com/calebeverett/metadata-dataframe\">https://www.kaggle.com/calebeverett/metadata-dataframe</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 768707,
      "author_name": "maralski",
      "author_url": "",
      "post_date": "03/11/2020 05:22:30",
      "content": "<p>It appears your model may be too confident. If you use 20 bins it would be more evident. Notice not much going on in the middle. You also have severe bias to predicting fake. Here is mine for comparison.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F389345%2F18e71790830104f5032e05eb08f34b09%2Fmodel.png?generation=1583903823209986&amp;alt=media\" alt=\"\"></p>\n\n<p>I suggest just use 1 frame and 1 face per video to start with and under sample the fake videos.</p>",
      "votes": null,
      "replies": [
        {
          "id": 768749,
          "author_name": "calebeverett",
          "author_url": "",
          "post_date": "03/11/2020 06:25:59",
          "content": "<p>I'll give that a try. Just out of curiosity, if you don't mind sharing what is the difference between your lb loss and loss on the 400?</p>\n\n<p>I noticed the bias and couldn't think of a reason why that would be the case. Every model I've trained has ended up with the same bias, even though I'm balancing the training data. Maybe it has something to do with the starting weights.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 768762,
          "author_name": "maralski",
          "author_url": "",
          "post_date": "03/11/2020 06:44:41",
          "content": "<p>About 0.14 on 4000 balanced samples in validation set.</p>\n\n<p>You are balancing by class but you might also be repeating the same images to the extent that there are naturally occurring clusters of similar images that can bias the model towards the dominant clusters. This is more likely to happen if you sample multiple frames per video. To reduce the likelihood sample one frame per real video and one per fake video. Then under sample the fake videos to match the real. From there you can look at more complicated approaches.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 768766,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "03/11/2020 06:50:52",
          "content": "<p><a href=\"/maralski\">@maralski</a> I also faced that problem (skewed predictions toward 1). But according to your reason, my predictions should skew towards 0 instead, because more real samples are used (I currently used, say, 5 frames for 1 specific real video if that real video has 5 fake versions, and each fake has 1 frame). How do you explain this (predictions still skew towards 1)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 768778,
          "author_name": "maralski",
          "author_url": "",
          "post_date": "03/11/2020 07:08:10",
          "content": "<p><a href=\"/khahuras\">@khahuras</a> What does the distribution look like and how are you choosing frames?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 772754,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "03/15/2020 21:38:36",
      "content": "<p>NO. THIS IS NOT AS BAD AS YOU THINK. Too confident is a good thing for log loss. Our best subs looks just like that. Notice: If overfitted, it should be all 1.0s and 0.0s.</p>",
      "votes": null,
      "replies": [
        {
          "id": 774838,
          "author_name": "maralski",
          "author_url": "",
          "post_date": "03/16/2020 01:31:02",
          "content": "<p>So you're not clipping then?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 774843,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "03/16/2020 01:49:35",
          "content": "<p>I am clipping, but with a very small value(0.01,0.09).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 773812,
      "author_name": "debanga",
      "author_url": "",
      "post_date": "03/15/2020 23:43:54",
      "content": "<p>Just for reference, here is our 0.30003LB histogram (20 bins) for 400 test videos. I usually check it every time and gives me a rough idea about how the model will perform after submission.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2279276%2Fc242d12df997fe52374606836fde25cf%2FScreen%20Shot%202020-03-15%20at%207.42.28%20PM.png?generation=1584315794339994&amp;alt=media\" alt=\"\"></p>\n\n<p>P.S. To match your 10 bin histogram:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2279276%2F3016dbbb523dc3e916ec720b304f76da%2FScreen%20Shot%202020-03-15%20at%207.46.34%20PM.png?generation=1584316041881590&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "768582": "I trained a Resnet50 model starting with vggface2 weights using ~40,000 videos, extracting faces from 30 frames per video with facenet_pytorch. I created a validation set of ~1,000 videos such that there were no overlapping original videos or fake versions thereof. I created separate datasets for the real and fake faces and sampled them evenly while training. I augmented the real videos with a random horizontal flip, thinking that there were enough fake videos without augmentation.\n\nI was able to get the validation loss down to ~0.25. I got a similar loss when I calculated it for the 400 test videos. However, when I submitted it, I ended up with a loss of 0.97 against the 4,000 public test set. In reading through some of the discussion threads, others noted differences in validation and test scores, but not nearly as large.\n\nHere is a histogram of my predictions on the 400 test videos.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2F5a8d702117d43d54f78454ba57438568%2F2020-03-10_18-41-49.png?generation=1583891005206513&amp;alt=media)\n\nFor comparison sake, here is a histogram of the predictions from the [Xception Binary Classifier - Inference notebook](https://www.kaggle.com/greatgamedota/xception-binary-classifier-inference), which had a leader board score of 0.55.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2F345fb8076ed27611b2efc6dfd38b45ea%2F2020-03-10_18-49-04.png?generation=1583891377484329&amp;alt=media)\n\nThe predictions are obviously much more distributed throughout the range. Is the polarization in the predictions from my model indicative of over-training or have I gone wrong someplace else?",
    "768643": "Not sure over-fitting versus log loss not being intuitive versus accuracy. You can try to clip your prediction e.g. max 0.05 or .95, or take square root of multiple predictions and average them.  Log loss can highly penalize confident but wrong predictions, e.g. definitely don't want 0 and 1s.  ~.69 is for guessing all .5s.",
    "768645": "The first one looks more\"correct\" to me.\nIs it possible that you are switching the labels in your model?",
    "768663": "There is no label for 400 test videos, how did you calculate the loss?",
    "768665": "Thanks - I'll go back through and check to make sure. One thing is that my model predicts real, i.e. 1 = real and 0 = fake, just as a legacy of the way I initially put the data frame together that I then used to write records to disk. After the final sigmoid activation, I was just subtracting it from 1 to get the complementary probability. I didn't think that was messing up the results, especially since my local log loss and confusion matrix were looking decent, but I'll go back and check and in any event make sure I'm not making an error in conforming the predictions.\n\nI've struggled with getting my prediction routine to actually produce a score, so it's also possible something is still amiss there. The difference between 0.25 locally and .97 on the leader board struck me as possible the prediction routine on submit isn't working right.",
    "768669": "Thanks - I was messing around with the predictions and noted as somebody else did that the 0.69 with all 0.5s means that the public test set is perfectly balanced. That may not be the case for the leader board. My model, for whatever reason does a better on the fakes than it does on the real ones. My local 0.25 on the 50/50 test videos is basically getting all of the fakes right but missing about a quarter of the real. So that would be another possibility, that the leader board test set is skewed towards real videos.\n\n**EDIT**: This isn't right - I think the leaderboard test set likely is balanced 50/50 because the submit score comes back at 0.69 with 0.5 for all videos.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F781502%2Fb54e61aaf5ccd3be886b6e09489767f1%2F2020-03-10_21-52-21.png?generation=1583902389166759&amp;alt=media)",
    "768672": "They are a subset of the training videos. I can put up the data frame I constructed from the meta data json files if you'd like it. Then you can just look up the label by the video name.",
    "768676": "https://www.kaggle.com/calebeverett/metadata-dataframe",
    "768707": "It appears your model may be too confident. If you use 20 bins it would be more evident. Notice not much going on in the middle. You also have severe bias to predicting fake. Here is mine for comparison.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F389345%2F18e71790830104f5032e05eb08f34b09%2Fmodel.png?generation=1583903823209986&amp;alt=media)\n\nI suggest just use 1 frame and 1 face per video to start with and under sample the fake videos.",
    "768749": "I'll give that a try. Just out of curiosity, if you don't mind sharing what is the difference between your lb loss and loss on the 400?\n\nI noticed the bias and couldn't think of a reason why that would be the case. Every model I've trained has ended up with the same bias, even though I'm balancing the training data. Maybe it has something to do with the starting weights.",
    "768762": "About 0.14 on 4000 balanced samples in validation set.\n\nYou are balancing by class but you might also be repeating the same images to the extent that there are naturally occurring clusters of similar images that can bias the model towards the dominant clusters. This is more likely to happen if you sample multiple frames per video. To reduce the likelihood sample one frame per real video and one per fake video. Then under sample the fake videos to match the real. From there you can look at more complicated approaches.",
    "768766": "maralski I also faced that problem (skewed predictions toward 1). But according to your reason, my predictions should skew towards 0 instead, because more real samples are used (I currently used, say, 5 frames for 1 specific real video if that real video has 5 fake versions, and each fake has 1 frame). How do you explain this (predictions still skew towards 1)?",
    "768778": "khahuras What does the distribution look like and how are you choosing frames?",
    "772754": "NO. THIS IS NOT AS BAD AS YOU THINK. Too confident is a good thing for log loss. Our best subs looks just like that. Notice: If overfitted, it should be all 1.0s and 0.0s.",
    "773812": "Just for reference, here is our 0.30003LB histogram (20 bins) for 400 test videos. I usually check it every time and gives me a rough idea about how the model will perform after submission.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2279276%2Fc242d12df997fe52374606836fde25cf%2FScreen%20Shot%202020-03-15%20at%207.42.28%20PM.png?generation=1584315794339994&amp;alt=media)\n\nP.S. To match your 10 bin histogram:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2279276%2F3016dbbb523dc3e916ec720b304f76da%2FScreen%20Shot%202020-03-15%20at%207.46.34%20PM.png?generation=1584316041881590&amp;alt=media)",
    "774838": "So you're not clipping then?",
    "774843": "I am clipping, but with a very small value(0.01,0.09)."
  },
  "source": "meta"
}