{
  "id": 134936,
  "title": "Analyzing Errors (the numbers)",
  "url": "/competitions/deepfake-detection-challenge/discussion/134936",
  "author_name": "",
  "post_date": "2020-03-11T07:40:53.118962400Z",
  "votes": 14,
  "comment_count": 1,
  "views": 0,
  "content": "<p>One of the most important parts of analyzing problems like this is analyzing the errors your model is making and potentially narrowing in on addressing the issues directly and focusing on the part of the pipeline where performance is being degraded. In this competition, a lot of us are reliant on external facial recognition models and we are fighting against several different deepfake methods so we need to be cognizant of mismatches in performance. </p>\n\n<p>The first step of doing this is inspecting which videos the model is getting wrong and where the majority of the loss is coming from. </p>\n\n<p>Preliminary numbers showed that the log loss on the 10769 validation videos was .11978. Significantly better than LB results, but that is an entirely different topic. </p>\n\n<p>First off, taking a look at the confusion matrix (threshold set at .5) we can see that there are quite a lot of true positives (the model is correctly identifying fakes as fake). There are very few false positives (model deemed video fake when it was real). Slightly more false negatives (model deemed video real when it was fake) and then a large number of true negatives (model deemed video real and it was real). \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Fb31355489e2ec2854ec0b24115adfa43%2Fdownload.png?generation=1583910142469658&amp;alt=media\" alt=\"\"></p>\n\n<p>But the metric we care about isn't really accuracy it is log loss so let's focus on that. </p>\n\n<p>Here is the overall loss histogram\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Ff38cd9a169633460eb09e69cfeacdd0b%2Fdownload%20(1\" alt=\"\">.png?generation=1583910451835914&amp;alt=media)</p>\n\n<p>We can see that the majority of validation samples we have very low log loss, the majority is below 0.2 but there is a small tail going all the way out to 3.1. That is coherent with the high accuracy we saw. So then the key is looking at the importance of the correctly classified samples vs the incorrectly classified samples. Is it more important for our majority to be even closer to the correct answer or to minimize the impact of the incorrect points with high loss?</p>\n\n<p>We can split the loss into the various TP, TN, FP, FN bins and analyze how much each is contributing. We can see that the average loss numbers of each bin show us the ones the model is getting correct have minuscule losses and the FN have a much higher mean of 1.359 and the FP follows it with 1.144. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F34a08a7bc61f635fb691d692d8d00123%2Fdownload%20(2\" alt=\"\">.png?generation=1583911813069510&amp;alt=media)</p>\n\n<p>That is interesting to know that our errors are on a different order of magnitude for the ones we get right vs the ones we get wrong, but not expected given it is log loss and the values rapidly scale when your model makes confident incorrect predictions. </p>\n\n<p>A more interesting analysis, though, is the overall sum of their contributions rather than averages as we care about the macro score more than anything and if we are to attack a bin in particular and minimize it' error we probably want to focus on the largest one as that is where there is the most room to grow. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Ff571b884236379e7551da1dea30b0a5a%2Fdownload%20(3\" alt=\"\">.png?generation=1583912078793596&amp;alt=media)</p>\n\n<p>Interestingly this shows a very different picture now. The true positives and true negatives end up being the largest contributors to overall loss (taking note that this is in the scenario where we have very high accuracy on validation, this does not hold true for the leaderboard). This is because while we had 88 false positives we 6726 True positives so even though the 88 false positives have a very high average log loss they overall contribute less. </p>\n\n<p>Given these numbers it is likely massively profitable on the leaderboard in order to create a model that has a greater focus on accuracy and generalization to the leaderboard because just a handful of confident incorrect answers can be nearly as powerful as literally thousands of correct classifications. </p>",
  "messages": [
    {
      "id": "768795",
      "postDate": "03/11/2020 07:40:53",
      "content": "<p>One of the most important parts of analyzing problems like this is analyzing the errors your model is making and potentially narrowing in on addressing the issues directly and focusing on the part of the pipeline where performance is being degraded. In this competition, a lot of us are reliant on external facial recognition models and we are fighting against several different deepfake methods so we need to be cognizant of mismatches in performance. </p>\n\n<p>The first step of doing this is inspecting which videos the model is getting wrong and where the majority of the loss is coming from. </p>\n\n<p>Preliminary numbers showed that the log loss on the 10769 validation videos was .11978. Significantly better than LB results, but that is an entirely different topic. </p>\n\n<p>First off, taking a look at the confusion matrix (threshold set at .5) we can see that there are quite a lot of true positives (the model is correctly identifying fakes as fake). There are very few false positives (model deemed video fake when it was real). Slightly more false negatives (model deemed video real when it was fake) and then a large number of true negatives (model deemed video real and it was real). \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Fb31355489e2ec2854ec0b24115adfa43%2Fdownload.png?generation=1583910142469658&amp;alt=media\" alt=\"\"></p>\n\n<p>But the metric we care about isn't really accuracy it is log loss so let's focus on that. </p>\n\n<p>Here is the overall loss histogram\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Ff38cd9a169633460eb09e69cfeacdd0b%2Fdownload%20(1\" alt=\"\">.png?generation=1583910451835914&amp;alt=media)</p>\n\n<p>We can see that the majority of validation samples we have very low log loss, the majority is below 0.2 but there is a small tail going all the way out to 3.1. That is coherent with the high accuracy we saw. So then the key is looking at the importance of the correctly classified samples vs the incorrectly classified samples. Is it more important for our majority to be even closer to the correct answer or to minimize the impact of the incorrect points with high loss?</p>\n\n<p>We can split the loss into the various TP, TN, FP, FN bins and analyze how much each is contributing. We can see that the average loss numbers of each bin show us the ones the model is getting correct have minuscule losses and the FN have a much higher mean of 1.359 and the FP follows it with 1.144. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F34a08a7bc61f635fb691d692d8d00123%2Fdownload%20(2\" alt=\"\">.png?generation=1583911813069510&amp;alt=media)</p>\n\n<p>That is interesting to know that our errors are on a different order of magnitude for the ones we get right vs the ones we get wrong, but not expected given it is log loss and the values rapidly scale when your model makes confident incorrect predictions. </p>\n\n<p>A more interesting analysis, though, is the overall sum of their contributions rather than averages as we care about the macro score more than anything and if we are to attack a bin in particular and minimize it' error we probably want to focus on the largest one as that is where there is the most room to grow. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Ff571b884236379e7551da1dea30b0a5a%2Fdownload%20(3\" alt=\"\">.png?generation=1583912078793596&amp;alt=media)</p>\n\n<p>Interestingly this shows a very different picture now. The true positives and true negatives end up being the largest contributors to overall loss (taking note that this is in the scenario where we have very high accuracy on validation, this does not hold true for the leaderboard). This is because while we had 88 false positives we 6726 True positives so even though the 88 false positives have a very high average log loss they overall contribute less. </p>\n\n<p>Given these numbers it is likely massively profitable on the leaderboard in order to create a model that has a greater focus on accuracy and generalization to the leaderboard because just a handful of confident incorrect answers can be nearly as powerful as literally thousands of correct classifications. </p>",
      "rawMarkdown": "One of the most important parts of analyzing problems like this is analyzing the errors your model is making and potentially narrowing in on addressing the issues directly and focusing on the part of the pipeline where performance is being degraded. In this competition, a lot of us are reliant on external facial recognition models and we are fighting against several different deepfake methods so we need to be cognizant of mismatches in performance. \n\nThe first step of doing this is inspecting which videos the model is getting wrong and where the majority of the loss is coming from. \n\nPreliminary numbers showed that the log loss on the 10769 validation videos was .11978. Significantly better than LB results, but that is an entirely different topic. \n\nFirst off, taking a look at the confusion matrix (threshold set at .5) we can see that there are quite a lot of true positives (the model is correctly identifying fakes as fake). There are very few false positives (model deemed video fake when it was real). Slightly more false negatives (model deemed video real when it was fake) and then a large number of true negatives (model deemed video real and it was real). \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Fb31355489e2ec2854ec0b24115adfa43%2Fdownload.png?generation=1583910142469658&amp;alt=media)\n\nBut the metric we care about isn't really accuracy it is log loss so let's focus on that. \n\nHere is the overall loss histogram\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Ff38cd9a169633460eb09e69cfeacdd0b%2Fdownload%20(1).png?generation=1583910451835914&amp;alt=media)\n\nWe can see that the majority of validation samples we have very low log loss, the majority is below 0.2 but there is a small tail going all the way out to 3.1. That is coherent with the high accuracy we saw. So then the key is looking at the importance of the correctly classified samples vs the incorrectly classified samples. Is it more important for our majority to be even closer to the correct answer or to minimize the impact of the incorrect points with high loss?\n\nWe can split the loss into the various TP, TN, FP, FN bins and analyze how much each is contributing. We can see that the average loss numbers of each bin show us the ones the model is getting correct have minuscule losses and the FN have a much higher mean of 1.359 and the FP follows it with 1.144. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F34a08a7bc61f635fb691d692d8d00123%2Fdownload%20(2).png?generation=1583911813069510&amp;alt=media)\n\nThat is interesting to know that our errors are on a different order of magnitude for the ones we get right vs the ones we get wrong, but not expected given it is log loss and the values rapidly scale when your model makes confident incorrect predictions. \n\nA more interesting analysis, though, is the overall sum of their contributions rather than averages as we care about the macro score more than anything and if we are to attack a bin in particular and minimize it' error we probably want to focus on the largest one as that is where there is the most room to grow. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Ff571b884236379e7551da1dea30b0a5a%2Fdownload%20(3).png?generation=1583912078793596&amp;alt=media)\n\nInterestingly this shows a very different picture now. The true positives and true negatives end up being the largest contributors to overall loss (taking note that this is in the scenario where we have very high accuracy on validation, this does not hold true for the leaderboard). This is because while we had 88 false positives we 6726 True positives so even though the 88 false positives have a very high average log loss they overall contribute less. \n\nGiven these numbers it is likely massively profitable on the leaderboard in order to create a model that has a greater focus on accuracy and generalization to the leaderboard because just a handful of confident incorrect answers can be nearly as powerful as literally thousands of correct classifications.",
      "votes": null
    },
    {
      "id": "768917",
      "postDate": "03/11/2020 10:56:53",
      "content": "<p>Nice. Really thanks for this beautiful post ryches. </p>",
      "rawMarkdown": "Nice. Really thanks for this beautiful post ryches.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 768917,
      "author_name": "khahuras",
      "author_url": "",
      "post_date": "03/11/2020 10:56:53",
      "content": "<p>Nice. Really thanks for this beautiful post ryches. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "768795": "One of the most important parts of analyzing problems like this is analyzing the errors your model is making and potentially narrowing in on addressing the issues directly and focusing on the part of the pipeline where performance is being degraded. In this competition, a lot of us are reliant on external facial recognition models and we are fighting against several different deepfake methods so we need to be cognizant of mismatches in performance. \n\nThe first step of doing this is inspecting which videos the model is getting wrong and where the majority of the loss is coming from. \n\nPreliminary numbers showed that the log loss on the 10769 validation videos was .11978. Significantly better than LB results, but that is an entirely different topic. \n\nFirst off, taking a look at the confusion matrix (threshold set at .5) we can see that there are quite a lot of true positives (the model is correctly identifying fakes as fake). There are very few false positives (model deemed video fake when it was real). Slightly more false negatives (model deemed video real when it was fake) and then a large number of true negatives (model deemed video real and it was real). \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Fb31355489e2ec2854ec0b24115adfa43%2Fdownload.png?generation=1583910142469658&amp;alt=media)\n\nBut the metric we care about isn't really accuracy it is log loss so let's focus on that. \n\nHere is the overall loss histogram\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Ff38cd9a169633460eb09e69cfeacdd0b%2Fdownload%20(1).png?generation=1583910451835914&amp;alt=media)\n\nWe can see that the majority of validation samples we have very low log loss, the majority is below 0.2 but there is a small tail going all the way out to 3.1. That is coherent with the high accuracy we saw. So then the key is looking at the importance of the correctly classified samples vs the incorrectly classified samples. Is it more important for our majority to be even closer to the correct answer or to minimize the impact of the incorrect points with high loss?\n\nWe can split the loss into the various TP, TN, FP, FN bins and analyze how much each is contributing. We can see that the average loss numbers of each bin show us the ones the model is getting correct have minuscule losses and the FN have a much higher mean of 1.359 and the FP follows it with 1.144. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F34a08a7bc61f635fb691d692d8d00123%2Fdownload%20(2).png?generation=1583911813069510&amp;alt=media)\n\nThat is interesting to know that our errors are on a different order of magnitude for the ones we get right vs the ones we get wrong, but not expected given it is log loss and the values rapidly scale when your model makes confident incorrect predictions. \n\nA more interesting analysis, though, is the overall sum of their contributions rather than averages as we care about the macro score more than anything and if we are to attack a bin in particular and minimize it' error we probably want to focus on the largest one as that is where there is the most room to grow. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Ff571b884236379e7551da1dea30b0a5a%2Fdownload%20(3).png?generation=1583912078793596&amp;alt=media)\n\nInterestingly this shows a very different picture now. The true positives and true negatives end up being the largest contributors to overall loss (taking note that this is in the scenario where we have very high accuracy on validation, this does not hold true for the leaderboard). This is because while we had 88 false positives we 6726 True positives so even though the 88 false positives have a very high average log loss they overall contribute less. \n\nGiven these numbers it is likely massively profitable on the leaderboard in order to create a model that has a greater focus on accuracy and generalization to the leaderboard because just a handful of confident incorrect answers can be nearly as powerful as literally thousands of correct classifications.",
    "768917": "Nice. Really thanks for this beautiful post ryches."
  },
  "source": "meta"
}