{
  "id": 145886,
  "title": "Thoughts about this competition",
  "url": "/competitions/deepfake-detection-challenge/discussion/145886",
  "author_name": "Moshel",
  "post_date": "2020-04-24T23:34:58.474000",
  "votes": 6,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I would like to offer my thoughts on this competition and its usefulness. This is meant as constructive criticism, of things that they might want to look into in the future if there is ever DFDC V2 competition. </p>\n\n<p>As pointed out elsewhere, the results on the private LB were very disappointing from \"deepfake detection\" point of view. detection rate is ~60%.</p>\n\n<p>However, I think this was not Kagglers fault, but competition rules and dataset fault. \nWe got a fully synthetic dataset, no organic videos at all and we were expected to generalise well enough to detect organic fakes. Moreover, the (pretty vague) competition rules pretty much forbidden us from using any external data.</p>\n\n<p>I assume the idea was for us to \"grasp the essence\" of fake so it can generalise to any method, but this obviously didn't work too well.</p>\n\n<p>Because of the relatively small scenarios and actors and the cross appearance in all folders, it was very difficult to build stable CV, so many (if not most) used LB as indicator for model success, something that failed generally as a strategy for the private set.</p>\n\n<p>The dataset weakness also caused more perceptive models to overfit very quickly and recognise the actors faces and not the deep fake modifications.</p>\n\n<p>The low number of submissions per day (counting also fails) made people go to less advanced method. I would expect that at least toward the end of the competition number of submission per day will go up to allow honing and experimenting. </p>\n\n<p>Takeaway:\n1. If you want generalised results, give us generalised dataset\n2. Limiting external data that much does not help. Also, rules should be very clear. When you invest that much in a competition, it is best to invest a little in verifying the rules wording.\n3. Low number of submissions was very damaging IMHO for the complexity and creativity of solutions. Naturally, if we had good stable CV, this wouldn't be needed.\n4. A proper validation set should have been supplied. If you count the number of posts about problematic CV-LB gap, its probably 90% of the discussions.</p>",
  "messages": [
    {
      "id": 819846,
      "postDate": "2020-04-24T23:34:58.473Z",
      "content": "<p>I would like to offer my thoughts on this competition and its usefulness. This is meant as constructive criticism, of things that they might want to look into in the future if there is ever DFDC V2 competition. </p>\n\n<p>As pointed out elsewhere, the results on the private LB were very disappointing from \"deepfake detection\" point of view. detection rate is ~60%.</p>\n\n<p>However, I think this was not Kagglers fault, but competition rules and dataset fault. \nWe got a fully synthetic dataset, no organic videos at all and we were expected to generalise well enough to detect organic fakes. Moreover, the (pretty vague) competition rules pretty much forbidden us from using any external data.</p>\n\n<p>I assume the idea was for us to \"grasp the essence\" of fake so it can generalise to any method, but this obviously didn't work too well.</p>\n\n<p>Because of the relatively small scenarios and actors and the cross appearance in all folders, it was very difficult to build stable CV, so many (if not most) used LB as indicator for model success, something that failed generally as a strategy for the private set.</p>\n\n<p>The dataset weakness also caused more perceptive models to overfit very quickly and recognise the actors faces and not the deep fake modifications.</p>\n\n<p>The low number of submissions per day (counting also fails) made people go to less advanced method. I would expect that at least toward the end of the competition number of submission per day will go up to allow honing and experimenting. </p>\n\n<p>Takeaway:\n1. If you want generalised results, give us generalised dataset\n2. Limiting external data that much does not help. Also, rules should be very clear. When you invest that much in a competition, it is best to invest a little in verifying the rules wording.\n3. Low number of submissions was very damaging IMHO for the complexity and creativity of solutions. Naturally, if we had good stable CV, this wouldn't be needed.\n4. A proper validation set should have been supplied. If you count the number of posts about problematic CV-LB gap, its probably 90% of the discussions.</p>",
      "rawMarkdown": "I would like to offer my thoughts on this competition and its usefulness. This is meant as constructive criticism, of things that they might want to look into in the future if there is ever DFDC V2 competition. \n\nAs pointed out elsewhere, the results on the private LB were very disappointing from \"deepfake detection\" point of view. detection rate is ~60%.\n\nHowever, I think this was not Kagglers fault, but competition rules and dataset fault. \nWe got a fully synthetic dataset, no organic videos at all and we were expected to generalise well enough to detect organic fakes. Moreover, the (pretty vague) competition rules pretty much forbidden us from using any external data.\n\nI assume the idea was for us to \"grasp the essence\" of fake so it can generalise to any method, but this obviously didn't work too well.\n\nBecause of the relatively small scenarios and actors and the cross appearance in all folders, it was very difficult to build stable CV, so many (if not most) used LB as indicator for model success, something that failed generally as a strategy for the private set.\n\nThe dataset weakness also caused more perceptive models to overfit very quickly and recognise the actors faces and not the deep fake modifications.\n\nThe low number of submissions per day (counting also fails) made people go to less advanced method. I would expect that at least toward the end of the competition number of submission per day will go up to allow honing and experimenting. \n\nTakeaway:\n1. If you want generalised results, give us generalised dataset\n2. Limiting external data that much does not help. Also, rules should be very clear. When you invest that much in a competition, it is best to invest a little in verifying the rules wording.\n3. Low number of submissions was very damaging IMHO for the complexity and creativity of solutions. Naturally, if we had good stable CV, this wouldn't be needed.\n4. A proper validation set should have been supplied. If you count the number of posts about problematic CV-LB gap, its probably 90% of the discussions.\n",
      "votes": 6
    },
    {
      "id": 819861,
      "postDate": "2020-04-25T00:03:46.910Z",
      "content": "<p>I agree with you\nit's good to aim at creating model that grasps the essence of fake from limited data.\nbut, If organizer are seeking for really good model, they should explain it and provide suitable validation set.</p>",
      "rawMarkdown": "I agree with you\nit's good to aim at creating model that grasps the essence of fake from limited data.\nbut, If organizer are seeking for really good model, they should explain it and provide suitable validation set.\n\n",
      "votes": 1
    },
    {
      "id": 820250,
      "postDate": "2020-04-25T09:21:04.413Z",
      "content": "<p>What about building a validation data set with eternal data? Is it forbiden? I suspect some of the participants have done it.</p>",
      "rawMarkdown": "What about building a validation data set with eternal data? Is it forbiden? I suspect some of the participants have done it.",
      "replies": [
        {
          "id": 820346,
          "postDate": "2020-04-25T11:08:34.617Z",
          "content": "<p>It is complicated to decide if this is ok or not. Probably not. The license of videos in YouTube and terms and conditions of YouTube basically prevent you from using these videos and there is no dataset with open license of fake videos. </p>",
          "rawMarkdown": "It is complicated to decide if this is ok or not. Probably not. The license of videos in YouTube and terms and conditions of YouTube basically prevent you from using these videos and there is no dataset with open license of fake videos. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 820007,
      "postDate": "2020-04-25T04:57:13.180Z",
      "content": "<p>Yes, Agree</p>",
      "rawMarkdown": "Yes, Agree"
    },
    {
      "id": 819876,
      "postDate": "2020-04-25T00:31:05.577Z",
      "content": "<p>How did you come up with 60% detection rate?\nThe winning private score was 0.42320. \nAssuming the model outputs correct probabilities (I know it's not totally correct), a model with an accuracy of about 85% would get 85% of the predictions correct and 15% wrong. With this assumption, the average log loss would be:\n-(0.85 * log(0.85) + (1 - 0.85) * log(1 - 0.85)) = 0.4227 which is close to the score.\nSo, my guess for the average detection accuracy would be around 85%.</p>\n\n<p>Am I wrong?</p>",
      "rawMarkdown": "How did you come up with 60% detection rate?\nThe winning private score was 0.42320. \nAssuming the model outputs correct probabilities (I know it's not totally correct), a model with an accuracy of about 85% would get 85% of the predictions correct and 15% wrong. With this assumption, the average log loss would be:\n-(0.85 * log(0.85) + (1 - 0.85) * log(1 - 0.85)) = 0.4227 which is close to the score.\nSo, my guess for the average detection accuracy would be around 85%.\n\nAm I wrong?",
      "replies": [
        {
          "id": 819898,
          "postDate": "2020-04-25T01:11:01.823Z",
          "content": "<p>You can see some calculations in this thread <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145749\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145749</a>.\nAlthough it is very difficult to get estimate with logloss, this is also my estimate based on the cv runs i did during the competition (logloss to accuracy). Not saying it is 100% but probably more realistic than 1/0.</p>",
          "rawMarkdown": "You can see some calculations in this thread https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145749.\nAlthough it is very difficult to get estimate with logloss, this is also my estimate based on the cv runs i did during the competition (logloss to accuracy). Not saying it is 100% but probably more realistic than 1/0."
        }
      ]
    },
    {
      "id": 820235,
      "postDate": "2020-04-25T09:06:41.133Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 819861,
      "author_name": "あほ",
      "author_url": "",
      "post_date": "2020-04-25T00:03:46.910000",
      "content": "<p>I agree with you\nit's good to aim at creating model that grasps the essence of fake from limited data.\nbut, If organizer are seeking for really good model, they should explain it and provide suitable validation set.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 820250,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2020-04-25T09:21:04.413000",
      "content": "<p>What about building a validation data set with eternal data? Is it forbiden? I suspect some of the participants have done it.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 820346,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2020-04-25T11:08:34.617000",
          "content": "<p>It is complicated to decide if this is ok or not. Probably not. The license of videos in YouTube and terms and conditions of YouTube basically prevent you from using these videos and there is no dataset with open license of fake videos. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 820007,
      "author_name": "Eswar Chand",
      "author_url": "",
      "post_date": "2020-04-25T04:57:13.180000",
      "content": "<p>Yes, Agree</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 819876,
      "author_name": "Ces Bertino",
      "author_url": "",
      "post_date": "2020-04-25T00:31:05.577000",
      "content": "<p>How did you come up with 60% detection rate?\nThe winning private score was 0.42320. \nAssuming the model outputs correct probabilities (I know it's not totally correct), a model with an accuracy of about 85% would get 85% of the predictions correct and 15% wrong. With this assumption, the average log loss would be:\n-(0.85 * log(0.85) + (1 - 0.85) * log(1 - 0.85)) = 0.4227 which is close to the score.\nSo, my guess for the average detection accuracy would be around 85%.</p>\n\n<p>Am I wrong?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 819898,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2020-04-25T01:11:01.823000",
          "content": "<p>You can see some calculations in this thread <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145749\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145749</a>.\nAlthough it is very difficult to get estimate with logloss, this is also my estimate based on the cv runs i did during the competition (logloss to accuracy). Not saying it is 100% but probably more realistic than 1/0.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 820235,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-25T09:06:41.133000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "819846": "I would like to offer my thoughts on this competition and its usefulness. This is meant as constructive criticism, of things that they might want to look into in the future if there is ever DFDC V2 competition. \n\nAs pointed out elsewhere, the results on the private LB were very disappointing from \"deepfake detection\" point of view. detection rate is ~60%.\n\nHowever, I think this was not Kagglers fault, but competition rules and dataset fault. \nWe got a fully synthetic dataset, no organic videos at all and we were expected to generalise well enough to detect organic fakes. Moreover, the (pretty vague) competition rules pretty much forbidden us from using any external data.\n\nI assume the idea was for us to \"grasp the essence\" of fake so it can generalise to any method, but this obviously didn't work too well.\n\nBecause of the relatively small scenarios and actors and the cross appearance in all folders, it was very difficult to build stable CV, so many (if not most) used LB as indicator for model success, something that failed generally as a strategy for the private set.\n\nThe dataset weakness also caused more perceptive models to overfit very quickly and recognise the actors faces and not the deep fake modifications.\n\nThe low number of submissions per day (counting also fails) made people go to less advanced method. I would expect that at least toward the end of the competition number of submission per day will go up to allow honing and experimenting. \n\nTakeaway:\n1. If you want generalised results, give us generalised dataset\n2. Limiting external data that much does not help. Also, rules should be very clear. When you invest that much in a competition, it is best to invest a little in verifying the rules wording.\n3. Low number of submissions was very damaging IMHO for the complexity and creativity of solutions. Naturally, if we had good stable CV, this wouldn't be needed.\n4. A proper validation set should have been supplied. If you count the number of posts about problematic CV-LB gap, its probably 90% of the discussions.\n",
    "819861": "I agree with you\nit's good to aim at creating model that grasps the essence of fake from limited data.\nbut, If organizer are seeking for really good model, they should explain it and provide suitable validation set.\n\n",
    "820250": "What about building a validation data set with eternal data? Is it forbiden? I suspect some of the participants have done it.",
    "820007": "Yes, Agree",
    "819876": "How did you come up with 60% detection rate?\nThe winning private score was 0.42320. \nAssuming the model outputs correct probabilities (I know it's not totally correct), a model with an accuracy of about 85% would get 85% of the predictions correct and 15% wrong. With this assumption, the average log loss would be:\n-(0.85 * log(0.85) + (1 - 0.85) * log(1 - 0.85)) = 0.4227 which is close to the score.\nSo, my guess for the average detection accuracy would be around 85%.\n\nAm I wrong?",
    "820235": ""
  }
}