{
  "id": 133421,
  "title": "Reminder from the organizers",
  "url": "/competitions/deepfake-detection-challenge/discussion/133421",
  "author_name": "",
  "post_date": "2020-03-02T19:48:06.479712500Z",
  "votes": 17,
  "comment_count": 10,
  "views": 0,
  "content": "<p>As the competition is getting into its final stages, we want to remind the participants of an important point:</p>\n\n<p>As mentioned in the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started\">Getting Started section</a>: <em>“The private test set contains videos with a similar format and nature as the Training and Public Validation/Test Sets, but are real, organic videos with and without deepfakes”</em>. Some shuffle is expected in the leaderboard after scoring of the private dataset, specially for those methods that overfitted to the public test set. Always keeping in mind the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/rules\">external data policy</a> of the competition, we encourage participants to stress test their algorithm against as many different types of deepfakes as possible.</p>\n\n<p>Good luck!</p>",
  "messages": [
    {
      "id": "761638",
      "postDate": "03/02/2020 19:48:06",
      "content": "<p>As the competition is getting into its final stages, we want to remind the participants of an important point:</p>\n\n<p>As mentioned in the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started\">Getting Started section</a>: <em>“The private test set contains videos with a similar format and nature as the Training and Public Validation/Test Sets, but are real, organic videos with and without deepfakes”</em>. Some shuffle is expected in the leaderboard after scoring of the private dataset, specially for those methods that overfitted to the public test set. Always keeping in mind the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/rules\">external data policy</a> of the competition, we encourage participants to stress test their algorithm against as many different types of deepfakes as possible.</p>\n\n<p>Good luck!</p>",
      "rawMarkdown": "As the competition is getting into its final stages, we want to remind the participants of an important point:\n\nAs mentioned in the [Getting Started section](https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started): *“The private test set contains videos with a similar format and nature as the Training and Public Validation/Test Sets, but are real, organic videos with and without deepfakes”*. Some shuffle is expected in the leaderboard after scoring of the private dataset, specially for those methods that overfitted to the public test set. Always keeping in mind the [external data policy](https://www.kaggle.com/c/deepfake-detection-challenge/rules) of the competition, we encourage participants to stress test their algorithm against as many different types of deepfakes as possible.\n\nGood luck!",
      "votes": null
    },
    {
      "id": "761651",
      "postDate": "03/02/2020 20:22:15",
      "content": "<p>Most of the real external DeepFake data is either:\n- For research only, non commercial. Accoriding to <a href=\"/juliaelliott\">@juliaelliott</a> we are not allowed to use them\n- with unknown license like videos from youtube - an open question, unanswered\nThere was no official answer to a lot of  questions from participants regarding clear definition of rules. </p>",
      "rawMarkdown": "Most of the real external DeepFake data is either:\n- For research only, non commercial. Accoriding to @juliaelliott we are not allowed to use them\n- with unknown license like videos from youtube - an open question, unanswered\nThere was no official answer to a lot of  questions from participants regarding clear definition of rules.",
      "votes": null
    },
    {
      "id": "761661",
      "postDate": "03/02/2020 20:35:57",
      "content": "<p><a href=\"/cristiancanton\">@cristiancanton</a> </p>\n\n<p>&gt; External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.</p>\n\n<p>As I understand the above words: the participants are allowed to use <strong>ANY</strong> data (but not code) as long as anyone can get access to it and the link to it is shared in this <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203\">thread</a>.</p>\n\n<p>&gt; Use of Open Source. Unless otherwise stated in the Specific Competition Rules above, if open-source code is used in the model to generate the Submission, then you must only use open source code licensed under an Open Source Initiative-approved license (see www.opensource.org) that in no event limits commercial use of such code or model containing or depending on such code.</p>\n\n<p>As I understand, we can use only open-source packages that allow commercial use.</p>\n\n<p><a href=\"/cristiancanton\">@cristiancanton</a> Could you please get aligned with <a href=\"/addisonhoward\">@addisonhoward</a> and <a href=\"/juliaelliott\">@juliaelliott</a> on what external data is allowed and what is not and publish a list of the dataset and pretrained models from External Data thread that is allowed to use?</p>",
      "rawMarkdown": "cristiancanton \n\n&gt; External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.\n\nAs I understand the above words: the participants are allowed to use **ANY** data (but not code) as long as anyone can get access to it and the link to it is shared in this [thread](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203).\n\n\n&gt; Use of Open Source. Unless otherwise stated in the Specific Competition Rules above, if open-source code is used in the model to generate the Submission, then you must only use open source code licensed under an Open Source Initiative-approved license (see www.opensource.org) that in no event limits commercial use of such code or model containing or depending on such code.\n\nAs I understand, we can use only open-source packages that allow commercial use.\n\n@cristiancanton Could you please get aligned with @addisonhoward and @juliaelliott on what external data is allowed and what is not and publish a list of the dataset and pretrained models from[ External Data thread]((https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203)) that is allowed to use?",
      "votes": null
    },
    {
      "id": "761772",
      "postDate": "03/02/2020 23:18:55",
      "content": "<p>The policy is very clear. Please, check on that <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203\">thread</a> for data that has been already shared. Unfortunately, compiling a list of datasets that you can use is too much of an ask :) I'd recommend caution, reading thoroughly the legal terms of external data and open source packages and, if in case of doubt, err on the side of caution.</p>",
      "rawMarkdown": "The policy is very clear. Please, check on that [thread](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203) for data that has been already shared. Unfortunately, compiling a list of datasets that you can use is too much of an ask :) I'd recommend caution, reading thoroughly the legal terms of external data and open source packages and, if in case of doubt, err on the side of caution.",
      "votes": null
    },
    {
      "id": "761774",
      "postDate": "03/02/2020 23:20:18",
      "content": "<p>In case of doubt, do not use it. If the license is unknown, you should assume it is the most restrictive hence do not use it.</p>",
      "rawMarkdown": "In case of doubt, do not use it. If the license is unknown, you should assume it is the most restrictive hence do not use it.",
      "votes": null
    },
    {
      "id": "761795",
      "postDate": "03/03/2020 00:17:50",
      "content": "<p><a href=\"/cristiancanton\">@cristiancanton</a> </p>\n\n<p>Are the sponsors or Kaggle going to decide, regarding the winning teams: \"this guy used FaceForensics++, but that's OK / a disqualification\"?</p>\n\n<p>The thing about ImageNet, FaceForensics++ and InsightFace (RetinaFace, ArcFace) is that almost everyone is using some of them, and I'd be willing to bet that the top-scoring teams used most of them.</p>\n\n<p>So, I would love to see your and Kaggle's input regarding these two questions: \n<a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/133316\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/133316</a></p>",
      "rawMarkdown": "cristiancanton \n\nAre the sponsors or Kaggle going to decide, regarding the winning teams: \"this guy used FaceForensics++, but that's OK / a disqualification\"?\n\nThe thing about ImageNet, FaceForensics++ and InsightFace (RetinaFace, ArcFace) is that almost everyone is using some of them, and I'd be willing to bet that the top-scoring teams used most of them.\n\nSo, I would love to see your and Kaggle's input regarding these two questions: \nhttps://www.kaggle.com/c/deepfake-detection-challenge/discussion/133316",
      "votes": null
    },
    {
      "id": "761797",
      "postDate": "03/03/2020 00:21:23",
      "content": "<p>It is ironic that you encourage to stress test against different types of deepfakes, given that all relevant datasets are not allowed due to license restrictions. So do you mean that you encourage us to generate different types of deepfakes from source code?</p>",
      "rawMarkdown": "It is ironic that you encourage to stress test against different types of deepfakes, given that all relevant datasets are not allowed due to license restrictions. So do you mean that you encourage us to generate different types of deepfakes from source code?",
      "votes": null
    },
    {
      "id": "762076",
      "postDate": "03/03/2020 07:27:28",
      "content": "<p>To ensure Fairness and Transparency, upon Competition ended and preliminary Private leaderboard score and place has been calculated, can you ask the top 10 Kernels to fully disclose the External dataset they used (together with their kernel code) for Public scrutiny before the confirmation of any award? This will allowed the Kaggle community to help organizer in catching any form of cheating or overlook.\n<a href=\"/cristiancanton\">@cristiancanton</a> </p>",
      "rawMarkdown": "To ensure Fairness and Transparency, upon Competition ended and preliminary Private leaderboard score and place has been calculated, can you ask the top 10 Kernels to fully disclose the External dataset they used (together with their kernel code) for Public scrutiny before the confirmation of any award? This will allowed the Kaggle community to help organizer in catching any form of cheating or overlook.\n@cristiancanton",
      "votes": null
    },
    {
      "id": "762244",
      "postDate": "03/03/2020 11:06:09",
      "content": "<p>The policy is not very clear, that's why people are commenting here. \nI believe it won't be too much to ask to compile a list of datasets AFTER the merger deadline (after which no more external datasets can be disclosured). Unless it's unclear for the organizers as well which datasets can be used from those mentioned in the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203\">disclosure thread</a></p>",
      "rawMarkdown": "The policy is not very clear, that's why people are commenting here. \nI believe it won't be too much to ask to compile a list of datasets AFTER the merger deadline (after which no more external datasets can be disclosured). Unless it's unclear for the organizers as well which datasets can be used from those mentioned in the [disclosure thread](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203)",
      "votes": null
    },
    {
      "id": "762840",
      "postDate": "03/03/2020 21:10:46",
      "content": "<p><a href=\"/cristiancanton\">@cristiancanton</a> There are a lot of datasets that mentioned that they can only be used in non-commercial activities. Is this competition considered a commercial activity? in my opinion this is a non-commercial activity, with a prize for the best models. Could you please comment on this, as there are a lot of other competitors that have the same question. Thank you</p>",
      "rawMarkdown": "cristiancanton There are a lot of datasets that mentioned that they can only be used in non-commercial activities. Is this competition considered a commercial activity? in my opinion this is a non-commercial activity, with a prize for the best models. Could you please comment on this, as there are a lot of other competitors that have the same question. Thank you",
      "votes": null
    },
    {
      "id": "762843",
      "postDate": "03/03/2020 21:13:37",
      "content": "<p>It looks that most and maybe 100% of external datasets referenced cannot be used in this competition for our training for the reasons you've explained. Currently, I'm not able to find one allowed for all purposes including commercial. Did someone find one?</p>\n\n<p>However, they could be used for local testing/hold-out to see behavior of our models on unseen data but not to improve learning directly.</p>",
      "rawMarkdown": "It looks that most and maybe 100% of external datasets referenced cannot be used in this competition for our training for the reasons you've explained. Currently, I'm not able to find one allowed for all purposes including commercial. Did someone find one?\n\nHowever, they could be used for local testing/hold-out to see behavior of our models on unseen data but not to improve learning directly.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 761651,
      "author_name": "selimsef",
      "author_url": "",
      "post_date": "03/02/2020 20:22:15",
      "content": "<p>Most of the real external DeepFake data is either:\n- For research only, non commercial. Accoriding to <a href=\"/juliaelliott\">@juliaelliott</a> we are not allowed to use them\n- with unknown license like videos from youtube - an open question, unanswered\nThere was no official answer to a lot of  questions from participants regarding clear definition of rules. </p>",
      "votes": null,
      "replies": [
        {
          "id": 761774,
          "author_name": "cristiancanton",
          "author_url": "",
          "post_date": "03/02/2020 23:20:18",
          "content": "<p>In case of doubt, do not use it. If the license is unknown, you should assume it is the most restrictive hence do not use it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 762843,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "03/03/2020 21:13:37",
          "content": "<p>It looks that most and maybe 100% of external datasets referenced cannot be used in this competition for our training for the reasons you've explained. Currently, I'm not able to find one allowed for all purposes including commercial. Did someone find one?</p>\n\n<p>However, they could be used for local testing/hold-out to see behavior of our models on unseen data but not to improve learning directly.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 761661,
      "author_name": "iglovikov",
      "author_url": "",
      "post_date": "03/02/2020 20:35:57",
      "content": "<p><a href=\"/cristiancanton\">@cristiancanton</a> </p>\n\n<p>&gt; External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.</p>\n\n<p>As I understand the above words: the participants are allowed to use <strong>ANY</strong> data (but not code) as long as anyone can get access to it and the link to it is shared in this <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203\">thread</a>.</p>\n\n<p>&gt; Use of Open Source. Unless otherwise stated in the Specific Competition Rules above, if open-source code is used in the model to generate the Submission, then you must only use open source code licensed under an Open Source Initiative-approved license (see www.opensource.org) that in no event limits commercial use of such code or model containing or depending on such code.</p>\n\n<p>As I understand, we can use only open-source packages that allow commercial use.</p>\n\n<p><a href=\"/cristiancanton\">@cristiancanton</a> Could you please get aligned with <a href=\"/addisonhoward\">@addisonhoward</a> and <a href=\"/juliaelliott\">@juliaelliott</a> on what external data is allowed and what is not and publish a list of the dataset and pretrained models from External Data thread that is allowed to use?</p>",
      "votes": null,
      "replies": [
        {
          "id": 761772,
          "author_name": "cristiancanton",
          "author_url": "",
          "post_date": "03/02/2020 23:18:55",
          "content": "<p>The policy is very clear. Please, check on that <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203\">thread</a> for data that has been already shared. Unfortunately, compiling a list of datasets that you can use is too much of an ask :) I'd recommend caution, reading thoroughly the legal terms of external data and open source packages and, if in case of doubt, err on the side of caution.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 762244,
          "author_name": "rafiko1",
          "author_url": "",
          "post_date": "03/03/2020 11:06:09",
          "content": "<p>The policy is not very clear, that's why people are commenting here. \nI believe it won't be too much to ask to compile a list of datasets AFTER the merger deadline (after which no more external datasets can be disclosured). Unless it's unclear for the organizers as well which datasets can be used from those mentioned in the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203\">disclosure thread</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 761795,
      "author_name": "olegtrott",
      "author_url": "",
      "post_date": "03/03/2020 00:17:50",
      "content": "<p><a href=\"/cristiancanton\">@cristiancanton</a> </p>\n\n<p>Are the sponsors or Kaggle going to decide, regarding the winning teams: \"this guy used FaceForensics++, but that's OK / a disqualification\"?</p>\n\n<p>The thing about ImageNet, FaceForensics++ and InsightFace (RetinaFace, ArcFace) is that almost everyone is using some of them, and I'd be willing to bet that the top-scoring teams used most of them.</p>\n\n<p>So, I would love to see your and Kaggle's input regarding these two questions: \n<a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/133316\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/133316</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 761797,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "03/03/2020 00:21:23",
      "content": "<p>It is ironic that you encourage to stress test against different types of deepfakes, given that all relevant datasets are not allowed due to license restrictions. So do you mean that you encourage us to generate different types of deepfakes from source code?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 762076,
      "author_name": "chewkokwahibrainai",
      "author_url": "",
      "post_date": "03/03/2020 07:27:28",
      "content": "<p>To ensure Fairness and Transparency, upon Competition ended and preliminary Private leaderboard score and place has been calculated, can you ask the top 10 Kernels to fully disclose the External dataset they used (together with their kernel code) for Public scrutiny before the confirmation of any award? This will allowed the Kaggle community to help organizer in catching any form of cheating or overlook.\n<a href=\"/cristiancanton\">@cristiancanton</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 762840,
      "author_name": "ngcferreira",
      "author_url": "",
      "post_date": "03/03/2020 21:10:46",
      "content": "<p><a href=\"/cristiancanton\">@cristiancanton</a> There are a lot of datasets that mentioned that they can only be used in non-commercial activities. Is this competition considered a commercial activity? in my opinion this is a non-commercial activity, with a prize for the best models. Could you please comment on this, as there are a lot of other competitors that have the same question. Thank you</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "761638": "As the competition is getting into its final stages, we want to remind the participants of an important point:\n\nAs mentioned in the [Getting Started section](https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started): *“The private test set contains videos with a similar format and nature as the Training and Public Validation/Test Sets, but are real, organic videos with and without deepfakes”*. Some shuffle is expected in the leaderboard after scoring of the private dataset, specially for those methods that overfitted to the public test set. Always keeping in mind the [external data policy](https://www.kaggle.com/c/deepfake-detection-challenge/rules) of the competition, we encourage participants to stress test their algorithm against as many different types of deepfakes as possible.\n\nGood luck!",
    "761651": "Most of the real external DeepFake data is either:\n- For research only, non commercial. Accoriding to @juliaelliott we are not allowed to use them\n- with unknown license like videos from youtube - an open question, unanswered\nThere was no official answer to a lot of  questions from participants regarding clear definition of rules.",
    "761661": "cristiancanton \n\n&gt; External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.\n\nAs I understand the above words: the participants are allowed to use **ANY** data (but not code) as long as anyone can get access to it and the link to it is shared in this [thread](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203).\n\n\n&gt; Use of Open Source. Unless otherwise stated in the Specific Competition Rules above, if open-source code is used in the model to generate the Submission, then you must only use open source code licensed under an Open Source Initiative-approved license (see www.opensource.org) that in no event limits commercial use of such code or model containing or depending on such code.\n\nAs I understand, we can use only open-source packages that allow commercial use.\n\n@cristiancanton Could you please get aligned with @addisonhoward and @juliaelliott on what external data is allowed and what is not and publish a list of the dataset and pretrained models from[ External Data thread]((https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203)) that is allowed to use?",
    "761772": "The policy is very clear. Please, check on that [thread](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203) for data that has been already shared. Unfortunately, compiling a list of datasets that you can use is too much of an ask :) I'd recommend caution, reading thoroughly the legal terms of external data and open source packages and, if in case of doubt, err on the side of caution.",
    "761774": "In case of doubt, do not use it. If the license is unknown, you should assume it is the most restrictive hence do not use it.",
    "761795": "cristiancanton \n\nAre the sponsors or Kaggle going to decide, regarding the winning teams: \"this guy used FaceForensics++, but that's OK / a disqualification\"?\n\nThe thing about ImageNet, FaceForensics++ and InsightFace (RetinaFace, ArcFace) is that almost everyone is using some of them, and I'd be willing to bet that the top-scoring teams used most of them.\n\nSo, I would love to see your and Kaggle's input regarding these two questions: \nhttps://www.kaggle.com/c/deepfake-detection-challenge/discussion/133316",
    "761797": "It is ironic that you encourage to stress test against different types of deepfakes, given that all relevant datasets are not allowed due to license restrictions. So do you mean that you encourage us to generate different types of deepfakes from source code?",
    "762076": "To ensure Fairness and Transparency, upon Competition ended and preliminary Private leaderboard score and place has been calculated, can you ask the top 10 Kernels to fully disclose the External dataset they used (together with their kernel code) for Public scrutiny before the confirmation of any award? This will allowed the Kaggle community to help organizer in catching any form of cheating or overlook.\n@cristiancanton",
    "762244": "The policy is not very clear, that's why people are commenting here. \nI believe it won't be too much to ask to compile a list of datasets AFTER the merger deadline (after which no more external datasets can be disclosured). Unless it's unclear for the organizers as well which datasets can be used from those mentioned in the [disclosure thread](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203)",
    "762840": "cristiancanton There are a lot of datasets that mentioned that they can only be used in non-commercial activities. Is this competition considered a commercial activity? in my opinion this is a non-commercial activity, with a prize for the best models. Could you please comment on this, as there are a lot of other competitors that have the same question. Thank you",
    "762843": "It looks that most and maybe 100% of external datasets referenced cannot be used in this competition for our training for the reasons you've explained. Currently, I'm not able to find one allowed for all purposes including commercial. Did someone find one?\n\nHowever, they could be used for local testing/hold-out to see behavior of our models on unseen data but not to improve learning directly."
  },
  "source": "meta"
}