{
  "id": 133316,
  "title": "The rules seem vague. Kaggle needs to clarify them ASAP",
  "url": "/competitions/deepfake-detection-challenge/discussion/133316",
  "author_name": "Oleg Trott",
  "post_date": "2020-03-02T02:59:34.851000",
  "votes": 26,
  "comment_count": 36,
  "views": 0,
  "content": "<p>... and then enforce them accordingly.</p>\n\n<p>The rules say:</p>\n\n<blockquote>\n  <p>C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is <strong><em>available</em></strong> to use by all participants of the competition <strong><em>for purposes of the competition</em></strong> at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.</p>\n</blockquote>\n\n<p>However, it's not at all clear what \"available for the purposes of the competition\" means.</p>\n\n<ul>\n<li>ImageNet</li>\n<li>FaceForensics++ dataset and </li>\n<li>InsightFace (RetinaFace, ArcFace) annotations <em>and</em> pretrained models.</li>\n</ul>\n\n<p>are all available for non-commercial/research use only. FaceForensics++ additionally requires a sign-up.</p>\n\n<ul>\n<li><p>Are we allowed to use them directly?</p></li>\n<li><p>Are we allowed to use third-party models trained on the above?</p></li>\n</ul>\n\n<p>@juliaelliott wrote earlier:</p>\n\n<blockquote>\n  <p>@vladislavleketush and others - Your use of external data should conform to the external data specification in the competition's rules:</p>\n  \n  <blockquote>\n    <p>you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.</p>\n  </blockquote>\n  \n  <p>Therefore, licenses which place restrictions on datasets' use (whether by purpose, affiliation, cost or other restrictive means) would be in violation of this rule.</p>\n</blockquote>\n\n<p>But I don't think this answers these questions.</p>\n\n<p>Is our purpose \"commercial\" rather than \"research\" here? (If so, what about all of the patents?!)</p>\n\n<p>By the way, pretrained models might not be restricted by the license of the data on which they are trained:  <a href=\"https://towardsdatascience.com/the-most-important-supreme-court-decision-for-data-science-and-machine-learning-44cfc1c1bcaf\">https://towardsdatascience.com/the-most-important-supreme-court-decision-for-data-science-and-machine-learning-44cfc1c1bcaf</a></p>\n\n<p>We need clarity here, soon.</p>",
  "messages": [
    {
      "id": 761009,
      "postDate": "2020-03-02T02:59:34.853Z",
      "content": "<p>... and then enforce them accordingly.</p>\n\n<p>The rules say:</p>\n\n<blockquote>\n  <p>C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is <strong><em>available</em></strong> to use by all participants of the competition <strong><em>for purposes of the competition</em></strong> at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.</p>\n</blockquote>\n\n<p>However, it's not at all clear what \"available for the purposes of the competition\" means.</p>\n\n<ul>\n<li>ImageNet</li>\n<li>FaceForensics++ dataset and </li>\n<li>InsightFace (RetinaFace, ArcFace) annotations <em>and</em> pretrained models.</li>\n</ul>\n\n<p>are all available for non-commercial/research use only. FaceForensics++ additionally requires a sign-up.</p>\n\n<ul>\n<li><p>Are we allowed to use them directly?</p></li>\n<li><p>Are we allowed to use third-party models trained on the above?</p></li>\n</ul>\n\n<p>@juliaelliott wrote earlier:</p>\n\n<blockquote>\n  <p>@vladislavleketush and others - Your use of external data should conform to the external data specification in the competition's rules:</p>\n  \n  <blockquote>\n    <p>you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.</p>\n  </blockquote>\n  \n  <p>Therefore, licenses which place restrictions on datasets' use (whether by purpose, affiliation, cost or other restrictive means) would be in violation of this rule.</p>\n</blockquote>\n\n<p>But I don't think this answers these questions.</p>\n\n<p>Is our purpose \"commercial\" rather than \"research\" here? (If so, what about all of the patents?!)</p>\n\n<p>By the way, pretrained models might not be restricted by the license of the data on which they are trained:  <a href=\"https://towardsdatascience.com/the-most-important-supreme-court-decision-for-data-science-and-machine-learning-44cfc1c1bcaf\">https://towardsdatascience.com/the-most-important-supreme-court-decision-for-data-science-and-machine-learning-44cfc1c1bcaf</a></p>\n\n<p>We need clarity here, soon.</p>",
      "rawMarkdown": "... and then enforce them accordingly.\n\nThe rules say:\n\n&gt; C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is ***available*** to use by all participants of the competition ***for purposes of the competition*** at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.\n\nHowever, it's not at all clear what \"available for the purposes of the competition\" means.\n\n- ImageNet\n- FaceForensics++ dataset and \n- InsightFace (RetinaFace, ArcFace) annotations *and* pretrained models.\n\nare all available for non-commercial/research use only. FaceForensics++ additionally requires a sign-up.\n\n- Are we allowed to use them directly?\n\n- Are we allowed to use third-party models trained on the above?\n\n@juliaelliott wrote earlier:\n\n&gt;@vladislavleketush and others - Your use of external data should conform to the external data specification in the competition's rules:\n\n&gt;&gt;you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.\n\n&gt;Therefore, licenses which place restrictions on datasets' use (whether by purpose, affiliation, cost or other restrictive means) would be in violation of this rule.\n\nBut I don't think this answers these questions.\n\nIs our purpose \"commercial\" rather than \"research\" here? (If so, what about all of the patents?!)\n\nBy the way, pretrained models might not be restricted by the license of the data on which they are trained:  https://towardsdatascience.com/the-most-important-supreme-court-decision-for-data-science-and-machine-learning-44cfc1c1bcaf\n\nWe need clarity here, soon.\n",
      "votes": 26
    },
    {
      "id": 762900,
      "postDate": "2020-03-03T23:29:22.583Z",
      "content": "<p>I'm sorry that this continues to be unclear despite multiple responses I've made. I will make another final statement on the matter, in the hopes this puts it to rest. This position is in alignment with the host's wishes.</p>\n\n<p>Regurgitating the rules, the external data requirement is that any dataset you use (for training or otherwise that contributes to your solution/submission) must be:\n1. available to use by <strong>all participants</strong>\n2. for the purposes of the competition - more specifically meaning that the solution must be possible to license in accordance with the WINNER LICENSE requirements stipulated in the rules \"You hereby license and will license your winning Submission and the source code used to generate the Submission under an Open Source Initiative-approved license (see www.opensource.org) that in no event limits commercial use of such code or model containing or depending on such code...\"\n3. at no cost to the other participants\n4. posted to the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203\">Official External Data thread</a> by the deadline (March 3, 2020 @ 11:59pm UTC)</p>\n\n<p>I have made multiple clarifications that <strong>datasets with \"non-commercial only\" or \"academic use only\" or \"research use only\" provisions are prohibited</strong>, because they inherently break both:\n- #1: Use is limited to those who are capable of furnishing an academic/research affiliation, therefore not available for use by all participants.\n- #2: Use of a commercially-restricted dataset in your solution will render your solution unable to meet the host's licensing requirements, because it would limit commercial use of such code.</p>\n\n<p>We are not in a position to be posting a comprehensive list of every \"approved\" dataset.</p>\n\n<p>As far as enforceability, the host has the right to review solutions (in particular those in winning standing) and disqualify solutions which violate this rule.</p>",
      "rawMarkdown": "I'm sorry that this continues to be unclear despite multiple responses I've made. I will make another final statement on the matter, in the hopes this puts it to rest. This position is in alignment with the host's wishes.\n\nRegurgitating the rules, the external data requirement is that any dataset you use (for training or otherwise that contributes to your solution/submission) must be:\n1. available to use by **all participants**\n2. for the purposes of the competition - more specifically meaning that the solution must be possible to license in accordance with the WINNER LICENSE requirements stipulated in the rules \"You hereby license and will license your winning Submission and the source code used to generate the Submission under an Open Source Initiative-approved license (see www.opensource.org) that in no event limits commercial use of such code or model containing or depending on such code...\"\n3. at no cost to the other participants\n4. posted to the [Official External Data thread](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203) by the deadline (March 3, 2020 @ 11:59pm UTC)\n\nI have made multiple clarifications that **datasets with \"non-commercial only\" or \"academic use only\" or \"research use only\" provisions are prohibited**, because they inherently break both:\n- #1: Use is limited to those who are capable of furnishing an academic/research affiliation, therefore not available for use by all participants.\n- #2: Use of a commercially-restricted dataset in your solution will render your solution unable to meet the host's licensing requirements, because it would limit commercial use of such code.\n\nWe are not in a position to be posting a comprehensive list of every \"approved\" dataset.\n\nAs far as enforceability, the host has the right to review solutions (in particular those in winning standing) and disqualify solutions which violate this rule.",
      "votes": 11,
      "replies": [
        {
          "id": 762926,
          "postDate": "2020-03-04T00:01:38.323Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> </p>\n\n<p>Thanks for replying.</p>\n\n<blockquote>\n  <p>for training or otherwise that contributes to your solution/submission</p>\n</blockquote>\n\n<p>In the interest of clarity, can you give a binary answer to whether models pretrained by third-parties on ImageNet are prohibited?</p>",
          "rawMarkdown": "@juliaelliott \n\nThanks for replying.\n\n&gt; for training or otherwise that contributes to your solution/submission\n\nIn the interest of clarity, can you give a binary answer to whether models pretrained by third-parties on ImageNet are prohibited?\n\n\n",
          "votes": 4
        },
        {
          "id": 763121,
          "postDate": "2020-03-04T06:45:58.533Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> \n&gt; 2: Use of a commercially-restricted dataset in your solution will render your solution unable to meet the host's licensing requirements, because it would limit commercial use of such code.</p>\n\n<p>I have a doubt with this statement. It lacks the consideration about the distinction between dataset itself and derived model. The interpretation about the license above can give the less chance for host to get a good model, e.g we miss ImageNet derived pre-trained model. That's the concern.</p>",
          "rawMarkdown": "@juliaelliott \n&gt; 2: Use of a commercially-restricted dataset in your solution will render your solution unable to meet the host's licensing requirements, because it would limit commercial use of such code.\n\nI have a doubt with this statement. It lacks the consideration about the distinction between dataset itself and derived model. The interpretation about the license above can give the less chance for host to get a good model, e.g we miss ImageNet derived pre-trained model. That's the concern.\n",
          "votes": 2
        },
        {
          "id": 763149,
          "postDate": "2020-03-04T07:26:58.450Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Thank you for making this clear.</p>",
          "rawMarkdown": "@juliaelliott Thank you for making this clear.",
          "votes": -3
        },
        {
          "id": 763731,
          "postDate": "2020-03-04T19:40:07.507Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> still it doesn't cover all of the cases</p>\n\n<p>Specifically a lot of competitors posted links to youtube.com and other websites that don't have dataset per se but can be scraped/collected. </p>\n\n<p>Therefore it is possible that :</p>\n\n<ul>\n<li>Competitor A: plays fair and safe and uses only allowed dataset (which basically means no external  deepfake data). Develops a good approach and models.</li>\n<li>Competitor B: uses a lot of data from youtube and other websites. Doesn't develop anything, let's say just uses some kernel with a lot of scraped data.</li>\n<li>I'm pretty sure Competitor B will win on the private set. </li>\n</ul>\n\n<p>I asked this question multiple times without any specific answer. Are competitors allowed to use data from youtube and other websites with CC license?</p>",
          "rawMarkdown": "@juliaelliott still it doesn't cover all of the cases\n\nSpecifically a lot of competitors posted links to youtube.com and other websites that don't have dataset per se but can be scraped/collected. \n\nTherefore it is possible that :\n\n- Competitor A: plays fair and safe and uses only allowed dataset (which basically means no external  deepfake data). Develops a good approach and models.\n- Competitor B: uses a lot of data from youtube and other websites. Doesn't develop anything, let's say just uses some kernel with a lot of scraped data.\n- I'm pretty sure Competitor B will win on the private set. \n\nI asked this question multiple times without any specific answer. Are competitors allowed to use data from youtube and other websites with CC license?\n\n\n\n\n",
          "votes": 20
        },
        {
          "id": 764857,
          "postDate": "2020-03-06T01:44:35.600Z",
          "content": "<p>Another edge case: the media is public domain, but it's collected in a database licensed under a \"non-commercial\" license, as posted here: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203#748358\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203#748358</a> by <a href=\"/harangdev\">@harangdev</a> </p>",
          "rawMarkdown": "Another edge case: the media is public domain, but it's collected in a database licensed under a \"non-commercial\" license, as posted here: https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203#748358 by @harangdev "
        },
        {
          "id": 781620,
          "postDate": "2020-03-21T13:59:14.450Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> any news? A lot of people reported youtube as their external data. Could you please clarify if that's legit? Just ignoring this question multiple times doesn't help at all. </p>",
          "rawMarkdown": "@juliaelliott any news? A lot of people reported youtube as their external data. Could you please clarify if that's legit? Just ignoring this question multiple times doesn't help at all. ",
          "votes": 9
        }
      ]
    },
    {
      "id": 761177,
      "postDate": "2020-03-02T08:15:55.733Z",
      "content": "<p><a href=\"/skylord\">@skylord</a> I would not be so sure.</p>\n\n<p>An answer from <a href=\"/juliaelliott\">@juliaelliott</a> regarding CC BY-NC</p>\n\n<blockquote>\n  <p>I’ve answered the question about BY-NC not being available for use by all (non-commercial use) and therefore violating the requirement that external data be available for use by all participants.</p>\n</blockquote>\n\n<p>Regarding FF+ etc the answer was also clear:</p>\n\n<blockquote>\n  <p>Edit: Also looking through all major fake videos datasets they all seem to be restricted for non-commercial research only: FaceForensics, FaceForensics++, DeepFakes Detection, both versions of Celeb-DF and still unpublished DeeperForensics-1.0. Is it ok to use them?</p>\n</blockquote>\n\n<p>An answer from <a href=\"/juliaelliott\">@juliaelliott</a> </p>\n\n<blockquote>\n  <p>Yes, if there are restrictions on a dataset’s use, it would violate the competition rules.</p>\n</blockquote>\n\n<p>Still it sounds weird that we are basically doing research but cannot use datasets for research only.</p>",
      "rawMarkdown": "@skylord I would not be so sure.\n\nAn answer from @juliaelliott regarding CC BY-NC\n&gt; I’ve answered the question about BY-NC not being available for use by all (non-commercial use) and therefore violating the requirement that external data be available for use by all participants.\n\nRegarding FF+ etc the answer was also clear:\n&gt; Edit: Also looking through all major fake videos datasets they all seem to be restricted for non-commercial research only: FaceForensics, FaceForensics++, DeepFakes Detection, both versions of Celeb-DF and still unpublished DeeperForensics-1.0. Is it ok to use them?\n\nAn answer from @juliaelliott \n&gt; Yes, if there are restrictions on a dataset’s use, it would violate the competition rules.\n\nStill it sounds weird that we are basically doing research but cannot use datasets for research only.\n",
      "votes": 10,
      "replies": [
        {
          "id": 761315,
          "postDate": "2020-03-02T11:36:39.683Z",
          "content": "<p>It seems the output of this will be used commercially in one way or the other ... </p>",
          "rawMarkdown": "It seems the output of this will be used commercially in one way or the other ... "
        },
        {
          "id": 761821,
          "postDate": "2020-03-03T00:43:59.413Z",
          "content": "<blockquote>\n  <p><strong>Selim Seferbekov wrote:</strong></p>\n  \n  <p>An answer from <a href=\"/juliaelliott\">@juliaelliott</a> </p>\n  \n  <blockquote>\n    <p>Yes, if there are restrictions on a dataset’s use, it would violate the competition rules.</p>\n  </blockquote>\n</blockquote>\n\n<p>This would imply that ImageNet is forbidden, but everyone uses it, and Kaggle is OK with it.</p>",
          "rawMarkdown": "&gt; **Selim Seferbekov wrote:**\n&gt; \n&gt; \n&gt; An answer from @juliaelliott \n&gt;&gt;Yes, if there are restrictions on a dataset’s use, it would violate the competition rules.\n&gt; \n\nThis would imply that ImageNet is forbidden, but everyone uses it, and Kaggle is OK with it.",
          "votes": 1
        },
        {
          "id": 762204,
          "postDate": "2020-03-03T10:19:37.110Z",
          "content": "<blockquote>\n  <p>This would imply that ImageNet is forbidden, but everyone uses it, and Kaggle is OK with it.</p>\n</blockquote>\n\n<p>This is not entirely true. Everyone uses models that are pretrained on ImageNet, which is different from using ImageNet directly. (In addition, ImageNet is available on Kaggle these days.)</p>\n\n<p>I think the <em>spirit</em> of the External Data rules is that the playing field should be equal for all participants. FaceForensics++ is off limits because not everyone can download it. Open source models that are pretrained on ImageNet should be fine (if the license allows it).</p>",
          "rawMarkdown": "&gt; This would imply that ImageNet is forbidden, but everyone uses it, and Kaggle is OK with it.\n\nThis is not entirely true. Everyone uses models that are pretrained on ImageNet, which is different from using ImageNet directly. (In addition, ImageNet is available on Kaggle these days.)\n\nI think the *spirit* of the External Data rules is that the playing field should be equal for all participants. FaceForensics++ is off limits because not everyone can download it. Open source models that are pretrained on ImageNet should be fine (if the license allows it).",
          "votes": 3
        },
        {
          "id": 762225,
          "postDate": "2020-03-03T10:40:19.097Z",
          "content": "<p>I thought that FF++ data can be downloaded after filling a google form ? They send you instructions for downloading the data. So the data is available for use. </p>\n\n<p>I think the organizers should list down all the datasets mentioned in the external data thread &amp; create a Yes/No table. This will remove the confusion. </p>",
          "rawMarkdown": "I thought that FF++ data can be downloaded after filling a google form ? They send you instructions for downloading the data. So the data is available for use. \n\nI think the organizers should list down all the datasets mentioned in the external data thread &amp; create a Yes/No table. This will remove the confusion. "
        },
        {
          "id": 762231,
          "postDate": "2020-03-03T10:50:49.640Z",
          "content": "<p><a href=\"/humananalog\">@humananalog</a> </p>\n\n<p>&gt; ImageNet, which is different from using ImageNet directly</p>\n\n<p>They are potentially different, hence the two separate questions in the original post.</p>\n\n<p>The reason I mention ImageNet is that it's analogous to the other two.</p>\n\n<p><a href=\"/skylord\">@skylord</a> </p>\n\n<p>It's impractical to go through all the things mentioned in that thread, but they should rule on the big three.</p>\n\n<p>Whether we are allowed to use FaceForensics++, or models trained on it is totally unclear.</p>",
          "rawMarkdown": "@humananalog \n\n&gt; ImageNet, which is different from using ImageNet directly\n\nThey are potentially different, hence the two separate questions in the original post.\n\nThe reason I mention ImageNet is that it's analogous to the other two.\n\n@skylord \n\nIt's impractical to go through all the things mentioned in that thread, but they should rule on the big three.\n\nWhether we are allowed to use FaceForensics++, or models trained on it is totally unclear.\n"
        },
        {
          "id": 762569,
          "postDate": "2020-03-03T16:06:34.373Z",
          "content": "<p>There was a bit of sarcasm in that....  🙂 <br>\nThere are around 350+ discussion threads, each with multiple links. Its not practical to read through each of them. </p>\n\n<p>But considering the <em>spirit of kaggle</em>  - which is open discussions, sharing of code kernels &amp; easily available datasets. I think as long as kagglers can <strong>access</strong> the data w/o any cost. (<em>freely available</em>) it should be ok. But again, it's just my reading! </p>",
          "rawMarkdown": "There was a bit of sarcasm in that....  🙂  \nThere are around 350+ discussion threads, each with multiple links. Its not practical to read through each of them. \n\nBut considering the *spirit of kaggle*  - which is open discussions, sharing of code kernels &amp; easily available datasets. I think as long as kagglers can **access** the data w/o any cost. (*freely available*) it should be ok. But again, it's just my reading! \n"
        },
        {
          "id": 762604,
          "postDate": "2020-03-03T16:28:15.157Z",
          "content": "<blockquote>\n  <p>I thought that FF++ data can be downloaded after filling a google form ? They send you instructions for downloading the data. So the data is available for use.</p>\n</blockquote>\n\n<p>Do they accept everyone? And if not, are you allowed to give the data to that person, for example by uploading it as a dataset on Kaggle?</p>",
          "rawMarkdown": "&gt; I thought that FF++ data can be downloaded after filling a google form ? They send you instructions for downloading the data. So the data is available for use.\n\nDo they accept everyone? And if not, are you allowed to give the data to that person, for example by uploading it as a dataset on Kaggle?\n"
        },
        {
          "id": 762646,
          "postDate": "2020-03-03T17:16:12.167Z",
          "content": "<p><a href=\"/humananalog\">@humananalog</a> \nWhile filling the google form, I wrote that I was taking part in the DFDC competition. within a day or two of filling it, I received an acceptance email. </p>\n\n<p>Edit: (a clarification)\nI believe you are referring to point 1, in the TOS. <a href=\"http://kaldir.vc.in.tum.de/FaceForensics/webpage/FaceForensics_TOS.pdf\">FaceForensics TOS</a></p>",
          "rawMarkdown": "@humananalog \nWhile filling the google form, I wrote that I was taking part in the DFDC competition. within a day or two of filling it, I received an acceptance email. \n\nEdit: (a clarification)\nI believe you are referring to point 1, in the TOS. [FaceForensics TOS](http://kaldir.vc.in.tum.de/FaceForensics/webpage/FaceForensics_TOS.pdf)\n "
        },
        {
          "id": 762651,
          "postDate": "2020-03-03T17:24:38.227Z",
          "content": "<p>You should also check point 4 &amp;6 \n```\n4. Researcher may provide research associates and colleagues with access to the Database provided that they first agree\nto be bound by these terms and conditions.\n.\n.\n.</p>\n\n<ol>\n<li>If Researcher is employed by a for-profit, commercial entity, Researcher's employer shall also be bound by these terms\nand conditions, and Researcher hereby represents that he or she is fully authorized to enter into this agreement on behalf\nof such employer.\n```</li>\n</ol>",
          "rawMarkdown": "You should also check point 4 &amp;6 \n```\n4. Researcher may provide research associates and colleagues with access to the Database provided that they first agree\nto be bound by these terms and conditions.\n.\n.\n.\n\n6. If Researcher is employed by a for-profit, commercial entity, Researcher's employer shall also be bound by these terms\nand conditions, and Researcher hereby represents that he or she is fully authorized to enter into this agreement on behalf\nof such employer.\n```\n"
        },
        {
          "id": 762656,
          "postDate": "2020-03-03T17:34:13.757Z",
          "content": "<p>This is a pretty clear answer: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203#744289\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203#744289</a></p>",
          "rawMarkdown": "This is a pretty clear answer: https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203#744289\n"
        },
        {
          "id": 762685,
          "postDate": "2020-03-03T18:00:46.063Z",
          "content": "<p>okay interesting!... I seem to have missed it. Thanks for pointing out. </p>",
          "rawMarkdown": "okay interesting!... I seem to have missed it. Thanks for pointing out. "
        },
        {
          "id": 762711,
          "postDate": "2020-03-03T18:45:07.123Z",
          "content": "<p>But then. How do we generalize to real world data ( like a deepfaked YouTube video) ? </p>",
          "rawMarkdown": "But then. How do we generalize to real world data ( like a deepfaked YouTube video) ? \n"
        },
        {
          "id": 762714,
          "postDate": "2020-03-03T18:47:59.720Z",
          "content": "<blockquote>\n  <p>Likewise, the winners' licensing terms require that the solution be <strong>open sourced</strong> that that in no event limits commercial use of such <strong>code</strong> or model containing or depending on such <strong>code</strong>.</p>\n</blockquote>\n\n<p><a href=\"/humananalog\">@humananalog</a> images are not code, it explicitly says you need to provide the code to generate the submission, not the data. They don't even define submission in the rules. It's not clear if it even must include the training code.</p>",
          "rawMarkdown": "&gt; Likewise, the winners' licensing terms require that the solution be **open sourced** that that in no event limits commercial use of such **code** or model containing or depending on such **code**.\n\n@humananalog images are not code, it explicitly says you need to provide the code to generate the submission, not the data. They don't even define submission in the rules. It's not clear if it even must include the training code.",
          "votes": 1
        },
        {
          "id": 762888,
          "postDate": "2020-03-03T22:52:13.230Z",
          "content": "<p><a href=\"/humananalog\">@humananalog</a> </p>\n\n<blockquote>\n  <p>ImageNet is available for download on Kaggle.</p>\n</blockquote>\n\n<p>... for non-commercial use only.</p>\n\n<p>From <a href=\"https://www.kaggle.com/c/imagenet-object-localization-challenge/rules\">https://www.kaggle.com/c/imagenet-object-localization-challenge/rules</a>:</p>\n\n<blockquote>\n  <p>Participant shall use the Database only for non-commercial research and educational purposes</p>\n</blockquote>\n\n<p>If the organizers can use models trained by others on ImageNet, what policy provisions stop them from using models trained by others on, say, FaceForensics++ or InsightFace?</p>",
          "rawMarkdown": "@humananalog \n\n&gt; ImageNet is available for download on Kaggle.\n\n... for non-commercial use only.\n\nFrom https://www.kaggle.com/c/imagenet-object-localization-challenge/rules:\n\n&gt; Participant shall use the Database only for non-commercial research and educational purposes\n\nIf the organizers can use models trained by others on ImageNet, what policy provisions stop them from using models trained by others on, say, FaceForensics++ or InsightFace?\n",
          "votes": 1
        },
        {
          "id": 762920,
          "postDate": "2020-03-03T23:57:56.030Z",
          "content": "<p><a href=\"/olegtrott\">@olegtrott</a> There are a few different things at play here:</p>\n\n<ol>\n<li>For some of these datasets the question is: can everyone use them? When I said ImageNet is available for download on Kaggle, that's what I was referring to. You don't need to fill out a form in order to download ImageNet, as opposed to FaceForensics++ for example.</li>\n<li>The other question is: can the dataset be used for non-commercial purposes? As you point out, ImageNet can't.</li>\n<li>However, it should be fine to use a model that is trained on ImageNet, if it is freely available to everyone and has a license that allows commercial usage. This is not the same thing as training on ImageNet directly.</li>\n</ol>\n\n<blockquote>\n  <p>If the organizers can use models trained by others on ImageNet, what policy provisions stop them from using models trained by others on, say, FaceForensics++ or InsightFace?</p>\n</blockquote>\n\n<p>As I understand the rules, a dataset such as FaceForensics++ is off-limits for this competition because it violates points 1 and 2 above. So any submissions using that dataset would be disqualified.</p>\n\n<p>However, if someone has trained a model on FaceForensics++, and makes it available under a license that is compatible with the competition, to everyone in the competition, then anyone will be able to use such a model in their own submissions (at least, if they had done so before the external data deadline).</p>",
          "rawMarkdown": "@olegtrott There are a few different things at play here:\n\n1. For some of these datasets the question is: can everyone use them? When I said ImageNet is available for download on Kaggle, that's what I was referring to. You don't need to fill out a form in order to download ImageNet, as opposed to FaceForensics++ for example.\n2. The other question is: can the dataset be used for non-commercial purposes? As you point out, ImageNet can't.\n3. However, it should be fine to use a model that is trained on ImageNet, if it is freely available to everyone and has a license that allows commercial usage. This is not the same thing as training on ImageNet directly.\n\n&gt; If the organizers can use models trained by others on ImageNet, what policy provisions stop them from using models trained by others on, say, FaceForensics++ or InsightFace?\n\nAs I understand the rules, a dataset such as FaceForensics++ is off-limits for this competition because it violates points 1 and 2 above. So any submissions using that dataset would be disqualified.\n\nHowever, if someone has trained a model on FaceForensics++, and makes it available under a license that is compatible with the competition, to everyone in the competition, then anyone will be able to use such a model in their own submissions (at least, if they had done so before the external data deadline)."
        },
        {
          "id": 762936,
          "postDate": "2020-03-04T00:26:01.493Z",
          "content": "<p><a href=\"/humananalog\">@humananalog</a> </p>\n\n<p>Kaggle does not appear to make a distinction between direct use of data and indirect one (see Julia's comment here). </p>\n\n<blockquote>\n  <p>any dataset you use (for training or <strong><em>OTHERWISE</em></strong> that contributes to your solution/submission) must be</p>\n</blockquote>",
          "rawMarkdown": "@humananalog \n\nKaggle does not appear to make a distinction between direct use of data and indirect one (see Julia's comment here). \n\n&gt; any dataset you use (for training or ***OTHERWISE*** that contributes to your solution/submission) must be",
          "votes": 1
        },
        {
          "id": 763306,
          "postDate": "2020-03-04T10:49:39.447Z",
          "content": "<p><a href=\"/olegtrott\">@olegtrott</a> That's a very wide interpretation of the rules. I think our disagreement comes from the fact that I say \"using a model trained on a dataset\" is totally unrelated to the original dataset, so under this interpretation \"any dataset you use\" does not mean you're using ImageNet if you're using a model trained on ImageNet (i.e it only covers the model not the dataset is was originally trained on). While you appear to believe that \"using a model trained on ImageNet\" by extension involves \"using ImageNet\".</p>\n\n<p>I guess we don't know how the competition organizers will interpret this wording. Perhaps it is vague on purpose.</p>",
          "rawMarkdown": "@olegtrott That's a very wide interpretation of the rules. I think our disagreement comes from the fact that I say \"using a model trained on a dataset\" is totally unrelated to the original dataset, so under this interpretation \"any dataset you use\" does not mean you're using ImageNet if you're using a model trained on ImageNet (i.e it only covers the model not the dataset is was originally trained on). While you appear to believe that \"using a model trained on ImageNet\" by extension involves \"using ImageNet\".\n\nI guess we don't know how the competition organizers will interpret this wording. Perhaps it is vague on purpose."
        },
        {
          "id": 767034,
          "postDate": "2020-03-09T04:13:53.280Z",
          "content": "<p>FF++ also makes pretrained models available: <a href=\"https://github.com/ondyari/FaceForensics/tree/master/classification\">https://github.com/ondyari/FaceForensics/tree/master/classification</a></p>\n\n<p>If ImageNet pretrained models are allowed (as they have always been on Kaggle), are FF++ pretrained models allowed? This is unclear.</p>",
          "rawMarkdown": "FF++ also makes pretrained models available: https://github.com/ondyari/FaceForensics/tree/master/classification\n\nIf ImageNet pretrained models are allowed (as they have always been on Kaggle), are FF++ pretrained models allowed? This is unclear."
        },
        {
          "id": 767212,
          "postDate": "2020-03-09T10:05:20.050Z",
          "content": "<p>Since the creators of that FaceForensics repo do not specifically state that you <em>can</em> use their pretrained models for non-research purposes, you'll have to assume that you can't. You can also ask them, of course. (This point has actually been clarified by the competition organizer: if there is no license, or the license is unclear, you cannot use it.)</p>",
          "rawMarkdown": "Since the creators of that FaceForensics repo do not specifically state that you *can* use their pretrained models for non-research purposes, you'll have to assume that you can't. You can also ask them, of course. (This point has actually been clarified by the competition organizer: if there is no license, or the license is unclear, you cannot use it.)"
        },
        {
          "id": 767282,
          "postDate": "2020-03-09T12:40:00.147Z",
          "content": "<p>&gt; the creators of that FaceForensics repo do not specifically state that you can use their pretrained models</p>\n\n<p>Few do: </p>\n\n<ul>\n<li><a href=\"https://discuss.pytorch.org/t/pre-trained-models-license/38647/3\">https://discuss.pytorch.org/t/pre-trained-models-license/38647/3</a></li>\n<li><a href=\"https://github.com/tensorflow/tensorflow/issues/27534\">https://github.com/tensorflow/tensorflow/issues/27534</a></li>\n</ul>",
          "rawMarkdown": "&gt; the creators of that FaceForensics repo do not specifically state that you can use their pretrained models\n\nFew do: \n\n* https://discuss.pytorch.org/t/pre-trained-models-license/38647/3\n* https://github.com/tensorflow/tensorflow/issues/27534\n\n"
        }
      ]
    },
    {
      "id": 761314,
      "postDate": "2020-03-02T11:35:38.043Z",
      "content": "<p>Well, the system is broken ... all the Non-Commercial / EULAs are protected under educational laws that makes it ok if you use copyrighted material as long as its for research and what not (imagenet is basically a massive web crawl .. so of course it contains copyrighted images) ... however the recent supreme court decision in the US for the landmark case of Writer's Guild vs. Google makes it ok ... and all of this is meaningless given the terms of this competition ... and to be frank, Kaggle can't police each and every researcher out of 2000 ... they stated their terms and you can contact the authors of the datasets you want to use to clarify their EULAs ... It's a weird time in ML TBH .. with Big Tech patenting stuff like dropout and the likes ... I guess if anyone goes to the podium will have to go through due diligence with the legal guys to clear their work 👀  </p>",
      "rawMarkdown": "Well, the system is broken ... all the Non-Commercial / EULAs are protected under educational laws that makes it ok if you use copyrighted material as long as its for research and what not (imagenet is basically a massive web crawl .. so of course it contains copyrighted images) ... however the recent supreme court decision in the US for the landmark case of Writer's Guild vs. Google makes it ok ... and all of this is meaningless given the terms of this competition ... and to be frank, Kaggle can't police each and every researcher out of 2000 ... they stated their terms and you can contact the authors of the datasets you want to use to clarify their EULAs ... It's a weird time in ML TBH .. with Big Tech patenting stuff like dropout and the likes ... I guess if anyone goes to the podium will have to go through due diligence with the legal guys to clear their work 👀  ",
      "votes": 5,
      "replies": [
        {
          "id": 761785,
          "postDate": "2020-03-02T23:39:29.100Z",
          "content": "<blockquote>\n  <p><strong>mjsML wrote:</strong></p>\n  \n  <p>Kaggle can't police each and every researcher out of 2000 </p>\n</blockquote>\n\n<p>They do verify the top teams' solutions (Or at least that was my experience with the TSA competition).</p>\n\n<p>I'm asking about just these 3 datasets that almost everyone seems to be using. Kaggle will have to make a judgement on them anyway.</p>",
          "rawMarkdown": "&gt; **mjsML wrote:**\n&gt; \n&gt; Kaggle can't police each and every researcher out of 2000 \n\nThey do verify the top teams' solutions (Or at least that was my experience with the TSA competition).\n\nI'm asking about just these 3 datasets that almost everyone seems to be using. Kaggle will have to make a judgement on them anyway.\n"
        },
        {
          "id": 765346,
          "postDate": "2020-03-06T14:14:33.207Z",
          "content": "<p><a href=\"/olegtrott\">@olegtrott</a> if you go through the DFDC EULA and the terms you'd see that it's quite well layed out ... they will audit everything very careful so I would not worry about that 😃 </p>",
          "rawMarkdown": "@olegtrott if you go through the DFDC EULA and the terms you'd see that it's quite well layed out ... they will audit everything very careful so I would not worry about that 😃 "
        }
      ]
    },
    {
      "id": 761290,
      "postDate": "2020-03-02T11:08:19.570Z",
      "content": "<p>This paragraph is the only reference to external data in the competition rules, and it doesn't mention any issue with non-commercial data.\n&gt; C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at <strong><em>no cost</em></strong> to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.</p>\n\n<p>But yeah, it would be very nice to hear it directly and clearly from the organizers.\nThanks for bringing it up again</p>",
      "rawMarkdown": "This paragraph is the only reference to external data in the competition rules, and it doesn't mention any issue with non-commercial data.\n&gt; C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at ***no cost*** to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.\n\nBut yeah, it would be very nice to hear it directly and clearly from the organizers.\nThanks for bringing it up again",
      "votes": 1
    },
    {
      "id": 761167,
      "postDate": "2020-03-02T07:53:42.933Z",
      "content": "<p>Just my take - \nThere is a <strong>no cost</strong> phrase in the conditions. There shouldn't be any issue as long as the data is available free of cost (or via a sign-up) ..and the dataset is declared publicly.</p>\n\n<p>*for purposes of the competition at <strong>no cost</strong> *</p>\n\n<p>Somewhat related. \nA <a href=\"https://www.extremetech.com/extreme/306575-new-tool-generates-every-possible-melody-for-public-domain-use\">pair of programmer-musicians</a> made public around 80mn melodies. Apparently they don't want musicians to be harassed by frivolous law-suits.   </p>\n\n<p>Link to <a href=\"https://github.com/allthemusicllc/atm-cli\">Github repo</a>\nLink to <a href=\"https://web.archive.org/web/20200301210148/https://archive.org/download/allthemusicllc-datasets\">Download the melodies</a></p>",
      "rawMarkdown": "Just my take - \nThere is a **no cost** phrase in the conditions. There shouldn't be any issue as long as the data is available free of cost (or via a sign-up) ..and the dataset is declared publicly.\n\n*for purposes of the competition at **no cost** *\n\nSomewhat related. \nA [pair of programmer-musicians](https://www.extremetech.com/extreme/306575-new-tool-generates-every-possible-melody-for-public-domain-use) made public around 80mn melodies. Apparently they don't want musicians to be harassed by frivolous law-suits.   \n\nLink to [Github repo](https://github.com/allthemusicllc/atm-cli)\nLink to [Download the melodies](https://web.archive.org/web/20200301210148/https://archive.org/download/allthemusicllc-datasets)",
      "votes": 1
    },
    {
      "id": 762846,
      "postDate": "2020-03-03T21:16:15.793Z",
      "content": "<p><a href=\"/cristiancanton\">@cristiancanton</a> <a href=\"/juliaelliott\">@juliaelliott</a> Can one of you please give a final answer on this so that this gets clear to everyone?</p>",
      "rawMarkdown": "@cristiancanton @juliaelliott Can one of you please give a final answer on this so that this gets clear to everyone?"
    },
    {
      "id": 761965,
      "postDate": "2020-03-03T04:40:57.193Z",
      "content": "<p>Nearly all the architectures use dropout which is patented. It is not really possible if patented things are banned for this competition.</p>",
      "rawMarkdown": "Nearly all the architectures use dropout which is patented. It is not really possible if patented things are banned for this competition."
    },
    {
      "id": 762658,
      "postDate": "2020-03-03T17:36:14.083Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 762900,
      "author_name": "Julia Elliott",
      "author_url": "",
      "post_date": "2020-03-03T23:29:22.583000",
      "content": "<p>I'm sorry that this continues to be unclear despite multiple responses I've made. I will make another final statement on the matter, in the hopes this puts it to rest. This position is in alignment with the host's wishes.</p>\n\n<p>Regurgitating the rules, the external data requirement is that any dataset you use (for training or otherwise that contributes to your solution/submission) must be:\n1. available to use by <strong>all participants</strong>\n2. for the purposes of the competition - more specifically meaning that the solution must be possible to license in accordance with the WINNER LICENSE requirements stipulated in the rules \"You hereby license and will license your winning Submission and the source code used to generate the Submission under an Open Source Initiative-approved license (see www.opensource.org) that in no event limits commercial use of such code or model containing or depending on such code...\"\n3. at no cost to the other participants\n4. posted to the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203\">Official External Data thread</a> by the deadline (March 3, 2020 @ 11:59pm UTC)</p>\n\n<p>I have made multiple clarifications that <strong>datasets with \"non-commercial only\" or \"academic use only\" or \"research use only\" provisions are prohibited</strong>, because they inherently break both:\n- #1: Use is limited to those who are capable of furnishing an academic/research affiliation, therefore not available for use by all participants.\n- #2: Use of a commercially-restricted dataset in your solution will render your solution unable to meet the host's licensing requirements, because it would limit commercial use of such code.</p>\n\n<p>We are not in a position to be posting a comprehensive list of every \"approved\" dataset.</p>\n\n<p>As far as enforceability, the host has the right to review solutions (in particular those in winning standing) and disqualify solutions which violate this rule.</p>",
      "votes": 11,
      "replies": [
        {
          "id": 762926,
          "author_name": "Oleg Trott",
          "author_url": "",
          "post_date": "2020-03-04T00:01:38.323000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> </p>\n\n<p>Thanks for replying.</p>\n\n<blockquote>\n  <p>for training or otherwise that contributes to your solution/submission</p>\n</blockquote>\n\n<p>In the interest of clarity, can you give a binary answer to whether models pretrained by third-parties on ImageNet are prohibited?</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 763121,
          "author_name": "AkiraSosa",
          "author_url": "",
          "post_date": "2020-03-04T06:45:58.533000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> \n&gt; 2: Use of a commercially-restricted dataset in your solution will render your solution unable to meet the host's licensing requirements, because it would limit commercial use of such code.</p>\n\n<p>I have a doubt with this statement. It lacks the consideration about the distinction between dataset itself and derived model. The interpretation about the license above can give the less chance for host to get a good model, e.g we miss ImageNet derived pre-trained model. That's the concern.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 763149,
          "author_name": "Nuno Ferreira",
          "author_url": "",
          "post_date": "2020-03-04T07:26:58.450000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Thank you for making this clear.</p>",
          "votes": -3,
          "replies": []
        },
        {
          "id": 763731,
          "author_name": "Selim Seferbekov",
          "author_url": "",
          "post_date": "2020-03-04T19:40:07.507000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> still it doesn't cover all of the cases</p>\n\n<p>Specifically a lot of competitors posted links to youtube.com and other websites that don't have dataset per se but can be scraped/collected. </p>\n\n<p>Therefore it is possible that :</p>\n\n<ul>\n<li>Competitor A: plays fair and safe and uses only allowed dataset (which basically means no external  deepfake data). Develops a good approach and models.</li>\n<li>Competitor B: uses a lot of data from youtube and other websites. Doesn't develop anything, let's say just uses some kernel with a lot of scraped data.</li>\n<li>I'm pretty sure Competitor B will win on the private set. </li>\n</ul>\n\n<p>I asked this question multiple times without any specific answer. Are competitors allowed to use data from youtube and other websites with CC license?</p>",
          "votes": 20,
          "replies": []
        },
        {
          "id": 764857,
          "author_name": "Oleg Trott",
          "author_url": "",
          "post_date": "2020-03-06T01:44:35.600000",
          "content": "<p>Another edge case: the media is public domain, but it's collected in a database licensed under a \"non-commercial\" license, as posted here: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203#748358\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203#748358</a> by <a href=\"/harangdev\">@harangdev</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 781620,
          "author_name": "Selim Seferbekov",
          "author_url": "",
          "post_date": "2020-03-21T13:59:14.450000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> any news? A lot of people reported youtube as their external data. Could you please clarify if that's legit? Just ignoring this question multiple times doesn't help at all. </p>",
          "votes": 9,
          "replies": []
        }
      ]
    },
    {
      "id": 761177,
      "author_name": "Selim Seferbekov",
      "author_url": "",
      "post_date": "2020-03-02T08:15:55.733000",
      "content": "<p><a href=\"/skylord\">@skylord</a> I would not be so sure.</p>\n\n<p>An answer from <a href=\"/juliaelliott\">@juliaelliott</a> regarding CC BY-NC</p>\n\n<blockquote>\n  <p>I’ve answered the question about BY-NC not being available for use by all (non-commercial use) and therefore violating the requirement that external data be available for use by all participants.</p>\n</blockquote>\n\n<p>Regarding FF+ etc the answer was also clear:</p>\n\n<blockquote>\n  <p>Edit: Also looking through all major fake videos datasets they all seem to be restricted for non-commercial research only: FaceForensics, FaceForensics++, DeepFakes Detection, both versions of Celeb-DF and still unpublished DeeperForensics-1.0. Is it ok to use them?</p>\n</blockquote>\n\n<p>An answer from <a href=\"/juliaelliott\">@juliaelliott</a> </p>\n\n<blockquote>\n  <p>Yes, if there are restrictions on a dataset’s use, it would violate the competition rules.</p>\n</blockquote>\n\n<p>Still it sounds weird that we are basically doing research but cannot use datasets for research only.</p>",
      "votes": 10,
      "replies": [
        {
          "id": 761315,
          "author_name": "mjsML",
          "author_url": "",
          "post_date": "2020-03-02T11:36:39.683000",
          "content": "<p>It seems the output of this will be used commercially in one way or the other ... </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 761821,
          "author_name": "Oleg Trott",
          "author_url": "",
          "post_date": "2020-03-03T00:43:59.413000",
          "content": "<blockquote>\n  <p><strong>Selim Seferbekov wrote:</strong></p>\n  \n  <p>An answer from <a href=\"/juliaelliott\">@juliaelliott</a> </p>\n  \n  <blockquote>\n    <p>Yes, if there are restrictions on a dataset’s use, it would violate the competition rules.</p>\n  </blockquote>\n</blockquote>\n\n<p>This would imply that ImageNet is forbidden, but everyone uses it, and Kaggle is OK with it.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 762204,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2020-03-03T10:19:37.110000",
          "content": "<blockquote>\n  <p>This would imply that ImageNet is forbidden, but everyone uses it, and Kaggle is OK with it.</p>\n</blockquote>\n\n<p>This is not entirely true. Everyone uses models that are pretrained on ImageNet, which is different from using ImageNet directly. (In addition, ImageNet is available on Kaggle these days.)</p>\n\n<p>I think the <em>spirit</em> of the External Data rules is that the playing field should be equal for all participants. FaceForensics++ is off limits because not everyone can download it. Open source models that are pretrained on ImageNet should be fine (if the license allows it).</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 762225,
          "author_name": "SkyLord",
          "author_url": "",
          "post_date": "2020-03-03T10:40:19.097000",
          "content": "<p>I thought that FF++ data can be downloaded after filling a google form ? They send you instructions for downloading the data. So the data is available for use. </p>\n\n<p>I think the organizers should list down all the datasets mentioned in the external data thread &amp; create a Yes/No table. This will remove the confusion. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762231,
          "author_name": "Oleg Trott",
          "author_url": "",
          "post_date": "2020-03-03T10:50:49.640000",
          "content": "<p><a href=\"/humananalog\">@humananalog</a> </p>\n\n<p>&gt; ImageNet, which is different from using ImageNet directly</p>\n\n<p>They are potentially different, hence the two separate questions in the original post.</p>\n\n<p>The reason I mention ImageNet is that it's analogous to the other two.</p>\n\n<p><a href=\"/skylord\">@skylord</a> </p>\n\n<p>It's impractical to go through all the things mentioned in that thread, but they should rule on the big three.</p>\n\n<p>Whether we are allowed to use FaceForensics++, or models trained on it is totally unclear.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762569,
          "author_name": "SkyLord",
          "author_url": "",
          "post_date": "2020-03-03T16:06:34.373000",
          "content": "<p>There was a bit of sarcasm in that....  🙂 <br>\nThere are around 350+ discussion threads, each with multiple links. Its not practical to read through each of them. </p>\n\n<p>But considering the <em>spirit of kaggle</em>  - which is open discussions, sharing of code kernels &amp; easily available datasets. I think as long as kagglers can <strong>access</strong> the data w/o any cost. (<em>freely available</em>) it should be ok. But again, it's just my reading! </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762604,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2020-03-03T16:28:15.157000",
          "content": "<blockquote>\n  <p>I thought that FF++ data can be downloaded after filling a google form ? They send you instructions for downloading the data. So the data is available for use.</p>\n</blockquote>\n\n<p>Do they accept everyone? And if not, are you allowed to give the data to that person, for example by uploading it as a dataset on Kaggle?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762646,
          "author_name": "SkyLord",
          "author_url": "",
          "post_date": "2020-03-03T17:16:12.167000",
          "content": "<p><a href=\"/humananalog\">@humananalog</a> \nWhile filling the google form, I wrote that I was taking part in the DFDC competition. within a day or two of filling it, I received an acceptance email. </p>\n\n<p>Edit: (a clarification)\nI believe you are referring to point 1, in the TOS. <a href=\"http://kaldir.vc.in.tum.de/FaceForensics/webpage/FaceForensics_TOS.pdf\">FaceForensics TOS</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762651,
          "author_name": "SkyLord",
          "author_url": "",
          "post_date": "2020-03-03T17:24:38.227000",
          "content": "<p>You should also check point 4 &amp;6 \n```\n4. Researcher may provide research associates and colleagues with access to the Database provided that they first agree\nto be bound by these terms and conditions.\n.\n.\n.</p>\n\n<ol>\n<li>If Researcher is employed by a for-profit, commercial entity, Researcher's employer shall also be bound by these terms\nand conditions, and Researcher hereby represents that he or she is fully authorized to enter into this agreement on behalf\nof such employer.\n```</li>\n</ol>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762656,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2020-03-03T17:34:13.757000",
          "content": "<p>This is a pretty clear answer: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203#744289\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203#744289</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762685,
          "author_name": "SkyLord",
          "author_url": "",
          "post_date": "2020-03-03T18:00:46.063000",
          "content": "<p>okay interesting!... I seem to have missed it. Thanks for pointing out. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762711,
          "author_name": "SkyLord",
          "author_url": "",
          "post_date": "2020-03-03T18:45:07.123000",
          "content": "<p>But then. How do we generalize to real world data ( like a deepfaked YouTube video) ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762714,
          "author_name": "Henrique Mendonça",
          "author_url": "",
          "post_date": "2020-03-03T18:47:59.720000",
          "content": "<blockquote>\n  <p>Likewise, the winners' licensing terms require that the solution be <strong>open sourced</strong> that that in no event limits commercial use of such <strong>code</strong> or model containing or depending on such <strong>code</strong>.</p>\n</blockquote>\n\n<p><a href=\"/humananalog\">@humananalog</a> images are not code, it explicitly says you need to provide the code to generate the submission, not the data. They don't even define submission in the rules. It's not clear if it even must include the training code.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 762888,
          "author_name": "Oleg Trott",
          "author_url": "",
          "post_date": "2020-03-03T22:52:13.230000",
          "content": "<p><a href=\"/humananalog\">@humananalog</a> </p>\n\n<blockquote>\n  <p>ImageNet is available for download on Kaggle.</p>\n</blockquote>\n\n<p>... for non-commercial use only.</p>\n\n<p>From <a href=\"https://www.kaggle.com/c/imagenet-object-localization-challenge/rules\">https://www.kaggle.com/c/imagenet-object-localization-challenge/rules</a>:</p>\n\n<blockquote>\n  <p>Participant shall use the Database only for non-commercial research and educational purposes</p>\n</blockquote>\n\n<p>If the organizers can use models trained by others on ImageNet, what policy provisions stop them from using models trained by others on, say, FaceForensics++ or InsightFace?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 762920,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2020-03-03T23:57:56.030000",
          "content": "<p><a href=\"/olegtrott\">@olegtrott</a> There are a few different things at play here:</p>\n\n<ol>\n<li>For some of these datasets the question is: can everyone use them? When I said ImageNet is available for download on Kaggle, that's what I was referring to. You don't need to fill out a form in order to download ImageNet, as opposed to FaceForensics++ for example.</li>\n<li>The other question is: can the dataset be used for non-commercial purposes? As you point out, ImageNet can't.</li>\n<li>However, it should be fine to use a model that is trained on ImageNet, if it is freely available to everyone and has a license that allows commercial usage. This is not the same thing as training on ImageNet directly.</li>\n</ol>\n\n<blockquote>\n  <p>If the organizers can use models trained by others on ImageNet, what policy provisions stop them from using models trained by others on, say, FaceForensics++ or InsightFace?</p>\n</blockquote>\n\n<p>As I understand the rules, a dataset such as FaceForensics++ is off-limits for this competition because it violates points 1 and 2 above. So any submissions using that dataset would be disqualified.</p>\n\n<p>However, if someone has trained a model on FaceForensics++, and makes it available under a license that is compatible with the competition, to everyone in the competition, then anyone will be able to use such a model in their own submissions (at least, if they had done so before the external data deadline).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762936,
          "author_name": "Oleg Trott",
          "author_url": "",
          "post_date": "2020-03-04T00:26:01.493000",
          "content": "<p><a href=\"/humananalog\">@humananalog</a> </p>\n\n<p>Kaggle does not appear to make a distinction between direct use of data and indirect one (see Julia's comment here). </p>\n\n<blockquote>\n  <p>any dataset you use (for training or <strong><em>OTHERWISE</em></strong> that contributes to your solution/submission) must be</p>\n</blockquote>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 763306,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2020-03-04T10:49:39.447000",
          "content": "<p><a href=\"/olegtrott\">@olegtrott</a> That's a very wide interpretation of the rules. I think our disagreement comes from the fact that I say \"using a model trained on a dataset\" is totally unrelated to the original dataset, so under this interpretation \"any dataset you use\" does not mean you're using ImageNet if you're using a model trained on ImageNet (i.e it only covers the model not the dataset is was originally trained on). While you appear to believe that \"using a model trained on ImageNet\" by extension involves \"using ImageNet\".</p>\n\n<p>I guess we don't know how the competition organizers will interpret this wording. Perhaps it is vague on purpose.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 767034,
          "author_name": "Oleg Trott",
          "author_url": "",
          "post_date": "2020-03-09T04:13:53.280000",
          "content": "<p>FF++ also makes pretrained models available: <a href=\"https://github.com/ondyari/FaceForensics/tree/master/classification\">https://github.com/ondyari/FaceForensics/tree/master/classification</a></p>\n\n<p>If ImageNet pretrained models are allowed (as they have always been on Kaggle), are FF++ pretrained models allowed? This is unclear.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 767212,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2020-03-09T10:05:20.050000",
          "content": "<p>Since the creators of that FaceForensics repo do not specifically state that you <em>can</em> use their pretrained models for non-research purposes, you'll have to assume that you can't. You can also ask them, of course. (This point has actually been clarified by the competition organizer: if there is no license, or the license is unclear, you cannot use it.)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 767282,
          "author_name": "Oleg Trott",
          "author_url": "",
          "post_date": "2020-03-09T12:40:00.147000",
          "content": "<p>&gt; the creators of that FaceForensics repo do not specifically state that you can use their pretrained models</p>\n\n<p>Few do: </p>\n\n<ul>\n<li><a href=\"https://discuss.pytorch.org/t/pre-trained-models-license/38647/3\">https://discuss.pytorch.org/t/pre-trained-models-license/38647/3</a></li>\n<li><a href=\"https://github.com/tensorflow/tensorflow/issues/27534\">https://github.com/tensorflow/tensorflow/issues/27534</a></li>\n</ul>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 761314,
      "author_name": "mjsML",
      "author_url": "",
      "post_date": "2020-03-02T11:35:38.043000",
      "content": "<p>Well, the system is broken ... all the Non-Commercial / EULAs are protected under educational laws that makes it ok if you use copyrighted material as long as its for research and what not (imagenet is basically a massive web crawl .. so of course it contains copyrighted images) ... however the recent supreme court decision in the US for the landmark case of Writer's Guild vs. Google makes it ok ... and all of this is meaningless given the terms of this competition ... and to be frank, Kaggle can't police each and every researcher out of 2000 ... they stated their terms and you can contact the authors of the datasets you want to use to clarify their EULAs ... It's a weird time in ML TBH .. with Big Tech patenting stuff like dropout and the likes ... I guess if anyone goes to the podium will have to go through due diligence with the legal guys to clear their work 👀  </p>",
      "votes": 5,
      "replies": [
        {
          "id": 761785,
          "author_name": "Oleg Trott",
          "author_url": "",
          "post_date": "2020-03-02T23:39:29.100000",
          "content": "<blockquote>\n  <p><strong>mjsML wrote:</strong></p>\n  \n  <p>Kaggle can't police each and every researcher out of 2000 </p>\n</blockquote>\n\n<p>They do verify the top teams' solutions (Or at least that was my experience with the TSA competition).</p>\n\n<p>I'm asking about just these 3 datasets that almost everyone seems to be using. Kaggle will have to make a judgement on them anyway.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 765346,
          "author_name": "mjsML",
          "author_url": "",
          "post_date": "2020-03-06T14:14:33.207000",
          "content": "<p><a href=\"/olegtrott\">@olegtrott</a> if you go through the DFDC EULA and the terms you'd see that it's quite well layed out ... they will audit everything very careful so I would not worry about that 😃 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 761290,
      "author_name": "Henrique Mendonça",
      "author_url": "",
      "post_date": "2020-03-02T11:08:19.570000",
      "content": "<p>This paragraph is the only reference to external data in the competition rules, and it doesn't mention any issue with non-commercial data.\n&gt; C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at <strong><em>no cost</em></strong> to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.</p>\n\n<p>But yeah, it would be very nice to hear it directly and clearly from the organizers.\nThanks for bringing it up again</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 761167,
      "author_name": "SkyLord",
      "author_url": "",
      "post_date": "2020-03-02T07:53:42.933000",
      "content": "<p>Just my take - \nThere is a <strong>no cost</strong> phrase in the conditions. There shouldn't be any issue as long as the data is available free of cost (or via a sign-up) ..and the dataset is declared publicly.</p>\n\n<p>*for purposes of the competition at <strong>no cost</strong> *</p>\n\n<p>Somewhat related. \nA <a href=\"https://www.extremetech.com/extreme/306575-new-tool-generates-every-possible-melody-for-public-domain-use\">pair of programmer-musicians</a> made public around 80mn melodies. Apparently they don't want musicians to be harassed by frivolous law-suits.   </p>\n\n<p>Link to <a href=\"https://github.com/allthemusicllc/atm-cli\">Github repo</a>\nLink to <a href=\"https://web.archive.org/web/20200301210148/https://archive.org/download/allthemusicllc-datasets\">Download the melodies</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 762846,
      "author_name": "Nuno Ferreira",
      "author_url": "",
      "post_date": "2020-03-03T21:16:15.793000",
      "content": "<p><a href=\"/cristiancanton\">@cristiancanton</a> <a href=\"/juliaelliott\">@juliaelliott</a> Can one of you please give a final answer on this so that this gets clear to everyone?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 761965,
      "author_name": "Shangqiu Li",
      "author_url": "",
      "post_date": "2020-03-03T04:40:57.193000",
      "content": "<p>Nearly all the architectures use dropout which is patented. It is not really possible if patented things are banned for this competition.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 762658,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-03T17:36:14.083000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "761009": "... and then enforce them accordingly.\n\nThe rules say:\n\n&gt; C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is ***available*** to use by all participants of the competition ***for purposes of the competition*** at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.\n\nHowever, it's not at all clear what \"available for the purposes of the competition\" means.\n\n- ImageNet\n- FaceForensics++ dataset and \n- InsightFace (RetinaFace, ArcFace) annotations *and* pretrained models.\n\nare all available for non-commercial/research use only. FaceForensics++ additionally requires a sign-up.\n\n- Are we allowed to use them directly?\n\n- Are we allowed to use third-party models trained on the above?\n\n@juliaelliott wrote earlier:\n\n&gt;@vladislavleketush and others - Your use of external data should conform to the external data specification in the competition's rules:\n\n&gt;&gt;you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.\n\n&gt;Therefore, licenses which place restrictions on datasets' use (whether by purpose, affiliation, cost or other restrictive means) would be in violation of this rule.\n\nBut I don't think this answers these questions.\n\nIs our purpose \"commercial\" rather than \"research\" here? (If so, what about all of the patents?!)\n\nBy the way, pretrained models might not be restricted by the license of the data on which they are trained:  https://towardsdatascience.com/the-most-important-supreme-court-decision-for-data-science-and-machine-learning-44cfc1c1bcaf\n\nWe need clarity here, soon.\n",
    "762900": "I'm sorry that this continues to be unclear despite multiple responses I've made. I will make another final statement on the matter, in the hopes this puts it to rest. This position is in alignment with the host's wishes.\n\nRegurgitating the rules, the external data requirement is that any dataset you use (for training or otherwise that contributes to your solution/submission) must be:\n1. available to use by **all participants**\n2. for the purposes of the competition - more specifically meaning that the solution must be possible to license in accordance with the WINNER LICENSE requirements stipulated in the rules \"You hereby license and will license your winning Submission and the source code used to generate the Submission under an Open Source Initiative-approved license (see www.opensource.org) that in no event limits commercial use of such code or model containing or depending on such code...\"\n3. at no cost to the other participants\n4. posted to the [Official External Data thread](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203) by the deadline (March 3, 2020 @ 11:59pm UTC)\n\nI have made multiple clarifications that **datasets with \"non-commercial only\" or \"academic use only\" or \"research use only\" provisions are prohibited**, because they inherently break both:\n- #1: Use is limited to those who are capable of furnishing an academic/research affiliation, therefore not available for use by all participants.\n- #2: Use of a commercially-restricted dataset in your solution will render your solution unable to meet the host's licensing requirements, because it would limit commercial use of such code.\n\nWe are not in a position to be posting a comprehensive list of every \"approved\" dataset.\n\nAs far as enforceability, the host has the right to review solutions (in particular those in winning standing) and disqualify solutions which violate this rule.",
    "761177": "@skylord I would not be so sure.\n\nAn answer from @juliaelliott regarding CC BY-NC\n&gt; I’ve answered the question about BY-NC not being available for use by all (non-commercial use) and therefore violating the requirement that external data be available for use by all participants.\n\nRegarding FF+ etc the answer was also clear:\n&gt; Edit: Also looking through all major fake videos datasets they all seem to be restricted for non-commercial research only: FaceForensics, FaceForensics++, DeepFakes Detection, both versions of Celeb-DF and still unpublished DeeperForensics-1.0. Is it ok to use them?\n\nAn answer from @juliaelliott \n&gt; Yes, if there are restrictions on a dataset’s use, it would violate the competition rules.\n\nStill it sounds weird that we are basically doing research but cannot use datasets for research only.\n",
    "761314": "Well, the system is broken ... all the Non-Commercial / EULAs are protected under educational laws that makes it ok if you use copyrighted material as long as its for research and what not (imagenet is basically a massive web crawl .. so of course it contains copyrighted images) ... however the recent supreme court decision in the US for the landmark case of Writer's Guild vs. Google makes it ok ... and all of this is meaningless given the terms of this competition ... and to be frank, Kaggle can't police each and every researcher out of 2000 ... they stated their terms and you can contact the authors of the datasets you want to use to clarify their EULAs ... It's a weird time in ML TBH .. with Big Tech patenting stuff like dropout and the likes ... I guess if anyone goes to the podium will have to go through due diligence with the legal guys to clear their work 👀  ",
    "761290": "This paragraph is the only reference to external data in the competition rules, and it doesn't mention any issue with non-commercial data.\n&gt; C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at ***no cost*** to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.\n\nBut yeah, it would be very nice to hear it directly and clearly from the organizers.\nThanks for bringing it up again",
    "761167": "Just my take - \nThere is a **no cost** phrase in the conditions. There shouldn't be any issue as long as the data is available free of cost (or via a sign-up) ..and the dataset is declared publicly.\n\n*for purposes of the competition at **no cost** *\n\nSomewhat related. \nA [pair of programmer-musicians](https://www.extremetech.com/extreme/306575-new-tool-generates-every-possible-melody-for-public-domain-use) made public around 80mn melodies. Apparently they don't want musicians to be harassed by frivolous law-suits.   \n\nLink to [Github repo](https://github.com/allthemusicllc/atm-cli)\nLink to [Download the melodies](https://web.archive.org/web/20200301210148/https://archive.org/download/allthemusicllc-datasets)",
    "762846": "@cristiancanton @juliaelliott Can one of you please give a final answer on this so that this gets clear to everyone?",
    "761965": "Nearly all the architectures use dropout which is patented. It is not really possible if patented things are banned for this competition.",
    "762658": ""
  }
}