{
  "id": 158244,
  "title": "Question about the logic used to decide what is valid or not",
  "url": "/competitions/deepfake-detection-challenge/discussion/158244",
  "author_name": "",
  "post_date": "2020-06-13T14:46:14.996516Z",
  "votes": 28,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Kaggle and Facebook:</p>\n\n<p>Could you explain your logic for deciding when a dataset needs individual permission?\nBecause there doesn't seem to be any. And I think this is important for all Kagglers, including for future competitions.</p>\n\n<p>As @aakashnain  mentioned here (<a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983</a>), there are several points to consider, which following your decision all\ncompetitors should be eliminated.\nEvery computer vision competition is common to use pre-trained models in other datasets and here wasn't different.\nBoth models trained on ImageNet and the MTCNN face detector (used a lot here) also have data that doesn't follow the standards you want.</p>\n\n<p>So I have a few questions:</p>\n\n<p>1 - Why is this rule being applied only in some cases?</p>\n\n<p>2 - If you had this concern at the beginning of the competition, why allow external datasets?\nIf you didn't, wouldn't that be a mistake from you that you are passing on to the competitors?</p>\n\n<p>3 - how should we behave in future situations?</p>\n\n<p>I think these questions are important to answer since competitors spend hundreds of hours working on something without any guarantees.</p>",
  "messages": [
    {
      "id": "884678",
      "postDate": "06/13/2020 14:46:14",
      "content": "<p>Kaggle and Facebook:</p>\n\n<p>Could you explain your logic for deciding when a dataset needs individual permission?\nBecause there doesn't seem to be any. And I think this is important for all Kagglers, including for future competitions.</p>\n\n<p>As @aakashnain  mentioned here (<a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983</a>), there are several points to consider, which following your decision all\ncompetitors should be eliminated.\nEvery computer vision competition is common to use pre-trained models in other datasets and here wasn't different.\nBoth models trained on ImageNet and the MTCNN face detector (used a lot here) also have data that doesn't follow the standards you want.</p>\n\n<p>So I have a few questions:</p>\n\n<p>1 - Why is this rule being applied only in some cases?</p>\n\n<p>2 - If you had this concern at the beginning of the competition, why allow external datasets?\nIf you didn't, wouldn't that be a mistake from you that you are passing on to the competitors?</p>\n\n<p>3 - how should we behave in future situations?</p>\n\n<p>I think these questions are important to answer since competitors spend hundreds of hours working on something without any guarantees.</p>",
      "rawMarkdown": "Kaggle and Facebook:\n\nCould you explain your logic for deciding when a dataset needs individual permission?\nBecause there doesn't seem to be any. And I think this is important for all Kagglers, including for future competitions.\n\nAs @aakashnain  mentioned here (https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983), there are several points to consider, which following your decision all\ncompetitors should be eliminated.\nEvery computer vision competition is common to use pre-trained models in other datasets and here wasn't different.\nBoth models trained on ImageNet and the MTCNN face detector (used a lot here) also have data that doesn't follow the standards you want.\n\nSo I have a few questions:\n\n1 - Why is this rule being applied only in some cases?\n\n2 - If you had this concern at the beginning of the competition, why allow external datasets?\nIf you didn't, wouldn't that be a mistake from you that you are passing on to the competitors?\n\n3 - how should we behave in future situations?\n\nI think these questions are important to answer since competitors spend hundreds of hours working on something without any guarantees.",
      "votes": null
    },
    {
      "id": "884685",
      "postDate": "06/13/2020 14:55:38",
      "content": "<p>I second you, Its really sad to work for hours and then they say we did not want it, I believe there should be a guarantee for the work competitors put here. Kaggle and Facebook both need to answer this, otherwise, should all the competitors depend on the decision you make over your own mistake?</p>",
      "rawMarkdown": "I second you, Its really sad to work for hours and then they say we did not want it, I believe there should be a guarantee for the work competitors put here. Kaggle and Facebook both need to answer this, otherwise, should all the competitors depend on the decision you make over your own mistake?",
      "votes": null
    },
    {
      "id": "888517",
      "postDate": "06/16/2020 12:15:21",
      "content": "<p>Unfortunately, it seems we won't have an answer</p>",
      "rawMarkdown": "Unfortunately, it seems we won't have an answer",
      "votes": null
    },
    {
      "id": "888532",
      "postDate": "06/16/2020 12:23:53",
      "content": "<p>Yep, I saw one comment by Julia, she said they will post something about this issue, Iet me find that for you.</p>\n\n<p><a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983#887946\">Here</a></p>\n\n<p>Let's wait until we get one.</p>",
      "rawMarkdown": "Yep, I saw one comment by Julia, she said they will post something about this issue, Iet me find that for you.\n\n[Here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983#887946)\n\nLet's wait until we get one.",
      "votes": null
    },
    {
      "id": "889086",
      "postDate": "06/16/2020 18:50:37",
      "content": "<p>We are unable to provide more pointed guidance on the topic of pre-trained models because licensing issues surrounding pre-trained models can vary on a case-by-case basis. A quick search on the topic will give varied opinions on the interplay between a model and the contents on which it was trained. As a result, we leave judgments as to which licenses are acceptable to the hosts in the hands of the hosts, and universal statements on pre-trained model eligibility cannot be made. </p>\n\n<p>In this competition, pre-trained models licensed for commercial use were stated as permitted and no team was disqualified for using them. In future competitions, we will urge hosts to clarify how they expect pretrained models to be licensed, and if there are any expected deviations from their competition’s general rules on external data use.</p>\n\n<p>Our clarification and response regarding the disqualification that was made can be read in <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983#889085\">this other thread</a>.</p>",
      "rawMarkdown": "We are unable to provide more pointed guidance on the topic of pre-trained models because licensing issues surrounding pre-trained models can vary on a case-by-case basis. A quick search on the topic will give varied opinions on the interplay between a model and the contents on which it was trained. As a result, we leave judgments as to which licenses are acceptable to the hosts in the hands of the hosts, and universal statements on pre-trained model eligibility cannot be made. \n\nIn this competition, pre-trained models licensed for commercial use were stated as permitted and no team was disqualified for using them. In future competitions, we will urge hosts to clarify how they expect pretrained models to be licensed, and if there are any expected deviations from their competition’s general rules on external data use.\n\nOur clarification and response regarding the disqualification that was made can be read in [this other thread](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983#889085).",
      "votes": null
    },
    {
      "id": "889968",
      "postDate": "06/17/2020 08:26:02",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> I have two points to make regarding this statement <code>As a result, we leave judgments as to which licenses are acceptable to the hosts in the hands of the hosts, and universal statements on pre-trained model eligibility cannot be made.</code></p>\n\n<ol>\n<li>Why ImageNet/MTCNN is permitted in the first place if the datasets used to train them doesn't comply with the licensing required by the host? The host should pick a side. It is totally misleading to say that we can selectively pick things for usage and make it compliant with the licensing. The choice is binary. </li>\n<li>If you want to avoid such situations in future, then I suggest complete ban on using pretrained weights and external data. By no means, any researcher, ML engineer, or a data scientist can verify every other thing. Also, the host didn't answer our question: Why is it okay to use ImageNet/COCO, etc if we don't have the consent from the source for each sample?</li>\n</ol>",
      "rawMarkdown": "juliaelliott I have two points to make regarding this statement `As a result, we leave judgments as to which licenses are acceptable to the hosts in the hands of the hosts, and universal statements on pre-trained model eligibility cannot be made. `\n\n1. Why ImageNet/MTCNN is permitted in the first place if the datasets used to train them doesn't comply with the licensing required by the host? The host should pick a side. It is totally misleading to say that we can selectively pick things for usage and make it compliant with the licensing. The choice is binary. \n2. If you want to avoid such situations in future, then I suggest complete ban on using pretrained weights and external data. By no means, any researcher, ML engineer, or a data scientist can verify every other thing. Also, the host didn't answer our question: Why is it okay to use ImageNet/COCO, etc if we don't have the consent from the source for each sample?",
      "votes": null
    },
    {
      "id": "890169",
      "postDate": "06/17/2020 11:03:38",
      "content": "<blockquote>\n  <p>Why is it okay to use ImageNet/COCO, etc if we don't have the consent from the source for each sample?</p>\n</blockquote>\n\n<p>There is an important difference between \"using pre-trained models trained on dataset X\" and \"using dataset X to train your own models\". Kaggle seems to allow the former, probably because everyone is doing that and there are no legal rulings that say you can't.</p>\n\n<p>Whether the latter is allowed depends on the license terms of that dataset. ImageNet can only be used for non-commercial purposes, IIRC.</p>\n\n<p>So you <em>can</em> use a model that was pre-trained on ImageNet but you can't train your own model on ImageNet.</p>\n\n<p>(I guess this particular interpretation would mean that the creator of, say, MTCNN can't actually use their own model in a Kaggle competition because they had to agree to the license terms of the original dataset. But anyone else who is using MTCNN doesn't have to follow these license terms, because they're not using the original dataset. Yes, this sounds weird. I am not a lawyer and my interpretation may be wrong.)</p>",
      "rawMarkdown": "&gt; Why is it okay to use ImageNet/COCO, etc if we don't have the consent from the source for each sample?\n\nThere is an important difference between \"using pre-trained models trained on dataset X\" and \"using dataset X to train your own models\". Kaggle seems to allow the former, probably because everyone is doing that and there are no legal rulings that say you can't.\n\nWhether the latter is allowed depends on the license terms of that dataset. ImageNet can only be used for non-commercial purposes, IIRC.\n\nSo you *can* use a model that was pre-trained on ImageNet but you can't train your own model on ImageNet.\n\n(I guess this particular interpretation would mean that the creator of, say, MTCNN can't actually use their own model in a Kaggle competition because they had to agree to the license terms of the original dataset. But anyone else who is using MTCNN doesn't have to follow these license terms, because they're not using the original dataset. Yes, this sounds weird. I am not a lawyer and my interpretation may be wrong.)",
      "votes": null
    },
    {
      "id": "890501",
      "postDate": "06/17/2020 14:34:50",
      "content": "<p>I interpret it slightly different. Was the model using the \"mislicensed\" data available for everyone to use. If it were it would fall under the same realms as ImageNet and MTCNN.</p>\n\n<p>The reality is they went with the majority. Instead of four teams challenging the decision they have one to deal with. Either way, doesn't make it right. The lack of empathy shown from the host and sponsor will serve to aggravate and alienate their users.</p>\n\n<p>In these situations great companies find a middle ground and compromise. Allowing them to keep their 7th place submission as a compromise is embarrassing to ever Kaggle contestant now and in the future.</p>\n\n<p>Put yourself in their shoes. How would you feel if you knowingly did not cheat and all of a sudden that $500K was replaced with $0 + thousands of hours of your time wasted + your bills + a 7th place standing? I know I would be absolutely depressed!</p>\n\n<p>This could happen to you and anyone else using this platform in the future. I am certain they will come to their senses and find a middle ground.</p>",
      "rawMarkdown": "I interpret it slightly different. Was the model using the \"mislicensed\" data available for everyone to use. If it were it would fall under the same realms as ImageNet and MTCNN.\n\nThe reality is they went with the majority. Instead of four teams challenging the decision they have one to deal with. Either way, doesn't make it right. The lack of empathy shown from the host and sponsor will serve to aggravate and alienate their users.\n\nIn these situations great companies find a middle ground and compromise. Allowing them to keep their 7th place submission as a compromise is embarrassing to ever Kaggle contestant now and in the future.\n\nPut yourself in their shoes. How would you feel if you knowingly did not cheat and all of a sudden that $500K was replaced with $0 + thousands of hours of your time wasted + your bills + a 7th place standing? I know I would be absolutely depressed!\n\nThis could happen to you and anyone else using this platform in the future. I am certain they will come to their senses and find a middle ground.",
      "votes": null
    },
    {
      "id": "890536",
      "postDate": "06/17/2020 14:52:37",
      "content": "<blockquote>\n  <p>This could happen to you and anyone else using this platform in the future.</p>\n</blockquote>\n\n<p>I wasn't saying I agreed with the decision. ;-) Just that equating the use of YouTube videos with the use of ImageNet pre-trained models is not an appropriate comparison.</p>",
      "rawMarkdown": "&gt; This could happen to you and anyone else using this platform in the future.\n\nI wasn't saying I agreed with the decision. ;-) Just that equating the use of YouTube videos with the use of ImageNet pre-trained models is not an appropriate comparison.",
      "votes": null
    },
    {
      "id": "891019",
      "postDate": "06/17/2020 20:42:29",
      "content": "<p><a href=\"/humananalog\">@humananalog</a> under your logic though, if the the youtube (or other external data) could be trained into a model, released under a permissive open-source license on github, pip, etc and then used as part of the end solution, it would suddenly be okay. That makes no sense.</p>",
      "rawMarkdown": "humananalog under your logic though, if the the youtube (or other external data) could be trained into a model, released under a permissive open-source license on github, pip, etc and then used as part of the end solution, it would suddenly be okay. That makes no sense.",
      "votes": null
    },
    {
      "id": "891617",
      "postDate": "06/18/2020 10:14:34",
      "content": "<p><a href=\"/rwightman\">@rwightman</a> Yes, if the logic I outlined was followed that would be the conclusion <em>if</em> we're just talking about the legal aspects of it (disclaimer: I am not a lawyer, laws vary depending on where you live).</p>\n\n<p>That said, I think this is legally a gray area. AFAIK, training a model on a dataset does not make the model a derivative work of the dataset, so the rights owners of the dataset have no say in how the model can be used. It's a legally gray area because maybe it <em>should</em> be a derivative work but no one has challenged this in court yet (to my knowledge).</p>\n\n<p>The lawyers involved in this competition probably think using ImageNet-based models is fine, even if the courts at some point decide that models indeed are derivative works of their datasets, because using ImageNet-based models is a widespread practice already and that would probably be exempt of such a ruling. (Of course, I'm speculating here.)</p>\n\n<p>For other datasets and models based on them, such as the hypothetical one you mentioned, the lawyers might play it safer and not allow those. Not because it would currently be illegal but it might be in the future.</p>\n\n<p>Of course, there is also the matter of ethical concerns. For this competition, apparently, a CC license is not enough: Unless you have explicit permission from everyone in your dataset, the dataset cannot be used. For ImageNet / COCO etc, this wouldn't apply because those datasets have been in widespread use for a long time now and it would be unreasonable to fix this retroactively. </p>\n\n<p>(Just a CC license would not be enough because people who licensed their data as CC probably didn't realize it might be used to train things like face detection algorithms, and might not actually agree to that kind of use. You can argue that it doesn't matter that they didn't give explicit consent for this, since the CC license allows you to do whatever you want with these images. Again, this is something that could be challenged in court but to my knowledge hasn't. That said, asking for explicit permission feels like the right thing to do since it's a touchy subject.)</p>",
      "rawMarkdown": "rwightman Yes, if the logic I outlined was followed that would be the conclusion *if* we're just talking about the legal aspects of it (disclaimer: I am not a lawyer, laws vary depending on where you live).\n\nThat said, I think this is legally a gray area. AFAIK, training a model on a dataset does not make the model a derivative work of the dataset, so the rights owners of the dataset have no say in how the model can be used. It's a legally gray area because maybe it *should* be a derivative work but no one has challenged this in court yet (to my knowledge).\n\nThe lawyers involved in this competition probably think using ImageNet-based models is fine, even if the courts at some point decide that models indeed are derivative works of their datasets, because using ImageNet-based models is a widespread practice already and that would probably be exempt of such a ruling. (Of course, I'm speculating here.)\n\nFor other datasets and models based on them, such as the hypothetical one you mentioned, the lawyers might play it safer and not allow those. Not because it would currently be illegal but it might be in the future.\n\nOf course, there is also the matter of ethical concerns. For this competition, apparently, a CC license is not enough: Unless you have explicit permission from everyone in your dataset, the dataset cannot be used. For ImageNet / COCO etc, this wouldn't apply because those datasets have been in widespread use for a long time now and it would be unreasonable to fix this retroactively. \n\n(Just a CC license would not be enough because people who licensed their data as CC probably didn't realize it might be used to train things like face detection algorithms, and might not actually agree to that kind of use. You can argue that it doesn't matter that they didn't give explicit consent for this, since the CC license allows you to do whatever you want with these images. Again, this is something that could be challenged in court but to my knowledge hasn't. That said, asking for explicit permission feels like the right thing to do since it's a touchy subject.)",
      "votes": null
    },
    {
      "id": "891633",
      "postDate": "06/18/2020 10:44:49",
      "content": "<p><a href=\"/humananalog\">@humananalog</a> this is a bit of a legal flip flop that is really going the extra mile to be cautious. Facebook had \"organic\" deepfakes in their hidden dataset. Do you think they got consent for using these from the people appearing in it? I doubt it. The allow that as it is a \"hidden\" thing. <br>\nBesides, this logic means that training a model on imagenet would make it illegal, or even using a subset of it. same with coco and open images. I seriously doubt that google went and got consent from anyone appearing in their 1.5M scraped images dataset, and we produced models based on this in a competition. They trust the flickr license.\nAnyways, beside this point, I think that in such a case, with such language, hidden so well under documentation clause, the winning team should have gotten a fair chance to \"fix\" their model and get a new submission with the offending videos removed. It rubs against the grain of most people here how a team that worked so hard can be disqualified without proper explanation and without being given a chance to either fix the problem or just not get prize money and keep the 1st place. I have seen a bit of injustice being done here, but it was mostly the other way around, kaggle turning a blind eye on minor offences. This is something new and a bit worrying.</p>",
      "rawMarkdown": "humananalog this is a bit of a legal flip flop that is really going the extra mile to be cautious. Facebook had \"organic\" deepfakes in their hidden dataset. Do you think they got consent for using these from the people appearing in it? I doubt it. The allow that as it is a \"hidden\" thing.  \nBesides, this logic means that training a model on imagenet would make it illegal, or even using a subset of it. same with coco and open images. I seriously doubt that google went and got consent from anyone appearing in their 1.5M scraped images dataset, and we produced models based on this in a competition. They trust the flickr license.\nAnyways, beside this point, I think that in such a case, with such language, hidden so well under documentation clause, the winning team should have gotten a fair chance to \"fix\" their model and get a new submission with the offending videos removed. It rubs against the grain of most people here how a team that worked so hard can be disqualified without proper explanation and without being given a chance to either fix the problem or just not get prize money and keep the 1st place. I have seen a bit of injustice being done here, but it was mostly the other way around, kaggle turning a blind eye on minor offences. This is something new and a bit worrying.",
      "votes": null
    },
    {
      "id": "893514",
      "postDate": "06/19/2020 17:12:57",
      "content": "<p>The logic is: there is no logic.</p>",
      "rawMarkdown": "The logic is: there is no logic.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 884685,
      "author_name": "sarques",
      "author_url": "",
      "post_date": "06/13/2020 14:55:38",
      "content": "<p>I second you, Its really sad to work for hours and then they say we did not want it, I believe there should be a guarantee for the work competitors put here. Kaggle and Facebook both need to answer this, otherwise, should all the competitors depend on the decision you make over your own mistake?</p>",
      "votes": null,
      "replies": [
        {
          "id": 888517,
          "author_name": "igormunizims",
          "author_url": "",
          "post_date": "06/16/2020 12:15:21",
          "content": "<p>Unfortunately, it seems we won't have an answer</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 888532,
          "author_name": "sarques",
          "author_url": "",
          "post_date": "06/16/2020 12:23:53",
          "content": "<p>Yep, I saw one comment by Julia, she said they will post something about this issue, Iet me find that for you.</p>\n\n<p><a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983#887946\">Here</a></p>\n\n<p>Let's wait until we get one.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 889086,
      "author_name": "juliaelliott",
      "author_url": "",
      "post_date": "06/16/2020 18:50:37",
      "content": "<p>We are unable to provide more pointed guidance on the topic of pre-trained models because licensing issues surrounding pre-trained models can vary on a case-by-case basis. A quick search on the topic will give varied opinions on the interplay between a model and the contents on which it was trained. As a result, we leave judgments as to which licenses are acceptable to the hosts in the hands of the hosts, and universal statements on pre-trained model eligibility cannot be made. </p>\n\n<p>In this competition, pre-trained models licensed for commercial use were stated as permitted and no team was disqualified for using them. In future competitions, we will urge hosts to clarify how they expect pretrained models to be licensed, and if there are any expected deviations from their competition’s general rules on external data use.</p>\n\n<p>Our clarification and response regarding the disqualification that was made can be read in <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983#889085\">this other thread</a>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 889968,
      "author_name": "aakashnain",
      "author_url": "",
      "post_date": "06/17/2020 08:26:02",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> I have two points to make regarding this statement <code>As a result, we leave judgments as to which licenses are acceptable to the hosts in the hands of the hosts, and universal statements on pre-trained model eligibility cannot be made.</code></p>\n\n<ol>\n<li>Why ImageNet/MTCNN is permitted in the first place if the datasets used to train them doesn't comply with the licensing required by the host? The host should pick a side. It is totally misleading to say that we can selectively pick things for usage and make it compliant with the licensing. The choice is binary. </li>\n<li>If you want to avoid such situations in future, then I suggest complete ban on using pretrained weights and external data. By no means, any researcher, ML engineer, or a data scientist can verify every other thing. Also, the host didn't answer our question: Why is it okay to use ImageNet/COCO, etc if we don't have the consent from the source for each sample?</li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 890169,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "06/17/2020 11:03:38",
          "content": "<blockquote>\n  <p>Why is it okay to use ImageNet/COCO, etc if we don't have the consent from the source for each sample?</p>\n</blockquote>\n\n<p>There is an important difference between \"using pre-trained models trained on dataset X\" and \"using dataset X to train your own models\". Kaggle seems to allow the former, probably because everyone is doing that and there are no legal rulings that say you can't.</p>\n\n<p>Whether the latter is allowed depends on the license terms of that dataset. ImageNet can only be used for non-commercial purposes, IIRC.</p>\n\n<p>So you <em>can</em> use a model that was pre-trained on ImageNet but you can't train your own model on ImageNet.</p>\n\n<p>(I guess this particular interpretation would mean that the creator of, say, MTCNN can't actually use their own model in a Kaggle competition because they had to agree to the license terms of the original dataset. But anyone else who is using MTCNN doesn't have to follow these license terms, because they're not using the original dataset. Yes, this sounds weird. I am not a lawyer and my interpretation may be wrong.)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 890501,
          "author_name": "maralski",
          "author_url": "",
          "post_date": "06/17/2020 14:34:50",
          "content": "<p>I interpret it slightly different. Was the model using the \"mislicensed\" data available for everyone to use. If it were it would fall under the same realms as ImageNet and MTCNN.</p>\n\n<p>The reality is they went with the majority. Instead of four teams challenging the decision they have one to deal with. Either way, doesn't make it right. The lack of empathy shown from the host and sponsor will serve to aggravate and alienate their users.</p>\n\n<p>In these situations great companies find a middle ground and compromise. Allowing them to keep their 7th place submission as a compromise is embarrassing to ever Kaggle contestant now and in the future.</p>\n\n<p>Put yourself in their shoes. How would you feel if you knowingly did not cheat and all of a sudden that $500K was replaced with $0 + thousands of hours of your time wasted + your bills + a 7th place standing? I know I would be absolutely depressed!</p>\n\n<p>This could happen to you and anyone else using this platform in the future. I am certain they will come to their senses and find a middle ground.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 890536,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "06/17/2020 14:52:37",
          "content": "<blockquote>\n  <p>This could happen to you and anyone else using this platform in the future.</p>\n</blockquote>\n\n<p>I wasn't saying I agreed with the decision. ;-) Just that equating the use of YouTube videos with the use of ImageNet pre-trained models is not an appropriate comparison.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 891019,
          "author_name": "rwightman",
          "author_url": "",
          "post_date": "06/17/2020 20:42:29",
          "content": "<p><a href=\"/humananalog\">@humananalog</a> under your logic though, if the the youtube (or other external data) could be trained into a model, released under a permissive open-source license on github, pip, etc and then used as part of the end solution, it would suddenly be okay. That makes no sense.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 891617,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "06/18/2020 10:14:34",
          "content": "<p><a href=\"/rwightman\">@rwightman</a> Yes, if the logic I outlined was followed that would be the conclusion <em>if</em> we're just talking about the legal aspects of it (disclaimer: I am not a lawyer, laws vary depending on where you live).</p>\n\n<p>That said, I think this is legally a gray area. AFAIK, training a model on a dataset does not make the model a derivative work of the dataset, so the rights owners of the dataset have no say in how the model can be used. It's a legally gray area because maybe it <em>should</em> be a derivative work but no one has challenged this in court yet (to my knowledge).</p>\n\n<p>The lawyers involved in this competition probably think using ImageNet-based models is fine, even if the courts at some point decide that models indeed are derivative works of their datasets, because using ImageNet-based models is a widespread practice already and that would probably be exempt of such a ruling. (Of course, I'm speculating here.)</p>\n\n<p>For other datasets and models based on them, such as the hypothetical one you mentioned, the lawyers might play it safer and not allow those. Not because it would currently be illegal but it might be in the future.</p>\n\n<p>Of course, there is also the matter of ethical concerns. For this competition, apparently, a CC license is not enough: Unless you have explicit permission from everyone in your dataset, the dataset cannot be used. For ImageNet / COCO etc, this wouldn't apply because those datasets have been in widespread use for a long time now and it would be unreasonable to fix this retroactively. </p>\n\n<p>(Just a CC license would not be enough because people who licensed their data as CC probably didn't realize it might be used to train things like face detection algorithms, and might not actually agree to that kind of use. You can argue that it doesn't matter that they didn't give explicit consent for this, since the CC license allows you to do whatever you want with these images. Again, this is something that could be challenged in court but to my knowledge hasn't. That said, asking for explicit permission feels like the right thing to do since it's a touchy subject.)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 891633,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "06/18/2020 10:44:49",
          "content": "<p><a href=\"/humananalog\">@humananalog</a> this is a bit of a legal flip flop that is really going the extra mile to be cautious. Facebook had \"organic\" deepfakes in their hidden dataset. Do you think they got consent for using these from the people appearing in it? I doubt it. The allow that as it is a \"hidden\" thing. <br>\nBesides, this logic means that training a model on imagenet would make it illegal, or even using a subset of it. same with coco and open images. I seriously doubt that google went and got consent from anyone appearing in their 1.5M scraped images dataset, and we produced models based on this in a competition. They trust the flickr license.\nAnyways, beside this point, I think that in such a case, with such language, hidden so well under documentation clause, the winning team should have gotten a fair chance to \"fix\" their model and get a new submission with the offending videos removed. It rubs against the grain of most people here how a team that worked so hard can be disqualified without proper explanation and without being given a chance to either fix the problem or just not get prize money and keep the 1st place. I have seen a bit of injustice being done here, but it was mostly the other way around, kaggle turning a blind eye on minor offences. This is something new and a bit worrying.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 893514,
      "author_name": "hiankovski",
      "author_url": "",
      "post_date": "06/19/2020 17:12:57",
      "content": "<p>The logic is: there is no logic.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "884678": "Kaggle and Facebook:\n\nCould you explain your logic for deciding when a dataset needs individual permission?\nBecause there doesn't seem to be any. And I think this is important for all Kagglers, including for future competitions.\n\nAs @aakashnain  mentioned here (https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983), there are several points to consider, which following your decision all\ncompetitors should be eliminated.\nEvery computer vision competition is common to use pre-trained models in other datasets and here wasn't different.\nBoth models trained on ImageNet and the MTCNN face detector (used a lot here) also have data that doesn't follow the standards you want.\n\nSo I have a few questions:\n\n1 - Why is this rule being applied only in some cases?\n\n2 - If you had this concern at the beginning of the competition, why allow external datasets?\nIf you didn't, wouldn't that be a mistake from you that you are passing on to the competitors?\n\n3 - how should we behave in future situations?\n\nI think these questions are important to answer since competitors spend hundreds of hours working on something without any guarantees.",
    "884685": "I second you, Its really sad to work for hours and then they say we did not want it, I believe there should be a guarantee for the work competitors put here. Kaggle and Facebook both need to answer this, otherwise, should all the competitors depend on the decision you make over your own mistake?",
    "888517": "Unfortunately, it seems we won't have an answer",
    "888532": "Yep, I saw one comment by Julia, she said they will post something about this issue, Iet me find that for you.\n\n[Here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983#887946)\n\nLet's wait until we get one.",
    "889086": "We are unable to provide more pointed guidance on the topic of pre-trained models because licensing issues surrounding pre-trained models can vary on a case-by-case basis. A quick search on the topic will give varied opinions on the interplay between a model and the contents on which it was trained. As a result, we leave judgments as to which licenses are acceptable to the hosts in the hands of the hosts, and universal statements on pre-trained model eligibility cannot be made. \n\nIn this competition, pre-trained models licensed for commercial use were stated as permitted and no team was disqualified for using them. In future competitions, we will urge hosts to clarify how they expect pretrained models to be licensed, and if there are any expected deviations from their competition’s general rules on external data use.\n\nOur clarification and response regarding the disqualification that was made can be read in [this other thread](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/157983#889085).",
    "889968": "juliaelliott I have two points to make regarding this statement `As a result, we leave judgments as to which licenses are acceptable to the hosts in the hands of the hosts, and universal statements on pre-trained model eligibility cannot be made. `\n\n1. Why ImageNet/MTCNN is permitted in the first place if the datasets used to train them doesn't comply with the licensing required by the host? The host should pick a side. It is totally misleading to say that we can selectively pick things for usage and make it compliant with the licensing. The choice is binary. \n2. If you want to avoid such situations in future, then I suggest complete ban on using pretrained weights and external data. By no means, any researcher, ML engineer, or a data scientist can verify every other thing. Also, the host didn't answer our question: Why is it okay to use ImageNet/COCO, etc if we don't have the consent from the source for each sample?",
    "890169": "&gt; Why is it okay to use ImageNet/COCO, etc if we don't have the consent from the source for each sample?\n\nThere is an important difference between \"using pre-trained models trained on dataset X\" and \"using dataset X to train your own models\". Kaggle seems to allow the former, probably because everyone is doing that and there are no legal rulings that say you can't.\n\nWhether the latter is allowed depends on the license terms of that dataset. ImageNet can only be used for non-commercial purposes, IIRC.\n\nSo you *can* use a model that was pre-trained on ImageNet but you can't train your own model on ImageNet.\n\n(I guess this particular interpretation would mean that the creator of, say, MTCNN can't actually use their own model in a Kaggle competition because they had to agree to the license terms of the original dataset. But anyone else who is using MTCNN doesn't have to follow these license terms, because they're not using the original dataset. Yes, this sounds weird. I am not a lawyer and my interpretation may be wrong.)",
    "890501": "I interpret it slightly different. Was the model using the \"mislicensed\" data available for everyone to use. If it were it would fall under the same realms as ImageNet and MTCNN.\n\nThe reality is they went with the majority. Instead of four teams challenging the decision they have one to deal with. Either way, doesn't make it right. The lack of empathy shown from the host and sponsor will serve to aggravate and alienate their users.\n\nIn these situations great companies find a middle ground and compromise. Allowing them to keep their 7th place submission as a compromise is embarrassing to ever Kaggle contestant now and in the future.\n\nPut yourself in their shoes. How would you feel if you knowingly did not cheat and all of a sudden that $500K was replaced with $0 + thousands of hours of your time wasted + your bills + a 7th place standing? I know I would be absolutely depressed!\n\nThis could happen to you and anyone else using this platform in the future. I am certain they will come to their senses and find a middle ground.",
    "890536": "&gt; This could happen to you and anyone else using this platform in the future.\n\nI wasn't saying I agreed with the decision. ;-) Just that equating the use of YouTube videos with the use of ImageNet pre-trained models is not an appropriate comparison.",
    "891019": "humananalog under your logic though, if the the youtube (or other external data) could be trained into a model, released under a permissive open-source license on github, pip, etc and then used as part of the end solution, it would suddenly be okay. That makes no sense.",
    "891617": "rwightman Yes, if the logic I outlined was followed that would be the conclusion *if* we're just talking about the legal aspects of it (disclaimer: I am not a lawyer, laws vary depending on where you live).\n\nThat said, I think this is legally a gray area. AFAIK, training a model on a dataset does not make the model a derivative work of the dataset, so the rights owners of the dataset have no say in how the model can be used. It's a legally gray area because maybe it *should* be a derivative work but no one has challenged this in court yet (to my knowledge).\n\nThe lawyers involved in this competition probably think using ImageNet-based models is fine, even if the courts at some point decide that models indeed are derivative works of their datasets, because using ImageNet-based models is a widespread practice already and that would probably be exempt of such a ruling. (Of course, I'm speculating here.)\n\nFor other datasets and models based on them, such as the hypothetical one you mentioned, the lawyers might play it safer and not allow those. Not because it would currently be illegal but it might be in the future.\n\nOf course, there is also the matter of ethical concerns. For this competition, apparently, a CC license is not enough: Unless you have explicit permission from everyone in your dataset, the dataset cannot be used. For ImageNet / COCO etc, this wouldn't apply because those datasets have been in widespread use for a long time now and it would be unreasonable to fix this retroactively. \n\n(Just a CC license would not be enough because people who licensed their data as CC probably didn't realize it might be used to train things like face detection algorithms, and might not actually agree to that kind of use. You can argue that it doesn't matter that they didn't give explicit consent for this, since the CC license allows you to do whatever you want with these images. Again, this is something that could be challenged in court but to my knowledge hasn't. That said, asking for explicit permission feels like the right thing to do since it's a touchy subject.)",
    "891633": "humananalog this is a bit of a legal flip flop that is really going the extra mile to be cautious. Facebook had \"organic\" deepfakes in their hidden dataset. Do you think they got consent for using these from the people appearing in it? I doubt it. The allow that as it is a \"hidden\" thing.  \nBesides, this logic means that training a model on imagenet would make it illegal, or even using a subset of it. same with coco and open images. I seriously doubt that google went and got consent from anyone appearing in their 1.5M scraped images dataset, and we produced models based on this in a competition. They trust the flickr license.\nAnyways, beside this point, I think that in such a case, with such language, hidden so well under documentation clause, the winning team should have gotten a fair chance to \"fix\" their model and get a new submission with the offending videos removed. It rubs against the grain of most people here how a team that worked so hard can be disqualified without proper explanation and without being given a chance to either fix the problem or just not get prize money and keep the 1st place. I have seen a bit of injustice being done here, but it was mostly the other way around, kaggle turning a blind eye on minor offences. This is something new and a bit worrying.",
    "893514": "The logic is: there is no logic."
  },
  "source": "meta"
}