{
  "id": 131121,
  "title": "Is using imagenet pretrained models allowed?",
  "url": "/competitions/deepfake-detection-challenge/discussion/131121",
  "author_name": "",
  "post_date": "2020-02-18T09:46:42.779086500Z",
  "votes": 10,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi,\nit is being discussed here: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203#749071\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203#749071</a> .\nBut I wanted to highlight it in its own discussion topic because it seems important.</p>\n\n<p>The competition rules says the license of anything(code/library/dataset/pretrained models) we are using should be available for commercial use.</p>\n\n<p>And it looks like imagenet is for non-commercial only.\nAlso note that many pre-trained models in library model zoos(including torchvision.models) are trained using imagenet if i am not mistaken.</p>\n\n<p>So is using imagenet pretrained models allowed or not?\nCan we have a clarification on this?</p>\n\n<p><a href=\"/juliaelliott\">@juliaelliott</a> </p>\n\n<p>Edit: Also note that many of the datasets and pre-trained models claimed in \"External Data Disclosure Thread\" are either imagenet based or trained on top of imagenet pre-trained models.</p>",
  "messages": [
    {
      "id": "749101",
      "postDate": "02/18/2020 09:46:42",
      "content": "<p>Hi,\nit is being discussed here: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203#749071\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203#749071</a> .\nBut I wanted to highlight it in its own discussion topic because it seems important.</p>\n\n<p>The competition rules says the license of anything(code/library/dataset/pretrained models) we are using should be available for commercial use.</p>\n\n<p>And it looks like imagenet is for non-commercial only.\nAlso note that many pre-trained models in library model zoos(including torchvision.models) are trained using imagenet if i am not mistaken.</p>\n\n<p>So is using imagenet pretrained models allowed or not?\nCan we have a clarification on this?</p>\n\n<p><a href=\"/juliaelliott\">@juliaelliott</a> </p>\n\n<p>Edit: Also note that many of the datasets and pre-trained models claimed in \"External Data Disclosure Thread\" are either imagenet based or trained on top of imagenet pre-trained models.</p>",
      "rawMarkdown": "Hi,\nit is being discussed here: https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203#749071 .\nBut I wanted to highlight it in its own discussion topic because it seems important.\n\nThe competition rules says the license of anything(code/library/dataset/pretrained models) we are using should be available for commercial use.\n\nAnd it looks like imagenet is for non-commercial only.\nAlso note that many pre-trained models in library model zoos(including torchvision.models) are trained using imagenet if i am not mistaken.\n\nSo is using imagenet pretrained models allowed or not?\nCan we have a clarification on this?\n\n@juliaelliott \n\nEdit: Also note that many of the datasets and pre-trained models claimed in \"External Data Disclosure Thread\" are either imagenet based or trained on top of imagenet pre-trained models.",
      "votes": null
    },
    {
      "id": "749147",
      "postDate": "02/18/2020 11:13:15",
      "content": "<p>I think it’s grey zone. Everything is based on everything. And something of those everything must be non-commercial. So it can be overridden, otherwise the competition is useless. </p>",
      "rawMarkdown": "I think it’s grey zone. Everything is based on everything. And something of those everything must be non-commercial. So it can be overridden, otherwise the competition is useless.",
      "votes": null
    },
    {
      "id": "749682",
      "postDate": "02/18/2020 21:17:46",
      "content": "<p>Exactly. Even dropout is patented!</p>",
      "rawMarkdown": "Exactly. Even dropout is patented!",
      "votes": null
    },
    {
      "id": "749726",
      "postDate": "02/18/2020 21:53:55",
      "content": "<p>I hope you are right. But it is stated many times by <a href=\"/juliaelliott\">@juliaelliott</a>  in \"External Data Disclosure Thread\" that if a dataset has non-commercial only license; we are not allowed to use it. \"I’ve answered the question about BY-NC not being available for use by all (non-commercial use) and therefore violating the requirement that external data be available for use by all participants.\"</p>\n\n<p>So I would prefer an official clarification on this one.</p>",
      "rawMarkdown": "I hope you are right. But it is stated many times by @juliaelliott  in \"External Data Disclosure Thread\" that if a dataset has non-commercial only license; we are not allowed to use it. \"I’ve answered the question about BY-NC not being available for use by all (non-commercial use) and therefore violating the requirement that external data be available for use by all participants.\"\n\nSo I would prefer an official clarification on this one.",
      "votes": null
    },
    {
      "id": "749863",
      "postDate": "02/19/2020 00:20:47",
      "content": "<p>If she did not reply, you got the answer, don’t complicate things. </p>",
      "rawMarkdown": "If she did not reply, you got the answer, don’t complicate things.",
      "votes": null
    },
    {
      "id": "751767",
      "postDate": "02/20/2020 13:50:42",
      "content": "<p>I think it was a grey area until this <a href=\"https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,_Inc\">https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,_Inc</a>. (at least in the US) ... point being you can't publish / alter then republish the datasets .. machine learning weights should be fine with this precedent legally ... Kaggle EULA / TS is something else though.</p>",
      "rawMarkdown": "I think it was a grey area until this https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,_Inc. (at least in the US) ... point being you can't publish / alter then republish the datasets .. machine learning weights should be fine with this precedent legally ... Kaggle EULA / TS is something else though.",
      "votes": null
    },
    {
      "id": "751862",
      "postDate": "02/20/2020 15:24:32",
      "content": "<p>I'm not a lawyer but usually when you make a derivative work of an existing work, you need to have permission from the original copyright holder (depending on what you want to do with it). </p>\n\n<p>The interesting question is: is the result of training a model a derivative work of the dataset, in the sense as understood by copyright law? </p>\n\n<p>If yes, then in the case of training on ImageNet, the model's trained weights would be covered by the ImageNet license.</p>\n\n<p>If no, then the ImageNet license would be irrelevant to the trained model.</p>\n\n<p>But that is a question for the courts, and I don't think this particular issue has been tried yet.</p>",
      "rawMarkdown": "I'm not a lawyer but usually when you make a derivative work of an existing work, you need to have permission from the original copyright holder (depending on what you want to do with it). \n\nThe interesting question is: is the result of training a model a derivative work of the dataset, in the sense as understood by copyright law? \n\nIf yes, then in the case of training on ImageNet, the model's trained weights would be covered by the ImageNet license.\n\nIf no, then the ImageNet license would be irrelevant to the trained model.\n\nBut that is a question for the courts, and I don't think this particular issue has been tried yet.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 749147,
      "author_name": "khahuras",
      "author_url": "",
      "post_date": "02/18/2020 11:13:15",
      "content": "<p>I think it’s grey zone. Everything is based on everything. And something of those everything must be non-commercial. So it can be overridden, otherwise the competition is useless. </p>",
      "votes": null,
      "replies": [
        {
          "id": 749682,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "02/18/2020 21:17:46",
          "content": "<p>Exactly. Even dropout is patented!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749726,
          "author_name": "emrebayram",
          "author_url": "",
          "post_date": "02/18/2020 21:53:55",
          "content": "<p>I hope you are right. But it is stated many times by <a href=\"/juliaelliott\">@juliaelliott</a>  in \"External Data Disclosure Thread\" that if a dataset has non-commercial only license; we are not allowed to use it. \"I’ve answered the question about BY-NC not being available for use by all (non-commercial use) and therefore violating the requirement that external data be available for use by all participants.\"</p>\n\n<p>So I would prefer an official clarification on this one.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749863,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "02/19/2020 00:20:47",
          "content": "<p>If she did not reply, you got the answer, don’t complicate things. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 751767,
      "author_name": "ma7moud",
      "author_url": "",
      "post_date": "02/20/2020 13:50:42",
      "content": "<p>I think it was a grey area until this <a href=\"https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,_Inc\">https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,_Inc</a>. (at least in the US) ... point being you can't publish / alter then republish the datasets .. machine learning weights should be fine with this precedent legally ... Kaggle EULA / TS is something else though.</p>",
      "votes": null,
      "replies": [
        {
          "id": 751862,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "02/20/2020 15:24:32",
          "content": "<p>I'm not a lawyer but usually when you make a derivative work of an existing work, you need to have permission from the original copyright holder (depending on what you want to do with it). </p>\n\n<p>The interesting question is: is the result of training a model a derivative work of the dataset, in the sense as understood by copyright law? </p>\n\n<p>If yes, then in the case of training on ImageNet, the model's trained weights would be covered by the ImageNet license.</p>\n\n<p>If no, then the ImageNet license would be irrelevant to the trained model.</p>\n\n<p>But that is a question for the courts, and I don't think this particular issue has been tried yet.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "749101": "Hi,\nit is being discussed here: https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121203#749071 .\nBut I wanted to highlight it in its own discussion topic because it seems important.\n\nThe competition rules says the license of anything(code/library/dataset/pretrained models) we are using should be available for commercial use.\n\nAnd it looks like imagenet is for non-commercial only.\nAlso note that many pre-trained models in library model zoos(including torchvision.models) are trained using imagenet if i am not mistaken.\n\nSo is using imagenet pretrained models allowed or not?\nCan we have a clarification on this?\n\n@juliaelliott \n\nEdit: Also note that many of the datasets and pre-trained models claimed in \"External Data Disclosure Thread\" are either imagenet based or trained on top of imagenet pre-trained models.",
    "749147": "I think it’s grey zone. Everything is based on everything. And something of those everything must be non-commercial. So it can be overridden, otherwise the competition is useless.",
    "749682": "Exactly. Even dropout is patented!",
    "749726": "I hope you are right. But it is stated many times by @juliaelliott  in \"External Data Disclosure Thread\" that if a dataset has non-commercial only license; we are not allowed to use it. \"I’ve answered the question about BY-NC not being available for use by all (non-commercial use) and therefore violating the requirement that external data be available for use by all participants.\"\n\nSo I would prefer an official clarification on this one.",
    "749863": "If she did not reply, you got the answer, don’t complicate things.",
    "751767": "I think it was a grey area until this https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,_Inc. (at least in the US) ... point being you can't publish / alter then republish the datasets .. machine learning weights should be fine with this precedent legally ... Kaggle EULA / TS is something else though.",
    "751862": "I'm not a lawyer but usually when you make a derivative work of an existing work, you need to have permission from the original copyright holder (depending on what you want to do with it). \n\nThe interesting question is: is the result of training a model a derivative work of the dataset, in the sense as understood by copyright law? \n\nIf yes, then in the case of training on ImageNet, the model's trained weights would be covered by the ImageNet license.\n\nIf no, then the ImageNet license would be irrelevant to the trained model.\n\nBut that is a question for the courts, and I don't think this particular issue has been tried yet."
  },
  "source": "meta"
}