{
  "id": 172233,
  "title": "Publishing datasets and gcs_paths",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/172233",
  "author_name": "",
  "post_date": "2020-08-04T08:04:26.426898200Z",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi,\nI just published a notebook [here ] (<a href=\"https://www.kaggle.com/wrrosa/siim-utility-package)and\">https://www.kaggle.com/wrrosa/siim-utility-package)and</a> a package in python concerning the use of gcs addresses outside of Kaggle. \nDue to already existing discussions and using these addresses in colab, it seems to be a useful solution.\nTherefore, I have a question for the organizers ( <a href=\"/jwebermsk\">@jwebermsk</a>  and <a href=\"/juliaelliott\">@juliaelliott</a> ): is it legal to publish these addresses outside the Kaggle (e.g. on github)?</p>",
  "messages": [
    {
      "id": "957316",
      "postDate": "08/04/2020 08:04:26",
      "content": "<p>Hi,\nI just published a notebook [here ] (<a href=\"https://www.kaggle.com/wrrosa/siim-utility-package)and\">https://www.kaggle.com/wrrosa/siim-utility-package)and</a> a package in python concerning the use of gcs addresses outside of Kaggle. \nDue to already existing discussions and using these addresses in colab, it seems to be a useful solution.\nTherefore, I have a question for the organizers ( <a href=\"/jwebermsk\">@jwebermsk</a>  and <a href=\"/juliaelliott\">@juliaelliott</a> ): is it legal to publish these addresses outside the Kaggle (e.g. on github)?</p>",
      "rawMarkdown": "Hi,\nI just published a notebook [here ] (https://www.kaggle.com/wrrosa/siim-utility-package)and a package in python concerning the use of gcs addresses outside of Kaggle. \nDue to already existing discussions and using these addresses in colab, it seems to be a useful solution.\nTherefore, I have a question for the organizers ( @jwebermsk  and @juliaelliott ): is it legal to publish these addresses outside the Kaggle (e.g. on github)?",
      "votes": null
    },
    {
      "id": "957435",
      "postDate": "08/04/2020 09:48:47",
      "content": "<p>Interesting question! Maybe its better to address someone directly - or they may oversee the question.</p>",
      "rawMarkdown": "Interesting question! Maybe its better to address someone directly - or they may oversee the question.",
      "votes": null
    },
    {
      "id": "957452",
      "postDate": "08/04/2020 10:18:02",
      "content": "<p>Hi <a href=\"/wrrosa\">@wrrosa</a>, thank you for your work! How would you update the GCS paths in your package? If I understand correctly, you still need to use <code>KaggleDatasets()</code> to do it once a path expires?</p>",
      "rawMarkdown": "Hi @wrrosa, thank you for your work! How would you update the GCS paths in your package? If I understand correctly, you still need to use `KaggleDatasets()` to do it once a path expires?",
      "votes": null
    },
    {
      "id": "957470",
      "postDate": "08/04/2020 10:44:15",
      "content": "<p><a href=\"/kozodoi\">@kozodoi</a> Yes, it's done with kaggle api. Update looks like this: \n1. (API) Generate a list of interesting datasets \n2. (API) Generate and push to kaggle notebooks with these inputs, retrieve gcs_paths using KaggleDatasets() and save results to csv. In fact, this step is realized with 10 notebooks - due to the long response time and httpErrors.\n3. Merge results with previous dictionary and upload to github</p>",
      "rawMarkdown": "kozodoi Yes, it's done with kaggle api. Update looks like this: \n1. (API) Generate a list of interesting datasets \n2. (API) Generate and push to kaggle notebooks with these inputs, retrieve gcs_paths using KaggleDatasets() and save results to csv. In fact, this step is realized with 10 notebooks - due to the long response time and httpErrors.\n3. Merge results with previous dictionary and upload to github",
      "votes": null
    },
    {
      "id": "957942",
      "postDate": "08/04/2020 16:34:58",
      "content": "<p>Hi <a href=\"/wrrosa\">@wrrosa</a>, thanks for checking in on whether this is permissible. In he case of competition data sources, you must ensure you are adhering to the rules of the competition about use of the dataset provided for that competition. For example, most competition rules state in section 7:</p>\n\n<blockquote>\n  <p>B. Data Security. You agree to use reasonable and suitable measures to prevent persons who have not formally agreed to these Rules from gaining access to the Competition Data. You agree not to transmit, duplicate, publish, redistribute or otherwise provide or make available the Competition Data to any party not participating in the Competition. You agree to notify Kaggle immediately upon learning of any possible unauthorized transmission of or unauthorized access to the Competition Data and agree to work with Kaggle to rectify any unauthorized transmission or access.</p>\n</blockquote>\n\n<p>Therefore, if your notebook can be interpreted as disseminating the dataset to others who have not formally agreed to the competition's rules, then that would be prohibited.</p>\n\n<p>For other datasets (not competition-related), the license(s) under which they are available would be shared by the dataset publisher. Many are shared with more open licenses.</p>",
      "rawMarkdown": "Hi @wrrosa, thanks for checking in on whether this is permissible. In he case of competition data sources, you must ensure you are adhering to the rules of the competition about use of the dataset provided for that competition. For example, most competition rules state in section 7:\n\n&gt; B. Data Security. You agree to use reasonable and suitable measures to prevent persons who have not formally agreed to these Rules from gaining access to the Competition Data. You agree not to transmit, duplicate, publish, redistribute or otherwise provide or make available the Competition Data to any party not participating in the Competition. You agree to notify Kaggle immediately upon learning of any possible unauthorized transmission of or unauthorized access to the Competition Data and agree to work with Kaggle to rectify any unauthorized transmission or access.\n\nTherefore, if your notebook can be interpreted as disseminating the dataset to others who have not formally agreed to the competition's rules, then that would be prohibited.\n\nFor other datasets (not competition-related), the license(s) under which they are available would be shared by the dataset publisher. Many are shared with more open licenses.",
      "votes": null
    },
    {
      "id": "959109",
      "postDate": "08/05/2020 10:59:33",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Thank you very much for your comprehensive answer! You've opened my eyes to many important aspects of publishing gcs_paths.\nI realized how important this is - not just for this competition, so I have a few more questions:\n1. Is publishing <strong>kaggle dataset gcs_path</strong> equivalent to publishing <strong>kaggle dataset</strong> (at the time of publication)? I assume the answer is <strong>positive</strong>.\n2. Is publishing notebook on kaggle (licence <strong>Apache 2.0</strong>) like <a href=\"https://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses\">here</a> <em>can be interpreted as disseminating the dataset to others who have not formally agreed to the competition's rules</em>? Note that this notebook, as all public notebooks on kaggle, is visible to others, not even kaggle users.\n3. Is publishing dataset on kaggle (licence <strong>Unknown</strong>) like <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">here</a> <em>can be interpreted as disseminating the dataset to others who have not formally agreed to the competition's rules</em>? Note that this dataset, as all public datasets on kaggle, is visible to others.</p>\n\n<p>For clarity, I just want to know what it looks like formally - I understand that informally it may work differently.</p>",
      "rawMarkdown": "juliaelliott Thank you very much for your comprehensive answer! You've opened my eyes to many important aspects of publishing gcs_paths.\nI realized how important this is - not just for this competition, so I have a few more questions:\n1. Is publishing **kaggle dataset gcs_path** equivalent to publishing **kaggle dataset** (at the time of publication)? I assume the answer is **positive**.\n2. Is publishing notebook on kaggle (licence **Apache 2.0**) like [here](https://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses) *can be interpreted as disseminating the dataset to others who have not formally agreed to the competition's rules*? Note that this notebook, as all public notebooks on kaggle, is visible to others, not even kaggle users.\n3. Is publishing dataset on kaggle (licence **Unknown**) like [here](https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images) *can be interpreted as disseminating the dataset to others who have not formally agreed to the competition's rules*? Note that this dataset, as all public datasets on kaggle, is visible to others.\n\nFor clarity, I just want to know what it looks like formally - I understand that informally it may work differently.",
      "votes": null
    },
    {
      "id": "959536",
      "postDate": "08/05/2020 17:12:49",
      "content": "<p><a href=\"/wrrosa\">@wrrosa</a> Appreciate your interest in clarifying this matter. Due to the public notebook sharing capability within our platform for a competition, we recognize that disclosure/sharing of gcs paths could happen on a spectrum -- despite there being some measures in place to mitigate its impact (i.e. timed access paths). As such, we acknowledge that there are ways that paths could come into the hands of people who have not accepted the rules. </p>\n\n<p>The intent of the rule is to discourage the more overt methods of doing this on that spectrum, hence, the rules specify to \"use reasonable and suitable measures...\" I'll also add that these are rules that you are accepting with the host of the competition (not Kaggle), and enforcement of violations would be in their hands. But, if you ask us whether activity that involves explicitly sharing a gcs path to a dataset is prohibited, we must say yes and cannot encourage the redistribution of the dataset.</p>",
      "rawMarkdown": "wrrosa Appreciate your interest in clarifying this matter. Due to the public notebook sharing capability within our platform for a competition, we recognize that disclosure/sharing of gcs paths could happen on a spectrum -- despite there being some measures in place to mitigate its impact (i.e. timed access paths). As such, we acknowledge that there are ways that paths could come into the hands of people who have not accepted the rules. \n\nThe intent of the rule is to discourage the more overt methods of doing this on that spectrum, hence, the rules specify to \"use reasonable and suitable measures...\" I'll also add that these are rules that you are accepting with the host of the competition (not Kaggle), and enforcement of violations would be in their hands. But, if you ask us whether activity that involves explicitly sharing a gcs path to a dataset is prohibited, we must say yes and cannot encourage the redistribution of the dataset.",
      "votes": null
    },
    {
      "id": "960485",
      "postDate": "08/06/2020 12:38:32",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Thank you very much!\nTherefore, I changed the title of the thread to a more appropriate one.\nAfter all, I developed some workaround to this situation (presented in the same notebook <a href=\"https://www.kaggle.com/wrrosa/siim-utility-package\">here</a>):\n1. Configure kaggle API client\n2. Push notebook to kaggle with requested input and retrieve gcs_path\nImho, such a solution seems to be satisfactory for both sides (kaggle/colab users and Kaggle Team/Competition Organizers). Once again, thank you for your answers.</p>",
      "rawMarkdown": "juliaelliott Thank you very much!\nTherefore, I changed the title of the thread to a more appropriate one.\nAfter all, I developed some workaround to this situation (presented in the same notebook [here](https://www.kaggle.com/wrrosa/siim-utility-package)):\n1. Configure kaggle API client\n2. Push notebook to kaggle with requested input and retrieve gcs_path\nImho, such a solution seems to be satisfactory for both sides (kaggle/colab users and Kaggle Team/Competition Organizers). Once again, thank you for your answers.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 957435,
      "author_name": "romanweilguny",
      "author_url": "",
      "post_date": "08/04/2020 09:48:47",
      "content": "<p>Interesting question! Maybe its better to address someone directly - or they may oversee the question.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 957452,
      "author_name": "kozodoi",
      "author_url": "",
      "post_date": "08/04/2020 10:18:02",
      "content": "<p>Hi <a href=\"/wrrosa\">@wrrosa</a>, thank you for your work! How would you update the GCS paths in your package? If I understand correctly, you still need to use <code>KaggleDatasets()</code> to do it once a path expires?</p>",
      "votes": null,
      "replies": [
        {
          "id": 957470,
          "author_name": "wrrosa",
          "author_url": "",
          "post_date": "08/04/2020 10:44:15",
          "content": "<p><a href=\"/kozodoi\">@kozodoi</a> Yes, it's done with kaggle api. Update looks like this: \n1. (API) Generate a list of interesting datasets \n2. (API) Generate and push to kaggle notebooks with these inputs, retrieve gcs_paths using KaggleDatasets() and save results to csv. In fact, this step is realized with 10 notebooks - due to the long response time and httpErrors.\n3. Merge results with previous dictionary and upload to github</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 957942,
      "author_name": "juliaelliott",
      "author_url": "",
      "post_date": "08/04/2020 16:34:58",
      "content": "<p>Hi <a href=\"/wrrosa\">@wrrosa</a>, thanks for checking in on whether this is permissible. In he case of competition data sources, you must ensure you are adhering to the rules of the competition about use of the dataset provided for that competition. For example, most competition rules state in section 7:</p>\n\n<blockquote>\n  <p>B. Data Security. You agree to use reasonable and suitable measures to prevent persons who have not formally agreed to these Rules from gaining access to the Competition Data. You agree not to transmit, duplicate, publish, redistribute or otherwise provide or make available the Competition Data to any party not participating in the Competition. You agree to notify Kaggle immediately upon learning of any possible unauthorized transmission of or unauthorized access to the Competition Data and agree to work with Kaggle to rectify any unauthorized transmission or access.</p>\n</blockquote>\n\n<p>Therefore, if your notebook can be interpreted as disseminating the dataset to others who have not formally agreed to the competition's rules, then that would be prohibited.</p>\n\n<p>For other datasets (not competition-related), the license(s) under which they are available would be shared by the dataset publisher. Many are shared with more open licenses.</p>",
      "votes": null,
      "replies": [
        {
          "id": 959109,
          "author_name": "wrrosa",
          "author_url": "",
          "post_date": "08/05/2020 10:59:33",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Thank you very much for your comprehensive answer! You've opened my eyes to many important aspects of publishing gcs_paths.\nI realized how important this is - not just for this competition, so I have a few more questions:\n1. Is publishing <strong>kaggle dataset gcs_path</strong> equivalent to publishing <strong>kaggle dataset</strong> (at the time of publication)? I assume the answer is <strong>positive</strong>.\n2. Is publishing notebook on kaggle (licence <strong>Apache 2.0</strong>) like <a href=\"https://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses\">here</a> <em>can be interpreted as disseminating the dataset to others who have not formally agreed to the competition's rules</em>? Note that this notebook, as all public notebooks on kaggle, is visible to others, not even kaggle users.\n3. Is publishing dataset on kaggle (licence <strong>Unknown</strong>) like <a href=\"https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images\">here</a> <em>can be interpreted as disseminating the dataset to others who have not formally agreed to the competition's rules</em>? Note that this dataset, as all public datasets on kaggle, is visible to others.</p>\n\n<p>For clarity, I just want to know what it looks like formally - I understand that informally it may work differently.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 959536,
          "author_name": "juliaelliott",
          "author_url": "",
          "post_date": "08/05/2020 17:12:49",
          "content": "<p><a href=\"/wrrosa\">@wrrosa</a> Appreciate your interest in clarifying this matter. Due to the public notebook sharing capability within our platform for a competition, we recognize that disclosure/sharing of gcs paths could happen on a spectrum -- despite there being some measures in place to mitigate its impact (i.e. timed access paths). As such, we acknowledge that there are ways that paths could come into the hands of people who have not accepted the rules. </p>\n\n<p>The intent of the rule is to discourage the more overt methods of doing this on that spectrum, hence, the rules specify to \"use reasonable and suitable measures...\" I'll also add that these are rules that you are accepting with the host of the competition (not Kaggle), and enforcement of violations would be in their hands. But, if you ask us whether activity that involves explicitly sharing a gcs path to a dataset is prohibited, we must say yes and cannot encourage the redistribution of the dataset.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 960485,
          "author_name": "wrrosa",
          "author_url": "",
          "post_date": "08/06/2020 12:38:32",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Thank you very much!\nTherefore, I changed the title of the thread to a more appropriate one.\nAfter all, I developed some workaround to this situation (presented in the same notebook <a href=\"https://www.kaggle.com/wrrosa/siim-utility-package\">here</a>):\n1. Configure kaggle API client\n2. Push notebook to kaggle with requested input and retrieve gcs_path\nImho, such a solution seems to be satisfactory for both sides (kaggle/colab users and Kaggle Team/Competition Organizers). Once again, thank you for your answers.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "957316": "Hi,\nI just published a notebook [here ] (https://www.kaggle.com/wrrosa/siim-utility-package)and a package in python concerning the use of gcs addresses outside of Kaggle. \nDue to already existing discussions and using these addresses in colab, it seems to be a useful solution.\nTherefore, I have a question for the organizers ( @jwebermsk  and @juliaelliott ): is it legal to publish these addresses outside the Kaggle (e.g. on github)?",
    "957435": "Interesting question! Maybe its better to address someone directly - or they may oversee the question.",
    "957452": "Hi @wrrosa, thank you for your work! How would you update the GCS paths in your package? If I understand correctly, you still need to use `KaggleDatasets()` to do it once a path expires?",
    "957470": "kozodoi Yes, it's done with kaggle api. Update looks like this: \n1. (API) Generate a list of interesting datasets \n2. (API) Generate and push to kaggle notebooks with these inputs, retrieve gcs_paths using KaggleDatasets() and save results to csv. In fact, this step is realized with 10 notebooks - due to the long response time and httpErrors.\n3. Merge results with previous dictionary and upload to github",
    "957942": "Hi @wrrosa, thanks for checking in on whether this is permissible. In he case of competition data sources, you must ensure you are adhering to the rules of the competition about use of the dataset provided for that competition. For example, most competition rules state in section 7:\n\n&gt; B. Data Security. You agree to use reasonable and suitable measures to prevent persons who have not formally agreed to these Rules from gaining access to the Competition Data. You agree not to transmit, duplicate, publish, redistribute or otherwise provide or make available the Competition Data to any party not participating in the Competition. You agree to notify Kaggle immediately upon learning of any possible unauthorized transmission of or unauthorized access to the Competition Data and agree to work with Kaggle to rectify any unauthorized transmission or access.\n\nTherefore, if your notebook can be interpreted as disseminating the dataset to others who have not formally agreed to the competition's rules, then that would be prohibited.\n\nFor other datasets (not competition-related), the license(s) under which they are available would be shared by the dataset publisher. Many are shared with more open licenses.",
    "959109": "juliaelliott Thank you very much for your comprehensive answer! You've opened my eyes to many important aspects of publishing gcs_paths.\nI realized how important this is - not just for this competition, so I have a few more questions:\n1. Is publishing **kaggle dataset gcs_path** equivalent to publishing **kaggle dataset** (at the time of publication)? I assume the answer is **positive**.\n2. Is publishing notebook on kaggle (licence **Apache 2.0**) like [here](https://www.kaggle.com/graf10a/siim-show-gcs-bucket-addresses) *can be interpreted as disseminating the dataset to others who have not formally agreed to the competition's rules*? Note that this notebook, as all public notebooks on kaggle, is visible to others, not even kaggle users.\n3. Is publishing dataset on kaggle (licence **Unknown**) like [here](https://www.kaggle.com/cdeotte/512x512-melanoma-tfrecords-70k-images) *can be interpreted as disseminating the dataset to others who have not formally agreed to the competition's rules*? Note that this dataset, as all public datasets on kaggle, is visible to others.\n\nFor clarity, I just want to know what it looks like formally - I understand that informally it may work differently.",
    "959536": "wrrosa Appreciate your interest in clarifying this matter. Due to the public notebook sharing capability within our platform for a competition, we recognize that disclosure/sharing of gcs paths could happen on a spectrum -- despite there being some measures in place to mitigate its impact (i.e. timed access paths). As such, we acknowledge that there are ways that paths could come into the hands of people who have not accepted the rules. \n\nThe intent of the rule is to discourage the more overt methods of doing this on that spectrum, hence, the rules specify to \"use reasonable and suitable measures...\" I'll also add that these are rules that you are accepting with the host of the competition (not Kaggle), and enforcement of violations would be in their hands. But, if you ask us whether activity that involves explicitly sharing a gcs path to a dataset is prohibited, we must say yes and cannot encourage the redistribution of the dataset.",
    "960485": "juliaelliott Thank you very much!\nTherefore, I changed the title of the thread to a more appropriate one.\nAfter all, I developed some workaround to this situation (presented in the same notebook [here](https://www.kaggle.com/wrrosa/siim-utility-package)):\n1. Configure kaggle API client\n2. Push notebook to kaggle with requested input and retrieve gcs_path\nImho, such a solution seems to be satisfactory for both sides (kaggle/colab users and Kaggle Team/Competition Organizers). Once again, thank you for your answers."
  },
  "source": "meta"
}