{
  "id": 234915,
  "title": "noobie question",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/234915",
  "author_name": "",
  "post_date": "2021-04-26T19:18:26.107930100Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Now that we have only 14 days to go, I have tons of idea and little TPU time left (yes, my ideas are extremely resource intensive, due to the deep model choice I made). Suppose I train the models on Colab but don't have enough TPU left, is it sufficient if I post and make the models public at least one week before the end of the contest? Yes, I realize others can benefit from them, but I will still have some kept private.</p>",
  "messages": [
    {
      "id": "1285323",
      "postDate": "04/26/2021 19:18:26",
      "content": "<p>Now that we have only 14 days to go, I have tons of idea and little TPU time left (yes, my ideas are extremely resource intensive, due to the deep model choice I made). Suppose I train the models on Colab but don't have enough TPU left, is it sufficient if I post and make the models public at least one week before the end of the contest? Yes, I realize others can benefit from them, but I will still have some kept private.</p>",
      "rawMarkdown": "Now that we have only 14 days to go, I have tons of idea and little TPU time left (yes, my ideas are extremely resource intensive, due to the deep model choice I made). Suppose I train the models on Colab but don't have enough TPU left, is it sufficient if I post and make the models public at least one week before the end of the contest? Yes, I realize others can benefit from them, but I will still have some kept private.",
      "votes": null
    },
    {
      "id": "1285368",
      "postDate": "04/26/2021 21:09:22",
      "content": "<p>I just want to check you know you can train on colab and just test/submit on Kaggle, loading it as an external dataset?</p>\n<p>What I do to greatly speed up committing is if only 5 files are present (IE it's a commit, not a submission) it only processes the smallest TIFF file, and skips the others. That way commits are a few minutes, even if a submission is much much longer.</p>",
      "rawMarkdown": "I just want to check you know you can train on colab and just test/submit on Kaggle, loading it as an external dataset?\n\nWhat I do to greatly speed up committing is if only 5 files are present (IE it's a commit, not a submission) it only processes the smallest TIFF file, and skips the others. That way commits are a few minutes, even if a submission is much much longer.",
      "votes": null
    },
    {
      "id": "1285383",
      "postDate": "04/26/2021 21:59:50",
      "content": "<p>yes, I'm aware. thanks!</p>",
      "rawMarkdown": "yes, I'm aware. thanks!",
      "votes": null
    },
    {
      "id": "1285539",
      "postDate": "04/27/2021 04:02:36",
      "content": "<p>There is no requirement to make your models public before the competition ends.  Only the prize winners are required to make their solutions public, open source after it ends.  And if you were to publish high scoring notebooks in the last week of the competition that would be frowned upon, a message banner will appear asking people not to do that soon. </p>\n<p>There is a requirement for external datasets to be public and shared access via the forum. But that is for additional images data, hand labelled data, not models or pretrained models.  You can review the rules page for more info.  </p>",
      "rawMarkdown": "There is no requirement to make your models public before the competition ends.  Only the prize winners are required to make their solutions public, open source after it ends.  And if you were to publish high scoring notebooks in the last week of the competition that would be frowned upon, a message banner will appear asking people not to do that soon. \n\nThere is a requirement for external datasets to be public and shared access via the forum. But that is for additional images data, hand labelled data, not models or pretrained models.  You can review the rules page for more info.",
      "votes": null
    },
    {
      "id": "1285578",
      "postDate": "04/27/2021 05:30:38",
      "content": "<p>this is what the rules say. I tend to agree with your statement, i.e. the models (*.h5 files) trained externally and imported as external datasets don't fall into the category below:</p>\n<p>C. External Data. You may use data other than the Competition Data (“External Data”) <strong>to develop and test your models and Submissions</strong>. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.</p>",
      "rawMarkdown": "this is what the rules say. I tend to agree with your statement, i.e. the models (*.h5 files) trained externally and imported as external datasets don't fall into the category below:\n\nC. External Data. You may use data other than the Competition Data (“External Data”) **to develop and test your models and Submissions**. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.",
      "votes": null
    },
    {
      "id": "1285928",
      "postDate": "04/27/2021 11:44:49",
      "content": "<blockquote>\n  <p>There is a requirement for external datasets to be public and shared access via the forum. But that is for additional images data, hand labelled data, not models or pretrained models. </p>\n</blockquote>\n<p>External datasets only need to be available for everyone. Not models.</p>\n<p>According to <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/227616#1254955\" target=\"_blank\">this thread</a> and Kaggle staff answer, hand labels is allowed but do not need to be made public:</p>\n<p><a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/227616#1256154\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/227616#1256154</a></p>\n<blockquote>\n  <p>The data itself and notifying of its existence is considered publicly available external data. The hand-labeling performed by any individual is not required to be made public.</p>\n</blockquote>",
      "rawMarkdown": "> There is a requirement for external datasets to be public and shared access via the forum. But that is for additional images data, hand labelled data, not models or pretrained models. \n\nExternal datasets only need to be available for everyone. Not models.\n\nAccording to [this thread](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/227616#1254955) and Kaggle staff answer, hand labels is allowed but do not need to be made public:\n\t\nhttps://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/227616#1256154\n\n> The data itself and notifying of its existence is considered publicly available external data. The hand-labeling performed by any individual is not required to be made public.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1285368,
      "author_name": "jamesphoward",
      "author_url": "",
      "post_date": "04/26/2021 21:09:22",
      "content": "<p>I just want to check you know you can train on colab and just test/submit on Kaggle, loading it as an external dataset?</p>\n<p>What I do to greatly speed up committing is if only 5 files are present (IE it's a commit, not a submission) it only processes the smallest TIFF file, and skips the others. That way commits are a few minutes, even if a submission is much much longer.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1285383,
          "author_name": "andrasferenczi",
          "author_url": "",
          "post_date": "04/26/2021 21:59:50",
          "content": "<p>yes, I'm aware. thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1285539,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "04/27/2021 04:02:36",
      "content": "<p>There is no requirement to make your models public before the competition ends.  Only the prize winners are required to make their solutions public, open source after it ends.  And if you were to publish high scoring notebooks in the last week of the competition that would be frowned upon, a message banner will appear asking people not to do that soon. </p>\n<p>There is a requirement for external datasets to be public and shared access via the forum. But that is for additional images data, hand labelled data, not models or pretrained models.  You can review the rules page for more info.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1285578,
          "author_name": "andrasferenczi",
          "author_url": "",
          "post_date": "04/27/2021 05:30:38",
          "content": "<p>this is what the rules say. I tend to agree with your statement, i.e. the models (*.h5 files) trained externally and imported as external datasets don't fall into the category below:</p>\n<p>C. External Data. You may use data other than the Competition Data (“External Data”) <strong>to develop and test your models and Submissions</strong>. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1285928,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "04/27/2021 11:44:49",
          "content": "<blockquote>\n  <p>There is a requirement for external datasets to be public and shared access via the forum. But that is for additional images data, hand labelled data, not models or pretrained models. </p>\n</blockquote>\n<p>External datasets only need to be available for everyone. Not models.</p>\n<p>According to <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/227616#1254955\" target=\"_blank\">this thread</a> and Kaggle staff answer, hand labels is allowed but do not need to be made public:</p>\n<p><a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/227616#1256154\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/227616#1256154</a></p>\n<blockquote>\n  <p>The data itself and notifying of its existence is considered publicly available external data. The hand-labeling performed by any individual is not required to be made public.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1285323": "Now that we have only 14 days to go, I have tons of idea and little TPU time left (yes, my ideas are extremely resource intensive, due to the deep model choice I made). Suppose I train the models on Colab but don't have enough TPU left, is it sufficient if I post and make the models public at least one week before the end of the contest? Yes, I realize others can benefit from them, but I will still have some kept private.",
    "1285368": "I just want to check you know you can train on colab and just test/submit on Kaggle, loading it as an external dataset?\n\nWhat I do to greatly speed up committing is if only 5 files are present (IE it's a commit, not a submission) it only processes the smallest TIFF file, and skips the others. That way commits are a few minutes, even if a submission is much much longer.",
    "1285383": "yes, I'm aware. thanks!",
    "1285539": "There is no requirement to make your models public before the competition ends.  Only the prize winners are required to make their solutions public, open source after it ends.  And if you were to publish high scoring notebooks in the last week of the competition that would be frowned upon, a message banner will appear asking people not to do that soon. \n\nThere is a requirement for external datasets to be public and shared access via the forum. But that is for additional images data, hand labelled data, not models or pretrained models.  You can review the rules page for more info.",
    "1285578": "this is what the rules say. I tend to agree with your statement, i.e. the models (*.h5 files) trained externally and imported as external datasets don't fall into the category below:\n\nC. External Data. You may use data other than the Competition Data (“External Data”) **to develop and test your models and Submissions**. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.",
    "1285928": "> There is a requirement for external datasets to be public and shared access via the forum. But that is for additional images data, hand labelled data, not models or pretrained models. \n\nExternal datasets only need to be available for everyone. Not models.\n\nAccording to [this thread](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/227616#1254955) and Kaggle staff answer, hand labels is allowed but do not need to be made public:\n\t\nhttps://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/227616#1256154\n\n> The data itself and notifying of its existence is considered publicly available external data. The hand-labeling performed by any individual is not required to be made public."
  },
  "source": "meta"
}