{
  "id": 197365,
  "title": "Can we use offline trained models for the final submission?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/197365",
  "author_name": "",
  "post_date": "2020-11-16T00:40:49.805679Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>With the limited resources provided by the Kaggle kernel, it's really time-consuming to process the data and train the model.<br>\nCurrently, I do the data processing and feature generation on spark. Then train the model on a deluxe-configured computer. After this, I uploaded the trained model and statistics to the kernel for batch inference.<br>\nI want to know whether I can still do this for the final submission?</p>",
  "messages": [
    {
      "id": "1079351",
      "postDate": "11/16/2020 00:40:49",
      "content": "<p>With the limited resources provided by the Kaggle kernel, it's really time-consuming to process the data and train the model.<br>\nCurrently, I do the data processing and feature generation on spark. Then train the model on a deluxe-configured computer. After this, I uploaded the trained model and statistics to the kernel for batch inference.<br>\nI want to know whether I can still do this for the final submission?</p>",
      "rawMarkdown": "With the limited resources provided by the Kaggle kernel, it's really time-consuming to process the data and train the model.\nCurrently, I do the data processing and feature generation on spark. Then train the model on a deluxe-configured computer. After this, I uploaded the trained model and statistics to the kernel for batch inference.\nI want to know whether I can still do this for the final submission?",
      "votes": null
    },
    {
      "id": "1079373",
      "postDate": "11/16/2020 01:22:19",
      "content": "<p>yes, this is allowed from <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/overview/code-requirements\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/overview/code-requirements</a> </p>\n<blockquote>\n  <p>Freely &amp; publicly available external data is allowed, including pre-trained models</p>\n</blockquote>",
      "rawMarkdown": "yes, this is allowed from https://www.kaggle.com/c/riiid-test-answer-prediction/overview/code-requirements \n> Freely & publicly available external data is allowed, including pre-trained models",
      "votes": null
    },
    {
      "id": "1079587",
      "postDate": "11/16/2020 09:09:50",
      "content": "<p>Got it, thanks <a href=\"https://www.kaggle.com/rashmibanthia\" target=\"_blank\">@rashmibanthia</a> </p>",
      "rawMarkdown": "Got it, thanks @rashmibanthia",
      "votes": null
    },
    {
      "id": "1079621",
      "postDate": "11/16/2020 10:13:12",
      "content": "<p>Actually, it is an interesting question. There are two statements in the Rules (<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/overview/code-requirements\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/overview/code-requirements</a>) that might influence this:</p>\n<ul>\n<li>Please note that for this competition training is not required in Notebooks.</li>\n<li>Freely &amp; publicly available external data is allowed, including pre-trained models.</li>\n</ul>\n<p>The second one allows us to use external data only if it is \"freely and publicly available\", therefore, strictly following this rule, you have to make a kind of <em>public</em> dataset with your trained model if you wish to use it during submission.</p>\n<p>The first one is (as I understand it) looks a bit more permissive with respect to the considered scenario. At least, I understand \"the spirit of this law\" that a model trained by a team is not classified as \"external data\". But still, if we follow \"the letter of the law\" a model trained by a team is somewhat external to the inference kernel, therefore, should be made public to be used.</p>\n<p>P.S. Please, sorry if I'm bringing some confusion to this discussion. I'm relatively new to Kaggle, so some things obvious for the community might not be as obvious to me.</p>",
      "rawMarkdown": "Actually, it is an interesting question. There are two statements in the Rules (https://www.kaggle.com/c/riiid-test-answer-prediction/overview/code-requirements) that might influence this:\n- Please note that for this competition training is not required in Notebooks.\n- Freely & publicly available external data is allowed, including pre-trained models.\n\nThe second one allows us to use external data only if it is \"freely and publicly available\", therefore, strictly following this rule, you have to make a kind of *public* dataset with your trained model if you wish to use it during submission.\n\nThe first one is (as I understand it) looks a bit more permissive with respect to the considered scenario. At least, I understand \"the spirit of this law\" that a model trained by a team is not classified as \"external data\". But still, if we follow \"the letter of the law\" a model trained by a team is somewhat external to the inference kernel, therefore, should be made public to be used.\n\nP.S. Please, sorry if I'm bringing some confusion to this discussion. I'm relatively new to Kaggle, so some things obvious for the community might not be as obvious to me.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1079373,
      "author_name": "rashmibanthia",
      "author_url": "",
      "post_date": "11/16/2020 01:22:19",
      "content": "<p>yes, this is allowed from <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/overview/code-requirements\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/overview/code-requirements</a> </p>\n<blockquote>\n  <p>Freely &amp; publicly available external data is allowed, including pre-trained models</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1079587,
      "author_name": "wuwenmin",
      "author_url": "",
      "post_date": "11/16/2020 09:09:50",
      "content": "<p>Got it, thanks <a href=\"https://www.kaggle.com/rashmibanthia\" target=\"_blank\">@rashmibanthia</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1079621,
      "author_name": "ponomarevav",
      "author_url": "",
      "post_date": "11/16/2020 10:13:12",
      "content": "<p>Actually, it is an interesting question. There are two statements in the Rules (<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/overview/code-requirements\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/overview/code-requirements</a>) that might influence this:</p>\n<ul>\n<li>Please note that for this competition training is not required in Notebooks.</li>\n<li>Freely &amp; publicly available external data is allowed, including pre-trained models.</li>\n</ul>\n<p>The second one allows us to use external data only if it is \"freely and publicly available\", therefore, strictly following this rule, you have to make a kind of <em>public</em> dataset with your trained model if you wish to use it during submission.</p>\n<p>The first one is (as I understand it) looks a bit more permissive with respect to the considered scenario. At least, I understand \"the spirit of this law\" that a model trained by a team is not classified as \"external data\". But still, if we follow \"the letter of the law\" a model trained by a team is somewhat external to the inference kernel, therefore, should be made public to be used.</p>\n<p>P.S. Please, sorry if I'm bringing some confusion to this discussion. I'm relatively new to Kaggle, so some things obvious for the community might not be as obvious to me.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1079351": "With the limited resources provided by the Kaggle kernel, it's really time-consuming to process the data and train the model.\nCurrently, I do the data processing and feature generation on spark. Then train the model on a deluxe-configured computer. After this, I uploaded the trained model and statistics to the kernel for batch inference.\nI want to know whether I can still do this for the final submission?",
    "1079373": "yes, this is allowed from https://www.kaggle.com/c/riiid-test-answer-prediction/overview/code-requirements \n> Freely & publicly available external data is allowed, including pre-trained models",
    "1079587": "Got it, thanks @rashmibanthia",
    "1079621": "Actually, it is an interesting question. There are two statements in the Rules (https://www.kaggle.com/c/riiid-test-answer-prediction/overview/code-requirements) that might influence this:\n- Please note that for this competition training is not required in Notebooks.\n- Freely & publicly available external data is allowed, including pre-trained models.\n\nThe second one allows us to use external data only if it is \"freely and publicly available\", therefore, strictly following this rule, you have to make a kind of *public* dataset with your trained model if you wish to use it during submission.\n\nThe first one is (as I understand it) looks a bit more permissive with respect to the considered scenario. At least, I understand \"the spirit of this law\" that a model trained by a team is not classified as \"external data\". But still, if we follow \"the letter of the law\" a model trained by a team is somewhat external to the inference kernel, therefore, should be made public to be used.\n\nP.S. Please, sorry if I'm bringing some confusion to this discussion. I'm relatively new to Kaggle, so some things obvious for the community might not be as obvious to me."
  },
  "source": "meta"
}