{
  "id": 141160,
  "title": "What's the purpose of limiting notebook time to 3 hours -- this limit doesn't seem relevant?",
  "url": "/competitions/jigsaw-multilingual-toxic-comment-classification/discussion/141160",
  "author_name": "",
  "post_date": "2020-04-04T23:29:44.151743500Z",
  "votes": 6,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Code Requirements say:</p>\n\n<blockquote>\n  <p>Submission Notebook run-time per session capped at 3 hours</p>\n</blockquote>\n\n<p>But the test data is fully visible, and external data is allowed. So one can just train a model, or an ensemble of models, for 1000 hours, then spend another 10 hours calculating the predictions, then upload them into a \"kernel\" that does nothing, and submit the predictions instantly.</p>\n\n<p>The 3 hour limit doesn't seem relevant. Am I misunderstanding the rules?</p>",
  "messages": [
    {
      "id": "797815",
      "postDate": "04/04/2020 23:29:44",
      "content": "<p>Code Requirements say:</p>\n\n<blockquote>\n  <p>Submission Notebook run-time per session capped at 3 hours</p>\n</blockquote>\n\n<p>But the test data is fully visible, and external data is allowed. So one can just train a model, or an ensemble of models, for 1000 hours, then spend another 10 hours calculating the predictions, then upload them into a \"kernel\" that does nothing, and submit the predictions instantly.</p>\n\n<p>The 3 hour limit doesn't seem relevant. Am I misunderstanding the rules?</p>",
      "rawMarkdown": "Code Requirements say:\n\n&gt; Submission Notebook run-time per session capped at 3 hours\n\nBut the test data is fully visible, and external data is allowed. So one can just train a model, or an ensemble of models, for 1000 hours, then spend another 10 hours calculating the predictions, then upload them into a \"kernel\" that does nothing, and submit the predictions instantly.\n\nThe 3 hour limit doesn't seem relevant. Am I misunderstanding the rules?",
      "votes": null
    },
    {
      "id": "797841",
      "postDate": "04/05/2020 00:13:15",
      "content": "<p>He will not be able to do that in the private test set. So, The thing to spend 10 hours calculating the predictions will be useless also, that thing does limit on number of models because model number of big models will take more time in the inference.</p>",
      "rawMarkdown": "He will not be able to do that in the private test set. So, The thing to spend 10 hours calculating the predictions will be useless also, that thing does limit on number of models because model number of big models will take more time in the inference.",
      "votes": null
    },
    {
      "id": "797845",
      "postDate": "04/05/2020 00:23:29",
      "content": "<blockquote>\n  <p><strong>Harshit Sheoran wrote:</strong></p>\n  \n  <p>He will not be able to do that in the private test set. </p>\n</blockquote>\n\n<p>There is a private test set? The \"code requirements\" also say:</p>\n\n<blockquote>\n  <p>Your code will not be re-run, as there is no completely hidden test set.</p>\n</blockquote>",
      "rawMarkdown": "&gt; **Harshit Sheoran wrote:**\n&gt; \n&gt; He will not be able to do that in the private test set. \n\nThere is a private test set? The \"code requirements\" also say:\n\n&gt; Your code will not be re-run, as there is no completely hidden test set.",
      "votes": null
    },
    {
      "id": "797849",
      "postDate": "04/05/2020 00:28:00",
      "content": "<p>There sure is, as the rule their is talking about the code will not be re-run on the public test set as the public test set it not hidden.</p>",
      "rawMarkdown": "There sure is, as the rule their is talking about the code will not be re-run on the public test set as the public test set it not hidden.",
      "votes": null
    },
    {
      "id": "797862",
      "postDate": "04/05/2020 00:35:44",
      "content": "<p>The test set is divided into \"public\" and \"private\", but they are both visible:</p>\n\n<blockquote>\n  <p>Your code will not be re-run, as there is no completely hidden test set.</p>\n</blockquote>\n\n<p>So you can calculate all predictions outside of the kernel. That's how I understood it.</p>",
      "rawMarkdown": "The test set is divided into \"public\" and \"private\", but they are both visible:\n\n&gt; Your code will not be re-run, as there is no completely hidden test set.\n\nSo you can calculate all predictions outside of the kernel. That's how I understood it.",
      "votes": null
    },
    {
      "id": "797866",
      "postDate": "04/05/2020 00:41:26",
      "content": "<p>Now I am not sure myself after the line \"but they both are visible\".</p>",
      "rawMarkdown": "Now I am not sure myself after the line \"but they both are visible\".",
      "votes": null
    },
    {
      "id": "801018",
      "postDate": "04/08/2020 00:37:29",
      "content": "<p><strong>There is no hidden test set. We are not re-runnning code.</strong> This is set up like Kaggle's traditional predictions submission competition, except that the submission.csv must originate from a notebook output. </p>\n\n<p>But yes, you could train entirely offline. The constraint is not intended to be a code-enforceable constraint as it is in a synchronous code re-run set-up. This design is solely intended to make building a model through our TPU integration and then submitting from that notebook more seamless.</p>\n\n<p>However, any winning submission still must adhere to the requirement that it not just simply be a hand-labeled test set. Winners' code will be verified and violators will be disqualified.</p>",
      "rawMarkdown": "**There is no hidden test set. We are not re-runnning code.** This is set up like Kaggle's traditional predictions submission competition, except that the submission.csv must originate from a notebook output. \n\nBut yes, you could train entirely offline. The constraint is not intended to be a code-enforceable constraint as it is in a synchronous code re-run set-up. This design is solely intended to make building a model through our TPU integration and then submitting from that notebook more seamless.\n\nHowever, any winning submission still must adhere to the requirement that it not just simply be a hand-labeled test set. Winners' code will be verified and violators will be disqualified.",
      "votes": null
    },
    {
      "id": "805366",
      "postDate": "04/12/2020 15:53:08",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Can we make predictions offline for all individual models and load the predictions as datasets in a notebook and ensemble for the final submission in a notebook? Please help clarify. </p>",
      "rawMarkdown": "juliaelliott Can we make predictions offline for all individual models and load the predictions as datasets in a notebook and ensemble for the final submission in a notebook? Please help clarify.",
      "votes": null
    },
    {
      "id": "805376",
      "postDate": "04/12/2020 16:05:05",
      "content": "<p>I guess from the context of OP and Julia the answer is yes, but just want to make sure my understanding is correct. Thanks.</p>",
      "rawMarkdown": "I guess from the context of OP and Julia the answer is yes, but just want to make sure my understanding is correct. Thanks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 797841,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "04/05/2020 00:13:15",
      "content": "<p>He will not be able to do that in the private test set. So, The thing to spend 10 hours calculating the predictions will be useless also, that thing does limit on number of models because model number of big models will take more time in the inference.</p>",
      "votes": null,
      "replies": [
        {
          "id": 797845,
          "author_name": "olegtrott",
          "author_url": "",
          "post_date": "04/05/2020 00:23:29",
          "content": "<blockquote>\n  <p><strong>Harshit Sheoran wrote:</strong></p>\n  \n  <p>He will not be able to do that in the private test set. </p>\n</blockquote>\n\n<p>There is a private test set? The \"code requirements\" also say:</p>\n\n<blockquote>\n  <p>Your code will not be re-run, as there is no completely hidden test set.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 797849,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "04/05/2020 00:28:00",
          "content": "<p>There sure is, as the rule their is talking about the code will not be re-run on the public test set as the public test set it not hidden.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 797862,
          "author_name": "olegtrott",
          "author_url": "",
          "post_date": "04/05/2020 00:35:44",
          "content": "<p>The test set is divided into \"public\" and \"private\", but they are both visible:</p>\n\n<blockquote>\n  <p>Your code will not be re-run, as there is no completely hidden test set.</p>\n</blockquote>\n\n<p>So you can calculate all predictions outside of the kernel. That's how I understood it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 797866,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "04/05/2020 00:41:26",
          "content": "<p>Now I am not sure myself after the line \"but they both are visible\".</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 801018,
      "author_name": "juliaelliott",
      "author_url": "",
      "post_date": "04/08/2020 00:37:29",
      "content": "<p><strong>There is no hidden test set. We are not re-runnning code.</strong> This is set up like Kaggle's traditional predictions submission competition, except that the submission.csv must originate from a notebook output. </p>\n\n<p>But yes, you could train entirely offline. The constraint is not intended to be a code-enforceable constraint as it is in a synchronous code re-run set-up. This design is solely intended to make building a model through our TPU integration and then submitting from that notebook more seamless.</p>\n\n<p>However, any winning submission still must adhere to the requirement that it not just simply be a hand-labeled test set. Winners' code will be verified and violators will be disqualified.</p>",
      "votes": null,
      "replies": [
        {
          "id": 805366,
          "author_name": "luohongchen1993",
          "author_url": "",
          "post_date": "04/12/2020 15:53:08",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> Can we make predictions offline for all individual models and load the predictions as datasets in a notebook and ensemble for the final submission in a notebook? Please help clarify. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 805376,
          "author_name": "luohongchen1993",
          "author_url": "",
          "post_date": "04/12/2020 16:05:05",
          "content": "<p>I guess from the context of OP and Julia the answer is yes, but just want to make sure my understanding is correct. Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "797815": "Code Requirements say:\n\n&gt; Submission Notebook run-time per session capped at 3 hours\n\nBut the test data is fully visible, and external data is allowed. So one can just train a model, or an ensemble of models, for 1000 hours, then spend another 10 hours calculating the predictions, then upload them into a \"kernel\" that does nothing, and submit the predictions instantly.\n\nThe 3 hour limit doesn't seem relevant. Am I misunderstanding the rules?",
    "797841": "He will not be able to do that in the private test set. So, The thing to spend 10 hours calculating the predictions will be useless also, that thing does limit on number of models because model number of big models will take more time in the inference.",
    "797845": "&gt; **Harshit Sheoran wrote:**\n&gt; \n&gt; He will not be able to do that in the private test set. \n\nThere is a private test set? The \"code requirements\" also say:\n\n&gt; Your code will not be re-run, as there is no completely hidden test set.",
    "797849": "There sure is, as the rule their is talking about the code will not be re-run on the public test set as the public test set it not hidden.",
    "797862": "The test set is divided into \"public\" and \"private\", but they are both visible:\n\n&gt; Your code will not be re-run, as there is no completely hidden test set.\n\nSo you can calculate all predictions outside of the kernel. That's how I understood it.",
    "797866": "Now I am not sure myself after the line \"but they both are visible\".",
    "801018": "**There is no hidden test set. We are not re-runnning code.** This is set up like Kaggle's traditional predictions submission competition, except that the submission.csv must originate from a notebook output. \n\nBut yes, you could train entirely offline. The constraint is not intended to be a code-enforceable constraint as it is in a synchronous code re-run set-up. This design is solely intended to make building a model through our TPU integration and then submitting from that notebook more seamless.\n\nHowever, any winning submission still must adhere to the requirement that it not just simply be a hand-labeled test set. Winners' code will be verified and violators will be disqualified.",
    "805366": "juliaelliott Can we make predictions offline for all individual models and load the predictions as datasets in a notebook and ensemble for the final submission in a notebook? Please help clarify.",
    "805376": "I guess from the context of OP and Julia the answer is yes, but just want to make sure my understanding is correct. Thanks."
  },
  "source": "meta"
}