{
  "id": 71177,
  "title": "How to make sure that model won't be trained from scratch, at the time of \"private\" test score validation ?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/71177",
  "author_name": "",
  "post_date": "2018-11-11T04:06:48.112848100Z",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>According to the competition rules, \nFollowing the final submission deadline for the competition, your kernel code will be re-run on a privately-held test set that is not provided to you. It is your model's score against this private test set that will determine your ranking on the private leaderboard and final standing in the competition.</p>\n\n<p>My concern is, the training script/notebook consists of model training part . So, at the time of re testing with new \"private\" test data, won't the script start from the scratch and retrain the model again? I am not sure, how this part is handling.</p>",
  "messages": [
    {
      "id": "419001",
      "postDate": "11/11/2018 04:06:48",
      "content": "<p>According to the competition rules, \nFollowing the final submission deadline for the competition, your kernel code will be re-run on a privately-held test set that is not provided to you. It is your model's score against this private test set that will determine your ranking on the private leaderboard and final standing in the competition.</p>\n\n<p>My concern is, the training script/notebook consists of model training part . So, at the time of re testing with new \"private\" test data, won't the script start from the scratch and retrain the model again? I am not sure, how this part is handling.</p>",
      "rawMarkdown": "According to the competition rules, \nFollowing the final submission deadline for the competition, your kernel code will be re-run on a privately-held test set that is not provided to you. It is your model's score against this private test set that will determine your ranking on the private leaderboard and final standing in the competition.\n\nMy concern is, the training script/notebook consists of model training part . So, at the time of re testing with new \"private\" test data, won't the script start from the scratch and retrain the model again? I am not sure, how this part is handling.",
      "votes": null
    },
    {
      "id": "419318",
      "postDate": "11/11/2018 17:46:47",
      "content": "<p>Kaggle only compares our predictions against the actual values if I am not wrong. For Public LB they compare it against a small part of the actual values and for the final part, they score it against the complete actual values that we are predicting. They don't run our scripts, evaluation is done on our submissions.</p>",
      "rawMarkdown": "Kaggle only compares our predictions against the actual values if I am not wrong. For Public LB they compare it against a small part of the actual values and for the final part, they score it against the complete actual values that we are predicting. They don't run our scripts, evaluation is done on our submissions.",
      "votes": null
    },
    {
      "id": "419321",
      "postDate": "11/11/2018 17:54:06",
      "content": "<blockquote>\n  <p>They don't run our scripts</p>\n</blockquote>\n\n<p>That's not what the competition description says.</p>\n\n<p>Obviously they don't run scripts on the private test data, otherwise one could output it in a file and then have a look at it.</p>",
      "rawMarkdown": "&gt; They don't run our scripts\n\nThat's not what the competition description says.\n\nObviously they don't run scripts on the private test data, otherwise one could output it in a file and then have a look at it.",
      "votes": null
    },
    {
      "id": "419325",
      "postDate": "11/11/2018 18:03:02",
      "content": "<p>That is what I meant too. but I am going to read the description again.</p>",
      "rawMarkdown": "That is what I meant too. but I am going to read the description again.",
      "votes": null
    },
    {
      "id": "419333",
      "postDate": "11/11/2018 18:33:08",
      "content": "<p>Why is it a problem that your model gets trained from scratch?  That's the purpose of kernel competitions: everything must be run from start to end.</p>",
      "rawMarkdown": "Why is it a problem that your model gets trained from scratch?  That's the purpose of kernel competitions: everything must be run from start to end.",
      "votes": null
    },
    {
      "id": "419457",
      "postDate": "11/12/2018 02:23:56",
      "content": "<p>As per Kaggle <code>your kernel code will be re-run on a privately-held test set that is not provided to you.</code> . Now, in a kernel, I will start like this\n```\n1.  load the train data and preprocess it\n2. create a model using preprocessed data\n3. Evaluate it on validation data ( which is a part of train data and this step is not necessary though )\n4. Make predictions on test data and submit it ( second level evaluation, this test data (private) is different).</p>\n\n<p>```\nNow, if they want to run the kernel code on privately held test data, that basically means they to predict the values right. To, predict those, they need a model, which they can be obtained only through a fresh model training right ?  Hope I make sense. That is why I asked, does the model training start from scratch.</p>",
      "rawMarkdown": "As per Kaggle ```your kernel code will be re-run on a privately-held test set that is not provided to you.``` . Now, in a kernel, I will start like this\n```\n1.  load the train data and preprocess it\n2. create a model using preprocessed data\n3. Evaluate it on validation data ( which is a part of train data and this step is not necessary though )\n4. Make predictions on test data and submit it ( second level evaluation, this test data (private) is different).\n\n```\nNow, if they want to run the kernel code on privately held test data, that basically means they to predict the values right. To, predict those, they need a model, which they can be obtained only through a fresh model training right ?  Hope I make sense. That is why I asked, does the model training start from scratch.",
      "votes": null
    },
    {
      "id": "419458",
      "postDate": "11/12/2018 02:30:37",
      "content": "<p>:-) I don't have any problem. Isn't it better if we save the model/models we have trained via the kernel and provide a new script to load the models and make the prediction? The randomness factor (intialization) might or might not help Neural networks quite a bit right. </p>",
      "rawMarkdown": ":-) I don't have any problem. Isn't it better if we save the model/models we have trained via the kernel and provide a new script to load the models and make the prediction? The randomness factor (intialization) might or might not help Neural networks quite a bit right.",
      "votes": null
    },
    {
      "id": "419608",
      "postDate": "11/12/2018 09:38:32",
      "content": "<p>If you allow saving and reusing, then you give room to building complex model ensembles, which I believe is not what the organizers want when running kernel competitions.</p>",
      "rawMarkdown": "If you allow saving and reusing, then you give room to building complex model ensembles, which I believe is not what the organizers want when running kernel competitions.",
      "votes": null
    },
    {
      "id": "419702",
      "postDate": "11/12/2018 12:27:37",
      "content": "<p>I totally agree :-) . </p>",
      "rawMarkdown": "I totally agree :-) .",
      "votes": null
    },
    {
      "id": "419738",
      "postDate": "11/12/2018 13:48:41",
      "content": "<p>IMHO, they will rerun the script as is, but they'll replace the current test data by the new, private, test data.</p>",
      "rawMarkdown": "IMHO, they will rerun the script as is, but they'll replace the current test data by the new, private, test data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 419318,
      "author_name": "satian",
      "author_url": "",
      "post_date": "11/11/2018 17:46:47",
      "content": "<p>Kaggle only compares our predictions against the actual values if I am not wrong. For Public LB they compare it against a small part of the actual values and for the final part, they score it against the complete actual values that we are predicting. They don't run our scripts, evaluation is done on our submissions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 419321,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/11/2018 17:54:06",
          "content": "<blockquote>\n  <p>They don't run our scripts</p>\n</blockquote>\n\n<p>That's not what the competition description says.</p>\n\n<p>Obviously they don't run scripts on the private test data, otherwise one could output it in a file and then have a look at it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419325,
          "author_name": "satian",
          "author_url": "",
          "post_date": "11/11/2018 18:03:02",
          "content": "<p>That is what I meant too. but I am going to read the description again.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419457,
          "author_name": "s4sarath",
          "author_url": "",
          "post_date": "11/12/2018 02:23:56",
          "content": "<p>As per Kaggle <code>your kernel code will be re-run on a privately-held test set that is not provided to you.</code> . Now, in a kernel, I will start like this\n```\n1.  load the train data and preprocess it\n2. create a model using preprocessed data\n3. Evaluate it on validation data ( which is a part of train data and this step is not necessary though )\n4. Make predictions on test data and submit it ( second level evaluation, this test data (private) is different).</p>\n\n<p>```\nNow, if they want to run the kernel code on privately held test data, that basically means they to predict the values right. To, predict those, they need a model, which they can be obtained only through a fresh model training right ?  Hope I make sense. That is why I asked, does the model training start from scratch.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419738,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/12/2018 13:48:41",
          "content": "<p>IMHO, they will rerun the script as is, but they'll replace the current test data by the new, private, test data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 419333,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "11/11/2018 18:33:08",
      "content": "<p>Why is it a problem that your model gets trained from scratch?  That's the purpose of kernel competitions: everything must be run from start to end.</p>",
      "votes": null,
      "replies": [
        {
          "id": 419458,
          "author_name": "s4sarath",
          "author_url": "",
          "post_date": "11/12/2018 02:30:37",
          "content": "<p>:-) I don't have any problem. Isn't it better if we save the model/models we have trained via the kernel and provide a new script to load the models and make the prediction? The randomness factor (intialization) might or might not help Neural networks quite a bit right. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419608,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/12/2018 09:38:32",
          "content": "<p>If you allow saving and reusing, then you give room to building complex model ensembles, which I believe is not what the organizers want when running kernel competitions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419702,
          "author_name": "s4sarath",
          "author_url": "",
          "post_date": "11/12/2018 12:27:37",
          "content": "<p>I totally agree :-) . </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "419001": "According to the competition rules, \nFollowing the final submission deadline for the competition, your kernel code will be re-run on a privately-held test set that is not provided to you. It is your model's score against this private test set that will determine your ranking on the private leaderboard and final standing in the competition.\n\nMy concern is, the training script/notebook consists of model training part . So, at the time of re testing with new \"private\" test data, won't the script start from the scratch and retrain the model again? I am not sure, how this part is handling.",
    "419318": "Kaggle only compares our predictions against the actual values if I am not wrong. For Public LB they compare it against a small part of the actual values and for the final part, they score it against the complete actual values that we are predicting. They don't run our scripts, evaluation is done on our submissions.",
    "419321": "&gt; They don't run our scripts\n\nThat's not what the competition description says.\n\nObviously they don't run scripts on the private test data, otherwise one could output it in a file and then have a look at it.",
    "419325": "That is what I meant too. but I am going to read the description again.",
    "419333": "Why is it a problem that your model gets trained from scratch?  That's the purpose of kernel competitions: everything must be run from start to end.",
    "419457": "As per Kaggle ```your kernel code will be re-run on a privately-held test set that is not provided to you.``` . Now, in a kernel, I will start like this\n```\n1.  load the train data and preprocess it\n2. create a model using preprocessed data\n3. Evaluate it on validation data ( which is a part of train data and this step is not necessary though )\n4. Make predictions on test data and submit it ( second level evaluation, this test data (private) is different).\n\n```\nNow, if they want to run the kernel code on privately held test data, that basically means they to predict the values right. To, predict those, they need a model, which they can be obtained only through a fresh model training right ?  Hope I make sense. That is why I asked, does the model training start from scratch.",
    "419458": ":-) I don't have any problem. Isn't it better if we save the model/models we have trained via the kernel and provide a new script to load the models and make the prediction? The randomness factor (intialization) might or might not help Neural networks quite a bit right.",
    "419608": "If you allow saving and reusing, then you give room to building complex model ensembles, which I believe is not what the organizers want when running kernel competitions.",
    "419702": "I totally agree :-) .",
    "419738": "IMHO, they will rerun the script as is, but they'll replace the current test data by the new, private, test data."
  },
  "source": "meta"
}