{
  "id": 17983,
  "title": "Submit model before end competition?",
  "url": "/competitions/second-annual-data-science-bowl/discussion/17983",
  "author_name": "",
  "post_date": "2015-12-17T14:40:26.857Z",
  "votes": 1,
  "comment_count": 7,
  "views": 1315,
  "content": "<blockquote>\n  <p>Two weeks before the final deadline, you will submit your model to Kaggle. At this point, the second stage of the competition starts. Kaggle will release the final test dataset, on which you will run your models. The final standings are based on this final test set.</p>\n</blockquote>\n\n<p>To me, this part of the rules are rather vague. What is meant with model? Just the code, or code+parameters? If a bug pops up on the final test data, does that mean that you cannot fix your stupid divide-by-zero? Can you do an unsupervised training on the test data before evaluating the test data on your final model? Do you need to pre-process/clean up all data automatically?</p>\n\n<p>To me this distinction seems arbitrary. I'd like to know where the judges will draw the line more or less, and what is meant by 'model'. I always understood the model to be a mathematical abstraction, distinct from code or parameters.</p>",
  "messages": [
    {
      "id": "101859",
      "postDate": "12/17/2015 14:40:26",
      "content": "<blockquote>\n  <p>Two weeks before the final deadline, you will submit your model to Kaggle. At this point, the second stage of the competition starts. Kaggle will release the final test dataset, on which you will run your models. The final standings are based on this final test set.</p>\n</blockquote>\n\n<p>To me, this part of the rules are rather vague. What is meant with model? Just the code, or code+parameters? If a bug pops up on the final test data, does that mean that you cannot fix your stupid divide-by-zero? Can you do an unsupervised training on the test data before evaluating the test data on your final model? Do you need to pre-process/clean up all data automatically?</p>\n\n<p>To me this distinction seems arbitrary. I'd like to know where the judges will draw the line more or less, and what is meant by 'model'. I always understood the model to be a mathematical abstraction, distinct from code or parameters.</p>",
      "rawMarkdown": "> Two weeks before the final deadline, you will submit your model to Kaggle. At this point, the second stage of the competition starts. Kaggle will release the final test dataset, on which you will run your models. The final standings are based on this final test set.\r\n\r\nTo me, this part of the rules are rather vague. What is meant with model? Just the code, or code+parameters? If a bug pops up on the final test data, does that mean that you cannot fix your stupid divide-by-zero? Can you do an unsupervised training on the test data before evaluating the test data on your final model? Do you need to pre-process/clean up all data automatically?\r\n\r\nTo me this distinction seems arbitrary. I'd like to know where the judges will draw the line more or less, and what is meant by 'model'. I always understood the model to be a mathematical abstraction, distinct from code or parameters.",
      "votes": null
    },
    {
      "id": "101866",
      "postDate": "12/17/2015 15:06:58",
      "content": "<p>Great questions!</p>\n\n<blockquote>\n  <p>What is meant with model? </p>\n</blockquote>\n\n<p>&quot;Model&quot; is code that turns an MRI into two volume numbers. It includes parameters, seeds, references to external libraries, etc. We have some rough best practices here (these are not mandatory, but help you and the host) -  </p>\n\n<p><a href=\"https://www.kaggle.com/wiki/ModelSubmissionBestPractices\">https://www.kaggle.com/wiki/ModelSubmissionBestPractices</a></p>\n\n<blockquote>\n  <p>If a bug pops up on the final test data, does that mean that you cannot fix your stupid divide-by-zero?</p>\n</blockquote>\n\n<p>We make reasonable allowances for bug fixes like this. Still, you should make every effort to build error handling into your pipeline. The less you need to fix when the new test set comes out, the less chance you will need to justify your changes.</p>\n\n<blockquote>\n  <p>Can you do an unsupervised training on the test data before evaluating the test data on your final model?</p>\n</blockquote>\n\n<p>Yes. Semi-supervised learning is permitted.</p>\n\n<blockquote>\n  <p>Do you need to pre-process/clean up all data automatically?</p>\n</blockquote>\n\n<p>Yes.</p>",
      "rawMarkdown": "Great questions!\r\n\r\n> What is meant with model? \r\n\r\n\"Model\" is code that turns an MRI into two volume numbers. It includes parameters, seeds, references to external libraries, etc. We have some rough best practices here (these are not mandatory, but help you and the host) -  \r\n\r\nhttps://www.kaggle.com/wiki/ModelSubmissionBestPractices\r\n\r\n> If a bug pops up on the final test data, does that mean that you cannot fix your stupid divide-by-zero?\r\n\r\nWe make reasonable allowances for bug fixes like this. Still, you should make every effort to build error handling into your pipeline. The less you need to fix when the new test set comes out, the less chance you will need to justify your changes.\r\n\r\n> Can you do an unsupervised training on the test data before evaluating the test data on your final model?\r\n\r\nYes. Semi-supervised learning is permitted.\r\n\r\n> Do you need to pre-process/clean up all data automatically?\r\n\r\nYes.",
      "votes": null
    },
    {
      "id": "101908",
      "postDate": "12/17/2015 19:51:54",
      "content": "<p>Thanks, I was wondering about how this worked as well. So that final time window is not there to let you tweak your model further using the test set, correct? Instead, it's just to give teams an adequate chance to actually run and upload new submissions based on the test set using their already submitted models?</p>",
      "rawMarkdown": "Thanks, I was wondering about how this worked as well. So that final time window is not there to let you tweak your model further using the test set, correct? Instead, it's just to give teams an adequate chance to actually run and upload new submissions based on the test set using their already submitted models?",
      "votes": null
    },
    {
      "id": "108657",
      "postDate": "02/19/2016 00:01:28",
      "content": "<p>I have another great question.</p>\n\n<p>How do I send the model files ? \nI need to put in a public repository, like Github, and share the link or Kaggle will open a space to upload ?</p>",
      "rawMarkdown": "I have another great question.\r\n\r\nHow do I send the model files ? \r\nI need to put in a public repository, like Github, and share the link or Kaggle will open a space to upload ?",
      "votes": null
    },
    {
      "id": "108669",
      "postDate": "02/19/2016 02:17:55",
      "content": "<p><a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/model\">https://www.kaggle.com/c/second-annual-data-science-bowl/model</a></p>",
      "rawMarkdown": "https://www.kaggle.com/c/second-annual-data-science-bowl/model",
      "votes": null
    },
    {
      "id": "108719",
      "postDate": "02/19/2016 12:18:37",
      "content": "<p>Thanks, I had interpreted this link sends the CSV file, because the page header.</p>\n\n<p>Is there a maximum limit to the size of each file or all files ? \nI did not find it in the informations.</p>",
      "rawMarkdown": "Thanks, I had interpreted this link sends the CSV file, because the page header.\r\n\r\nIs there a maximum limit to the size of each file or all files ? \r\nI did not find it in the informations.",
      "votes": null
    },
    {
      "id": "108722",
      "postDate": "02/19/2016 13:34:57",
      "content": "<p>The model uploader can handle large files. I don't know the limit offhand, but it is larger than your model should be. You don't need to include the dataset as part of the model.</p>",
      "rawMarkdown": "The model uploader can handle large files. I don't know the limit offhand, but it is larger than your model should be. You don't need to include the dataset as part of the model.",
      "votes": null
    },
    {
      "id": "108729",
      "postDate": "02/19/2016 14:19:14",
      "content": "<p>Thanks for the answer.</p>",
      "rawMarkdown": "Thanks for the answer.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 101866,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "12/17/2015 15:06:58",
      "content": "<p>Great questions!</p>\n\n<blockquote>\n  <p>What is meant with model? </p>\n</blockquote>\n\n<p>&quot;Model&quot; is code that turns an MRI into two volume numbers. It includes parameters, seeds, references to external libraries, etc. We have some rough best practices here (these are not mandatory, but help you and the host) -  </p>\n\n<p><a href=\"https://www.kaggle.com/wiki/ModelSubmissionBestPractices\">https://www.kaggle.com/wiki/ModelSubmissionBestPractices</a></p>\n\n<blockquote>\n  <p>If a bug pops up on the final test data, does that mean that you cannot fix your stupid divide-by-zero?</p>\n</blockquote>\n\n<p>We make reasonable allowances for bug fixes like this. Still, you should make every effort to build error handling into your pipeline. The less you need to fix when the new test set comes out, the less chance you will need to justify your changes.</p>\n\n<blockquote>\n  <p>Can you do an unsupervised training on the test data before evaluating the test data on your final model?</p>\n</blockquote>\n\n<p>Yes. Semi-supervised learning is permitted.</p>\n\n<blockquote>\n  <p>Do you need to pre-process/clean up all data automatically?</p>\n</blockquote>\n\n<p>Yes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 101908,
      "author_name": "bgalbraith",
      "author_url": "",
      "post_date": "12/17/2015 19:51:54",
      "content": "<p>Thanks, I was wondering about how this worked as well. So that final time window is not there to let you tweak your model further using the test set, correct? Instead, it's just to give teams an adequate chance to actually run and upload new submissions based on the test set using their already submitted models?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108657,
      "author_name": "alvaroosvaldo",
      "author_url": "",
      "post_date": "02/19/2016 00:01:28",
      "content": "<p>I have another great question.</p>\n\n<p>How do I send the model files ? \nI need to put in a public repository, like Github, and share the link or Kaggle will open a space to upload ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108669,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "02/19/2016 02:17:55",
      "content": "<p><a href=\"https://www.kaggle.com/c/second-annual-data-science-bowl/model\">https://www.kaggle.com/c/second-annual-data-science-bowl/model</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108719,
      "author_name": "alvaroosvaldo",
      "author_url": "",
      "post_date": "02/19/2016 12:18:37",
      "content": "<p>Thanks, I had interpreted this link sends the CSV file, because the page header.</p>\n\n<p>Is there a maximum limit to the size of each file or all files ? \nI did not find it in the informations.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108722,
      "author_name": "wcukierski",
      "author_url": "",
      "post_date": "02/19/2016 13:34:57",
      "content": "<p>The model uploader can handle large files. I don't know the limit offhand, but it is larger than your model should be. You don't need to include the dataset as part of the model.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 108729,
      "author_name": "alvaroosvaldo",
      "author_url": "",
      "post_date": "02/19/2016 14:19:14",
      "content": "<p>Thanks for the answer.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "101859": "> Two weeks before the final deadline, you will submit your model to Kaggle. At this point, the second stage of the competition starts. Kaggle will release the final test dataset, on which you will run your models. The final standings are based on this final test set.\r\n\r\nTo me, this part of the rules are rather vague. What is meant with model? Just the code, or code+parameters? If a bug pops up on the final test data, does that mean that you cannot fix your stupid divide-by-zero? Can you do an unsupervised training on the test data before evaluating the test data on your final model? Do you need to pre-process/clean up all data automatically?\r\n\r\nTo me this distinction seems arbitrary. I'd like to know where the judges will draw the line more or less, and what is meant by 'model'. I always understood the model to be a mathematical abstraction, distinct from code or parameters.",
    "101866": "Great questions!\r\n\r\n> What is meant with model? \r\n\r\n\"Model\" is code that turns an MRI into two volume numbers. It includes parameters, seeds, references to external libraries, etc. We have some rough best practices here (these are not mandatory, but help you and the host) -  \r\n\r\nhttps://www.kaggle.com/wiki/ModelSubmissionBestPractices\r\n\r\n> If a bug pops up on the final test data, does that mean that you cannot fix your stupid divide-by-zero?\r\n\r\nWe make reasonable allowances for bug fixes like this. Still, you should make every effort to build error handling into your pipeline. The less you need to fix when the new test set comes out, the less chance you will need to justify your changes.\r\n\r\n> Can you do an unsupervised training on the test data before evaluating the test data on your final model?\r\n\r\nYes. Semi-supervised learning is permitted.\r\n\r\n> Do you need to pre-process/clean up all data automatically?\r\n\r\nYes.",
    "101908": "Thanks, I was wondering about how this worked as well. So that final time window is not there to let you tweak your model further using the test set, correct? Instead, it's just to give teams an adequate chance to actually run and upload new submissions based on the test set using their already submitted models?",
    "108657": "I have another great question.\r\n\r\nHow do I send the model files ? \r\nI need to put in a public repository, like Github, and share the link or Kaggle will open a space to upload ?",
    "108669": "https://www.kaggle.com/c/second-annual-data-science-bowl/model",
    "108719": "Thanks, I had interpreted this link sends the CSV file, because the page header.\r\n\r\nIs there a maximum limit to the size of each file or all files ? \r\nI did not find it in the informations.",
    "108722": "The model uploader can handle large files. I don't know the limit offhand, but it is larger than your model should be. You don't need to include the dataset as part of the model.",
    "108729": "Thanks for the answer."
  },
  "source": "meta"
}