{
  "id": 75965,
  "title": "Is offline preprocessing allowed?",
  "url": "/competitions/NFL-Punt-Analytics-Competition/discussion/75965",
  "author_name": "",
  "post_date": "2018-12-28T02:40:45.520788100Z",
  "votes": 4,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I've done a lot of data preprocessing/cleansing offline on my personal machine. Instead of having to run this preprocessing again through the kaggle kernel (which would be very slow and limited to the disk space of the kernel VM) - I was hoping to just upload the processed data as a separate dataset- which I would link to in my submission kernel.</p>\n\n<p>I would reveal all my preprocessing scripts so the analysis would be reproducible.</p>\n\n<p>The rules state that <code>Your kernels should be easy to understand and the analysis should be reproducible.</code> </p>\n\n<ul>\n<li>Would this approach meet the kernel requirement?</li>\n<li>Are others having the same issue? If so- how do you plan to deal with it?</li>\n</ul>",
  "messages": [
    {
      "id": "446407",
      "postDate": "12/28/2018 02:40:45",
      "content": "<p>I've done a lot of data preprocessing/cleansing offline on my personal machine. Instead of having to run this preprocessing again through the kaggle kernel (which would be very slow and limited to the disk space of the kernel VM) - I was hoping to just upload the processed data as a separate dataset- which I would link to in my submission kernel.</p>\n\n<p>I would reveal all my preprocessing scripts so the analysis would be reproducible.</p>\n\n<p>The rules state that <code>Your kernels should be easy to understand and the analysis should be reproducible.</code> </p>\n\n<ul>\n<li>Would this approach meet the kernel requirement?</li>\n<li>Are others having the same issue? If so- how do you plan to deal with it?</li>\n</ul>",
      "rawMarkdown": "I've done a lot of data preprocessing/cleansing offline on my personal machine. Instead of having to run this preprocessing again through the kaggle kernel (which would be very slow and limited to the disk space of the kernel VM) - I was hoping to just upload the processed data as a separate dataset- which I would link to in my submission kernel.\n\nI would reveal all my preprocessing scripts so the analysis would be reproducible.\n\nThe rules state that `Your kernels should be easy to understand and the analysis should be reproducible.` \n\n- Would this approach meet the kernel requirement?\n- Are others having the same issue? If so- how do you plan to deal with it?",
      "votes": null
    },
    {
      "id": "446995",
      "postDate": "12/29/2018 00:59:02",
      "content": "<p>Haven't read the whole rules but am annoyed by the kernel requirement as well. I'll probably end up doing what you're doing too but commenting to see if we get an official reply</p>",
      "rawMarkdown": "Haven't read the whole rules but am annoyed by the kernel requirement as well. I'll probably end up doing what you're doing too but commenting to see if we get an official reply",
      "votes": null
    },
    {
      "id": "447092",
      "postDate": "12/29/2018 05:55:12",
      "content": "<p>I wish my personal machine was better than a Kaggle kernel in more aspects than SSD space...</p>",
      "rawMarkdown": "I wish my personal machine was better than a Kaggle kernel in more aspects than SSD space...",
      "votes": null
    },
    {
      "id": "447295",
      "postDate": "12/29/2018 15:27:52",
      "content": "<p>As long as complete generating code is provided I think there is no problem with using some of the most time consuming process as a linked dataset. </p>\n\n<p>I  will possibly do that for a subset of the calculations related to NGS data that required some heavy paralelized processing.</p>",
      "rawMarkdown": "As long as complete generating code is provided I think there is no problem with using some of the most time consuming process as a linked dataset. \n\nI  will possibly do that for a subset of the calculations related to NGS data that required some heavy paralelized processing.",
      "votes": null
    },
    {
      "id": "447359",
      "postDate": "12/29/2018 17:55:22",
      "content": "<p>If you need more resources you could always spin up a Google Compute Engine or AWS virtual machine. It might cost a little money but if you only use it for heavy computation and turn it off when you're not using it the cost is pretty minimal. </p>",
      "rawMarkdown": "If you need more resources you could always spin up a Google Compute Engine or AWS virtual machine. It might cost a little money but if you only use it for heavy computation and turn it off when you're not using it the cost is pretty minimal.",
      "votes": null
    },
    {
      "id": "449285",
      "postDate": "01/02/2019 22:45:05",
      "content": "<p>You should probably do all of your work in Kernels because it's really the only way to guarantee we have the same environment for reproducibility. With a Kernel, I can fork it and run it flawlessly every time. I would be amazed if you sent me a script for processing a dataset that ran perfectly on the first try. </p>",
      "rawMarkdown": "You should probably do all of your work in Kernels because it's really the only way to guarantee we have the same environment for reproducibility. With a Kernel, I can fork it and run it flawlessly every time. I would be amazed if you sent me a script for processing a dataset that ran perfectly on the first try.",
      "votes": null
    },
    {
      "id": "449322",
      "postDate": "01/03/2019 00:43:16",
      "content": "<p>Noob question, but is it possible to use KNIME (similar to Alteryx or RapidMiner) for preprocessing and attach the workflow in the Kernel?</p>",
      "rawMarkdown": "Noob question, but is it possible to use KNIME (similar to Alteryx or RapidMiner) for preprocessing and attach the workflow in the Kernel?",
      "votes": null
    },
    {
      "id": "449831",
      "postDate": "01/03/2019 20:28:09",
      "content": "<p>Can you export the workflows to Python or R and run it in a kernel? This is the first time I've heard of KNIME.</p>",
      "rawMarkdown": "Can you export the workflows to Python or R and run it in a kernel? This is the first time I've heard of KNIME.",
      "votes": null
    },
    {
      "id": "449904",
      "postDate": "01/03/2019 23:39:31",
      "content": "<p>KNIME is an open source data workbench that lets you quickly process data in a visual format, instead of scripts. I can likely translate necessary steps into Python - but if it's all the same, I would much rather submit a copy of the workflow as it can easily be opened and run in a free open-source download of KNIME (and it's my current strong suit as I progress in Python). There is an interesting Kernel mirroring KNIME and Python workflows on the Titanic data set if you're interested: <a href=\"https://www.kaggle.com/grayrat/scripting-a-model-from-knime-workflow-to-python\">https://www.kaggle.com/grayrat/scripting-a-model-from-knime-workflow-to-python</a> </p>",
      "rawMarkdown": "KNIME is an open source data workbench that lets you quickly process data in a visual format, instead of scripts. I can likely translate necessary steps into Python - but if it's all the same, I would much rather submit a copy of the workflow as it can easily be opened and run in a free open-source download of KNIME (and it's my current strong suit as I progress in Python). There is an interesting Kernel mirroring KNIME and Python workflows on the Titanic data set if you're interested: https://www.kaggle.com/grayrat/scripting-a-model-from-knime-workflow-to-python",
      "votes": null
    },
    {
      "id": "450380",
      "postDate": "01/04/2019 20:24:09",
      "content": "<p>You should definitely covert to R or Python if you can, because:</p>\n\n<ol>\n<li>A kitten is born every time someone converts a workflow to a Kernel</li>\n<li>We have to validate hundreds of submissions, so \"easily reproducible\" is a factor</li>\n</ol>",
      "rawMarkdown": "You should definitely covert to R or Python if you can, because:\n\n1. A kitten is born every time someone converts a workflow to a Kernel\n2. We have to validate hundreds of submissions, so \"easily reproducible\" is a factor",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 446995,
      "author_name": "mhoch2",
      "author_url": "",
      "post_date": "12/29/2018 00:59:02",
      "content": "<p>Haven't read the whole rules but am annoyed by the kernel requirement as well. I'll probably end up doing what you're doing too but commenting to see if we get an official reply</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 447092,
      "author_name": "reidtc",
      "author_url": "",
      "post_date": "12/29/2018 05:55:12",
      "content": "<p>I wish my personal machine was better than a Kaggle kernel in more aspects than SSD space...</p>",
      "votes": null,
      "replies": [
        {
          "id": 447359,
          "author_name": "robikscube",
          "author_url": "",
          "post_date": "12/29/2018 17:55:22",
          "content": "<p>If you need more resources you could always spin up a Google Compute Engine or AWS virtual machine. It might cost a little money but if you only use it for heavy computation and turn it off when you're not using it the cost is pretty minimal. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 447295,
      "author_name": "miguelpm",
      "author_url": "",
      "post_date": "12/29/2018 15:27:52",
      "content": "<p>As long as complete generating code is provided I think there is no problem with using some of the most time consuming process as a linked dataset. </p>\n\n<p>I  will possibly do that for a subset of the calculations related to NGS data that required some heavy paralelized processing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 449285,
      "author_name": "crawford",
      "author_url": "",
      "post_date": "01/02/2019 22:45:05",
      "content": "<p>You should probably do all of your work in Kernels because it's really the only way to guarantee we have the same environment for reproducibility. With a Kernel, I can fork it and run it flawlessly every time. I would be amazed if you sent me a script for processing a dataset that ran perfectly on the first try. </p>",
      "votes": null,
      "replies": [
        {
          "id": 449322,
          "author_name": "rightmirec",
          "author_url": "",
          "post_date": "01/03/2019 00:43:16",
          "content": "<p>Noob question, but is it possible to use KNIME (similar to Alteryx or RapidMiner) for preprocessing and attach the workflow in the Kernel?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 449831,
          "author_name": "crawford",
          "author_url": "",
          "post_date": "01/03/2019 20:28:09",
          "content": "<p>Can you export the workflows to Python or R and run it in a kernel? This is the first time I've heard of KNIME.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 449904,
          "author_name": "rightmirec",
          "author_url": "",
          "post_date": "01/03/2019 23:39:31",
          "content": "<p>KNIME is an open source data workbench that lets you quickly process data in a visual format, instead of scripts. I can likely translate necessary steps into Python - but if it's all the same, I would much rather submit a copy of the workflow as it can easily be opened and run in a free open-source download of KNIME (and it's my current strong suit as I progress in Python). There is an interesting Kernel mirroring KNIME and Python workflows on the Titanic data set if you're interested: <a href=\"https://www.kaggle.com/grayrat/scripting-a-model-from-knime-workflow-to-python\">https://www.kaggle.com/grayrat/scripting-a-model-from-knime-workflow-to-python</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 450380,
          "author_name": "crawford",
          "author_url": "",
          "post_date": "01/04/2019 20:24:09",
          "content": "<p>You should definitely covert to R or Python if you can, because:</p>\n\n<ol>\n<li>A kitten is born every time someone converts a workflow to a Kernel</li>\n<li>We have to validate hundreds of submissions, so \"easily reproducible\" is a factor</li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "446407": "I've done a lot of data preprocessing/cleansing offline on my personal machine. Instead of having to run this preprocessing again through the kaggle kernel (which would be very slow and limited to the disk space of the kernel VM) - I was hoping to just upload the processed data as a separate dataset- which I would link to in my submission kernel.\n\nI would reveal all my preprocessing scripts so the analysis would be reproducible.\n\nThe rules state that `Your kernels should be easy to understand and the analysis should be reproducible.` \n\n- Would this approach meet the kernel requirement?\n- Are others having the same issue? If so- how do you plan to deal with it?",
    "446995": "Haven't read the whole rules but am annoyed by the kernel requirement as well. I'll probably end up doing what you're doing too but commenting to see if we get an official reply",
    "447092": "I wish my personal machine was better than a Kaggle kernel in more aspects than SSD space...",
    "447295": "As long as complete generating code is provided I think there is no problem with using some of the most time consuming process as a linked dataset. \n\nI  will possibly do that for a subset of the calculations related to NGS data that required some heavy paralelized processing.",
    "447359": "If you need more resources you could always spin up a Google Compute Engine or AWS virtual machine. It might cost a little money but if you only use it for heavy computation and turn it off when you're not using it the cost is pretty minimal.",
    "449285": "You should probably do all of your work in Kernels because it's really the only way to guarantee we have the same environment for reproducibility. With a Kernel, I can fork it and run it flawlessly every time. I would be amazed if you sent me a script for processing a dataset that ran perfectly on the first try.",
    "449322": "Noob question, but is it possible to use KNIME (similar to Alteryx or RapidMiner) for preprocessing and attach the workflow in the Kernel?",
    "449831": "Can you export the workflows to Python or R and run it in a kernel? This is the first time I've heard of KNIME.",
    "449904": "KNIME is an open source data workbench that lets you quickly process data in a visual format, instead of scripts. I can likely translate necessary steps into Python - but if it's all the same, I would much rather submit a copy of the workflow as it can easily be opened and run in a free open-source download of KNIME (and it's my current strong suit as I progress in Python). There is an interesting Kernel mirroring KNIME and Python workflows on the Titanic data set if you're interested: https://www.kaggle.com/grayrat/scripting-a-model-from-knime-workflow-to-python",
    "450380": "You should definitely covert to R or Python if you can, because:\n\n1. A kitten is born every time someone converts a workflow to a Kernel\n2. We have to validate hundreds of submissions, so \"easily reproducible\" is a factor"
  },
  "source": "meta"
}