{
  "id": 39750,
  "title": "How to submit and how to work with data",
  "url": "/competitions/kkbox-churn-prediction-challenge/discussion/39750",
  "author_name": "",
  "post_date": "2017-09-20T12:25:45.411977200Z",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hello! I am a novice in machine learning. It's my first kaggle competition. And I want to ask you some (stupid) questions. </p>\n\n<p>1) My laptop has not much RAM. I noticed that the dataset for this competition is more than 6GB. How will I be able to work with this data if my computer has, for example, only 4GB RAM? Is there some trick to read data by chunks? </p>\n\n<p>2) I have heard about kernels. If I have correct understanding, I can work with data not on my local computer but on the server which is provided by Kaggle. It this correct? </p>\n\n<p>3) If I will work with data on the Kaggle's server, how then I will be able to submit my solution (predictions)? I tried to do this in Titanic competition and was not able to upload submission from kernel. </p>\n\n<p>I'm sorry for such questions but I really need help to start. </p>",
  "messages": [
    {
      "id": "222871",
      "postDate": "09/20/2017 12:25:45",
      "content": "<p>Hello! I am a novice in machine learning. It's my first kaggle competition. And I want to ask you some (stupid) questions. </p>\n\n<p>1) My laptop has not much RAM. I noticed that the dataset for this competition is more than 6GB. How will I be able to work with this data if my computer has, for example, only 4GB RAM? Is there some trick to read data by chunks? </p>\n\n<p>2) I have heard about kernels. If I have correct understanding, I can work with data not on my local computer but on the server which is provided by Kaggle. It this correct? </p>\n\n<p>3) If I will work with data on the Kaggle's server, how then I will be able to submit my solution (predictions)? I tried to do this in Titanic competition and was not able to upload submission from kernel. </p>\n\n<p>I'm sorry for such questions but I really need help to start. </p>",
      "rawMarkdown": "Hello! I am a novice in machine learning. It's my first kaggle competition. And I want to ask you some (stupid) questions. \n\n1) My laptop has not much RAM. I noticed that the dataset for this competition is more than 6GB. How will I be able to work with this data if my computer has, for example, only 4GB RAM? Is there some trick to read data by chunks? \n\n2) I have heard about kernels. If I have correct understanding, I can work with data not on my local computer but on the server which is provided by Kaggle. It this correct? \n\n3) If I will work with data on the Kaggle's server, how then I will be able to submit my solution (predictions)? I tried to do this in Titanic competition and was not able to upload submission from kernel. \n\nI'm sorry for such questions but I really need help to start.",
      "votes": null
    },
    {
      "id": "222930",
      "postDate": "09/20/2017 16:13:47",
      "content": "<p>Just to answer some of your questions.</p>\n\n<p>1) My laptop has not much RAM. I noticed that the dataset for this competition is more than 6GB. How will I be able to work with this data if my computer has, for example, only 4GB RAM? Is there some trick to read data by chunks?</p>\n\n<p>Yes, generators are a good starting point. There is a lot of examples on how to implement one in previous competitions and in competitions that are still going on. If you go through the forums and kernels of previous competitions you'll find a lot of code and examples.</p>\n\n<p>2) I have heard about kernels. If I have correct understanding, I can work with data not on my local computer but on the server which is provided by Kaggle. It this correct?</p>\n\n<p>Yes that is correct. The kernels essentially are docker environments which when press the button to create a new kernel it comes pre-configured with the data usually in the path <code>/input/data/</code></p>\n\n<p>3) If I will work with data on the Kaggle's server, how then I will be able to submit my solution (predictions)? I tried to do this in Titanic competition and was not able to upload submission from kernel.</p>\n\n<p>I don't know about that because I never had this need before. But, I'm sure that if your search a little bit in the forums of also previous competitions you'll always find sth that might help you out.</p>",
      "rawMarkdown": "Just to answer some of your questions.\n\n1) My laptop has not much RAM. I noticed that the dataset for this competition is more than 6GB. How will I be able to work with this data if my computer has, for example, only 4GB RAM? Is there some trick to read data by chunks?\n\nYes, generators are a good starting point. There is a lot of examples on how to implement one in previous competitions and in competitions that are still going on. If you go through the forums and kernels of previous competitions you'll find a lot of code and examples.\n\n2) I have heard about kernels. If I have correct understanding, I can work with data not on my local computer but on the server which is provided by Kaggle. It this correct?\n\nYes that is correct. The kernels essentially are docker environments which when press the button to create a new kernel it comes pre-configured with the data usually in the path `/input/data/`\n\n3) If I will work with data on the Kaggle's server, how then I will be able to submit my solution (predictions)? I tried to do this in Titanic competition and was not able to upload submission from kernel.\n\nI don't know about that because I never had this need before. But, I'm sure that if your search a little bit in the forums of also previous competitions you'll always find sth that might help you out.",
      "votes": null
    },
    {
      "id": "222939",
      "postDate": "09/20/2017 16:36:15",
      "content": "<p>Thank you for the explanation! I have already found how to use and get output from the kernels. One last question: what is the benefit to run all the code on a local machine compare to the Kaggle's servers? </p>",
      "rawMarkdown": "Thank you for the explanation! I have already found how to use and get output from the kernels. One last question: what is the benefit to run all the code on a local machine compare to the Kaggle's servers?",
      "votes": null
    },
    {
      "id": "222969",
      "postDate": "09/20/2017 18:27:58",
      "content": "<p>Kernels timeout on the Kaggle servers. I don't remember what the limit is (maybe 20-30 minutes?) --but you will be limited.</p>\n\n<p>I recommend using an online service like Microsoft's Azure ML Studio, AWS, Google, IBM, etc. Having some cloud experience is always a plus.</p>",
      "rawMarkdown": "Kernels timeout on the Kaggle servers. I don't remember what the limit is (maybe 20-30 minutes?) --but you will be limited.\n\nI recommend using an online service like Microsoft's Azure ML Studio, AWS, Google, IBM, etc. Having some cloud experience is always a plus.",
      "votes": null
    },
    {
      "id": "222984",
      "postDate": "09/20/2017 19:13:19",
      "content": "<p>Regarding point 3: if you save your solution as CSV within the kernel, and run the kernel afterwards, you'll be able to submit/download the file. See this kernel for example: <a href=\"https://www.kaggle.com/opanichev/mean-baseline-lb-0-30786\">https://www.kaggle.com/opanichev/mean-baseline-lb-0-30786</a></p>",
      "rawMarkdown": "Regarding point 3: if you save your solution as CSV within the kernel, and run the kernel afterwards, you'll be able to submit/download the file. See this kernel for example: https://www.kaggle.com/opanichev/mean-baseline-lb-0-30786",
      "votes": null
    },
    {
      "id": "222995",
      "postDate": "09/20/2017 20:06:09",
      "content": "<p>Yes, I have seen this kernel and it helped me to understand how to use kernels. Thank you</p>",
      "rawMarkdown": "Yes, I have seen this kernel and it helped me to understand how to use kernels. Thank you",
      "votes": null
    },
    {
      "id": "222996",
      "postDate": "09/20/2017 20:06:50",
      "content": "<p>Are there some free platforms?</p>",
      "rawMarkdown": "Are there some free platforms?",
      "votes": null
    },
    {
      "id": "222998",
      "postDate": "09/20/2017 20:14:37",
      "content": "<p>Can somebody give me an idea of how to merge files to create unite data frame? In particular, I'm interested in merging data frames where the first dataframe has column with unique id for each row, and the second data frame has more than one row with the same ids. For example, in this competition, data frames <em>transactions</em> and *users_log* for each user have several rows with corresponding information. How can I create the data frame where the user's id will be unique for each row but also <strong>all</strong> information about this user from other files will be included? </p>",
      "rawMarkdown": "Can somebody give me an idea of how to merge files to create unite data frame? In particular, I'm interested in merging data frames where the first dataframe has column with unique id for each row, and the second data frame has more than one row with the same ids. For example, in this competition, data frames *transactions* and *users_log* for each user have several rows with corresponding information. How can I create the data frame where the user's id will be unique for each row but also **all** information about this user from other files will be included?",
      "votes": null
    },
    {
      "id": "223133",
      "postDate": "09/21/2017 07:48:38",
      "content": "<p>I just noticed <a href=\"https://www.kaggle.com/product-feedback/39790\">this</a> thread. Kernel resources have been increased significantly. This may fulfill your needs.</p>",
      "rawMarkdown": "I just noticed [this][1] thread. Kernel resources have been increased significantly. This may fulfill your needs.\n\n\n  [1]: https://www.kaggle.com/product-feedback/39790",
      "votes": null
    },
    {
      "id": "246347",
      "postDate": "11/21/2017 01:23:07",
      "content": "<p>Have you figured this out?</p>",
      "rawMarkdown": "Have you figured this out?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 222930,
      "author_name": "kirk86",
      "author_url": "",
      "post_date": "09/20/2017 16:13:47",
      "content": "<p>Just to answer some of your questions.</p>\n\n<p>1) My laptop has not much RAM. I noticed that the dataset for this competition is more than 6GB. How will I be able to work with this data if my computer has, for example, only 4GB RAM? Is there some trick to read data by chunks?</p>\n\n<p>Yes, generators are a good starting point. There is a lot of examples on how to implement one in previous competitions and in competitions that are still going on. If you go through the forums and kernels of previous competitions you'll find a lot of code and examples.</p>\n\n<p>2) I have heard about kernels. If I have correct understanding, I can work with data not on my local computer but on the server which is provided by Kaggle. It this correct?</p>\n\n<p>Yes that is correct. The kernels essentially are docker environments which when press the button to create a new kernel it comes pre-configured with the data usually in the path <code>/input/data/</code></p>\n\n<p>3) If I will work with data on the Kaggle's server, how then I will be able to submit my solution (predictions)? I tried to do this in Titanic competition and was not able to upload submission from kernel.</p>\n\n<p>I don't know about that because I never had this need before. But, I'm sure that if your search a little bit in the forums of also previous competitions you'll always find sth that might help you out.</p>",
      "votes": null,
      "replies": [
        {
          "id": 222984,
          "author_name": "kevinbonnes",
          "author_url": "",
          "post_date": "09/20/2017 19:13:19",
          "content": "<p>Regarding point 3: if you save your solution as CSV within the kernel, and run the kernel afterwards, you'll be able to submit/download the file. See this kernel for example: <a href=\"https://www.kaggle.com/opanichev/mean-baseline-lb-0-30786\">https://www.kaggle.com/opanichev/mean-baseline-lb-0-30786</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 222995,
          "author_name": "fesenkod",
          "author_url": "",
          "post_date": "09/20/2017 20:06:09",
          "content": "<p>Yes, I have seen this kernel and it helped me to understand how to use kernels. Thank you</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 222939,
      "author_name": "fesenkod",
      "author_url": "",
      "post_date": "09/20/2017 16:36:15",
      "content": "<p>Thank you for the explanation! I have already found how to use and get output from the kernels. One last question: what is the benefit to run all the code on a local machine compare to the Kaggle's servers? </p>",
      "votes": null,
      "replies": [
        {
          "id": 222969,
          "author_name": "puremath86",
          "author_url": "",
          "post_date": "09/20/2017 18:27:58",
          "content": "<p>Kernels timeout on the Kaggle servers. I don't remember what the limit is (maybe 20-30 minutes?) --but you will be limited.</p>\n\n<p>I recommend using an online service like Microsoft's Azure ML Studio, AWS, Google, IBM, etc. Having some cloud experience is always a plus.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 222996,
          "author_name": "fesenkod",
          "author_url": "",
          "post_date": "09/20/2017 20:06:50",
          "content": "<p>Are there some free platforms?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 222998,
      "author_name": "fesenkod",
      "author_url": "",
      "post_date": "09/20/2017 20:14:37",
      "content": "<p>Can somebody give me an idea of how to merge files to create unite data frame? In particular, I'm interested in merging data frames where the first dataframe has column with unique id for each row, and the second data frame has more than one row with the same ids. For example, in this competition, data frames <em>transactions</em> and *users_log* for each user have several rows with corresponding information. How can I create the data frame where the user's id will be unique for each row but also <strong>all</strong> information about this user from other files will be included? </p>",
      "votes": null,
      "replies": [
        {
          "id": 246347,
          "author_name": "davischumacher",
          "author_url": "",
          "post_date": "11/21/2017 01:23:07",
          "content": "<p>Have you figured this out?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 223133,
      "author_name": "kevinbonnes",
      "author_url": "",
      "post_date": "09/21/2017 07:48:38",
      "content": "<p>I just noticed <a href=\"https://www.kaggle.com/product-feedback/39790\">this</a> thread. Kernel resources have been increased significantly. This may fulfill your needs.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "222871": "Hello! I am a novice in machine learning. It's my first kaggle competition. And I want to ask you some (stupid) questions. \n\n1) My laptop has not much RAM. I noticed that the dataset for this competition is more than 6GB. How will I be able to work with this data if my computer has, for example, only 4GB RAM? Is there some trick to read data by chunks? \n\n2) I have heard about kernels. If I have correct understanding, I can work with data not on my local computer but on the server which is provided by Kaggle. It this correct? \n\n3) If I will work with data on the Kaggle's server, how then I will be able to submit my solution (predictions)? I tried to do this in Titanic competition and was not able to upload submission from kernel. \n\nI'm sorry for such questions but I really need help to start.",
    "222930": "Just to answer some of your questions.\n\n1) My laptop has not much RAM. I noticed that the dataset for this competition is more than 6GB. How will I be able to work with this data if my computer has, for example, only 4GB RAM? Is there some trick to read data by chunks?\n\nYes, generators are a good starting point. There is a lot of examples on how to implement one in previous competitions and in competitions that are still going on. If you go through the forums and kernels of previous competitions you'll find a lot of code and examples.\n\n2) I have heard about kernels. If I have correct understanding, I can work with data not on my local computer but on the server which is provided by Kaggle. It this correct?\n\nYes that is correct. The kernels essentially are docker environments which when press the button to create a new kernel it comes pre-configured with the data usually in the path `/input/data/`\n\n3) If I will work with data on the Kaggle's server, how then I will be able to submit my solution (predictions)? I tried to do this in Titanic competition and was not able to upload submission from kernel.\n\nI don't know about that because I never had this need before. But, I'm sure that if your search a little bit in the forums of also previous competitions you'll always find sth that might help you out.",
    "222939": "Thank you for the explanation! I have already found how to use and get output from the kernels. One last question: what is the benefit to run all the code on a local machine compare to the Kaggle's servers?",
    "222969": "Kernels timeout on the Kaggle servers. I don't remember what the limit is (maybe 20-30 minutes?) --but you will be limited.\n\nI recommend using an online service like Microsoft's Azure ML Studio, AWS, Google, IBM, etc. Having some cloud experience is always a plus.",
    "222984": "Regarding point 3: if you save your solution as CSV within the kernel, and run the kernel afterwards, you'll be able to submit/download the file. See this kernel for example: https://www.kaggle.com/opanichev/mean-baseline-lb-0-30786",
    "222995": "Yes, I have seen this kernel and it helped me to understand how to use kernels. Thank you",
    "222996": "Are there some free platforms?",
    "222998": "Can somebody give me an idea of how to merge files to create unite data frame? In particular, I'm interested in merging data frames where the first dataframe has column with unique id for each row, and the second data frame has more than one row with the same ids. For example, in this competition, data frames *transactions* and *users_log* for each user have several rows with corresponding information. How can I create the data frame where the user's id will be unique for each row but also **all** information about this user from other files will be included?",
    "223133": "I just noticed [this][1] thread. Kernel resources have been increased significantly. This may fulfill your needs.\n\n\n  [1]: https://www.kaggle.com/product-feedback/39790",
    "246347": "Have you figured this out?"
  },
  "source": "meta"
}