{
  "id": 68460,
  "title": "How do you read test_set.csv ?",
  "url": "/competitions/PLAsTiCC-2018/discussion/68460",
  "author_name": "",
  "post_date": "2018-10-13T01:05:46.980040200Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello all,\nI tried to read test_set.csv in kaggle kernel, but I could not do that because of the error message such as “your kernel looks like dead. It restarts automatically”. \nThe kernel used 100% memory, so I think that the kernel could not read test_set.csv into memory completely.</p>\n\n<p>How do you read test_set.csv ?\nCould anyone read test_set.csv in kaggle kernels ?</p>",
  "messages": [
    {
      "id": "403154",
      "postDate": "10/13/2018 01:05:46",
      "content": "<p>Hello all,\nI tried to read test_set.csv in kaggle kernel, but I could not do that because of the error message such as “your kernel looks like dead. It restarts automatically”. \nThe kernel used 100% memory, so I think that the kernel could not read test_set.csv into memory completely.</p>\n\n<p>How do you read test_set.csv ?\nCould anyone read test_set.csv in kaggle kernels ?</p>",
      "rawMarkdown": "Hello all,\nI tried to read test_set.csv in kaggle kernel, but I could not do that because of the error message such as “your kernel looks like dead. It restarts automatically”. \nThe kernel used 100% memory, so I think that the kernel could not read test_set.csv into memory completely.\n\nHow do you read test_set.csv ?\nCould anyone read test_set.csv in kaggle kernels ?",
      "votes": null
    },
    {
      "id": "403174",
      "postDate": "10/13/2018 02:46:28",
      "content": "<p>Set <code>chunksize</code> <br>\n<a href=\"http://pandas.pydata.org/pandas-docs/stable/io.html#iterating-through-files-chunk-by-chunk\">http://pandas.pydata.org/pandas-docs/stable/io.html#iterating-through-files-chunk-by-chunk</a> <br>\nAnd due to 6 hours limit of kernel running time, maybe it is better separate test.csv into n chunks and use n kernels to process them in parallel</p>",
      "rawMarkdown": "Set ```chunksize```  \nhttp://pandas.pydata.org/pandas-docs/stable/io.html#iterating-through-files-chunk-by-chunk  \nAnd due to 6 hours limit of kernel running time, maybe it is better separate test.csv into n chunks and use n kernels to process them in parallel",
      "votes": null
    },
    {
      "id": "403321",
      "postDate": "10/13/2018 09:56:16",
      "content": "<p>Thank you for your reply !  I will try it.</p>",
      "rawMarkdown": "Thank you for your reply !  I will try it.",
      "votes": null
    },
    {
      "id": "406446",
      "postDate": "10/19/2018 09:00:24",
      "content": "<p>Another solution maybe <a href=\"http://docs.dask.org/en/latest/dataframe.html\">dask dataframe</a>. I am also trying this package.</p>",
      "rawMarkdown": "Another solution maybe [dask dataframe][1]. I am also trying this package.\n  [1]: http://docs.dask.org/en/latest/dataframe.html",
      "votes": null
    },
    {
      "id": "408371",
      "postDate": "10/22/2018 18:32:17",
      "content": "<p>Check <a href=\"https://www.kaggle.com/alexfir/fast-test-set-reading\">this</a> kernel, maybe this helps.</p>",
      "rawMarkdown": "Check [this][1] kernel, maybe this helps.\n\n\n  [1]: https://www.kaggle.com/alexfir/fast-test-set-reading",
      "votes": null
    },
    {
      "id": "408855",
      "postDate": "10/23/2018 14:45:07",
      "content": "<p>Thank you for sharing kernels !  I forked your kernels and I am reading them :)\nMy another idea was that:</p>\n\n<ol>\n<li><p>read chunk of data by using \"nrows\" and \"skiprows\"</p></li>\n<li><p>process (ex. predict a target) for the chunk</p></li>\n<li><p>delete the chunk</p></li>\n<li><p>repeat No.1 to No.3</p></li>\n</ol>\n\n<p>(these process may be able to code using \"chunksize\" in the reply above ...)</p>",
      "rawMarkdown": "Thank you for sharing kernels !  I forked your kernels and I am reading them :)\nMy another idea was that:\n\n 1.  read chunk of data by using \"nrows\" and \"skiprows\"\n\n 2. process (ex. predict a target) for the chunk\n\n 3. delete the chunk\n\n 4. repeat No.1 to No.3\n\n(these process may be able to code using \"chunksize\" in the reply above ...)",
      "votes": null
    },
    {
      "id": "410334",
      "postDate": "10/25/2018 21:23:41",
      "content": "<p>If you are trying to read the test set on your local machine you might take a look at the discussion in <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69392#410333\">this thread</a>. </p>",
      "rawMarkdown": "If you are trying to read the test set on your local machine you might take a look at the discussion in [this thread][1]. \n\n  [1]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69392#410333",
      "votes": null
    },
    {
      "id": "412628",
      "postDate": "10/30/2018 14:17:41",
      "content": "<p>Thank you for your reply. I read the test set on kaggle kernel, but it may be effective that I prepare my own machine :)</p>",
      "rawMarkdown": "Thank you for your reply. I read the test set on kaggle kernel, but it may be effective that I prepare my own machine :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 403174,
      "author_name": "johnfarrell",
      "author_url": "",
      "post_date": "10/13/2018 02:46:28",
      "content": "<p>Set <code>chunksize</code> <br>\n<a href=\"http://pandas.pydata.org/pandas-docs/stable/io.html#iterating-through-files-chunk-by-chunk\">http://pandas.pydata.org/pandas-docs/stable/io.html#iterating-through-files-chunk-by-chunk</a> <br>\nAnd due to 6 hours limit of kernel running time, maybe it is better separate test.csv into n chunks and use n kernels to process them in parallel</p>",
      "votes": null,
      "replies": [
        {
          "id": 403321,
          "author_name": "yk0711",
          "author_url": "",
          "post_date": "10/13/2018 09:56:16",
          "content": "<p>Thank you for your reply !  I will try it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 406446,
      "author_name": "yk0711",
      "author_url": "",
      "post_date": "10/19/2018 09:00:24",
      "content": "<p>Another solution maybe <a href=\"http://docs.dask.org/en/latest/dataframe.html\">dask dataframe</a>. I am also trying this package.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 408371,
      "author_name": "alexfir",
      "author_url": "",
      "post_date": "10/22/2018 18:32:17",
      "content": "<p>Check <a href=\"https://www.kaggle.com/alexfir/fast-test-set-reading\">this</a> kernel, maybe this helps.</p>",
      "votes": null,
      "replies": [
        {
          "id": 408855,
          "author_name": "yk0711",
          "author_url": "",
          "post_date": "10/23/2018 14:45:07",
          "content": "<p>Thank you for sharing kernels !  I forked your kernels and I am reading them :)\nMy another idea was that:</p>\n\n<ol>\n<li><p>read chunk of data by using \"nrows\" and \"skiprows\"</p></li>\n<li><p>process (ex. predict a target) for the chunk</p></li>\n<li><p>delete the chunk</p></li>\n<li><p>repeat No.1 to No.3</p></li>\n</ol>\n\n<p>(these process may be able to code using \"chunksize\" in the reply above ...)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 410334,
      "author_name": "smallyellowduck",
      "author_url": "",
      "post_date": "10/25/2018 21:23:41",
      "content": "<p>If you are trying to read the test set on your local machine you might take a look at the discussion in <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69392#410333\">this thread</a>. </p>",
      "votes": null,
      "replies": [
        {
          "id": 412628,
          "author_name": "yk0711",
          "author_url": "",
          "post_date": "10/30/2018 14:17:41",
          "content": "<p>Thank you for your reply. I read the test set on kaggle kernel, but it may be effective that I prepare my own machine :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "403154": "Hello all,\nI tried to read test_set.csv in kaggle kernel, but I could not do that because of the error message such as “your kernel looks like dead. It restarts automatically”. \nThe kernel used 100% memory, so I think that the kernel could not read test_set.csv into memory completely.\n\nHow do you read test_set.csv ?\nCould anyone read test_set.csv in kaggle kernels ?",
    "403174": "Set ```chunksize```  \nhttp://pandas.pydata.org/pandas-docs/stable/io.html#iterating-through-files-chunk-by-chunk  \nAnd due to 6 hours limit of kernel running time, maybe it is better separate test.csv into n chunks and use n kernels to process them in parallel",
    "403321": "Thank you for your reply !  I will try it.",
    "406446": "Another solution maybe [dask dataframe][1]. I am also trying this package.\n  [1]: http://docs.dask.org/en/latest/dataframe.html",
    "408371": "Check [this][1] kernel, maybe this helps.\n\n\n  [1]: https://www.kaggle.com/alexfir/fast-test-set-reading",
    "408855": "Thank you for sharing kernels !  I forked your kernels and I am reading them :)\nMy another idea was that:\n\n 1.  read chunk of data by using \"nrows\" and \"skiprows\"\n\n 2. process (ex. predict a target) for the chunk\n\n 3. delete the chunk\n\n 4. repeat No.1 to No.3\n\n(these process may be able to code using \"chunksize\" in the reply above ...)",
    "410334": "If you are trying to read the test set on your local machine you might take a look at the discussion in [this thread][1]. \n\n  [1]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69392#410333",
    "412628": "Thank you for your reply. I read the test set on kaggle kernel, but it may be effective that I prepare my own machine :)"
  },
  "source": "meta"
}