{
  "id": 79389,
  "title": "How much time are you guys spending on preprocessing?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/79389",
  "author_name": "",
  "post_date": "2019-02-03T13:53:20.564528600Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I guess a lot of people are still using forks from the public kernels.  How much time are you spending on the load_and_prec() function? I'm spending like a 1000 seconds on it thinking if I should skip it.</p>",
  "messages": [
    {
      "id": "465566",
      "postDate": "02/03/2019 13:53:20",
      "content": "<p>I guess a lot of people are still using forks from the public kernels.  How much time are you spending on the load_and_prec() function? I'm spending like a 1000 seconds on it thinking if I should skip it.</p>",
      "rawMarkdown": "I guess a lot of people are still using forks from the public kernels.  How much time are you spending on the load_and_prec() function? I'm spending like a 1000 seconds on it thinking if I should skip it.",
      "votes": null
    },
    {
      "id": "465687",
      "postDate": "02/03/2019 19:20:47",
      "content": "<p>You could try multiprocessing, which lowers my preprocessing time from 400 to 330 seconds. It doesn't help a lot since we only have 2 CPU cores available on GPU kernels. <br>\n(The code is quite comprehensible here: <a href=\"http://www.racketracer.com/2016/07/06/pandas-in-parallel/\">http://www.racketracer.com/2016/07/06/pandas-in-parallel/</a>)</p>\n\n<p>And for people who use PyTorch, the parameter 'num_workers' of DataLoader should save you some time as well.</p>",
      "rawMarkdown": "You could try multiprocessing, which lowers my preprocessing time from 400 to 330 seconds. It doesn't help a lot since we only have 2 CPU cores available on GPU kernels.  \n(The code is quite comprehensible here: http://www.racketracer.com/2016/07/06/pandas-in-parallel/)\n\nAnd for people who use PyTorch, the parameter 'num_workers' of DataLoader should save you some time as well.",
      "votes": null
    },
    {
      "id": "465770",
      "postDate": "02/04/2019 01:18:20",
      "content": "<p>Thanks. I'll see what changes I can make</p>",
      "rawMarkdown": "Thanks. I'll see what changes I can make",
      "votes": null
    },
    {
      "id": "466109",
      "postDate": "02/04/2019 17:44:03",
      "content": "<p>num_workers &gt; 0 causes the excecution hang for a long time....i encounted this bug both on my local server and the kernel....</p>",
      "rawMarkdown": "num_workers &gt; 0 causes the excecution hang for a long time....i encounted this bug both on my local server and the kernel....",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 465687,
      "author_name": "jerrykuo7727",
      "author_url": "",
      "post_date": "02/03/2019 19:20:47",
      "content": "<p>You could try multiprocessing, which lowers my preprocessing time from 400 to 330 seconds. It doesn't help a lot since we only have 2 CPU cores available on GPU kernels. <br>\n(The code is quite comprehensible here: <a href=\"http://www.racketracer.com/2016/07/06/pandas-in-parallel/\">http://www.racketracer.com/2016/07/06/pandas-in-parallel/</a>)</p>\n\n<p>And for people who use PyTorch, the parameter 'num_workers' of DataLoader should save you some time as well.</p>",
      "votes": null,
      "replies": [
        {
          "id": 465770,
          "author_name": "icemekaveli",
          "author_url": "",
          "post_date": "02/04/2019 01:18:20",
          "content": "<p>Thanks. I'll see what changes I can make</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 466109,
          "author_name": "tiopon",
          "author_url": "",
          "post_date": "02/04/2019 17:44:03",
          "content": "<p>num_workers &gt; 0 causes the excecution hang for a long time....i encounted this bug both on my local server and the kernel....</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "465566": "I guess a lot of people are still using forks from the public kernels.  How much time are you spending on the load_and_prec() function? I'm spending like a 1000 seconds on it thinking if I should skip it.",
    "465687": "You could try multiprocessing, which lowers my preprocessing time from 400 to 330 seconds. It doesn't help a lot since we only have 2 CPU cores available on GPU kernels.  \n(The code is quite comprehensible here: http://www.racketracer.com/2016/07/06/pandas-in-parallel/)\n\nAnd for people who use PyTorch, the parameter 'num_workers' of DataLoader should save you some time as well.",
    "465770": "Thanks. I'll see what changes I can make",
    "466109": "num_workers &gt; 0 causes the excecution hang for a long time....i encounted this bug both on my local server and the kernel...."
  },
  "source": "meta"
}