{
  "id": 62534,
  "title": "n_jobs ignored in python dbscan ",
  "url": "/competitions/trackml-particle-identification/discussion/62534",
  "author_name": "",
  "post_date": "2018-08-02T21:52:29.640543700Z",
  "votes": 2,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I tried to set n_jobs = -1 or n_jobs= n (a number) in my 8-cpu-core computer, but there is always one cpu working. I googled the issue but didn't find any solutions. I don't know if anyone has similar observation or suggestions. Thanks a lot.</p>",
  "messages": [
    {
      "id": "365548",
      "postDate": "08/02/2018 21:52:29",
      "content": "<p>I tried to set n_jobs = -1 or n_jobs= n (a number) in my 8-cpu-core computer, but there is always one cpu working. I googled the issue but didn't find any solutions. I don't know if anyone has similar observation or suggestions. Thanks a lot.</p>",
      "rawMarkdown": "I tried to set n_jobs = -1 or n_jobs= n (a number) in my 8-cpu-core computer, but there is always one cpu working. I googled the issue but didn't find any solutions. I don't know if anyone has similar observation or suggestions. Thanks a lot.",
      "votes": null
    },
    {
      "id": "365567",
      "postDate": "08/02/2018 23:03:10",
      "content": "<p>If using the kd_tree algorithm n_jobs will be ignored as parallelism has not been implemented for kd_tree yet. There is an issue open on GitHub for it <a href=\"https://github.com/scikit-learn/scikit-learn/issues/8003\">https://github.com/scikit-learn/scikit-learn/issues/8003</a></p>",
      "rawMarkdown": "If using the kd_tree algorithm n_jobs will be ignored as parallelism has not been implemented for kd_tree yet. There is an issue open on GitHub for it https://github.com/scikit-learn/scikit-learn/issues/8003",
      "votes": null
    },
    {
      "id": "365591",
      "postDate": "08/03/2018 01:11:51",
      "content": "<p>Thanks @Jack. I did try {algorithm=\"brute\"}, however it didn't work. </p>",
      "rawMarkdown": "Thanks @Jack. I did try {algorithm=\"brute\"}, however it didn't work.",
      "votes": null
    },
    {
      "id": "365641",
      "postDate": "08/03/2018 04:57:13",
      "content": "<p>Maybe try HDBSCAN? I've heard good things about it, both for the effectiveness and the speed. The API doc is at <a href=\"https://hdbscan.readthedocs.io/en/latest/api.html\">https://hdbscan.readthedocs.io/en/latest/api.html</a> </p>",
      "rawMarkdown": "Maybe try HDBSCAN? I've heard good things about it, both for the effectiveness and the speed. The API doc is at https://hdbscan.readthedocs.io/en/latest/api.html",
      "votes": null
    },
    {
      "id": "365646",
      "postDate": "08/03/2018 05:11:04",
      "content": "<p>Probably you don't need it. You should run a lot of dbscans. So you can parallelize them, not the dbscan itself.</p>",
      "rawMarkdown": "Probably you don't need it. You should run a lot of dbscans. So you can parallelize them, not the dbscan itself.",
      "votes": null
    },
    {
      "id": "365648",
      "postDate": "08/03/2018 05:19:29",
      "content": "<p>I am actually doing that. But very often I have to just wait the events that take time significantly longer than others to finish when most of cpus are idle. </p>",
      "rawMarkdown": "I am actually doing that. But very often I have to just wait the events that take time significantly longer than others to finish when most of cpus are idle.",
      "votes": null
    },
    {
      "id": "365650",
      "postDate": "08/03/2018 05:22:21",
      "content": "<p>Do you run a lot of dbscans inside one event? Like 1000 or even more, like many of us do.</p>",
      "rawMarkdown": "Do you run a lot of dbscans inside one event? Like 1000 or even more, like many of us do.",
      "votes": null
    },
    {
      "id": "365652",
      "postDate": "08/03/2018 05:28:27",
      "content": "<p>I indeed run a lot of dbscans in each event, 1800 currently.</p>",
      "rawMarkdown": "I indeed run a lot of dbscans in each event, 1800 currently.",
      "votes": null
    },
    {
      "id": "365653",
      "postDate": "08/03/2018 05:40:48",
      "content": "<p>So you can parallelize these 1800 inside one event.</p>",
      "rawMarkdown": "So you can parallelize these 1800 inside one event.",
      "votes": null
    },
    {
      "id": "365654",
      "postDate": "08/03/2018 05:44:41",
      "content": "<p>I see. Thanks @Sergey Zlobin</p>",
      "rawMarkdown": "I see. Thanks @Sergey Zlobin",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 365567,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "08/02/2018 23:03:10",
      "content": "<p>If using the kd_tree algorithm n_jobs will be ignored as parallelism has not been implemented for kd_tree yet. There is an issue open on GitHub for it <a href=\"https://github.com/scikit-learn/scikit-learn/issues/8003\">https://github.com/scikit-learn/scikit-learn/issues/8003</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 365591,
          "author_name": "ybwu01",
          "author_url": "",
          "post_date": "08/03/2018 01:11:51",
          "content": "<p>Thanks @Jack. I did try {algorithm=\"brute\"}, however it didn't work. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 365641,
      "author_name": "jpmiller",
      "author_url": "",
      "post_date": "08/03/2018 04:57:13",
      "content": "<p>Maybe try HDBSCAN? I've heard good things about it, both for the effectiveness and the speed. The API doc is at <a href=\"https://hdbscan.readthedocs.io/en/latest/api.html\">https://hdbscan.readthedocs.io/en/latest/api.html</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 365646,
      "author_name": "sergeyzlobin",
      "author_url": "",
      "post_date": "08/03/2018 05:11:04",
      "content": "<p>Probably you don't need it. You should run a lot of dbscans. So you can parallelize them, not the dbscan itself.</p>",
      "votes": null,
      "replies": [
        {
          "id": 365648,
          "author_name": "ybwu01",
          "author_url": "",
          "post_date": "08/03/2018 05:19:29",
          "content": "<p>I am actually doing that. But very often I have to just wait the events that take time significantly longer than others to finish when most of cpus are idle. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 365650,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "08/03/2018 05:22:21",
          "content": "<p>Do you run a lot of dbscans inside one event? Like 1000 or even more, like many of us do.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 365652,
          "author_name": "ybwu01",
          "author_url": "",
          "post_date": "08/03/2018 05:28:27",
          "content": "<p>I indeed run a lot of dbscans in each event, 1800 currently.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 365653,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "08/03/2018 05:40:48",
          "content": "<p>So you can parallelize these 1800 inside one event.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 365654,
          "author_name": "ybwu01",
          "author_url": "",
          "post_date": "08/03/2018 05:44:41",
          "content": "<p>I see. Thanks @Sergey Zlobin</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "365548": "I tried to set n_jobs = -1 or n_jobs= n (a number) in my 8-cpu-core computer, but there is always one cpu working. I googled the issue but didn't find any solutions. I don't know if anyone has similar observation or suggestions. Thanks a lot.",
    "365567": "If using the kd_tree algorithm n_jobs will be ignored as parallelism has not been implemented for kd_tree yet. There is an issue open on GitHub for it https://github.com/scikit-learn/scikit-learn/issues/8003",
    "365591": "Thanks @Jack. I did try {algorithm=\"brute\"}, however it didn't work.",
    "365641": "Maybe try HDBSCAN? I've heard good things about it, both for the effectiveness and the speed. The API doc is at https://hdbscan.readthedocs.io/en/latest/api.html",
    "365646": "Probably you don't need it. You should run a lot of dbscans. So you can parallelize them, not the dbscan itself.",
    "365648": "I am actually doing that. But very often I have to just wait the events that take time significantly longer than others to finish when most of cpus are idle.",
    "365650": "Do you run a lot of dbscans inside one event? Like 1000 or even more, like many of us do.",
    "365652": "I indeed run a lot of dbscans in each event, 1800 currently.",
    "365653": "So you can parallelize these 1800 inside one event.",
    "365654": "I see. Thanks @Sergey Zlobin"
  },
  "source": "meta"
}