{
  "id": 61960,
  "title": "RAM usage for dbscan",
  "url": "/competitions/trackml-particle-identification/discussion/61960",
  "author_name": "",
  "post_date": "2018-07-25T15:20:33.264709400Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>How much ram should I need for helix unrolling and zshifting? I have 52gb and almost all of it is used up when I run ~20000 iterations of dbscan. Is there some way to optimize dbscan to use less ram?</p>",
  "messages": [
    {
      "id": "362053",
      "postDate": "07/25/2018 15:20:33",
      "content": "<p>How much ram should I need for helix unrolling and zshifting? I have 52gb and almost all of it is used up when I run ~20000 iterations of dbscan. Is there some way to optimize dbscan to use less ram?</p>",
      "rawMarkdown": "How much ram should I need for helix unrolling and zshifting? I have 52gb and almost all of it is used up when I run ~20000 iterations of dbscan. Is there some way to optimize dbscan to use less ram?",
      "votes": null
    },
    {
      "id": "362095",
      "postDate": "07/25/2018 17:12:15",
      "content": "<p>cut the iteration job , save the result to disk and load back when selecting candidate tracks</p>",
      "rawMarkdown": "cut the iteration job , save the result to disk and load back when selecting candidate tracks",
      "votes": null
    },
    {
      "id": "362310",
      "postDate": "07/26/2018 06:24:36",
      "content": "<p>I don't have any memory issues. Probably a memory leak in your code, not in dbscan?</p>",
      "rawMarkdown": "I don't have any memory issues. Probably a memory leak in your code, not in dbscan?",
      "votes": null
    },
    {
      "id": "362317",
      "postDate": "07/26/2018 06:42:10",
      "content": "<p>I run in few GB.  You probably keep all the clusters from each DBSCAN.  Just do the math, and see why your memory will explode.  You have to merge on the fly if you run that many DBSCAN.</p>",
      "rawMarkdown": "I run in few GB.  You probably keep all the clusters from each DBSCAN.  Just do the math, and see why your memory will explode.  You have to merge on the fly if you run that many DBSCAN.",
      "votes": null
    },
    {
      "id": "362671",
      "postDate": "07/26/2018 21:24:44",
      "content": "<p>Update:  </p>\n\n<p>Saving the iterations to disk, then loading the iterations for merging, reduces ram usage to 2-3 GB. However, merging track candidates as they are generated sounds like a better idea.</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Update:  \n\nSaving the iterations to disk, then loading the iterations for merging, reduces ram usage to 2-3 GB. However, merging track candidates as they are generated sounds like a better idea.\n\nThanks",
      "votes": null
    },
    {
      "id": "363845",
      "postDate": "07/30/2018 06:19:16",
      "content": "<p>I was wrong. I tried to use more pairs (several thousands) and hot Memory Error.</p>",
      "rawMarkdown": "I was wrong. I tried to use more pairs (several thousands) and hot Memory Error.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 362095,
      "author_name": "atom1231",
      "author_url": "",
      "post_date": "07/25/2018 17:12:15",
      "content": "<p>cut the iteration job , save the result to disk and load back when selecting candidate tracks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 362310,
      "author_name": "sergeyzlobin",
      "author_url": "",
      "post_date": "07/26/2018 06:24:36",
      "content": "<p>I don't have any memory issues. Probably a memory leak in your code, not in dbscan?</p>",
      "votes": null,
      "replies": [
        {
          "id": 363845,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "07/30/2018 06:19:16",
          "content": "<p>I was wrong. I tried to use more pairs (several thousands) and hot Memory Error.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 362317,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "07/26/2018 06:42:10",
      "content": "<p>I run in few GB.  You probably keep all the clusters from each DBSCAN.  Just do the math, and see why your memory will explode.  You have to merge on the fly if you run that many DBSCAN.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 362671,
      "author_name": "bkkaggle",
      "author_url": "",
      "post_date": "07/26/2018 21:24:44",
      "content": "<p>Update:  </p>\n\n<p>Saving the iterations to disk, then loading the iterations for merging, reduces ram usage to 2-3 GB. However, merging track candidates as they are generated sounds like a better idea.</p>\n\n<p>Thanks</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "362053": "How much ram should I need for helix unrolling and zshifting? I have 52gb and almost all of it is used up when I run ~20000 iterations of dbscan. Is there some way to optimize dbscan to use less ram?",
    "362095": "cut the iteration job , save the result to disk and load back when selecting candidate tracks",
    "362310": "I don't have any memory issues. Probably a memory leak in your code, not in dbscan?",
    "362317": "I run in few GB.  You probably keep all the clusters from each DBSCAN.  Just do the math, and see why your memory will explode.  You have to merge on the fly if you run that many DBSCAN.",
    "362671": "Update:  \n\nSaving the iterations to disk, then loading the iterations for merging, reduces ram usage to 2-3 GB. However, merging track candidates as they are generated sounds like a better idea.\n\nThanks",
    "363845": "I was wrong. I tried to use more pairs (several thousands) and hot Memory Error."
  },
  "source": "meta"
}