{
  "id": 57354,
  "title": "metric learning + clustering",
  "url": "/competitions/trackml-particle-identification/discussion/57354",
  "author_name": "",
  "post_date": "2018-05-22T23:40:46.952430500Z",
  "votes": 10,
  "comment_count": 11,
  "views": 0,
  "content": "<p>the idea is to learn transformation functions for func(hit), func(pairs) or func(triplets) etc .... where  distance of func(x1,) , func(x2) is small if and only if x1,x2 belong to the same track.</p>\n\n<p>We can then apply DBSCAN or other clustering  algorithms on func(x), instead of the original x</p>\n\n<p>Below shows  an example results using triple los.  Single hit is used as input.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/332286/9498/triple_loss.png\" alt=\"enter image description here\"></p>",
  "messages": [
    {
      "id": "332286",
      "postDate": "05/22/2018 23:40:46",
      "content": "<p>the idea is to learn transformation functions for func(hit), func(pairs) or func(triplets) etc .... where  distance of func(x1,) , func(x2) is small if and only if x1,x2 belong to the same track.</p>\n\n<p>We can then apply DBSCAN or other clustering  algorithms on func(x), instead of the original x</p>\n\n<p>Below shows  an example results using triple los.  Single hit is used as input.</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/332286/9498/triple_loss.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "the idea is to learn transformation functions for func(hit), func(pairs) or func(triplets) etc .... where  distance of func(x1,) , func(x2) is small if and only if x1,x2 belong to the same track.\n\nWe can then apply DBSCAN or other clustering  algorithms on func(x), instead of the original x\n\nBelow shows  an example results using triple los.  Single hit is used as input.\n\n   ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/332286/9498/triple_loss.png",
      "votes": null
    },
    {
      "id": "332632",
      "postDate": "05/23/2018 13:38:52",
      "content": "<p>Could you explain your idea more?\nI understand for pairs you can call function with vectors and then group similar vectors into clusters.\nBut how it is going to work with \"func(hit)\"? \n\"example results using triple los. Single hit is used as input.\" - does it mean function is \"func(triplets)\" and you color each hit separately? Please explain if possible :)</p>",
      "rawMarkdown": "Could you explain your idea more?\nI understand for pairs you can call function with vectors and then group similar vectors into clusters.\nBut how it is going to work with \"func(hit)\"? \n\"example results using triple los. Single hit is used as input.\" - does it mean function is \"func(triplets)\" and you color each hit separately? Please explain if possible :)",
      "votes": null
    },
    {
      "id": "332640",
      "postDate": "05/23/2018 13:50:57",
      "content": "<p>current dbscan transform (x,y,z) to (x/d,y/d,z/r) where d=(x*<em>2+y</em>*2+z*<em>2)</em>*0.5 and r=(x*<em>2+y</em>*2)**0.5.</p>\n\n<p>assume that there is a better transform function and we want to find it by metric learning. let the transform function be func (x,y,z) = func (hit).</p>\n\n<p>to construct triple loss, say:</p>\n\n<ol>\n<li><p>select a hit where its particle_id !=0. call this x.</p></li>\n<li><p>find another hit with the same particle_id as  x above, call this x+.</p></li>\n<li><p>find another hit with the different particle_id as  x above, call this x-.</p></li>\n</ol>\n\n<p>then you can formulate triple loss with (x,x+,x-)</p>\n\n<hr>\n\n<p>Now this results may not be good enough. assume given a hit x, we use kdtree to find some neighbors neigbour, n0 n1 ...nk.  Instead of 3 dim (x,y,z) we add another feature called the \"forward motion\", vi = ni-x. now a hit point is described by (x,y,z,v0), (x,y,z,v1), ...(x,y,z,vk). We treat each of these as separate data point. we now want to cluster this. note that v0 has 3 dim and func (x,y,z,v)  can be respresented by  func (hit,neighbour), aka func (hit pair)</p>",
      "rawMarkdown": "current dbscan transform (x,y,z) to (x/d,y/d,z/r) where d=(x**2+y**2+z**2)**0.5 and r=(x**2+y**2)**0.5.\n\nassume that there is a better transform function and we want to find it by metric learning. let the transform function be func (x,y,z) = func (hit).\n\n\nto construct triple loss, say:\n\n 1. select a hit where its particle_id !=0. call this x.\n \n 2. find another hit with the same particle_id as  x above, call this x+.\n\n 3. find another hit with the different particle_id as  x above, call this x-.\n\nthen you can formulate triple loss with (x,x+,x-)\n\n---\n\nNow this results may not be good enough. assume given a hit x, we use kdtree to find some neighbors neigbour, n0 n1 ...nk.  Instead of 3 dim (x,y,z) we add another feature called the \"forward motion\", vi = ni-x. now a hit point is described by (x,y,z,v0), (x,y,z,v1), ...(x,y,z,vk). We treat each of these as separate data point. we now want to cluster this. note that v0 has 3 dim and func (x,y,z,v)  can be respresented by  func (hit,neighbour), aka func (hit pair)",
      "votes": null
    },
    {
      "id": "332842",
      "postDate": "05/23/2018 21:51:22",
      "content": "<p>Any success with this so far? I have been trying a similar method to find a transform, but have mostly succeeded in running out of RAM and freezing my machine ;)</p>",
      "rawMarkdown": "Any success with this so far? I have been trying a similar method to find a transform, but have mostly succeeded in running out of RAM and freezing my machine ;)",
      "votes": null
    },
    {
      "id": "333895",
      "postDate": "05/26/2018 01:54:33",
      "content": "<p>Would learning the transformation be similar to how a CNN learns kernel/filter matrices?</p>",
      "rawMarkdown": "Would learning the transformation be similar to how a CNN learns kernel/filter matrices?",
      "votes": null
    },
    {
      "id": "333972",
      "postDate": "05/26/2018 07:01:09",
      "content": "<p>They're of the same spirit, but that's about it. Direct kernel learning is transductive (cannot be used on new 'unobserved' objects). Also, DKL guarantees constraint satisfaction. However, learning a generating function via metric learning is inductive and does not guarantee constraint satisfaction.</p>",
      "rawMarkdown": "They're of the same spirit, but that's about it. Direct kernel learning is transductive (cannot be used on new 'unobserved' objects). Also, DKL guarantees constraint satisfaction. However, learning a generating function via metric learning is inductive and does not guarantee constraint satisfaction.",
      "votes": null
    },
    {
      "id": "334034",
      "postDate": "05/26/2018 11:03:11",
      "content": "<p>Some progress, but too early to call it success.</p>\n\n<p>Classification rates for pairs of hits with simple Random Forest looks initially promising, at 75-90%.  I also found a <a href=\"https://arxiv.org/pdf/1201.0610.pdf\">Random Forest Distance paper</a> with an extended approach that gets excellent results, lending some credibility to this approach. </p>\n\n<p>I needed a solution to avoid running out of memory, so I wrote a quick-and-dirty function for sparse pairwise distances in that can be used with sklearn DBSCAN with metric='precomputed'. See <code>pairwise_distance_sparse_fast</code> in <a href=\"https://github.com/jonnor/datascience-master/blob/master/trackml/Particles.ipynb\">my notebook</a>.</p>\n\n<p>However runtimes for calculating distances are too long to be productive, estimated 1 hour per event with my metric. Some approaches for fixing this:</p>\n\n<ol>\n<li>Use a more optimized Random Forest implementation in C++, based on <a href=\"https://github.com/jonnor/emtrees\">emtrees</a>.</li>\n<li>Try to partition the hits to reduce number of pairwise combinations. It seems generally unlikely that hits on opposite sites of sensor part of the same track for instance. Maybe some basic partitioning rules can be learned?</li>\n<li>Implement a radius search and use that to precompute the distances passed to sklearn DBSCAN.</li>\n<li>Use a custom DBSCAN algorithm implementation with the custom metric, using radius search to avoid computing so many distances.</li>\n</ol>",
      "rawMarkdown": "Some progress, but too early to call it success.\n\nClassification rates for pairs of hits with simple Random Forest looks initially promising, at 75-90%.  I also found a [Random Forest Distance paper][1] with an extended approach that gets excellent results, lending some credibility to this approach. \n\nI needed a solution to avoid running out of memory, so I wrote a quick-and-dirty function for sparse pairwise distances in that can be used with sklearn DBSCAN with metric='precomputed'. See `pairwise_distance_sparse_fast` in [my notebook][2].\n\nHowever runtimes for calculating distances are too long to be productive, estimated 1 hour per event with my metric. Some approaches for fixing this:\n\n1. Use a more optimized Random Forest implementation in C++, based on [emtrees][3].\n2. Try to partition the hits to reduce number of pairwise combinations. It seems generally unlikely that hits on opposite sites of sensor part of the same track for instance. Maybe some basic partitioning rules can be learned?\n3. Implement a radius search and use that to precompute the distances passed to sklearn DBSCAN.\n4. Use a custom DBSCAN algorithm implementation with the custom metric, using radius search to avoid computing so many distances.\n\n  [1]: https://arxiv.org/pdf/1201.0610.pdf\n  [2]: https://github.com/jonnor/datascience-master/blob/master/trackml/Particles.ipynb\n  [3]: https://github.com/jonnor/emtrees",
      "votes": null
    },
    {
      "id": "334246",
      "postDate": "05/26/2018 19:00:55",
      "content": "<p>Would this still be the case when the labels are taken into account and the problem is reframed as a semi-supervised learning problem where first you use the labels to somehow train a custom metric which could then be used with DBSCAN on some un labeled data or the test set?</p>",
      "rawMarkdown": "Would this still be the case when the labels are taken into account and the problem is reframed as a semi-supervised learning problem where first you use the labels to somehow train a custom metric which could then be used with DBSCAN on some un labeled data or the test set?",
      "votes": null
    },
    {
      "id": "335164",
      "postDate": "05/29/2018 10:20:52",
      "content": "<p>...  sklearn DBSCAN with metric='precomputed' ...</p>\n\n<p>There are two ways to use DBSCAN+machine learning</p>\n\n<ol>\n<li><p>input data ---&gt; DBSCAN with learned distance function</p></li>\n<li><p>input data ---&gt; transformed data ---&gt; DBSCAN with usual L2 distance function</p></li>\n</ol>\n\n<p>I think method 2 is easier to implement. All \"non linear distance function\" can be re-implemented as \"non linear feature + linear distance\". </p>\n\n<p>Referring to  <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/57643\">https://www.kaggle.com/c/trackml-particle-identification/discussion/57643</a>, you many want to split all hits into \"smaller cones\" and perform \"modified DBSCAN\" separately on it.</p>",
      "rawMarkdown": "...  sklearn DBSCAN with metric='precomputed' ...\n\nThere are two ways to use DBSCAN+machine learning\n\n1.  input data ---&gt; DBSCAN with learned distance function\n\n2. input data ---&gt; transformed data ---&gt; DBSCAN with usual L2 distance function\n\nI think method 2 is easier to implement. All \"non linear distance function\" can be re-implemented as \"non linear feature + linear distance\". \n\n\nReferring to  https://www.kaggle.com/c/trackml-particle-identification/discussion/57643, you many want to split all hits into \"smaller cones\" and perform \"modified DBSCAN\" separately on it.",
      "votes": null
    },
    {
      "id": "335320",
      "postDate": "05/29/2018 15:22:15",
      "content": "<p>You mention selecting a hit where particle_id!=0; is it a bad idea to just filter all of those hits out as noise from the get-go? Removes detector noise, and reduces your working dataset by several thousands right out of the gate if you do that. </p>",
      "rawMarkdown": "You mention selecting a hit where particle_id!=0; is it a bad idea to just filter all of those hits out as noise from the get-go? Removes detector noise, and reduces your working dataset by several thousands right out of the gate if you do that.",
      "votes": null
    },
    {
      "id": "335376",
      "postDate": "05/29/2018 16:54:44",
      "content": "<p>For exploratory analysis, you could... But the test set doesn't have particle IDs, so you won't be able to do this for your submission. </p>\n\n<p>If you train a model with the cleaned data, the model will probably perform worse on the test set where you won't be able to remove the noise hits. </p>",
      "rawMarkdown": "For exploratory analysis, you could... But the test set doesn't have particle IDs, so you won't be able to do this for your submission. \n\nIf you train a model with the cleaned data, the model will probably perform worse on the test set where you won't be able to remove the noise hits.",
      "votes": null
    },
    {
      "id": "376746",
      "postDate": "08/28/2018 02:34:28",
      "content": "<p>Hi, I am trying metric-learning based clustering on high-dimensional data. Could u pls tell me what metric learning method u use? MMC, ITML, or  lmnn? In my practice, it seemed only MMC is suited for semi-supervised clustering. However, it only work when dimension of data is less than 10. <br>\nAny idea about the metric-learning algorithm? or any suggestion about the clustering algorithm for high-dimension data(several hundred)?\nThanks!</p>",
      "rawMarkdown": "Hi, I am trying metric-learning based clustering on high-dimensional data. Could u pls tell me what metric learning method u use? MMC, ITML, or  lmnn? In my practice, it seemed only MMC is suited for semi-supervised clustering. However, it only work when dimension of data is less than 10.  \nAny idea about the metric-learning algorithm? or any suggestion about the clustering algorithm for high-dimension data(several hundred)?\nThanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 332632,
      "author_name": "jacekpoplawski",
      "author_url": "",
      "post_date": "05/23/2018 13:38:52",
      "content": "<p>Could you explain your idea more?\nI understand for pairs you can call function with vectors and then group similar vectors into clusters.\nBut how it is going to work with \"func(hit)\"? \n\"example results using triple los. Single hit is used as input.\" - does it mean function is \"func(triplets)\" and you color each hit separately? Please explain if possible :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 332640,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "05/23/2018 13:50:57",
          "content": "<p>current dbscan transform (x,y,z) to (x/d,y/d,z/r) where d=(x*<em>2+y</em>*2+z*<em>2)</em>*0.5 and r=(x*<em>2+y</em>*2)**0.5.</p>\n\n<p>assume that there is a better transform function and we want to find it by metric learning. let the transform function be func (x,y,z) = func (hit).</p>\n\n<p>to construct triple loss, say:</p>\n\n<ol>\n<li><p>select a hit where its particle_id !=0. call this x.</p></li>\n<li><p>find another hit with the same particle_id as  x above, call this x+.</p></li>\n<li><p>find another hit with the different particle_id as  x above, call this x-.</p></li>\n</ol>\n\n<p>then you can formulate triple loss with (x,x+,x-)</p>\n\n<hr>\n\n<p>Now this results may not be good enough. assume given a hit x, we use kdtree to find some neighbors neigbour, n0 n1 ...nk.  Instead of 3 dim (x,y,z) we add another feature called the \"forward motion\", vi = ni-x. now a hit point is described by (x,y,z,v0), (x,y,z,v1), ...(x,y,z,vk). We treat each of these as separate data point. we now want to cluster this. note that v0 has 3 dim and func (x,y,z,v)  can be respresented by  func (hit,neighbour), aka func (hit pair)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 335320,
          "author_name": "ssoder",
          "author_url": "",
          "post_date": "05/29/2018 15:22:15",
          "content": "<p>You mention selecting a hit where particle_id!=0; is it a bad idea to just filter all of those hits out as noise from the get-go? Removes detector noise, and reduces your working dataset by several thousands right out of the gate if you do that. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 335376,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "05/29/2018 16:54:44",
          "content": "<p>For exploratory analysis, you could... But the test set doesn't have particle IDs, so you won't be able to do this for your submission. </p>\n\n<p>If you train a model with the cleaned data, the model will probably perform worse on the test set where you won't be able to remove the noise hits. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 332842,
      "author_name": "cosmoton",
      "author_url": "",
      "post_date": "05/23/2018 21:51:22",
      "content": "<p>Any success with this so far? I have been trying a similar method to find a transform, but have mostly succeeded in running out of RAM and freezing my machine ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 334034,
          "author_name": "jonnor",
          "author_url": "",
          "post_date": "05/26/2018 11:03:11",
          "content": "<p>Some progress, but too early to call it success.</p>\n\n<p>Classification rates for pairs of hits with simple Random Forest looks initially promising, at 75-90%.  I also found a <a href=\"https://arxiv.org/pdf/1201.0610.pdf\">Random Forest Distance paper</a> with an extended approach that gets excellent results, lending some credibility to this approach. </p>\n\n<p>I needed a solution to avoid running out of memory, so I wrote a quick-and-dirty function for sparse pairwise distances in that can be used with sklearn DBSCAN with metric='precomputed'. See <code>pairwise_distance_sparse_fast</code> in <a href=\"https://github.com/jonnor/datascience-master/blob/master/trackml/Particles.ipynb\">my notebook</a>.</p>\n\n<p>However runtimes for calculating distances are too long to be productive, estimated 1 hour per event with my metric. Some approaches for fixing this:</p>\n\n<ol>\n<li>Use a more optimized Random Forest implementation in C++, based on <a href=\"https://github.com/jonnor/emtrees\">emtrees</a>.</li>\n<li>Try to partition the hits to reduce number of pairwise combinations. It seems generally unlikely that hits on opposite sites of sensor part of the same track for instance. Maybe some basic partitioning rules can be learned?</li>\n<li>Implement a radius search and use that to precompute the distances passed to sklearn DBSCAN.</li>\n<li>Use a custom DBSCAN algorithm implementation with the custom metric, using radius search to avoid computing so many distances.</li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 335164,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "05/29/2018 10:20:52",
          "content": "<p>...  sklearn DBSCAN with metric='precomputed' ...</p>\n\n<p>There are two ways to use DBSCAN+machine learning</p>\n\n<ol>\n<li><p>input data ---&gt; DBSCAN with learned distance function</p></li>\n<li><p>input data ---&gt; transformed data ---&gt; DBSCAN with usual L2 distance function</p></li>\n</ol>\n\n<p>I think method 2 is easier to implement. All \"non linear distance function\" can be re-implemented as \"non linear feature + linear distance\". </p>\n\n<p>Referring to  <a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/57643\">https://www.kaggle.com/c/trackml-particle-identification/discussion/57643</a>, you many want to split all hits into \"smaller cones\" and perform \"modified DBSCAN\" separately on it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 333895,
      "author_name": "jackvial",
      "author_url": "",
      "post_date": "05/26/2018 01:54:33",
      "content": "<p>Would learning the transformation be similar to how a CNN learns kernel/filter matrices?</p>",
      "votes": null,
      "replies": [
        {
          "id": 333972,
          "author_name": "puremath86",
          "author_url": "",
          "post_date": "05/26/2018 07:01:09",
          "content": "<p>They're of the same spirit, but that's about it. Direct kernel learning is transductive (cannot be used on new 'unobserved' objects). Also, DKL guarantees constraint satisfaction. However, learning a generating function via metric learning is inductive and does not guarantee constraint satisfaction.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 334246,
          "author_name": "jackvial",
          "author_url": "",
          "post_date": "05/26/2018 19:00:55",
          "content": "<p>Would this still be the case when the labels are taken into account and the problem is reframed as a semi-supervised learning problem where first you use the labels to somehow train a custom metric which could then be used with DBSCAN on some un labeled data or the test set?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 376746,
      "author_name": "davidwang1113",
      "author_url": "",
      "post_date": "08/28/2018 02:34:28",
      "content": "<p>Hi, I am trying metric-learning based clustering on high-dimensional data. Could u pls tell me what metric learning method u use? MMC, ITML, or  lmnn? In my practice, it seemed only MMC is suited for semi-supervised clustering. However, it only work when dimension of data is less than 10. <br>\nAny idea about the metric-learning algorithm? or any suggestion about the clustering algorithm for high-dimension data(several hundred)?\nThanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "332286": "the idea is to learn transformation functions for func(hit), func(pairs) or func(triplets) etc .... where  distance of func(x1,) , func(x2) is small if and only if x1,x2 belong to the same track.\n\nWe can then apply DBSCAN or other clustering  algorithms on func(x), instead of the original x\n\nBelow shows  an example results using triple los.  Single hit is used as input.\n\n   ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/332286/9498/triple_loss.png",
    "332632": "Could you explain your idea more?\nI understand for pairs you can call function with vectors and then group similar vectors into clusters.\nBut how it is going to work with \"func(hit)\"? \n\"example results using triple los. Single hit is used as input.\" - does it mean function is \"func(triplets)\" and you color each hit separately? Please explain if possible :)",
    "332640": "current dbscan transform (x,y,z) to (x/d,y/d,z/r) where d=(x**2+y**2+z**2)**0.5 and r=(x**2+y**2)**0.5.\n\nassume that there is a better transform function and we want to find it by metric learning. let the transform function be func (x,y,z) = func (hit).\n\n\nto construct triple loss, say:\n\n 1. select a hit where its particle_id !=0. call this x.\n \n 2. find another hit with the same particle_id as  x above, call this x+.\n\n 3. find another hit with the different particle_id as  x above, call this x-.\n\nthen you can formulate triple loss with (x,x+,x-)\n\n---\n\nNow this results may not be good enough. assume given a hit x, we use kdtree to find some neighbors neigbour, n0 n1 ...nk.  Instead of 3 dim (x,y,z) we add another feature called the \"forward motion\", vi = ni-x. now a hit point is described by (x,y,z,v0), (x,y,z,v1), ...(x,y,z,vk). We treat each of these as separate data point. we now want to cluster this. note that v0 has 3 dim and func (x,y,z,v)  can be respresented by  func (hit,neighbour), aka func (hit pair)",
    "332842": "Any success with this so far? I have been trying a similar method to find a transform, but have mostly succeeded in running out of RAM and freezing my machine ;)",
    "333895": "Would learning the transformation be similar to how a CNN learns kernel/filter matrices?",
    "333972": "They're of the same spirit, but that's about it. Direct kernel learning is transductive (cannot be used on new 'unobserved' objects). Also, DKL guarantees constraint satisfaction. However, learning a generating function via metric learning is inductive and does not guarantee constraint satisfaction.",
    "334034": "Some progress, but too early to call it success.\n\nClassification rates for pairs of hits with simple Random Forest looks initially promising, at 75-90%.  I also found a [Random Forest Distance paper][1] with an extended approach that gets excellent results, lending some credibility to this approach. \n\nI needed a solution to avoid running out of memory, so I wrote a quick-and-dirty function for sparse pairwise distances in that can be used with sklearn DBSCAN with metric='precomputed'. See `pairwise_distance_sparse_fast` in [my notebook][2].\n\nHowever runtimes for calculating distances are too long to be productive, estimated 1 hour per event with my metric. Some approaches for fixing this:\n\n1. Use a more optimized Random Forest implementation in C++, based on [emtrees][3].\n2. Try to partition the hits to reduce number of pairwise combinations. It seems generally unlikely that hits on opposite sites of sensor part of the same track for instance. Maybe some basic partitioning rules can be learned?\n3. Implement a radius search and use that to precompute the distances passed to sklearn DBSCAN.\n4. Use a custom DBSCAN algorithm implementation with the custom metric, using radius search to avoid computing so many distances.\n\n  [1]: https://arxiv.org/pdf/1201.0610.pdf\n  [2]: https://github.com/jonnor/datascience-master/blob/master/trackml/Particles.ipynb\n  [3]: https://github.com/jonnor/emtrees",
    "334246": "Would this still be the case when the labels are taken into account and the problem is reframed as a semi-supervised learning problem where first you use the labels to somehow train a custom metric which could then be used with DBSCAN on some un labeled data or the test set?",
    "335164": "...  sklearn DBSCAN with metric='precomputed' ...\n\nThere are two ways to use DBSCAN+machine learning\n\n1.  input data ---&gt; DBSCAN with learned distance function\n\n2. input data ---&gt; transformed data ---&gt; DBSCAN with usual L2 distance function\n\nI think method 2 is easier to implement. All \"non linear distance function\" can be re-implemented as \"non linear feature + linear distance\". \n\n\nReferring to  https://www.kaggle.com/c/trackml-particle-identification/discussion/57643, you many want to split all hits into \"smaller cones\" and perform \"modified DBSCAN\" separately on it.",
    "335320": "You mention selecting a hit where particle_id!=0; is it a bad idea to just filter all of those hits out as noise from the get-go? Removes detector noise, and reduces your working dataset by several thousands right out of the gate if you do that.",
    "335376": "For exploratory analysis, you could... But the test set doesn't have particle IDs, so you won't be able to do this for your submission. \n\nIf you train a model with the cleaned data, the model will probably perform worse on the test set where you won't be able to remove the noise hits.",
    "376746": "Hi, I am trying metric-learning based clustering on high-dimensional data. Could u pls tell me what metric learning method u use? MMC, ITML, or  lmnn? In my practice, it seemed only MMC is suited for semi-supervised clustering. However, it only work when dimension of data is less than 10.  \nAny idea about the metric-learning algorithm? or any suggestion about the clustering algorithm for high-dimension data(several hundred)?\nThanks!"
  },
  "source": "meta"
}