{
  "id": 58261,
  "title": "pytorch stater kit for deep metric learning",
  "url": "/competitions/trackml-particle-identification/discussion/58261",
  "author_name": "",
  "post_date": "2018-06-05T07:55:25.745657700Z",
  "votes": 10,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Note:</p>\n\n<ul>\n<li><p>this is not completed training code. It is meant to show the feasibility of using deep metric learning + dbscan.</p></li>\n<li><p>no submission made to LB yet. </p></li>\n<li><p>the results shows fitting a train sample (not on test sample). If we can fit a train sample, then we can generalize it on more train data.</p></li>\n<li><p>the attached code perform the following tasks:</p>\n\n<ul><li><p>input: xyz + transformation (e.g. r/d, sin(a), cos(a), ...)</p></li>\n<li><p>learned a embedded features using triplet loss</p></li>\n<li><p>build kdtree on embedded features and retrieve neighbors of a query hit </p></li></ul>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/338504/9570/Slide1.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/338504/9569/Slide2.png\" alt=\"enter image description here\"></p></li>\n</ul>",
  "messages": [
    {
      "id": "338504",
      "postDate": "06/05/2018 07:55:25",
      "content": "<p>Note:</p>\n\n<ul>\n<li><p>this is not completed training code. It is meant to show the feasibility of using deep metric learning + dbscan.</p></li>\n<li><p>no submission made to LB yet. </p></li>\n<li><p>the results shows fitting a train sample (not on test sample). If we can fit a train sample, then we can generalize it on more train data.</p></li>\n<li><p>the attached code perform the following tasks:</p>\n\n<ul><li><p>input: xyz + transformation (e.g. r/d, sin(a), cos(a), ...)</p></li>\n<li><p>learned a embedded features using triplet loss</p></li>\n<li><p>build kdtree on embedded features and retrieve neighbors of a query hit </p></li></ul>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/338504/9570/Slide1.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/338504/9569/Slide2.png\" alt=\"enter image description here\"></p></li>\n</ul>",
      "rawMarkdown": "Note:\n\n- this is not completed training code. It is meant to show the feasibility of using deep metric learning + dbscan.\n\n- no submission made to LB yet. \n\n- the results shows fitting a train sample (not on test sample). If we can fit a train sample, then we can generalize it on more train data.\n\n- the attached code perform the following tasks:\n\n     - input: xyz + transformation (e.g. r/d, sin(a), cos(a), ...)\n\n     - learned a embedded features using triplet loss\n\n     - build kdtree on embedded features and retrieve neighbors of a query hit \n\n   \n\n\n   ![enter image description here][1]\n\n   ![enter image description here][2]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/338504/9570/Slide1.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/338504/9569/Slide2.png",
      "votes": null
    },
    {
      "id": "340826",
      "postDate": "06/10/2018 12:24:39",
      "content": "<p>results for running DBSCAN on embedded features on train data:</p>\n\n<p>for half of the volume only (z&gt;0), trained on particles with momentum greater than 1:</p>\n\n<ul>\n<li><p>dbscan, eps=1: [ 0  000001029] score : 0.26149985   </p></li>\n<li><p>dbscan, eps=1.05: [ 0  000001029] score : 0.26531884   </p></li>\n</ul>\n\n<hr>\n\n<pre><code>    net.set_mode('test')\n    with torch.no_grad():\n        inputs = torch.from_numpy(hits).cuda()\n        features = net.forward( inputs )\n\n    #show results---\n    xyz    = df[['x','y','z']].values\n    features = features.data.cpu().numpy()\n\n    _,labels = dbscan(features,\n       eps=1,\n       min_samples=1,\n       algorithm='auto',\n       n_jobs=-1)\n\n    data_dir = '/root/share/project/kaggle/cern/data/__download__/train_100_events'\n    event_id= '000001029'\n    particles = pd.read_csv(data_dir + '/event%s-particles.csv'%event_id)\n    hits  = pd.read_csv(data_dir + '/event%s-hits.csv'%event_id)\n    cells = pd.read_csv(data_dir + '/event%s-cells.csv'%event_id)\n    truth = pd.read_csv(data_dir + '/event%s-truth.csv'%event_id)\n\n    track_id = np.zeros(len(hits),np.int32)\n    track_id[np.where(hits.z.values&amp;gt;0)]  = labels\n    submission = pd.DataFrame(columns=['event_id', 'hit_id', 'track_id'],\n        data=np.column_stack(([int(event_id),]*len(hits), hits.hit_id.values, track_id))\n    ).astype(int)\n\n    score = score_event(truth, submission)\n    print('[%2d  %s] score : %0.8f   %0.1f min'%(0, event_id, score, (0)/60))\n</code></pre>",
      "rawMarkdown": "results for running DBSCAN on embedded features on train data:\n\nfor half of the volume only (z&gt;0), trained on particles with momentum greater than 1:\n\n   - dbscan, eps=1: [ 0  000001029] score : 0.26149985   \n\n   - dbscan, eps=1.05: [ 0  000001029] score : 0.26531884   \n\n\n---\n\n\n        net.set_mode('test')\n        with torch.no_grad():\n            inputs = torch.from_numpy(hits).cuda()\n            features = net.forward( inputs )\n\n        #show results---\n        xyz    = df[['x','y','z']].values\n        features = features.data.cpu().numpy()\n \n        _,labels = dbscan(features,\n           eps=1,\n           min_samples=1,\n           algorithm='auto',\n           n_jobs=-1)\n\n        data_dir = '/root/share/project/kaggle/cern/data/__download__/train_100_events'\n        event_id= '000001029'\n        particles = pd.read_csv(data_dir + '/event%s-particles.csv'%event_id)\n        hits  = pd.read_csv(data_dir + '/event%s-hits.csv'%event_id)\n        cells = pd.read_csv(data_dir + '/event%s-cells.csv'%event_id)\n        truth = pd.read_csv(data_dir + '/event%s-truth.csv'%event_id)\n\n        track_id = np.zeros(len(hits),np.int32)\n        track_id[np.where(hits.z.values&gt;0)]  = labels\n        submission = pd.DataFrame(columns=['event_id', 'hit_id', 'track_id'],\n            data=np.column_stack(([int(event_id),]*len(hits), hits.hit_id.values, track_id))\n        ).astype(int)\n\n        score = score_event(truth, submission)\n        print('[%2d  %s] score : %0.8f   %0.1f min'%(0, event_id, score, (0)/60))",
      "votes": null
    },
    {
      "id": "341556",
      "postDate": "06/11/2018 19:39:53",
      "content": "<p>How long does it approximately take to train to get to 0.26+ ?\nYou use 1 as the separation parameter in the triplet loss I take it?</p>",
      "rawMarkdown": "How long does it approximately take to train to get to 0.26+ ?\nYou use 1 as the separation parameter in the triplet loss I take it?",
      "votes": null
    },
    {
      "id": "341709",
      "postDate": "06/12/2018 05:38:21",
      "content": "<p>This is interesting, it does not seem as good (yet) as dbscan on unrolled helix, but you run dbscan only once, not hundreds times.  The only caveat is you use a single event to validate, hard to compare with other approaches that way.</p>",
      "rawMarkdown": "This is interesting, it does not seem as good (yet) as dbscan on unrolled helix, but you run dbscan only once, not hundreds times.  The only caveat is you use a single event to validate, hard to compare with other approaches that way.",
      "votes": null
    },
    {
      "id": "341764",
      "postDate": "06/12/2018 08:18:01",
      "content": "<p>\"the only caveat is you use a single event to validate\" ...</p>\n\n<p>this experiment shows the approximate upper bound. I estimate the performance using huge train data, model tuning etc would be  about  -10% to 10% better or worse.</p>\n\n<p>It is not known if we train on other particles (e.g. momentum of certain range) and combine the results, would it be better or not.</p>\n\n<p>One thing to note is that track extension do not improve results here:</p>\n\n<p>metric_learning+dbscan+extend_track improves only from 0.005 to 0.010 (either the metric learning tracks are already well extended or they are broken and skip some hits in the middle) </p>",
      "rawMarkdown": "\"the only caveat is you use a single event to validate\" ...\n\nthis experiment shows the approximate upper bound. I estimate the performance using huge train data, model tuning etc would be  about  -10% to 10% better or worse.\n\nIt is not known if we train on other particles (e.g. momentum of certain range) and combine the results, would it be better or not.\n\nOne thing to note is that track extension do not improve results here:\n\nmetric_learning+dbscan+extend_track improves only from 0.005 to 0.010 (either the metric learning tracks are already well extended or they are broken and skip some hits in the middle)",
      "votes": null
    },
    {
      "id": "341798",
      "postDate": "06/12/2018 09:39:35",
      "content": "<p>I think for siamese networks with triplet loss, how you sample your positive and negative examples are as important as hyperparameter tuning etc.  Need to sample negative examples that are hard(not just random) to get good convergence.</p>",
      "rawMarkdown": "I think for siamese networks with triplet loss, how you sample your positive and negative examples are as important as hyperparameter tuning etc.  Need to sample negative examples that are hard(not just random) to get good convergence.",
      "votes": null
    },
    {
      "id": "341854",
      "postDate": "06/12/2018 11:50:14",
      "content": "<p>sampling quite straight forward in this case. just sample in KNN neighbours</p>",
      "rawMarkdown": "sampling quite straight forward in this case. just sample in KNN neighbours",
      "votes": null
    },
    {
      "id": "346897",
      "postDate": "06/22/2018 17:16:42",
      "content": "<p>for those who are using pytorch, here is pytorch compatible KNN and dbscan on gpu:</p>\n\n<p><a href=\"https://github.com/facebookresearch/faiss\">https://github.com/facebookresearch/faiss</a></p>\n\n<p><a href=\"https://github.com/karthikv2k/gpu_dbscan/blob/master/gpu_dbscan/gpu_dbscan.py\">https://github.com/karthikv2k/gpu_dbscan/blob/master/gpu_dbscan/gpu_dbscan.py</a></p>\n\n<p>\"Faiss is a library for efficient similarity search and clustering of dense vectors. It contains algorithms that search in sets of vectors of any size, up to ones that possibly do not fit in RAM. It also contains supporting code for evaluation and parameter tuning. Faiss is written in C++ with complete wrappers for Python/numpy. Some of the most useful algorithms are implemented on the GPU. It is developed by Facebook AI Research.\"</p>",
      "rawMarkdown": "for those who are using pytorch, here is pytorch compatible KNN and dbscan on gpu:\n\nhttps://github.com/facebookresearch/faiss\n\nhttps://github.com/karthikv2k/gpu_dbscan/blob/master/gpu_dbscan/gpu_dbscan.py\n\n\"Faiss is a library for efficient similarity search and clustering of dense vectors. It contains algorithms that search in sets of vectors of any size, up to ones that possibly do not fit in RAM. It also contains supporting code for evaluation and parameter tuning. Faiss is written in C++ with complete wrappers for Python/numpy. Some of the most useful algorithms are implemented on the GPU. It is developed by Facebook AI Research.\"",
      "votes": null
    },
    {
      "id": "346898",
      "postDate": "06/22/2018 17:22:08",
      "content": "<p>Excellent - have you tried it?  Does it provide effective parallelization for the clustering you are doing?</p>",
      "rawMarkdown": "Excellent - have you tried it?  Does it provide effective parallelization for the clustering you are doing?",
      "votes": null
    },
    {
      "id": "346912",
      "postDate": "06/22/2018 18:09:52",
      "content": "<p>i haven't tried. The dbscan code doesn't seems to be fast. </p>",
      "rawMarkdown": "i haven't tried. The dbscan code doesn't seems to be fast.",
      "votes": null
    },
    {
      "id": "347034",
      "postDate": "06/23/2018 03:45:08",
      "content": "<p>Thanks for sharing, this looks interesting.  I don't think it depend on pytorch at all actually. The only link is that it is in the pytorch channel when installing via conda.</p>",
      "rawMarkdown": "Thanks for sharing, this looks interesting.  I don't think it depend on pytorch at all actually. The only link is that it is in the pytorch channel when installing via conda.",
      "votes": null
    },
    {
      "id": "347039",
      "postDate": "06/23/2018 03:58:20",
      "content": "<p>Interesting stuff .Thanks for sharing Heng</p>",
      "rawMarkdown": "Interesting stuff .Thanks for sharing Heng",
      "votes": null
    },
    {
      "id": "347046",
      "postDate": "06/23/2018 04:12:45",
      "content": "<p>there is one more here:</p>\n\n<p><a href=\"https://github.com/ghoulsblade/CudaDBClustering/tree/master/src\">https://github.com/ghoulsblade/CudaDBClustering/tree/master/src</a></p>",
      "rawMarkdown": "there is one more here:\n\nhttps://github.com/ghoulsblade/CudaDBClustering/tree/master/src",
      "votes": null
    },
    {
      "id": "347052",
      "postDate": "06/23/2018 04:32:46",
      "content": "<p>check these as well</p>\n\n<p><a href=\"https://github.com/ZihengChen/ImageAlogrithm3D\">https://github.com/ZihengChen/ImageAlogrithm3D</a></p>\n\n<p><a href=\"https://galleryziheng.wordpress.com/2017/12/08/gpu-acceleration-of-imaging-algorithm/\">https://galleryziheng.wordpress.com/2017/12/08/gpu-acceleration-of-imaging-algorithm/</a></p>\n\n<p><a href=\"https://galleryziheng.wordpress.com/2017/08/27/get-python-clustering-work/\">https://galleryziheng.wordpress.com/2017/08/27/get-python-clustering-work/</a></p>",
      "rawMarkdown": "check these as well\n\nhttps://github.com/ZihengChen/ImageAlogrithm3D\n\nhttps://galleryziheng.wordpress.com/2017/12/08/gpu-acceleration-of-imaging-algorithm/\n\nhttps://galleryziheng.wordpress.com/2017/08/27/get-python-clustering-work/",
      "votes": null
    },
    {
      "id": "347912",
      "postDate": "06/25/2018 16:54:16",
      "content": "<p>@Heng, can you please share (if possible), the feature extraction pytorch code you have used for this?</p>",
      "rawMarkdown": "Heng, can you please share (if possible), the feature extraction pytorch code you have used for this?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 340826,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/10/2018 12:24:39",
      "content": "<p>results for running DBSCAN on embedded features on train data:</p>\n\n<p>for half of the volume only (z&gt;0), trained on particles with momentum greater than 1:</p>\n\n<ul>\n<li><p>dbscan, eps=1: [ 0  000001029] score : 0.26149985   </p></li>\n<li><p>dbscan, eps=1.05: [ 0  000001029] score : 0.26531884   </p></li>\n</ul>\n\n<hr>\n\n<pre><code>    net.set_mode('test')\n    with torch.no_grad():\n        inputs = torch.from_numpy(hits).cuda()\n        features = net.forward( inputs )\n\n    #show results---\n    xyz    = df[['x','y','z']].values\n    features = features.data.cpu().numpy()\n\n    _,labels = dbscan(features,\n       eps=1,\n       min_samples=1,\n       algorithm='auto',\n       n_jobs=-1)\n\n    data_dir = '/root/share/project/kaggle/cern/data/__download__/train_100_events'\n    event_id= '000001029'\n    particles = pd.read_csv(data_dir + '/event%s-particles.csv'%event_id)\n    hits  = pd.read_csv(data_dir + '/event%s-hits.csv'%event_id)\n    cells = pd.read_csv(data_dir + '/event%s-cells.csv'%event_id)\n    truth = pd.read_csv(data_dir + '/event%s-truth.csv'%event_id)\n\n    track_id = np.zeros(len(hits),np.int32)\n    track_id[np.where(hits.z.values&amp;gt;0)]  = labels\n    submission = pd.DataFrame(columns=['event_id', 'hit_id', 'track_id'],\n        data=np.column_stack(([int(event_id),]*len(hits), hits.hit_id.values, track_id))\n    ).astype(int)\n\n    score = score_event(truth, submission)\n    print('[%2d  %s] score : %0.8f   %0.1f min'%(0, event_id, score, (0)/60))\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 341709,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/12/2018 05:38:21",
          "content": "<p>This is interesting, it does not seem as good (yet) as dbscan on unrolled helix, but you run dbscan only once, not hundreds times.  The only caveat is you use a single event to validate, hard to compare with other approaches that way.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 341764,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "06/12/2018 08:18:01",
          "content": "<p>\"the only caveat is you use a single event to validate\" ...</p>\n\n<p>this experiment shows the approximate upper bound. I estimate the performance using huge train data, model tuning etc would be  about  -10% to 10% better or worse.</p>\n\n<p>It is not known if we train on other particles (e.g. momentum of certain range) and combine the results, would it be better or not.</p>\n\n<p>One thing to note is that track extension do not improve results here:</p>\n\n<p>metric_learning+dbscan+extend_track improves only from 0.005 to 0.010 (either the metric learning tracks are already well extended or they are broken and skip some hits in the middle) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 341798,
          "author_name": "makahana",
          "author_url": "",
          "post_date": "06/12/2018 09:39:35",
          "content": "<p>I think for siamese networks with triplet loss, how you sample your positive and negative examples are as important as hyperparameter tuning etc.  Need to sample negative examples that are hard(not just random) to get good convergence.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 341854,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "06/12/2018 11:50:14",
          "content": "<p>sampling quite straight forward in this case. just sample in KNN neighbours</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 347912,
          "author_name": "bhartha",
          "author_url": "",
          "post_date": "06/25/2018 16:54:16",
          "content": "<p>@Heng, can you please share (if possible), the feature extraction pytorch code you have used for this?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 341556,
      "author_name": "makahana",
      "author_url": "",
      "post_date": "06/11/2018 19:39:53",
      "content": "<p>How long does it approximately take to train to get to 0.26+ ?\nYou use 1 as the separation parameter in the triplet loss I take it?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 346897,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/22/2018 17:16:42",
      "content": "<p>for those who are using pytorch, here is pytorch compatible KNN and dbscan on gpu:</p>\n\n<p><a href=\"https://github.com/facebookresearch/faiss\">https://github.com/facebookresearch/faiss</a></p>\n\n<p><a href=\"https://github.com/karthikv2k/gpu_dbscan/blob/master/gpu_dbscan/gpu_dbscan.py\">https://github.com/karthikv2k/gpu_dbscan/blob/master/gpu_dbscan/gpu_dbscan.py</a></p>\n\n<p>\"Faiss is a library for efficient similarity search and clustering of dense vectors. It contains algorithms that search in sets of vectors of any size, up to ones that possibly do not fit in RAM. It also contains supporting code for evaluation and parameter tuning. Faiss is written in C++ with complete wrappers for Python/numpy. Some of the most useful algorithms are implemented on the GPU. It is developed by Facebook AI Research.\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 346898,
          "author_name": "johnhsweeney",
          "author_url": "",
          "post_date": "06/22/2018 17:22:08",
          "content": "<p>Excellent - have you tried it?  Does it provide effective parallelization for the clustering you are doing?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 346912,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "06/22/2018 18:09:52",
          "content": "<p>i haven't tried. The dbscan code doesn't seems to be fast. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 347034,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/23/2018 03:45:08",
          "content": "<p>Thanks for sharing, this looks interesting.  I don't think it depend on pytorch at all actually. The only link is that it is in the pytorch channel when installing via conda.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 347039,
      "author_name": "pavansanagapati",
      "author_url": "",
      "post_date": "06/23/2018 03:58:20",
      "content": "<p>Interesting stuff .Thanks for sharing Heng</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 347046,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/23/2018 04:12:45",
      "content": "<p>there is one more here:</p>\n\n<p><a href=\"https://github.com/ghoulsblade/CudaDBClustering/tree/master/src\">https://github.com/ghoulsblade/CudaDBClustering/tree/master/src</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 347052,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/23/2018 04:32:46",
      "content": "<p>check these as well</p>\n\n<p><a href=\"https://github.com/ZihengChen/ImageAlogrithm3D\">https://github.com/ZihengChen/ImageAlogrithm3D</a></p>\n\n<p><a href=\"https://galleryziheng.wordpress.com/2017/12/08/gpu-acceleration-of-imaging-algorithm/\">https://galleryziheng.wordpress.com/2017/12/08/gpu-acceleration-of-imaging-algorithm/</a></p>\n\n<p><a href=\"https://galleryziheng.wordpress.com/2017/08/27/get-python-clustering-work/\">https://galleryziheng.wordpress.com/2017/08/27/get-python-clustering-work/</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "338504": "Note:\n\n- this is not completed training code. It is meant to show the feasibility of using deep metric learning + dbscan.\n\n- no submission made to LB yet. \n\n- the results shows fitting a train sample (not on test sample). If we can fit a train sample, then we can generalize it on more train data.\n\n- the attached code perform the following tasks:\n\n     - input: xyz + transformation (e.g. r/d, sin(a), cos(a), ...)\n\n     - learned a embedded features using triplet loss\n\n     - build kdtree on embedded features and retrieve neighbors of a query hit \n\n   \n\n\n   ![enter image description here][1]\n\n   ![enter image description here][2]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/338504/9570/Slide1.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/338504/9569/Slide2.png",
    "340826": "results for running DBSCAN on embedded features on train data:\n\nfor half of the volume only (z&gt;0), trained on particles with momentum greater than 1:\n\n   - dbscan, eps=1: [ 0  000001029] score : 0.26149985   \n\n   - dbscan, eps=1.05: [ 0  000001029] score : 0.26531884   \n\n\n---\n\n\n        net.set_mode('test')\n        with torch.no_grad():\n            inputs = torch.from_numpy(hits).cuda()\n            features = net.forward( inputs )\n\n        #show results---\n        xyz    = df[['x','y','z']].values\n        features = features.data.cpu().numpy()\n \n        _,labels = dbscan(features,\n           eps=1,\n           min_samples=1,\n           algorithm='auto',\n           n_jobs=-1)\n\n        data_dir = '/root/share/project/kaggle/cern/data/__download__/train_100_events'\n        event_id= '000001029'\n        particles = pd.read_csv(data_dir + '/event%s-particles.csv'%event_id)\n        hits  = pd.read_csv(data_dir + '/event%s-hits.csv'%event_id)\n        cells = pd.read_csv(data_dir + '/event%s-cells.csv'%event_id)\n        truth = pd.read_csv(data_dir + '/event%s-truth.csv'%event_id)\n\n        track_id = np.zeros(len(hits),np.int32)\n        track_id[np.where(hits.z.values&gt;0)]  = labels\n        submission = pd.DataFrame(columns=['event_id', 'hit_id', 'track_id'],\n            data=np.column_stack(([int(event_id),]*len(hits), hits.hit_id.values, track_id))\n        ).astype(int)\n\n        score = score_event(truth, submission)\n        print('[%2d  %s] score : %0.8f   %0.1f min'%(0, event_id, score, (0)/60))",
    "341556": "How long does it approximately take to train to get to 0.26+ ?\nYou use 1 as the separation parameter in the triplet loss I take it?",
    "341709": "This is interesting, it does not seem as good (yet) as dbscan on unrolled helix, but you run dbscan only once, not hundreds times.  The only caveat is you use a single event to validate, hard to compare with other approaches that way.",
    "341764": "\"the only caveat is you use a single event to validate\" ...\n\nthis experiment shows the approximate upper bound. I estimate the performance using huge train data, model tuning etc would be  about  -10% to 10% better or worse.\n\nIt is not known if we train on other particles (e.g. momentum of certain range) and combine the results, would it be better or not.\n\nOne thing to note is that track extension do not improve results here:\n\nmetric_learning+dbscan+extend_track improves only from 0.005 to 0.010 (either the metric learning tracks are already well extended or they are broken and skip some hits in the middle)",
    "341798": "I think for siamese networks with triplet loss, how you sample your positive and negative examples are as important as hyperparameter tuning etc.  Need to sample negative examples that are hard(not just random) to get good convergence.",
    "341854": "sampling quite straight forward in this case. just sample in KNN neighbours",
    "346897": "for those who are using pytorch, here is pytorch compatible KNN and dbscan on gpu:\n\nhttps://github.com/facebookresearch/faiss\n\nhttps://github.com/karthikv2k/gpu_dbscan/blob/master/gpu_dbscan/gpu_dbscan.py\n\n\"Faiss is a library for efficient similarity search and clustering of dense vectors. It contains algorithms that search in sets of vectors of any size, up to ones that possibly do not fit in RAM. It also contains supporting code for evaluation and parameter tuning. Faiss is written in C++ with complete wrappers for Python/numpy. Some of the most useful algorithms are implemented on the GPU. It is developed by Facebook AI Research.\"",
    "346898": "Excellent - have you tried it?  Does it provide effective parallelization for the clustering you are doing?",
    "346912": "i haven't tried. The dbscan code doesn't seems to be fast.",
    "347034": "Thanks for sharing, this looks interesting.  I don't think it depend on pytorch at all actually. The only link is that it is in the pytorch channel when installing via conda.",
    "347039": "Interesting stuff .Thanks for sharing Heng",
    "347046": "there is one more here:\n\nhttps://github.com/ghoulsblade/CudaDBClustering/tree/master/src",
    "347052": "check these as well\n\nhttps://github.com/ZihengChen/ImageAlogrithm3D\n\nhttps://galleryziheng.wordpress.com/2017/12/08/gpu-acceleration-of-imaging-algorithm/\n\nhttps://galleryziheng.wordpress.com/2017/08/27/get-python-clustering-work/",
    "347912": "Heng, can you please share (if possible), the feature extraction pytorch code you have used for this?"
  },
  "source": "meta"
}