{
  "id": 59615,
  "title": "Word2Vec style embeddings for Track-ml",
  "url": "/competitions/trackml-particle-identification/discussion/59615",
  "author_name": "",
  "post_date": "2018-06-25T02:41:59.590911800Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>Since Word2Vec and other word embedding architectures are able to capture representations of words and group similar words close together, wouldn't training an unsupervised Word2Vec embedding on hits capture embeddings that would group hits from the same track together? </p>\n\n<p>Also, would the way you train Word2Vec (trying to keep the distance between the embeddings of two words that appear close together in the dataset small while keeping the distance between the embeddings of two words that appear far away in the dataset large) be transferable to track-ml?</p>\n\n<p>Thanks</p>",
  "messages": [
    {
      "id": "347664",
      "postDate": "06/25/2018 02:41:59",
      "content": "<p>Hi,</p>\n\n<p>Since Word2Vec and other word embedding architectures are able to capture representations of words and group similar words close together, wouldn't training an unsupervised Word2Vec embedding on hits capture embeddings that would group hits from the same track together? </p>\n\n<p>Also, would the way you train Word2Vec (trying to keep the distance between the embeddings of two words that appear close together in the dataset small while keeping the distance between the embeddings of two words that appear far away in the dataset large) be transferable to track-ml?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi,\n\nSince Word2Vec and other word embedding architectures are able to capture representations of words and group similar words close together, wouldn't training an unsupervised Word2Vec embedding on hits capture embeddings that would group hits from the same track together? \n\nAlso, would the way you train Word2Vec (trying to keep the distance between the embeddings of two words that appear close together in the dataset small while keeping the distance between the embeddings of two words that appear far away in the dataset large) be transferable to track-ml?\n\nThanks",
      "votes": null
    },
    {
      "id": "347677",
      "postDate": "06/25/2018 03:24:55",
      "content": "<p>This is interesting, but you'd want a parametric word embedding, as you won't get the exact same hit from one event to another one.  And if you think of it this comes closer to some of the approaches Heng is documenting on learning metric.</p>",
      "rawMarkdown": "This is interesting, but you'd want a parametric word embedding, as you won't get the exact same hit from one event to another one.  And if you think of it this comes closer to some of the approaches Heng is documenting on learning metric.",
      "votes": null
    },
    {
      "id": "347679",
      "postDate": "06/25/2018 03:25:29",
      "content": "<p>the problem of applying supervised learning to the trackML problem is the preprocessing of the data for input.</p>\n\n<p>you cannot simply classify individual hits. you need to find a way to get a track candidate (a group of hits), either in end-to-end or offline method.</p>\n\n<p>As an analogy, you are now given letters (not words). </p>\n\n<p>Once you figure a way to generate track candidates, the supervised learning is quite quite straight forward and many methods (including embedding) are possible.</p>",
      "rawMarkdown": "the problem of applying supervised learning to the trackML problem is the preprocessing of the data for input.\n\nyou cannot simply classify individual hits. you need to find a way to get a track candidate (a group of hits), either in end-to-end or offline method.\n\nAs an analogy, you are now given letters (not words). \n\nOnce you figure a way to generate track candidates, the supervised learning is quite quite straight forward and many methods (including embedding) are possible.",
      "votes": null
    },
    {
      "id": "347691",
      "postDate": "06/25/2018 04:33:18",
      "content": "<p>Here is something I tried that is basically unsupervised: learn an embedding vector for each event and each hit. For training, sample 4 hits and calculate a score from the inner product of their embedding vectors, and the target of this score is some metric of how well they fit on a helix. You can even calculate the distribution of the helix parameters and include this likelihood in your objective.</p>\n\n<p>I got it working on a small number of tracks. Biggest challenge of scaling up is to find the right distribution to sample from, otherwise it won't train at all (took me a lot of work to go from 3 tracks to 10 tracks, for example). Have thought about seeding the embedding with a clustering from other algorithms, but don't have time to pursue it atm.</p>",
      "rawMarkdown": "Here is something I tried that is basically unsupervised: learn an embedding vector for each event and each hit. For training, sample 4 hits and calculate a score from the inner product of their embedding vectors, and the target of this score is some metric of how well they fit on a helix. You can even calculate the distribution of the helix parameters and include this likelihood in your objective.\n\nI got it working on a small number of tracks. Biggest challenge of scaling up is to find the right distribution to sample from, otherwise it won't train at all (took me a lot of work to go from 3 tracks to 10 tracks, for example). Have thought about seeding the embedding with a clustering from other algorithms, but don't have time to pursue it atm.",
      "votes": null
    },
    {
      "id": "347803",
      "postDate": "06/25/2018 13:08:05",
      "content": "<p>I took the 2.4 Gcell to create a static bit array skeleton, selected the 500 khit to remove the unhit cells, and selected the closest one and farthest three hits from the highest pT track and a number of random \"background\" hits. I did this for 1024 events to create a workable sample. From this, 128 discriminators were created with four signal hits and four background hits. Each \"ring\" (stolen from quant lattice) of eight hits was convolved. In a sense, this is like a CNN...</p>\n\n<p>The AUC per discriminator ranged between .498 and .502, i.e. it is difficult.</p>\n\n<p>I enclose the result from 1024 CV events. As can be seen, some information can be extracted. However, it is difficult. The test time is O(n), which is the beauty of it. However, with my implementation, the train time would be over 350 years for 8850 events and 5000 tracks / event.</p>\n\n<p>Perhaps somebody will get some epiphanies from this...</p>",
      "rawMarkdown": "I took the 2.4 Gcell to create a static bit array skeleton, selected the 500 khit to remove the unhit cells, and selected the closest one and farthest three hits from the highest pT track and a number of random \"background\" hits. I did this for 1024 events to create a workable sample. From this, 128 discriminators were created with four signal hits and four background hits. Each \"ring\" (stolen from quant lattice) of eight hits was convolved. In a sense, this is like a CNN...\n\nThe AUC per discriminator ranged between .498 and .502, i.e. it is difficult.\n\nI enclose the result from 1024 CV events. As can be seen, some information can be extracted. However, it is difficult. The test time is O(n), which is the beauty of it. However, with my implementation, the train time would be over 350 years for 8850 events and 5000 tracks / event.\n\nPerhaps somebody will get some epiphanies from this...",
      "votes": null
    },
    {
      "id": "348184",
      "postDate": "06/26/2018 08:04:49",
      "content": "<p>I think, the problem of sequences can be dealt with words as well as letters, both are interchangeable, as far as a model is concerned. A seq of letters is word &amp; a seq of words is a sentence/context. A model can be made to learn both words and contexts.\nNow coming to the problem at hand a seq/cluster of hits is a track. it should not necessarily matter whether you treat the hits as words or letter. I might be wrong, please correct me.</p>",
      "rawMarkdown": "I think, the problem of sequences can be dealt with words as well as letters, both are interchangeable, as far as a model is concerned. A seq of letters is word &amp; a seq of words is a sentence/context. A model can be made to learn both words and contexts.\nNow coming to the problem at hand a seq/cluster of hits is a track. it should not necessarily matter whether you treat the hits as words or letter. I might be wrong, please correct me.",
      "votes": null
    },
    {
      "id": "348185",
      "postDate": "06/26/2018 08:09:07",
      "content": "<p>Let me expand on my answer.  If we are using some coordinates for the hits, then we need a parametric model, as it is very unlikely to get the exact same hit again in a different event.  If we use a finite number of 'letters' to represent hits, for instance volume, layer and module, then we are back to sequence prediction, and embeddings can be relevant.</p>",
      "rawMarkdown": "Let me expand on my answer.  If we are using some coordinates for the hits, then we need a parametric model, as it is very unlikely to get the exact same hit again in a different event.  If we use a finite number of 'letters' to represent hits, for instance volume, layer and module, then we are back to sequence prediction, and embeddings can be relevant.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 347677,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "06/25/2018 03:24:55",
      "content": "<p>This is interesting, but you'd want a parametric word embedding, as you won't get the exact same hit from one event to another one.  And if you think of it this comes closer to some of the approaches Heng is documenting on learning metric.</p>",
      "votes": null,
      "replies": [
        {
          "id": 348185,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/26/2018 08:09:07",
          "content": "<p>Let me expand on my answer.  If we are using some coordinates for the hits, then we need a parametric model, as it is very unlikely to get the exact same hit again in a different event.  If we use a finite number of 'letters' to represent hits, for instance volume, layer and module, then we are back to sequence prediction, and embeddings can be relevant.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 347679,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/25/2018 03:25:29",
      "content": "<p>the problem of applying supervised learning to the trackML problem is the preprocessing of the data for input.</p>\n\n<p>you cannot simply classify individual hits. you need to find a way to get a track candidate (a group of hits), either in end-to-end or offline method.</p>\n\n<p>As an analogy, you are now given letters (not words). </p>\n\n<p>Once you figure a way to generate track candidates, the supervised learning is quite quite straight forward and many methods (including embedding) are possible.</p>",
      "votes": null,
      "replies": [
        {
          "id": 348184,
          "author_name": "ramankishore",
          "author_url": "",
          "post_date": "06/26/2018 08:04:49",
          "content": "<p>I think, the problem of sequences can be dealt with words as well as letters, both are interchangeable, as far as a model is concerned. A seq of letters is word &amp; a seq of words is a sentence/context. A model can be made to learn both words and contexts.\nNow coming to the problem at hand a seq/cluster of hits is a track. it should not necessarily matter whether you treat the hits as words or letter. I might be wrong, please correct me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 347691,
      "author_name": "sangxia",
      "author_url": "",
      "post_date": "06/25/2018 04:33:18",
      "content": "<p>Here is something I tried that is basically unsupervised: learn an embedding vector for each event and each hit. For training, sample 4 hits and calculate a score from the inner product of their embedding vectors, and the target of this score is some metric of how well they fit on a helix. You can even calculate the distribution of the helix parameters and include this likelihood in your objective.</p>\n\n<p>I got it working on a small number of tracks. Biggest challenge of scaling up is to find the right distribution to sample from, otherwise it won't train at all (took me a lot of work to go from 3 tracks to 10 tracks, for example). Have thought about seeding the embedding with a clustering from other algorithms, but don't have time to pursue it atm.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 347803,
      "author_name": "glimmung",
      "author_url": "",
      "post_date": "06/25/2018 13:08:05",
      "content": "<p>I took the 2.4 Gcell to create a static bit array skeleton, selected the 500 khit to remove the unhit cells, and selected the closest one and farthest three hits from the highest pT track and a number of random \"background\" hits. I did this for 1024 events to create a workable sample. From this, 128 discriminators were created with four signal hits and four background hits. Each \"ring\" (stolen from quant lattice) of eight hits was convolved. In a sense, this is like a CNN...</p>\n\n<p>The AUC per discriminator ranged between .498 and .502, i.e. it is difficult.</p>\n\n<p>I enclose the result from 1024 CV events. As can be seen, some information can be extracted. However, it is difficult. The test time is O(n), which is the beauty of it. However, with my implementation, the train time would be over 350 years for 8850 events and 5000 tracks / event.</p>\n\n<p>Perhaps somebody will get some epiphanies from this...</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "347664": "Hi,\n\nSince Word2Vec and other word embedding architectures are able to capture representations of words and group similar words close together, wouldn't training an unsupervised Word2Vec embedding on hits capture embeddings that would group hits from the same track together? \n\nAlso, would the way you train Word2Vec (trying to keep the distance between the embeddings of two words that appear close together in the dataset small while keeping the distance between the embeddings of two words that appear far away in the dataset large) be transferable to track-ml?\n\nThanks",
    "347677": "This is interesting, but you'd want a parametric word embedding, as you won't get the exact same hit from one event to another one.  And if you think of it this comes closer to some of the approaches Heng is documenting on learning metric.",
    "347679": "the problem of applying supervised learning to the trackML problem is the preprocessing of the data for input.\n\nyou cannot simply classify individual hits. you need to find a way to get a track candidate (a group of hits), either in end-to-end or offline method.\n\nAs an analogy, you are now given letters (not words). \n\nOnce you figure a way to generate track candidates, the supervised learning is quite quite straight forward and many methods (including embedding) are possible.",
    "347691": "Here is something I tried that is basically unsupervised: learn an embedding vector for each event and each hit. For training, sample 4 hits and calculate a score from the inner product of their embedding vectors, and the target of this score is some metric of how well they fit on a helix. You can even calculate the distribution of the helix parameters and include this likelihood in your objective.\n\nI got it working on a small number of tracks. Biggest challenge of scaling up is to find the right distribution to sample from, otherwise it won't train at all (took me a lot of work to go from 3 tracks to 10 tracks, for example). Have thought about seeding the embedding with a clustering from other algorithms, but don't have time to pursue it atm.",
    "347803": "I took the 2.4 Gcell to create a static bit array skeleton, selected the 500 khit to remove the unhit cells, and selected the closest one and farthest three hits from the highest pT track and a number of random \"background\" hits. I did this for 1024 events to create a workable sample. From this, 128 discriminators were created with four signal hits and four background hits. Each \"ring\" (stolen from quant lattice) of eight hits was convolved. In a sense, this is like a CNN...\n\nThe AUC per discriminator ranged between .498 and .502, i.e. it is difficult.\n\nI enclose the result from 1024 CV events. As can be seen, some information can be extracted. However, it is difficult. The test time is O(n), which is the beauty of it. However, with my implementation, the train time would be over 350 years for 8850 events and 5000 tracks / event.\n\nPerhaps somebody will get some epiphanies from this...",
    "348184": "I think, the problem of sequences can be dealt with words as well as letters, both are interchangeable, as far as a model is concerned. A seq of letters is word &amp; a seq of words is a sentence/context. A model can be made to learn both words and contexts.\nNow coming to the problem at hand a seq/cluster of hits is a track. it should not necessarily matter whether you treat the hits as words or letter. I might be wrong, please correct me.",
    "348185": "Let me expand on my answer.  If we are using some coordinates for the hits, then we need a parametric model, as it is very unlikely to get the exact same hit again in a different event.  If we use a finite number of 'letters' to represent hits, for instance volume, layer and module, then we are back to sequence prediction, and embeddings can be relevant."
  },
  "source": "meta"
}