{
  "id": 56646,
  "title": "Supervised Learning",
  "url": "/competitions/trackml-particle-identification/discussion/56646",
  "author_name": "",
  "post_date": "2018-05-12T15:26:50.495451700Z",
  "votes": 6,
  "comment_count": 13,
  "views": 0,
  "content": "<p>This competition looks fun, topic is interesting, however, I understand it is mostly geometry not physics.</p>\n\n<p>I looked into existing ideas and if I am correct they are using Unsupervised Learning.</p>\n\n<p>We have huge training data, but with Unsupervised Learning all we can do with this data is to test our algorithms, not train them.</p>\n\n<p>How can we define our ML task here?</p>\n\n<p>We have large number of points and we want to group them into tracks.\nWe have large number of data rows and we want to classify them into large number of classes.</p>\n\n<p>We have hits: h1, h2, h3, h4, h5, ...\nFor each hit we have track: ht1, ht2, ht3, ht4, ht5,, ...\nIf h1 and h3 are on same track, then ht1 must be equal to ht3</p>\n\n<p>Let's define our X as h1, h2, h3, h4, h5, ....\nThen our Y is ht1, ht2, ht3, ht4, ht5, ... right?</p>\n\n<p>But what is value of Y here? It is not just 0 or 1, it is class id, but we have large number of classes!\nFor instance - we have 93680 hits and 7700 particles.\nIt means that our input vector is 93680, our output vector is 93680 and each output can have one of 7700 values.\nWith one hot encoding it will be 721336000 outputs.</p>\n\n<p>So let's try different approach.</p>\n\n<p>Again:\nX is h1, h2, h3, h4, h5, ....\nY is ht1, ht2, ht3, ht4, ht5, ...\nbut this time let's make binary classification</p>\n\n<p>So we have 93680 inputs and 93680 outputs. \n(each input is composed from at least 3 numbers; x, y, z)\nOutput 1 - hit is on this specific track\nOutput 0 - hit is not on this specific track\nBut which track is chosen one? I think we can assume it is track of \"h1\", so \"ht1\" is always 1.\n(then we can rotate our vector 93680  times to find all tracks)</p>\n\n<p>OK, we defined problem. Now can it be solved by any machine learning algorithm?</p>\n\n<p>How can we check is h1 and h2 on same track? Probably we can't, but we can with more points. So machine learning algorithm can learn somehow relation of multiple points. For instance it can find vector between h1 and h2, then between h1 and h3 then compare them and calc probability.</p>\n\n<p>I think it may work but not with tree algorithm (like xgboost) but with deep learning.</p>\n\n<p>But wait, let's continue our thinking.</p>\n\n<p>One event has 93680 hits but another one has 120939 hits.</p>\n\n<p>So we can't just use vector of constant size... But we have to! So should we set it to maximum size of event and fill with invalid data at end?</p>\n\n<p>Or can we optimize size of out input/output data?</p>\n\n<p>We can do following:</p>\n\n<pre><code>for hit1 in event:\n  for hit2 in event:\n     for hit3 in event:\n         classify(hit1, hit2, hit3)\n</code></pre>\n\n<p>but we have 822130284032000 iterations here, for 3 hits</p>\n\n<p>For my only conclusion is that we need some way to limit size of data before using it in any Supervised Learning algorithm. So first we have to do some preprocessing then perform Supervised Learning on smaller subsets.</p>\n\n<p>What are your thoughts and ideas? Of course you can tell me why I am wrong. </p>",
  "messages": [
    {
      "id": "327831",
      "postDate": "05/12/2018 15:26:50",
      "content": "<p>This competition looks fun, topic is interesting, however, I understand it is mostly geometry not physics.</p>\n\n<p>I looked into existing ideas and if I am correct they are using Unsupervised Learning.</p>\n\n<p>We have huge training data, but with Unsupervised Learning all we can do with this data is to test our algorithms, not train them.</p>\n\n<p>How can we define our ML task here?</p>\n\n<p>We have large number of points and we want to group them into tracks.\nWe have large number of data rows and we want to classify them into large number of classes.</p>\n\n<p>We have hits: h1, h2, h3, h4, h5, ...\nFor each hit we have track: ht1, ht2, ht3, ht4, ht5,, ...\nIf h1 and h3 are on same track, then ht1 must be equal to ht3</p>\n\n<p>Let's define our X as h1, h2, h3, h4, h5, ....\nThen our Y is ht1, ht2, ht3, ht4, ht5, ... right?</p>\n\n<p>But what is value of Y here? It is not just 0 or 1, it is class id, but we have large number of classes!\nFor instance - we have 93680 hits and 7700 particles.\nIt means that our input vector is 93680, our output vector is 93680 and each output can have one of 7700 values.\nWith one hot encoding it will be 721336000 outputs.</p>\n\n<p>So let's try different approach.</p>\n\n<p>Again:\nX is h1, h2, h3, h4, h5, ....\nY is ht1, ht2, ht3, ht4, ht5, ...\nbut this time let's make binary classification</p>\n\n<p>So we have 93680 inputs and 93680 outputs. \n(each input is composed from at least 3 numbers; x, y, z)\nOutput 1 - hit is on this specific track\nOutput 0 - hit is not on this specific track\nBut which track is chosen one? I think we can assume it is track of \"h1\", so \"ht1\" is always 1.\n(then we can rotate our vector 93680  times to find all tracks)</p>\n\n<p>OK, we defined problem. Now can it be solved by any machine learning algorithm?</p>\n\n<p>How can we check is h1 and h2 on same track? Probably we can't, but we can with more points. So machine learning algorithm can learn somehow relation of multiple points. For instance it can find vector between h1 and h2, then between h1 and h3 then compare them and calc probability.</p>\n\n<p>I think it may work but not with tree algorithm (like xgboost) but with deep learning.</p>\n\n<p>But wait, let's continue our thinking.</p>\n\n<p>One event has 93680 hits but another one has 120939 hits.</p>\n\n<p>So we can't just use vector of constant size... But we have to! So should we set it to maximum size of event and fill with invalid data at end?</p>\n\n<p>Or can we optimize size of out input/output data?</p>\n\n<p>We can do following:</p>\n\n<pre><code>for hit1 in event:\n  for hit2 in event:\n     for hit3 in event:\n         classify(hit1, hit2, hit3)\n</code></pre>\n\n<p>but we have 822130284032000 iterations here, for 3 hits</p>\n\n<p>For my only conclusion is that we need some way to limit size of data before using it in any Supervised Learning algorithm. So first we have to do some preprocessing then perform Supervised Learning on smaller subsets.</p>\n\n<p>What are your thoughts and ideas? Of course you can tell me why I am wrong. </p>",
      "rawMarkdown": "This competition looks fun, topic is interesting, however, I understand it is mostly geometry not physics.\n\nI looked into existing ideas and if I am correct they are using Unsupervised Learning.\n\nWe have huge training data, but with Unsupervised Learning all we can do with this data is to test our algorithms, not train them.\n\nHow can we define our ML task here?\n\nWe have large number of points and we want to group them into tracks.\nWe have large number of data rows and we want to classify them into large number of classes.\n\nWe have hits: h1, h2, h3, h4, h5, ...\nFor each hit we have track: ht1, ht2, ht3, ht4, ht5,, ...\nIf h1 and h3 are on same track, then ht1 must be equal to ht3\n\nLet's define our X as h1, h2, h3, h4, h5, ....\nThen our Y is ht1, ht2, ht3, ht4, ht5, ... right?\n\nBut what is value of Y here? It is not just 0 or 1, it is class id, but we have large number of classes!\nFor instance - we have 93680 hits and 7700 particles.\nIt means that our input vector is 93680, our output vector is 93680 and each output can have one of 7700 values.\nWith one hot encoding it will be 721336000 outputs.\n\nSo let's try different approach.\n\nAgain:\nX is h1, h2, h3, h4, h5, ....\nY is ht1, ht2, ht3, ht4, ht5, ...\nbut this time let's make binary classification\n\nSo we have 93680 inputs and 93680 outputs. \n(each input is composed from at least 3 numbers; x, y, z)\nOutput 1 - hit is on this specific track\nOutput 0 - hit is not on this specific track\nBut which track is chosen one? I think we can assume it is track of \"h1\", so \"ht1\" is always 1.\n(then we can rotate our vector 93680  times to find all tracks)\n\nOK, we defined problem. Now can it be solved by any machine learning algorithm?\n\nHow can we check is h1 and h2 on same track? Probably we can't, but we can with more points. So machine learning algorithm can learn somehow relation of multiple points. For instance it can find vector between h1 and h2, then between h1 and h3 then compare them and calc probability.\n\nI think it may work but not with tree algorithm (like xgboost) but with deep learning.\n\nBut wait, let's continue our thinking.\n\nOne event has 93680 hits but another one has 120939 hits.\n\nSo we can't just use vector of constant size... But we have to! So should we set it to maximum size of event and fill with invalid data at end?\n\nOr can we optimize size of out input/output data?\n\nWe can do following:\n\n    for hit1 in event:\n      for hit2 in event:\n         for hit3 in event:\n             classify(hit1, hit2, hit3)\n\nbut we have 822130284032000 iterations here, for 3 hits\n\nFor my only conclusion is that we need some way to limit size of data before using it in any Supervised Learning algorithm. So first we have to do some preprocessing then perform Supervised Learning on smaller subsets.\n\nWhat are your thoughts and ideas? Of course you can tell me why I am wrong.",
      "votes": null
    },
    {
      "id": "330024",
      "postDate": "05/17/2018 21:14:22",
      "content": "<p>Maybe it's time to try some regression.</p>",
      "rawMarkdown": "Maybe it's time to try some regression.",
      "votes": null
    },
    {
      "id": "330117",
      "postDate": "05/18/2018 04:21:05",
      "content": "<p>GANs can be used to find latent spaces based on valid data. GANs sample from \"random noise\" to \"generate\" valid data. \"random noise\" =&gt; sets of hits, \"generated\" data =&gt; tracks.</p>\n\n<p>So theoretically, one could create some kind of GAN that given a vector of hits would be able to filter out hits that are not part of a track and only return the hits part of a track. Unsupervised!</p>\n\n<p>From a supervised learning point of view I would imagine the use of auto-encoders to learn a latent space. Similar to GANs you could use it to \"de-noise\" the set of hits and filter out \"interesting\" hits. </p>\n\n<p>I can imagine RNN used to learn sequences of ... volumes, layers, modules, where to look for hits from the same track.</p>\n\n<p>There are some strict rules in regards to the relation between hits and tracks. Like:\n- \"particles do not interact with each others. A particle trajectory is not influenced in any way by other close-by particles\"</p>\n\n<p>Some \"rules\" might not be that obvious. </p>\n\n<p>But applying that rules on the hits would automatically filter out hits that have no way of being on the same track - drastically reducing the problem, or let's say the \"set\" of points one has to look for tracks.</p>\n\n<p>I would start by building a set of NNs  trained to filter out hits that would be part of \"particle_id = 0\" (noise). </p>\n\n<p>I think one could use ML to have it learn geometry and properly identify helices points out of set of points. </p>\n\n<p>Obviously there's a long way from theory to practice ;-)</p>",
      "rawMarkdown": "GANs can be used to find latent spaces based on valid data. GANs sample from \"random noise\" to \"generate\" valid data. \"random noise\" =&gt; sets of hits, \"generated\" data =&gt; tracks.\n\nSo theoretically, one could create some kind of GAN that given a vector of hits would be able to filter out hits that are not part of a track and only return the hits part of a track. Unsupervised!\n\nFrom a supervised learning point of view I would imagine the use of auto-encoders to learn a latent space. Similar to GANs you could use it to \"de-noise\" the set of hits and filter out \"interesting\" hits. \n\nI can imagine RNN used to learn sequences of ... volumes, layers, modules, where to look for hits from the same track.\n\nThere are some strict rules in regards to the relation between hits and tracks. Like:\n- \"particles do not interact with each others. A particle trajectory is not influenced in any way by other close-by particles\"\n\nSome \"rules\" might not be that obvious. \n\nBut applying that rules on the hits would automatically filter out hits that have no way of being on the same track - drastically reducing the problem, or let's say the \"set\" of points one has to look for tracks.\n\nI would start by building a set of NNs  trained to filter out hits that would be part of \"particle_id = 0\" (noise). \n\nI think one could use ML to have it learn geometry and properly identify helices points out of set of points. \n\nObviously there's a long way from theory to practice ;-)",
      "votes": null
    },
    {
      "id": "330423",
      "postDate": "05/18/2018 19:20:15",
      "content": "<p>Due to the nature of the problem, a brute force approach with all of the data is too much computation... current public kernels show usage of clustering to relatively quickly label some of the particles, and it has been shown (<a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/56580\">https://www.kaggle.com/c/trackml-particle-identification/discussion/56580</a>) that many of the paths are relatively simple. Starting with relatively quick steps to reduce the total amount of inputs would allow you to implement a final approach to more robustly calculate the noisy tracks or ones with a less intuitively modeled trajectory. </p>\n\n<p>Possibly something like finding all exactly linear tracks first, then tracks following a perfect helix, then finding nearly linear, and finally calculating the noisy tracks with a more robust classifier that can tolerate noise?</p>\n\n<p>The current high scoring public kernels cluster points relatively quickly and might be a good starting point? Possibly cluster points then find a way to assign a confidence level to the output, and remove the high confidence tracks from consideration?</p>",
      "rawMarkdown": "Due to the nature of the problem, a brute force approach with all of the data is too much computation... current public kernels show usage of clustering to relatively quickly label some of the particles, and it has been shown (https://www.kaggle.com/c/trackml-particle-identification/discussion/56580) that many of the paths are relatively simple. Starting with relatively quick steps to reduce the total amount of inputs would allow you to implement a final approach to more robustly calculate the noisy tracks or ones with a less intuitively modeled trajectory. \n\nPossibly something like finding all exactly linear tracks first, then tracks following a perfect helix, then finding nearly linear, and finally calculating the noisy tracks with a more robust classifier that can tolerate noise?\n\nThe current high scoring public kernels cluster points relatively quickly and might be a good starting point? Possibly cluster points then find a way to assign a confidence level to the output, and remove the high confidence tracks from consideration?",
      "votes": null
    },
    {
      "id": "332038",
      "postDate": "05/22/2018 11:48:07",
      "content": "<p>I'm thinking to try supervised learning to learn a (pairwise) distance metric function, and then use that metric for clustering.  With 130k hits **2 there are 16.9e9 hit combinations. Maybe some combinations can be eliminated beforehand, for instance hits with positive Z might not need to be combined with hits with negative Z. But some sort of incremental learning will likely be needed to learn across events anyway, so might as well use it inside each event?\nBut I think the first step is to see if hit combinations can be classified correctly as same-track/not-same-track.</p>",
      "rawMarkdown": "I'm thinking to try supervised learning to learn a (pairwise) distance metric function, and then use that metric for clustering.  With 130k hits **2 there are 16.9e9 hit combinations. Maybe some combinations can be eliminated beforehand, for instance hits with positive Z might not need to be combined with hits with negative Z. But some sort of incremental learning will likely be needed to learn across events anyway, so might as well use it inside each event?\nBut I think the first step is to see if hit combinations can be classified correctly as same-track/not-same-track.",
      "votes": null
    },
    {
      "id": "332040",
      "postDate": "05/22/2018 11:50:14",
      "content": "<p>A challenge in the hit-combinations space is the sparsity of same-track=1 compared to not-same-track. Anyone got a good approach for dealing with that?</p>",
      "rawMarkdown": "A challenge in the hit-combinations space is the sparsity of same-track=1 compared to not-same-track. Anyone got a good approach for dealing with that?",
      "votes": null
    },
    {
      "id": "332068",
      "postDate": "05/22/2018 12:59:54",
      "content": "<p>Good luck with that! However, I assumed that pair of hits is not enough to learn, because any two points could create straight line or something almost straight. But maybe you will find some useful feature this way.</p>",
      "rawMarkdown": "Good luck with that! However, I assumed that pair of hits is not enough to learn, because any two points could create straight line or something almost straight. But maybe you will find some useful feature this way.",
      "votes": null
    },
    {
      "id": "332073",
      "postDate": "05/22/2018 13:02:13",
      "content": "<blockquote>\n  <p>So theoretically, one could create some kind of GAN that given a\n  vector of hits would be able to filter out hits that are not part of a\n  track and only return the hits part of a track. Unsupervised!</p>\n</blockquote>\n\n<p>still - you need to define input somehow, should we take largest possible size, then copy event hits there and fill empty space with.... what? zeros?</p>",
      "rawMarkdown": "&gt; So theoretically, one could create some kind of GAN that given a\n&gt; vector of hits would be able to filter out hits that are not part of a\n&gt; track and only return the hits part of a track. Unsupervised!\n\nstill - you need to define input somehow, should we take largest possible size, then copy event hits there and fill empty space with.... what? zeros?",
      "votes": null
    },
    {
      "id": "332074",
      "postDate": "05/22/2018 13:03:46",
      "content": "<p>Probably clustering should be first step and supervised learning should be second step, because for supervised learning we need to filter-out lots of data.</p>",
      "rawMarkdown": "Probably clustering should be first step and supervised learning should be second step, because for supervised learning we need to filter-out lots of data.",
      "votes": null
    },
    {
      "id": "332171",
      "postDate": "05/22/2018 17:22:13",
      "content": "<p>What about learning the parameters of the helix that best approximates the particle track? Hits with similar helix-parameters end up in the same track. From what I understood, it should be possible to compute the theoretical helix parameters from the particle-files. I am not yet sure, however, whether this would be able to learn anything useful (assuming the inputs would be (x, y, z) from a single hit and the outputs are the helix parameters, which depends on your preferred parametrisation), especially for the track-origins.</p>",
      "rawMarkdown": "What about learning the parameters of the helix that best approximates the particle track? Hits with similar helix-parameters end up in the same track. From what I understood, it should be possible to compute the theoretical helix parameters from the particle-files. I am not yet sure, however, whether this would be able to learn anything useful (assuming the inputs would be (x, y, z) from a single hit and the outputs are the helix parameters, which depends on your preferred parametrisation), especially for the track-origins.",
      "votes": null
    },
    {
      "id": "332175",
      "postDate": "05/22/2018 17:30:31",
      "content": "<p>You don't need Supervised Learning or any Machine Learning to check if point is on some theoretical track or even straight line - it is simple math. There are infinite number of tracks you can assign to any hit. The problem here is that you need relations between points to find tracks. And number of points and tracks is large. \nCorrect me if I misunderstood you.</p>",
      "rawMarkdown": "You don't need Supervised Learning or any Machine Learning to check if point is on some theoretical track or even straight line - it is simple math. There are infinite number of tracks you can assign to any hit. The problem here is that you need relations between points to find tracks. And number of points and tracks is large. \nCorrect me if I misunderstood you.",
      "votes": null
    },
    {
      "id": "332177",
      "postDate": "05/22/2018 17:35:42",
      "content": "<p>I don't think the issue is with input, as it can be defined as being N number of \"sampled\" points. \nAnyway a GAN is not enough to solve the problem - it can be used maybe as a filtering mechanism for hits - use the GAN to train the discriminator into filtering out noise. Training the generator to generate valid tracks might not do any good. </p>",
      "rawMarkdown": "I don't think the issue is with input, as it can be defined as being N number of \"sampled\" points. \nAnyway a GAN is not enough to solve the problem - it can be used maybe as a filtering mechanism for hits - use the GAN to train the discriminator into filtering out noise. Training the generator to generate valid tracks might not do any good.",
      "votes": null
    },
    {
      "id": "332181",
      "postDate": "05/22/2018 17:39:23",
      "content": "<p>It was a rather impulsive idea that I wanted to share. The more I think about it, the less sense it seems to make, so probably you understand me better than I do. In a simple feed-forward network this would probably make little sense, but maybe using a sequence-to-sequence method, the model might be able to find relations between the different hits?</p>",
      "rawMarkdown": "It was a rather impulsive idea that I wanted to share. The more I think about it, the less sense it seems to make, so probably you understand me better than I do. In a simple feed-forward network this would probably make little sense, but maybe using a sequence-to-sequence method, the model might be able to find relations between the different hits?",
      "votes": null
    },
    {
      "id": "3517280",
      "postDate": "08/26/2026 16:30:11",
      "content": "<p>**Interesting tracking problem! **</p>\n<p>**🤔 Curious how different ML models would approach this. **</p>\n<p>This guide gives a quick look at where Classification vs. Regression fits: \n<a href=\"https://mlguidance.blogspot.com/2026/08/what-are-models-of-supervised-learning.html\" target=\"_blank\">https://mlguidance.blogspot.com/2026/08/what-are-models-of-supervised-learning.html</a></p>",
      "rawMarkdown": "**Interesting tracking problem! **\n\n**🤔 Curious how different ML models would approach this. **\n\nThis guide gives a quick look at where Classification vs. Regression fits: \nhttps://mlguidance.blogspot.com/2026/08/what-are-models-of-supervised-learning.html",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3517280,
      "author_name": "kaustubh994",
      "author_url": "",
      "post_date": "08/26/2026 16:30:11",
      "content": "<p>**Interesting tracking problem! **</p>\n<p>**🤔 Curious how different ML models would approach this. **</p>\n<p>This guide gives a quick look at where Classification vs. Regression fits: \n<a href=\"https://mlguidance.blogspot.com/2026/08/what-are-models-of-supervised-learning.html\" target=\"_blank\">https://mlguidance.blogspot.com/2026/08/what-are-models-of-supervised-learning.html</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 330024,
      "author_name": "maximantonovich",
      "author_url": "",
      "post_date": "05/17/2018 21:14:22",
      "content": "<p>Maybe it's time to try some regression.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 330117,
      "author_name": "profetul",
      "author_url": "",
      "post_date": "05/18/2018 04:21:05",
      "content": "<p>GANs can be used to find latent spaces based on valid data. GANs sample from \"random noise\" to \"generate\" valid data. \"random noise\" =&gt; sets of hits, \"generated\" data =&gt; tracks.</p>\n\n<p>So theoretically, one could create some kind of GAN that given a vector of hits would be able to filter out hits that are not part of a track and only return the hits part of a track. Unsupervised!</p>\n\n<p>From a supervised learning point of view I would imagine the use of auto-encoders to learn a latent space. Similar to GANs you could use it to \"de-noise\" the set of hits and filter out \"interesting\" hits. </p>\n\n<p>I can imagine RNN used to learn sequences of ... volumes, layers, modules, where to look for hits from the same track.</p>\n\n<p>There are some strict rules in regards to the relation between hits and tracks. Like:\n- \"particles do not interact with each others. A particle trajectory is not influenced in any way by other close-by particles\"</p>\n\n<p>Some \"rules\" might not be that obvious. </p>\n\n<p>But applying that rules on the hits would automatically filter out hits that have no way of being on the same track - drastically reducing the problem, or let's say the \"set\" of points one has to look for tracks.</p>\n\n<p>I would start by building a set of NNs  trained to filter out hits that would be part of \"particle_id = 0\" (noise). </p>\n\n<p>I think one could use ML to have it learn geometry and properly identify helices points out of set of points. </p>\n\n<p>Obviously there's a long way from theory to practice ;-)</p>",
      "votes": null,
      "replies": [
        {
          "id": 332073,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "05/22/2018 13:02:13",
          "content": "<blockquote>\n  <p>So theoretically, one could create some kind of GAN that given a\n  vector of hits would be able to filter out hits that are not part of a\n  track and only return the hits part of a track. Unsupervised!</p>\n</blockquote>\n\n<p>still - you need to define input somehow, should we take largest possible size, then copy event hits there and fill empty space with.... what? zeros?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 332177,
          "author_name": "profetul",
          "author_url": "",
          "post_date": "05/22/2018 17:35:42",
          "content": "<p>I don't think the issue is with input, as it can be defined as being N number of \"sampled\" points. \nAnyway a GAN is not enough to solve the problem - it can be used maybe as a filtering mechanism for hits - use the GAN to train the discriminator into filtering out noise. Training the generator to generate valid tracks might not do any good. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 330423,
      "author_name": "macfarll",
      "author_url": "",
      "post_date": "05/18/2018 19:20:15",
      "content": "<p>Due to the nature of the problem, a brute force approach with all of the data is too much computation... current public kernels show usage of clustering to relatively quickly label some of the particles, and it has been shown (<a href=\"https://www.kaggle.com/c/trackml-particle-identification/discussion/56580\">https://www.kaggle.com/c/trackml-particle-identification/discussion/56580</a>) that many of the paths are relatively simple. Starting with relatively quick steps to reduce the total amount of inputs would allow you to implement a final approach to more robustly calculate the noisy tracks or ones with a less intuitively modeled trajectory. </p>\n\n<p>Possibly something like finding all exactly linear tracks first, then tracks following a perfect helix, then finding nearly linear, and finally calculating the noisy tracks with a more robust classifier that can tolerate noise?</p>\n\n<p>The current high scoring public kernels cluster points relatively quickly and might be a good starting point? Possibly cluster points then find a way to assign a confidence level to the output, and remove the high confidence tracks from consideration?</p>",
      "votes": null,
      "replies": [
        {
          "id": 332074,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "05/22/2018 13:03:46",
          "content": "<p>Probably clustering should be first step and supervised learning should be second step, because for supervised learning we need to filter-out lots of data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 332038,
      "author_name": "jonnor",
      "author_url": "",
      "post_date": "05/22/2018 11:48:07",
      "content": "<p>I'm thinking to try supervised learning to learn a (pairwise) distance metric function, and then use that metric for clustering.  With 130k hits **2 there are 16.9e9 hit combinations. Maybe some combinations can be eliminated beforehand, for instance hits with positive Z might not need to be combined with hits with negative Z. But some sort of incremental learning will likely be needed to learn across events anyway, so might as well use it inside each event?\nBut I think the first step is to see if hit combinations can be classified correctly as same-track/not-same-track.</p>",
      "votes": null,
      "replies": [
        {
          "id": 332040,
          "author_name": "jonnor",
          "author_url": "",
          "post_date": "05/22/2018 11:50:14",
          "content": "<p>A challenge in the hit-combinations space is the sparsity of same-track=1 compared to not-same-track. Anyone got a good approach for dealing with that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 332068,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "05/22/2018 12:59:54",
          "content": "<p>Good luck with that! However, I assumed that pair of hits is not enough to learn, because any two points could create straight line or something almost straight. But maybe you will find some useful feature this way.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 332171,
      "author_name": "mrtsjolder",
      "author_url": "",
      "post_date": "05/22/2018 17:22:13",
      "content": "<p>What about learning the parameters of the helix that best approximates the particle track? Hits with similar helix-parameters end up in the same track. From what I understood, it should be possible to compute the theoretical helix parameters from the particle-files. I am not yet sure, however, whether this would be able to learn anything useful (assuming the inputs would be (x, y, z) from a single hit and the outputs are the helix parameters, which depends on your preferred parametrisation), especially for the track-origins.</p>",
      "votes": null,
      "replies": [
        {
          "id": 332175,
          "author_name": "jacekpoplawski",
          "author_url": "",
          "post_date": "05/22/2018 17:30:31",
          "content": "<p>You don't need Supervised Learning or any Machine Learning to check if point is on some theoretical track or even straight line - it is simple math. There are infinite number of tracks you can assign to any hit. The problem here is that you need relations between points to find tracks. And number of points and tracks is large. \nCorrect me if I misunderstood you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 332181,
          "author_name": "mrtsjolder",
          "author_url": "",
          "post_date": "05/22/2018 17:39:23",
          "content": "<p>It was a rather impulsive idea that I wanted to share. The more I think about it, the less sense it seems to make, so probably you understand me better than I do. In a simple feed-forward network this would probably make little sense, but maybe using a sequence-to-sequence method, the model might be able to find relations between the different hits?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "327831": "This competition looks fun, topic is interesting, however, I understand it is mostly geometry not physics.\n\nI looked into existing ideas and if I am correct they are using Unsupervised Learning.\n\nWe have huge training data, but with Unsupervised Learning all we can do with this data is to test our algorithms, not train them.\n\nHow can we define our ML task here?\n\nWe have large number of points and we want to group them into tracks.\nWe have large number of data rows and we want to classify them into large number of classes.\n\nWe have hits: h1, h2, h3, h4, h5, ...\nFor each hit we have track: ht1, ht2, ht3, ht4, ht5,, ...\nIf h1 and h3 are on same track, then ht1 must be equal to ht3\n\nLet's define our X as h1, h2, h3, h4, h5, ....\nThen our Y is ht1, ht2, ht3, ht4, ht5, ... right?\n\nBut what is value of Y here? It is not just 0 or 1, it is class id, but we have large number of classes!\nFor instance - we have 93680 hits and 7700 particles.\nIt means that our input vector is 93680, our output vector is 93680 and each output can have one of 7700 values.\nWith one hot encoding it will be 721336000 outputs.\n\nSo let's try different approach.\n\nAgain:\nX is h1, h2, h3, h4, h5, ....\nY is ht1, ht2, ht3, ht4, ht5, ...\nbut this time let's make binary classification\n\nSo we have 93680 inputs and 93680 outputs. \n(each input is composed from at least 3 numbers; x, y, z)\nOutput 1 - hit is on this specific track\nOutput 0 - hit is not on this specific track\nBut which track is chosen one? I think we can assume it is track of \"h1\", so \"ht1\" is always 1.\n(then we can rotate our vector 93680  times to find all tracks)\n\nOK, we defined problem. Now can it be solved by any machine learning algorithm?\n\nHow can we check is h1 and h2 on same track? Probably we can't, but we can with more points. So machine learning algorithm can learn somehow relation of multiple points. For instance it can find vector between h1 and h2, then between h1 and h3 then compare them and calc probability.\n\nI think it may work but not with tree algorithm (like xgboost) but with deep learning.\n\nBut wait, let's continue our thinking.\n\nOne event has 93680 hits but another one has 120939 hits.\n\nSo we can't just use vector of constant size... But we have to! So should we set it to maximum size of event and fill with invalid data at end?\n\nOr can we optimize size of out input/output data?\n\nWe can do following:\n\n    for hit1 in event:\n      for hit2 in event:\n         for hit3 in event:\n             classify(hit1, hit2, hit3)\n\nbut we have 822130284032000 iterations here, for 3 hits\n\nFor my only conclusion is that we need some way to limit size of data before using it in any Supervised Learning algorithm. So first we have to do some preprocessing then perform Supervised Learning on smaller subsets.\n\nWhat are your thoughts and ideas? Of course you can tell me why I am wrong.",
    "330024": "Maybe it's time to try some regression.",
    "330117": "GANs can be used to find latent spaces based on valid data. GANs sample from \"random noise\" to \"generate\" valid data. \"random noise\" =&gt; sets of hits, \"generated\" data =&gt; tracks.\n\nSo theoretically, one could create some kind of GAN that given a vector of hits would be able to filter out hits that are not part of a track and only return the hits part of a track. Unsupervised!\n\nFrom a supervised learning point of view I would imagine the use of auto-encoders to learn a latent space. Similar to GANs you could use it to \"de-noise\" the set of hits and filter out \"interesting\" hits. \n\nI can imagine RNN used to learn sequences of ... volumes, layers, modules, where to look for hits from the same track.\n\nThere are some strict rules in regards to the relation between hits and tracks. Like:\n- \"particles do not interact with each others. A particle trajectory is not influenced in any way by other close-by particles\"\n\nSome \"rules\" might not be that obvious. \n\nBut applying that rules on the hits would automatically filter out hits that have no way of being on the same track - drastically reducing the problem, or let's say the \"set\" of points one has to look for tracks.\n\nI would start by building a set of NNs  trained to filter out hits that would be part of \"particle_id = 0\" (noise). \n\nI think one could use ML to have it learn geometry and properly identify helices points out of set of points. \n\nObviously there's a long way from theory to practice ;-)",
    "330423": "Due to the nature of the problem, a brute force approach with all of the data is too much computation... current public kernels show usage of clustering to relatively quickly label some of the particles, and it has been shown (https://www.kaggle.com/c/trackml-particle-identification/discussion/56580) that many of the paths are relatively simple. Starting with relatively quick steps to reduce the total amount of inputs would allow you to implement a final approach to more robustly calculate the noisy tracks or ones with a less intuitively modeled trajectory. \n\nPossibly something like finding all exactly linear tracks first, then tracks following a perfect helix, then finding nearly linear, and finally calculating the noisy tracks with a more robust classifier that can tolerate noise?\n\nThe current high scoring public kernels cluster points relatively quickly and might be a good starting point? Possibly cluster points then find a way to assign a confidence level to the output, and remove the high confidence tracks from consideration?",
    "332038": "I'm thinking to try supervised learning to learn a (pairwise) distance metric function, and then use that metric for clustering.  With 130k hits **2 there are 16.9e9 hit combinations. Maybe some combinations can be eliminated beforehand, for instance hits with positive Z might not need to be combined with hits with negative Z. But some sort of incremental learning will likely be needed to learn across events anyway, so might as well use it inside each event?\nBut I think the first step is to see if hit combinations can be classified correctly as same-track/not-same-track.",
    "332040": "A challenge in the hit-combinations space is the sparsity of same-track=1 compared to not-same-track. Anyone got a good approach for dealing with that?",
    "332068": "Good luck with that! However, I assumed that pair of hits is not enough to learn, because any two points could create straight line or something almost straight. But maybe you will find some useful feature this way.",
    "332073": "&gt; So theoretically, one could create some kind of GAN that given a\n&gt; vector of hits would be able to filter out hits that are not part of a\n&gt; track and only return the hits part of a track. Unsupervised!\n\nstill - you need to define input somehow, should we take largest possible size, then copy event hits there and fill empty space with.... what? zeros?",
    "332074": "Probably clustering should be first step and supervised learning should be second step, because for supervised learning we need to filter-out lots of data.",
    "332171": "What about learning the parameters of the helix that best approximates the particle track? Hits with similar helix-parameters end up in the same track. From what I understood, it should be possible to compute the theoretical helix parameters from the particle-files. I am not yet sure, however, whether this would be able to learn anything useful (assuming the inputs would be (x, y, z) from a single hit and the outputs are the helix parameters, which depends on your preferred parametrisation), especially for the track-origins.",
    "332175": "You don't need Supervised Learning or any Machine Learning to check if point is on some theoretical track or even straight line - it is simple math. There are infinite number of tracks you can assign to any hit. The problem here is that you need relations between points to find tracks. And number of points and tracks is large. \nCorrect me if I misunderstood you.",
    "332177": "I don't think the issue is with input, as it can be defined as being N number of \"sampled\" points. \nAnyway a GAN is not enough to solve the problem - it can be used maybe as a filtering mechanism for hits - use the GAN to train the discriminator into filtering out noise. Training the generator to generate valid tracks might not do any good.",
    "332181": "It was a rather impulsive idea that I wanted to share. The more I think about it, the less sense it seems to make, so probably you understand me better than I do. In a simple feed-forward network this would probably make little sense, but maybe using a sequence-to-sequence method, the model might be able to find relations between the different hits?",
    "3517280": "**Interesting tracking problem! **\n\n**🤔 Curious how different ML models would approach this. **\n\nThis guide gives a quick look at where Classification vs. Regression fits: \nhttps://mlguidance.blogspot.com/2026/08/what-are-models-of-supervised-learning.html"
  },
  "source": "meta"
}