{
  "id": 60899,
  "title": "Solving the real problem, cell and timing data, changing direction",
  "url": "/competitions/trackml-particle-identification/discussion/60899",
  "author_name": "",
  "post_date": "2018-07-11T17:51:14.772424200Z",
  "votes": 1,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Dear TrackML Team:</p>\n\n<p>I just joined this group recently.  My first question to myself was, \"Why are they not working together to design better detectors and signal processors?\"  and \"If they invert this model all they will do is reverse engineer the simulator and not the real thing.  It is a good game, but not the real thing.  The real problem is that the information coming from the experiments, as they are using it now, is ambiguous.  That probably requires some new signal processing hardware, better data management and better algorithms.\"</p>\n\n<p>With this amount of data, what you are getting is about what you can get.  But if you want nearly perfect tracking, you will need to get some more very specific clues.</p>\n\n<p><strong>The first clue is timing</strong>.  Someone here said sub-nanosecond timing is \"impossible\".  But I know that the ADC technology has progressed to 10 Gsps (giga samples per second) with good resolution.  I am not selling these chips, I have been tracking their progress for application to high sampling rate gravimeter arrays.  There are plenty of technologies to time stamp some layers of the z,x and y axes to help resolve sequences of events.  The more ways you can break down the timing of the hits from a collision, the faster your algorithms.  Investing in getting timing data reduces the time cost of the trajectory and particle identification.   I would aim to be able to process and classify all the data and not lose any.</p>\n\n<p>Recommendation: Run the \"truth\" simulation and give the TrackML team the timings - as though you could do that in the real machine.  I think you are going to see this problem is almost trivial to unravel.  Then go back and figure how to get that time information ( or some proxies ) from the real equipment.  Fix the real machines. Stop playing a game you cannot win.</p>\n\n<p><a href=\"http://www.analog.com/en/products/analog-to-digital-converters/standard-adc/high-speed-ad-10msps/ad9213.html\">http://www.analog.com/en/products/analog-to-digital-converters/standard-adc/high-speed-ad-10msps/ad9213.html</a></p>\n\n<p><strong>The second clue is cell hit geometry</strong>.  For many of the hits, the cell geometry, regardless of time sequence, gives tight constraints on the possible trajectories.  The particle is going one way or another along the line of cells hit.  </p>\n\n<p>Recommendation:  Explicitly provide the cell locations and geometries.  A more true test here, is full cell information, better organized and documented.  Run the simulation and break the hits into cell level data.  Test using cell data, not averaged hit data.</p>\n\n<p>Until you can modify the sensors to get timing, you can probably get much closer by using the cell hit geometric constraints.  There might be many possible cell trajectories for a hit based on cell data, but this space of trajectories, when compared to the whole space on independent hits is much smaller.  The hit turns from a point into a beam with a fairly narrow range of likely directions. When stitching together trajectories from hits, most of the hits have multiple cell points and therefore beams.</p>\n\n<p><strong>Changing from a game to a real problem:</strong></p>\n\n<p>I am not criticizing.  Personally, I like playing with lots of numbers and inverting models.  But I joined because I thought I might see some real data.  I would like to solve for the magnetic field structure and temporal variations.  I would like to see a massive dataset of low energy electron positron events, and other particle-antiparticle events.  I would like to look at the entire dataset.  I was dismayed that most of the data is being ignored (not used even when it is gathered), thrown away without deep examination, and not shared particularly well - (For the world at large, it is the more common data that is more important than things the specialists care about).</p>\n\n<p>If this many people are going to expend time to solve this problem, it is not for the money or the fame.  Would you all consider changing direction and solve the real problem?  Kaggle is throwing money at lots of problems, for their own purposes;  hoping to get some great solutions, to further advertise themselves.  This is not a criticism.  I review these sorts of things, and I am simply telling you that most collaborative websites are self-serving.  It is a phase they all go through until they understand their true goals.  I have a pretty good idea what they should be doing, but one thing at a time.</p>\n\n<p>You are probably not investing much into this, since it is such a small amount of money and fame at stake.  But if you (collectively) solve a real problem, the stakes are much higher.</p>\n\n<p>Change this from a puzzle for a prize, to a real effort to drastically reduce the cost of the next generation of accelerator upgrades.  Rather than thousands go after billions.</p>\n\n<p>I apologize for throwing this in your laps on my first hello, but I have a lot on my plate and not much time left.</p>\n\n<p>I am a senior mathematical statistician.  I spent my life on global problems - international development, global economic and social modeling, famine early warning, clean air, alternative fuels, global climate change, nonprofits, industry modeling, technology modeling, business intelligence, Y2K, and for the last 20 years the relation between society and the Internet.  My personal interests are dna genealogy, 3D technologies, and gravimeter imaging arrays.   </p>\n\n<p>Sincere regards,\nRichard Collins, The Internet Foundation</p>",
  "messages": [
    {
      "id": "355450",
      "postDate": "07/11/2018 17:51:14",
      "content": "<p>Dear TrackML Team:</p>\n\n<p>I just joined this group recently.  My first question to myself was, \"Why are they not working together to design better detectors and signal processors?\"  and \"If they invert this model all they will do is reverse engineer the simulator and not the real thing.  It is a good game, but not the real thing.  The real problem is that the information coming from the experiments, as they are using it now, is ambiguous.  That probably requires some new signal processing hardware, better data management and better algorithms.\"</p>\n\n<p>With this amount of data, what you are getting is about what you can get.  But if you want nearly perfect tracking, you will need to get some more very specific clues.</p>\n\n<p><strong>The first clue is timing</strong>.  Someone here said sub-nanosecond timing is \"impossible\".  But I know that the ADC technology has progressed to 10 Gsps (giga samples per second) with good resolution.  I am not selling these chips, I have been tracking their progress for application to high sampling rate gravimeter arrays.  There are plenty of technologies to time stamp some layers of the z,x and y axes to help resolve sequences of events.  The more ways you can break down the timing of the hits from a collision, the faster your algorithms.  Investing in getting timing data reduces the time cost of the trajectory and particle identification.   I would aim to be able to process and classify all the data and not lose any.</p>\n\n<p>Recommendation: Run the \"truth\" simulation and give the TrackML team the timings - as though you could do that in the real machine.  I think you are going to see this problem is almost trivial to unravel.  Then go back and figure how to get that time information ( or some proxies ) from the real equipment.  Fix the real machines. Stop playing a game you cannot win.</p>\n\n<p><a href=\"http://www.analog.com/en/products/analog-to-digital-converters/standard-adc/high-speed-ad-10msps/ad9213.html\">http://www.analog.com/en/products/analog-to-digital-converters/standard-adc/high-speed-ad-10msps/ad9213.html</a></p>\n\n<p><strong>The second clue is cell hit geometry</strong>.  For many of the hits, the cell geometry, regardless of time sequence, gives tight constraints on the possible trajectories.  The particle is going one way or another along the line of cells hit.  </p>\n\n<p>Recommendation:  Explicitly provide the cell locations and geometries.  A more true test here, is full cell information, better organized and documented.  Run the simulation and break the hits into cell level data.  Test using cell data, not averaged hit data.</p>\n\n<p>Until you can modify the sensors to get timing, you can probably get much closer by using the cell hit geometric constraints.  There might be many possible cell trajectories for a hit based on cell data, but this space of trajectories, when compared to the whole space on independent hits is much smaller.  The hit turns from a point into a beam with a fairly narrow range of likely directions. When stitching together trajectories from hits, most of the hits have multiple cell points and therefore beams.</p>\n\n<p><strong>Changing from a game to a real problem:</strong></p>\n\n<p>I am not criticizing.  Personally, I like playing with lots of numbers and inverting models.  But I joined because I thought I might see some real data.  I would like to solve for the magnetic field structure and temporal variations.  I would like to see a massive dataset of low energy electron positron events, and other particle-antiparticle events.  I would like to look at the entire dataset.  I was dismayed that most of the data is being ignored (not used even when it is gathered), thrown away without deep examination, and not shared particularly well - (For the world at large, it is the more common data that is more important than things the specialists care about).</p>\n\n<p>If this many people are going to expend time to solve this problem, it is not for the money or the fame.  Would you all consider changing direction and solve the real problem?  Kaggle is throwing money at lots of problems, for their own purposes;  hoping to get some great solutions, to further advertise themselves.  This is not a criticism.  I review these sorts of things, and I am simply telling you that most collaborative websites are self-serving.  It is a phase they all go through until they understand their true goals.  I have a pretty good idea what they should be doing, but one thing at a time.</p>\n\n<p>You are probably not investing much into this, since it is such a small amount of money and fame at stake.  But if you (collectively) solve a real problem, the stakes are much higher.</p>\n\n<p>Change this from a puzzle for a prize, to a real effort to drastically reduce the cost of the next generation of accelerator upgrades.  Rather than thousands go after billions.</p>\n\n<p>I apologize for throwing this in your laps on my first hello, but I have a lot on my plate and not much time left.</p>\n\n<p>I am a senior mathematical statistician.  I spent my life on global problems - international development, global economic and social modeling, famine early warning, clean air, alternative fuels, global climate change, nonprofits, industry modeling, technology modeling, business intelligence, Y2K, and for the last 20 years the relation between society and the Internet.  My personal interests are dna genealogy, 3D technologies, and gravimeter imaging arrays.   </p>\n\n<p>Sincere regards,\nRichard Collins, The Internet Foundation</p>",
      "rawMarkdown": "Dear TrackML Team:\n\nI just joined this group recently.  My first question to myself was, \"Why are they not working together to design better detectors and signal processors?\"  and \"If they invert this model all they will do is reverse engineer the simulator and not the real thing.  It is a good game, but not the real thing.  The real problem is that the information coming from the experiments, as they are using it now, is ambiguous.  That probably requires some new signal processing hardware, better data management and better algorithms.\"\n\nWith this amount of data, what you are getting is about what you can get.  But if you want nearly perfect tracking, you will need to get some more very specific clues.\n\n**The first clue is timing**.  Someone here said sub-nanosecond timing is \"impossible\".  But I know that the ADC technology has progressed to 10 Gsps (giga samples per second) with good resolution.  I am not selling these chips, I have been tracking their progress for application to high sampling rate gravimeter arrays.  There are plenty of technologies to time stamp some layers of the z,x and y axes to help resolve sequences of events.  The more ways you can break down the timing of the hits from a collision, the faster your algorithms.  Investing in getting timing data reduces the time cost of the trajectory and particle identification.   I would aim to be able to process and classify all the data and not lose any.\n\nRecommendation: Run the \"truth\" simulation and give the TrackML team the timings - as though you could do that in the real machine.  I think you are going to see this problem is almost trivial to unravel.  Then go back and figure how to get that time information ( or some proxies ) from the real equipment.  Fix the real machines. Stop playing a game you cannot win.\n\nhttp://www.analog.com/en/products/analog-to-digital-converters/standard-adc/high-speed-ad-10msps/ad9213.html\n\n**The second clue is cell hit geometry**.  For many of the hits, the cell geometry, regardless of time sequence, gives tight constraints on the possible trajectories.  The particle is going one way or another along the line of cells hit.  \n\nRecommendation:  Explicitly provide the cell locations and geometries.  A more true test here, is full cell information, better organized and documented.  Run the simulation and break the hits into cell level data.  Test using cell data, not averaged hit data.\n\nUntil you can modify the sensors to get timing, you can probably get much closer by using the cell hit geometric constraints.  There might be many possible cell trajectories for a hit based on cell data, but this space of trajectories, when compared to the whole space on independent hits is much smaller.  The hit turns from a point into a beam with a fairly narrow range of likely directions. When stitching together trajectories from hits, most of the hits have multiple cell points and therefore beams.\n\n**Changing from a game to a real problem:**\n\nI am not criticizing.  Personally, I like playing with lots of numbers and inverting models.  But I joined because I thought I might see some real data.  I would like to solve for the magnetic field structure and temporal variations.  I would like to see a massive dataset of low energy electron positron events, and other particle-antiparticle events.  I would like to look at the entire dataset.  I was dismayed that most of the data is being ignored (not used even when it is gathered), thrown away without deep examination, and not shared particularly well - (For the world at large, it is the more common data that is more important than things the specialists care about).\n\nIf this many people are going to expend time to solve this problem, it is not for the money or the fame.  Would you all consider changing direction and solve the real problem?  Kaggle is throwing money at lots of problems, for their own purposes;  hoping to get some great solutions, to further advertise themselves.  This is not a criticism.  I review these sorts of things, and I am simply telling you that most collaborative websites are self-serving.  It is a phase they all go through until they understand their true goals.  I have a pretty good idea what they should be doing, but one thing at a time.\n\nYou are probably not investing much into this, since it is such a small amount of money and fame at stake.  But if you (collectively) solve a real problem, the stakes are much higher.\n\nChange this from a puzzle for a prize, to a real effort to drastically reduce the cost of the next generation of accelerator upgrades.  Rather than thousands go after billions.\n\nI apologize for throwing this in your laps on my first hello, but I have a lot on my plate and not much time left.\n\nI am a senior mathematical statistician.  I spent my life on global problems - international development, global economic and social modeling, famine early warning, clean air, alternative fuels, global climate change, nonprofits, industry modeling, technology modeling, business intelligence, Y2K, and for the last 20 years the relation between society and the Internet.  My personal interests are dna genealogy, 3D technologies, and gravimeter imaging arrays.   \n\nSincere regards,\nRichard Collins, The Internet Foundation",
      "votes": null
    },
    {
      "id": "355461",
      "postDate": "07/11/2018 18:20:52",
      "content": "<p>cell geometry is provided.</p>",
      "rawMarkdown": "cell geometry is provided.",
      "votes": null
    },
    {
      "id": "355509",
      "postDate": "07/11/2018 20:59:23",
      "content": "<p>Dear CPMP,</p>\n\n<p>Yes.  Sorry, I am very tired.  I was writing quickly to see if anyone is listening.  I was thinking more along the lines of having the test provide the cell locations directly, rather than making everyone look them up.  Collapsing into \"hits\" is something you did to shorten the test files, but the data ultimately is available, or could be available, at the cell level.   I would provide one \"cell hits file\" that contains all the information needed for cell-level resolution of the hits.  I do not think you are testing people's ability to join data files and look up things.  Rather it is to test new ways to get to the physics?  I took liberties writing quickly to say \"better cell data\", but I was really thinking of getting all the information into an array of cell objects with ALL the necessary information.  </p>\n\n<p>It becomes a partnership of the test data providers continually improving the readability and usefulness of the test data, and the others trying to find the best and fastest algorithms for classifying into tracks, particles, particle events and vacuum processes (I lump all the quark gluon and many electromagnetic events into vacuum properties.  That is my personal view of things.)</p>\n\n<p>What did you think about running the simulation, so as to provide timing data?  Assuming picosecond resolution, for instance?  Or a range of time resolutions on hits?  I think the people gathered here can solve that easily, if they have already tried to solve your current problem, and the results might better inform decision-making.  Your test fairly faithfully represents the present situation and its data limitations.  The problem I think you ought to solve is how to change the detectors and signal processing to capture enough information to resolve the particles and events, and not throw away so much good information.  So I am suggesting you run the time resolved simulation and let these people look at it, while everyone has these things in mind.  It won't affect the contest much, but might provide some new insights toward better algorithms.  More, I think it will move everyone toward a solution to reduce the cost of upgrades to equipment and get to where you can capture all the data.</p>\n\n<p>I think the memory for the event data is too small, relative to the full size of the data stream.  This is a vast simplification, but it seems you are trying to represent a many gigabyte event with a few megabytes of data.  No matter how much you want to resolve it, there is not enough to do the filtering precisely and efficiently.  So you are pretty much guaranteed to lose.  If you compare the difficulty of capturing and filter a gigabyte per event, to the huge cost of unraveling poorly documented hit data after the event, I think you might see what I am aiming at.</p>\n\n<p>Richard Collins</p>",
      "rawMarkdown": "Dear CPMP,\n\nYes.  Sorry, I am very tired.  I was writing quickly to see if anyone is listening.  I was thinking more along the lines of having the test provide the cell locations directly, rather than making everyone look them up.  Collapsing into \"hits\" is something you did to shorten the test files, but the data ultimately is available, or could be available, at the cell level.   I would provide one \"cell hits file\" that contains all the information needed for cell-level resolution of the hits.  I do not think you are testing people's ability to join data files and look up things.  Rather it is to test new ways to get to the physics?  I took liberties writing quickly to say \"better cell data\", but I was really thinking of getting all the information into an array of cell objects with ALL the necessary information.  \n\nIt becomes a partnership of the test data providers continually improving the readability and usefulness of the test data, and the others trying to find the best and fastest algorithms for classifying into tracks, particles, particle events and vacuum processes (I lump all the quark gluon and many electromagnetic events into vacuum properties.  That is my personal view of things.)\n\nWhat did you think about running the simulation, so as to provide timing data?  Assuming picosecond resolution, for instance?  Or a range of time resolutions on hits?  I think the people gathered here can solve that easily, if they have already tried to solve your current problem, and the results might better inform decision-making.  Your test fairly faithfully represents the present situation and its data limitations.  The problem I think you ought to solve is how to change the detectors and signal processing to capture enough information to resolve the particles and events, and not throw away so much good information.  So I am suggesting you run the time resolved simulation and let these people look at it, while everyone has these things in mind.  It won't affect the contest much, but might provide some new insights toward better algorithms.  More, I think it will move everyone toward a solution to reduce the cost of upgrades to equipment and get to where you can capture all the data.\n\nI think the memory for the event data is too small, relative to the full size of the data stream.  This is a vast simplification, but it seems you are trying to represent a many gigabyte event with a few megabytes of data.  No matter how much you want to resolve it, there is not enough to do the filtering precisely and efficiently.  So you are pretty much guaranteed to lose.  If you compare the difficulty of capturing and filter a gigabyte per event, to the huge cost of unraveling poorly documented hit data after the event, I think you might see what I am aiming at.\n\nRichard Collins",
      "votes": null
    },
    {
      "id": "355532",
      "postDate": "07/11/2018 23:06:53",
      "content": "<p>It looks as if the simulation went \"We have all these particleIds, what cells did they hit and what value would it yield?\" So, in eventId 1000, hitIds 14589, 95, 604, 5, 13 &amp; 4 (with six particleIds) have the same cellId (305/5) and similar values (.25-.37). Does the cell contain sub cells?</p>\n\n<p>I still have no answer from the organizers on this one...</p>",
      "rawMarkdown": "It looks as if the simulation went \"We have all these particleIds, what cells did they hit and what value would it yield?\" So, in eventId 1000, hitIds 14589, 95, 604, 5, 13 &amp; 4 (with six particleIds) have the same cellId (305/5) and similar values (.25-.37). Does the cell contain sub cells?\n\nI still have no answer from the organizers on this one...",
      "votes": null
    },
    {
      "id": "355554",
      "postDate": "07/12/2018 00:51:37",
      "content": "<p>Dear Glimmung,</p>\n\n<p>I expect that they ran the simulation, it generated those six particles, the particles went through cell 305-5, and all deposited similar amounts of charge/energy.  </p>\n\n<p>If the particles all had nearly identical initial properties, and went through the one point, then hopefully they are sufficiently different elsewhere to unravel them.  But, in any case, where there are two or more cells in one sensor, their signals can say something about the energy and direction through that cell.  That is a lot easier to use as a constraint than a huge number of completely independent \"hits\" with no direction or energy attached. </p>\n\n<p>In your example, from Test_100 by the way, the cell 305-5 has 23 hits associated with it (I sorted by ch0 then ch1).   These are the hit numbers 14589, 14595, 14604, 14605, 14613, 14614, 72232, 82386, 109437, 110299, 110383, 110753, 110952, 111081, 112450, 113187, 114008, 114311, 115134, 115614, 115854, 116161, 116540.</p>\n\n<p>In the hits file these correspond to many particles, the ones all starting with 6395 have very similar properties, and are the ones you are probably referring to.  These are 639511834298179000, 639511834298167000, 639511834298171000, 639511834298187000, 639511834298175000, 639511834298183000.</p>\n\n<p>If you propagate these particles from the origin through the magnetic field, and they all have to go through that one cell, then that tells you a bit about their mass, type, charge and the time they go through each cell. </p>\n\n<p>These other particles : 968286014512562000, 734097665658191000, 238692841851723000, 864702054851936000, 589971757343965000, 734106118170611000, 499903269489868000, 117094071364751000, 864698000402808000, 292762425642450000, 797156306778587000, 418853731921035000 all go through cell 305-5 as well.</p>\n\n<p>I presume the cell identifier is unique and not used in different sensors.  I get the impression that providing the cell data is an afterthought.</p>\n\n<p>I would start with the hits having more than one cell, work out the constraints on the trajectories, assign those particles first, then use another method for the other tracks.  My guess is the multi-cell hits are going to be more interesting any way.  It is nice the constraints are tighter.</p>\n\n<p>Richard</p>",
      "rawMarkdown": "Dear Glimmung,\n\nI expect that they ran the simulation, it generated those six particles, the particles went through cell 305-5, and all deposited similar amounts of charge/energy.  \n\nIf the particles all had nearly identical initial properties, and went through the one point, then hopefully they are sufficiently different elsewhere to unravel them.  But, in any case, where there are two or more cells in one sensor, their signals can say something about the energy and direction through that cell.  That is a lot easier to use as a constraint than a huge number of completely independent \"hits\" with no direction or energy attached. \n\nIn your example, from Test_100 by the way, the cell 305-5 has 23 hits associated with it (I sorted by ch0 then ch1).   These are the hit numbers 14589, 14595, 14604, 14605, 14613, 14614, 72232, 82386, 109437, 110299, 110383, 110753, 110952, 111081, 112450, 113187, 114008, 114311, 115134, 115614, 115854, 116161, 116540.\n\nIn the hits file these correspond to many particles, the ones all starting with 6395 have very similar properties, and are the ones you are probably referring to.  These are 639511834298179000, 639511834298167000, 639511834298171000, 639511834298187000, 639511834298175000, 639511834298183000.\n\nIf you propagate these particles from the origin through the magnetic field, and they all have to go through that one cell, then that tells you a bit about their mass, type, charge and the time they go through each cell. \n\nThese other particles : 968286014512562000, 734097665658191000, 238692841851723000, 864702054851936000, 589971757343965000, 734106118170611000, 499903269489868000, 117094071364751000, 864698000402808000, 292762425642450000, 797156306778587000, 418853731921035000 all go through cell 305-5 as well.\n\nI presume the cell identifier is unique and not used in different sensors.  I get the impression that providing the cell data is an afterthought.\n\nI would start with the hits having more than one cell, work out the constraints on the trajectories, assign those particles first, then use another method for the other tracks.  My guess is the multi-cell hits are going to be more interesting any way.  It is nice the constraints are tighter.\n\nRichard",
      "votes": null
    },
    {
      "id": "355570",
      "postDate": "07/12/2018 02:23:44",
      "content": "<p>A cell does not contain a sub cell.  And this looks like close to reality actually.  There is a limit to the resolution you can have for hit locations. Two interesting kernels look at cell geometry: </p>\n\n<p><a href=\"https://www.kaggle.com/asalzburger/pixel-detector-cells\">https://www.kaggle.com/asalzburger/pixel-detector-cells</a></p>\n\n<p><a href=\"https://www.kaggle.com/jakubguzowski/calculate-angle-of-incidence-based-on-cells-csv\">https://www.kaggle.com/jakubguzowski/calculate-angle-of-incidence-based-on-cells-csv</a></p>",
      "rawMarkdown": "A cell does not contain a sub cell.  And this looks like close to reality actually.  There is a limit to the resolution you can have for hit locations. Two interesting kernels look at cell geometry: \n\nhttps://www.kaggle.com/asalzburger/pixel-detector-cells\n\nhttps://www.kaggle.com/jakubguzowski/calculate-angle-of-incidence-based-on-cells-csv",
      "votes": null
    },
    {
      "id": "355571",
      "postDate": "07/12/2018 02:26:46",
      "content": "<p>Hi, you write to me as if I was part of the organizer team.  I am not, and I do have my concerns about the setup of this competition, because of bugs in the simulator.   But for the rest it looks like they really tried to simulate something close to the forthcoming detectors at LHC.</p>",
      "rawMarkdown": "Hi, you write to me as if I was part of the organizer team.  I am not, and I do have my concerns about the setup of this competition, because of bugs in the simulator.   But for the rest it looks like they really tried to simulate something close to the forthcoming detectors at LHC.",
      "votes": null
    },
    {
      "id": "355591",
      "postDate": "07/12/2018 03:22:46",
      "content": "<p>CPMP,</p>\n\n<p>I am new to this group.  I was writing out of respect for your progress in the competition, and I was interested in your opinions.  I was also writing for anyone who might read it now, and in the future.</p>\n\n<p>That aside, I agree with you that the simulation is close, but there are many approximations and hidden assumptions.  In a real world problem, anyone would have complete documentation on the instruments, sensors, processing, models and assumptions.  This is a bit artificial.  And, as I said, I think it will only yield a reverse engineered model of the organizer's simulation.  And I doubt that the results are going to do more than confirm the organizers' understanding of things, and not likely reveal new insights into the real world.  That is not a criticism, just a fact - data from a model will only reveal the model.  We might go so far as to completely reengineer the model, but it is an exercise.  Perhaps to find new employees or associates or ideas.</p>\n\n<p>The organizers could produce truth data with timings, then the modelling would test the feasibility of a different approach to instrumenting the experiments.  A way to test the effectiveness and feasibility of fast algorithms, using timed hits, not asynchronous hits.  I think the results will be orders of magnitude faster for classification and calibration, which might well justify the cost of paying for better hardware and signal processing chains.  </p>\n\n<p>Pehaps the whole group might be interested in tackling some harder problems, related to this one.  You all are close to some good stuff, but aimed in the wrong direction slightly.  You all can expend a lot of energy solving this artificial and ambiguous problem, or specify and solve a problem with more realistic constraints, that will provide the field with useful guidance.  I would aim for 100% and accept 99.995% as a first cut.  The harder the problem, the more you learn.</p>\n\n<p>Richard</p>",
      "rawMarkdown": "CPMP,\n\nI am new to this group.  I was writing out of respect for your progress in the competition, and I was interested in your opinions.  I was also writing for anyone who might read it now, and in the future.\n\nThat aside, I agree with you that the simulation is close, but there are many approximations and hidden assumptions.  In a real world problem, anyone would have complete documentation on the instruments, sensors, processing, models and assumptions.  This is a bit artificial.  And, as I said, I think it will only yield a reverse engineered model of the organizer's simulation.  And I doubt that the results are going to do more than confirm the organizers' understanding of things, and not likely reveal new insights into the real world.  That is not a criticism, just a fact - data from a model will only reveal the model.  We might go so far as to completely reengineer the model, but it is an exercise.  Perhaps to find new employees or associates or ideas.\n\nThe organizers could produce truth data with timings, then the modelling would test the feasibility of a different approach to instrumenting the experiments.  A way to test the effectiveness and feasibility of fast algorithms, using timed hits, not asynchronous hits.  I think the results will be orders of magnitude faster for classification and calibration, which might well justify the cost of paying for better hardware and signal processing chains.  \n\nPehaps the whole group might be interested in tackling some harder problems, related to this one.  You all are close to some good stuff, but aimed in the wrong direction slightly.  You all can expend a lot of energy solving this artificial and ambiguous problem, or specify and solve a problem with more realistic constraints, that will provide the field with useful guidance.  I would aim for 100% and accept 99.995% as a first cut.  The harder the problem, the more you learn.\n\nRichard",
      "votes": null
    },
    {
      "id": "355727",
      "postDate": "07/12/2018 09:04:35",
      "content": "<p>cell channel id is local to a given volume_id,layer_id,module_id. This is indeed the case in your example. What has happened is that a particle has hit the module, interacted with it and created a spray of secondary particles. Each of these particles causes a different hit in the same cell. One can tell by going from hitid in cell file to hitid,particleid in truth file to particle in particles file and see that the origin coordinates vx,vy,vz of the particle 4.45153,-30.7167,-597.5 correspond to the coordinate of the hits. \nWhen several particles go through the same pixel (=cell) we consider each particle to be measured independently. This is a simplification, justified by 1) this allows to have a one to one correspondance between measured hit and  true hit 2) this is a rare occurence, a few per mille of the hits.</p>",
      "rawMarkdown": "cell channel id is local to a given volume_id,layer_id,module_id. This is indeed the case in your example. What has happened is that a particle has hit the module, interacted with it and created a spray of secondary particles. Each of these particles causes a different hit in the same cell. One can tell by going from hitid in cell file to hitid,particleid in truth file to particle in particles file and see that the origin coordinates vx,vy,vz of the particle 4.45153,-30.7167,-597.5 correspond to the coordinate of the hits. \nWhen several particles go through the same pixel (=cell) we consider each particle to be measured independently. This is a simplification, justified by 1) this allows to have a one to one correspondance between measured hit and  true hit 2) this is a rare occurence, a few per mille of the hits.",
      "votes": null
    },
    {
      "id": "355734",
      "postDate": "07/12/2018 09:16:13",
      "content": "<p>Thank you, David; much appreciated! </p>\n\n<p>So, in real life, we would get only one read-out from the cell?</p>\n\n<p>What would the value be? Sum of non-zero weights = .855? Their average = .285? Sum of all = 1.82? Average of all = .303? Or is the value of no &gt;ahem&lt; value?</p>\n\n<p>Edit: I see now. Apologies. I was lazy. I see in particles that vR and vZ are way out. I will exclude them!</p>",
      "rawMarkdown": "Thank you, David; much appreciated! \n\nSo, in real life, we would get only one read-out from the cell?\n\nWhat would the value be? Sum of non-zero weights = .855? Their average = .285? Sum of all = 1.82? Average of all = .303? Or is the value of no &gt;ahem&lt; value?\n\nEdit: I see now. Apologies. I was lazy. I see in particles that vR and vZ are way out. I will exclude them!",
      "votes": null
    },
    {
      "id": "355737",
      "postDate": "07/12/2018 09:17:39",
      "content": "<p>The response is linear to a good approximation, so this would be the sum.</p>",
      "rawMarkdown": "The response is linear to a good approximation, so this would be the sum.",
      "votes": null
    },
    {
      "id": "355739",
      "postDate": "07/12/2018 09:19:18",
      "content": "<p>See my answer above. Be careful that cell channel number is local, one has to use volume_id,layer_id,module_id to get a unique numbering. Also be careful that particleid is a i64 (you should rarely have multiple of 1000)</p>",
      "rawMarkdown": "See my answer above. Be careful that cell channel number is local, one has to use volume_id,layer_id,module_id to get a unique numbering. Also be careful that particleid is a i64 (you should rarely have multiple of 1000)",
      "votes": null
    },
    {
      "id": "355754",
      "postDate": "07/12/2018 09:38:12",
      "content": "<p>Hi Richard, \n(answer here for your two posts)\n - yes ps precision timing is possible but one has yet to demonstrate that 1) this is possible with the very fine granularity necessary to get the spatial resolution we need 2) with no or reasonable increase of the price tag\n - when designing a challenge one has to chose one specific problem (here pattern recognition/connecting the dots) for particle physics, more precisely for future LHC tracker, balance the simplification of the problem (so that it can be addressed to a large community with little or even no background in the domain) while maintaining sufficient complexity so that the best solutions to the challenge remain relevant to the domain\n- we expect/hope beyond this challenge to foster long term activity/collaboration/publications on pattern recognition/connecting the dots for particle physics</p>",
      "rawMarkdown": "Hi Richard, \n(answer here for your two posts)\n - yes ps precision timing is possible but one has yet to demonstrate that 1) this is possible with the very fine granularity necessary to get the spatial resolution we need 2) with no or reasonable increase of the price tag\n - when designing a challenge one has to chose one specific problem (here pattern recognition/connecting the dots) for particle physics, more precisely for future LHC tracker, balance the simplification of the problem (so that it can be addressed to a large community with little or even no background in the domain) while maintaining sufficient complexity so that the best solutions to the challenge remain relevant to the domain\n- we expect/hope beyond this challenge to foster long term activity/collaboration/publications on pattern recognition/connecting the dots for particle physics",
      "votes": null
    },
    {
      "id": "355764",
      "postDate": "07/12/2018 09:57:41",
      "content": "<p>Excellent! Thank you!</p>\n\n<p>While I have you on the line, is it something similar with event 1012, hits 17156, 9 &amp; 75? Cell 69/409.\nVolume 8, layer 2 &amp; module 92. vZ looks fairly normal at 1.82 - 6.15 for three particles?</p>",
      "rawMarkdown": "Excellent! Thank you!\n\nWhile I have you on the line, is it something similar with event 1012, hits 17156, 9 &amp; 75? Cell 69/409.\nVolume 8, layer 2 &amp; module 92. vZ looks fairly normal at 1.82 - 6.15 for three particles?",
      "votes": null
    },
    {
      "id": "355777",
      "postDate": "07/12/2018 10:46:59",
      "content": "<p>3 particles are indeed coming from origin, looks more like an accidental coincidence (than particle interaction as in your first example)</p>",
      "rawMarkdown": "3 particles are indeed coming from origin, looks more like an accidental coincidence (than particle interaction as in your first example)",
      "votes": null
    },
    {
      "id": "355797",
      "postDate": "07/12/2018 11:27:06",
      "content": "<p>Richard,</p>\n\n<p>I understand your point.  In many of the Kaggle competitions I entered some participants (including me from time to time) express doubts about how the competition is defined.  Sometimes the benefit to the sponsor is really unclear.</p>\n\n<p>However, I'd rather spend my energy on solving the problem as defined by the sponsor than trying to have the sponsor change the problem definition.  After all, they are the ones paying, and if they think their way is right, fine with me.  </p>",
      "rawMarkdown": "Richard,\n\nI understand your point.  In many of the Kaggle competitions I entered some participants (including me from time to time) express doubts about how the competition is defined.  Sometimes the benefit to the sponsor is really unclear.\n\nHowever, I'd rather spend my energy on solving the problem as defined by the sponsor than trying to have the sponsor change the problem definition.  After all, they are the ones paying, and if they think their way is right, fine with me.",
      "votes": null
    },
    {
      "id": "355832",
      "postDate": "07/12/2018 12:29:57",
      "content": "<p>David, </p>\n\n<p>Thank you for explaining; it is as I suspected.  I am trying to find ways to make collaborative websites more effective.  Even though this is a competition, collectively everyone is trying hard to solve this  problem, and some of them are interested in the problem(s) which caused you to pose it in the first place.  Here I think this whole group would gain by trying the timing problem.  Sort of an \"optional problem\".</p>\n\n<p>What I am looking at is the cost of <em>not</em> collecting all the data using an 85% filter algorithm.  If you significantly increase the collision rate, and data rate, yet continue to have such poor identification, then the cost to all the subsequent analysis is very high.  If you can make precise identifications, you will not miss, or misinterpret, rare events so often.  Alternatively, one can simply ask, what information would provide 99.9% accuracy in identifying tracks and particles and try to estimate the cost and benefits of doing so.  Or 97%, or 99.999%.</p>\n\n<p>If, for instance, the z axis were instrumented alone, then the hits will be separated (by time) into a fairly large number, k, of clusters.  For even n^2 problems, the sum of k \"n/k\" problems is k times smaller than one \"n\" problem.  I think this is actually one of those n^2.5 problems so the gain in speed is going to be about k^(1.5) = n^2.5/[k*(n/k)^2.5], where there are k \"n/k\" problems rather then one \"n\" problem.  Sorry it hard to write this here.</p>\n\n<p>For ten clusters that is a factor of 31.6.  If you can break x, y, z data into 10 groups each, k will be 1000, and the speedup about 1000^(1.5) or about 30,000.  That gives you a lot of room to run more agressive algorithms.  Also, the time gives you an ordered set, so it is faster still, since several events in a particular order do not have to be compared without order.</p>\n\n<p>LHC throws out about 99% of the data now, and expects 10x more?   How many events is it missing because of ambiguity or incomplete classification?</p>\n\n<p>10 Gsps is about 3 centimeter resolution.  If you break z, alone, into 100 identifiable groups by timing, then the speedup (depending on your classification algorithm efficiency) is roughly 100^1.5 =  1000 times faster.</p>\n\n<p>I was explaining this to my older neighbor yesterday.  If she first sorts her 1000 piece puzzle into 10 piles roughly by color and texture, then solves each individual piles separately, she can do it about 10 times faster, depending on her efficiency with processing the smaller piles.  I kept the math simple.  She is trading space for speed, and having to implement two strategies - sorting, and then matching and fitting.  Manual puzzle identification and sorting has a lot of fixed costs.  I am trying to find ways to reduce that in signal processing chains.  </p>\n\n<p>I spent a lot of time this year looking at \"compile to silicon\" methods.  You are proposing to test the speed of these contest algorithms on an ordinary PC.  Have you tried the same on an FPGA?  The usual speedup is about 100, and they are cheaper, smaller and use less energy. That is why the cloud guys are all moving that way.  I think they have room for improvement since I usually get 300 and sometime more.</p>\n\n<p>I apologize this is so crudely described here.  It is hard to write general rule for comparing algorithms.</p>\n\n<p>Richard</p>",
      "rawMarkdown": "David, \n\nThank you for explaining; it is as I suspected.  I am trying to find ways to make collaborative websites more effective.  Even though this is a competition, collectively everyone is trying hard to solve this  problem, and some of them are interested in the problem(s) which caused you to pose it in the first place.  Here I think this whole group would gain by trying the timing problem.  Sort of an \"optional problem\".\n\nWhat I am looking at is the cost of *not* collecting all the data using an 85% filter algorithm.  If you significantly increase the collision rate, and data rate, yet continue to have such poor identification, then the cost to all the subsequent analysis is very high.  If you can make precise identifications, you will not miss, or misinterpret, rare events so often.  Alternatively, one can simply ask, what information would provide 99.9% accuracy in identifying tracks and particles and try to estimate the cost and benefits of doing so.  Or 97%, or 99.999%.\n\nIf, for instance, the z axis were instrumented alone, then the hits will be separated (by time) into a fairly large number, k, of clusters.  For even n^2 problems, the sum of k \"n/k\" problems is k times smaller than one \"n\" problem.  I think this is actually one of those n^2.5 problems so the gain in speed is going to be about k^(1.5) = n^2.5/[k*(n/k)^2.5], where there are k \"n/k\" problems rather then one \"n\" problem.  Sorry it hard to write this here.\n\nFor ten clusters that is a factor of 31.6.  If you can break x, y, z data into 10 groups each, k will be 1000, and the speedup about 1000^(1.5) or about 30,000.  That gives you a lot of room to run more agressive algorithms.  Also, the time gives you an ordered set, so it is faster still, since several events in a particular order do not have to be compared without order.\n\nLHC throws out about 99% of the data now, and expects 10x more?   How many events is it missing because of ambiguity or incomplete classification?\n\n10 Gsps is about 3 centimeter resolution.  If you break z, alone, into 100 identifiable groups by timing, then the speedup (depending on your classification algorithm efficiency) is roughly 100^1.5 =  1000 times faster.\n\nI was explaining this to my older neighbor yesterday.  If she first sorts her 1000 piece puzzle into 10 piles roughly by color and texture, then solves each individual piles separately, she can do it about 10 times faster, depending on her efficiency with processing the smaller piles.  I kept the math simple.  She is trading space for speed, and having to implement two strategies - sorting, and then matching and fitting.  Manual puzzle identification and sorting has a lot of fixed costs.  I am trying to find ways to reduce that in signal processing chains.  \n\nI spent a lot of time this year looking at \"compile to silicon\" methods.  You are proposing to test the speed of these contest algorithms on an ordinary PC.  Have you tried the same on an FPGA?  The usual speedup is about 100, and they are cheaper, smaller and use less energy. That is why the cloud guys are all moving that way.  I think they have room for improvement since I usually get 300 and sometime more.\n\nI apologize this is so crudely described here.  It is hard to write general rule for comparing algorithms.\n\nRichard",
      "votes": null
    },
    {
      "id": "355841",
      "postDate": "07/12/2018 12:43:12",
      "content": "<p>Richard, I think many of the points you make are valid and I guess such points are being discussed in the research community. However, they would be beyond the scope of a kaggle competition, I think.</p>\n\n<p>In my opinion, the organizers did a good job of simplifying just enough so that the problem becomes approachable for non-specialists while still remaining interesting. The barrier of entry is already higher than in the typical kaggle competition, I think. (By that I mean the amount of work needed until machine learning techniques can be applied, not the difficulty of the challenge, which is very high also in the other competitions, I assume.)</p>\n\n<p>Let's not forget that the organizers made quite a big and risky investment of their time for setting up this competition <em>in addition to</em>, not as a replacement for, their work in the usual scientific context.</p>",
      "rawMarkdown": "Richard, I think many of the points you make are valid and I guess such points are being discussed in the research community. However, they would be beyond the scope of a kaggle competition, I think.\n\nIn my opinion, the organizers did a good job of simplifying just enough so that the problem becomes approachable for non-specialists while still remaining interesting. The barrier of entry is already higher than in the typical kaggle competition, I think. (By that I mean the amount of work needed until machine learning techniques can be applied, not the difficulty of the challenge, which is very high also in the other competitions, I assume.)\n\nLet's not forget that the organizers made quite a big and risky investment of their time for setting up this competition *in addition to*, not as a replacement for, their work in the usual scientific context.",
      "votes": null
    },
    {
      "id": "355866",
      "postDate": "07/12/2018 13:19:32",
      "content": "<ul>\n<li>there is no point discussing the merit of timing information, the technology to have it while preserving spatial information and within budget is not there</li>\n<li>we are actually throwing away ~39999/40000 of the data (or to be more specific, of the proton bunch collisions) . This is is a complex issue unrelated to this challenge.</li>\n<li>we are using already FPGA, and custom ASIC even in some cases. The point of  running on ordinary PC core is that:\n<ul><li>it is the bulk of our resources</li>\n<li>it is (much) more portable</li>\n<li>more important, it does not require rare competences</li></ul></li>\n</ul>",
      "rawMarkdown": "there is no point discussing the merit of timing information, the technology to have it while preserving spatial information and within budget is not there\n- we are actually throwing away ~39999/40000 of the data (or to be more specific, of the proton bunch collisions) . This is is a complex issue unrelated to this challenge.\n- we are using already FPGA, and custom ASIC even in some cases. The point of  running on ordinary PC core is that:\n  - it is the bulk of our resources\n  - it is (much) more portable\n  - more important, it does not require rare competences",
      "votes": null
    },
    {
      "id": "355889",
      "postDate": "07/12/2018 13:44:42",
      "content": "<p>Edwin,</p>\n\n<p>Yes, I have a great respect for the organizers, and appreciate the time they invested.  But, too, a competition of this sort can sometimes absorb scarce resources that would otherwise be working on critical problems in the real world.  That is why I do not like toy problems, and try to find ways to structure cooperative groups who can apply non-specialists to highly specialized problems in the real world, with a good chance of solving them.</p>\n\n<p>The \"non-specialist\" part of this could be improved a bit, if you allow that part of what happens in a Kaggle kind of environment is that people get introduced to new areas of research.   There are a whole range of \"fast learning\" methods that can be applied.  That means still more expense in setting up these competition/training grounds/collaborative research groups.  But, I hope that Kaggle is constantly monitoring the methods for the whole, and then investing that back into improving the process.</p>\n\n<p>If the organizers' investment is high, then solving the real problem can offer still greater rewards.  Say you invest three man years to set up this thing, but it saves $10 billion in the short term, and millions of man years eventually when the new information is absorbed by society.  That is a good trade.  Society might benefit that much, and not specific individuals.  I judge by society as a whole, and not so much the individuals.  Realistically, if the group cuts the cost of an LHC upgrade by a factor of 10, I expect the organizers and participants might be personally and professionally rewarded. </p>\n\n<p>I apologize if I am a bit vague here.  I am trying to restructure the Internet as a whole, including things like collaborative and topic sites.  I have to use methods from hundreds of disciplines, and they do not all fit neatly together.  I am new here and still talking to strangers.  I am reviewing CERN as a whole, and opendata.cern.ch particularly.  I came here because I think this group has a high probability of changing CERN, with just relatively tiny adjustments.  Doing it with non-specialists offers a positive outlook for so many of the outstanding problems in the world today.  I know how to take any problem and structure it for non-specialists.  And non-specialists can improve their methods to become even more effective at solving any problem.  Machine learning pales compared to getting human groups to learn and solve problems faster.  Perhaps they should make a new discipline called HL, human learning.   Or HML for human machine learning.  When a person with computers can do in days what it takes others years to do, that changes the whole game.  And it is a lot of fun.</p>\n\n<p>Richard   </p>",
      "rawMarkdown": "Edwin,\n\nYes, I have a great respect for the organizers, and appreciate the time they invested.  But, too, a competition of this sort can sometimes absorb scarce resources that would otherwise be working on critical problems in the real world.  That is why I do not like toy problems, and try to find ways to structure cooperative groups who can apply non-specialists to highly specialized problems in the real world, with a good chance of solving them.\n\nThe \"non-specialist\" part of this could be improved a bit, if you allow that part of what happens in a Kaggle kind of environment is that people get introduced to new areas of research.   There are a whole range of \"fast learning\" methods that can be applied.  That means still more expense in setting up these competition/training grounds/collaborative research groups.  But, I hope that Kaggle is constantly monitoring the methods for the whole, and then investing that back into improving the process.\n\nIf the organizers' investment is high, then solving the real problem can offer still greater rewards.  Say you invest three man years to set up this thing, but it saves $10 billion in the short term, and millions of man years eventually when the new information is absorbed by society.  That is a good trade.  Society might benefit that much, and not specific individuals.  I judge by society as a whole, and not so much the individuals.  Realistically, if the group cuts the cost of an LHC upgrade by a factor of 10, I expect the organizers and participants might be personally and professionally rewarded. \n\nI apologize if I am a bit vague here.  I am trying to restructure the Internet as a whole, including things like collaborative and topic sites.  I have to use methods from hundreds of disciplines, and they do not all fit neatly together.  I am new here and still talking to strangers.  I am reviewing CERN as a whole, and opendata.cern.ch particularly.  I came here because I think this group has a high probability of changing CERN, with just relatively tiny adjustments.  Doing it with non-specialists offers a positive outlook for so many of the outstanding problems in the world today.  I know how to take any problem and structure it for non-specialists.  And non-specialists can improve their methods to become even more effective at solving any problem.  Machine learning pales compared to getting human groups to learn and solve problems faster.  Perhaps they should make a new discipline called HL, human learning.   Or HML for human machine learning.  When a person with computers can do in days what it takes others years to do, that changes the whole game.  And it is a lot of fun.\n\nRichard",
      "votes": null
    },
    {
      "id": "355897",
      "postDate": "07/12/2018 13:54:56",
      "content": "<p>David,</p>\n\n<p>Thanks for all your hard work!  My general comments, just getting started, are not a criticism.  More along the lines of finding what others are doing and thinking.  You have focused onto a good problem here.  Maybe at some point people can use real data to test algorithms in the next phase.   Then a hardware competition as an adjunct (or side bet) could benefit LHC, and you can still get done what you want to do.</p>\n\n<p>I will not mention this again.  I was thinking of offering a $1000 prize to anyone who can implement the fastest, low cost,  good algorithm, regardless of hardware.  That will help LHC, and is very timely.  Your \"non specialists\" are actually deep in a wide range of competencies.  And many of these things are actually easy to pick up with a short introduction.</p>\n\n<p>Thanks again.  I have something else to do now.</p>\n\n<p>Richard</p>",
      "rawMarkdown": "David,\n\nThanks for all your hard work!  My general comments, just getting started, are not a criticism.  More along the lines of finding what others are doing and thinking.  You have focused onto a good problem here.  Maybe at some point people can use real data to test algorithms in the next phase.   Then a hardware competition as an adjunct (or side bet) could benefit LHC, and you can still get done what you want to do.\n\nI will not mention this again.  I was thinking of offering a $1000 prize to anyone who can implement the fastest, low cost,  good algorithm, regardless of hardware.  That will help LHC, and is very timely.  Your \"non specialists\" are actually deep in a wide range of competencies.  And many of these things are actually easy to pick up with a short introduction.\n\nThanks again.  I have something else to do now.\n\nRichard",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 355461,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "07/11/2018 18:20:52",
      "content": "<p>cell geometry is provided.</p>",
      "votes": null,
      "replies": [
        {
          "id": 355532,
          "author_name": "glimmung",
          "author_url": "",
          "post_date": "07/11/2018 23:06:53",
          "content": "<p>It looks as if the simulation went \"We have all these particleIds, what cells did they hit and what value would it yield?\" So, in eventId 1000, hitIds 14589, 95, 604, 5, 13 &amp; 4 (with six particleIds) have the same cellId (305/5) and similar values (.25-.37). Does the cell contain sub cells?</p>\n\n<p>I still have no answer from the organizers on this one...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355570,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "07/12/2018 02:23:44",
          "content": "<p>A cell does not contain a sub cell.  And this looks like close to reality actually.  There is a limit to the resolution you can have for hit locations. Two interesting kernels look at cell geometry: </p>\n\n<p><a href=\"https://www.kaggle.com/asalzburger/pixel-detector-cells\">https://www.kaggle.com/asalzburger/pixel-detector-cells</a></p>\n\n<p><a href=\"https://www.kaggle.com/jakubguzowski/calculate-angle-of-incidence-based-on-cells-csv\">https://www.kaggle.com/jakubguzowski/calculate-angle-of-incidence-based-on-cells-csv</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355727,
          "author_name": "droussea",
          "author_url": "",
          "post_date": "07/12/2018 09:04:35",
          "content": "<p>cell channel id is local to a given volume_id,layer_id,module_id. This is indeed the case in your example. What has happened is that a particle has hit the module, interacted with it and created a spray of secondary particles. Each of these particles causes a different hit in the same cell. One can tell by going from hitid in cell file to hitid,particleid in truth file to particle in particles file and see that the origin coordinates vx,vy,vz of the particle 4.45153,-30.7167,-597.5 correspond to the coordinate of the hits. \nWhen several particles go through the same pixel (=cell) we consider each particle to be measured independently. This is a simplification, justified by 1) this allows to have a one to one correspondance between measured hit and  true hit 2) this is a rare occurence, a few per mille of the hits.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355734,
          "author_name": "glimmung",
          "author_url": "",
          "post_date": "07/12/2018 09:16:13",
          "content": "<p>Thank you, David; much appreciated! </p>\n\n<p>So, in real life, we would get only one read-out from the cell?</p>\n\n<p>What would the value be? Sum of non-zero weights = .855? Their average = .285? Sum of all = 1.82? Average of all = .303? Or is the value of no &gt;ahem&lt; value?</p>\n\n<p>Edit: I see now. Apologies. I was lazy. I see in particles that vR and vZ are way out. I will exclude them!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355737,
          "author_name": "droussea",
          "author_url": "",
          "post_date": "07/12/2018 09:17:39",
          "content": "<p>The response is linear to a good approximation, so this would be the sum.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355764,
          "author_name": "glimmung",
          "author_url": "",
          "post_date": "07/12/2018 09:57:41",
          "content": "<p>Excellent! Thank you!</p>\n\n<p>While I have you on the line, is it something similar with event 1012, hits 17156, 9 &amp; 75? Cell 69/409.\nVolume 8, layer 2 &amp; module 92. vZ looks fairly normal at 1.82 - 6.15 for three particles?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355777,
          "author_name": "droussea",
          "author_url": "",
          "post_date": "07/12/2018 10:46:59",
          "content": "<p>3 particles are indeed coming from origin, looks more like an accidental coincidence (than particle interaction as in your first example)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 355509,
      "author_name": "richardcollins",
      "author_url": "",
      "post_date": "07/11/2018 20:59:23",
      "content": "<p>Dear CPMP,</p>\n\n<p>Yes.  Sorry, I am very tired.  I was writing quickly to see if anyone is listening.  I was thinking more along the lines of having the test provide the cell locations directly, rather than making everyone look them up.  Collapsing into \"hits\" is something you did to shorten the test files, but the data ultimately is available, or could be available, at the cell level.   I would provide one \"cell hits file\" that contains all the information needed for cell-level resolution of the hits.  I do not think you are testing people's ability to join data files and look up things.  Rather it is to test new ways to get to the physics?  I took liberties writing quickly to say \"better cell data\", but I was really thinking of getting all the information into an array of cell objects with ALL the necessary information.  </p>\n\n<p>It becomes a partnership of the test data providers continually improving the readability and usefulness of the test data, and the others trying to find the best and fastest algorithms for classifying into tracks, particles, particle events and vacuum processes (I lump all the quark gluon and many electromagnetic events into vacuum properties.  That is my personal view of things.)</p>\n\n<p>What did you think about running the simulation, so as to provide timing data?  Assuming picosecond resolution, for instance?  Or a range of time resolutions on hits?  I think the people gathered here can solve that easily, if they have already tried to solve your current problem, and the results might better inform decision-making.  Your test fairly faithfully represents the present situation and its data limitations.  The problem I think you ought to solve is how to change the detectors and signal processing to capture enough information to resolve the particles and events, and not throw away so much good information.  So I am suggesting you run the time resolved simulation and let these people look at it, while everyone has these things in mind.  It won't affect the contest much, but might provide some new insights toward better algorithms.  More, I think it will move everyone toward a solution to reduce the cost of upgrades to equipment and get to where you can capture all the data.</p>\n\n<p>I think the memory for the event data is too small, relative to the full size of the data stream.  This is a vast simplification, but it seems you are trying to represent a many gigabyte event with a few megabytes of data.  No matter how much you want to resolve it, there is not enough to do the filtering precisely and efficiently.  So you are pretty much guaranteed to lose.  If you compare the difficulty of capturing and filter a gigabyte per event, to the huge cost of unraveling poorly documented hit data after the event, I think you might see what I am aiming at.</p>\n\n<p>Richard Collins</p>",
      "votes": null,
      "replies": [
        {
          "id": 355571,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "07/12/2018 02:26:46",
          "content": "<p>Hi, you write to me as if I was part of the organizer team.  I am not, and I do have my concerns about the setup of this competition, because of bugs in the simulator.   But for the rest it looks like they really tried to simulate something close to the forthcoming detectors at LHC.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355591,
          "author_name": "richardcollins",
          "author_url": "",
          "post_date": "07/12/2018 03:22:46",
          "content": "<p>CPMP,</p>\n\n<p>I am new to this group.  I was writing out of respect for your progress in the competition, and I was interested in your opinions.  I was also writing for anyone who might read it now, and in the future.</p>\n\n<p>That aside, I agree with you that the simulation is close, but there are many approximations and hidden assumptions.  In a real world problem, anyone would have complete documentation on the instruments, sensors, processing, models and assumptions.  This is a bit artificial.  And, as I said, I think it will only yield a reverse engineered model of the organizer's simulation.  And I doubt that the results are going to do more than confirm the organizers' understanding of things, and not likely reveal new insights into the real world.  That is not a criticism, just a fact - data from a model will only reveal the model.  We might go so far as to completely reengineer the model, but it is an exercise.  Perhaps to find new employees or associates or ideas.</p>\n\n<p>The organizers could produce truth data with timings, then the modelling would test the feasibility of a different approach to instrumenting the experiments.  A way to test the effectiveness and feasibility of fast algorithms, using timed hits, not asynchronous hits.  I think the results will be orders of magnitude faster for classification and calibration, which might well justify the cost of paying for better hardware and signal processing chains.  </p>\n\n<p>Pehaps the whole group might be interested in tackling some harder problems, related to this one.  You all are close to some good stuff, but aimed in the wrong direction slightly.  You all can expend a lot of energy solving this artificial and ambiguous problem, or specify and solve a problem with more realistic constraints, that will provide the field with useful guidance.  I would aim for 100% and accept 99.995% as a first cut.  The harder the problem, the more you learn.</p>\n\n<p>Richard</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355797,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "07/12/2018 11:27:06",
          "content": "<p>Richard,</p>\n\n<p>I understand your point.  In many of the Kaggle competitions I entered some participants (including me from time to time) express doubts about how the competition is defined.  Sometimes the benefit to the sponsor is really unclear.</p>\n\n<p>However, I'd rather spend my energy on solving the problem as defined by the sponsor than trying to have the sponsor change the problem definition.  After all, they are the ones paying, and if they think their way is right, fine with me.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 355554,
      "author_name": "richardcollins",
      "author_url": "",
      "post_date": "07/12/2018 00:51:37",
      "content": "<p>Dear Glimmung,</p>\n\n<p>I expect that they ran the simulation, it generated those six particles, the particles went through cell 305-5, and all deposited similar amounts of charge/energy.  </p>\n\n<p>If the particles all had nearly identical initial properties, and went through the one point, then hopefully they are sufficiently different elsewhere to unravel them.  But, in any case, where there are two or more cells in one sensor, their signals can say something about the energy and direction through that cell.  That is a lot easier to use as a constraint than a huge number of completely independent \"hits\" with no direction or energy attached. </p>\n\n<p>In your example, from Test_100 by the way, the cell 305-5 has 23 hits associated with it (I sorted by ch0 then ch1).   These are the hit numbers 14589, 14595, 14604, 14605, 14613, 14614, 72232, 82386, 109437, 110299, 110383, 110753, 110952, 111081, 112450, 113187, 114008, 114311, 115134, 115614, 115854, 116161, 116540.</p>\n\n<p>In the hits file these correspond to many particles, the ones all starting with 6395 have very similar properties, and are the ones you are probably referring to.  These are 639511834298179000, 639511834298167000, 639511834298171000, 639511834298187000, 639511834298175000, 639511834298183000.</p>\n\n<p>If you propagate these particles from the origin through the magnetic field, and they all have to go through that one cell, then that tells you a bit about their mass, type, charge and the time they go through each cell. </p>\n\n<p>These other particles : 968286014512562000, 734097665658191000, 238692841851723000, 864702054851936000, 589971757343965000, 734106118170611000, 499903269489868000, 117094071364751000, 864698000402808000, 292762425642450000, 797156306778587000, 418853731921035000 all go through cell 305-5 as well.</p>\n\n<p>I presume the cell identifier is unique and not used in different sensors.  I get the impression that providing the cell data is an afterthought.</p>\n\n<p>I would start with the hits having more than one cell, work out the constraints on the trajectories, assign those particles first, then use another method for the other tracks.  My guess is the multi-cell hits are going to be more interesting any way.  It is nice the constraints are tighter.</p>\n\n<p>Richard</p>",
      "votes": null,
      "replies": [
        {
          "id": 355739,
          "author_name": "droussea",
          "author_url": "",
          "post_date": "07/12/2018 09:19:18",
          "content": "<p>See my answer above. Be careful that cell channel number is local, one has to use volume_id,layer_id,module_id to get a unique numbering. Also be careful that particleid is a i64 (you should rarely have multiple of 1000)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 355754,
      "author_name": "droussea",
      "author_url": "",
      "post_date": "07/12/2018 09:38:12",
      "content": "<p>Hi Richard, \n(answer here for your two posts)\n - yes ps precision timing is possible but one has yet to demonstrate that 1) this is possible with the very fine granularity necessary to get the spatial resolution we need 2) with no or reasonable increase of the price tag\n - when designing a challenge one has to chose one specific problem (here pattern recognition/connecting the dots) for particle physics, more precisely for future LHC tracker, balance the simplification of the problem (so that it can be addressed to a large community with little or even no background in the domain) while maintaining sufficient complexity so that the best solutions to the challenge remain relevant to the domain\n- we expect/hope beyond this challenge to foster long term activity/collaboration/publications on pattern recognition/connecting the dots for particle physics</p>",
      "votes": null,
      "replies": [
        {
          "id": 355832,
          "author_name": "richardcollins",
          "author_url": "",
          "post_date": "07/12/2018 12:29:57",
          "content": "<p>David, </p>\n\n<p>Thank you for explaining; it is as I suspected.  I am trying to find ways to make collaborative websites more effective.  Even though this is a competition, collectively everyone is trying hard to solve this  problem, and some of them are interested in the problem(s) which caused you to pose it in the first place.  Here I think this whole group would gain by trying the timing problem.  Sort of an \"optional problem\".</p>\n\n<p>What I am looking at is the cost of <em>not</em> collecting all the data using an 85% filter algorithm.  If you significantly increase the collision rate, and data rate, yet continue to have such poor identification, then the cost to all the subsequent analysis is very high.  If you can make precise identifications, you will not miss, or misinterpret, rare events so often.  Alternatively, one can simply ask, what information would provide 99.9% accuracy in identifying tracks and particles and try to estimate the cost and benefits of doing so.  Or 97%, or 99.999%.</p>\n\n<p>If, for instance, the z axis were instrumented alone, then the hits will be separated (by time) into a fairly large number, k, of clusters.  For even n^2 problems, the sum of k \"n/k\" problems is k times smaller than one \"n\" problem.  I think this is actually one of those n^2.5 problems so the gain in speed is going to be about k^(1.5) = n^2.5/[k*(n/k)^2.5], where there are k \"n/k\" problems rather then one \"n\" problem.  Sorry it hard to write this here.</p>\n\n<p>For ten clusters that is a factor of 31.6.  If you can break x, y, z data into 10 groups each, k will be 1000, and the speedup about 1000^(1.5) or about 30,000.  That gives you a lot of room to run more agressive algorithms.  Also, the time gives you an ordered set, so it is faster still, since several events in a particular order do not have to be compared without order.</p>\n\n<p>LHC throws out about 99% of the data now, and expects 10x more?   How many events is it missing because of ambiguity or incomplete classification?</p>\n\n<p>10 Gsps is about 3 centimeter resolution.  If you break z, alone, into 100 identifiable groups by timing, then the speedup (depending on your classification algorithm efficiency) is roughly 100^1.5 =  1000 times faster.</p>\n\n<p>I was explaining this to my older neighbor yesterday.  If she first sorts her 1000 piece puzzle into 10 piles roughly by color and texture, then solves each individual piles separately, she can do it about 10 times faster, depending on her efficiency with processing the smaller piles.  I kept the math simple.  She is trading space for speed, and having to implement two strategies - sorting, and then matching and fitting.  Manual puzzle identification and sorting has a lot of fixed costs.  I am trying to find ways to reduce that in signal processing chains.  </p>\n\n<p>I spent a lot of time this year looking at \"compile to silicon\" methods.  You are proposing to test the speed of these contest algorithms on an ordinary PC.  Have you tried the same on an FPGA?  The usual speedup is about 100, and they are cheaper, smaller and use less energy. That is why the cloud guys are all moving that way.  I think they have room for improvement since I usually get 300 and sometime more.</p>\n\n<p>I apologize this is so crudely described here.  It is hard to write general rule for comparing algorithms.</p>\n\n<p>Richard</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355866,
          "author_name": "droussea",
          "author_url": "",
          "post_date": "07/12/2018 13:19:32",
          "content": "<ul>\n<li>there is no point discussing the merit of timing information, the technology to have it while preserving spatial information and within budget is not there</li>\n<li>we are actually throwing away ~39999/40000 of the data (or to be more specific, of the proton bunch collisions) . This is is a complex issue unrelated to this challenge.</li>\n<li>we are using already FPGA, and custom ASIC even in some cases. The point of  running on ordinary PC core is that:\n<ul><li>it is the bulk of our resources</li>\n<li>it is (much) more portable</li>\n<li>more important, it does not require rare competences</li></ul></li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 355897,
          "author_name": "richardcollins",
          "author_url": "",
          "post_date": "07/12/2018 13:54:56",
          "content": "<p>David,</p>\n\n<p>Thanks for all your hard work!  My general comments, just getting started, are not a criticism.  More along the lines of finding what others are doing and thinking.  You have focused onto a good problem here.  Maybe at some point people can use real data to test algorithms in the next phase.   Then a hardware competition as an adjunct (or side bet) could benefit LHC, and you can still get done what you want to do.</p>\n\n<p>I will not mention this again.  I was thinking of offering a $1000 prize to anyone who can implement the fastest, low cost,  good algorithm, regardless of hardware.  That will help LHC, and is very timely.  Your \"non specialists\" are actually deep in a wide range of competencies.  And many of these things are actually easy to pick up with a short introduction.</p>\n\n<p>Thanks again.  I have something else to do now.</p>\n\n<p>Richard</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 355841,
      "author_name": "edwinst",
      "author_url": "",
      "post_date": "07/12/2018 12:43:12",
      "content": "<p>Richard, I think many of the points you make are valid and I guess such points are being discussed in the research community. However, they would be beyond the scope of a kaggle competition, I think.</p>\n\n<p>In my opinion, the organizers did a good job of simplifying just enough so that the problem becomes approachable for non-specialists while still remaining interesting. The barrier of entry is already higher than in the typical kaggle competition, I think. (By that I mean the amount of work needed until machine learning techniques can be applied, not the difficulty of the challenge, which is very high also in the other competitions, I assume.)</p>\n\n<p>Let's not forget that the organizers made quite a big and risky investment of their time for setting up this competition <em>in addition to</em>, not as a replacement for, their work in the usual scientific context.</p>",
      "votes": null,
      "replies": [
        {
          "id": 355889,
          "author_name": "richardcollins",
          "author_url": "",
          "post_date": "07/12/2018 13:44:42",
          "content": "<p>Edwin,</p>\n\n<p>Yes, I have a great respect for the organizers, and appreciate the time they invested.  But, too, a competition of this sort can sometimes absorb scarce resources that would otherwise be working on critical problems in the real world.  That is why I do not like toy problems, and try to find ways to structure cooperative groups who can apply non-specialists to highly specialized problems in the real world, with a good chance of solving them.</p>\n\n<p>The \"non-specialist\" part of this could be improved a bit, if you allow that part of what happens in a Kaggle kind of environment is that people get introduced to new areas of research.   There are a whole range of \"fast learning\" methods that can be applied.  That means still more expense in setting up these competition/training grounds/collaborative research groups.  But, I hope that Kaggle is constantly monitoring the methods for the whole, and then investing that back into improving the process.</p>\n\n<p>If the organizers' investment is high, then solving the real problem can offer still greater rewards.  Say you invest three man years to set up this thing, but it saves $10 billion in the short term, and millions of man years eventually when the new information is absorbed by society.  That is a good trade.  Society might benefit that much, and not specific individuals.  I judge by society as a whole, and not so much the individuals.  Realistically, if the group cuts the cost of an LHC upgrade by a factor of 10, I expect the organizers and participants might be personally and professionally rewarded. </p>\n\n<p>I apologize if I am a bit vague here.  I am trying to restructure the Internet as a whole, including things like collaborative and topic sites.  I have to use methods from hundreds of disciplines, and they do not all fit neatly together.  I am new here and still talking to strangers.  I am reviewing CERN as a whole, and opendata.cern.ch particularly.  I came here because I think this group has a high probability of changing CERN, with just relatively tiny adjustments.  Doing it with non-specialists offers a positive outlook for so many of the outstanding problems in the world today.  I know how to take any problem and structure it for non-specialists.  And non-specialists can improve their methods to become even more effective at solving any problem.  Machine learning pales compared to getting human groups to learn and solve problems faster.  Perhaps they should make a new discipline called HL, human learning.   Or HML for human machine learning.  When a person with computers can do in days what it takes others years to do, that changes the whole game.  And it is a lot of fun.</p>\n\n<p>Richard   </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "355450": "Dear TrackML Team:\n\nI just joined this group recently.  My first question to myself was, \"Why are they not working together to design better detectors and signal processors?\"  and \"If they invert this model all they will do is reverse engineer the simulator and not the real thing.  It is a good game, but not the real thing.  The real problem is that the information coming from the experiments, as they are using it now, is ambiguous.  That probably requires some new signal processing hardware, better data management and better algorithms.\"\n\nWith this amount of data, what you are getting is about what you can get.  But if you want nearly perfect tracking, you will need to get some more very specific clues.\n\n**The first clue is timing**.  Someone here said sub-nanosecond timing is \"impossible\".  But I know that the ADC technology has progressed to 10 Gsps (giga samples per second) with good resolution.  I am not selling these chips, I have been tracking their progress for application to high sampling rate gravimeter arrays.  There are plenty of technologies to time stamp some layers of the z,x and y axes to help resolve sequences of events.  The more ways you can break down the timing of the hits from a collision, the faster your algorithms.  Investing in getting timing data reduces the time cost of the trajectory and particle identification.   I would aim to be able to process and classify all the data and not lose any.\n\nRecommendation: Run the \"truth\" simulation and give the TrackML team the timings - as though you could do that in the real machine.  I think you are going to see this problem is almost trivial to unravel.  Then go back and figure how to get that time information ( or some proxies ) from the real equipment.  Fix the real machines. Stop playing a game you cannot win.\n\nhttp://www.analog.com/en/products/analog-to-digital-converters/standard-adc/high-speed-ad-10msps/ad9213.html\n\n**The second clue is cell hit geometry**.  For many of the hits, the cell geometry, regardless of time sequence, gives tight constraints on the possible trajectories.  The particle is going one way or another along the line of cells hit.  \n\nRecommendation:  Explicitly provide the cell locations and geometries.  A more true test here, is full cell information, better organized and documented.  Run the simulation and break the hits into cell level data.  Test using cell data, not averaged hit data.\n\nUntil you can modify the sensors to get timing, you can probably get much closer by using the cell hit geometric constraints.  There might be many possible cell trajectories for a hit based on cell data, but this space of trajectories, when compared to the whole space on independent hits is much smaller.  The hit turns from a point into a beam with a fairly narrow range of likely directions. When stitching together trajectories from hits, most of the hits have multiple cell points and therefore beams.\n\n**Changing from a game to a real problem:**\n\nI am not criticizing.  Personally, I like playing with lots of numbers and inverting models.  But I joined because I thought I might see some real data.  I would like to solve for the magnetic field structure and temporal variations.  I would like to see a massive dataset of low energy electron positron events, and other particle-antiparticle events.  I would like to look at the entire dataset.  I was dismayed that most of the data is being ignored (not used even when it is gathered), thrown away without deep examination, and not shared particularly well - (For the world at large, it is the more common data that is more important than things the specialists care about).\n\nIf this many people are going to expend time to solve this problem, it is not for the money or the fame.  Would you all consider changing direction and solve the real problem?  Kaggle is throwing money at lots of problems, for their own purposes;  hoping to get some great solutions, to further advertise themselves.  This is not a criticism.  I review these sorts of things, and I am simply telling you that most collaborative websites are self-serving.  It is a phase they all go through until they understand their true goals.  I have a pretty good idea what they should be doing, but one thing at a time.\n\nYou are probably not investing much into this, since it is such a small amount of money and fame at stake.  But if you (collectively) solve a real problem, the stakes are much higher.\n\nChange this from a puzzle for a prize, to a real effort to drastically reduce the cost of the next generation of accelerator upgrades.  Rather than thousands go after billions.\n\nI apologize for throwing this in your laps on my first hello, but I have a lot on my plate and not much time left.\n\nI am a senior mathematical statistician.  I spent my life on global problems - international development, global economic and social modeling, famine early warning, clean air, alternative fuels, global climate change, nonprofits, industry modeling, technology modeling, business intelligence, Y2K, and for the last 20 years the relation between society and the Internet.  My personal interests are dna genealogy, 3D technologies, and gravimeter imaging arrays.   \n\nSincere regards,\nRichard Collins, The Internet Foundation",
    "355461": "cell geometry is provided.",
    "355509": "Dear CPMP,\n\nYes.  Sorry, I am very tired.  I was writing quickly to see if anyone is listening.  I was thinking more along the lines of having the test provide the cell locations directly, rather than making everyone look them up.  Collapsing into \"hits\" is something you did to shorten the test files, but the data ultimately is available, or could be available, at the cell level.   I would provide one \"cell hits file\" that contains all the information needed for cell-level resolution of the hits.  I do not think you are testing people's ability to join data files and look up things.  Rather it is to test new ways to get to the physics?  I took liberties writing quickly to say \"better cell data\", but I was really thinking of getting all the information into an array of cell objects with ALL the necessary information.  \n\nIt becomes a partnership of the test data providers continually improving the readability and usefulness of the test data, and the others trying to find the best and fastest algorithms for classifying into tracks, particles, particle events and vacuum processes (I lump all the quark gluon and many electromagnetic events into vacuum properties.  That is my personal view of things.)\n\nWhat did you think about running the simulation, so as to provide timing data?  Assuming picosecond resolution, for instance?  Or a range of time resolutions on hits?  I think the people gathered here can solve that easily, if they have already tried to solve your current problem, and the results might better inform decision-making.  Your test fairly faithfully represents the present situation and its data limitations.  The problem I think you ought to solve is how to change the detectors and signal processing to capture enough information to resolve the particles and events, and not throw away so much good information.  So I am suggesting you run the time resolved simulation and let these people look at it, while everyone has these things in mind.  It won't affect the contest much, but might provide some new insights toward better algorithms.  More, I think it will move everyone toward a solution to reduce the cost of upgrades to equipment and get to where you can capture all the data.\n\nI think the memory for the event data is too small, relative to the full size of the data stream.  This is a vast simplification, but it seems you are trying to represent a many gigabyte event with a few megabytes of data.  No matter how much you want to resolve it, there is not enough to do the filtering precisely and efficiently.  So you are pretty much guaranteed to lose.  If you compare the difficulty of capturing and filter a gigabyte per event, to the huge cost of unraveling poorly documented hit data after the event, I think you might see what I am aiming at.\n\nRichard Collins",
    "355532": "It looks as if the simulation went \"We have all these particleIds, what cells did they hit and what value would it yield?\" So, in eventId 1000, hitIds 14589, 95, 604, 5, 13 &amp; 4 (with six particleIds) have the same cellId (305/5) and similar values (.25-.37). Does the cell contain sub cells?\n\nI still have no answer from the organizers on this one...",
    "355554": "Dear Glimmung,\n\nI expect that they ran the simulation, it generated those six particles, the particles went through cell 305-5, and all deposited similar amounts of charge/energy.  \n\nIf the particles all had nearly identical initial properties, and went through the one point, then hopefully they are sufficiently different elsewhere to unravel them.  But, in any case, where there are two or more cells in one sensor, their signals can say something about the energy and direction through that cell.  That is a lot easier to use as a constraint than a huge number of completely independent \"hits\" with no direction or energy attached. \n\nIn your example, from Test_100 by the way, the cell 305-5 has 23 hits associated with it (I sorted by ch0 then ch1).   These are the hit numbers 14589, 14595, 14604, 14605, 14613, 14614, 72232, 82386, 109437, 110299, 110383, 110753, 110952, 111081, 112450, 113187, 114008, 114311, 115134, 115614, 115854, 116161, 116540.\n\nIn the hits file these correspond to many particles, the ones all starting with 6395 have very similar properties, and are the ones you are probably referring to.  These are 639511834298179000, 639511834298167000, 639511834298171000, 639511834298187000, 639511834298175000, 639511834298183000.\n\nIf you propagate these particles from the origin through the magnetic field, and they all have to go through that one cell, then that tells you a bit about their mass, type, charge and the time they go through each cell. \n\nThese other particles : 968286014512562000, 734097665658191000, 238692841851723000, 864702054851936000, 589971757343965000, 734106118170611000, 499903269489868000, 117094071364751000, 864698000402808000, 292762425642450000, 797156306778587000, 418853731921035000 all go through cell 305-5 as well.\n\nI presume the cell identifier is unique and not used in different sensors.  I get the impression that providing the cell data is an afterthought.\n\nI would start with the hits having more than one cell, work out the constraints on the trajectories, assign those particles first, then use another method for the other tracks.  My guess is the multi-cell hits are going to be more interesting any way.  It is nice the constraints are tighter.\n\nRichard",
    "355570": "A cell does not contain a sub cell.  And this looks like close to reality actually.  There is a limit to the resolution you can have for hit locations. Two interesting kernels look at cell geometry: \n\nhttps://www.kaggle.com/asalzburger/pixel-detector-cells\n\nhttps://www.kaggle.com/jakubguzowski/calculate-angle-of-incidence-based-on-cells-csv",
    "355571": "Hi, you write to me as if I was part of the organizer team.  I am not, and I do have my concerns about the setup of this competition, because of bugs in the simulator.   But for the rest it looks like they really tried to simulate something close to the forthcoming detectors at LHC.",
    "355591": "CPMP,\n\nI am new to this group.  I was writing out of respect for your progress in the competition, and I was interested in your opinions.  I was also writing for anyone who might read it now, and in the future.\n\nThat aside, I agree with you that the simulation is close, but there are many approximations and hidden assumptions.  In a real world problem, anyone would have complete documentation on the instruments, sensors, processing, models and assumptions.  This is a bit artificial.  And, as I said, I think it will only yield a reverse engineered model of the organizer's simulation.  And I doubt that the results are going to do more than confirm the organizers' understanding of things, and not likely reveal new insights into the real world.  That is not a criticism, just a fact - data from a model will only reveal the model.  We might go so far as to completely reengineer the model, but it is an exercise.  Perhaps to find new employees or associates or ideas.\n\nThe organizers could produce truth data with timings, then the modelling would test the feasibility of a different approach to instrumenting the experiments.  A way to test the effectiveness and feasibility of fast algorithms, using timed hits, not asynchronous hits.  I think the results will be orders of magnitude faster for classification and calibration, which might well justify the cost of paying for better hardware and signal processing chains.  \n\nPehaps the whole group might be interested in tackling some harder problems, related to this one.  You all are close to some good stuff, but aimed in the wrong direction slightly.  You all can expend a lot of energy solving this artificial and ambiguous problem, or specify and solve a problem with more realistic constraints, that will provide the field with useful guidance.  I would aim for 100% and accept 99.995% as a first cut.  The harder the problem, the more you learn.\n\nRichard",
    "355727": "cell channel id is local to a given volume_id,layer_id,module_id. This is indeed the case in your example. What has happened is that a particle has hit the module, interacted with it and created a spray of secondary particles. Each of these particles causes a different hit in the same cell. One can tell by going from hitid in cell file to hitid,particleid in truth file to particle in particles file and see that the origin coordinates vx,vy,vz of the particle 4.45153,-30.7167,-597.5 correspond to the coordinate of the hits. \nWhen several particles go through the same pixel (=cell) we consider each particle to be measured independently. This is a simplification, justified by 1) this allows to have a one to one correspondance between measured hit and  true hit 2) this is a rare occurence, a few per mille of the hits.",
    "355734": "Thank you, David; much appreciated! \n\nSo, in real life, we would get only one read-out from the cell?\n\nWhat would the value be? Sum of non-zero weights = .855? Their average = .285? Sum of all = 1.82? Average of all = .303? Or is the value of no &gt;ahem&lt; value?\n\nEdit: I see now. Apologies. I was lazy. I see in particles that vR and vZ are way out. I will exclude them!",
    "355737": "The response is linear to a good approximation, so this would be the sum.",
    "355739": "See my answer above. Be careful that cell channel number is local, one has to use volume_id,layer_id,module_id to get a unique numbering. Also be careful that particleid is a i64 (you should rarely have multiple of 1000)",
    "355754": "Hi Richard, \n(answer here for your two posts)\n - yes ps precision timing is possible but one has yet to demonstrate that 1) this is possible with the very fine granularity necessary to get the spatial resolution we need 2) with no or reasonable increase of the price tag\n - when designing a challenge one has to chose one specific problem (here pattern recognition/connecting the dots) for particle physics, more precisely for future LHC tracker, balance the simplification of the problem (so that it can be addressed to a large community with little or even no background in the domain) while maintaining sufficient complexity so that the best solutions to the challenge remain relevant to the domain\n- we expect/hope beyond this challenge to foster long term activity/collaboration/publications on pattern recognition/connecting the dots for particle physics",
    "355764": "Excellent! Thank you!\n\nWhile I have you on the line, is it something similar with event 1012, hits 17156, 9 &amp; 75? Cell 69/409.\nVolume 8, layer 2 &amp; module 92. vZ looks fairly normal at 1.82 - 6.15 for three particles?",
    "355777": "3 particles are indeed coming from origin, looks more like an accidental coincidence (than particle interaction as in your first example)",
    "355797": "Richard,\n\nI understand your point.  In many of the Kaggle competitions I entered some participants (including me from time to time) express doubts about how the competition is defined.  Sometimes the benefit to the sponsor is really unclear.\n\nHowever, I'd rather spend my energy on solving the problem as defined by the sponsor than trying to have the sponsor change the problem definition.  After all, they are the ones paying, and if they think their way is right, fine with me.",
    "355832": "David, \n\nThank you for explaining; it is as I suspected.  I am trying to find ways to make collaborative websites more effective.  Even though this is a competition, collectively everyone is trying hard to solve this  problem, and some of them are interested in the problem(s) which caused you to pose it in the first place.  Here I think this whole group would gain by trying the timing problem.  Sort of an \"optional problem\".\n\nWhat I am looking at is the cost of *not* collecting all the data using an 85% filter algorithm.  If you significantly increase the collision rate, and data rate, yet continue to have such poor identification, then the cost to all the subsequent analysis is very high.  If you can make precise identifications, you will not miss, or misinterpret, rare events so often.  Alternatively, one can simply ask, what information would provide 99.9% accuracy in identifying tracks and particles and try to estimate the cost and benefits of doing so.  Or 97%, or 99.999%.\n\nIf, for instance, the z axis were instrumented alone, then the hits will be separated (by time) into a fairly large number, k, of clusters.  For even n^2 problems, the sum of k \"n/k\" problems is k times smaller than one \"n\" problem.  I think this is actually one of those n^2.5 problems so the gain in speed is going to be about k^(1.5) = n^2.5/[k*(n/k)^2.5], where there are k \"n/k\" problems rather then one \"n\" problem.  Sorry it hard to write this here.\n\nFor ten clusters that is a factor of 31.6.  If you can break x, y, z data into 10 groups each, k will be 1000, and the speedup about 1000^(1.5) or about 30,000.  That gives you a lot of room to run more agressive algorithms.  Also, the time gives you an ordered set, so it is faster still, since several events in a particular order do not have to be compared without order.\n\nLHC throws out about 99% of the data now, and expects 10x more?   How many events is it missing because of ambiguity or incomplete classification?\n\n10 Gsps is about 3 centimeter resolution.  If you break z, alone, into 100 identifiable groups by timing, then the speedup (depending on your classification algorithm efficiency) is roughly 100^1.5 =  1000 times faster.\n\nI was explaining this to my older neighbor yesterday.  If she first sorts her 1000 piece puzzle into 10 piles roughly by color and texture, then solves each individual piles separately, she can do it about 10 times faster, depending on her efficiency with processing the smaller piles.  I kept the math simple.  She is trading space for speed, and having to implement two strategies - sorting, and then matching and fitting.  Manual puzzle identification and sorting has a lot of fixed costs.  I am trying to find ways to reduce that in signal processing chains.  \n\nI spent a lot of time this year looking at \"compile to silicon\" methods.  You are proposing to test the speed of these contest algorithms on an ordinary PC.  Have you tried the same on an FPGA?  The usual speedup is about 100, and they are cheaper, smaller and use less energy. That is why the cloud guys are all moving that way.  I think they have room for improvement since I usually get 300 and sometime more.\n\nI apologize this is so crudely described here.  It is hard to write general rule for comparing algorithms.\n\nRichard",
    "355841": "Richard, I think many of the points you make are valid and I guess such points are being discussed in the research community. However, they would be beyond the scope of a kaggle competition, I think.\n\nIn my opinion, the organizers did a good job of simplifying just enough so that the problem becomes approachable for non-specialists while still remaining interesting. The barrier of entry is already higher than in the typical kaggle competition, I think. (By that I mean the amount of work needed until machine learning techniques can be applied, not the difficulty of the challenge, which is very high also in the other competitions, I assume.)\n\nLet's not forget that the organizers made quite a big and risky investment of their time for setting up this competition *in addition to*, not as a replacement for, their work in the usual scientific context.",
    "355866": "there is no point discussing the merit of timing information, the technology to have it while preserving spatial information and within budget is not there\n- we are actually throwing away ~39999/40000 of the data (or to be more specific, of the proton bunch collisions) . This is is a complex issue unrelated to this challenge.\n- we are using already FPGA, and custom ASIC even in some cases. The point of  running on ordinary PC core is that:\n  - it is the bulk of our resources\n  - it is (much) more portable\n  - more important, it does not require rare competences",
    "355889": "Edwin,\n\nYes, I have a great respect for the organizers, and appreciate the time they invested.  But, too, a competition of this sort can sometimes absorb scarce resources that would otherwise be working on critical problems in the real world.  That is why I do not like toy problems, and try to find ways to structure cooperative groups who can apply non-specialists to highly specialized problems in the real world, with a good chance of solving them.\n\nThe \"non-specialist\" part of this could be improved a bit, if you allow that part of what happens in a Kaggle kind of environment is that people get introduced to new areas of research.   There are a whole range of \"fast learning\" methods that can be applied.  That means still more expense in setting up these competition/training grounds/collaborative research groups.  But, I hope that Kaggle is constantly monitoring the methods for the whole, and then investing that back into improving the process.\n\nIf the organizers' investment is high, then solving the real problem can offer still greater rewards.  Say you invest three man years to set up this thing, but it saves $10 billion in the short term, and millions of man years eventually when the new information is absorbed by society.  That is a good trade.  Society might benefit that much, and not specific individuals.  I judge by society as a whole, and not so much the individuals.  Realistically, if the group cuts the cost of an LHC upgrade by a factor of 10, I expect the organizers and participants might be personally and professionally rewarded. \n\nI apologize if I am a bit vague here.  I am trying to restructure the Internet as a whole, including things like collaborative and topic sites.  I have to use methods from hundreds of disciplines, and they do not all fit neatly together.  I am new here and still talking to strangers.  I am reviewing CERN as a whole, and opendata.cern.ch particularly.  I came here because I think this group has a high probability of changing CERN, with just relatively tiny adjustments.  Doing it with non-specialists offers a positive outlook for so many of the outstanding problems in the world today.  I know how to take any problem and structure it for non-specialists.  And non-specialists can improve their methods to become even more effective at solving any problem.  Machine learning pales compared to getting human groups to learn and solve problems faster.  Perhaps they should make a new discipline called HL, human learning.   Or HML for human machine learning.  When a person with computers can do in days what it takes others years to do, that changes the whole game.  And it is a lot of fun.\n\nRichard",
    "355897": "David,\n\nThanks for all your hard work!  My general comments, just getting started, are not a criticism.  More along the lines of finding what others are doing and thinking.  You have focused onto a good problem here.  Maybe at some point people can use real data to test algorithms in the next phase.   Then a hardware competition as an adjunct (or side bet) could benefit LHC, and you can still get done what you want to do.\n\nI will not mention this again.  I was thinking of offering a $1000 prize to anyone who can implement the fastest, low cost,  good algorithm, regardless of hardware.  That will help LHC, and is very timely.  Your \"non specialists\" are actually deep in a wide range of competencies.  And many of these things are actually easy to pick up with a short introduction.\n\nThanks again.  I have something else to do now.\n\nRichard"
  },
  "source": "meta"
}