{
  "id": 58614,
  "title": "Competition heating up",
  "url": "/competitions/trackml-particle-identification/discussion/58614",
  "author_name": "CPMP",
  "post_date": "2018-06-11T08:41:33.221000",
  "votes": 4,
  "comment_count": 19,
  "views": 0,
  "content": "<p>The top 3 leaders all improved by about 0.04 in the past day.  Impressive.</p>",
  "messages": [
    {
      "id": 341225,
      "postDate": "2018-06-11T08:41:33.220Z",
      "content": "<p>The top 3 leaders all improved by about 0.04 in the past day.  Impressive.</p>",
      "rawMarkdown": "The top 3 leaders all improved by about 0.04 in the past day.  Impressive.",
      "votes": 4
    },
    {
      "id": 341257,
      "postDate": "2018-06-11T10:06:26.107Z",
      "content": "<p>my estimation is public LB of &gt;0.80 at the end of the challenge. There are 2 more months to go. </p>",
      "rawMarkdown": "my estimation is public LB of &gt;0.80 at the end of the challenge. There are 2 more months to go. ",
      "votes": 1,
      "replies": [
        {
          "id": 341269,
          "postDate": "2018-06-11T10:31:17.107Z",
          "content": "<p>I agree, even think 0.9 is reachable.</p>",
          "rawMarkdown": "I agree, even think 0.9 is reachable.",
          "votes": 1
        },
        {
          "id": 341477,
          "postDate": "2018-06-11T16:34:34.193Z",
          "content": "<p>The contest overview document says that \"typically, at least 90% of the true tracks should be recovered.\"  Although I'm not sure how that statement translates into our scoring metric since the metric incorporates weights we can't see (in the test set).</p>",
          "rawMarkdown": "The contest overview document says that \"typically, at least 90% of the true tracks should be recovered.\"  Although I'm not sure how that statement translates into our scoring metric since the metric incorporates weights we can't see (in the test set)."
        },
        {
          "id": 341481,
          "postDate": "2018-06-11T16:52:05.917Z",
          "content": "<blockquote>\n  <p>The contest overview document says that \"typically, at least 90% of the true tracks should be recovered.\" </p>\n</blockquote>\n\n<p>I asked almost a week ago if this was wishful thinking or not, and I am still waiting for the answer.  This makes me think it is wishful thinking.</p>",
          "rawMarkdown": "&gt; The contest overview document says that \"typically, at least 90% of the true tracks should be recovered.\" \n\nI asked almost a week ago if this was wishful thinking or not, and I am still waiting for the answer.  This makes me think it is wishful thinking."
        },
        {
          "id": 341506,
          "postDate": "2018-06-11T17:33:42.263Z",
          "content": "<p>I assume that the current methodology, Kalman filtering, has been developed over the past decade or more.  That gives the &gt;90% coverage.</p>\n\n<p>We know that the team at CERN has been evaluating ML methodology for several years.  Have they published anything conclusive?</p>\n\n<p>As smart as the Kaggle community is, it may be overly optimistic to think we can achieve the same level of performance with 3 months to develop a new method.</p>\n\n<p>I assume what the organizers really hope for is a collection of new ideas that they never considered before...</p>",
          "rawMarkdown": "I assume that the current methodology, Kalman filtering, has been developed over the past decade or more.  That gives the &gt;90% coverage.\n\nWe know that the team at CERN has been evaluating ML methodology for several years.  Have they published anything conclusive?\n\nAs smart as the Kaggle community is, it may be overly optimistic to think we can achieve the same level of performance with 3 months to develop a new method.\n\nI assume what the organizers really hope for is a collection of new ideas that they never considered before...",
          "votes": 5
        },
        {
          "id": 341581,
          "postDate": "2018-06-11T20:27:44.620Z",
          "content": "<blockquote>\n  <p>That gives the &gt;90% coverage.</p>\n</blockquote>\n\n<p>Not on this data.</p>",
          "rawMarkdown": "&gt; That gives the &gt;90% coverage.\n\nNot on this data."
        },
        {
          "id": 341590,
          "postDate": "2018-06-11T20:50:37.003Z",
          "content": "<p>I don't think there is anything different or special about this data vs historical.  I thought the issue is simply in the volume of data generated.  The data collected will outgrow their ability to process it without more $$ to spend  on compute, or lower processing cost per unit of data (i.e. more efficient algorithm)</p>",
          "rawMarkdown": "I don't think there is anything different or special about this data vs historical.  I thought the issue is simply in the volume of data generated.  The data collected will outgrow their ability to process it without more $$ to spend  on compute, or lower processing cost per unit of data (i.e. more efficient algorithm)"
        },
        {
          "id": 341704,
          "postDate": "2018-06-12T05:32:51.067Z",
          "content": "<p>There is, this data is way denser in events than what we see on published papers.</p>",
          "rawMarkdown": "There is, this data is way denser in events than what we see on published papers."
        },
        {
          "id": 341705,
          "postDate": "2018-06-12T05:34:05.607Z",
          "content": "<p>And if I am wrong it is very simple to prove it: the organizers tell us what they get on this data with their best algorithms.  This could motivate us further.</p>",
          "rawMarkdown": "And if I am wrong it is very simple to prove it: the organizers tell us what they get on this data with their best algorithms.  This could motivate us further.",
          "votes": 2
        },
        {
          "id": 342628,
          "postDate": "2018-06-13T20:14:40.960Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 342664,
          "postDate": "2018-06-13T21:12:39.597Z",
          "content": "<p>Whatever.</p>",
          "rawMarkdown": "Whatever.\n",
          "votes": -1
        },
        {
          "id": 343210,
          "postDate": "2018-06-14T21:56:04.450Z",
          "content": "<p>Hello everybody ,</p>\n\n<p>We have seen a short talk at Paris ML meetup on this challenge, expectation of 0,8 of \"accuracy\" seems the target. Effectively classical approach (Kalman filtering and other techniques) are better but not scalable.\nNext monday, the second part of the challenge ; speed optimization will be explained in a dedicated workshop, Hors-serie session of Paris ML.\nSome materials will be published may be</p>",
          "rawMarkdown": "Hello everybody ,\n \nWe have seen a short talk at Paris ML meetup on this challenge, expectation of 0,8 of \"accuracy\" seems the target. Effectively classical approach (Kalman filtering and other techniques) are better but not scalable.\nNext monday, the second part of the challenge ; speed optimization will be explained in a dedicated workshop, Hors-serie session of Paris ML.\nSome materials will be published may be",
          "votes": 5
        },
        {
          "id": 343328,
          "postDate": "2018-06-15T06:28:25.387Z",
          "content": "<p>@Bruno16, thank you, but it does not answer the question: what is the score of current CERN algorithms on this data?  </p>\n\n<p>Having this information could be motivating.  For instance, in DSB, the organizers have shared their performance up front in the leaderboard.</p>",
          "rawMarkdown": "@Bruno16, thank you, but it does not answer the question: what is the score of current CERN algorithms on this data?  \n\nHaving this information could be motivating.  For instance, in DSB, the organizers have shared their performance up front in the leaderboard."
        },
        {
          "id": 343700,
          "postDate": "2018-06-15T21:21:52.310Z",
          "content": "<p>I have a hunch their current systems won't perform well on the dataset, the increased density might make the dataset in this competition too 'out of sample', and it could also have issues with performance. </p>\n\n<p>Either way, I'm sure their algorithms will have lots to learn from this competition, and we can stand to learn a lot from them as well. </p>",
          "rawMarkdown": "I have a hunch their current systems won't perform well on the dataset, the increased density might make the dataset in this competition too 'out of sample', and it could also have issues with performance. \n\nEither way, I'm sure their algorithms will have lots to learn from this competition, and we can stand to learn a lot from them as well. "
        },
        {
          "id": 345107,
          "postDate": "2018-06-19T08:05:15.820Z",
          "content": "<p>@Heng, your prediction is already true, let's see if mine will become true...</p>",
          "rawMarkdown": "@Heng, your prediction is already true, let's see if mine will become true..."
        },
        {
          "id": 345108,
          "postDate": "2018-06-19T08:06:36.970Z",
          "content": "<p>@CPMP</p>\n\n<p>I think i should raise the bar. The top rank results could reach 0.9 :)</p>",
          "rawMarkdown": "@CPMP\n\nI think i should raise the bar. The top rank results could reach 0.9 :)",
          "votes": 2
        }
      ]
    },
    {
      "id": 345070,
      "postDate": "2018-06-19T06:34:40.957Z",
      "content": "<p>Wow! The first place is &gt; 0.8 now!!!\nOutrunner did it with the only one entry.</p>",
      "rawMarkdown": "Wow! The first place is &gt; 0.8 now!!!\nOutrunner did it with the only one entry."
    },
    {
      "id": 341656,
      "postDate": "2018-06-12T03:26:55.453Z",
      "content": "<p>Wow, yes! First score &gt; 0.7 was just submitted</p>",
      "rawMarkdown": "Wow, yes! First score &gt; 0.7 was just submitted",
      "replies": [
        {
          "id": 341658,
          "postDate": "2018-06-12T03:38:35.293Z",
          "content": "<p>i believed supervised learning based has now been used in the challenge. And I expect further improvement. </p>",
          "rawMarkdown": "i believed supervised learning based has now been used in the challenge. And I expect further improvement. ",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 341257,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-06-11T10:06:26.107000",
      "content": "<p>my estimation is public LB of &gt;0.80 at the end of the challenge. There are 2 more months to go. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 341269,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-11T10:31:17.107000",
          "content": "<p>I agree, even think 0.9 is reachable.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 341477,
          "author_name": "Michael Maguire",
          "author_url": "",
          "post_date": "2018-06-11T16:34:34.193000",
          "content": "<p>The contest overview document says that \"typically, at least 90% of the true tracks should be recovered.\"  Although I'm not sure how that statement translates into our scoring metric since the metric incorporates weights we can't see (in the test set).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 341481,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-11T16:52:05.917000",
          "content": "<blockquote>\n  <p>The contest overview document says that \"typically, at least 90% of the true tracks should be recovered.\" </p>\n</blockquote>\n\n<p>I asked almost a week ago if this was wishful thinking or not, and I am still waiting for the answer.  This makes me think it is wishful thinking.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 341506,
          "author_name": "John Sweeney",
          "author_url": "",
          "post_date": "2018-06-11T17:33:42.263000",
          "content": "<p>I assume that the current methodology, Kalman filtering, has been developed over the past decade or more.  That gives the &gt;90% coverage.</p>\n\n<p>We know that the team at CERN has been evaluating ML methodology for several years.  Have they published anything conclusive?</p>\n\n<p>As smart as the Kaggle community is, it may be overly optimistic to think we can achieve the same level of performance with 3 months to develop a new method.</p>\n\n<p>I assume what the organizers really hope for is a collection of new ideas that they never considered before...</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 341581,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-11T20:27:44.620000",
          "content": "<blockquote>\n  <p>That gives the &gt;90% coverage.</p>\n</blockquote>\n\n<p>Not on this data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 341590,
          "author_name": "John Sweeney",
          "author_url": "",
          "post_date": "2018-06-11T20:50:37.003000",
          "content": "<p>I don't think there is anything different or special about this data vs historical.  I thought the issue is simply in the volume of data generated.  The data collected will outgrow their ability to process it without more $$ to spend  on compute, or lower processing cost per unit of data (i.e. more efficient algorithm)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 341704,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-12T05:32:51.067000",
          "content": "<p>There is, this data is way denser in events than what we see on published papers.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 341705,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-12T05:34:05.607000",
          "content": "<p>And if I am wrong it is very simple to prove it: the organizers tell us what they get on this data with their best algorithms.  This could motivate us further.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 342628,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-06-13T20:14:40.960000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 342664,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-13T21:12:39.597000",
          "content": "<p>Whatever.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 343210,
          "author_name": "Bruno16",
          "author_url": "",
          "post_date": "2018-06-14T21:56:04.450000",
          "content": "<p>Hello everybody ,</p>\n\n<p>We have seen a short talk at Paris ML meetup on this challenge, expectation of 0,8 of \"accuracy\" seems the target. Effectively classical approach (Kalman filtering and other techniques) are better but not scalable.\nNext monday, the second part of the challenge ; speed optimization will be explained in a dedicated workshop, Hors-serie session of Paris ML.\nSome materials will be published may be</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 343328,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-15T06:28:25.387000",
          "content": "<p>@Bruno16, thank you, but it does not answer the question: what is the score of current CERN algorithms on this data?  </p>\n\n<p>Having this information could be motivating.  For instance, in DSB, the organizers have shared their performance up front in the leaderboard.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 343700,
          "author_name": "macfarll",
          "author_url": "",
          "post_date": "2018-06-15T21:21:52.310000",
          "content": "<p>I have a hunch their current systems won't perform well on the dataset, the increased density might make the dataset in this competition too 'out of sample', and it could also have issues with performance. </p>\n\n<p>Either way, I'm sure their algorithms will have lots to learn from this competition, and we can stand to learn a lot from them as well. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 345107,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-06-19T08:05:15.820000",
          "content": "<p>@Heng, your prediction is already true, let's see if mine will become true...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 345108,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-06-19T08:06:36.970000",
          "content": "<p>@CPMP</p>\n\n<p>I think i should raise the bar. The top rank results could reach 0.9 :)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 345070,
      "author_name": "Sergey Zlobin",
      "author_url": "",
      "post_date": "2018-06-19T06:34:40.957000",
      "content": "<p>Wow! The first place is &gt; 0.8 now!!!\nOutrunner did it with the only one entry.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 341656,
      "author_name": "Matt",
      "author_url": "",
      "post_date": "2018-06-12T03:26:55.453000",
      "content": "<p>Wow, yes! First score &gt; 0.7 was just submitted</p>",
      "votes": 0,
      "replies": [
        {
          "id": 341658,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-06-12T03:38:35.293000",
          "content": "<p>i believed supervised learning based has now been used in the challenge. And I expect further improvement. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "341225": "The top 3 leaders all improved by about 0.04 in the past day.  Impressive.",
    "341257": "my estimation is public LB of &gt;0.80 at the end of the challenge. There are 2 more months to go. ",
    "345070": "Wow! The first place is &gt; 0.8 now!!!\nOutrunner did it with the only one entry.",
    "341656": "Wow, yes! First score &gt; 0.7 was just submitted"
  }
}