{
  "id": 56172,
  "title": "computation power",
  "url": "/competitions/trackml-particle-identification/discussion/56172",
  "author_name": "Jacek Poplawski",
  "post_date": "2018-05-07T09:36:23.701000",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>This competition looks very cool, but I wonder - can we even start without big cluster of computers? It looks very computation intensive task. What are your thoughts?</p>",
  "messages": [
    {
      "id": 324415,
      "postDate": "2018-05-07T17:20:25.900Z",
      "content": "<p>You can definitely start with a standard computer, for instance with the DB scan kernel. For going further, it all depends on the method you choose. Two important ideas: </p>\n\n<ul>\n<li>The dataset is very large. It does not mean that you have to use all\nof it! For instance, if you plan to learn a model, you might be able\nto learn it only from a fraction of the data. It's up to you.   </li>\n<li>The ultimate goal is to have fast inference (applying the model to the\ntest data) on standard processors, although in the accuracy phase,\nthere is no consideration of speed.</li>\n</ul>",
      "rawMarkdown": "You can definitely start with a standard computer, for instance with the DB scan kernel. For going further, it all depends on the method you choose. Two important ideas: \n \n\n - The dataset is very large. It does not mean that you have to use all\n   of it! For instance, if you plan to learn a model, you might be able\n   to learn it only from a fraction of the data. It's up to you.   \n - The ultimate goal is to have fast inference (applying the model to the\n   test data) on standard processors, although in the accuracy phase,\n   there is no consideration of speed.\n",
      "votes": 3,
      "replies": [
        {
          "id": 324420,
          "postDate": "2018-05-07T17:24:12.527Z",
          "content": "<p>so there is still a hope :) thank you</p>",
          "rawMarkdown": "so there is still a hope :) thank you"
        }
      ]
    },
    {
      "id": 325363,
      "postDate": "2018-05-08T10:52:19.637Z",
      "content": "<p>@Jacek, I think, 10 events for training is enough to reach the score ~0.6, so we can start with one computer and a few good ideas. </p>",
      "rawMarkdown": "@Jacek, I think, 10 events for training is enough to reach the score ~0.6, so we can start with one computer and a few good ideas. ",
      "votes": 4,
      "replies": [
        {
          "id": 325366,
          "postDate": "2018-05-08T10:53:54.040Z",
          "content": "<p>Thanks Grzegorz - and congratulations on your LB score :)</p>",
          "rawMarkdown": "Thanks Grzegorz - and congratulations on your LB score :)"
        },
        {
          "id": 330743,
          "postDate": "2018-05-19T15:52:48.283Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 330841,
          "postDate": "2018-05-19T20:32:48.217Z",
          "content": "<p>@Ram, My score 0.42 (~0.2 above the benchmark) was obtained using one personal computer, one simple idea, one event and run time below one hour, so I think the score 0.6 was rather a safe approximation for one computer, a few good ideas and ten events. The emphasis is still on the number and the quality of ideas (feature engineering) and not on the number of events.</p>",
          "rawMarkdown": "@Ram, My score 0.42 (~0.2 above the benchmark) was obtained using one personal computer, one simple idea, one event and run time below one hour, so I think the score 0.6 was rather a safe approximation for one computer, a few good ideas and ten events. The emphasis is still on the number and the quality of ideas (feature engineering) and not on the number of events.",
          "votes": 7
        }
      ]
    },
    {
      "id": 324157,
      "postDate": "2018-05-07T09:36:23.700Z",
      "content": "<p>This competition looks very cool, but I wonder - can we even start without big cluster of computers? It looks very computation intensive task. What are your thoughts?</p>",
      "rawMarkdown": "This competition looks very cool, but I wonder - can we even start without big cluster of computers? It looks very computation intensive task. What are your thoughts?",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 324415,
      "author_name": "CecileGermain",
      "author_url": "",
      "post_date": "2018-05-07T17:20:25.900000",
      "content": "<p>You can definitely start with a standard computer, for instance with the DB scan kernel. For going further, it all depends on the method you choose. Two important ideas: </p>\n\n<ul>\n<li>The dataset is very large. It does not mean that you have to use all\nof it! For instance, if you plan to learn a model, you might be able\nto learn it only from a fraction of the data. It's up to you.   </li>\n<li>The ultimate goal is to have fast inference (applying the model to the\ntest data) on standard processors, although in the accuracy phase,\nthere is no consideration of speed.</li>\n</ul>",
      "votes": 3,
      "replies": [
        {
          "id": 324420,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2018-05-07T17:24:12.527000",
          "content": "<p>so there is still a hope :) thank you</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 325363,
      "author_name": "Grzegorz Sionkowski",
      "author_url": "",
      "post_date": "2018-05-08T10:52:19.637000",
      "content": "<p>@Jacek, I think, 10 events for training is enough to reach the score ~0.6, so we can start with one computer and a few good ideas. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 325366,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2018-05-08T10:53:54.040000",
          "content": "<p>Thanks Grzegorz - and congratulations on your LB score :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 330743,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-05-19T15:52:48.283000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 330841,
          "author_name": "Grzegorz Sionkowski",
          "author_url": "",
          "post_date": "2018-05-19T20:32:48.217000",
          "content": "<p>@Ram, My score 0.42 (~0.2 above the benchmark) was obtained using one personal computer, one simple idea, one event and run time below one hour, so I think the score 0.6 was rather a safe approximation for one computer, a few good ideas and ten events. The emphasis is still on the number and the quality of ideas (feature engineering) and not on the number of events.</p>",
          "votes": 7,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "324415": "You can definitely start with a standard computer, for instance with the DB scan kernel. For going further, it all depends on the method you choose. Two important ideas: \n \n\n - The dataset is very large. It does not mean that you have to use all\n   of it! For instance, if you plan to learn a model, you might be able\n   to learn it only from a fraction of the data. It's up to you.   \n - The ultimate goal is to have fast inference (applying the model to the\n   test data) on standard processors, although in the accuracy phase,\n   there is no consideration of speed.\n",
    "325363": "@Jacek, I think, 10 events for training is enough to reach the score ~0.6, so we can start with one computer and a few good ideas. ",
    "324157": "This competition looks very cool, but I wonder - can we even start without big cluster of computers? It looks very computation intensive task. What are your thoughts?"
  }
}