{
  "id": 67107,
  "title": "Welcome!",
  "url": "/competitions/PLAsTiCC-2018/discussion/67107",
  "author_name": "",
  "post_date": "2018-09-28T21:37:20.064867800Z",
  "votes": 21,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Welcome to the PLAsTiCC Astronomical Classification challenge. You'll need to classify astronomical events based on censored time series data (for any given object, the telescope isn't pointed at it most of the time). The methods used in the competition will help inform how astronomical data is processed in the real world. The data are simulated, but this is a real problem facing the astronomical community. We're excited to see what you can come up with!</p>\n\n<p>Happy Kaggling!</p>",
  "messages": [
    {
      "id": "395577",
      "postDate": "09/28/2018 21:37:20",
      "content": "<p>Welcome to the PLAsTiCC Astronomical Classification challenge. You'll need to classify astronomical events based on censored time series data (for any given object, the telescope isn't pointed at it most of the time). The methods used in the competition will help inform how astronomical data is processed in the real world. The data are simulated, but this is a real problem facing the astronomical community. We're excited to see what you can come up with!</p>\n\n<p>Happy Kaggling!</p>",
      "rawMarkdown": "Welcome to the PLAsTiCC Astronomical Classification challenge. You'll need to classify astronomical events based on censored time series data (for any given object, the telescope isn't pointed at it most of the time). The methods used in the competition will help inform how astronomical data is processed in the real world. The data are simulated, but this is a real problem facing the astronomical community. We're excited to see what you can come up with!\n\nHappy Kaggling!",
      "votes": null
    },
    {
      "id": "395619",
      "postDate": "09/29/2018 01:38:27",
      "content": "<p>This data is vary interesting!!\nI am curious ,the unique target value in training_set_metadata.csv :92, 88, 42, 90, 65, 16, 67, 95, 62, 15, 52,  6, 64, 53\nbut class_99 is in sample_submission.csv column?</p>",
      "rawMarkdown": "This data is vary interesting!!\nI am curious ,the unique target value in training_set_metadata.csv :92, 88, 42, 90, 65, 16, 67, 95, 62, 15, 52,  6, 64, 53\nbut class_99 is in sample_submission.csv column?",
      "votes": null
    },
    {
      "id": "395627",
      "postDate": "09/29/2018 02:31:06",
      "content": "<p>Hi there!</p>\n\n<p>This is described in the starter kernel a bit: <a href=\"https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit\">https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit</a></p>\n\n<p>The short version is that we expect that LSST will discover new classes of astrophysical sources that we've never seen before. Because we've never seen them before, they can't be represented in the training set. However, you might be able to find some objects in the test that don't quite fit the bill when considered against the training classes. These might be some of the most scientifically interesting results as well! Think of it as anomaly detection - label these kinds of objects with class_99 if you are confident they're not representative of something in the training set.</p>",
      "rawMarkdown": "Hi there!\n\nThis is described in the starter kernel a bit: https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit\n\nThe short version is that we expect that LSST will discover new classes of astrophysical sources that we've never seen before. Because we've never seen them before, they can't be represented in the training set. However, you might be able to find some objects in the test that don't quite fit the bill when considered against the training classes. These might be some of the most scientifically interesting results as well! Think of it as anomaly detection - label these kinds of objects with class_99 if you are confident they're not representative of something in the training set.",
      "votes": null
    },
    {
      "id": "399128",
      "postDate": "10/05/2018 09:04:37",
      "content": "<p>Class 99 is the most strange објect in Universe:)</p>",
      "rawMarkdown": "Class 99 is the most strange објect in Universe:)",
      "votes": null
    },
    {
      "id": "417466",
      "postDate": "11/08/2018 10:19:03",
      "content": "<p>Can you produce an additional dataset that is sorted by object_id and mjd - most people are aggregating on chunks which means people who don't have huge RAM are at a distinct disadvantage which is unfair</p>",
      "rawMarkdown": "Can you produce an additional dataset that is sorted by object_id and mjd - most people are aggregating on chunks which means people who don't have huge RAM are at a distinct disadvantage which is unfair",
      "votes": null
    },
    {
      "id": "426882",
      "postDate": "11/24/2018 04:13:01",
      "content": "<p>While I enjoy the challenge so far, I got to the point this evening where I almost said forget it.  </p>\n\n<p>Not sure why you believe a huge test data set makes sense, when as I understand it, the data is fake.  It's looking like at least 8 hours to run the prediction with my current model (this is after a long time running the model).   I understand that the telescope will generate huge amounts of data, but pretty sure your going to be using a huge amount of computing power.</p>\n\n<p>With a model with no predictive power, the process to create a submission was pretty slow but within my pain tolerance.  One a decent model (I think) created than 8 hours to see a leader board score is stupid.  It looks like my main challenge in this challenge will be figuring out how to cut that 8 hours (at least I hope its only 8) of prediction time down.</p>\n\n<p>As noted by Scirpus, folks with small ram have a disadvantage.  I have 64GB and I feel like that's not even close enough.</p>\n\n<p>Most of these Kaggle challenges end up with lots of teams jumping in near the last weeks - my feeble computer predicts this challenge with have record low participation - I will correlate it to the size of the test data set.</p>\n\n<p>I don't mind models that take days to run - that's mostly in my control.  But test sets that takes a full day to process MAKES NO SENSE TO ME.</p>",
      "rawMarkdown": "While I enjoy the challenge so far, I got to the point this evening where I almost said forget it.  \n\nNot sure why you believe a huge test data set makes sense, when as I understand it, the data is fake.  It's looking like at least 8 hours to run the prediction with my current model (this is after a long time running the model).   I understand that the telescope will generate huge amounts of data, but pretty sure your going to be using a huge amount of computing power.\n\nWith a model with no predictive power, the process to create a submission was pretty slow but within my pain tolerance.  One a decent model (I think) created than 8 hours to see a leader board score is stupid.  It looks like my main challenge in this challenge will be figuring out how to cut that 8 hours (at least I hope its only 8) of prediction time down.\n\nAs noted by Scirpus, folks with small ram have a disadvantage.  I have 64GB and I feel like that's not even close enough.\n\nMost of these Kaggle challenges end up with lots of teams jumping in near the last weeks - my feeble computer predicts this challenge with have record low participation - I will correlate it to the size of the test data set.\n\nI don't mind models that take days to run - that's mostly in my control.  But test sets that takes a full day to process MAKES NO SENSE TO ME.",
      "votes": null
    },
    {
      "id": "427743",
      "postDate": "11/26/2018 04:07:59",
      "content": "<p>Hi, my process runs in less than 6MB and it takes 20 minutes to create the submission that scores 0.801.  Uploading the result to Kaggle takes longer actually with my slow home internet connection ;)</p>\n\n<p>This is with lightgbm on a 20 core Xeon machine.  On a 4 core i7 it takes less than one hour.  And for NN models it is even faster.</p>\n\n<p>Key is to work with chunks of the test data, see Olivier's kernel for an example.  Key is also to not recompute features each time you generate a submission: store your features from a run to reuse them in the next run.</p>",
      "rawMarkdown": "Hi, my process runs in less than 6MB and it takes 20 minutes to create the submission that scores 0.801.  Uploading the result to Kaggle takes longer actually with my slow home internet connection ;)\n\nThis is with lightgbm on a 20 core Xeon machine.  On a 4 core i7 it takes less than one hour.  And for NN models it is even faster.\n\nKey is to work with chunks of the test data, see Olivier's kernel for an example.  Key is also to not recompute features each time you generate a submission: store your features from a run to reuse them in the next run.",
      "votes": null
    },
    {
      "id": "427750",
      "postDate": "11/26/2018 04:25:57",
      "content": "<p>Thanks for info - will look into chunks.</p>\n\n<p>Still think the test set is evil over-kill.  Perhaps if the data was real I would feel better about it - but fake data created by some astronomy simulation seems like a huge waste of resources (not as bad as bitcoin mining, but getting there :)</p>\n\n<p>But perhaps the challenge is actually an evaluation of that simulation - a huge category 99 result would indicate the sim needs lots of work?</p>",
      "rawMarkdown": "Thanks for info - will look into chunks.\n\nStill think the test set is evil over-kill.  Perhaps if the data was real I would feel better about it - but fake data created by some astronomy simulation seems like a huge waste of resources (not as bad as bitcoin mining, but getting there :)\n\nBut perhaps the challenge is actually an evaluation of that simulation - a huge category 99 result would indicate the sim needs lots of work?",
      "votes": null
    },
    {
      "id": "428844",
      "postDate": "11/28/2018 00:52:26",
      "content": "<p>Its difficult to load the data to pd dataframe even with high performance laptop</p>",
      "rawMarkdown": "Its difficult to load the data to pd dataframe even with high performance laptop",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 395619,
      "author_name": "yentianbao",
      "author_url": "",
      "post_date": "09/29/2018 01:38:27",
      "content": "<p>This data is vary interesting!!\nI am curious ,the unique target value in training_set_metadata.csv :92, 88, 42, 90, 65, 16, 67, 95, 62, 15, 52,  6, 64, 53\nbut class_99 is in sample_submission.csv column?</p>",
      "votes": null,
      "replies": [
        {
          "id": 395627,
          "author_name": "gsnarayan",
          "author_url": "",
          "post_date": "09/29/2018 02:31:06",
          "content": "<p>Hi there!</p>\n\n<p>This is described in the starter kernel a bit: <a href=\"https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit\">https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit</a></p>\n\n<p>The short version is that we expect that LSST will discover new classes of astrophysical sources that we've never seen before. Because we've never seen them before, they can't be represented in the training set. However, you might be able to find some objects in the test that don't quite fit the bill when considered against the training classes. These might be some of the most scientifically interesting results as well! Think of it as anomaly detection - label these kinds of objects with class_99 if you are confident they're not representative of something in the training set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 399128,
      "author_name": "mykras",
      "author_url": "",
      "post_date": "10/05/2018 09:04:37",
      "content": "<p>Class 99 is the most strange објect in Universe:)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 417466,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "11/08/2018 10:19:03",
      "content": "<p>Can you produce an additional dataset that is sorted by object_id and mjd - most people are aggregating on chunks which means people who don't have huge RAM are at a distinct disadvantage which is unfair</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 426882,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "11/24/2018 04:13:01",
      "content": "<p>While I enjoy the challenge so far, I got to the point this evening where I almost said forget it.  </p>\n\n<p>Not sure why you believe a huge test data set makes sense, when as I understand it, the data is fake.  It's looking like at least 8 hours to run the prediction with my current model (this is after a long time running the model).   I understand that the telescope will generate huge amounts of data, but pretty sure your going to be using a huge amount of computing power.</p>\n\n<p>With a model with no predictive power, the process to create a submission was pretty slow but within my pain tolerance.  One a decent model (I think) created than 8 hours to see a leader board score is stupid.  It looks like my main challenge in this challenge will be figuring out how to cut that 8 hours (at least I hope its only 8) of prediction time down.</p>\n\n<p>As noted by Scirpus, folks with small ram have a disadvantage.  I have 64GB and I feel like that's not even close enough.</p>\n\n<p>Most of these Kaggle challenges end up with lots of teams jumping in near the last weeks - my feeble computer predicts this challenge with have record low participation - I will correlate it to the size of the test data set.</p>\n\n<p>I don't mind models that take days to run - that's mostly in my control.  But test sets that takes a full day to process MAKES NO SENSE TO ME.</p>",
      "votes": null,
      "replies": [
        {
          "id": 427743,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/26/2018 04:07:59",
          "content": "<p>Hi, my process runs in less than 6MB and it takes 20 minutes to create the submission that scores 0.801.  Uploading the result to Kaggle takes longer actually with my slow home internet connection ;)</p>\n\n<p>This is with lightgbm on a 20 core Xeon machine.  On a 4 core i7 it takes less than one hour.  And for NN models it is even faster.</p>\n\n<p>Key is to work with chunks of the test data, see Olivier's kernel for an example.  Key is also to not recompute features each time you generate a submission: store your features from a run to reuse them in the next run.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 427750,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "11/26/2018 04:25:57",
          "content": "<p>Thanks for info - will look into chunks.</p>\n\n<p>Still think the test set is evil over-kill.  Perhaps if the data was real I would feel better about it - but fake data created by some astronomy simulation seems like a huge waste of resources (not as bad as bitcoin mining, but getting there :)</p>\n\n<p>But perhaps the challenge is actually an evaluation of that simulation - a huge category 99 result would indicate the sim needs lots of work?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 428844,
      "author_name": "srinibujjay",
      "author_url": "",
      "post_date": "11/28/2018 00:52:26",
      "content": "<p>Its difficult to load the data to pd dataframe even with high performance laptop</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "395577": "Welcome to the PLAsTiCC Astronomical Classification challenge. You'll need to classify astronomical events based on censored time series data (for any given object, the telescope isn't pointed at it most of the time). The methods used in the competition will help inform how astronomical data is processed in the real world. The data are simulated, but this is a real problem facing the astronomical community. We're excited to see what you can come up with!\n\nHappy Kaggling!",
    "395619": "This data is vary interesting!!\nI am curious ,the unique target value in training_set_metadata.csv :92, 88, 42, 90, 65, 16, 67, 95, 62, 15, 52,  6, 64, 53\nbut class_99 is in sample_submission.csv column?",
    "395627": "Hi there!\n\nThis is described in the starter kernel a bit: https://www.kaggle.com/michaelapers/the-plasticc-astronomy-starter-kit\n\nThe short version is that we expect that LSST will discover new classes of astrophysical sources that we've never seen before. Because we've never seen them before, they can't be represented in the training set. However, you might be able to find some objects in the test that don't quite fit the bill when considered against the training classes. These might be some of the most scientifically interesting results as well! Think of it as anomaly detection - label these kinds of objects with class_99 if you are confident they're not representative of something in the training set.",
    "399128": "Class 99 is the most strange објect in Universe:)",
    "417466": "Can you produce an additional dataset that is sorted by object_id and mjd - most people are aggregating on chunks which means people who don't have huge RAM are at a distinct disadvantage which is unfair",
    "426882": "While I enjoy the challenge so far, I got to the point this evening where I almost said forget it.  \n\nNot sure why you believe a huge test data set makes sense, when as I understand it, the data is fake.  It's looking like at least 8 hours to run the prediction with my current model (this is after a long time running the model).   I understand that the telescope will generate huge amounts of data, but pretty sure your going to be using a huge amount of computing power.\n\nWith a model with no predictive power, the process to create a submission was pretty slow but within my pain tolerance.  One a decent model (I think) created than 8 hours to see a leader board score is stupid.  It looks like my main challenge in this challenge will be figuring out how to cut that 8 hours (at least I hope its only 8) of prediction time down.\n\nAs noted by Scirpus, folks with small ram have a disadvantage.  I have 64GB and I feel like that's not even close enough.\n\nMost of these Kaggle challenges end up with lots of teams jumping in near the last weeks - my feeble computer predicts this challenge with have record low participation - I will correlate it to the size of the test data set.\n\nI don't mind models that take days to run - that's mostly in my control.  But test sets that takes a full day to process MAKES NO SENSE TO ME.",
    "427743": "Hi, my process runs in less than 6MB and it takes 20 minutes to create the submission that scores 0.801.  Uploading the result to Kaggle takes longer actually with my slow home internet connection ;)\n\nThis is with lightgbm on a 20 core Xeon machine.  On a 4 core i7 it takes less than one hour.  And for NN models it is even faster.\n\nKey is to work with chunks of the test data, see Olivier's kernel for an example.  Key is also to not recompute features each time you generate a submission: store your features from a run to reuse them in the next run.",
    "427750": "Thanks for info - will look into chunks.\n\nStill think the test set is evil over-kill.  Perhaps if the data was real I would feel better about it - but fake data created by some astronomy simulation seems like a huge waste of resources (not as bad as bitcoin mining, but getting there :)\n\nBut perhaps the challenge is actually an evaluation of that simulation - a huge category 99 result would indicate the sim needs lots of work?",
    "428844": "Its difficult to load the data to pd dataframe even with high performance laptop"
  },
  "source": "meta"
}