{
  "id": 74846,
  "title": "Enter the LSST Workshop Challenge ",
  "url": "/competitions/PLAsTiCC-2018/discussion/74846",
  "author_name": "",
  "post_date": "2018-12-16T14:38:23.571454600Z",
  "votes": 14,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hi Kagglers</p>\n\n<p>Thank you for all your hard work participating in the PLAsTiCC challenge! While the challenge closes on the 17th of December (tomorrow!), we wanted to bring your attention to the LSST Workshop part of the challenge, which remains open until January 15th 2019 (see details: <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018#Prizes\">https://www.kaggle.com/c/PLAsTiCC-2018#Prizes</a>).</p>\n\n<p>As part of this challenge, we will be rewarding teams that have solutions/algorithms of scientific interest in the following areas:</p>\n\n<ul>\n<li>codes that yield high efficiency at class 99-specific classifications</li>\n<li>novel methods that classify one type of object with high precision (purity) and high recall (completeness), particularly on objects where others have struggled</li>\n<li>models that can generate new light curves from the existing data </li>\n<li>methods that work well on classifying light curves that have large flux uncertainties</li>\n</ul>\n\n<p>In order to enter the LSST Workshop part of the challenge, we ask that you submit a Google form here: <a href=\"https://goo.gl/forms/EjzSeBE1tYPHGVCx1\">https://goo.gl/forms/EjzSeBE1tYPHGVCx1</a> with details of the code you wish to submit for this challenge.</p>\n\n<p>Thanks again for a really exciting challenge, and best of luck for this final push!</p>\n\n<p>The PLAsTiCC Team</p>",
  "messages": [
    {
      "id": "439862",
      "postDate": "12/16/2018 14:38:23",
      "content": "<p>Hi Kagglers</p>\n\n<p>Thank you for all your hard work participating in the PLAsTiCC challenge! While the challenge closes on the 17th of December (tomorrow!), we wanted to bring your attention to the LSST Workshop part of the challenge, which remains open until January 15th 2019 (see details: <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018#Prizes\">https://www.kaggle.com/c/PLAsTiCC-2018#Prizes</a>).</p>\n\n<p>As part of this challenge, we will be rewarding teams that have solutions/algorithms of scientific interest in the following areas:</p>\n\n<ul>\n<li>codes that yield high efficiency at class 99-specific classifications</li>\n<li>novel methods that classify one type of object with high precision (purity) and high recall (completeness), particularly on objects where others have struggled</li>\n<li>models that can generate new light curves from the existing data </li>\n<li>methods that work well on classifying light curves that have large flux uncertainties</li>\n</ul>\n\n<p>In order to enter the LSST Workshop part of the challenge, we ask that you submit a Google form here: <a href=\"https://goo.gl/forms/EjzSeBE1tYPHGVCx1\">https://goo.gl/forms/EjzSeBE1tYPHGVCx1</a> with details of the code you wish to submit for this challenge.</p>\n\n<p>Thanks again for a really exciting challenge, and best of luck for this final push!</p>\n\n<p>The PLAsTiCC Team</p>",
      "rawMarkdown": "Hi Kagglers\n\nThank you for all your hard work participating in the PLAsTiCC challenge! While the challenge closes on the 17th of December (tomorrow!), we wanted to bring your attention to the LSST Workshop part of the challenge, which remains open until January 15th 2019 (see details: https://www.kaggle.com/c/PLAsTiCC-2018#Prizes).\n\nAs part of this challenge, we will be rewarding teams that have solutions/algorithms of scientific interest in the following areas:\n\n - codes that yield high efficiency at class 99-specific classifications\n - novel methods that classify one type of object with high precision (purity) and high recall (completeness), particularly on objects where others have struggled\n - models that can generate new light curves from the existing data \n - methods that work well on classifying light curves that have large flux uncertainties\n\nIn order to enter the LSST Workshop part of the challenge, we ask that you submit a Google form here: https://goo.gl/forms/EjzSeBE1tYPHGVCx1 with details of the code you wish to submit for this challenge.\n\nThanks again for a really exciting challenge, and best of luck for this final push!\n\nThe PLAsTiCC Team",
      "votes": null
    },
    {
      "id": "439971",
      "postDate": "12/16/2018 18:36:29",
      "content": "<blockquote>\n  <p>codes that yield high efficiency at class 99-specific classifications</p>\n</blockquote>\n\n<p>Does this mean you will disclose target for test data after the competition ends?</p>",
      "rawMarkdown": "&gt; codes that yield high efficiency at class 99-specific classifications\n\nDoes this mean you will disclose target for test data after the competition ends?",
      "votes": null
    },
    {
      "id": "440604",
      "postDate": "12/17/2018 19:06:14",
      "content": "<p>Yes, we plan on releasing the data post challenge (i.e. after Jan 15th). </p>",
      "rawMarkdown": "Yes, we plan on releasing the data post challenge (i.e. after Jan 15th).",
      "votes": null
    },
    {
      "id": "441015",
      "postDate": "12/18/2018 07:25:03",
      "content": "<p>Then how can we know about this:</p>\n\n<blockquote>\n  <p>codes that yield high efficiency at class 99-specific classifications</p>\n</blockquote>",
      "rawMarkdown": "Then how can we know about this:\n&gt; codes that yield high efficiency at class 99-specific classifications",
      "votes": null
    },
    {
      "id": "441355",
      "postDate": "12/18/2018 15:23:03",
      "content": "<p>Kaggle is set up purely for supervised classification challenges, but the real research questions we work on involve a more complicated combination of problems. Astronomers have tried unsupervised learning - anomaly detection and clustering techniques - for finding objects we've never seen before. We're interested to see what techniques the Kaggle community can come up with. </p>\n\n<p>It's an open research question, and yes, you don't get the truth tables aren't provided for participants to measure their performance. That said, there's already information to help design something like this. We've said that none of the class 99 objects are in the training set. People have found ways to augment the training set, and people have tried to figure out what classes in the training set are likely to be confused with class 99.  </p>\n\n<p>That's the same starting point astronomers have!</p>\n\n<p>Cheers,</p>\n\n<p>-Gautham for the PLAsTiCC team</p>",
      "rawMarkdown": "Kaggle is set up purely for supervised classification challenges, but the real research questions we work on involve a more complicated combination of problems. Astronomers have tried unsupervised learning - anomaly detection and clustering techniques - for finding objects we've never seen before. We're interested to see what techniques the Kaggle community can come up with. \n\nIt's an open research question, and yes, you don't get the truth tables aren't provided for participants to measure their performance. That said, there's already information to help design something like this. We've said that none of the class 99 objects are in the training set. People have found ways to augment the training set, and people have tried to figure out what classes in the training set are likely to be confused with class 99.  \n\nThat's the same starting point astronomers have!\n\nCheers,\n\n-Gautham for the PLAsTiCC team",
      "votes": null
    },
    {
      "id": "441421",
      "postDate": "12/18/2018 16:40:54",
      "content": "<p>I am not clear enough obviously.  You say to submit code that are good at predicting class 99 but how could we know that given you don't give feedback on how our code do it, nor do you provide the ground truth so that we evaluate ourselves?  </p>",
      "rawMarkdown": "I am not clear enough obviously.  You say to submit code that are good at predicting class 99 but how could we know that given you don't give feedback on how our code do it, nor do you provide the ground truth so that we evaluate ourselves?",
      "votes": null
    },
    {
      "id": "441472",
      "postDate": "12/18/2018 17:28:46",
      "content": "<p>The idea here isn't to submit multiple entries and get constant feedback as you do with the LB to optimize, but rather use all of your familiarity with the data gained through the contest to sketch a method that seems reasonable approach to try.</p>\n\n<p>For example, and I'm not in any way endorsing this as something to try, but as a possible entry:</p>\n\n<ol>\n<li>there is already ground truth for all the other classes but 99 in the training set</li>\n<li>people have come up with various ways to augment the training set </li>\n<li>one could use an augmented training set to classify the test set into everything but class 99, and remove the highest confidence events, adding them into an expanded training set</li>\n<li>see if there are objects that remain that have systematically low scores for all classes (even if they sum to 1) or are being classified with a high possibility of a few different classes </li>\n<li>try clustering these, and maybe using visual inspection to see if you can find features that discriminate between them and the other classes</li>\n</ol>\n\n<p>or who knows, maybe the GP modeling you've been trying can be applied to the test set, and you can cluster on GP interpolated light curves directly. Maybe that could work on the well-sampled DDF light curves. </p>\n\n<p>There's no shortage of methods to try, and that's what we want to reward --- <em>novel methods</em> --- approaches we didn't think about. Our sense was that the leaderboard rewarded one metric and it helped answers one big question, but there are plenty of other questions that are interesting scientifically and are more open-ended than can be encapsulated in Kaggle's framework. The nature of these questions means that you don't get feedback and there is no ground-truth. That's OK! </p>\n\n<p>You folks came up with really wonderful and creative ideas. We spent a good deal of time on slack last night after the contest ended discussing methods as contestants posted on the discussion forum about them, making plots for ourselves, and being absolutely blown away by the amount of complexity that many people managed to incorporate in such a short amount of time. We've learned a ton from this contest, and there's going to be a lot of work to do now that it is over to really make use of some of the new ideas. </p>\n\n<p>The workshop challenge is another avenue to be rewarded for some of those ideas, and you should feel free to experiment without the constraints of the leaderboard score. We're open to writing papers about these methods too (in addition to the rewards) with contributors as co-authors. </p>\n\n<p>Again, many thanks for all the hard work, and our best wishes,</p>\n\n<p>-Gautham for the PLAsTiCC team</p>",
      "rawMarkdown": "The idea here isn't to submit multiple entries and get constant feedback as you do with the LB to optimize, but rather use all of your familiarity with the data gained through the contest to sketch a method that seems reasonable approach to try.\n\nFor example, and I'm not in any way endorsing this as something to try, but as a possible entry:\n\n 1. there is already ground truth for all the other classes but 99 in the training set\n 2. people have come up with various ways to augment the training set \n 3. one could use an augmented training set to classify the test set into everything but class 99, and remove the highest confidence events, adding them into an expanded training set\n 4. see if there are objects that remain that have systematically low scores for all classes (even if they sum to 1) or are being classified with a high possibility of a few different classes \n 5. try clustering these, and maybe using visual inspection to see if you can find features that discriminate between them and the other classes\n\nor who knows, maybe the GP modeling you've been trying can be applied to the test set, and you can cluster on GP interpolated light curves directly. Maybe that could work on the well-sampled DDF light curves. \n\nThere's no shortage of methods to try, and that's what we want to reward --- *novel methods* --- approaches we didn't think about. Our sense was that the leaderboard rewarded one metric and it helped answers one big question, but there are plenty of other questions that are interesting scientifically and are more open-ended than can be encapsulated in Kaggle's framework. The nature of these questions means that you don't get feedback and there is no ground-truth. That's OK! \n\nYou folks came up with really wonderful and creative ideas. We spent a good deal of time on slack last night after the contest ended discussing methods as contestants posted on the discussion forum about them, making plots for ourselves, and being absolutely blown away by the amount of complexity that many people managed to incorporate in such a short amount of time. We've learned a ton from this contest, and there's going to be a lot of work to do now that it is over to really make use of some of the new ideas. \n\nThe workshop challenge is another avenue to be rewarded for some of those ideas, and you should feel free to experiment without the constraints of the leaderboard score. We're open to writing papers about these methods too (in addition to the rewards) with contributors as co-authors. \n\nAgain, many thanks for all the hard work, and our best wishes,\n\n-Gautham for the PLAsTiCC team",
      "votes": null
    },
    {
      "id": "441490",
      "postDate": "12/18/2018 17:49:01",
      "content": "<p>Thanks, makes sense.  Your dataset is one of the best I worked with. </p>",
      "rawMarkdown": "Thanks, makes sense.  Your dataset is one of the best I worked with.",
      "votes": null
    },
    {
      "id": "441541",
      "postDate": "12/18/2018 19:12:21",
      "content": "<p>I think that designing a good class 99 model would be a lot easier if we had a test set to evaluate our performance. We don't necessarily need to know what is a class 99 object, just some score on how the class 99 model does. Obviously people could tune models by iterating on the test set (like a lot of us did), but if you have a separate challenge, you can look at the code and figure out who did real predictions and who just tuned the model.</p>\n\n<p>That being said, I recently figured out how to measure your model's performance on class 99s by crafting special submissions. We know that each object is galactic or extragalactic, and that there is no chance of it crossing that boundary. To test class 99 performance, classify the 99s as you normally would, and then put known wrong predictions (a galactic class if extragalactic and vice versa) for the other classes. Because of the metric that was used, you will get a penalty of 16/18*log(1e-15) for the non-class 99 objects while maintaining your class 99 predictions. If you subtract off that penalty, you will get your class 99 score.</p>\n\n<p>I played around with this, and found that for the extragalactic class 99s, they are all supernova-like objects that look mostly like classes 42/62/52/95.</p>",
      "rawMarkdown": "I think that designing a good class 99 model would be a lot easier if we had a test set to evaluate our performance. We don't necessarily need to know what is a class 99 object, just some score on how the class 99 model does. Obviously people could tune models by iterating on the test set (like a lot of us did), but if you have a separate challenge, you can look at the code and figure out who did real predictions and who just tuned the model.\n\nThat being said, I recently figured out how to measure your model's performance on class 99s by crafting special submissions. We know that each object is galactic or extragalactic, and that there is no chance of it crossing that boundary. To test class 99 performance, classify the 99s as you normally would, and then put known wrong predictions (a galactic class if extragalactic and vice versa) for the other classes. Because of the metric that was used, you will get a penalty of 16/18*log(1e-15) for the non-class 99 objects while maintaining your class 99 predictions. If you subtract off that penalty, you will get your class 99 score.\n\nI played around with this, and found that for the extragalactic class 99s, they are all supernova-like objects that look mostly like classes 42/62/52/95.",
      "votes": null
    },
    {
      "id": "441553",
      "postDate": "12/18/2018 19:35:04",
      "content": "<blockquote>\n  <p>they are all supernova-like objects that look mostly like classes 42/62/52/95.</p>\n</blockquote>\n\n<p>Thanks, this explained one of my failed attempts.  </p>",
      "rawMarkdown": "&gt;  they are all supernova-like objects that look mostly like classes 42/62/52/95.\n\nThanks, this explained one of my failed attempts.",
      "votes": null
    },
    {
      "id": "441579",
      "postDate": "12/18/2018 20:00:28",
      "content": "<p>What is the deadline to submit code? Is that Jan. 15, or is that when you will announce the results? I need to clean up my code to make it usable by other people. With the holidays coming up, the timing is a bit tough.</p>",
      "rawMarkdown": "What is the deadline to submit code? Is that Jan. 15, or is that when you will announce the results? I need to clean up my code to make it usable by other people. With the holidays coming up, the timing is a bit tough.",
      "votes": null
    },
    {
      "id": "441639",
      "postDate": "12/18/2018 21:59:40",
      "content": "<p>Jan 15th is the deadline to enter. If it's any consolation, all the excellent work on this contest means that the PLAsTiCC team will also be working over these holidays, before AAS!</p>",
      "rawMarkdown": "Jan 15th is the deadline to enter. If it's any consolation, all the excellent work on this contest means that the PLAsTiCC team will also be working over these holidays, before AAS!",
      "votes": null
    },
    {
      "id": "442027",
      "postDate": "12/19/2018 11:40:20",
      "content": "<p>Had exactly the same problem... I initially reasoned that the Product(1-p) method should not work well because it is mostly measuring inter-supernova uncertainties while class 99 might be something arbitrary (I thought it could be generated from some signal sources unlike any natural objects, like something SETI might be interested in), but if it turns out that most class 99 are indeed also supernova classes, it explains the unreasonable success of the simple formula, because areas of high supernova confusion is exactly where class 99 is likely to hide.</p>",
      "rawMarkdown": "Had exactly the same problem... I initially reasoned that the Product(1-p) method should not work well because it is mostly measuring inter-supernova uncertainties while class 99 might be something arbitrary (I thought it could be generated from some signal sources unlike any natural objects, like something SETI might be interested in), but if it turns out that most class 99 are indeed also supernova classes, it explains the unreasonable success of the simple formula, because areas of high supernova confusion is exactly where class 99 is likely to hide.",
      "votes": null
    },
    {
      "id": "442261",
      "postDate": "12/19/2018 17:41:08",
      "content": "<p>I'm a bit confused regarding the format of the workshop. I read this timeline:</p>\n\n<pre><code>December 10, 2018 - Entry deadline. You must accept the competition rules before this date in order to compete.\n\nDecember 10, 2018 - Team Merger deadline. This is the last day participants may join or merge teams.\n\nDecember 17, 2018 - Final submission deadline.\n\nJanuary 15, 2019 - LSST Workshop entry deadline.\n\nFebruary 15, 2019 - LSST Workshop announcement.\n</code></pre>\n\n<p>All deadlines are at 11:59 PM UTC on the corresponding day unless otherwise noted. The competition organizers reserve the right to update the contest timeline if they deem it necessary.</p>\n\n<p>Is the LSST Workshop an official part of this kaggle competition, a separate kaggle competition, or something different altogether? Is January 15, 2019 the last day to submit code, or is February 15, 2019? Or is February 15, 2019 the start of the workshop?</p>",
      "rawMarkdown": "I'm a bit confused regarding the format of the workshop. I read this timeline:\n\n    December 10, 2018 - Entry deadline. You must accept the competition rules before this date in order to compete.\n\n    December 10, 2018 - Team Merger deadline. This is the last day participants may join or merge teams.\n\n    December 17, 2018 - Final submission deadline.\n\n    January 15, 2019 - LSST Workshop entry deadline.\n\n    February 15, 2019 - LSST Workshop announcement.\n\nAll deadlines are at 11:59 PM UTC on the corresponding day unless otherwise noted. The competition organizers reserve the right to update the contest timeline if they deem it necessary.\n\nIs the LSST Workshop an official part of this kaggle competition, a separate kaggle competition, or something different altogether? Is January 15, 2019 the last day to submit code, or is February 15, 2019? Or is February 15, 2019 the start of the workshop?",
      "votes": null
    },
    {
      "id": "442313",
      "postDate": "12/19/2018 19:28:18",
      "content": "<blockquote>\n  <p>If you subtract off that penalty, you will get your class 99 score.</p>\n</blockquote>\n\n<p>What a simple and clever method to get class 99 score!!!</p>",
      "rawMarkdown": "&gt; If you subtract off that penalty, you will get your class 99 score.\n\nWhat a simple and clever method to get class 99 score!!!",
      "votes": null
    },
    {
      "id": "443029",
      "postDate": "12/20/2018 22:53:27",
      "content": "<p>The workshop competition is separate from the Kaggle competition. Submissions in the work of a Github link to your group's working code are due Jan 15th, 2019, and winners will be announced Feb 15, on this forum as well as through email (submitted along with the Github link through the google doc in Renee's post).</p>",
      "rawMarkdown": "The workshop competition is separate from the Kaggle competition. Submissions in the work of a Github link to your group's working code are due Jan 15th, 2019, and winners will be announced Feb 15, on this forum as well as through email (submitted along with the Github link through the google doc in Renee's post).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 439971,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "12/16/2018 18:36:29",
      "content": "<blockquote>\n  <p>codes that yield high efficiency at class 99-specific classifications</p>\n</blockquote>\n\n<p>Does this mean you will disclose target for test data after the competition ends?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 440604,
      "author_name": "reneehlozek",
      "author_url": "",
      "post_date": "12/17/2018 19:06:14",
      "content": "<p>Yes, we plan on releasing the data post challenge (i.e. after Jan 15th). </p>",
      "votes": null,
      "replies": [
        {
          "id": 441015,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/18/2018 07:25:03",
          "content": "<p>Then how can we know about this:</p>\n\n<blockquote>\n  <p>codes that yield high efficiency at class 99-specific classifications</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441355,
          "author_name": "gsnarayan",
          "author_url": "",
          "post_date": "12/18/2018 15:23:03",
          "content": "<p>Kaggle is set up purely for supervised classification challenges, but the real research questions we work on involve a more complicated combination of problems. Astronomers have tried unsupervised learning - anomaly detection and clustering techniques - for finding objects we've never seen before. We're interested to see what techniques the Kaggle community can come up with. </p>\n\n<p>It's an open research question, and yes, you don't get the truth tables aren't provided for participants to measure their performance. That said, there's already information to help design something like this. We've said that none of the class 99 objects are in the training set. People have found ways to augment the training set, and people have tried to figure out what classes in the training set are likely to be confused with class 99.  </p>\n\n<p>That's the same starting point astronomers have!</p>\n\n<p>Cheers,</p>\n\n<p>-Gautham for the PLAsTiCC team</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441421,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/18/2018 16:40:54",
          "content": "<p>I am not clear enough obviously.  You say to submit code that are good at predicting class 99 but how could we know that given you don't give feedback on how our code do it, nor do you provide the ground truth so that we evaluate ourselves?  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441472,
          "author_name": "gsnarayan",
          "author_url": "",
          "post_date": "12/18/2018 17:28:46",
          "content": "<p>The idea here isn't to submit multiple entries and get constant feedback as you do with the LB to optimize, but rather use all of your familiarity with the data gained through the contest to sketch a method that seems reasonable approach to try.</p>\n\n<p>For example, and I'm not in any way endorsing this as something to try, but as a possible entry:</p>\n\n<ol>\n<li>there is already ground truth for all the other classes but 99 in the training set</li>\n<li>people have come up with various ways to augment the training set </li>\n<li>one could use an augmented training set to classify the test set into everything but class 99, and remove the highest confidence events, adding them into an expanded training set</li>\n<li>see if there are objects that remain that have systematically low scores for all classes (even if they sum to 1) or are being classified with a high possibility of a few different classes </li>\n<li>try clustering these, and maybe using visual inspection to see if you can find features that discriminate between them and the other classes</li>\n</ol>\n\n<p>or who knows, maybe the GP modeling you've been trying can be applied to the test set, and you can cluster on GP interpolated light curves directly. Maybe that could work on the well-sampled DDF light curves. </p>\n\n<p>There's no shortage of methods to try, and that's what we want to reward --- <em>novel methods</em> --- approaches we didn't think about. Our sense was that the leaderboard rewarded one metric and it helped answers one big question, but there are plenty of other questions that are interesting scientifically and are more open-ended than can be encapsulated in Kaggle's framework. The nature of these questions means that you don't get feedback and there is no ground-truth. That's OK! </p>\n\n<p>You folks came up with really wonderful and creative ideas. We spent a good deal of time on slack last night after the contest ended discussing methods as contestants posted on the discussion forum about them, making plots for ourselves, and being absolutely blown away by the amount of complexity that many people managed to incorporate in such a short amount of time. We've learned a ton from this contest, and there's going to be a lot of work to do now that it is over to really make use of some of the new ideas. </p>\n\n<p>The workshop challenge is another avenue to be rewarded for some of those ideas, and you should feel free to experiment without the constraints of the leaderboard score. We're open to writing papers about these methods too (in addition to the rewards) with contributors as co-authors. </p>\n\n<p>Again, many thanks for all the hard work, and our best wishes,</p>\n\n<p>-Gautham for the PLAsTiCC team</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441490,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/18/2018 17:49:01",
          "content": "<p>Thanks, makes sense.  Your dataset is one of the best I worked with. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441541,
          "author_name": "kyleboone",
          "author_url": "",
          "post_date": "12/18/2018 19:12:21",
          "content": "<p>I think that designing a good class 99 model would be a lot easier if we had a test set to evaluate our performance. We don't necessarily need to know what is a class 99 object, just some score on how the class 99 model does. Obviously people could tune models by iterating on the test set (like a lot of us did), but if you have a separate challenge, you can look at the code and figure out who did real predictions and who just tuned the model.</p>\n\n<p>That being said, I recently figured out how to measure your model's performance on class 99s by crafting special submissions. We know that each object is galactic or extragalactic, and that there is no chance of it crossing that boundary. To test class 99 performance, classify the 99s as you normally would, and then put known wrong predictions (a galactic class if extragalactic and vice versa) for the other classes. Because of the metric that was used, you will get a penalty of 16/18*log(1e-15) for the non-class 99 objects while maintaining your class 99 predictions. If you subtract off that penalty, you will get your class 99 score.</p>\n\n<p>I played around with this, and found that for the extragalactic class 99s, they are all supernova-like objects that look mostly like classes 42/62/52/95.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441553,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/18/2018 19:35:04",
          "content": "<blockquote>\n  <p>they are all supernova-like objects that look mostly like classes 42/62/52/95.</p>\n</blockquote>\n\n<p>Thanks, this explained one of my failed attempts.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442027,
          "author_name": "mithrillion",
          "author_url": "",
          "post_date": "12/19/2018 11:40:20",
          "content": "<p>Had exactly the same problem... I initially reasoned that the Product(1-p) method should not work well because it is mostly measuring inter-supernova uncertainties while class 99 might be something arbitrary (I thought it could be generated from some signal sources unlike any natural objects, like something SETI might be interested in), but if it turns out that most class 99 are indeed also supernova classes, it explains the unreasonable success of the simple formula, because areas of high supernova confusion is exactly where class 99 is likely to hide.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442313,
          "author_name": "sergeyzlobin",
          "author_url": "",
          "post_date": "12/19/2018 19:28:18",
          "content": "<blockquote>\n  <p>If you subtract off that penalty, you will get your class 99 score.</p>\n</blockquote>\n\n<p>What a simple and clever method to get class 99 score!!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 441579,
      "author_name": "kyleboone",
      "author_url": "",
      "post_date": "12/18/2018 20:00:28",
      "content": "<p>What is the deadline to submit code? Is that Jan. 15, or is that when you will announce the results? I need to clean up my code to make it usable by other people. With the holidays coming up, the timing is a bit tough.</p>",
      "votes": null,
      "replies": [
        {
          "id": 441639,
          "author_name": "gsnarayan",
          "author_url": "",
          "post_date": "12/18/2018 21:59:40",
          "content": "<p>Jan 15th is the deadline to enter. If it's any consolation, all the excellent work on this contest means that the PLAsTiCC team will also be working over these holidays, before AAS!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 442261,
      "author_name": "d8an1nj4",
      "author_url": "",
      "post_date": "12/19/2018 17:41:08",
      "content": "<p>I'm a bit confused regarding the format of the workshop. I read this timeline:</p>\n\n<pre><code>December 10, 2018 - Entry deadline. You must accept the competition rules before this date in order to compete.\n\nDecember 10, 2018 - Team Merger deadline. This is the last day participants may join or merge teams.\n\nDecember 17, 2018 - Final submission deadline.\n\nJanuary 15, 2019 - LSST Workshop entry deadline.\n\nFebruary 15, 2019 - LSST Workshop announcement.\n</code></pre>\n\n<p>All deadlines are at 11:59 PM UTC on the corresponding day unless otherwise noted. The competition organizers reserve the right to update the contest timeline if they deem it necessary.</p>\n\n<p>Is the LSST Workshop an official part of this kaggle competition, a separate kaggle competition, or something different altogether? Is January 15, 2019 the last day to submit code, or is February 15, 2019? Or is February 15, 2019 the start of the workshop?</p>",
      "votes": null,
      "replies": [
        {
          "id": 443029,
          "author_name": "gsnarayan",
          "author_url": "",
          "post_date": "12/20/2018 22:53:27",
          "content": "<p>The workshop competition is separate from the Kaggle competition. Submissions in the work of a Github link to your group's working code are due Jan 15th, 2019, and winners will be announced Feb 15, on this forum as well as through email (submitted along with the Github link through the google doc in Renee's post).</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "439862": "Hi Kagglers\n\nThank you for all your hard work participating in the PLAsTiCC challenge! While the challenge closes on the 17th of December (tomorrow!), we wanted to bring your attention to the LSST Workshop part of the challenge, which remains open until January 15th 2019 (see details: https://www.kaggle.com/c/PLAsTiCC-2018#Prizes).\n\nAs part of this challenge, we will be rewarding teams that have solutions/algorithms of scientific interest in the following areas:\n\n - codes that yield high efficiency at class 99-specific classifications\n - novel methods that classify one type of object with high precision (purity) and high recall (completeness), particularly on objects where others have struggled\n - models that can generate new light curves from the existing data \n - methods that work well on classifying light curves that have large flux uncertainties\n\nIn order to enter the LSST Workshop part of the challenge, we ask that you submit a Google form here: https://goo.gl/forms/EjzSeBE1tYPHGVCx1 with details of the code you wish to submit for this challenge.\n\nThanks again for a really exciting challenge, and best of luck for this final push!\n\nThe PLAsTiCC Team",
    "439971": "&gt; codes that yield high efficiency at class 99-specific classifications\n\nDoes this mean you will disclose target for test data after the competition ends?",
    "440604": "Yes, we plan on releasing the data post challenge (i.e. after Jan 15th).",
    "441015": "Then how can we know about this:\n&gt; codes that yield high efficiency at class 99-specific classifications",
    "441355": "Kaggle is set up purely for supervised classification challenges, but the real research questions we work on involve a more complicated combination of problems. Astronomers have tried unsupervised learning - anomaly detection and clustering techniques - for finding objects we've never seen before. We're interested to see what techniques the Kaggle community can come up with. \n\nIt's an open research question, and yes, you don't get the truth tables aren't provided for participants to measure their performance. That said, there's already information to help design something like this. We've said that none of the class 99 objects are in the training set. People have found ways to augment the training set, and people have tried to figure out what classes in the training set are likely to be confused with class 99.  \n\nThat's the same starting point astronomers have!\n\nCheers,\n\n-Gautham for the PLAsTiCC team",
    "441421": "I am not clear enough obviously.  You say to submit code that are good at predicting class 99 but how could we know that given you don't give feedback on how our code do it, nor do you provide the ground truth so that we evaluate ourselves?",
    "441472": "The idea here isn't to submit multiple entries and get constant feedback as you do with the LB to optimize, but rather use all of your familiarity with the data gained through the contest to sketch a method that seems reasonable approach to try.\n\nFor example, and I'm not in any way endorsing this as something to try, but as a possible entry:\n\n 1. there is already ground truth for all the other classes but 99 in the training set\n 2. people have come up with various ways to augment the training set \n 3. one could use an augmented training set to classify the test set into everything but class 99, and remove the highest confidence events, adding them into an expanded training set\n 4. see if there are objects that remain that have systematically low scores for all classes (even if they sum to 1) or are being classified with a high possibility of a few different classes \n 5. try clustering these, and maybe using visual inspection to see if you can find features that discriminate between them and the other classes\n\nor who knows, maybe the GP modeling you've been trying can be applied to the test set, and you can cluster on GP interpolated light curves directly. Maybe that could work on the well-sampled DDF light curves. \n\nThere's no shortage of methods to try, and that's what we want to reward --- *novel methods* --- approaches we didn't think about. Our sense was that the leaderboard rewarded one metric and it helped answers one big question, but there are plenty of other questions that are interesting scientifically and are more open-ended than can be encapsulated in Kaggle's framework. The nature of these questions means that you don't get feedback and there is no ground-truth. That's OK! \n\nYou folks came up with really wonderful and creative ideas. We spent a good deal of time on slack last night after the contest ended discussing methods as contestants posted on the discussion forum about them, making plots for ourselves, and being absolutely blown away by the amount of complexity that many people managed to incorporate in such a short amount of time. We've learned a ton from this contest, and there's going to be a lot of work to do now that it is over to really make use of some of the new ideas. \n\nThe workshop challenge is another avenue to be rewarded for some of those ideas, and you should feel free to experiment without the constraints of the leaderboard score. We're open to writing papers about these methods too (in addition to the rewards) with contributors as co-authors. \n\nAgain, many thanks for all the hard work, and our best wishes,\n\n-Gautham for the PLAsTiCC team",
    "441490": "Thanks, makes sense.  Your dataset is one of the best I worked with.",
    "441541": "I think that designing a good class 99 model would be a lot easier if we had a test set to evaluate our performance. We don't necessarily need to know what is a class 99 object, just some score on how the class 99 model does. Obviously people could tune models by iterating on the test set (like a lot of us did), but if you have a separate challenge, you can look at the code and figure out who did real predictions and who just tuned the model.\n\nThat being said, I recently figured out how to measure your model's performance on class 99s by crafting special submissions. We know that each object is galactic or extragalactic, and that there is no chance of it crossing that boundary. To test class 99 performance, classify the 99s as you normally would, and then put known wrong predictions (a galactic class if extragalactic and vice versa) for the other classes. Because of the metric that was used, you will get a penalty of 16/18*log(1e-15) for the non-class 99 objects while maintaining your class 99 predictions. If you subtract off that penalty, you will get your class 99 score.\n\nI played around with this, and found that for the extragalactic class 99s, they are all supernova-like objects that look mostly like classes 42/62/52/95.",
    "441553": "&gt;  they are all supernova-like objects that look mostly like classes 42/62/52/95.\n\nThanks, this explained one of my failed attempts.",
    "441579": "What is the deadline to submit code? Is that Jan. 15, or is that when you will announce the results? I need to clean up my code to make it usable by other people. With the holidays coming up, the timing is a bit tough.",
    "441639": "Jan 15th is the deadline to enter. If it's any consolation, all the excellent work on this contest means that the PLAsTiCC team will also be working over these holidays, before AAS!",
    "442027": "Had exactly the same problem... I initially reasoned that the Product(1-p) method should not work well because it is mostly measuring inter-supernova uncertainties while class 99 might be something arbitrary (I thought it could be generated from some signal sources unlike any natural objects, like something SETI might be interested in), but if it turns out that most class 99 are indeed also supernova classes, it explains the unreasonable success of the simple formula, because areas of high supernova confusion is exactly where class 99 is likely to hide.",
    "442261": "I'm a bit confused regarding the format of the workshop. I read this timeline:\n\n    December 10, 2018 - Entry deadline. You must accept the competition rules before this date in order to compete.\n\n    December 10, 2018 - Team Merger deadline. This is the last day participants may join or merge teams.\n\n    December 17, 2018 - Final submission deadline.\n\n    January 15, 2019 - LSST Workshop entry deadline.\n\n    February 15, 2019 - LSST Workshop announcement.\n\nAll deadlines are at 11:59 PM UTC on the corresponding day unless otherwise noted. The competition organizers reserve the right to update the contest timeline if they deem it necessary.\n\nIs the LSST Workshop an official part of this kaggle competition, a separate kaggle competition, or something different altogether? Is January 15, 2019 the last day to submit code, or is February 15, 2019? Or is February 15, 2019 the start of the workshop?",
    "442313": "&gt; If you subtract off that penalty, you will get your class 99 score.\n\nWhat a simple and clever method to get class 99 score!!!",
    "443029": "The workshop competition is separate from the Kaggle competition. Submissions in the work of a Github link to your group's working code are due Jan 15th, 2019, and winners will be announced Feb 15, on this forum as well as through email (submitted along with the Github link through the google doc in Renee's post)."
  },
  "source": "meta"
}