{
  "id": 78366,
  "title": "Results on the unblinded test set",
  "url": "/competitions/PLAsTiCC-2018/discussion/78366",
  "author_name": "",
  "post_date": "2019-01-22T23:12:02.385036700Z",
  "votes": 26,
  "comment_count": 2,
  "views": 0,
  "content": "<p>For those who haven't noticed, the PLAsTiCC team released the labels for the test set today and they can be found at: <a href=\"https://plasticc.org/data-release/\">https://plasticc.org/data-release/</a></p>\n\n<p>Here are some quick tests that I have done with these labels using the predictions that the top 5 teams released in this thread: <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75179\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75179</a></p>\n\n<p>I did all of my tests with the class 99 objects removed as nobody really did a proper prediction for these objects and instead just did different kinds of leaderboard probing. With the class 99 predictions removed and everything renormalized, we find that there is a bit of a shakeup of the leaderboard with Mike &amp; Silogram outperforming everyone else:</p>\n\n<ul>\n<li>Silogram &amp; Mike: 0.39777</li>\n<li>Kyle Boone: 0.43164</li>\n<li>AhmetErdem: 0.43984</li>\n<li>Major Tom: 0.44529</li>\n<li>CPMP: 0.45112</li>\n</ul>\n\n<hr>\n\n<ul>\n<li>Mamas' top 5 combination: 0.36937</li>\n</ul>\n\n<p>Another way to look at this is how much each object contributed to the final score. Here are the contributions to the final metric for each object multiplied by 100 (if you add up every row and divide by 100 you'll get the numbers listed above):</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/460066/11085/metric_contributions.png\" alt=\"metric contributions\"></p>\n\n<p>Classes 42, 52 and 62 were definitely the hardest to predict. I personally focused on classifying the different kinds of supernovae, and you can see that I did quite a bit worse than the competition on the Class 6 objects (TDEs) and class 88 objects (AGN). Some of the differences, like me doing better on classes 42 and 62 but worse on class 52 than the competition, are highly correlated and really just amount to effectively using slightly different weights for the different classes rather than anything real.</p>\n\n<p>Here is my confusion matrix for the test set with class 99 objects removed:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/460066/11087/confusion_matrix_kyle.png\" alt=\"kyle confusion matrix\"></p>\n\n<p>You can see that my model did not do very well on the class 52 objects which is likely due to it overfitting the small training set. I have attached the confusion matrices for the other top models as well.</p>",
  "messages": [
    {
      "id": "460066",
      "postDate": "01/22/2019 23:12:02",
      "content": "<p>For those who haven't noticed, the PLAsTiCC team released the labels for the test set today and they can be found at: <a href=\"https://plasticc.org/data-release/\">https://plasticc.org/data-release/</a></p>\n\n<p>Here are some quick tests that I have done with these labels using the predictions that the top 5 teams released in this thread: <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75179\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75179</a></p>\n\n<p>I did all of my tests with the class 99 objects removed as nobody really did a proper prediction for these objects and instead just did different kinds of leaderboard probing. With the class 99 predictions removed and everything renormalized, we find that there is a bit of a shakeup of the leaderboard with Mike &amp; Silogram outperforming everyone else:</p>\n\n<ul>\n<li>Silogram &amp; Mike: 0.39777</li>\n<li>Kyle Boone: 0.43164</li>\n<li>AhmetErdem: 0.43984</li>\n<li>Major Tom: 0.44529</li>\n<li>CPMP: 0.45112</li>\n</ul>\n\n<hr>\n\n<ul>\n<li>Mamas' top 5 combination: 0.36937</li>\n</ul>\n\n<p>Another way to look at this is how much each object contributed to the final score. Here are the contributions to the final metric for each object multiplied by 100 (if you add up every row and divide by 100 you'll get the numbers listed above):</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/460066/11085/metric_contributions.png\" alt=\"metric contributions\"></p>\n\n<p>Classes 42, 52 and 62 were definitely the hardest to predict. I personally focused on classifying the different kinds of supernovae, and you can see that I did quite a bit worse than the competition on the Class 6 objects (TDEs) and class 88 objects (AGN). Some of the differences, like me doing better on classes 42 and 62 but worse on class 52 than the competition, are highly correlated and really just amount to effectively using slightly different weights for the different classes rather than anything real.</p>\n\n<p>Here is my confusion matrix for the test set with class 99 objects removed:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/460066/11087/confusion_matrix_kyle.png\" alt=\"kyle confusion matrix\"></p>\n\n<p>You can see that my model did not do very well on the class 52 objects which is likely due to it overfitting the small training set. I have attached the confusion matrices for the other top models as well.</p>",
      "rawMarkdown": "For those who haven't noticed, the PLAsTiCC team released the labels for the test set today and they can be found at: https://plasticc.org/data-release/\n\nHere are some quick tests that I have done with these labels using the predictions that the top 5 teams released in this thread: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75179\n\nI did all of my tests with the class 99 objects removed as nobody really did a proper prediction for these objects and instead just did different kinds of leaderboard probing. With the class 99 predictions removed and everything renormalized, we find that there is a bit of a shakeup of the leaderboard with Mike &amp; Silogram outperforming everyone else:\n\n- Silogram &amp; Mike: 0.39777\n- Kyle Boone: 0.43164\n- AhmetErdem: 0.43984\n- Major Tom: 0.44529\n- CPMP: 0.45112\n\n---\n\n- Mamas' top 5 combination: 0.36937\n\nAnother way to look at this is how much each object contributed to the final score. Here are the contributions to the final metric for each object multiplied by 100 (if you add up every row and divide by 100 you'll get the numbers listed above):\n\n![metric contributions][1]\n\nClasses 42, 52 and 62 were definitely the hardest to predict. I personally focused on classifying the different kinds of supernovae, and you can see that I did quite a bit worse than the competition on the Class 6 objects (TDEs) and class 88 objects (AGN). Some of the differences, like me doing better on classes 42 and 62 but worse on class 52 than the competition, are highly correlated and really just amount to effectively using slightly different weights for the different classes rather than anything real.\n\nHere is my confusion matrix for the test set with class 99 objects removed:\n\n![kyle confusion matrix][2]\n\nYou can see that my model did not do very well on the class 52 objects which is likely due to it overfitting the small training set. I have attached the confusion matrices for the other top models as well.\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/460066/11085/metric_contributions.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/460066/11087/confusion_matrix_kyle.png",
      "votes": null
    },
    {
      "id": "460093",
      "postDate": "01/23/2019 01:47:10",
      "content": "<p>Kyle, Thanks for the interesting post! \nI've already noticed Silogram &amp; Mike should have been actually 1st without class99 using some tests, but I'm surprised at the big difference between Silogram &amp; Mike and Kyle. Their models are terribly good!</p>",
      "rawMarkdown": "Kyle, Thanks for the interesting post! \nI've already noticed Silogram &amp; Mike should have been actually 1st without class99 using some tests, but I'm surprised at the big difference between Silogram &amp; Mike and Kyle. Their models are terribly good!",
      "votes": null
    },
    {
      "id": "614486",
      "postDate": "08/31/2019 14:55:51",
      "content": "<p>Hello, Kyle!\nI tried to do the same things. I downloaded some of these subs and got another results. \nFor example, \nwith class 99 your final submission gives 0.6802563147863133\nLooking at LB that's right, I think. I used weight = 2 for class 99.\nBut when I removed 99 class objects (and renormalized probabalities), I got 0.4855977044371161. You got 0.43164.\nI attach my script.</p>",
      "rawMarkdown": "Hello, Kyle!\nI tried to do the same things. I downloaded some of these subs and got another results. \nFor example, \nwith class 99 your final submission gives 0.6802563147863133\nLooking at LB that's right, I think. I used weight = 2 for class 99.\nBut when I removed 99 class objects (and renormalized probabalities), I got 0.4855977044371161. You got 0.43164.\nI attach my script.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 460093,
      "author_name": "mamasinkgs",
      "author_url": "",
      "post_date": "01/23/2019 01:47:10",
      "content": "<p>Kyle, Thanks for the interesting post! \nI've already noticed Silogram &amp; Mike should have been actually 1st without class99 using some tests, but I'm surprised at the big difference between Silogram &amp; Mike and Kyle. Their models are terribly good!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 614486,
      "author_name": "sergeyzlobin",
      "author_url": "",
      "post_date": "08/31/2019 14:55:51",
      "content": "<p>Hello, Kyle!\nI tried to do the same things. I downloaded some of these subs and got another results. \nFor example, \nwith class 99 your final submission gives 0.6802563147863133\nLooking at LB that's right, I think. I used weight = 2 for class 99.\nBut when I removed 99 class objects (and renormalized probabalities), I got 0.4855977044371161. You got 0.43164.\nI attach my script.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "460066": "For those who haven't noticed, the PLAsTiCC team released the labels for the test set today and they can be found at: https://plasticc.org/data-release/\n\nHere are some quick tests that I have done with these labels using the predictions that the top 5 teams released in this thread: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/75179\n\nI did all of my tests with the class 99 objects removed as nobody really did a proper prediction for these objects and instead just did different kinds of leaderboard probing. With the class 99 predictions removed and everything renormalized, we find that there is a bit of a shakeup of the leaderboard with Mike &amp; Silogram outperforming everyone else:\n\n- Silogram &amp; Mike: 0.39777\n- Kyle Boone: 0.43164\n- AhmetErdem: 0.43984\n- Major Tom: 0.44529\n- CPMP: 0.45112\n\n---\n\n- Mamas' top 5 combination: 0.36937\n\nAnother way to look at this is how much each object contributed to the final score. Here are the contributions to the final metric for each object multiplied by 100 (if you add up every row and divide by 100 you'll get the numbers listed above):\n\n![metric contributions][1]\n\nClasses 42, 52 and 62 were definitely the hardest to predict. I personally focused on classifying the different kinds of supernovae, and you can see that I did quite a bit worse than the competition on the Class 6 objects (TDEs) and class 88 objects (AGN). Some of the differences, like me doing better on classes 42 and 62 but worse on class 52 than the competition, are highly correlated and really just amount to effectively using slightly different weights for the different classes rather than anything real.\n\nHere is my confusion matrix for the test set with class 99 objects removed:\n\n![kyle confusion matrix][2]\n\nYou can see that my model did not do very well on the class 52 objects which is likely due to it overfitting the small training set. I have attached the confusion matrices for the other top models as well.\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/460066/11085/metric_contributions.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/460066/11087/confusion_matrix_kyle.png",
    "460093": "Kyle, Thanks for the interesting post! \nI've already noticed Silogram &amp; Mike should have been actually 1st without class99 using some tests, but I'm surprised at the big difference between Silogram &amp; Mike and Kyle. Their models are terribly good!",
    "614486": "Hello, Kyle!\nI tried to do the same things. I downloaded some of these subs and got another results. \nFor example, \nwith class 99 your final submission gives 0.6802563147863133\nLooking at LB that's right, I think. I used weight = 2 for class 99.\nBut when I removed 99 class objects (and renormalized probabalities), I got 0.4855977044371161. You got 0.43164.\nI attach my script."
  },
  "source": "meta"
}