{
  "id": 67170,
  "title": "Where are the class weights?",
  "url": "/competitions/PLAsTiCC-2018/discussion/67170",
  "author_name": "",
  "post_date": "2018-09-29T14:48:34.698464700Z",
  "votes": 10,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I can't find the class weights for the evaluation metric? or it calculated by the class frequencies? Also, how does weighted work for class-99 when we don't have any in the dataset. This means the metric doesn't hold true from train to test?</p>",
  "messages": [
    {
      "id": "395896",
      "postDate": "09/29/2018 14:48:34",
      "content": "<p>I can't find the class weights for the evaluation metric? or it calculated by the class frequencies? Also, how does weighted work for class-99 when we don't have any in the dataset. This means the metric doesn't hold true from train to test?</p>",
      "rawMarkdown": "I can't find the class weights for the evaluation metric? or it calculated by the class frequencies? Also, how does weighted work for class-99 when we don't have any in the dataset. This means the metric doesn't hold true from train to test?",
      "votes": null
    },
    {
      "id": "395936",
      "postDate": "09/29/2018 16:12:39",
      "content": "<p>Kaggle's data scientists have asked that we don't disclose the exact weights in order to encourage people to work with the data and discourage trying to game the leaderboard. </p>\n\n<p>We can say that the weights differ from unity by no more than a number of O(0), so you can effectively treat all classes as equally weighted (as it says on the metrics page). In that situation, the outer summation of the metric reduces to the straight arithmetic mean. </p>\n\n<p>I'll also add that we've submitted a paper on the metric to the arXiv repository examining the properties of the metrics, including its behavior with different weights. This will be included with their public listing on Monday, and we'll update this thread with a link to that metrics paper as soon as it becomes available. </p>",
      "rawMarkdown": "Kaggle's data scientists have asked that we don't disclose the exact weights in order to encourage people to work with the data and discourage trying to game the leaderboard. \n\nWe can say that the weights differ from unity by no more than a number of O(0), so you can effectively treat all classes as equally weighted (as it says on the metrics page). In that situation, the outer summation of the metric reduces to the straight arithmetic mean. \n\nI'll also add that we've submitted a paper on the metric to the arXiv repository examining the properties of the metrics, including its behavior with different weights. This will be included with their public listing on Monday, and we'll update this thread with a link to that metrics paper as soon as it becomes available.",
      "votes": null
    },
    {
      "id": "395999",
      "postDate": "09/29/2018 19:57:15",
      "content": "<p>There is no way to optimize the real metric without knowing the weights a-priori. What will happen is people probing the LB to estimate the Public weights or doing equally weighted models. If organizer are happy with sub-optimum models, it is ok :-). </p>\n\n<p>Probing all 15 individual classes scores will require at least 15 submissions.\nProbing all / Public LB:</p>\n\n<p>99 / 30.701</p>\n\n<p>95 / ?</p>\n\n<p>92 / ?</p>\n\n<p>90 / 32.620</p>\n\n<p>88 / ?</p>\n\n<p>67 / ?</p>\n\n<p>65 / ?</p>\n\n<p>64 / ?</p>\n\n<p>62 / ?</p>\n\n<p>53 / ?</p>\n\n<p>52 / ?</p>\n\n<p>42 / ?</p>\n\n<p>16 / ?</p>\n\n<p>15 / ?</p>\n\n<p>6 / ?</p>",
      "rawMarkdown": "There is no way to optimize the real metric without knowing the weights a-priori. What will happen is people probing the LB to estimate the Public weights or doing equally weighted models. If organizer are happy with sub-optimum models, it is ok :-). \n\nProbing all 15 individual classes scores will require at least 15 submissions.\nProbing all / Public LB:\n\n99 / 30.701\n\n95 / ?\n\n92 / ?\n\n90 / 32.620\n\n88 / ?\n\n67 / ?\n\n65 / ?\n\n64 / ?\n\n62 / ?\n\n53 / ?\n\n52 / ?\n\n42 / ?\n\n16 / ?\n\n15 / ?\n\n6 / ?",
      "votes": null
    },
    {
      "id": "396025",
      "postDate": "09/29/2018 21:31:15",
      "content": "<p>Hi Giba,</p>\n\n<p>&gt;  If organizer are happy with sub-optimum models, it is ok :-). </p>\n\n<p>The metric is necessarily a compromise. </p>\n\n<p>Even within just the group of people that put PLAsTiCC together, there's a wide range of scientific interests. Within LSST, there are two large collaborations that are interested in variable and transient sources - the Dark Energy Science Collaboration (DESC) and the Transient and Variable Stars (TVS) collaboration - both of which are several times the size of our dev team. There's no one number that actually represents what we're all interested in, and there's no way to really weight things a priori as a) there are objects we will discover that we don't know about and can't say in advance how interesting these will be b) even for what we do know about, we all disagree on the weights and c) we're good Bayesians and the results of this challenge will actually help design the survey and change the weights anyway. </p>\n\n<p>What we're really interested in is seeing a range of approaches to tackle this problem, and maybe techniques we've not used before. This is also why there's a separate science prize, and we hope we can collaborate with contestants who submit interesting entries to reward their effort but also include them in the project and invite them to share their work in papers and at conferences. </p>\n\n<p>There's got to be a single yardstick to evaluate entries and hence there's the metric, but this Kaggle challenge will be a success if it results in a diverse range of ideas. Simply, we don't actually do science by evaluating a single number.</p>\n\n<p>The Kaggle data scientists suggested we don't disclose the weights to encourage this spirit. Yes, we're sure the LB can be probed, but I'm sure the weights won't remain secret for long, but playing with the data is hopefully more fun than just the LB!</p>\n\n<p>Best,</p>\n\n<p>-The PLAsTiCC Dev Team </p>",
      "rawMarkdown": "Hi Giba,\n\n&gt;  If organizer are happy with sub-optimum models, it is ok :-). \n\nThe metric is necessarily a compromise. \n\nEven within just the group of people that put PLAsTiCC together, there's a wide range of scientific interests. Within LSST, there are two large collaborations that are interested in variable and transient sources - the Dark Energy Science Collaboration (DESC) and the Transient and Variable Stars (TVS) collaboration - both of which are several times the size of our dev team. There's no one number that actually represents what we're all interested in, and there's no way to really weight things a priori as a) there are objects we will discover that we don't know about and can't say in advance how interesting these will be b) even for what we do know about, we all disagree on the weights and c) we're good Bayesians and the results of this challenge will actually help design the survey and change the weights anyway. \n\nWhat we're really interested in is seeing a range of approaches to tackle this problem, and maybe techniques we've not used before. This is also why there's a separate science prize, and we hope we can collaborate with contestants who submit interesting entries to reward their effort but also include them in the project and invite them to share their work in papers and at conferences. \n\nThere's got to be a single yardstick to evaluate entries and hence there's the metric, but this Kaggle challenge will be a success if it results in a diverse range of ideas. Simply, we don't actually do science by evaluating a single number.\n\nThe Kaggle data scientists suggested we don't disclose the weights to encourage this spirit. Yes, we're sure the LB can be probed, but I'm sure the weights won't remain secret for long, but playing with the data is hopefully more fun than just the LB!\n\nBest,\n\n-The PLAsTiCC Dev Team",
      "votes": null
    },
    {
      "id": "398140",
      "postDate": "10/03/2018 15:06:50",
      "content": "<blockquote>\n  <p>'ll also add that we've submitted a paper on the metric to the arXiv repository examining the properties of the metrics, including its behavior with different weights. This will be included with their public listing on Monday, and we'll update this thread with a link to that metrics paper as soon as it becomes available.</p>\n</blockquote>\n\n<p>I suppose this is the paper we are looking for:\n<a href=\"https://arxiv.org/abs/1810.00001\">https://arxiv.org/abs/1810.00001</a></p>",
      "rawMarkdown": "&gt; 'll also add that we've submitted a paper on the metric to the arXiv repository examining the properties of the metrics, including its behavior with different weights. This will be included with their public listing on Monday, and we'll update this thread with a link to that metrics paper as soon as it becomes available.\n\nI suppose this is the paper we are looking for:\nhttps://arxiv.org/abs/1810.00001",
      "votes": null
    },
    {
      "id": "398146",
      "postDate": "10/03/2018 15:14:34",
      "content": "<p>Yes, Renee Hlozek posted about it here: <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67376#\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67376#</a></p>",
      "rawMarkdown": "Yes, Renee Hlozek posted about it here: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67376#",
      "votes": null
    },
    {
      "id": "398155",
      "postDate": "10/03/2018 15:34:34",
      "content": "<p>and <a href=\"https://arxiv.org/abs/1809.11145\">https://arxiv.org/abs/1809.11145</a></p>",
      "rawMarkdown": "and https://arxiv.org/abs/1809.11145",
      "votes": null
    },
    {
      "id": "398430",
      "postDate": "10/04/2018 05:32:49",
      "content": "<p><em>...but playing with the data is hopefully more fun than just the LB!</em></p>\n\n<p>Of course ! But sometimes it is frustrating to see no progress between submissions even with a hard work in the meantime.</p>\n\n<p>This competition is really interesting.</p>",
      "rawMarkdown": "*...but playing with the data is hopefully more fun than just the LB!*\n\nOf course ! But sometimes it is frustrating to see no progress between submissions even with a hard work in the meantime.\n\nThis competition is really interesting.",
      "votes": null
    },
    {
      "id": "398656",
      "postDate": "10/04/2018 12:26:37",
      "content": "<p>Oh, we're definitely watching the leaderboard very closely and we're really happy to see the progress you folks have been making! It is really gratifying that you find this interesting - we had a lot of fun putting it together, but this was an experiment for us, and for Kaggle in many ways. It seems to be going well so far!</p>\n\n<p>We figured people would determine the weights fast, and we weren't concealing them out of some misplaced notion that they are secrets to be closely guarded.  Rather, this was to encourage people to look past the metric and the leaderboard. </p>\n\n<p>As we said in the reply to Giba, there are more prizes available than for just the top three entrants on the LB, and it might be that the highest performing submissions on the LB do not perform as well as other approaches when we look at different metrics - for example, the best submission on the LB might not be the best at identifying events in class 99. </p>\n\n<p>We're interested in novel ideas and approaches. We're happy to work with contestants to flesh these ideas out into scientific papers, have them come to conferences, and there are additional monetary rewards for this work. </p>\n\n<p>There are much more definitive statements we can make about what we're considering for the science prizes, and we will in a few weeks, but for now, we're trying not to bias the approaches you folks might take by specifying what other factors we'll consider - that just adds new metrics to optimize. At least early on, we feel it's better to encourage thinking outside the box.</p>\n\n<p>Best,</p>\n\n<p>-Gautham on behalf of the PLAsTiCC team</p>",
      "rawMarkdown": "Oh, we're definitely watching the leaderboard very closely and we're really happy to see the progress you folks have been making! It is really gratifying that you find this interesting - we had a lot of fun putting it together, but this was an experiment for us, and for Kaggle in many ways. It seems to be going well so far!\n\nWe figured people would determine the weights fast, and we weren't concealing them out of some misplaced notion that they are secrets to be closely guarded.  Rather, this was to encourage people to look past the metric and the leaderboard. \n\nAs we said in the reply to Giba, there are more prizes available than for just the top three entrants on the LB, and it might be that the highest performing submissions on the LB do not perform as well as other approaches when we look at different metrics - for example, the best submission on the LB might not be the best at identifying events in class 99. \n\nWe're interested in novel ideas and approaches. We're happy to work with contestants to flesh these ideas out into scientific papers, have them come to conferences, and there are additional monetary rewards for this work. \n\nThere are much more definitive statements we can make about what we're considering for the science prizes, and we will in a few weeks, but for now, we're trying not to bias the approaches you folks might take by specifying what other factors we'll consider - that just adds new metrics to optimize. At least early on, we feel it's better to encourage thinking outside the box.\n\nBest,\n\n-Gautham on behalf of the PLAsTiCC team",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 395936,
      "author_name": "gsnarayan",
      "author_url": "",
      "post_date": "09/29/2018 16:12:39",
      "content": "<p>Kaggle's data scientists have asked that we don't disclose the exact weights in order to encourage people to work with the data and discourage trying to game the leaderboard. </p>\n\n<p>We can say that the weights differ from unity by no more than a number of O(0), so you can effectively treat all classes as equally weighted (as it says on the metrics page). In that situation, the outer summation of the metric reduces to the straight arithmetic mean. </p>\n\n<p>I'll also add that we've submitted a paper on the metric to the arXiv repository examining the properties of the metrics, including its behavior with different weights. This will be included with their public listing on Monday, and we'll update this thread with a link to that metrics paper as soon as it becomes available. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 395999,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "09/29/2018 19:57:15",
      "content": "<p>There is no way to optimize the real metric without knowing the weights a-priori. What will happen is people probing the LB to estimate the Public weights or doing equally weighted models. If organizer are happy with sub-optimum models, it is ok :-). </p>\n\n<p>Probing all 15 individual classes scores will require at least 15 submissions.\nProbing all / Public LB:</p>\n\n<p>99 / 30.701</p>\n\n<p>95 / ?</p>\n\n<p>92 / ?</p>\n\n<p>90 / 32.620</p>\n\n<p>88 / ?</p>\n\n<p>67 / ?</p>\n\n<p>65 / ?</p>\n\n<p>64 / ?</p>\n\n<p>62 / ?</p>\n\n<p>53 / ?</p>\n\n<p>52 / ?</p>\n\n<p>42 / ?</p>\n\n<p>16 / ?</p>\n\n<p>15 / ?</p>\n\n<p>6 / ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 396025,
          "author_name": "gsnarayan",
          "author_url": "",
          "post_date": "09/29/2018 21:31:15",
          "content": "<p>Hi Giba,</p>\n\n<p>&gt;  If organizer are happy with sub-optimum models, it is ok :-). </p>\n\n<p>The metric is necessarily a compromise. </p>\n\n<p>Even within just the group of people that put PLAsTiCC together, there's a wide range of scientific interests. Within LSST, there are two large collaborations that are interested in variable and transient sources - the Dark Energy Science Collaboration (DESC) and the Transient and Variable Stars (TVS) collaboration - both of which are several times the size of our dev team. There's no one number that actually represents what we're all interested in, and there's no way to really weight things a priori as a) there are objects we will discover that we don't know about and can't say in advance how interesting these will be b) even for what we do know about, we all disagree on the weights and c) we're good Bayesians and the results of this challenge will actually help design the survey and change the weights anyway. </p>\n\n<p>What we're really interested in is seeing a range of approaches to tackle this problem, and maybe techniques we've not used before. This is also why there's a separate science prize, and we hope we can collaborate with contestants who submit interesting entries to reward their effort but also include them in the project and invite them to share their work in papers and at conferences. </p>\n\n<p>There's got to be a single yardstick to evaluate entries and hence there's the metric, but this Kaggle challenge will be a success if it results in a diverse range of ideas. Simply, we don't actually do science by evaluating a single number.</p>\n\n<p>The Kaggle data scientists suggested we don't disclose the weights to encourage this spirit. Yes, we're sure the LB can be probed, but I'm sure the weights won't remain secret for long, but playing with the data is hopefully more fun than just the LB!</p>\n\n<p>Best,</p>\n\n<p>-The PLAsTiCC Dev Team </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 398430,
          "author_name": "pfichou",
          "author_url": "",
          "post_date": "10/04/2018 05:32:49",
          "content": "<p><em>...but playing with the data is hopefully more fun than just the LB!</em></p>\n\n<p>Of course ! But sometimes it is frustrating to see no progress between submissions even with a hard work in the meantime.</p>\n\n<p>This competition is really interesting.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 398656,
          "author_name": "gsnarayan",
          "author_url": "",
          "post_date": "10/04/2018 12:26:37",
          "content": "<p>Oh, we're definitely watching the leaderboard very closely and we're really happy to see the progress you folks have been making! It is really gratifying that you find this interesting - we had a lot of fun putting it together, but this was an experiment for us, and for Kaggle in many ways. It seems to be going well so far!</p>\n\n<p>We figured people would determine the weights fast, and we weren't concealing them out of some misplaced notion that they are secrets to be closely guarded.  Rather, this was to encourage people to look past the metric and the leaderboard. </p>\n\n<p>As we said in the reply to Giba, there are more prizes available than for just the top three entrants on the LB, and it might be that the highest performing submissions on the LB do not perform as well as other approaches when we look at different metrics - for example, the best submission on the LB might not be the best at identifying events in class 99. </p>\n\n<p>We're interested in novel ideas and approaches. We're happy to work with contestants to flesh these ideas out into scientific papers, have them come to conferences, and there are additional monetary rewards for this work. </p>\n\n<p>There are much more definitive statements we can make about what we're considering for the science prizes, and we will in a few weeks, but for now, we're trying not to bias the approaches you folks might take by specifying what other factors we'll consider - that just adds new metrics to optimize. At least early on, we feel it's better to encourage thinking outside the box.</p>\n\n<p>Best,</p>\n\n<p>-Gautham on behalf of the PLAsTiCC team</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 398140,
      "author_name": "kauffmanan",
      "author_url": "",
      "post_date": "10/03/2018 15:06:50",
      "content": "<blockquote>\n  <p>'ll also add that we've submitted a paper on the metric to the arXiv repository examining the properties of the metrics, including its behavior with different weights. This will be included with their public listing on Monday, and we'll update this thread with a link to that metrics paper as soon as it becomes available.</p>\n</blockquote>\n\n<p>I suppose this is the paper we are looking for:\n<a href=\"https://arxiv.org/abs/1810.00001\">https://arxiv.org/abs/1810.00001</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 398146,
          "author_name": "gsnarayan",
          "author_url": "",
          "post_date": "10/03/2018 15:14:34",
          "content": "<p>Yes, Renee Hlozek posted about it here: <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67376#\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67376#</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 398155,
          "author_name": "reneehlozek",
          "author_url": "",
          "post_date": "10/03/2018 15:34:34",
          "content": "<p>and <a href=\"https://arxiv.org/abs/1809.11145\">https://arxiv.org/abs/1809.11145</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "395896": "I can't find the class weights for the evaluation metric? or it calculated by the class frequencies? Also, how does weighted work for class-99 when we don't have any in the dataset. This means the metric doesn't hold true from train to test?",
    "395936": "Kaggle's data scientists have asked that we don't disclose the exact weights in order to encourage people to work with the data and discourage trying to game the leaderboard. \n\nWe can say that the weights differ from unity by no more than a number of O(0), so you can effectively treat all classes as equally weighted (as it says on the metrics page). In that situation, the outer summation of the metric reduces to the straight arithmetic mean. \n\nI'll also add that we've submitted a paper on the metric to the arXiv repository examining the properties of the metrics, including its behavior with different weights. This will be included with their public listing on Monday, and we'll update this thread with a link to that metrics paper as soon as it becomes available.",
    "395999": "There is no way to optimize the real metric without knowing the weights a-priori. What will happen is people probing the LB to estimate the Public weights or doing equally weighted models. If organizer are happy with sub-optimum models, it is ok :-). \n\nProbing all 15 individual classes scores will require at least 15 submissions.\nProbing all / Public LB:\n\n99 / 30.701\n\n95 / ?\n\n92 / ?\n\n90 / 32.620\n\n88 / ?\n\n67 / ?\n\n65 / ?\n\n64 / ?\n\n62 / ?\n\n53 / ?\n\n52 / ?\n\n42 / ?\n\n16 / ?\n\n15 / ?\n\n6 / ?",
    "396025": "Hi Giba,\n\n&gt;  If organizer are happy with sub-optimum models, it is ok :-). \n\nThe metric is necessarily a compromise. \n\nEven within just the group of people that put PLAsTiCC together, there's a wide range of scientific interests. Within LSST, there are two large collaborations that are interested in variable and transient sources - the Dark Energy Science Collaboration (DESC) and the Transient and Variable Stars (TVS) collaboration - both of which are several times the size of our dev team. There's no one number that actually represents what we're all interested in, and there's no way to really weight things a priori as a) there are objects we will discover that we don't know about and can't say in advance how interesting these will be b) even for what we do know about, we all disagree on the weights and c) we're good Bayesians and the results of this challenge will actually help design the survey and change the weights anyway. \n\nWhat we're really interested in is seeing a range of approaches to tackle this problem, and maybe techniques we've not used before. This is also why there's a separate science prize, and we hope we can collaborate with contestants who submit interesting entries to reward their effort but also include them in the project and invite them to share their work in papers and at conferences. \n\nThere's got to be a single yardstick to evaluate entries and hence there's the metric, but this Kaggle challenge will be a success if it results in a diverse range of ideas. Simply, we don't actually do science by evaluating a single number.\n\nThe Kaggle data scientists suggested we don't disclose the weights to encourage this spirit. Yes, we're sure the LB can be probed, but I'm sure the weights won't remain secret for long, but playing with the data is hopefully more fun than just the LB!\n\nBest,\n\n-The PLAsTiCC Dev Team",
    "398140": "&gt; 'll also add that we've submitted a paper on the metric to the arXiv repository examining the properties of the metrics, including its behavior with different weights. This will be included with their public listing on Monday, and we'll update this thread with a link to that metrics paper as soon as it becomes available.\n\nI suppose this is the paper we are looking for:\nhttps://arxiv.org/abs/1810.00001",
    "398146": "Yes, Renee Hlozek posted about it here: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/67376#",
    "398155": "and https://arxiv.org/abs/1809.11145",
    "398430": "*...but playing with the data is hopefully more fun than just the LB!*\n\nOf course ! But sometimes it is frustrating to see no progress between submissions even with a hard work in the meantime.\n\nThis competition is really interesting.",
    "398656": "Oh, we're definitely watching the leaderboard very closely and we're really happy to see the progress you folks have been making! It is really gratifying that you find this interesting - we had a lot of fun putting it together, but this was an experiment for us, and for Kaggle in many ways. It seems to be going well so far!\n\nWe figured people would determine the weights fast, and we weren't concealing them out of some misplaced notion that they are secrets to be closely guarded.  Rather, this was to encourage people to look past the metric and the leaderboard. \n\nAs we said in the reply to Giba, there are more prizes available than for just the top three entrants on the LB, and it might be that the highest performing submissions on the LB do not perform as well as other approaches when we look at different metrics - for example, the best submission on the LB might not be the best at identifying events in class 99. \n\nWe're interested in novel ideas and approaches. We're happy to work with contestants to flesh these ideas out into scientific papers, have them come to conferences, and there are additional monetary rewards for this work. \n\nThere are much more definitive statements we can make about what we're considering for the science prizes, and we will in a few weeks, but for now, we're trying not to bias the approaches you folks might take by specifying what other factors we'll consider - that just adds new metrics to optimize. At least early on, we feel it's better to encourage thinking outside the box.\n\nBest,\n\n-Gautham on behalf of the PLAsTiCC team"
  },
  "source": "meta"
}