{
  "id": 20884,
  "title": "Score evaluation: too many half-integers?",
  "url": "/competitions/draper-satellite-image-chronology/discussion/20884",
  "author_name": "",
  "post_date": "2016-05-12T05:26:16.913Z",
  "votes": 1,
  "comment_count": 7,
  "views": 1290,
  "content": "<p>Hi there!</p>\n\n<p>I have question on score evaluation.\nIf a participant uses distinct numbers 1,2,3,4,5 in each row of submission file, then formula \n$$r_s = 1 - \\dfrac{6\\cdot SquareSum_s}{n(n^2-1)}$$ \ncan be used, and since n=5, then\n$$r_s = 1 - \\dfrac{SquareSum_s}{20}.$$\nIf number of tests is N (N ~ 0.17*274 ~ 44...49), then score is\n$$R = \\frac{1}{N} \\sum_{s=1}^{N} r_s = \\frac{1}{20N} \\sum_{s=1}^{N}(20-SquareSum_s).$$\nBut (!) SquareSum_s is integer in this case, \nso 20*N*R must be integer, right?</p>\n\n<p>When consider scores from 'raw data' archive and multiply each by 20*49 (use N=49), then one will get a lot of (almost) integer numbers, which is expectable; but <em>many half-integers too</em> (like 3.5, 10.5, ...).</p>\n\n<p>I used distinct 'days' 1,2,3,4,5 in each row of my draft submission file, and have got half-integers sometimes.</p>\n\n<p><strong>Question</strong>: <em>Why half-integers occur so often?</em>\n<br>a - I can't believe that so many participants use non-distinct 'days' in submission files.\n<br>b - maybe wrong my calculations? where?\n<br>c - maybe divide by 2N, not N (to get scores in [-0.5, 0.5])? but no, looking at current LB;\n<br>d - other reasons?  </p>",
  "messages": [
    {
      "id": "119664",
      "postDate": "05/12/2016 05:26:16",
      "content": "<p>Hi there!</p>\n\n<p>I have question on score evaluation.\nIf a participant uses distinct numbers 1,2,3,4,5 in each row of submission file, then formula \n$$r_s = 1 - \\dfrac{6\\cdot SquareSum_s}{n(n^2-1)}$$ \ncan be used, and since n=5, then\n$$r_s = 1 - \\dfrac{SquareSum_s}{20}.$$\nIf number of tests is N (N ~ 0.17*274 ~ 44...49), then score is\n$$R = \\frac{1}{N} \\sum_{s=1}^{N} r_s = \\frac{1}{20N} \\sum_{s=1}^{N}(20-SquareSum_s).$$\nBut (!) SquareSum_s is integer in this case, \nso 20*N*R must be integer, right?</p>\n\n<p>When consider scores from 'raw data' archive and multiply each by 20*49 (use N=49), then one will get a lot of (almost) integer numbers, which is expectable; but <em>many half-integers too</em> (like 3.5, 10.5, ...).</p>\n\n<p>I used distinct 'days' 1,2,3,4,5 in each row of my draft submission file, and have got half-integers sometimes.</p>\n\n<p><strong>Question</strong>: <em>Why half-integers occur so often?</em>\n<br>a - I can't believe that so many participants use non-distinct 'days' in submission files.\n<br>b - maybe wrong my calculations? where?\n<br>c - maybe divide by 2N, not N (to get scores in [-0.5, 0.5])? but no, looking at current LB;\n<br>d - other reasons?  </p>",
      "rawMarkdown": "Hi there!\r\n\r\nI have question on score evaluation.\r\nIf a participant uses distinct numbers 1,2,3,4,5 in each row of submission file, then formula \r\n$$r_s = 1 - \\dfrac{6\\cdot SquareSum_s}{n(n^2-1)}$$ \r\ncan be used, and since n=5, then\r\n$$r_s = 1 - \\dfrac{SquareSum_s}{20}.$$\r\nIf number of tests is N (N ~ 0.17*274 ~ 44...49), then score is\r\n$$R = \\frac{1}{N} \\sum_{s=1}^{N} r_s = \\frac{1}{20N} \\sum_{s=1}^{N}(20-SquareSum_s).$$\r\nBut (!) SquareSum_s is integer in this case, \r\nso 20*N*R must be integer, right?\r\n\r\nWhen consider scores from 'raw data' archive and multiply each by 20*49 (use N=49), then one will get a lot of (almost) integer numbers, which is expectable; but *many half-integers too* (like 3.5, 10.5, ...).\r\n\r\nI used distinct 'days' 1,2,3,4,5 in each row of my draft submission file, and have got half-integers sometimes.\r\n\r\n**Question**: *Why half-integers occur so often?*\r\n<br>a - I can't believe that so many participants use non-distinct 'days' in submission files.\r\n<br>b - maybe wrong my calculations? where?\r\n<br>c - maybe divide by 2N, not N (to get scores in [-0.5, 0.5])? but no, looking at current LB;\r\n<br>d - other reasons?",
      "votes": null
    },
    {
      "id": "119706",
      "postDate": "05/12/2016 12:04:27",
      "content": "<p>When assume that scores were calculated from N=98 testing entries, then all seems to be correct.</p>\n\n<p>But then words &quot;This leaderboard is calculated on approximately 17% of the test data&quot; need to be replaced with &quot;... 35% of the test data&quot;.</p>",
      "rawMarkdown": "When assume that scores were calculated from N=98 testing entries, then all seems to be correct.\r\n\r\nBut then words \"This leaderboard is calculated on approximately 17% of the test data\" need to be replaced with \"... 35% of the test data\".",
      "votes": null
    },
    {
      "id": "124052",
      "postDate": "06/15/2016 07:47:38",
      "content": "<p>My estimations show that N is divisible by 14, and  most probable value for N is 28. </p>\n\n<p>Then &quot;This leaderboard is calculated on approximately 17% of the test data&quot; need to be replaced with &quot;... 10% of the test data&quot; (or 11%).</p>",
      "rawMarkdown": "My estimations show that N is divisible by 14, and  most probable value for N is 28. \r\n\r\nThen \"This leaderboard is calculated on approximately 17% of the test data\" need to be replaced with \"... 10% of the test data\" (or 11%).",
      "votes": null
    },
    {
      "id": "124117",
      "postDate": "06/15/2016 17:10:36",
      "content": "<p>I wondered about that too.  My current hypothesis based on William's comment on this thread\nabout <a href=\"https://www.kaggle.com/c/draper-satellite-image-chronology/forums/t/20983/no-change-in-submission-score-despite-changes/120166#post120166\">no changes in leaderboard score</a> is that it probably is 17% of the data actually being used but recall that not all sets in the test data are being used. So if N=28 for the public leaderboard, then there are only 164 or 165 sets in the public/private leaderboard combined.  This would imply that around 85 sets are in neither.  I must admit that I have certain sets that I am hoping are in the 85 so to speak...</p>",
      "rawMarkdown": "I wondered about that too.  My current hypothesis based on William's comment on this thread\r\nabout [no changes in leaderboard score][1] is that it probably is 17% of the data actually being used but recall that not all sets in the test data are being used. So if N=28 for the public leaderboard, then there are only 164 or 165 sets in the public/private leaderboard combined.  This would imply that around 85 sets are in neither.  I must admit that I have certain sets that I am hoping are in the 85 so to speak...\r\n\r\n  [1]: https://www.kaggle.com/c/draper-satellite-image-chronology/forums/t/20983/no-change-in-submission-score-despite-changes/120166#post120166",
      "votes": null
    },
    {
      "id": "124144",
      "postDate": "06/15/2016 19:45:50",
      "content": "<p>Why N=28, wouldn't N=42 be more likely?</p>\n\n<p>William  said a small number of sets aren't part of the scoring. If N=28 then, as Chris notes, that leaves circa 85 sets omitted, which doesn't seem small to me.</p>\n\n<p>If N=42 and there is 17% coverage then the public/private leaderboard will contain between 241 to 254 sets, omitting anything between 20 to 33 sets (<em>as an aside still seems too many to be classed as small to me</em>).</p>\n\n<p>Anyhow when you also take into account that;</p>\n\n<ul>\n<li>there are 2 image clusters where the training and test data overlap</li>\n<li>theses overlap clusters contain 22 images</li>\n<li>my personal experience is that none of these images are in the public leaderboard</li>\n</ul>\n\n<p>Then I think that this points to N=42 and that the omitted sets are the 22 sets that geographically overlap the training data.</p>",
      "rawMarkdown": "Why N=28, wouldn't N=42 be more likely?\r\n\r\nWilliam  said a small number of sets aren't part of the scoring. If N=28 then, as Chris notes, that leaves circa 85 sets omitted, which doesn't seem small to me.\r\n\r\nIf N=42 and there is 17% coverage then the public/private leaderboard will contain between 241 to 254 sets, omitting anything between 20 to 33 sets (*as an aside still seems too many to be classed as small to me*).\r\n\r\nAnyhow when you also take into account that;\r\n\r\n - there are 2 image clusters where the training and test data overlap\r\n - theses overlap clusters contain 22 images\r\n - my personal experience is that none of these images are in the public leaderboard\r\n\r\nThen I think that this points to N=42 and that the omitted sets are the 22 sets that geographically overlap the training data.",
      "votes": null
    },
    {
      "id": "124348",
      "postDate": "06/17/2016 12:06:48",
      "content": "<p>@Chippy, Yes, the ratio 42/274 is closer to 0.17 than 28/274. <br>\nAnd I agree that none of these 22 (or 23) images are in the Public LB. </p>\n\n<p>My thought was: since 14| N, then I focused on 14, 28, 42, 56, ... and obtained contradictions for all these values but 28. (I hope my estimations were error-free).</p>\n\n<p>Well, I just assume that N=28 (don't claim by 100%).</p>\n\n<p>Non-direct way:\nfocus on any of 5 southern clusters; and try to evaluate number of images which change Private LB score; for each of these clusters I've got ratio \n$$\nr_{cluster} = \\dfrac{images\\; which \\; change\\;Private\\;LB\\; score}{total \\;images} &lt; 0.148;\n$$\nBut in the case of N=42  -- at least one of these clusters must have this ratio greater than\n(42/274) ~ 0.153 .</p>\n\n<p>And there is no guarantee that the rate of ignored images is the same for Private/Public LBs.\n<br>For example: <br> \n - Public set: 47 images (28 essential + 19 ignored);<br>\n - Private set: 227 images (135 essential + 92 ignored).<br>\nkeeps ratio (more-less).</p>\n\n<p>But it can be <br>\n - Public set: 47 images (28 essential + 19 ignored);<br>\n - Private set: 227 images (227 essential + 0 ignored).<br>\nwhy not?</p>\n\n<p>Here is my main confuse.\nI thought before, that set distribution is extremely easy:\n$$\nPublic Set (17\\%) + Private Set(83\\%) = 274,$$\nbut it seems that it is more complicated.</p>",
      "rawMarkdown": "Chippy, Yes, the ratio 42/274 is closer to 0.17 than 28/274. <br>\r\nAnd I agree that none of these 22 (or 23) images are in the Public LB. \r\n\r\nMy thought was: since 14| N, then I focused on 14, 28, 42, 56, ... and obtained contradictions for all these values but 28. (I hope my estimations were error-free).\r\n\r\nWell, I just assume that N=28 (don't claim by 100%).\r\n\r\nNon-direct way:\r\nfocus on any of 5 southern clusters; and try to evaluate number of images which change Private LB score; for each of these clusters I've got ratio \r\n$$\r\nr_{cluster} = \\dfrac{images\\; which \\; change\\;Private\\;LB\\; score}{total \\;images} < 0.148;\r\n$$\r\nBut in the case of N=42  -- at least one of these clusters must have this ratio greater than\r\n(42/274) ~ 0.153 .\r\n\r\n\r\nAnd there is no guarantee that the rate of ignored images is the same for Private/Public LBs.\r\n<br>For example: <br> \r\n - Public set: 47 images (28 essential + 19 ignored);<br>\r\n - Private set: 227 images (135 essential + 92 ignored).<br>\r\nkeeps ratio (more-less).\r\n\r\nBut it can be <br>\r\n - Public set: 47 images (28 essential + 19 ignored);<br>\r\n - Private set: 227 images (227 essential + 0 ignored).<br>\r\nwhy not?\r\n\r\nHere is my main confuse.\r\nI thought before, that set distribution is extremely easy:\r\n$$\r\nPublic Set (17\\%) + Private Set(83\\%) = 274,$$\r\nbut it seems that it is more complicated.",
      "votes": null
    },
    {
      "id": "124373",
      "postDate": "06/17/2016 16:23:19",
      "content": "<p>@Virtuoso, thanks for your follow up explanation. </p>\n\n<p>It's helpful to know that this is your conclusion based upon more extensive leaderboard feedback than I've been able to do. </p>\n\n<p>In my limited leaderboard feedback I've certainly seen strange behaviour in some clusters which would be consistent with your hypothesis. </p>\n\n<p>To date though I've thought this behaviour is been more likely to be due to errors in my submissions! Maybe they are right and I'm focusing my efforts in the wrong places? </p>",
      "rawMarkdown": "Virtuoso, thanks for your follow up explanation. \r\n\r\nIt's helpful to know that this is your conclusion based upon more extensive leaderboard feedback than I've been able to do. \r\n\r\nIn my limited leaderboard feedback I've certainly seen strange behaviour in some clusters which would be consistent with your hypothesis. \r\n\r\nTo date though I've thought this behaviour is been more likely to be due to errors in my submissions! Maybe they are right and I'm focusing my efforts in the wrong places?",
      "votes": null
    },
    {
      "id": "124655",
      "postDate": "06/20/2016 22:43:10",
      "content": "<p>@Virtuoso, from the change in my score from a chance submission in which I only changed the order of 1 set I too believe that indeed only 28 sets are contributing to the public leaderboard..!</p>",
      "rawMarkdown": "Virtuoso, from the change in my score from a chance submission in which I only changed the order of 1 set I too believe that indeed only 28 sets are contributing to the public leaderboard..!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 119706,
      "author_name": "virtuoso",
      "author_url": "",
      "post_date": "05/12/2016 12:04:27",
      "content": "<p>When assume that scores were calculated from N=98 testing entries, then all seems to be correct.</p>\n\n<p>But then words &quot;This leaderboard is calculated on approximately 17% of the test data&quot; need to be replaced with &quot;... 35% of the test data&quot;.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124052,
      "author_name": "virtuoso",
      "author_url": "",
      "post_date": "06/15/2016 07:47:38",
      "content": "<p>My estimations show that N is divisible by 14, and  most probable value for N is 28. </p>\n\n<p>Then &quot;This leaderboard is calculated on approximately 17% of the test data&quot; need to be replaced with &quot;... 10% of the test data&quot; (or 11%).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124117,
      "author_name": "hurlburt",
      "author_url": "",
      "post_date": "06/15/2016 17:10:36",
      "content": "<p>I wondered about that too.  My current hypothesis based on William's comment on this thread\nabout <a href=\"https://www.kaggle.com/c/draper-satellite-image-chronology/forums/t/20983/no-change-in-submission-score-despite-changes/120166#post120166\">no changes in leaderboard score</a> is that it probably is 17% of the data actually being used but recall that not all sets in the test data are being used. So if N=28 for the public leaderboard, then there are only 164 or 165 sets in the public/private leaderboard combined.  This would imply that around 85 sets are in neither.  I must admit that I have certain sets that I am hoping are in the 85 so to speak...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124144,
      "author_name": "nigelcarpenter",
      "author_url": "",
      "post_date": "06/15/2016 19:45:50",
      "content": "<p>Why N=28, wouldn't N=42 be more likely?</p>\n\n<p>William  said a small number of sets aren't part of the scoring. If N=28 then, as Chris notes, that leaves circa 85 sets omitted, which doesn't seem small to me.</p>\n\n<p>If N=42 and there is 17% coverage then the public/private leaderboard will contain between 241 to 254 sets, omitting anything between 20 to 33 sets (<em>as an aside still seems too many to be classed as small to me</em>).</p>\n\n<p>Anyhow when you also take into account that;</p>\n\n<ul>\n<li>there are 2 image clusters where the training and test data overlap</li>\n<li>theses overlap clusters contain 22 images</li>\n<li>my personal experience is that none of these images are in the public leaderboard</li>\n</ul>\n\n<p>Then I think that this points to N=42 and that the omitted sets are the 22 sets that geographically overlap the training data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124348,
      "author_name": "virtuoso",
      "author_url": "",
      "post_date": "06/17/2016 12:06:48",
      "content": "<p>@Chippy, Yes, the ratio 42/274 is closer to 0.17 than 28/274. <br>\nAnd I agree that none of these 22 (or 23) images are in the Public LB. </p>\n\n<p>My thought was: since 14| N, then I focused on 14, 28, 42, 56, ... and obtained contradictions for all these values but 28. (I hope my estimations were error-free).</p>\n\n<p>Well, I just assume that N=28 (don't claim by 100%).</p>\n\n<p>Non-direct way:\nfocus on any of 5 southern clusters; and try to evaluate number of images which change Private LB score; for each of these clusters I've got ratio \n$$\nr_{cluster} = \\dfrac{images\\; which \\; change\\;Private\\;LB\\; score}{total \\;images} &lt; 0.148;\n$$\nBut in the case of N=42  -- at least one of these clusters must have this ratio greater than\n(42/274) ~ 0.153 .</p>\n\n<p>And there is no guarantee that the rate of ignored images is the same for Private/Public LBs.\n<br>For example: <br> \n - Public set: 47 images (28 essential + 19 ignored);<br>\n - Private set: 227 images (135 essential + 92 ignored).<br>\nkeeps ratio (more-less).</p>\n\n<p>But it can be <br>\n - Public set: 47 images (28 essential + 19 ignored);<br>\n - Private set: 227 images (227 essential + 0 ignored).<br>\nwhy not?</p>\n\n<p>Here is my main confuse.\nI thought before, that set distribution is extremely easy:\n$$\nPublic Set (17\\%) + Private Set(83\\%) = 274,$$\nbut it seems that it is more complicated.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124373,
      "author_name": "nigelcarpenter",
      "author_url": "",
      "post_date": "06/17/2016 16:23:19",
      "content": "<p>@Virtuoso, thanks for your follow up explanation. </p>\n\n<p>It's helpful to know that this is your conclusion based upon more extensive leaderboard feedback than I've been able to do. </p>\n\n<p>In my limited leaderboard feedback I've certainly seen strange behaviour in some clusters which would be consistent with your hypothesis. </p>\n\n<p>To date though I've thought this behaviour is been more likely to be due to errors in my submissions! Maybe they are right and I'm focusing my efforts in the wrong places? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124655,
      "author_name": "nigelcarpenter",
      "author_url": "",
      "post_date": "06/20/2016 22:43:10",
      "content": "<p>@Virtuoso, from the change in my score from a chance submission in which I only changed the order of 1 set I too believe that indeed only 28 sets are contributing to the public leaderboard..!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "119664": "Hi there!\r\n\r\nI have question on score evaluation.\r\nIf a participant uses distinct numbers 1,2,3,4,5 in each row of submission file, then formula \r\n$$r_s = 1 - \\dfrac{6\\cdot SquareSum_s}{n(n^2-1)}$$ \r\ncan be used, and since n=5, then\r\n$$r_s = 1 - \\dfrac{SquareSum_s}{20}.$$\r\nIf number of tests is N (N ~ 0.17*274 ~ 44...49), then score is\r\n$$R = \\frac{1}{N} \\sum_{s=1}^{N} r_s = \\frac{1}{20N} \\sum_{s=1}^{N}(20-SquareSum_s).$$\r\nBut (!) SquareSum_s is integer in this case, \r\nso 20*N*R must be integer, right?\r\n\r\nWhen consider scores from 'raw data' archive and multiply each by 20*49 (use N=49), then one will get a lot of (almost) integer numbers, which is expectable; but *many half-integers too* (like 3.5, 10.5, ...).\r\n\r\nI used distinct 'days' 1,2,3,4,5 in each row of my draft submission file, and have got half-integers sometimes.\r\n\r\n**Question**: *Why half-integers occur so often?*\r\n<br>a - I can't believe that so many participants use non-distinct 'days' in submission files.\r\n<br>b - maybe wrong my calculations? where?\r\n<br>c - maybe divide by 2N, not N (to get scores in [-0.5, 0.5])? but no, looking at current LB;\r\n<br>d - other reasons?",
    "119706": "When assume that scores were calculated from N=98 testing entries, then all seems to be correct.\r\n\r\nBut then words \"This leaderboard is calculated on approximately 17% of the test data\" need to be replaced with \"... 35% of the test data\".",
    "124052": "My estimations show that N is divisible by 14, and  most probable value for N is 28. \r\n\r\nThen \"This leaderboard is calculated on approximately 17% of the test data\" need to be replaced with \"... 10% of the test data\" (or 11%).",
    "124117": "I wondered about that too.  My current hypothesis based on William's comment on this thread\r\nabout [no changes in leaderboard score][1] is that it probably is 17% of the data actually being used but recall that not all sets in the test data are being used. So if N=28 for the public leaderboard, then there are only 164 or 165 sets in the public/private leaderboard combined.  This would imply that around 85 sets are in neither.  I must admit that I have certain sets that I am hoping are in the 85 so to speak...\r\n\r\n  [1]: https://www.kaggle.com/c/draper-satellite-image-chronology/forums/t/20983/no-change-in-submission-score-despite-changes/120166#post120166",
    "124144": "Why N=28, wouldn't N=42 be more likely?\r\n\r\nWilliam  said a small number of sets aren't part of the scoring. If N=28 then, as Chris notes, that leaves circa 85 sets omitted, which doesn't seem small to me.\r\n\r\nIf N=42 and there is 17% coverage then the public/private leaderboard will contain between 241 to 254 sets, omitting anything between 20 to 33 sets (*as an aside still seems too many to be classed as small to me*).\r\n\r\nAnyhow when you also take into account that;\r\n\r\n - there are 2 image clusters where the training and test data overlap\r\n - theses overlap clusters contain 22 images\r\n - my personal experience is that none of these images are in the public leaderboard\r\n\r\nThen I think that this points to N=42 and that the omitted sets are the 22 sets that geographically overlap the training data.",
    "124348": "Chippy, Yes, the ratio 42/274 is closer to 0.17 than 28/274. <br>\r\nAnd I agree that none of these 22 (or 23) images are in the Public LB. \r\n\r\nMy thought was: since 14| N, then I focused on 14, 28, 42, 56, ... and obtained contradictions for all these values but 28. (I hope my estimations were error-free).\r\n\r\nWell, I just assume that N=28 (don't claim by 100%).\r\n\r\nNon-direct way:\r\nfocus on any of 5 southern clusters; and try to evaluate number of images which change Private LB score; for each of these clusters I've got ratio \r\n$$\r\nr_{cluster} = \\dfrac{images\\; which \\; change\\;Private\\;LB\\; score}{total \\;images} < 0.148;\r\n$$\r\nBut in the case of N=42  -- at least one of these clusters must have this ratio greater than\r\n(42/274) ~ 0.153 .\r\n\r\n\r\nAnd there is no guarantee that the rate of ignored images is the same for Private/Public LBs.\r\n<br>For example: <br> \r\n - Public set: 47 images (28 essential + 19 ignored);<br>\r\n - Private set: 227 images (135 essential + 92 ignored).<br>\r\nkeeps ratio (more-less).\r\n\r\nBut it can be <br>\r\n - Public set: 47 images (28 essential + 19 ignored);<br>\r\n - Private set: 227 images (227 essential + 0 ignored).<br>\r\nwhy not?\r\n\r\nHere is my main confuse.\r\nI thought before, that set distribution is extremely easy:\r\n$$\r\nPublic Set (17\\%) + Private Set(83\\%) = 274,$$\r\nbut it seems that it is more complicated.",
    "124373": "Virtuoso, thanks for your follow up explanation. \r\n\r\nIt's helpful to know that this is your conclusion based upon more extensive leaderboard feedback than I've been able to do. \r\n\r\nIn my limited leaderboard feedback I've certainly seen strange behaviour in some clusters which would be consistent with your hypothesis. \r\n\r\nTo date though I've thought this behaviour is been more likely to be due to errors in my submissions! Maybe they are right and I'm focusing my efforts in the wrong places?",
    "124655": "Virtuoso, from the change in my score from a chance submission in which I only changed the order of 1 set I too believe that indeed only 28 sets are contributing to the public leaderboard..!"
  },
  "source": "meta"
}