{
  "id": 175559,
  "title": "LB probing: our insights into the test data",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175559",
  "author_name": "",
  "post_date": "2020-08-18T15:07:53.366564400Z",
  "votes": 38,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello Kagglers, it was a beautiful competition, hopefully you learnt a lot! </p>\n<p>I want to share with you how our team reached second place on the public LB, and some conclusions that we got along the way. </p>\n<h3>Probing algorithm</h3>\n<p>With 3 digits for the public score initially, and then the switch to 4 digits one month before the competition end, it was interesting to get as much information as possible from that output. I have tried LB probing before in many competitions, mainly by applying mixed integer programming (MIP), but it was the first time I encountered AUC ROC metric. And it turns out it is especially hard for MIP to crack, because it is order-based. </p>\n<p>Additionally, a straightforward MIP formulation would result in about 10 million boolean variables, which in combination with a difficult problem is impossible to solve. It required a simplification and I used the following procedure:</p>\n<blockquote>\n  <ol>\n  <li>Split 690 test patients into few groups. For each group we define two variables that we want to find, - number of public ones and number of public zeros.</li>\n  <li>With MIP formulation generate 50-500 feasible solutions. Where each submission translates into two constraints.</li>\n  <li>Randomly generate 1000-2000 candidate submissions where each group has the same target value, and the values are different between groups. Select a submission which best separates the feasible solutions from the previous point.</li>\n  <li>Repeat 2-3 until only one feasible solution is left, - thus we found number of public ones and zeros for those groups.</li>\n  <li>Split several more groups into halves and repeat 2-4</li>\n  </ol>\n</blockquote>\n<p>The MIP problem I solved with Gurobi solver. Each already available submission is two constraint, - we know that Kaggle rounds down the scores, so we have upper and lower bound. The score is simply number of public zeros to the left of all public ones, divided by the total number of zeros multiplied by the total number of ones.</p>\n<h3>Final output</h3>\n<p>At the end I had 210 test patient groups, and knowing number of public ones and zeros for each group. If a group has a public one then this group contains only one test patient. This was the aim, - to know all test patients with public melanomas. I was splitting until I got there.</p>\n<h3>The submission scoring public 2nd</h3>\n<p>From the output above I know which patients have public melanomas and which do not. Here I used our best submission and set all non-public-melanoma targets to zero, and also scaled up targets for patients who have sum of targets lower than number of public melanomas for that patient. This submission, of course, is useless for the private. It scored 0.6334 on private, 0.9926 on public.</p>\n<h3>General conclusions from the probing</h3>\n<p>early into the probing I concluded that:</p>\n<ol>\n<li>There are overall 78 public ones and 3216 public zeros. Making it almost exactly 30% of the test.</li>\n<li>The public-private split is not by patient. That is, images of every test patient are partially in public and partially in private. More on that below.</li>\n</ol>\n<h3>Public-private split</h3>\n<p>For a competition whose know-how is (a quote from the description below)</p>\n<blockquote>\n  <p>Dermatologists could enhance their diagnostic accuracy if detection algorithms take into account contextual images within the same patient to determine which images represent a melanoma.</p>\n</blockquote>\n<p>I think it strange to not separate public-private by patient. On the other hand, given the competition framework, at least from the probing point of view, it was a correct decision. If the split was by patient, it would have been possible to completely resolve the split in 100-200 submissions. There would be no information leak between the public and the private, but it is still 3k additional data-points to train on. And in the current split it is much harder to probe, but there is more information leak.</p>\n<h3>Using the probing in our final submissions</h3>\n<p>We were not able to use the obtained information to improve the score, our final submissions do not include any usage of it. We tried the following, at the end nothing of it made its way into the final submissions:</p>\n<ol>\n<li>Adding this information as features. During the training we simulated 30% randomly into \"public\" and the rest into \"private\", calculated number of ones and zeros in the \"public\" and added these features to help the \"private\" training.</li>\n<li>Added the test data to the training as in pseudo-labeling and added a new objective, - is there a \"public\" melanoma for a patient. For the train data we used the same simulation from the previous point. For the test data we used the probing output.</li>\n<li>In post-processing, scale up targets for the test patients whose sum of targets is smaller than number of public ones.</li>\n</ol>\n<p>With each of the above points there is a different story how much it improved the CV, LB and why it failed at the end, but not too interesting to go into details.</p>\n<h3>Public-private split, analysis</h3>\n<p>Back to the public-private split. Knowing number of public ones and zeros for 210 groups of test patients I was interested to understand whether the split is completely random or there is something. Spoiler: it seems to be completely random. </p>\n<p>Basically I want to check this hypothesis of splitting logic:</p>\n<blockquote>\n  <p><strong>The hypothesis</strong>: the splitting is done by randomly permuting <code>range(10982)</code> and taking the first 3294 (which is 30%) indexes. </p>\n</blockquote>\n<p>In order to test that hypothesis I first considered Kolmogorov-Smirnov test, but I don't know how to adjust it to different underlying distribution from sample to sample (because each group is different in size). Instead I calculated log-likelihood of my probing result, and then randomly simulated <code>10000</code> spits by following the hypothesis and calculated their log-likelihoods as well. The result is the following graph:</p>\n<p><img src=\"https://imgur.com/GD8wCUk.png\" alt=\"log-likelihood of split\"></p>\n<p>From the graph we can conclude that we indeed can not reject the hypothesis. For comparison there is also a red line, which indicated the log-likelihood of a split where each test patient is separated exactly 30% to public and 70% to private. We see that it is highly improbable under the current hypothesis, however it maximizes likelihood. Please share if you have thoughts or comments on how to do this statistical test properly.</p>\n<p>For comparison test-train separation is not completely random. Specifically, all images with height 1920 are in test. They are probably newer images (because they don't have in-patient age differences) and the organizers wanted them all in the test.</p>\n<h3>General thoughts</h3>\n<p>Looking at this competition design retrospectively, I think it was possible to completely hack it, that is, to categorize each test sample into one of three categories: private, public one or public zero. Probably separating between private and public zeros for low probability is not possible (almost no submissions which cut there) but it is also much less important. The outline of the approach can include 3 steps:</p>\n<ol>\n<li>Find all test patients with public ones</li>\n<li>Find exact samples with public ones (can also be done without going through the first point)</li>\n<li>The point 2 above greatly simplifies the MIP formulation if completed. This will allow adding any and all submissions as constraints. Specifically, all the submissions from the public kernels. That amount of information should be enough to resolve most of the samples.</li>\n</ol>\n<p>And then there is a question how much it helps for the private LB. In this competition in practice it didn't help, but I think it was close, so if I were Kaggle I would consider designing competitions which are more LB-probing-proof in the future.</p>\n<p>Thanks for reading this excessively long post!</p>",
  "messages": [
    {
      "id": "975989",
      "postDate": "08/18/2020 15:07:53",
      "content": "<p>Hello Kagglers, it was a beautiful competition, hopefully you learnt a lot! </p>\n<p>I want to share with you how our team reached second place on the public LB, and some conclusions that we got along the way. </p>\n<h3>Probing algorithm</h3>\n<p>With 3 digits for the public score initially, and then the switch to 4 digits one month before the competition end, it was interesting to get as much information as possible from that output. I have tried LB probing before in many competitions, mainly by applying mixed integer programming (MIP), but it was the first time I encountered AUC ROC metric. And it turns out it is especially hard for MIP to crack, because it is order-based. </p>\n<p>Additionally, a straightforward MIP formulation would result in about 10 million boolean variables, which in combination with a difficult problem is impossible to solve. It required a simplification and I used the following procedure:</p>\n<blockquote>\n  <ol>\n  <li>Split 690 test patients into few groups. For each group we define two variables that we want to find, - number of public ones and number of public zeros.</li>\n  <li>With MIP formulation generate 50-500 feasible solutions. Where each submission translates into two constraints.</li>\n  <li>Randomly generate 1000-2000 candidate submissions where each group has the same target value, and the values are different between groups. Select a submission which best separates the feasible solutions from the previous point.</li>\n  <li>Repeat 2-3 until only one feasible solution is left, - thus we found number of public ones and zeros for those groups.</li>\n  <li>Split several more groups into halves and repeat 2-4</li>\n  </ol>\n</blockquote>\n<p>The MIP problem I solved with Gurobi solver. Each already available submission is two constraint, - we know that Kaggle rounds down the scores, so we have upper and lower bound. The score is simply number of public zeros to the left of all public ones, divided by the total number of zeros multiplied by the total number of ones.</p>\n<h3>Final output</h3>\n<p>At the end I had 210 test patient groups, and knowing number of public ones and zeros for each group. If a group has a public one then this group contains only one test patient. This was the aim, - to know all test patients with public melanomas. I was splitting until I got there.</p>\n<h3>The submission scoring public 2nd</h3>\n<p>From the output above I know which patients have public melanomas and which do not. Here I used our best submission and set all non-public-melanoma targets to zero, and also scaled up targets for patients who have sum of targets lower than number of public melanomas for that patient. This submission, of course, is useless for the private. It scored 0.6334 on private, 0.9926 on public.</p>\n<h3>General conclusions from the probing</h3>\n<p>early into the probing I concluded that:</p>\n<ol>\n<li>There are overall 78 public ones and 3216 public zeros. Making it almost exactly 30% of the test.</li>\n<li>The public-private split is not by patient. That is, images of every test patient are partially in public and partially in private. More on that below.</li>\n</ol>\n<h3>Public-private split</h3>\n<p>For a competition whose know-how is (a quote from the description below)</p>\n<blockquote>\n  <p>Dermatologists could enhance their diagnostic accuracy if detection algorithms take into account contextual images within the same patient to determine which images represent a melanoma.</p>\n</blockquote>\n<p>I think it strange to not separate public-private by patient. On the other hand, given the competition framework, at least from the probing point of view, it was a correct decision. If the split was by patient, it would have been possible to completely resolve the split in 100-200 submissions. There would be no information leak between the public and the private, but it is still 3k additional data-points to train on. And in the current split it is much harder to probe, but there is more information leak.</p>\n<h3>Using the probing in our final submissions</h3>\n<p>We were not able to use the obtained information to improve the score, our final submissions do not include any usage of it. We tried the following, at the end nothing of it made its way into the final submissions:</p>\n<ol>\n<li>Adding this information as features. During the training we simulated 30% randomly into \"public\" and the rest into \"private\", calculated number of ones and zeros in the \"public\" and added these features to help the \"private\" training.</li>\n<li>Added the test data to the training as in pseudo-labeling and added a new objective, - is there a \"public\" melanoma for a patient. For the train data we used the same simulation from the previous point. For the test data we used the probing output.</li>\n<li>In post-processing, scale up targets for the test patients whose sum of targets is smaller than number of public ones.</li>\n</ol>\n<p>With each of the above points there is a different story how much it improved the CV, LB and why it failed at the end, but not too interesting to go into details.</p>\n<h3>Public-private split, analysis</h3>\n<p>Back to the public-private split. Knowing number of public ones and zeros for 210 groups of test patients I was interested to understand whether the split is completely random or there is something. Spoiler: it seems to be completely random. </p>\n<p>Basically I want to check this hypothesis of splitting logic:</p>\n<blockquote>\n  <p><strong>The hypothesis</strong>: the splitting is done by randomly permuting <code>range(10982)</code> and taking the first 3294 (which is 30%) indexes. </p>\n</blockquote>\n<p>In order to test that hypothesis I first considered Kolmogorov-Smirnov test, but I don't know how to adjust it to different underlying distribution from sample to sample (because each group is different in size). Instead I calculated log-likelihood of my probing result, and then randomly simulated <code>10000</code> spits by following the hypothesis and calculated their log-likelihoods as well. The result is the following graph:</p>\n<p><img src=\"https://imgur.com/GD8wCUk.png\" alt=\"log-likelihood of split\"></p>\n<p>From the graph we can conclude that we indeed can not reject the hypothesis. For comparison there is also a red line, which indicated the log-likelihood of a split where each test patient is separated exactly 30% to public and 70% to private. We see that it is highly improbable under the current hypothesis, however it maximizes likelihood. Please share if you have thoughts or comments on how to do this statistical test properly.</p>\n<p>For comparison test-train separation is not completely random. Specifically, all images with height 1920 are in test. They are probably newer images (because they don't have in-patient age differences) and the organizers wanted them all in the test.</p>\n<h3>General thoughts</h3>\n<p>Looking at this competition design retrospectively, I think it was possible to completely hack it, that is, to categorize each test sample into one of three categories: private, public one or public zero. Probably separating between private and public zeros for low probability is not possible (almost no submissions which cut there) but it is also much less important. The outline of the approach can include 3 steps:</p>\n<ol>\n<li>Find all test patients with public ones</li>\n<li>Find exact samples with public ones (can also be done without going through the first point)</li>\n<li>The point 2 above greatly simplifies the MIP formulation if completed. This will allow adding any and all submissions as constraints. Specifically, all the submissions from the public kernels. That amount of information should be enough to resolve most of the samples.</li>\n</ol>\n<p>And then there is a question how much it helps for the private LB. In this competition in practice it didn't help, but I think it was close, so if I were Kaggle I would consider designing competitions which are more LB-probing-proof in the future.</p>\n<p>Thanks for reading this excessively long post!</p>",
      "rawMarkdown": "Hello Kagglers, it was a beautiful competition, hopefully you learnt a lot! \n\nI want to share with you how our team reached second place on the public LB, and some conclusions that we got along the way. \n\n### Probing algorithm\n\nWith 3 digits for the public score initially, and then the switch to 4 digits one month before the competition end, it was interesting to get as much information as possible from that output. I have tried LB probing before in many competitions, mainly by applying mixed integer programming (MIP), but it was the first time I encountered AUC ROC metric. And it turns out it is especially hard for MIP to crack, because it is order-based. \n\nAdditionally, a straightforward MIP formulation would result in about 10 million boolean variables, which in combination with a difficult problem is impossible to solve. It required a simplification and I used the following procedure:\n\n> 1. Split 690 test patients into few groups. For each group we define two variables that we want to find, - number of public ones and number of public zeros.\n2. With MIP formulation generate 50-500 feasible solutions. Where each submission translates into two constraints.\n3. Randomly generate 1000-2000 candidate submissions where each group has the same target value, and the values are different between groups. Select a submission which best separates the feasible solutions from the previous point.\n4. Repeat 2-3 until only one feasible solution is left, - thus we found number of public ones and zeros for those groups.\n5. Split several more groups into halves and repeat 2-4\n\nThe MIP problem I solved with Gurobi solver. Each already available submission is two constraint, - we know that Kaggle rounds down the scores, so we have upper and lower bound. The score is simply number of public zeros to the left of all public ones, divided by the total number of zeros multiplied by the total number of ones.\n\n### Final output\n\nAt the end I had 210 test patient groups, and knowing number of public ones and zeros for each group. If a group has a public one then this group contains only one test patient. This was the aim, - to know all test patients with public melanomas. I was splitting until I got there.\n\n### The submission scoring public 2nd\n\nFrom the output above I know which patients have public melanomas and which do not. Here I used our best submission and set all non-public-melanoma targets to zero, and also scaled up targets for patients who have sum of targets lower than number of public melanomas for that patient. This submission, of course, is useless for the private. It scored 0.6334 on private, 0.9926 on public.\n\n### General conclusions from the probing\n\nearly into the probing I concluded that:\n1. There are overall 78 public ones and 3216 public zeros. Making it almost exactly 30% of the test.\n2. The public-private split is not by patient. That is, images of every test patient are partially in public and partially in private. More on that below.\n\n### Public-private split\n\nFor a competition whose know-how is (a quote from the description below)\n\n> Dermatologists could enhance their diagnostic accuracy if detection algorithms take into account contextual images within the same patient to determine which images represent a melanoma.\n\nI think it strange to not separate public-private by patient. On the other hand, given the competition framework, at least from the probing point of view, it was a correct decision. If the split was by patient, it would have been possible to completely resolve the split in 100-200 submissions. There would be no information leak between the public and the private, but it is still 3k additional data-points to train on. And in the current split it is much harder to probe, but there is more information leak.\n\n### Using the probing in our final submissions\n\nWe were not able to use the obtained information to improve the score, our final submissions do not include any usage of it. We tried the following, at the end nothing of it made its way into the final submissions:\n1. Adding this information as features. During the training we simulated 30% randomly into \"public\" and the rest into \"private\", calculated number of ones and zeros in the \"public\" and added these features to help the \"private\" training.\n2. Added the test data to the training as in pseudo-labeling and added a new objective, - is there a \"public\" melanoma for a patient. For the train data we used the same simulation from the previous point. For the test data we used the probing output.\n3. In post-processing, scale up targets for the test patients whose sum of targets is smaller than number of public ones.\n\nWith each of the above points there is a different story how much it improved the CV, LB and why it failed at the end, but not too interesting to go into details.\n\n### Public-private split, analysis\n\nBack to the public-private split. Knowing number of public ones and zeros for 210 groups of test patients I was interested to understand whether the split is completely random or there is something. Spoiler: it seems to be completely random. \n\nBasically I want to check this hypothesis of splitting logic:\n\n> **The hypothesis**: the splitting is done by randomly permuting `range(10982)` and taking the first 3294 (which is 30%) indexes. \n\nIn order to test that hypothesis I first considered Kolmogorov-Smirnov test, but I don't know how to adjust it to different underlying distribution from sample to sample (because each group is different in size). Instead I calculated log-likelihood of my probing result, and then randomly simulated `10000` spits by following the hypothesis and calculated their log-likelihoods as well. The result is the following graph:\n\n![log-likelihood of split](https://imgur.com/GD8wCUk.png)\n\nFrom the graph we can conclude that we indeed can not reject the hypothesis. For comparison there is also a red line, which indicated the log-likelihood of a split where each test patient is separated exactly 30% to public and 70% to private. We see that it is highly improbable under the current hypothesis, however it maximizes likelihood. Please share if you have thoughts or comments on how to do this statistical test properly.\n\nFor comparison test-train separation is not completely random. Specifically, all images with height 1920 are in test. They are probably newer images (because they don't have in-patient age differences) and the organizers wanted them all in the test.\n\n### General thoughts\n\nLooking at this competition design retrospectively, I think it was possible to completely hack it, that is, to categorize each test sample into one of three categories: private, public one or public zero. Probably separating between private and public zeros for low probability is not possible (almost no submissions which cut there) but it is also much less important. The outline of the approach can include 3 steps:\n1. Find all test patients with public ones\n2. Find exact samples with public ones (can also be done without going through the first point)\n3. The point 2 above greatly simplifies the MIP formulation if completed. This will allow adding any and all submissions as constraints. Specifically, all the submissions from the public kernels. That amount of information should be enough to resolve most of the samples.\n\nAnd then there is a question how much it helps for the private LB. In this competition in practice it didn't help, but I think it was close, so if I were Kaggle I would consider designing competitions which are more LB-probing-proof in the future.\n\nThanks for reading this excessively long post!",
      "votes": null
    },
    {
      "id": "976122",
      "postDate": "08/18/2020 16:39:56",
      "content": "<p><a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> Will it be possible to share the Gurobi MIP model files</p>",
      "rawMarkdown": "zaharch Will it be possible to share the Gurobi MIP model files",
      "votes": null
    },
    {
      "id": "976318",
      "postDate": "08/18/2020 19:33:36",
      "content": "<p><a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> thanks for sharing.</p>\n<p>I wanted to have your thoughts on this: would it prevent LB probing if the public LB was not fixed (each time you make a NEW submission your score gets computed on a different random sample)? How would you hack this?</p>\n<p>I feel that LB would be safer with a moving random seed, but maybe I’m missing something.</p>",
      "rawMarkdown": "zaharch thanks for sharing.\n\nI wanted to have your thoughts on this: would it prevent LB probing if the public LB was not fixed (each time you make a NEW submission your score gets computed on a different random sample)? How would you hack this?\n\nI feel that LB would be safer with a moving random seed, but maybe I’m missing something.",
      "votes": null
    },
    {
      "id": "976334",
      "postDate": "08/18/2020 19:43:43",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> , answering you in a separate message to maximize horizontal space. The core of the model looks like this. Note that I do not attach the constants definitions used in the model, but I hope that is enough to get an impression how it looks like in general.</p>\n<pre><code>model = Model('Full')\n\nK = model.addVars(range(G), range(G), name=\"matrx\", vtype=GRB.CONTINUOUS, lb=0, ub=1500)\nV = model.addVars(range(S), name=\"value\", vtype=GRB.CONTINUOUS, lb=0)\nX = model.addVars(range(G), name=\"zeros\", vtype=GRB.INTEGER, lb=0, ub=2000)\nO = model.addVars(range(G), range(R), name=\"oness\", vtype=GRB.BINARY)\n\nmodel.setObjective(X[0], GRB.MINIMIZE)\n\nmodel.addConstr(((quicksum(r*O[i,r] for i in range(G) for r in range(R) if res_O[i] &lt; 0) +\n                  quicksum(res_O[i] for i in range(G) if res_O[i] &gt;= 0)) == 78), \"ones_sum\")\n\nmodel.addConstrs((X[i] &lt;= np.array(GRP_cnt[i]).sum() for i in range(G) if res_X[i] &lt; 0), \"zeros_lim\")\n\nmodel.addConstrs((X[i] &gt;= min_X[i] for i in range(G) if (((min_X[i] &gt;= 0) and (res_X[i] &lt; 0)))), \"min_X\")\nmodel.addConstrs((X[i] &lt;= max_X[i] for i in range(G) if (((max_X[i] &gt;= 0) and (res_X[i] &lt; 0)))), \"max_X\")\nmodel.addConstrs((X[i] == res_X[i] for i in range(G) if res_X[i] &gt;= 0), \"res_X\")\n\nmodel.addConstrs(((K[i1,i2] - r*X[i2] &gt;= -3000*(1-O[i1,r]))\n    for i1 in range(G) for i2 in range(G) for r in range(R) if res_O[i1] &lt; 0), \"ones_up\")\nmodel.addConstrs(((K[i1,i2] - r*X[i2] &lt;= 3000*(1-O[i1,r]))\n    for i1 in range(G) for i2 in range(G) for r in range(R) if res_O[i1] &lt; 0), \"ones_dn\")\nmodel.addConstrs(((K[i1,i2] == res_O[i1]*X[i2])\n    for i1 in range(G) for i2 in range(G) if res_O[i1] &gt;= 0), \"ones_eq\")\n\nmodel.addConstrs(((quicksum(O[i1,r] for r in range(R)) == 1) for i1 in range(G) if res_O[i1] &lt; 0), \"ones_sum\")\nmodel.addConstrs(((quicksum(O[i1,r] for r in range(R)) == 0) for i1 in range(G) if res_O[i1] &gt;= 0), \"ones_sum2\")\n\nmodel.addConstrs(((V[s] == quicksum(C[s,i1,i2]*K[i1,i2] \n    for i1 in range(G) for i2 in range(G))) for s in range(S)), \"sub_value\")\n\nmodel.addConstrs(((M*subs[s][1] &lt;= V[s]) for s in range(S)), \"sub_up\")\nmodel.addConstrs(((M*(subs[s][1]+0.0001) &gt;= V[s]) for s in range(S)), \"sub_dn\")\n</code></pre>",
      "rawMarkdown": "Hi @sirishks , answering you in a separate message to maximize horizontal space. The core of the model looks like this. Note that I do not attach the constants definitions used in the model, but I hope that is enough to get an impression how it looks like in general.\n\n```\nmodel = Model('Full')\n\nK = model.addVars(range(G), range(G), name=\"matrx\", vtype=GRB.CONTINUOUS, lb=0, ub=1500)\nV = model.addVars(range(S), name=\"value\", vtype=GRB.CONTINUOUS, lb=0)\nX = model.addVars(range(G), name=\"zeros\", vtype=GRB.INTEGER, lb=0, ub=2000)\nO = model.addVars(range(G), range(R), name=\"oness\", vtype=GRB.BINARY)\n\nmodel.setObjective(X[0], GRB.MINIMIZE)\n\nmodel.addConstr(((quicksum(r*O[i,r] for i in range(G) for r in range(R) if res_O[i] < 0) +\n                  quicksum(res_O[i] for i in range(G) if res_O[i] >= 0)) == 78), \"ones_sum\")\n\nmodel.addConstrs((X[i] <= np.array(GRP_cnt[i]).sum() for i in range(G) if res_X[i] < 0), \"zeros_lim\")\n\nmodel.addConstrs((X[i] >= min_X[i] for i in range(G) if (((min_X[i] >= 0) and (res_X[i] < 0)))), \"min_X\")\nmodel.addConstrs((X[i] <= max_X[i] for i in range(G) if (((max_X[i] >= 0) and (res_X[i] < 0)))), \"max_X\")\nmodel.addConstrs((X[i] == res_X[i] for i in range(G) if res_X[i] >= 0), \"res_X\")\n\nmodel.addConstrs(((K[i1,i2] - r*X[i2] >= -3000*(1-O[i1,r]))\n    for i1 in range(G) for i2 in range(G) for r in range(R) if res_O[i1] < 0), \"ones_up\")\nmodel.addConstrs(((K[i1,i2] - r*X[i2] <= 3000*(1-O[i1,r]))\n    for i1 in range(G) for i2 in range(G) for r in range(R) if res_O[i1] < 0), \"ones_dn\")\nmodel.addConstrs(((K[i1,i2] == res_O[i1]*X[i2])\n    for i1 in range(G) for i2 in range(G) if res_O[i1] >= 0), \"ones_eq\")\n\nmodel.addConstrs(((quicksum(O[i1,r] for r in range(R)) == 1) for i1 in range(G) if res_O[i1] < 0), \"ones_sum\")\nmodel.addConstrs(((quicksum(O[i1,r] for r in range(R)) == 0) for i1 in range(G) if res_O[i1] >= 0), \"ones_sum2\")\n\nmodel.addConstrs(((V[s] == quicksum(C[s,i1,i2]*K[i1,i2] \n    for i1 in range(G) for i2 in range(G))) for s in range(S)), \"sub_value\")\n\nmodel.addConstrs(((M*subs[s][1] <= V[s]) for s in range(S)), \"sub_up\")\nmodel.addConstrs(((M*(subs[s][1]+0.0001) >= V[s]) for s in range(S)), \"sub_dn\")\n```",
      "votes": null
    },
    {
      "id": "976348",
      "postDate": "08/18/2020 19:57:21",
      "content": "<p>First, you definitely want to leave the private part and never touch it, so I assume you imply each time taking 80-90% of the public part for scoring. That way it seems to be close to just adding some random noise to the scores. This will work more or less, but I think a cleaner and a simpler way to reduce the amount of information transferred is just show less decimal digits. In your approach, besides reduced information, there is also uncertainty. But for the probing it is generally not an issue, any uncertainty can be model as well, as long as there is information behind it. But it will definitely annoy people! So not sure what are the advantages of your idea.</p>",
      "rawMarkdown": "First, you definitely want to leave the private part and never touch it, so I assume you imply each time taking 80-90% of the public part for scoring. That way it seems to be close to just adding some random noise to the scores. This will work more or less, but I think a cleaner and a simpler way to reduce the amount of information transferred is just show less decimal digits. In your approach, besides reduced information, there is also uncertainty. But for the probing it is generally not an issue, any uncertainty can be model as well, as long as there is information behind it. But it will definitely annoy people! So not sure what are the advantages of your idea.",
      "votes": null
    },
    {
      "id": "976426",
      "postDate": "08/18/2020 20:53:29",
      "content": "<p>Hmm actually I was thinking of no «&nbsp;official private&nbsp;»  , it would be harder to probe but in fact you could probe final scores which is annoying. Well I guess it was a bad idea!</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hmm actually I was thinking of no « official private »  , it would be harder to probe but in fact you could probe final scores which is annoying. Well I guess it was a bad idea!\n\nThanks!",
      "votes": null
    },
    {
      "id": "977726",
      "postDate": "08/19/2020 17:08:54",
      "content": "<p>Hi, this is interesting, but I don't understand why you didn't get a roc-auc of 1.  </p>",
      "rawMarkdown": "Hi, this is interesting, but I don't understand why you didn't get a roc-auc of 1.",
      "votes": null
    },
    {
      "id": "981328",
      "postDate": "08/22/2020 11:26:17",
      "content": "<p>nice</p>",
      "rawMarkdown": "nice",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 976122,
      "author_name": "sirishks",
      "author_url": "",
      "post_date": "08/18/2020 16:39:56",
      "content": "<p><a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> Will it be possible to share the Gurobi MIP model files</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 976318,
      "author_name": "optimo",
      "author_url": "",
      "post_date": "08/18/2020 19:33:36",
      "content": "<p><a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> thanks for sharing.</p>\n<p>I wanted to have your thoughts on this: would it prevent LB probing if the public LB was not fixed (each time you make a NEW submission your score gets computed on a different random sample)? How would you hack this?</p>\n<p>I feel that LB would be safer with a moving random seed, but maybe I’m missing something.</p>",
      "votes": null,
      "replies": [
        {
          "id": 976348,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "08/18/2020 19:57:21",
          "content": "<p>First, you definitely want to leave the private part and never touch it, so I assume you imply each time taking 80-90% of the public part for scoring. That way it seems to be close to just adding some random noise to the scores. This will work more or less, but I think a cleaner and a simpler way to reduce the amount of information transferred is just show less decimal digits. In your approach, besides reduced information, there is also uncertainty. But for the probing it is generally not an issue, any uncertainty can be model as well, as long as there is information behind it. But it will definitely annoy people! So not sure what are the advantages of your idea.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 976426,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "08/18/2020 20:53:29",
          "content": "<p>Hmm actually I was thinking of no «&nbsp;official private&nbsp;»  , it would be harder to probe but in fact you could probe final scores which is annoying. Well I guess it was a bad idea!</p>\n<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 976334,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "08/18/2020 19:43:43",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> , answering you in a separate message to maximize horizontal space. The core of the model looks like this. Note that I do not attach the constants definitions used in the model, but I hope that is enough to get an impression how it looks like in general.</p>\n<pre><code>model = Model('Full')\n\nK = model.addVars(range(G), range(G), name=\"matrx\", vtype=GRB.CONTINUOUS, lb=0, ub=1500)\nV = model.addVars(range(S), name=\"value\", vtype=GRB.CONTINUOUS, lb=0)\nX = model.addVars(range(G), name=\"zeros\", vtype=GRB.INTEGER, lb=0, ub=2000)\nO = model.addVars(range(G), range(R), name=\"oness\", vtype=GRB.BINARY)\n\nmodel.setObjective(X[0], GRB.MINIMIZE)\n\nmodel.addConstr(((quicksum(r*O[i,r] for i in range(G) for r in range(R) if res_O[i] &lt; 0) +\n                  quicksum(res_O[i] for i in range(G) if res_O[i] &gt;= 0)) == 78), \"ones_sum\")\n\nmodel.addConstrs((X[i] &lt;= np.array(GRP_cnt[i]).sum() for i in range(G) if res_X[i] &lt; 0), \"zeros_lim\")\n\nmodel.addConstrs((X[i] &gt;= min_X[i] for i in range(G) if (((min_X[i] &gt;= 0) and (res_X[i] &lt; 0)))), \"min_X\")\nmodel.addConstrs((X[i] &lt;= max_X[i] for i in range(G) if (((max_X[i] &gt;= 0) and (res_X[i] &lt; 0)))), \"max_X\")\nmodel.addConstrs((X[i] == res_X[i] for i in range(G) if res_X[i] &gt;= 0), \"res_X\")\n\nmodel.addConstrs(((K[i1,i2] - r*X[i2] &gt;= -3000*(1-O[i1,r]))\n    for i1 in range(G) for i2 in range(G) for r in range(R) if res_O[i1] &lt; 0), \"ones_up\")\nmodel.addConstrs(((K[i1,i2] - r*X[i2] &lt;= 3000*(1-O[i1,r]))\n    for i1 in range(G) for i2 in range(G) for r in range(R) if res_O[i1] &lt; 0), \"ones_dn\")\nmodel.addConstrs(((K[i1,i2] == res_O[i1]*X[i2])\n    for i1 in range(G) for i2 in range(G) if res_O[i1] &gt;= 0), \"ones_eq\")\n\nmodel.addConstrs(((quicksum(O[i1,r] for r in range(R)) == 1) for i1 in range(G) if res_O[i1] &lt; 0), \"ones_sum\")\nmodel.addConstrs(((quicksum(O[i1,r] for r in range(R)) == 0) for i1 in range(G) if res_O[i1] &gt;= 0), \"ones_sum2\")\n\nmodel.addConstrs(((V[s] == quicksum(C[s,i1,i2]*K[i1,i2] \n    for i1 in range(G) for i2 in range(G))) for s in range(S)), \"sub_value\")\n\nmodel.addConstrs(((M*subs[s][1] &lt;= V[s]) for s in range(S)), \"sub_up\")\nmodel.addConstrs(((M*(subs[s][1]+0.0001) &gt;= V[s]) for s in range(S)), \"sub_dn\")\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 977726,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/19/2020 17:08:54",
      "content": "<p>Hi, this is interesting, but I don't understand why you didn't get a roc-auc of 1.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 981328,
      "author_name": "vinaypratap",
      "author_url": "",
      "post_date": "08/22/2020 11:26:17",
      "content": "<p>nice</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "975989": "Hello Kagglers, it was a beautiful competition, hopefully you learnt a lot! \n\nI want to share with you how our team reached second place on the public LB, and some conclusions that we got along the way. \n\n### Probing algorithm\n\nWith 3 digits for the public score initially, and then the switch to 4 digits one month before the competition end, it was interesting to get as much information as possible from that output. I have tried LB probing before in many competitions, mainly by applying mixed integer programming (MIP), but it was the first time I encountered AUC ROC metric. And it turns out it is especially hard for MIP to crack, because it is order-based. \n\nAdditionally, a straightforward MIP formulation would result in about 10 million boolean variables, which in combination with a difficult problem is impossible to solve. It required a simplification and I used the following procedure:\n\n> 1. Split 690 test patients into few groups. For each group we define two variables that we want to find, - number of public ones and number of public zeros.\n2. With MIP formulation generate 50-500 feasible solutions. Where each submission translates into two constraints.\n3. Randomly generate 1000-2000 candidate submissions where each group has the same target value, and the values are different between groups. Select a submission which best separates the feasible solutions from the previous point.\n4. Repeat 2-3 until only one feasible solution is left, - thus we found number of public ones and zeros for those groups.\n5. Split several more groups into halves and repeat 2-4\n\nThe MIP problem I solved with Gurobi solver. Each already available submission is two constraint, - we know that Kaggle rounds down the scores, so we have upper and lower bound. The score is simply number of public zeros to the left of all public ones, divided by the total number of zeros multiplied by the total number of ones.\n\n### Final output\n\nAt the end I had 210 test patient groups, and knowing number of public ones and zeros for each group. If a group has a public one then this group contains only one test patient. This was the aim, - to know all test patients with public melanomas. I was splitting until I got there.\n\n### The submission scoring public 2nd\n\nFrom the output above I know which patients have public melanomas and which do not. Here I used our best submission and set all non-public-melanoma targets to zero, and also scaled up targets for patients who have sum of targets lower than number of public melanomas for that patient. This submission, of course, is useless for the private. It scored 0.6334 on private, 0.9926 on public.\n\n### General conclusions from the probing\n\nearly into the probing I concluded that:\n1. There are overall 78 public ones and 3216 public zeros. Making it almost exactly 30% of the test.\n2. The public-private split is not by patient. That is, images of every test patient are partially in public and partially in private. More on that below.\n\n### Public-private split\n\nFor a competition whose know-how is (a quote from the description below)\n\n> Dermatologists could enhance their diagnostic accuracy if detection algorithms take into account contextual images within the same patient to determine which images represent a melanoma.\n\nI think it strange to not separate public-private by patient. On the other hand, given the competition framework, at least from the probing point of view, it was a correct decision. If the split was by patient, it would have been possible to completely resolve the split in 100-200 submissions. There would be no information leak between the public and the private, but it is still 3k additional data-points to train on. And in the current split it is much harder to probe, but there is more information leak.\n\n### Using the probing in our final submissions\n\nWe were not able to use the obtained information to improve the score, our final submissions do not include any usage of it. We tried the following, at the end nothing of it made its way into the final submissions:\n1. Adding this information as features. During the training we simulated 30% randomly into \"public\" and the rest into \"private\", calculated number of ones and zeros in the \"public\" and added these features to help the \"private\" training.\n2. Added the test data to the training as in pseudo-labeling and added a new objective, - is there a \"public\" melanoma for a patient. For the train data we used the same simulation from the previous point. For the test data we used the probing output.\n3. In post-processing, scale up targets for the test patients whose sum of targets is smaller than number of public ones.\n\nWith each of the above points there is a different story how much it improved the CV, LB and why it failed at the end, but not too interesting to go into details.\n\n### Public-private split, analysis\n\nBack to the public-private split. Knowing number of public ones and zeros for 210 groups of test patients I was interested to understand whether the split is completely random or there is something. Spoiler: it seems to be completely random. \n\nBasically I want to check this hypothesis of splitting logic:\n\n> **The hypothesis**: the splitting is done by randomly permuting `range(10982)` and taking the first 3294 (which is 30%) indexes. \n\nIn order to test that hypothesis I first considered Kolmogorov-Smirnov test, but I don't know how to adjust it to different underlying distribution from sample to sample (because each group is different in size). Instead I calculated log-likelihood of my probing result, and then randomly simulated `10000` spits by following the hypothesis and calculated their log-likelihoods as well. The result is the following graph:\n\n![log-likelihood of split](https://imgur.com/GD8wCUk.png)\n\nFrom the graph we can conclude that we indeed can not reject the hypothesis. For comparison there is also a red line, which indicated the log-likelihood of a split where each test patient is separated exactly 30% to public and 70% to private. We see that it is highly improbable under the current hypothesis, however it maximizes likelihood. Please share if you have thoughts or comments on how to do this statistical test properly.\n\nFor comparison test-train separation is not completely random. Specifically, all images with height 1920 are in test. They are probably newer images (because they don't have in-patient age differences) and the organizers wanted them all in the test.\n\n### General thoughts\n\nLooking at this competition design retrospectively, I think it was possible to completely hack it, that is, to categorize each test sample into one of three categories: private, public one or public zero. Probably separating between private and public zeros for low probability is not possible (almost no submissions which cut there) but it is also much less important. The outline of the approach can include 3 steps:\n1. Find all test patients with public ones\n2. Find exact samples with public ones (can also be done without going through the first point)\n3. The point 2 above greatly simplifies the MIP formulation if completed. This will allow adding any and all submissions as constraints. Specifically, all the submissions from the public kernels. That amount of information should be enough to resolve most of the samples.\n\nAnd then there is a question how much it helps for the private LB. In this competition in practice it didn't help, but I think it was close, so if I were Kaggle I would consider designing competitions which are more LB-probing-proof in the future.\n\nThanks for reading this excessively long post!",
    "976122": "zaharch Will it be possible to share the Gurobi MIP model files",
    "976318": "zaharch thanks for sharing.\n\nI wanted to have your thoughts on this: would it prevent LB probing if the public LB was not fixed (each time you make a NEW submission your score gets computed on a different random sample)? How would you hack this?\n\nI feel that LB would be safer with a moving random seed, but maybe I’m missing something.",
    "976334": "Hi @sirishks , answering you in a separate message to maximize horizontal space. The core of the model looks like this. Note that I do not attach the constants definitions used in the model, but I hope that is enough to get an impression how it looks like in general.\n\n```\nmodel = Model('Full')\n\nK = model.addVars(range(G), range(G), name=\"matrx\", vtype=GRB.CONTINUOUS, lb=0, ub=1500)\nV = model.addVars(range(S), name=\"value\", vtype=GRB.CONTINUOUS, lb=0)\nX = model.addVars(range(G), name=\"zeros\", vtype=GRB.INTEGER, lb=0, ub=2000)\nO = model.addVars(range(G), range(R), name=\"oness\", vtype=GRB.BINARY)\n\nmodel.setObjective(X[0], GRB.MINIMIZE)\n\nmodel.addConstr(((quicksum(r*O[i,r] for i in range(G) for r in range(R) if res_O[i] < 0) +\n                  quicksum(res_O[i] for i in range(G) if res_O[i] >= 0)) == 78), \"ones_sum\")\n\nmodel.addConstrs((X[i] <= np.array(GRP_cnt[i]).sum() for i in range(G) if res_X[i] < 0), \"zeros_lim\")\n\nmodel.addConstrs((X[i] >= min_X[i] for i in range(G) if (((min_X[i] >= 0) and (res_X[i] < 0)))), \"min_X\")\nmodel.addConstrs((X[i] <= max_X[i] for i in range(G) if (((max_X[i] >= 0) and (res_X[i] < 0)))), \"max_X\")\nmodel.addConstrs((X[i] == res_X[i] for i in range(G) if res_X[i] >= 0), \"res_X\")\n\nmodel.addConstrs(((K[i1,i2] - r*X[i2] >= -3000*(1-O[i1,r]))\n    for i1 in range(G) for i2 in range(G) for r in range(R) if res_O[i1] < 0), \"ones_up\")\nmodel.addConstrs(((K[i1,i2] - r*X[i2] <= 3000*(1-O[i1,r]))\n    for i1 in range(G) for i2 in range(G) for r in range(R) if res_O[i1] < 0), \"ones_dn\")\nmodel.addConstrs(((K[i1,i2] == res_O[i1]*X[i2])\n    for i1 in range(G) for i2 in range(G) if res_O[i1] >= 0), \"ones_eq\")\n\nmodel.addConstrs(((quicksum(O[i1,r] for r in range(R)) == 1) for i1 in range(G) if res_O[i1] < 0), \"ones_sum\")\nmodel.addConstrs(((quicksum(O[i1,r] for r in range(R)) == 0) for i1 in range(G) if res_O[i1] >= 0), \"ones_sum2\")\n\nmodel.addConstrs(((V[s] == quicksum(C[s,i1,i2]*K[i1,i2] \n    for i1 in range(G) for i2 in range(G))) for s in range(S)), \"sub_value\")\n\nmodel.addConstrs(((M*subs[s][1] <= V[s]) for s in range(S)), \"sub_up\")\nmodel.addConstrs(((M*(subs[s][1]+0.0001) >= V[s]) for s in range(S)), \"sub_dn\")\n```",
    "976348": "First, you definitely want to leave the private part and never touch it, so I assume you imply each time taking 80-90% of the public part for scoring. That way it seems to be close to just adding some random noise to the scores. This will work more or less, but I think a cleaner and a simpler way to reduce the amount of information transferred is just show less decimal digits. In your approach, besides reduced information, there is also uncertainty. But for the probing it is generally not an issue, any uncertainty can be model as well, as long as there is information behind it. But it will definitely annoy people! So not sure what are the advantages of your idea.",
    "976426": "Hmm actually I was thinking of no « official private »  , it would be harder to probe but in fact you could probe final scores which is annoying. Well I guess it was a bad idea!\n\nThanks!",
    "977726": "Hi, this is interesting, but I don't understand why you didn't get a roc-auc of 1.",
    "981328": "nice"
  },
  "source": "meta"
}