{
  "id": 76799,
  "title": "What's your estimate of shakeup?",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/76799",
  "author_name": "",
  "post_date": "2019-01-06T22:36:01.202929600Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>A way to get a sense of how LB behaves is to pick two submissions A, B with the roughly the same LB (+-~0.002), and calculate:</p>\n\n<ol>\n<li>if A is the ground truth what would be the macro F1 of B</li>\n<li>the other way around</li>\n</ol>\n\n<p>It would be great if you can report your results here!</p>",
  "messages": [
    {
      "id": "451343",
      "postDate": "01/06/2019 22:36:01",
      "content": "<p>A way to get a sense of how LB behaves is to pick two submissions A, B with the roughly the same LB (+-~0.002), and calculate:</p>\n\n<ol>\n<li>if A is the ground truth what would be the macro F1 of B</li>\n<li>the other way around</li>\n</ol>\n\n<p>It would be great if you can report your results here!</p>",
      "rawMarkdown": "A way to get a sense of how LB behaves is to pick two submissions A, B with the roughly the same LB (+-~0.002), and calculate:\n\n1. if A is the ground truth what would be the macro F1 of B\n2. the other way around\n\nIt would be great if you can report your results here!",
      "votes": null
    },
    {
      "id": "451351",
      "postDate": "01/06/2019 23:17:52",
      "content": "<p>How exactly are you going to estimate the shake-up if someone was to give you these two scores?\nE.g. if two my submissions are scored 0.92 and 0.95 respectively, and you don't know whether they are similar (e.g. single network with different thresholds and TTA) or not, what conclusion can you draw?</p>",
      "rawMarkdown": "How exactly are you going to estimate the shake-up if someone was to give you these two scores?\nE.g. if two my submissions are scored 0.92 and 0.95 respectively, and you don't know whether they are similar (e.g. single network with different thresholds and TTA) or not, what conclusion can you draw?",
      "votes": null
    },
    {
      "id": "451429",
      "postDate": "01/07/2019 04:54:02",
      "content": "<p>Related question:</p>\n\n<p>Two submissions with exactly the same training run - one using the leak postings to improve the public LB and the other without harvesting that improvement. Which would you ask them to use as your entry?</p>\n\n<p>My guess: does not matter, assuming the organizers posted comments are accurate. </p>",
      "rawMarkdown": "Related question:\n\nTwo submissions with exactly the same training run - one using the leak postings to improve the public LB and the other without harvesting that improvement. Which would you ask them to use as your entry?\n\nMy guess: does not matter, assuming the organizers posted comments are accurate.",
      "votes": null
    },
    {
      "id": "451588",
      "postDate": "01/07/2019 10:03:09",
      "content": "<p><a href=\"/petewills\">@petewills</a> agree though perhaps safer to include the leak. Shake-up could be huge in my view...very interested to know how much HPA data was necessary for a strong model - makes choosing the final submission pretty tricky.</p>",
      "rawMarkdown": "petewills agree though perhaps safer to include the leak. Shake-up could be huge in my view...very interested to know how much HPA data was necessary for a strong model - makes choosing the final submission pretty tricky.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 451351,
      "author_name": "hokmund",
      "author_url": "",
      "post_date": "01/06/2019 23:17:52",
      "content": "<p>How exactly are you going to estimate the shake-up if someone was to give you these two scores?\nE.g. if two my submissions are scored 0.92 and 0.95 respectively, and you don't know whether they are similar (e.g. single network with different thresholds and TTA) or not, what conclusion can you draw?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 451429,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "01/07/2019 04:54:02",
      "content": "<p>Related question:</p>\n\n<p>Two submissions with exactly the same training run - one using the leak postings to improve the public LB and the other without harvesting that improvement. Which would you ask them to use as your entry?</p>\n\n<p>My guess: does not matter, assuming the organizers posted comments are accurate. </p>",
      "votes": null,
      "replies": [
        {
          "id": 451588,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "01/07/2019 10:03:09",
          "content": "<p><a href=\"/petewills\">@petewills</a> agree though perhaps safer to include the leak. Shake-up could be huge in my view...very interested to know how much HPA data was necessary for a strong model - makes choosing the final submission pretty tricky.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "451343": "A way to get a sense of how LB behaves is to pick two submissions A, B with the roughly the same LB (+-~0.002), and calculate:\n\n1. if A is the ground truth what would be the macro F1 of B\n2. the other way around\n\nIt would be great if you can report your results here!",
    "451351": "How exactly are you going to estimate the shake-up if someone was to give you these two scores?\nE.g. if two my submissions are scored 0.92 and 0.95 respectively, and you don't know whether they are similar (e.g. single network with different thresholds and TTA) or not, what conclusion can you draw?",
    "451429": "Related question:\n\nTwo submissions with exactly the same training run - one using the leak postings to improve the public LB and the other without harvesting that improvement. Which would you ask them to use as your entry?\n\nMy guess: does not matter, assuming the organizers posted comments are accurate.",
    "451588": "petewills agree though perhaps safer to include the leak. Shake-up could be huge in my view...very interested to know how much HPA data was necessary for a strong model - makes choosing the final submission pretty tricky."
  },
  "source": "meta"
}