{
  "id": 419110,
  "title": "Sparse annotations in dataset2?",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/discussion/419110",
  "author_name": "",
  "post_date": "2023-06-24T08:07:49.060429400Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I noticed that if I train with dataset2, my LB is lower. Maybe someone asked/answered this question already.<br>\nDoes \"sparse\" mean not complete???<br>\nFrom the data section:<br>\n\"Dataset 2 comprises the remaining tiles from these same WSIs and contain <strong>sparse</strong> annotations that have not been expert reviewed.\"</p>",
  "messages": [
    {
      "id": "2315554",
      "postDate": "06/24/2023 08:07:49",
      "content": "<p>I noticed that if I train with dataset2, my LB is lower. Maybe someone asked/answered this question already.<br>\nDoes \"sparse\" mean not complete???<br>\nFrom the data section:<br>\n\"Dataset 2 comprises the remaining tiles from these same WSIs and contain <strong>sparse</strong> annotations that have not been expert reviewed.\"</p>",
      "rawMarkdown": "I noticed that if I train with dataset2, my LB is lower. Maybe someone asked/answered this question already.\nDoes \"sparse\" mean not complete???\nFrom the data section:\n\"Dataset 2 comprises the remaining tiles from these same WSIs and contain **sparse** annotations that have not been expert reviewed.\"",
      "votes": null
    },
    {
      "id": "2318757",
      "postDate": "06/26/2023 15:02:48",
      "content": "<p>Hello, yes the Dataset 2 annotations may be incomplete and have not gone through the final expert validation. </p>",
      "rawMarkdown": "Hello, yes the Dataset 2 annotations may be incomplete and have not gone through the final expert validation.",
      "votes": null
    },
    {
      "id": "2319224",
      "postDate": "06/27/2023 01:50:00",
      "content": "<p>\"Does \"sparse\" mean not complete???\"</p>\n<p>i understand what it means:</p>\n<ol>\n<li>split dataset1 into train and val. train a model</li>\n<li>apply model on val. then apply model on dataset.2</li>\n</ol>\n<p>The difference in results and observation tells you what is different<br>\ne.g. dataset.2 has many more \"fp\" than dataset.1 val. this implies that dataset.2 might have \"missing annotations\"</p>\n<hr>\n<p>also  use tsne to plot dataset1 train and val. train and dataset.2 using external model (e.g image or SAM model)<br>\nand your trained model.<br>\nThis shows the smiliarity distance of the dataset</p>",
      "rawMarkdown": "\"Does \"sparse\" mean not complete???\"\n\ni understand what it means:\n1. split dataset1 into train and val. train a model\n2. apply model on val. then apply model on dataset.2\n\nThe difference in results and observation tells you what is different\ne.g. dataset.2 has many more \"fp\" than dataset.1 val. this implies that dataset.2 might have \"missing annotations\"\n\n---\n\nalso  use tsne to plot dataset1 train and val. train and dataset.2 using external model (e.g image or SAM model)\nand your trained model.\nThis shows the smiliarity distance of the dataset",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2318757,
      "author_name": "yashvrdnjain",
      "author_url": "",
      "post_date": "06/26/2023 15:02:48",
      "content": "<p>Hello, yes the Dataset 2 annotations may be incomplete and have not gone through the final expert validation. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2319224,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "06/27/2023 01:50:00",
      "content": "<p>\"Does \"sparse\" mean not complete???\"</p>\n<p>i understand what it means:</p>\n<ol>\n<li>split dataset1 into train and val. train a model</li>\n<li>apply model on val. then apply model on dataset.2</li>\n</ol>\n<p>The difference in results and observation tells you what is different<br>\ne.g. dataset.2 has many more \"fp\" than dataset.1 val. this implies that dataset.2 might have \"missing annotations\"</p>\n<hr>\n<p>also  use tsne to plot dataset1 train and val. train and dataset.2 using external model (e.g image or SAM model)<br>\nand your trained model.<br>\nThis shows the smiliarity distance of the dataset</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2315554": "I noticed that if I train with dataset2, my LB is lower. Maybe someone asked/answered this question already.\nDoes \"sparse\" mean not complete???\nFrom the data section:\n\"Dataset 2 comprises the remaining tiles from these same WSIs and contain **sparse** annotations that have not been expert reviewed.\"",
    "2318757": "Hello, yes the Dataset 2 annotations may be incomplete and have not gone through the final expert validation.",
    "2319224": "\"Does \"sparse\" mean not complete???\"\n\ni understand what it means:\n1. split dataset1 into train and val. train a model\n2. apply model on val. then apply model on dataset.2\n\nThe difference in results and observation tells you what is different\ne.g. dataset.2 has many more \"fp\" than dataset.1 val. this implies that dataset.2 might have \"missing annotations\"\n\n---\n\nalso  use tsne to plot dataset1 train and val. train and dataset.2 using external model (e.g image or SAM model)\nand your trained model.\nThis shows the smiliarity distance of the dataset"
  },
  "source": "meta"
}