{
  "id": 417012,
  "title": "Test data annotation",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/discussion/417012",
  "author_name": "Sasha Mogilevskii",
  "post_date": "2023-06-13T20:45:04.365000",
  "votes": 17,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hello everyone. After studying the data from datasets 1 and 2, I noticed that the data is annotated in completely different ways. In some instances, the annotations follow the lines, while in other images, they deviate from them. I would like to inquire with the competition organizers about how the test data was annotated or selected. Are these strictly annotated data? Or is the annotation of a similar level? Th u </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6197543%2F9b3efffbbe418bfb299d90f66b0d79de%2FScreenshot_4.jpg?generation=1686689058146565&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6197543%2F47596f0b79ed5d8a8120969ee1411426%2FScreenshot_3.jpg?generation=1686689082173218&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6197543%2Fb2d5de8f842635d33d280d7405104789%2FScreenshot_2.jpg?generation=1686689091526052&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2301372,
      "postDate": "2023-06-13T20:45:04.367Z",
      "content": "<p>Hello everyone. After studying the data from datasets 1 and 2, I noticed that the data is annotated in completely different ways. In some instances, the annotations follow the lines, while in other images, they deviate from them. I would like to inquire with the competition organizers about how the test data was annotated or selected. Are these strictly annotated data? Or is the annotation of a similar level? Th u </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6197543%2F9b3efffbbe418bfb299d90f66b0d79de%2FScreenshot_4.jpg?generation=1686689058146565&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6197543%2F47596f0b79ed5d8a8120969ee1411426%2FScreenshot_3.jpg?generation=1686689082173218&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6197543%2Fb2d5de8f842635d33d280d7405104789%2FScreenshot_2.jpg?generation=1686689091526052&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hello everyone. After studying the data from datasets 1 and 2, I noticed that the data is annotated in completely different ways. In some instances, the annotations follow the lines, while in other images, they deviate from them. I would like to inquire with the competition organizers about how the test data was annotated or selected. Are these strictly annotated data? Or is the annotation of a similar level? Th u \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6197543%2F9b3efffbbe418bfb299d90f66b0d79de%2FScreenshot_4.jpg?generation=1686689058146565&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6197543%2F47596f0b79ed5d8a8120969ee1411426%2FScreenshot_3.jpg?generation=1686689082173218&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6197543%2Fb2d5de8f842635d33d280d7405104789%2FScreenshot_2.jpg?generation=1686689091526052&alt=media)",
      "votes": 17
    },
    {
      "id": 2301416,
      "postDate": "2023-06-13T22:39:24.207Z",
      "content": "<p>great finding, so probably it is one of the reasons why dilation boost both local CV and public lb score <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/416901\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/416901</a></p>",
      "rawMarkdown": "great finding, so probably it is one of the reasons why dilation boost both local CV and public lb score https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/416901",
      "votes": 5,
      "replies": [
        {
          "id": 2301489,
          "postDate": "2023-06-14T01:43:30.783Z",
          "content": "<p>So it could boost your LB, high probability that it harms private lb</p>",
          "rawMarkdown": "So it could boost your LB, high probability that it harms private lb",
          "replies": [
            {
              "id": 2301577,
              "postDate": "2023-06-14T03:29:07.080Z",
              "content": "<p>But public test set is also from dataset1, which means smaller contours. Why it benefits from dilation so much?</p>",
              "rawMarkdown": "But public test set is also from dataset1, which means smaller contours. Why it benefits from dilation so much?",
              "votes": 3
            },
            {
              "id": 2301606,
              "postDate": "2023-06-14T04:06:37.197Z",
              "content": "<p>dataset1 is split into public test and private test, I think it means they are sharing the same labeling rules.</p>",
              "rawMarkdown": "dataset1 is split into public test and private test, I think it means they are sharing the same labeling rules."
            },
            {
              "id": 2301668,
              "postDate": "2023-06-14T04:59:27.790Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 2301502,
          "postDate": "2023-06-14T01:58:41.307Z",
          "content": "<p>We should be sure by experimentation, but that may not be relevant to the diffrence of dataset.</p>\n<p>The reason is that dataset1 seems to fit the actual contours better, while dataset2 seems to create masks those are rougher than the actual contours.<br>\nIf dataset2 was used for the test data, then the hypothesis might fit, but in reality dataset1 is used.</p>\n<p>If your post-processing improves the score in all CVs (especially in dataset1) and in the test data, then optimization to loss would not mean optimization to metrics(mAP).</p>\n<p>By the way, your findings are really interesting and very surprising for me. Thanks for sharing!</p>",
          "rawMarkdown": "We should be sure by experimentation, but that may not be relevant to the diffrence of dataset.\n\nThe reason is that dataset1 seems to fit the actual contours better, while dataset2 seems to create masks those are rougher than the actual contours.\nIf dataset2 was used for the test data, then the hypothesis might fit, but in reality dataset1 is used.\n\nIf your post-processing improves the score in all CVs (especially in dataset1) and in the test data, then optimization to loss would not mean optimization to metrics(mAP).\n\nBy the way, your findings are really interesting and very surprising for me. Thanks for sharing!"
        }
      ]
    },
    {
      "id": 2304251,
      "postDate": "2023-06-15T21:03:59.187Z",
      "content": "<p>Hello,</p>\n<p>Thank you for your question regarding the annotations in the dataset. First, I would like to note that generation of manual annotations is a time consuming and expensive process. Some noise in the labels is to be expected with any such dataset. In this particular dataset, the borders of microvasculature structures are not always clearly distinguishable given the tissue thickness and resolution of the images. For example, when annotating peritubular capillaries, the endothelial cell membrane may be indistinguishable from the tubular epithelial membrane, leading to some overlap of annotation borders with tubular epithelium borders. One other thing to note is that often the borders of the annotations themselves can obscure important visual information used to distinguish the vessels from surrounding structures. It is helpful to have a side by side comparison of both annotated and un-annotated images to distinguish vessels. </p>\n<p>Best,</p>\n<ul>\n<li>Kate Gustilo, Research Analyst and Anatomist</li>\n</ul>",
      "rawMarkdown": "Hello,\n\nThank you for your question regarding the annotations in the dataset. First, I would like to note that generation of manual annotations is a time consuming and expensive process. Some noise in the labels is to be expected with any such dataset. In this particular dataset, the borders of microvasculature structures are not always clearly distinguishable given the tissue thickness and resolution of the images. For example, when annotating peritubular capillaries, the endothelial cell membrane may be indistinguishable from the tubular epithelial membrane, leading to some overlap of annotation borders with tubular epithelium borders. One other thing to note is that often the borders of the annotations themselves can obscure important visual information used to distinguish the vessels from surrounding structures. It is helpful to have a side by side comparison of both annotated and un-annotated images to distinguish vessels. \n\nBest,\n - Kate Gustilo, Research Analyst and Anatomist",
      "votes": 3,
      "replies": [
        {
          "id": 2304370,
          "postDate": "2023-06-16T00:19:34.850Z",
          "content": "<p>Thanks for reply!<br>\nI also respect the great efforts of the annotators, thank you all!</p>\n<p>The only thing that concerns me is that the annotation trends are very different even in dataset1. Perhaps this is unavoidable since multiple annotators are working on it.</p>\n<p>I just want to be clear about the distribution of test data.<br>\nAre separate annotators responsible for train, public and private testing respectively?<br>\nOr is each annotator randomly distributed to each dataset?<br>\nFor example, is annotator 1 responsible for both the train and the public and private test datasets?</p>\n<p>If the unknown annotator is only included in the private test, then it would be much more difficult to create a robust model. The post-processing such as dilation and erosion can cause huge shake-up.</p>\n<p>We look forward to your reply. Thanks!</p>",
          "rawMarkdown": "Thanks for reply!\nI also respect the great efforts of the annotators, thank you all!\n\nThe only thing that concerns me is that the annotation trends are very different even in dataset1. Perhaps this is unavoidable since multiple annotators are working on it.\n\nI just want to be clear about the distribution of test data.\nAre separate annotators responsible for train, public and private testing respectively?\nOr is each annotator randomly distributed to each dataset?\nFor example, is annotator 1 responsible for both the train and the public and private test datasets?\n\nIf the unknown annotator is only included in the private test, then it would be much more difficult to create a robust model. The post-processing such as dilation and erosion can cause huge shake-up.\n\nWe look forward to your reply. Thanks!",
          "votes": 2,
          "replies": [
            {
              "id": 2305579,
              "postDate": "2023-06-16T19:23:42.740Z",
              "content": "<p>YYama, </p>\n<p>The initial annotations were completed by multiple annotators, however all data was validated by the same expert. Thus, differences in annotation trends are minimal and have gone through the same validation process for the training, private and public test sets. I hope this is helpful.</p>\n<p>Best,<br>\nKate Gustilo</p>",
              "rawMarkdown": "YYama, \n\nThe initial annotations were completed by multiple annotators, however all data was validated by the same expert. Thus, differences in annotation trends are minimal and have gone through the same validation process for the training, private and public test sets. I hope this is helpful.\n\nBest,\nKate Gustilo",
              "votes": 5
            },
            {
              "id": 2308978,
              "postDate": "2023-06-19T10:31:02.853Z",
              "content": "<p>Thanks for the reply!</p>\n<p>So there is no bias in the annotation trend from one dataset to another.<br>\nVery good news!</p>",
              "rawMarkdown": "Thanks for the reply!\n\nSo there is no bias in the annotation trend from one dataset to another.\nVery good news!"
            }
          ]
        }
      ]
    },
    {
      "id": 2301431,
      "postDate": "2023-06-13T23:50:20.010Z",
      "content": "<p>\"The competition data comprises tiles extracted from five Whole Slide Images (WSI) split into two datasets. Tiles from Dataset 1 have annotations that have been expert reviewed. Dataset 2 comprises the remaining tiles from these same WSIs and contain sparse annotations that have not been expert reviewed.\"<br>\n<a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data</a></p>\n<p>This is how the dataset description states it. In other words, the annotations in dataset1 are of good accuracy and the annotations in dataset2 are of slightly poorer quality. This information may have an impact on the validation method.</p>\n<p>EDIT:<br>\nAccording to the description, only data from dataset1 is used for the test data. This is very important information.</p>",
      "rawMarkdown": "\"The competition data comprises tiles extracted from five Whole Slide Images (WSI) split into two datasets. Tiles from Dataset 1 have annotations that have been expert reviewed. Dataset 2 comprises the remaining tiles from these same WSIs and contain sparse annotations that have not been expert reviewed.\"\nhttps://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data\n\nThis is how the dataset description states it. In other words, the annotations in dataset1 are of good accuracy and the annotations in dataset2 are of slightly poorer quality. This information may have an impact on the validation method.\n\nEDIT:\nAccording to the description, only data from dataset1 is used for the test data. This is very important information.",
      "votes": 4,
      "replies": [
        {
          "id": 2302618,
          "postDate": "2023-06-14T16:49:29.300Z",
          "content": "<p>Yes, but i send img from data_1. For example this image from data_1 too. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6197543%2Fa8fc9e565fae02cb36b7107032eb9aa8%2F1.jpg?generation=1686761343902051&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6197543%2Fe02e1ca0b30bb65d400392732bcf4b78%2FScreenshot_2.jpg?generation=1686761353747613&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Yes, but i send img from data_1. For example this image from data_1 too. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6197543%2Fa8fc9e565fae02cb36b7107032eb9aa8%2F1.jpg?generation=1686761343902051&alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6197543%2Fe02e1ca0b30bb65d400392732bcf4b78%2FScreenshot_2.jpg?generation=1686761353747613&alt=media)",
          "votes": 2,
          "replies": [
            {
              "id": 2302845,
              "postDate": "2023-06-14T22:04:37.953Z",
              "content": "<p>Ah, I see. So both were from DATASET1. Thank you for sharing!</p>\n<p>Then I can only assume that they were annotated by separate annotators.<br>\nSome of them have gone beyond the extravascular membrane and some of them have even included tubular epithelial cells.</p>\n<p>I would like the host to disclose information about the annotator too. Are the test sets made by more than one annotator, and is there any bias between public and private annotators?</p>",
              "rawMarkdown": "Ah, I see. So both were from DATASET1. Thank you for sharing!\n\nThen I can only assume that they were annotated by separate annotators.\nSome of them have gone beyond the extravascular membrane and some of them have even included tubular epithelial cells.\n\nI would like the host to disclose information about the annotator too. Are the test sets made by more than one annotator, and is there any bias between public and private annotators?",
              "votes": 2
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2301416,
      "author_name": "Kostiantyn Maksymov",
      "author_url": "",
      "post_date": "2023-06-13T22:39:24.207000",
      "content": "<p>great finding, so probably it is one of the reasons why dilation boost both local CV and public lb score <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/416901\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/416901</a></p>",
      "votes": 5,
      "replies": [
        {
          "id": 2301489,
          "author_name": "Phat Tran",
          "author_url": "",
          "post_date": "2023-06-14T01:43:30.783000",
          "content": "<p>So it could boost your LB, high probability that it harms private lb</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2301577,
              "author_name": "Leon",
              "author_url": "",
              "post_date": "2023-06-14T03:29:07.080000",
              "content": "<p>But public test set is also from dataset1, which means smaller contours. Why it benefits from dilation so much?</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2301606,
              "author_name": "Yi Wu",
              "author_url": "",
              "post_date": "2023-06-14T04:06:37.197000",
              "content": "<p>dataset1 is split into public test and private test, I think it means they are sharing the same labeling rules.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2301668,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-06-14T04:59:27.790000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2301502,
          "author_name": "YYama",
          "author_url": "",
          "post_date": "2023-06-14T01:58:41.307000",
          "content": "<p>We should be sure by experimentation, but that may not be relevant to the diffrence of dataset.</p>\n<p>The reason is that dataset1 seems to fit the actual contours better, while dataset2 seems to create masks those are rougher than the actual contours.<br>\nIf dataset2 was used for the test data, then the hypothesis might fit, but in reality dataset1 is used.</p>\n<p>If your post-processing improves the score in all CVs (especially in dataset1) and in the test data, then optimization to loss would not mean optimization to metrics(mAP).</p>\n<p>By the way, your findings are really interesting and very surprising for me. Thanks for sharing!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2304251,
      "author_name": "Katherine Gustilo",
      "author_url": "",
      "post_date": "2023-06-15T21:03:59.187000",
      "content": "<p>Hello,</p>\n<p>Thank you for your question regarding the annotations in the dataset. First, I would like to note that generation of manual annotations is a time consuming and expensive process. Some noise in the labels is to be expected with any such dataset. In this particular dataset, the borders of microvasculature structures are not always clearly distinguishable given the tissue thickness and resolution of the images. For example, when annotating peritubular capillaries, the endothelial cell membrane may be indistinguishable from the tubular epithelial membrane, leading to some overlap of annotation borders with tubular epithelium borders. One other thing to note is that often the borders of the annotations themselves can obscure important visual information used to distinguish the vessels from surrounding structures. It is helpful to have a side by side comparison of both annotated and un-annotated images to distinguish vessels. </p>\n<p>Best,</p>\n<ul>\n<li>Kate Gustilo, Research Analyst and Anatomist</li>\n</ul>",
      "votes": 3,
      "replies": [
        {
          "id": 2304370,
          "author_name": "YYama",
          "author_url": "",
          "post_date": "2023-06-16T00:19:34.850000",
          "content": "<p>Thanks for reply!<br>\nI also respect the great efforts of the annotators, thank you all!</p>\n<p>The only thing that concerns me is that the annotation trends are very different even in dataset1. Perhaps this is unavoidable since multiple annotators are working on it.</p>\n<p>I just want to be clear about the distribution of test data.<br>\nAre separate annotators responsible for train, public and private testing respectively?<br>\nOr is each annotator randomly distributed to each dataset?<br>\nFor example, is annotator 1 responsible for both the train and the public and private test datasets?</p>\n<p>If the unknown annotator is only included in the private test, then it would be much more difficult to create a robust model. The post-processing such as dilation and erosion can cause huge shake-up.</p>\n<p>We look forward to your reply. Thanks!</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2305579,
              "author_name": "Katherine Gustilo",
              "author_url": "",
              "post_date": "2023-06-16T19:23:42.740000",
              "content": "<p>YYama, </p>\n<p>The initial annotations were completed by multiple annotators, however all data was validated by the same expert. Thus, differences in annotation trends are minimal and have gone through the same validation process for the training, private and public test sets. I hope this is helpful.</p>\n<p>Best,<br>\nKate Gustilo</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2308978,
              "author_name": "YYama",
              "author_url": "",
              "post_date": "2023-06-19T10:31:02.853000",
              "content": "<p>Thanks for the reply!</p>\n<p>So there is no bias in the annotation trend from one dataset to another.<br>\nVery good news!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2301431,
      "author_name": "YYama",
      "author_url": "",
      "post_date": "2023-06-13T23:50:20.010000",
      "content": "<p>\"The competition data comprises tiles extracted from five Whole Slide Images (WSI) split into two datasets. Tiles from Dataset 1 have annotations that have been expert reviewed. Dataset 2 comprises the remaining tiles from these same WSIs and contain sparse annotations that have not been expert reviewed.\"<br>\n<a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data\" target=\"_blank\">https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data</a></p>\n<p>This is how the dataset description states it. In other words, the annotations in dataset1 are of good accuracy and the annotations in dataset2 are of slightly poorer quality. This information may have an impact on the validation method.</p>\n<p>EDIT:<br>\nAccording to the description, only data from dataset1 is used for the test data. This is very important information.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2302618,
          "author_name": "Sasha Mogilevskii",
          "author_url": "",
          "post_date": "2023-06-14T16:49:29.300000",
          "content": "<p>Yes, but i send img from data_1. For example this image from data_1 too. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6197543%2Fa8fc9e565fae02cb36b7107032eb9aa8%2F1.jpg?generation=1686761343902051&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6197543%2Fe02e1ca0b30bb65d400392732bcf4b78%2FScreenshot_2.jpg?generation=1686761353747613&amp;alt=media\" alt=\"\"></p>",
          "votes": 2,
          "replies": [
            {
              "id": 2302845,
              "author_name": "YYama",
              "author_url": "",
              "post_date": "2023-06-14T22:04:37.953000",
              "content": "<p>Ah, I see. So both were from DATASET1. Thank you for sharing!</p>\n<p>Then I can only assume that they were annotated by separate annotators.<br>\nSome of them have gone beyond the extravascular membrane and some of them have even included tubular epithelial cells.</p>\n<p>I would like the host to disclose information about the annotator too. Are the test sets made by more than one annotator, and is there any bias between public and private annotators?</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2301372": "Hello everyone. After studying the data from datasets 1 and 2, I noticed that the data is annotated in completely different ways. In some instances, the annotations follow the lines, while in other images, they deviate from them. I would like to inquire with the competition organizers about how the test data was annotated or selected. Are these strictly annotated data? Or is the annotation of a similar level? Th u \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6197543%2F9b3efffbbe418bfb299d90f66b0d79de%2FScreenshot_4.jpg?generation=1686689058146565&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6197543%2F47596f0b79ed5d8a8120969ee1411426%2FScreenshot_3.jpg?generation=1686689082173218&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F6197543%2Fb2d5de8f842635d33d280d7405104789%2FScreenshot_2.jpg?generation=1686689091526052&alt=media)",
    "2301416": "great finding, so probably it is one of the reasons why dilation boost both local CV and public lb score https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/416901",
    "2304251": "Hello,\n\nThank you for your question regarding the annotations in the dataset. First, I would like to note that generation of manual annotations is a time consuming and expensive process. Some noise in the labels is to be expected with any such dataset. In this particular dataset, the borders of microvasculature structures are not always clearly distinguishable given the tissue thickness and resolution of the images. For example, when annotating peritubular capillaries, the endothelial cell membrane may be indistinguishable from the tubular epithelial membrane, leading to some overlap of annotation borders with tubular epithelium borders. One other thing to note is that often the borders of the annotations themselves can obscure important visual information used to distinguish the vessels from surrounding structures. It is helpful to have a side by side comparison of both annotated and un-annotated images to distinguish vessels. \n\nBest,\n - Kate Gustilo, Research Analyst and Anatomist",
    "2301431": "\"The competition data comprises tiles extracted from five Whole Slide Images (WSI) split into two datasets. Tiles from Dataset 1 have annotations that have been expert reviewed. Dataset 2 comprises the remaining tiles from these same WSIs and contain sparse annotations that have not been expert reviewed.\"\nhttps://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/data\n\nThis is how the dataset description states it. In other words, the annotations in dataset1 are of good accuracy and the annotations in dataset2 are of slightly poorer quality. This information may have an impact on the validation method.\n\nEDIT:\nAccording to the description, only data from dataset1 is used for the test data. This is very important information."
  }
}