{
  "id": 486937,
  "title": "Whole Slide Image & Fine-Grained Visual Semantic Interaction. Prompt Learning in Vision-Language Models. Diagnosis Prompts. ",
  "url": "/competitions/HyperLeaf2024/discussion/486937",
  "author_name": "Marília Prata",
  "post_date": "2024-03-26T22:53:45.525000",
  "votes": 14,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>WSI (Whole Slide Image) with FiVE</h1>\n<p>Posted on Kaggle on March 26, 2024.</p>\n<p>Generalizable Whole Slide Image Classification with Fine-Grained Visual-Semantic Interaction</p>\n<p>Authors: Hao Li, Ying Chen, Yifei Chen, Wenxian Yang, Bowen Ding, Yuchen Han, Liansheng Wang, Rongshan Yu - <a href=\"https://doi.org/10.48550/arXiv.2402.19326\" target=\"_blank\">https://doi.org/10.48550/arXiv.2402.19326</a></p>\n<p>\"How to extract useful information from raw pathological reports to construct WSI report pairs is a key issue\"</p>\n<p>\"How to craft prompts to make full use of this semantic information to guide fine-grained feature learning is a challenging task.\"</p>\n<p>\"The authors proposed a novel whole slide image classification method with Fine-grained Visual Semantic interaction termed as FiVE, which shows robust generalizability and efficiency in computation. Firstly, they obtained WSIs with non-standardized raw pathological reports from a public database. Collaborating with professional pathologists, they crafted a set of specialized prompts to standardize reports. Following this, the authors employed the large language model GPT-4 to automatically clean and standardize the raw report data. In addition, they proposed the Task-specific Fine-grained Semantic (TFS) Module, which utilizes manually designed prompts to direct visual attention to specific pathological areas while constructing Fine-Grained Guidance to enhance the semantic relevance of model features.</p>\n<p>\" The contributions of this paper are summarized as follows:</p>\n<p>• \"The authors pioneered the utilization of the available WSI diagnostic reports with fine-grained guidance. The obtained fine-rained description labels lead to improved supervision by discriminating the visual appearances more precisely.\"</p>\n<p>• They introduced a novel Task-specific Fine-grained Semantics (TFS) Module to provide fine-grained guidance for model training, substantially improving the model’s generalization capabilities.\"</p>\n<p>• The authors implemented a patch sampling strategy on visual instances during training to enhance computational efficiency without significantly compromising accuracy, thereby optimizing the model’s training process.\"</p>\n<p>PROMPT LEARNING in VISION-LANGUAGE MODELS</p>\n<p>Context Optimization (CoOp)</p>\n<p>\"Drawing inspiration from prompt learning in natural language processing, some studies have proposed adapting Vision-Language models through end-to-end training of prompt tokens. CoOp (Context Optimization) enhanced CLIP for few-shot ransfer by optimizing a continuous array of prompt vectors within its language branch. Co-CoOp identified CoOp’s suboptimal performance on new classes and tackled the generalization issue by conditioning prompts directly on image instances.\"</p>\n<p>\"It was advocated for optimizing diverse sets of prompts by understanding their distribution. MaPLe (Multi-modal prompt learning) investigated the effectiveness of multi-modal prompt learning in order to improve alignment between vision and language representations. It was adopted a set of learnable adaption prompts and prepend them to the word tokens at higher transformer layers, efficiently fine-tuning LLaMA with less cost.\"</p>\n<p>\"Furthermore, in the context of WSI (Whole Slide Image) classification, prompts function as valuable adjuncts, enriching contextual information and semantic interpretation. The strategic utilization of prompts substantially improved model performance.\"</p>\n<p>DIAGNOSIS PROMPTS</p>\n<p>\"The authors introduced Diagnosis Prompts to guide the aggregation of instance features into bag-level features. They computed the similarity between the instance features and the given manual prompts, utilizing the similarity scores as weights W for feature aggregation to improve the task-specific relevance of the features. Here they utilized the identical manual prompts as those used to standardize the raw data.\"</p>\n<p>\"In addition, manual-designed prompts may have some flaws, potentially failing to comprehensively capture the specific morphological characteristics of the lesion, and the model struggles to generalize towards unseen classes due to the late fusion through the transformer layers. Besides, fine-tuning the model may not always be feasible as it requires training a large number of parameters. Particularly in the case of low-data regimes, where the availability of training data like whole slide images is extremely limited.\"</p>\n<p>\"LLaMA-Adapter and LLaMA-Adapter-v2 explore the way to efficient fine-tuning of Language Models and Vision-Language Models respectively. These approaches introduced the Adaptation Prompt to gradually acquire instructional knowledge. They adopted zero-initialized attention with gating mechanisms to ensure stable training in the early stages. Inspired by these methods, they introduced learnable continuous diagnosis prompts Ul to enrich the context information and make their model have stronger transferability.\"</p>\n<p>\"Different with the traditional context learning prompts method, their approach pays attention to the acquisition of prior knowledge, similar to the methodology employed in Detection Transformer (DETR). The authors aimed to acquire a set of appropriate query values to streamline subsequent feature screening processes, and can also quickly transfer to other tasks by fine-tuning this set of queries.\"</p>\n<p>\"In conclusion: they introduced FiVE (Fine-grained Visual Semantic), a novel framework that demonstrates robust generalization and strong transferability for WSI (Whole Slide Image) classification. Their work pioneers the use of non-standardized pathological reports and corresponding WSIs from public databases to develop VLM.\"</p>\n<p>\"To harness the intricacies and diversity present in these reports, they introduced the Task-specific Fine-grained Semantics (TFS) module. This module reconstructs fine-grained labels and corresponding diagnosis prompts during the training phase, while introducing diagnosis prompts, thereby enhancing themantic relevance of its features. Furthermore, considering that pathological visual patterns are redundantly distributed across tissue slices, we sample a subset of visual patches during training. Their results demonstrate the robust generalizability and computational efficiency of our proposed framework, which also exhibits strong zero-shot performance and is readily adaptable to other tasks with fine-tuning lightly.\"</p>\n<p>\"Moreover, they observed that asthe maximum number of sampled patches increases, the model’s performance consistently improves until it plateaus. They aspire to provide empirical insights and contribute to AI pathology research through their methodology.\"</p>\n<p><a href=\"https://arxiv.org/abs/2402.19326\" target=\"_blank\">https://arxiv.org/abs/2402.19326</a></p>",
  "messages": [
    {
      "id": 2718049,
      "postDate": "2024-03-26T22:53:45.527Z",
      "content": "<h1>WSI (Whole Slide Image) with FiVE</h1>\n<p>Posted on Kaggle on March 26, 2024.</p>\n<p>Generalizable Whole Slide Image Classification with Fine-Grained Visual-Semantic Interaction</p>\n<p>Authors: Hao Li, Ying Chen, Yifei Chen, Wenxian Yang, Bowen Ding, Yuchen Han, Liansheng Wang, Rongshan Yu - <a href=\"https://doi.org/10.48550/arXiv.2402.19326\" target=\"_blank\">https://doi.org/10.48550/arXiv.2402.19326</a></p>\n<p>\"How to extract useful information from raw pathological reports to construct WSI report pairs is a key issue\"</p>\n<p>\"How to craft prompts to make full use of this semantic information to guide fine-grained feature learning is a challenging task.\"</p>\n<p>\"The authors proposed a novel whole slide image classification method with Fine-grained Visual Semantic interaction termed as FiVE, which shows robust generalizability and efficiency in computation. Firstly, they obtained WSIs with non-standardized raw pathological reports from a public database. Collaborating with professional pathologists, they crafted a set of specialized prompts to standardize reports. Following this, the authors employed the large language model GPT-4 to automatically clean and standardize the raw report data. In addition, they proposed the Task-specific Fine-grained Semantic (TFS) Module, which utilizes manually designed prompts to direct visual attention to specific pathological areas while constructing Fine-Grained Guidance to enhance the semantic relevance of model features.</p>\n<p>\" The contributions of this paper are summarized as follows:</p>\n<p>• \"The authors pioneered the utilization of the available WSI diagnostic reports with fine-grained guidance. The obtained fine-rained description labels lead to improved supervision by discriminating the visual appearances more precisely.\"</p>\n<p>• They introduced a novel Task-specific Fine-grained Semantics (TFS) Module to provide fine-grained guidance for model training, substantially improving the model’s generalization capabilities.\"</p>\n<p>• The authors implemented a patch sampling strategy on visual instances during training to enhance computational efficiency without significantly compromising accuracy, thereby optimizing the model’s training process.\"</p>\n<p>PROMPT LEARNING in VISION-LANGUAGE MODELS</p>\n<p>Context Optimization (CoOp)</p>\n<p>\"Drawing inspiration from prompt learning in natural language processing, some studies have proposed adapting Vision-Language models through end-to-end training of prompt tokens. CoOp (Context Optimization) enhanced CLIP for few-shot ransfer by optimizing a continuous array of prompt vectors within its language branch. Co-CoOp identified CoOp’s suboptimal performance on new classes and tackled the generalization issue by conditioning prompts directly on image instances.\"</p>\n<p>\"It was advocated for optimizing diverse sets of prompts by understanding their distribution. MaPLe (Multi-modal prompt learning) investigated the effectiveness of multi-modal prompt learning in order to improve alignment between vision and language representations. It was adopted a set of learnable adaption prompts and prepend them to the word tokens at higher transformer layers, efficiently fine-tuning LLaMA with less cost.\"</p>\n<p>\"Furthermore, in the context of WSI (Whole Slide Image) classification, prompts function as valuable adjuncts, enriching contextual information and semantic interpretation. The strategic utilization of prompts substantially improved model performance.\"</p>\n<p>DIAGNOSIS PROMPTS</p>\n<p>\"The authors introduced Diagnosis Prompts to guide the aggregation of instance features into bag-level features. They computed the similarity between the instance features and the given manual prompts, utilizing the similarity scores as weights W for feature aggregation to improve the task-specific relevance of the features. Here they utilized the identical manual prompts as those used to standardize the raw data.\"</p>\n<p>\"In addition, manual-designed prompts may have some flaws, potentially failing to comprehensively capture the specific morphological characteristics of the lesion, and the model struggles to generalize towards unseen classes due to the late fusion through the transformer layers. Besides, fine-tuning the model may not always be feasible as it requires training a large number of parameters. Particularly in the case of low-data regimes, where the availability of training data like whole slide images is extremely limited.\"</p>\n<p>\"LLaMA-Adapter and LLaMA-Adapter-v2 explore the way to efficient fine-tuning of Language Models and Vision-Language Models respectively. These approaches introduced the Adaptation Prompt to gradually acquire instructional knowledge. They adopted zero-initialized attention with gating mechanisms to ensure stable training in the early stages. Inspired by these methods, they introduced learnable continuous diagnosis prompts Ul to enrich the context information and make their model have stronger transferability.\"</p>\n<p>\"Different with the traditional context learning prompts method, their approach pays attention to the acquisition of prior knowledge, similar to the methodology employed in Detection Transformer (DETR). The authors aimed to acquire a set of appropriate query values to streamline subsequent feature screening processes, and can also quickly transfer to other tasks by fine-tuning this set of queries.\"</p>\n<p>\"In conclusion: they introduced FiVE (Fine-grained Visual Semantic), a novel framework that demonstrates robust generalization and strong transferability for WSI (Whole Slide Image) classification. Their work pioneers the use of non-standardized pathological reports and corresponding WSIs from public databases to develop VLM.\"</p>\n<p>\"To harness the intricacies and diversity present in these reports, they introduced the Task-specific Fine-grained Semantics (TFS) module. This module reconstructs fine-grained labels and corresponding diagnosis prompts during the training phase, while introducing diagnosis prompts, thereby enhancing themantic relevance of its features. Furthermore, considering that pathological visual patterns are redundantly distributed across tissue slices, we sample a subset of visual patches during training. Their results demonstrate the robust generalizability and computational efficiency of our proposed framework, which also exhibits strong zero-shot performance and is readily adaptable to other tasks with fine-tuning lightly.\"</p>\n<p>\"Moreover, they observed that asthe maximum number of sampled patches increases, the model’s performance consistently improves until it plateaus. They aspire to provide empirical insights and contribute to AI pathology research through their methodology.\"</p>\n<p><a href=\"https://arxiv.org/abs/2402.19326\" target=\"_blank\">https://arxiv.org/abs/2402.19326</a></p>",
      "rawMarkdown": "#WSI (Whole Slide Image) with FiVE\n\nPosted on Kaggle on March 26, 2024.\n\nGeneralizable Whole Slide Image Classification with Fine-Grained Visual-Semantic Interaction\n\nAuthors: Hao Li, Ying Chen, Yifei Chen, Wenxian Yang, Bowen Ding, Yuchen Han, Liansheng Wang, Rongshan Yu - https://doi.org/10.48550/arXiv.2402.19326\n\n\"How to extract useful information from raw pathological reports to construct WSI report pairs is a key issue\"\n\n\"How to craft prompts to make full use of this semantic information to guide fine-grained feature learning is a challenging task.\"\n\n\"The authors proposed a novel whole slide image classification method with Fine-grained Visual Semantic interaction termed as FiVE, which shows robust generalizability and efficiency in computation. Firstly, they obtained WSIs with non-standardized raw pathological reports from a public database. Collaborating with professional pathologists, they crafted a set of specialized prompts to standardize reports. Following this, the authors employed the large language model GPT-4 to automatically clean and standardize the raw report data. In addition, they proposed the Task-specific Fine-grained Semantic (TFS) Module, which utilizes manually designed prompts to direct visual attention to specific pathological areas while constructing Fine-Grained Guidance to enhance the semantic relevance of model features.\n\n\" The contributions of this paper are summarized as follows:\n\n• \"The authors pioneered the utilization of the available WSI diagnostic reports with fine-grained guidance. The obtained fine-rained description labels lead to improved supervision by discriminating the visual appearances more precisely.\"\n\n• They introduced a novel Task-specific Fine-grained Semantics (TFS) Module to provide fine-grained guidance for model training, substantially improving the model’s generalization capabilities.\"\n\n• The authors implemented a patch sampling strategy on visual instances during training to enhance computational efficiency without significantly compromising accuracy, thereby optimizing the model’s training process.\"\n\nPROMPT LEARNING in VISION-LANGUAGE MODELS\n\nContext Optimization (CoOp)\n\n\"Drawing inspiration from prompt learning in natural language processing, some studies have proposed adapting Vision-Language models through end-to-end training of prompt tokens. CoOp (Context Optimization) enhanced CLIP for few-shot ransfer by optimizing a continuous array of prompt vectors within its language branch. Co-CoOp identified CoOp’s suboptimal performance on new classes and tackled the generalization issue by conditioning prompts directly on image instances.\"\n\n\"It was advocated for optimizing diverse sets of prompts by understanding their distribution. MaPLe (Multi-modal prompt learning) investigated the effectiveness of multi-modal prompt learning in order to improve alignment between vision and language representations. It was adopted a set of learnable adaption prompts and prepend them to the word tokens at higher transformer layers, efficiently fine-tuning LLaMA with less cost.\"\n\n\"Furthermore, in the context of WSI (Whole Slide Image) classification, prompts function as valuable adjuncts, enriching contextual information and semantic interpretation. The strategic utilization of prompts substantially improved model performance.\"\n\nDIAGNOSIS PROMPTS\n\n\"The authors introduced Diagnosis Prompts to guide the aggregation of instance features into bag-level features. They computed the similarity between the instance features and the given manual prompts, utilizing the similarity scores as weights W for feature aggregation to improve the task-specific relevance of the features. Here they utilized the identical manual prompts as those used to standardize the raw data.\"\n\n\"In addition, manual-designed prompts may have some flaws, potentially failing to comprehensively capture the specific morphological characteristics of the lesion, and the model struggles to generalize towards unseen classes due to the late fusion through the transformer layers. Besides, fine-tuning the model may not always be feasible as it requires training a large number of parameters. Particularly in the case of low-data regimes, where the availability of training data like whole slide images is extremely limited.\"\n\n\"LLaMA-Adapter and LLaMA-Adapter-v2 explore the way to efficient fine-tuning of Language Models and Vision-Language Models respectively. These approaches introduced the Adaptation Prompt to gradually acquire instructional knowledge. They adopted zero-initialized attention with gating mechanisms to ensure stable training in the early stages. Inspired by these methods, they introduced learnable continuous diagnosis prompts Ul to enrich the context information and make their model have stronger transferability.\"\n\n\"Different with the traditional context learning prompts method, their approach pays attention to the acquisition of prior knowledge, similar to the methodology employed in Detection Transformer (DETR). The authors aimed to acquire a set of appropriate query values to streamline subsequent feature screening processes, and can also quickly transfer to other tasks by fine-tuning this set of queries.\"\n\n\"In conclusion: they introduced FiVE (Fine-grained Visual Semantic), a novel framework that demonstrates robust generalization and strong transferability for WSI (Whole Slide Image) classification. Their work pioneers the use of non-standardized pathological reports and corresponding WSIs from public databases to develop VLM.\"\n\n\"To harness the intricacies and diversity present in these reports, they introduced the Task-specific Fine-grained Semantics (TFS) module. This module reconstructs fine-grained labels and corresponding diagnosis prompts during the training phase, while introducing diagnosis prompts, thereby enhancing themantic relevance of its features. Furthermore, considering that pathological visual patterns are redundantly distributed across tissue slices, we sample a subset of visual patches during training. Their results demonstrate the robust generalizability and computational efficiency of our proposed framework, which also exhibits strong zero-shot performance and is readily adaptable to other tasks with fine-tuning lightly.\"\n\n\"Moreover, they observed that asthe maximum number of sampled patches increases, the model’s performance consistently improves until it plateaus. They aspire to provide empirical insights and contribute to AI pathology research through their methodology.\"\n\nhttps://arxiv.org/abs/2402.19326",
      "votes": 13
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2718049": "#WSI (Whole Slide Image) with FiVE\n\nPosted on Kaggle on March 26, 2024.\n\nGeneralizable Whole Slide Image Classification with Fine-Grained Visual-Semantic Interaction\n\nAuthors: Hao Li, Ying Chen, Yifei Chen, Wenxian Yang, Bowen Ding, Yuchen Han, Liansheng Wang, Rongshan Yu - https://doi.org/10.48550/arXiv.2402.19326\n\n\"How to extract useful information from raw pathological reports to construct WSI report pairs is a key issue\"\n\n\"How to craft prompts to make full use of this semantic information to guide fine-grained feature learning is a challenging task.\"\n\n\"The authors proposed a novel whole slide image classification method with Fine-grained Visual Semantic interaction termed as FiVE, which shows robust generalizability and efficiency in computation. Firstly, they obtained WSIs with non-standardized raw pathological reports from a public database. Collaborating with professional pathologists, they crafted a set of specialized prompts to standardize reports. Following this, the authors employed the large language model GPT-4 to automatically clean and standardize the raw report data. In addition, they proposed the Task-specific Fine-grained Semantic (TFS) Module, which utilizes manually designed prompts to direct visual attention to specific pathological areas while constructing Fine-Grained Guidance to enhance the semantic relevance of model features.\n\n\" The contributions of this paper are summarized as follows:\n\n• \"The authors pioneered the utilization of the available WSI diagnostic reports with fine-grained guidance. The obtained fine-rained description labels lead to improved supervision by discriminating the visual appearances more precisely.\"\n\n• They introduced a novel Task-specific Fine-grained Semantics (TFS) Module to provide fine-grained guidance for model training, substantially improving the model’s generalization capabilities.\"\n\n• The authors implemented a patch sampling strategy on visual instances during training to enhance computational efficiency without significantly compromising accuracy, thereby optimizing the model’s training process.\"\n\nPROMPT LEARNING in VISION-LANGUAGE MODELS\n\nContext Optimization (CoOp)\n\n\"Drawing inspiration from prompt learning in natural language processing, some studies have proposed adapting Vision-Language models through end-to-end training of prompt tokens. CoOp (Context Optimization) enhanced CLIP for few-shot ransfer by optimizing a continuous array of prompt vectors within its language branch. Co-CoOp identified CoOp’s suboptimal performance on new classes and tackled the generalization issue by conditioning prompts directly on image instances.\"\n\n\"It was advocated for optimizing diverse sets of prompts by understanding their distribution. MaPLe (Multi-modal prompt learning) investigated the effectiveness of multi-modal prompt learning in order to improve alignment between vision and language representations. It was adopted a set of learnable adaption prompts and prepend them to the word tokens at higher transformer layers, efficiently fine-tuning LLaMA with less cost.\"\n\n\"Furthermore, in the context of WSI (Whole Slide Image) classification, prompts function as valuable adjuncts, enriching contextual information and semantic interpretation. The strategic utilization of prompts substantially improved model performance.\"\n\nDIAGNOSIS PROMPTS\n\n\"The authors introduced Diagnosis Prompts to guide the aggregation of instance features into bag-level features. They computed the similarity between the instance features and the given manual prompts, utilizing the similarity scores as weights W for feature aggregation to improve the task-specific relevance of the features. Here they utilized the identical manual prompts as those used to standardize the raw data.\"\n\n\"In addition, manual-designed prompts may have some flaws, potentially failing to comprehensively capture the specific morphological characteristics of the lesion, and the model struggles to generalize towards unseen classes due to the late fusion through the transformer layers. Besides, fine-tuning the model may not always be feasible as it requires training a large number of parameters. Particularly in the case of low-data regimes, where the availability of training data like whole slide images is extremely limited.\"\n\n\"LLaMA-Adapter and LLaMA-Adapter-v2 explore the way to efficient fine-tuning of Language Models and Vision-Language Models respectively. These approaches introduced the Adaptation Prompt to gradually acquire instructional knowledge. They adopted zero-initialized attention with gating mechanisms to ensure stable training in the early stages. Inspired by these methods, they introduced learnable continuous diagnosis prompts Ul to enrich the context information and make their model have stronger transferability.\"\n\n\"Different with the traditional context learning prompts method, their approach pays attention to the acquisition of prior knowledge, similar to the methodology employed in Detection Transformer (DETR). The authors aimed to acquire a set of appropriate query values to streamline subsequent feature screening processes, and can also quickly transfer to other tasks by fine-tuning this set of queries.\"\n\n\"In conclusion: they introduced FiVE (Fine-grained Visual Semantic), a novel framework that demonstrates robust generalization and strong transferability for WSI (Whole Slide Image) classification. Their work pioneers the use of non-standardized pathological reports and corresponding WSIs from public databases to develop VLM.\"\n\n\"To harness the intricacies and diversity present in these reports, they introduced the Task-specific Fine-grained Semantics (TFS) module. This module reconstructs fine-grained labels and corresponding diagnosis prompts during the training phase, while introducing diagnosis prompts, thereby enhancing themantic relevance of its features. Furthermore, considering that pathological visual patterns are redundantly distributed across tissue slices, we sample a subset of visual patches during training. Their results demonstrate the robust generalizability and computational efficiency of our proposed framework, which also exhibits strong zero-shot performance and is readily adaptable to other tasks with fine-tuning lightly.\"\n\n\"Moreover, they observed that asthe maximum number of sampled patches increases, the model’s performance consistently improves until it plateaus. They aspire to provide empirical insights and contribute to AI pathology research through their methodology.\"\n\nhttps://arxiv.org/abs/2402.19326"
  }
}